Data Aggregation is the Real Risk of AI

Taylor Karl
Data Aggregation is the Real Risk of AI 1 0

Key Takeaways

  • Permissions vs. coverage: Correct access can still let AI create a combination no one evaluated

  • Aggregation vs. inference: Combining data is routine, but what it reveals may not be

  • Local approval, global exposure: Correct approvals can still leave combinations unreviewed

  • Audit scope shift: Reviews need a new question: what can this system combine?

  • Accountability vs. participation: Spotting aggregation risk differs from formally owning it


A data analyst pulls together a routine report. They've done this dozens of times: staffing numbers from HR, a few finance line items, an operations schedule. The company's new AI assistant handles the query in seconds, pulling from systems they've always had access to.

Halfway through reading the output, they stop. The three data sets, sitting side by side, spell out something nobody meant to announce yet: a site closure, weeks before anyone outside a small planning group was supposed to know.

Nothing about their access was wrong. Every system they queried, they were authorized to query. Every field they pulled, they'd pulled before. The permissions model worked exactly as designed, and that's what makes the moment worth examining.

Teams tend to treat permissions as the finish line for data security: get the access controls right, run the audit, check the box. Organizations have combined data across sources for decades, through spreadsheets, BI tools, and manual effort. What's changed isn't the possibility of combination. It's the friction required to do it.

An AI assistant can retrieve across authorized systems in a single interaction, without having any context that the resulting inference wasn't meant to be revealed yet.

The gap between correct access and safe outcome is the real issue. It's what happens once an AI assistant does the combining in a fraction of the time.

A Passed Permissions Audit Answers the Wrong Question

Ask most IT or data governance teams how they know their systems are secure, and the answer usually points to the last access audit. Reviewers check who has access to what, whether that access is still appropriate, and whether any accounts need to be trimmed back. Passing it usually means the organization considers itself protected.

It's only half true. An access audit answers a specific question: who is authorized to see this piece of information. It's a valuable question, and getting it right matters. But it doesn't answer the second question: what becomes knowable once several pieces of authorized information land in the same place at the same time.

Traditional access and data-protection reviews commonly examine controls such as:

  • Role-based access: Confirming each employee's access matches their current job function

  • Group membership accuracy: Verifying that shared folders and system groups reflect who actually needs them

  • Stale account cleanup: Removing access left over from role changes or departures

  • Data loss prevention flags: Confirming sensitive fields carry the right monitoring tags

Every one of those checks can pass cleanly, and an AI system can still make an inference none of the controls were built to catch. The audit confirms who can open which doors. It says nothing about what someone, or something, can build once several of those doors are open at once.

The gap between access and inference seems wide, until an AI assistant closes it in the time it takes to run a query. A question worth asking is what combinations can reveal.

Artificial Intelligence

How Individually Harmless Data Becomes a Sensitive Conclusion

Combined sources reveal more than any single one shows. HR staffing data shows headcount by location, standard reporting that dozens of managers see regularly. Finance forecasts show projected costs by region. Operations scheduling shows shift assignments for the next quarter, useful for planning but not sensitive in isolation.

Put these different data points together, and a pattern appears. One location shows a headcount reduction planned in the HR data. The same location shows a cost forecast dropping sharply in the finance data, and shift assignments thinning out in the ops data.

No single document says "this site is closing." The AI assistant connected three authorized data points and produced a conclusion none of them explicitly stated.

There's a distinct difference between aggregation and inference:

  • Aggregation is the act of combining multiple authorized sources into one view

  • Inference is a new fact from combined sources that no single source stated

Think of an office badge: it gets you into the building, the break room, and the file room. Track every swipe across all three, and a system can reconstruct your entire day.

AI introduces a shift. Organizations already aggregate data to build dashboards, forecasts, and reports for better decision-making. The risk shows up when combination produces an inference, a fact that may not be authorized for release yet. What makes a fact sensitive often isn't the fact itself. It's who can access it, when, and for what purpose. A combination like this can exist outside any one team's output.

What Happens When Multiple Departments' Data Meet

Enterprise governance is often organized around data boundaries. Each department's decisions get evaluated within its own scope, by its own reviewers, against its own policies. Every department involved in an AI query approved its own data use, and every one of those approvals was reasonable on its own terms.

For example, each department's approval, viewed side by side, shows exactly where the seam sits:

  • HR: Approved staffing data for internal headcount reporting

  • Finance: Approved forecast data for internal budget planning

  • Operations: Approved scheduling data for internal shift management

The problem is what happens at the seam between those decisions. A department's review typically covers its own use case. When an AI system pulls from several data sources at once, the resulting conclusion belongs to none of them individually and appears beyond each department's scope.

No one did anything wrong here. The same pattern holds for any department feeding a given AI query, each approving its own data responsibly. The exposure lives in the gap this creates: department-level approval doesn't guarantee anyone evaluated what the combination reveals.

It's structural, not a staffing problem or training issue inside any one function. Organizations build governance around departments, and AI systems can cross those boundaries when they retrieve and synthesize information across them.

Organizations need a way to evaluate risk at the combination level, not the source level alone.

Audit What Can Be Combined

Reviewing access more frequently within each department won't close this gap. What has to change is the practice itself: what gets reviewed, and what question the review is trying to answer.

A different kind of question needs asking during a review, one aimed at combination and inference rather than individual access.

A practical framework for that review includes four questions:

  • Retrieval scope: Which data sources can this AI system retrieve together, and for which users or roles?

  • New knowability: What sensitive facts become newly knowable once those sources are combined?

  • Boundary crossing: Which of those combinations cross established departmental governance lines?

  • Intentionality: Was this insight explicitly governed, or ungoverned by default?

Working through those four questions gives a review team something an access log can't: a picture of what the AI system can produce. A review team can act on that picture in a few ways: restricting which roles can retrieve a given combination, separating certain data sources from certain AI connectors, or requiring added approval before a high-risk combination reaches a query result.

A redefined review cycle should include this check, with an additional look outside that cadence whenever something changes the retrieval landscape. New AI tool rollouts, departmental restructures, or newly connected data sources can each open combinations that didn't exist the last time anyone checked.

Combination itself isn't a bad thing. Organizations run on analytics precisely because combining data creates legitimate new knowledge, better forecasts, sharper resourcing decisions, and faster problem detection. A good review asks whether what became newly knowable was something the organization chose to govern, or something that slipped past without anyone deciding either way.

Who runs the review is a separate question.

Separate Who Spots Aggregation Risk From Who Owns It

Spotting a risky combination and owning responsibility for it are two different roles. Data and analytics practitioners are well-positioned to notice when something like this happens, because they work directly with how systems retrieve and combine information. It's what allowed the analyst in this pattern to recognize what the combination was showing them.

The instinct for spotting risk is valuable, but formal accountability for aggregation risk belongs to a structure that spans several functions:

  • Data owners: Understand what each source contains and its approved use

  • Security and privacy teams: Assess what combinations create regulatory or confidentiality exposure

  • AI governance leads: Track what retrieval and synthesis capabilities exist across deployed systems

  • Business leadership: Decide which insights the organization is prepared to surface, and when

Contributing to the structure doesn't require formal ownership of it. What it requires is fluency: understanding how data moves and connects, how AI systems retrieve and synthesize it, how access controls function, and how departmental context shapes a given combination. Building this fluency takes deliberate effort, and it's exactly what makes credible participation possible.

The analyst didn't cause the exposure they found, and the responsibility for preventing it doesn't rest on them alone. What they offer is the fluency to recognize it and speak clearly about what they saw, the kind of contribution a governance conversation needs from more people across more functions.

No single role can see every aggregation risk. Effective governance depends on the right functions contributing their perspective to a clearly owned process.

A Better Question for the Next Audit

The analyst decided to close the report and flag it to their manager. The conversation that followed involved people from multiple departments to discuss what the AI system produced. Nothing malfunctioned, and the data permissions followed the rules, granted to the right people, for legitimate reasons. Regardless, the output still caught everyone off guard.

The permissions model was incomplete, not broken. What it leaves unaddressed is a different question: what becomes knowable when an AI system combines authorized data in seconds, work that used to take a person real time and effort to piece together manually. Departmental approval, however careful, never answered that second question either, because the space where one department's data meets another's can fall outside either department's review.

A better audit expands the question beyond "who can access this" to include what can be combined and what that combination might reveal.

Building Cross-Functional Fluency for Aggregation Risk

New Horizons partners with organizations to build fluency that connects data, AI systems, and governance into one coherent skill set instead of three disconnected ones.

Fluency like this is what turns a governance conversation from reactive to credible.

Reach out to New Horizons to build the skills your teams need to understand what AI systems can access, what those systems can infer, and how to govern the difference.

Print