The GDPR regulates data combination and inference more directly than most people realize. But when integrated data reveals collective, institutional or national-security intelligence rather than information about an identifiable person, much of the legal protection falls away.
A grocery list may not look like health information. Location data may not state a person’s religion. A collection of page views does not expressly declare a political affiliation, sexual orientation or financial problem. Combine those records, however, and they can reveal all of those things, and often more accurately than the person ever disclosed them.
That is the privacy problem created by data synthesis: separate pieces of information are integrated, compared or analyzed to produce knowledge that did not exist in any individual source. It also makes traditional descriptions of privacy (what was collected, where it is stored and who received it) increasingly incomplete.
The law has not ignored this problem. In fact, the European Union’s General Data Protection Regulation addresses it more directly than many summaries of the GDPR suggest. But existing rules remain strongest when synthesis produces information about identifiable individuals. They become far less coherent when the same capability reveals vulnerabilities involving communities, organizations, infrastructure or national security.
The GDPR expressly covers combination
The GDPR does not use “data synthesis” as a defined term. It does not need to. Article 4(2) defines the processing of personal data to include its “alignment or combination.” Joining databases, resolving two records to the same identity, constructing a profile and generating an inference can therefore all constitute regulated processing.
This matters because organizations cannot argue that every input was collected lawfully and treat the knowledge created by combining those inputs as legally neutral. The integrated use still needs a lawful basis. It must comply with purpose limitation, data minimization, transparency, accuracy, security and storage limits.
If information collected for one purpose is used to create a new analytical product, the organization must determine whether the new purpose is compatible with the original one. If it is not, it generally needs fresh consent or specific legal authority. The mere fact that a system can integrate the data does not establish that it may.
Combination can also change the legal character of information. Records that do not identify anyone in isolation may become personal data once linked with another dataset. Supposedly anonymous information loses its exemption if a person becomes reasonably identifiable. And ordinary personal data can become specially protected data if synthesis reveals health, ethnicity, religion, political opinions, trade-union membership, genetics, biometrics, sex life or sexual orientation.
The European Data Protection Board gives a simple example: grocery-purchase records combined with nutritional information may allow a company to infer someone’s health. Neither source necessarily contains a diagnosis. The synthesis creates health information.
European courts have begun drawing the line
The Court of Justice of the European Union has now applied these principles to real data-fusion practices.
In the 2024 case Schrems v. Meta, the Court held that Meta could not aggregate, analyze and process personal data gathered on and off Facebook, from users and third parties, for targeted advertising without limits on time or distinctions between types of data. The ruling was based on data minimization: possessing a large pool of information does not give a platform an unlimited right to synthesize it for every advertising purpose.
In an earlier Meta judgment, the Court found that collecting information from third-party sites and apps, linking it to a social-media account and using it can amount to processing sensitive data when the combination reveals a protected characteristic. And in a 2022 Lithuanian case, it held that information capable of indirectly revealing sexual orientation could receive the GDPR’s special-category protection.
The inference does not necessarily have to be correct to create a privacy problem. A system that categorizes someone as ill, gay, Muslim, heavily indebted or politically extreme can affect that person even when its conclusion is false. Privacy risk comes from the classification and its use, not simply its factual accuracy.
People also have rights in relation to the result. European Data Protection Board guidance says the GDPR right of access includes personal data inferred or derived by a service provider, not only information supplied directly by the individual.
The GDPR requires a data-protection impact assessment when processing is likely to create a high risk to people’s rights and freedoms. European regulators identify matching or combining datasets as one warning sign. That does not make every database join unlawful or automatically require an assessment, but it recognizes that integration can create a risk greater than the sum of its inputs.
Article 22 is narrower. It provides protection against certain decisions based solely on automated processing that have legal or similarly significant effects. It is not a general ban on profiling, inference or synthesis. A company may therefore be subject to the GDPR’s ordinary processing rules even when Article 22 does not apply.
Some laws are even more explicit
Other jurisdictions have adopted narrower laws that name the problem directly.
The EU’s Digital Markets Act restricts designated technology gatekeepers from combining or cross-using personal data across different services, or with third-party services, without the required user consent. Unlike the GDPR’s general principles, this rule targets cross-service integration itself, but only for companies designated as gatekeepers.
Washington State’s My Health My Data Act is unusually clear. It covers health information derived or extrapolated from non-health information, including proxy, inferred, emergent and algorithmic data. Its definition of collection expressly includes inferring and deriving data. A wellness app cannot escape the law merely because it inferred a health condition from shopping, location or behavioural records rather than receiving it from a doctor.
California’s privacy law treats inferences used to construct consumer profiles as personal information. Its regulator has pursued data brokers dealing in inferred profiles, and its new centralized deletion system extends requests to associated inferences.
Australia, meanwhile, has a statutory regime governing certain government data-matching programs. It requires written protocols, technical standards and privacy oversight when tax and benefits records are compared. Its scope is limited, but it demonstrates that law can regulate matching—not only each original database.
Canada recognizes inference, but not synthesis risk
Canada’s current federal private-sector law, PIPEDA, addresses synthesis indirectly. Organizations must identify their purposes, limit collection and use, and generally obtain consent before using personal information for a previously unidentified purpose. Its broad definition of personal information can encompass inferences about an identifiable person, but it contains no express general test for integration risk.
The federal government’s new Bill C-36 would make the point explicit. The proposed law defines personal information as information about an identifiable individual, “including information that is inferred about the individual.” It would also provide explanation rights for automated predictions, recommendations or decisions with legal or similarly significant effects. But the bill, currently at second reading, does not expressly regulate dataset combination or require a distinct synthesis assessment.
Quebec is further ahead. Its regime defines profiling, requires notice about technologies used to identify, locate or profile people, regulates exclusively automated decisions and requires privacy impact assessments for information-system projects involving personal information.
The law’s much larger blind spot
All of these regimes share a boundary: privacy law is principally concerned with natural persons. The GDPR expressly does not protect legal persons as data subjects. If synthesis reveals an individual soldier’s health or location, privacy law may apply. If it reveals a military unit’s readiness, a government’s procurement dependencies, a hospital network’s vulnerabilities or a community’s political susceptibility without identifying particular people, it may not.
We’ve written about the power of AI in pattern recognition and the dangerous data exhaust made possible by combining advertising data with AI prompts, email data where available, social media data and content, and search history. Data brokers in the US for example can go much further with current address address history, phone number and email history as well. The lack of notification and recourse is also a huge issue her. Security professionals understand this as aggregation risk or the “mosaic effect”: facts that appear harmless separately can disclose a sensitive capability or relationship when assembled. Canada’s Cyber Centre acknowledges that aggregation can change data sensitivity and risk. Access-to-information law sometimes permits information to be withheld where combined disclosures would create a demonstrable security risk.
But these are fragmented security practices and disclosure doctrines. They do not create a general, enforceable obligation to assess what a contractor, analytics platform or foreign-controlled vendor can learn by operating across several adjacent systems.
A modern framework should therefore evaluate more than the classification of each input. It should examine which datasets can be integrated, what new facts can be inferred, who can run cross-system queries, whether combined information must be reclassified, which vendors can access the synthesis layer, and what individual, collective, commercial or sovereign harm could follow.
Privacy law has begun to move from protecting data to protecting people from what data can reveal. The next step is harder: governing what integrated systems can know, including knowledge whose subject is not one identifiable person. Until that happens, the law will continue to protect the pieces more reliably than the picture they create.
What the law should require
The starting presumption should be separation. Data collected in different contexts should remain in separate systems unless combining it is demonstrably necessary for a specific, legitimate purpose and the additional context cannot reasonably be obtained in a less intrusive way. Synthesis should be an exceptional, documented act, not a default capability quietly enabled because the technology permits it.
Advertising personalization should not satisfy that necessity test. Showing someone a marginally more relevant advertisement is not sufficient justification for constructing a persistent, cross-context representation of their health, finances, relationships, movements, beliefs or vulnerabilities. The commercial value of an inference to an advertiser is not equivalent to necessity, and it should not override the individual’s interest in keeping unrelated parts of their life separate.
The law must also account for the fact that databases are frequently outdated, incomplete, incorrectly matched or simply wrong. Combining several records can make an error appear more authoritative while making its origin harder to identify. An inaccurate address, misidentified purchase or false association can become the foundation for increasingly consequential conclusions, with the affected person having little idea that the synthesized profile exists and almost no practical ability to challenge it.
Notice should therefore be triggered when synthesis occurs—not only when the original data are collected or when an automated decision is eventually made. People should be able to see which categories of information have been brought together, where those data came from, what was inferred, why the synthesis was conducted, how long the result will be retained and which organizations or systems can access it. This need not require disclosure of proprietary code, but it must reveal enough for a person to understand the profile being constructed about them.
People must also have enforceable rights to inspect, correct, contest and delete synthesized information. Where an inference is disputed, consequential use of it should be suspended until it is reviewed, and any correction should follow the information to vendors, affiliates and other downstream recipients. Without that recourse, transparency merely allows people to observe an error continuing to operate against them.
The governing principle should be simple: data collected separately should remain separate unless synthesis is necessary, proportionate, disclosed and contestable. Organizations should not acquire an unrestricted right to know more about a person simply because they possess enough disconnected information to calculate it.

