Can "anonymized" health data still identify you?
“Don’t worry, the data is anonymized.” Anyone who shares health data has heard this. The mechanism behind it is weaker than it sounds.
How de-identification is supposed to work
Under HIPAA, the central mechanism for enabling research on health records is de-identification: strip the 18 listed identifiers (names, addresses below state level, almost all dates, ID numbers and the like), or have an expert certify that the re-identification risk is “very small”, and the record is no longer protected health information. It can then be shared without the patient’s authorisation. The assumption is that without identifiers, the record no longer points back to a person.
The same assumption, written into French law
France puts it in a statute. Article L1461-2 of the code de la santé publique requires that data from the national health data system released to the public be processed into aggregated statistics, or into individual data constituted in such a way that direct or indirect identification of the persons concerned is impossible. Reuse may have neither the object nor the effect of identifying anyone, and the data are free.
The two limbs are not equivalent, though the article treats them as alternatives. An aggregate genuinely discards the individual: a count, a rate or a mean over a large enough group describes nobody in particular. A record that stays individual still describes one person, and what makes it re-identifiable is not the identifiers that were removed but the combination of what remains. The first limb can deliver what the article asks. Whether the second can is the question the rest of this page is about.
Why it fails
Removing direct identifiers is not enough. A systematic review of re-identification attacks found that about a quarter of records were re-identified on average, a third in the health-data attacks, almost always in datasets that had only been stripped of names and other direct identifiers rather than de-identified to a formal standard. Later work showed that 15 demographic attributes are enough to re-identify 99.98% of Americans. A birth date here, a postcode there, a rare diagnosis, an admission date: none of these is a name, but together they can narrow a record down to one person.
Breaches of protected health information reported to the US federal regulator since 2009 had, by the end of 2025, exposed the records of more than one billion people (people are counted once per breach, so the total is not a count of distinct persons). Once a de-identified dataset leaks, anyone holding other data about you can attempt the combination.
What French enforcement has actually found
Two French decisions of 2026 tested the claim directly, and both rejected it.
In February 2026 the Conseil d’État upheld three CNIL fines against companies in the Cegedim group, which ran two databases built from GP software and pharmacy systems: 13.4 million consultations against 4 million patient codes, and some 78 million customer identifiers from 8,500 pharmacies. The companies argued that a patient code made the records anonymous. The court set out the test. Data count as anonymised by pseudonymisation only where the risk of identification is insignificant, identification being unachievable in practice because it would demand a disproportionate effort in time, cost and manpower. Here it was achievable. Alongside the code sat age, sex, socio-professional category, pathologies, prescriptions, sick leave, vaccinations, the date and sometimes the exact time of each visit, and the prescriber’s ADELI or RPPS number, which a public search engine turns into a name. The regulator had picked out individual patients using an ordinary spreadsheet and the companies’ own code list, in little time and with few resources.
Three months later the CNIL fined IQVIA five million euros over two authorised health data warehouses, one fed by about 14,000 pharmacies and the other by several thousand doctors. IQVIA argued that the data were anonymous. The CNIL held that they were pseudonymous, because each patient carried a unique identifier that followed a care pathway, the records ran deep enough to include marital status, number of children, diagnosis, symptoms, allergies, weight, height, pulse, vaccinations and sick leave, and they could be combined with publicly available data.
Both decisions turn on the point this page has been making. What makes a record re-identifiable is not the identifier that was removed but the combination of what remains.
Pseudonymised is not anonymised, and the difference is legal
The two words describe different legal objects, and using them interchangeably is how the companies above lost.
The GDPR defines pseudonymisation at article 4(5) as processing personal data so that it “can no longer be attributed to a specific data subject without the use of additional information”, where that additional information is kept separately and protected. The key word is separately. The link still exists; it is held somewhere else.
Recital 26 draws the consequence. Pseudonymised data “should be considered to be information on an identifiable natural person”, so the whole regulation continues to apply to it. Only genuinely anonymous information falls outside, and the recital sets the test: account must be taken of “all the means reasonably likely to be used, such as singling out, either by the controller or by another person”, weighing “the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments”.
Three things follow. Pseudonymisation is a security measure, not an exit from data protection law. Anonymisation, if it succeeds, is an exit, which is why the standard for it is severe. And the test is not whether you can re-identify the record; it is whether anyone reasonably could, with the technology available now and as it develops.
Why health data resists anonymisation in particular
The difficulty is not that health records are sensitive. It is that they are precise, and precision is what identifies.
Rare disease makes this plain. A condition affecting five people in every 10,000 is already a near-unique attribute in a region; add an age band and a date of care and the group of candidates often falls to one. The diagnosis that makes the record worth studying is the same attribute that singles the person out, which is exactly the operation recital 26 names.
Continuous measurement does the same thing from the other direction. Physical activity recorded by accelerometer, aggregated into twenty-minute intervals and combined with demographics, allowed re-identification of 94.9% of adults and 87.4% of children in a national survey, and aggregating the data over longer periods did not materially reduce that. The measurements carried no name, no address and no date of birth. A week of ordinary movement was enough.
So the more useful a health record is, the harder it is to anonymise, and the fields a researcher most needs are the fields that identify. That is not a flaw in any particular method. It is the shape of the problem.
The trade-off
The standard response to re-identification risk is to strip more: coarser dates, broader regions, fewer fields. But every field removed also removes scientific value, and the usual institutional response, tightening rules and access controls, has slowed research without clearly improving privacy: two-thirds of US epidemiologists surveyed said the HIPAA Privacy Rule had made research more difficult while only a quarter thought it had enhanced participants’ privacy, and Finland’s secondary-use law was followed by an estimated 47% drop in new data permits in 2023. Removing fields asks one mechanism to deliver privacy and research value at the same time, and it delivers neither fully. That is the trade-off the techniques in the next section exist to manage.
What the modern techniques actually achieve
Stripping identifiers is the weakest form of the art, and criticising it proves little. Three serious approaches exist, and each delivers something real.
Differential privacy adds calibrated noise to the output of a computation, so that the result is almost unchanged whether or not any one person is in the input. Its strength is that the guarantee is formal and quantified rather than asserted, with a measurable budget that degrades as more questions are asked (Dwork and Roth, 2014). Its limit is the thing it protects: it protects answers. It does not hand a researcher an individual-level record to work with, which is what most clinical research needs.
Synthetic data generates artificial records from a model fitted to the real ones, with the promise of statistical realism and no real person inside. The first quantitative evaluation of that promise, across a range of state-of-the-art generative models, concluded that synthetic data “either does not prevent inference attacks or does not retain data utility”, and that it “does not provide a better tradeoff between privacy and utility than traditional anonymisation techniques”. Worse for a governance purpose, the trade-off is unpredictable in advance, because which signals survive the generation is not knowable beforehand.
Federated analysis leaves the records where they are and sends the computation to them, so only results travel. That genuinely removes the bulk transfer, and it is the most promising of the three for multi-centre work. Two caveats hold. The results that leave can themselves leak membership, so federation displaces the disclosure question rather than closing it, and a 2025 survey of the European Health Data Space found the approach still largely at a theoretical stage in practical healthcare settings.
None of the three is a trick, and none of them is useless. What none of them does is make an individual-level health record both fully informative and genuinely anonymous.
Why Health Data Safe asks for consent instead
Our team works with these techniques and knows them well. None of them ships as part of the platform today. A partner whose protocol calls for differential privacy, synthetic data or a federated design can have it built.
What none of them does at Health Data Safe is replace consent. Access to a data set always rests on the consent of the people described in it, given against a specific research protocol, whether the data in question is identified, pseudonymised or anonymised. We do not treat anonymisation as a route around asking.
Three properties explain the choice, and no technique above provides them.
Full fidelity. Nothing is stripped, blurred or synthesised, so the record keeps the precision that made it worth studying. The protection comes from governance rather than from degrading the data.
Revocability. A person can withdraw. Anonymisation is irreversible by design, which is its legal point: once data is genuinely anonymous it is nobody’s any more, and the person it came from has lost the standing to object to anything done with it afterwards. Consent keeps that standing alive.
An audit trail per access. Every access is logged with the identity of the accessor and the scope of their consent record. A person can see who read what, and when. A noise budget or a synthetic sample answers no such question.
Data stays under the patient’s control on open-source infrastructure, and a researcher’s access requires a time-limited, purpose-specific consent record.
See How it works for the mechanics, or read how this compares to the laws protecting health data in Europe. The French provisions quoted above sit in their full context on the health data law page for France.
Sources
- U.S. Code of Federal Regulations, 45 CFR § 164.514, Standard: de-identification of protected health information (Safe Harbor and Expert Determination). ecfr.gov
- Code de la santé publique, article L1461-2, mise à disposition du public des données du système national des données de santé. legifrance.gouv.fr
- Conseil d’État, 13 February 2026, société GERS and Cegedim Santé, nos 498628, 498629 and 498749, mentioned in the Lebon tables. conseil-etat.fr
- Commission nationale de l’informatique et des libertés, deliberation SAN-2026-008 of 26 May 2026, IQVIA OPERATIONS FRANCE. cnil.fr
- El Emam K, Jonker E, Arbuckle L, Malin B (2011). A systematic review of re-identification attacks on health data. PLoS ONE 6(12): e28071. doi:10.1371/journal.pone.0028071
- Rocher L, Hendrickx JM, de Montjoye Y-A (2019). Estimating the success of re-identifications in incomplete datasets using generative models. Nature Communications 10: 3069. doi:10.1038/s41467-019-10933-3
- U.S. Department of Health and Human Services, Office for Civil Rights. Breach Portal: breaches of unsecured protected health information affecting 500 or more individuals (2009 to present). ocrportal.hhs.gov
- Ness RB (2007). Influence of the HIPAA Privacy Rule on health research. JAMA 298(18): 2164–2170. doi:10.1001/jama.298.18.2164
- Regulation (EU) 2016/679 (General Data Protection Regulation), article 4(5) and recital 26. eur-lex.europa.eu
- Na L, Yang C, Lo C-C, Zhao F, Fukuoka Y, Aswani A (2018). Feasibility of reidentifying individuals in large national physical activity data sets from which protected health information has been removed with use of machine learning. JAMA Network Open 1(8): e186040. doi:10.1001/jamanetworkopen.2018.6040
- Dwork C, Roth A (2014). The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science 9(3-4): 211–487. doi:10.1561/0400000042
- Stadler T, Oprisanu B, Troncoso C (2022). Synthetic data: anonymisation groundhog day. 31st USENIX Security Symposium, 1451–1468. usenix.org
- van Drumpt S, Chawla K, Barbereau T, Spagnuelo D, van de Burgwal L (2025). Secondary use under the European Health Data Space: setting the scene and towards a research agenda on privacy-enhancing technologies. Frontiers in Digital Health 7: 1602101. doi:10.3389/fdgth.2025.1602101
- Brück O, Sanmark E, Ponkilainen V, et al. (2024). European health regulations reduce registry-based research. Health Research Policy and Systems 22: 135. doi:10.1186/s12961-024-01228-1