# Anonymise or pseudonymise: what the law says

> The difference is not a matter of wording: it decides whether the GDPR applies. Pseudonymised data remains personal data; only effective anonymisation, demonstrated and documented, takes data outside the scope of the regulation. France's highest administrative court made the point in February 2026 by upholding €1.8 million in fines imposed by the CNIL.

Source: https://smartagt.ai/en/ressources/anonymisation-pseudonymisation/
Published: 2026-10-08
Publisher: SmartAGT (NERVIAL LABS)

---
## Two concepts, two regimes {#definitions}

The [GDPR](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32016R0679) defines pseudonymisation in Article 4: processing in such a manner that the data “can no longer be attributed to a specific data subject without the use of additional information”, provided that this information is kept separately and protected. It does not define anonymisation, but Recital 26 sets its scope: the regulation does not apply to anonymous information, meaning information that no longer identifies a person, taking account of “all the means reasonably likely to be used”.

The same recital settles the status of pseudonymised data: data “which could be attributed to a natural person by the use of additional information should be considered to be information on an identifiable natural person”.

- **Pseudonymisation** Identifiers are replaced by pseudonyms, and the information needed to trace them back to individuals is kept separately and secured. The process is reversible for whoever holds that information: the data remains personal and the GDPR applies in full.
- **Anonymisation** No one can identify the individuals any more using means reasonably likely to be used. The result falls outside the scope of the GDPR; the operation that produces it is still processing of personal data.

*Pseudonymisation protects data without taking it outside the GDPR; anonymisation takes it outside, provided it is demonstrated.*

The regulation encourages pseudonymisation without ever making it a way out. Recital 28 notes that it “can reduce the risks to the data subjects concerned”, and several articles cite it as an appropriate measure: for data protection by design (Article 25), for the security of processing (Article 32), for research (Article 89) and when assessing whether further processing is compatible (Article 6(4)). It is a compliance tool, not an exemption.

## The Cegedim case: France's Conseil d'État sets the record straight {#cegedim}

In 2024, the CNIL, France's data protection authority, fined three companies that used data taken from software used by doctors and pharmacies: €800,000 for GERS and €200,000 for Santestat on 28 August, then [€800,000 for Cegedim Santé](https://www.cnil.fr/en/health-data-cegedim-sante-fined-eu800000) on 5 September, with the penalty made public. In these datasets, patients were identified only by a code. The companies therefore treated the data as non-personal, and processed it without the authorisation required for health data.

The [Conseil d'État, France's highest administrative court, dismissed their appeals on 13 February 2026](https://www.legifrance.gouv.fr/ceta/id/CETATEXT000053483461) (nos. 498628, 498629 and 498749). Drawing on the case law of the Court of Justice of the European Union, it held that data can only be regarded as made anonymous by pseudonymisation “if the risk of identification is insignificant, such identification being impossible in practice, in particular because it would require a disproportionate effort in terms of time, cost and manpower” (our translation).

It then applied that test to the facts. Besides age, sex, conditions and medicines, the data included the date, sometimes the exact time, of the consultation or purchase, together with information that located or identified the healthcare professionals involved: in GERS's pharmacy data, the prescribers' national professional identifiers, entered into an ordinary search engine, revealed who they were. It was therefore possible to trace care pathways and single out patients. According to the CNIL's findings, endorsed by the court, this singling out “required little time and few resources, in particular the use of a commonly used spreadsheet program” (our translation).

Three lessons follow:

- **Replacing the name is not enough.** Indirect identifiers (dates, places, occupations, rare procedures) often make it possible to find a person.
- **What the holder actually does is irrelevant.** The court held that the fact that the companies drew no inferences themselves “has no bearing” (our translation): what matters is what the data makes possible.
- **Calling data anonymous is a claim you must back up.** The GDPR requires the controller to be able to demonstrate compliance (Article 5(2)): whoever claims a dataset is anonymous must be able to prove it.

## Assessing anonymisation: three criteria {#criteria}

In 2014, in their Opinion 05/2014 on anonymisation techniques, the European data protection authorities, then sitting as the Article 29 Working Party, set three criteria for assessing anonymisation: singling out, linkability and inference. The European Data Protection Board (EDPB), which succeeded the Working Party, updates and reformulates them in its [draft Guidelines 02/2026 on anonymisation](https://www.edpb.europa.eu/public-consultations/guidelines-022026-on-anonymisation_en), adopted on 7 July 2026 and open for public consultation until 30 October 2026.

- **No record isolation** No record may contain a combination of values that is unique in the dataset. Sex, date of birth and postcode are often enough to make a record unique.
- **No linkage** No record may be matched, with certainty or high likelihood, to a record about the same person in another dataset.
- **No inference** The data must not allow any specific and meaningful conclusion to be drawn about a particular person, even once aggregated.

*The three criteria of the EDPB's draft Guidelines 02/2026, building on the 2014 opinion.*

If all three criteria are met, the data can be considered anonymous. If any one fails, the analysis must continue before reaching a conclusion. The EDPB describes two ways to carry out the assessment: a contextual approach, which examines the means actually available to each party that could access the data, and a simplified, more cautious approach, which disregards those differences in means.

The draft stresses three practical points:

- **Assessments age.** Recital 26 requires account to be taken of available technology “and technological developments”. The draft recommends safety margins and observes that AI, agentic AI in particular, is likely to reduce further the time and cost of re-identification techniques. A data leak may also require a dataset that was anonymous until then to be reassessed.
- **The demonstration must be documented.** The method, tests and results must be kept, to prove both that the operation complied with the GDPR and that it was effective.
- **Words must be accurate.** The draft asks controllers not to use terms such as “anonymous”, “de-identified” or “de-personalised” when individuals are still identifiable.

## “Anonymous to whom?”: the SRB judgment {#relative}

The Court of Justice has long assessed identifiability in the light of the means actually available: in 2016, in Breyer (C-582/14), it took account of information that a party can lawfully obtain from a third party; in 2024, in OC v Commission (C-479/22 P), it looked at the matter from the standpoint of the journalists for whom a press release was intended.

On 4 September 2025, in [case C-413/23 P](https://curia.europa.eu/site/upload/docs/application/pdf/2025-09/cp250107en.pdf), between the European Data Protection Supervisor and the Single Resolution Board (SRB), it applied this reasoning to pseudonymised data passed to a third party. The judgment interprets Regulation 2018/1725, which applies to EU institutions and whose concepts mirror the GDPR's. The SRB had passed pseudonymised comments from shareholders and creditors to an audit firm. The Court held that pseudonymised data “must not be regarded as constituting, in all cases and for every person, personal data”: depending on the circumstances, pseudonymisation may prevent persons other than the controller from identifying the data subject.

The same judgment also set a limit. As regards the obligation to provide information, identifiability “must be assessed at the time of collection of the data and from the point of view of the controller”. The SRB therefore had to inform people of the transfer to the audit firm, whether or not the data was personal for that firm.

The EDPB's draft guidelines, for their part, set an important rule for outsourcing: when one party processes data on behalf of another, whether the data is personal is assessed from the perspective of the controller. Data that is personal for you therefore remains personal for your processor, even if it only receives pseudonyms. The relative approach only benefits genuinely independent recipients.

> **Key point** For you, holding the key, pseudonymised data is always personal. Under the EDPB's draft guidelines, a processor handling it on your behalf is in the same position: the relative approach of the SRB judgment then only applies to an independent recipient with no reasonable means of re-identifying the individuals.

## The Digital Omnibus: a redefinition under debate {#omnibus}

On 19 November 2025, as part of a simplification package known as the Digital Omnibus, the European Commission proposed to supplement the definition of personal data. Information would not be personal for a party that cannot identify the person with the means it is reasonably likely to use, even if another party can. The proposal also leaves it to implementing acts to specify when pseudonymised data ceases to be personal for certain parties.

The EDPB and the European Data Protection Supervisor oppose this in their [Joint Opinion 2/2026](https://www.edpb.europa.eu/system/files/2026-02/edpb_edps_jointopinion_202602_digitalomnibus_en.pdf): in their view, the change would narrow the concept of personal data and would go well beyond codifying case law. They urge the co-legislators not to adopt it and object to an implementing act determining the scope of the GDPR. Negotiations continue in the Council and the European Parliament.

> **Key point** As at the date of this guide, nothing has been adopted. The applicable law remains the GDPR as interpreted by the Court of Justice and, in France, by the Conseil d'État.

## Techniques and what they produce {#techniques}

The name of a technique says nothing about its result: the three criteria decide. The guidance below draws on the 2014 opinion, the [EDPB's draft Guidelines 01/2025 on pseudonymisation](https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2025/guidelines-012025-pseudonymisation_en) (put to consultation in January 2025, with no final version published to date) and the 2026 draft on anonymisation.

| Technique | Principle | Usual result |
| --- | --- | --- |
| Token replacement | Each identifier becomes a token; a mapping table leads back to the original value | Pseudonymisation. The table itself contains personal data to be protected |
| Encryption | The value is unreadable without the key | Pseudonymisation and a security measure: encryption is reversible by nature |
| Hashing | A one-way function produces a digest of the value | Pseudonymisation only with a secret key kept separately (a keyed function such as HMAC). Without a key, hashing candidate values (a list of names or email addresses) is enough to find matches. Either way, the result remains personal data |
| Removing direct identifiers | Name, address and numbers are deleted | Usually still personal data: indirect identifiers remain |
| Generalisation | A year instead of a date, a region instead of an address | Anonymisation possible, to be checked against the three criteria |
| Aggregation | Only statistics about groups are released | Anonymisation often achieved, except for small groups or cross-referenced tables |
| Adding noise | Values are perturbed, for example through differential privacy | Anonymisation possible depending on the settings, to be assessed |
| Synthetic data and trained models | Data or a model reproduces the properties of the source | Not necessarily anonymous: well-chosen queries can extract information about a person; anonymity must be shown case by case |

For AI models, the EDPB stated in its [Opinion 28/2024](https://www.edpb.europa.eu/documents/opinion-of-the-board-art-64/opinion-282024-on-certain-data-protection-aspects-related-to_en) that a model trained on personal data cannot, in all cases, be considered anonymous: the assessment is made case by case.

## What the classification changes {#obligations}

| Obligation | Pseudonymised data | Anonymous data |
| --- | --- | --- |
| Legal basis (Articles 6 and 9) | Required for each processing operation | Not applicable to the result; required for the anonymisation itself |
| Informing individuals (Articles 13 and 14) | Required, including about recipients | Not applicable to the result; the planned anonymisation must be stated clearly |
| Data subject rights (access, erasure, objection) | Applicable; Article 11 may disapply Articles 15 to 20 where the controller shows it is not in a position to identify the person | Not applicable |
| Data protection impact assessment (Article 35) | Under the usual criteria; pseudonymisation is one of the risk-reduction measures | Not applicable to the result |
| Transfers outside the EEA (Chapter V) | Regulated; pseudonymisation can supplement the safeguards | Outside Chapter V of the GDPR |
| Security (Article 32) | Required, especially for the key or mapping table | Outside the GDPR, but anonymity is reassessed after an incident |
| Retention period | Limited to what the purpose requires | Not limited by the GDPR |

The gap is considerable, which explains the temptation to call a dataset anonymous too quickly. That is precisely what the CNIL penalised and the Conseil d'État upheld.

## Free text and generative AI {#ai}

The criteria were designed for tables. Free text, such as an email, meeting notes or a request sent to a language model, is even harder to anonymise: once names are removed, job titles, places, dates and events remain, and they often point to a single person. “The finance director at the Lyon site, on sick leave since March” contains no name and is still not anonymous.

For these uses, the realistic goal is pseudonymisation: replace direct identifiers with tokens before sending to the model, keep the mapping table on your side, then restore the original values in the answer. It is an effective data minimisation measure, but the data received by the provider remains personal when it acts on your behalf. The guide [Pseudonymising data before sending it to an LLM](~/ressources/anonymiser-donnees-llm/) details the method and its limits.

## A six-step approach {#approach}

- **Set the purpose** Will you ever need to trace the individuals, to reply to them, follow up or correct their data? If so, aim for pseudonymisation: it is sufficient, provided it is presented as such.
- **Map the data** Direct identifiers, indirect identifiers, special-category data, and what recipients already hold.
- **Choose the technique** Tokens and a protected mapping table to pseudonymise; generalisation, aggregation or adding noise to aim for anonymity.
- **Check the three criteria** Record isolation, linkage, inference, from the perspective of each party that could access the data, with a safety margin.
- **Document and reassess** Keep the demonstration and review it when the data, the recipients or re-identification techniques change.
- **Use the right word** Write “anonymous” only if the demonstration exists; otherwise, speak of personal data, pseudonymised where that is the case.

*The choice between pseudonymisation and anonymisation depends on the purpose; anonymity, for its part, must be demonstrated.*

The overall framework, linking these choices to data location and applicable law, is set out in the guide [Data sovereignty and generative AI](~/ressources/souverainete-donnees-ia/). The other decisions to document before a rollout are in the guide [GDPR and generative AI](~/ressources/rgpd-ia-generative/).

## Frequently asked questions {#faq}

### Is pseudonymised data still personal data?

For the controller holding the information needed to trace it back to individuals, yes, always. Since the SRB judgment of 4 September 2025, it may cease to be personal for an independent recipient with no reasonable means of re-identifying the individuals. For a processor handling it on the controller's behalf, it remains personal, according to the EDPB's draft guidelines.

### Does hashing an email address make it anonymous?

No. Without a secret key, anyone with a list of addresses can hash them in turn and compare the digests. With a secret key kept separately, the result is pseudonymised: it remains personal data.

### Can pseudonymised data be presented as anonymous?

No. In its 2026 draft guidelines, the EDPB asks controllers not to call data “anonymous” when it still allows individuals to be identified. The Cegedim case also shows that the classification is assessed on the data itself, not on the way it is presented.

### Is anonymisation a processing operation subject to the GDPR?

Yes. The operation applies to personal data: it requires a legal basis, clear information for individuals and documentation. Only its result, if it is genuinely anonymous, falls outside the scope of the regulation.

### Will the Digital Omnibus change the definition of personal data?

Perhaps, but nothing has been adopted yet. In November 2025 the Commission proposed writing a party-by-party assessment into the GDPR; the EDPB and the European Data Protection Supervisor oppose it, and negotiations continue. In the meantime, the applicable law remains the GDPR as interpreted by the Court of Justice.

### SmartAGT pseudonymises, and says so
In hybrid mode, SmartAGT's privacy vault replaces direct identifiers with tokens before data is sent to an external provider's model, then restores the original values in the answer. The mapping table stays in your infrastructure.
This is a pseudonymisation measure, and we present it as such: the data remains personal, for you and for the provider acting on your behalf. For the most sensitive uses, the model can run entirely within your infrastructure. Details are on the [Security](~/security/) page.
