Anonymise or pseudonymise data before sending it to an LLM
Stripping names from a prompt before sending it to an external model has become a reflex. You still need to know what you end up with: anonymous data, which falls outside the GDPR, or pseudonymised data, which stays inside it. The difference changes the risk analysis, and what you can promise your DPO.
Published September 22, 2026
Two notions the law keeps apart
The GDPR defines pseudonymisation in Article 4: processing such that data can no longer be attributed to a person without the use of additional information, kept separately and protected. Anonymous data, on the other hand, is information that no longer identifies a person, taking into account the means reasonably likely to be used; Recital 26 places it outside the regulation's scope.
Identifiers are replaced by tokens. The mapping exists and is kept apart. Reversible by design, and the data remains personal for whoever holds the key.
Any reasonable possibility of re-identification is removed, irreversibly. The data falls outside the GDPR, but the result is rarely usable for free text.
The CNIL, the French data protection authority, sums it up: pseudonymised data keeps its personal character, whereas anonymisation makes identification impossible in practice.
Why a prompt is hard to anonymise
European data protection authorities, in Opinion 05/2014 of the Article 29 Working Party, which the CNIL relies on, set three criteria to judge anonymisation: it must no longer be possible to single out a person, to link records about them across datasets, or to infer new information about them.
Free text often fails all three. Remove the name from an email and what remains is the job title, the department, the city, a date, an event. “The finance director at the Lyon site, on leave since March” points to a single person in most companies. A language model is precisely good at piecing together clues of this kind.
There is also a practical constraint. For an answer to be useful, the real names often have to go back in: a letter must be addressed to the right person, a case summary must name the right parties. Irreversible anonymisation makes that impossible. In most generative AI uses, pseudonymisation is therefore what is workable, with its limits.
What 2025 case law changed
On 4 September 2025, in case C-413/23 P, EDPS v SRB, the Court of Justice of the EU held that pseudonymised data must not be regarded, in all cases and for every person, as personal data. Depending on the circumstances, pseudonymisation can prevent a recipient without the key from identifying the data subject. The judgment concerns the regulation that applies to EU institutions, whose concepts mirror the GDPR's.
The same decision sets two limits. Whether a person is identifiable is assessed case by case, according to the means available to each party. And the controller's obligations, including informing people about recipients, are assessed from its own point of view, at the time of collection, and remain in full.
The EDPB's Guidelines 01/2025 on pseudonymisation, put to public consultation in 2025, stress what it brings: a measure that helps meet data minimisation, data protection by design and security.
For you, the controller holding the mapping table, pseudonymised data remains personal. The judgment opens a question about its status at the model provider; it does not exempt you from the impact assessment, nor from framing the transfer.
A pseudonymisation pipeline before sending
Pseudonymisation that works for generative AI follows a precise path, and the mapping table never leaves your perimeter.
- DetectFind the direct identifiers in the prompt and in the document context: names, email addresses, phone numbers, postal addresses, national insurance numbers, IBANs, contract numbers.
- Replace consistentlyThe same person gets the same token throughout the request, so the model can follow the reasoning.
- Keep the mapping locallyThe table is stored in a vault, within your perimeter, with restricted access.
- Send and restoreThe model works on tokens; the real values are put back into the answer when it returns, on your side.
- Log without the raw textLogs keep the pseudonymised version, never the original prompt in clear.
What automatic detection misses
No detector guarantees full coverage. In particular, it misses:
- Indirect identifiers: job title, department, place, date, rare event. These are what enable re-identification, and they match no known format.
- Special-category data spelled out: “following her cancer”, “union representative”. No name, but Article 9 GDPR data.
- Unexpected formats: typos, internal identifiers, rare or foreign names, text extracted from a scanned document.
That is why pseudonymisation is only one layer of protection. It is combined with a usage rule: special-category data and the most confidential files go to a model run locally, not to an external provider, even after pseudonymisation.
Choosing by use
| Use | Suitable technique | What the external provider sees |
|---|---|---|
| Generic drafting, monitoring, non-sensitive code | None, no personal data | The text as is |
| Customer correspondence, routine case summary | Pseudonymisation with restoration | Tokens and indirect identifiers |
| Statistics, test datasets, training | Anonymisation, checked against the three criteria | Aggregated or transformed data |
| Health, sensitive HR, litigation | Local model, nothing sent out | Nothing |
This table applies, for personal data, the mapping described in the guide Data sovereignty and generative AI. Questions of lawful basis and impact assessment are in the guide GDPR and generative AI, and exposure to foreign laws in the guide The Cloud Act and AI.
Frequently asked questions
Is pseudonymised data still personal data?
For whoever holds the mapping table, yes. Since the Court of Justice judgment of 4 September 2025, it may not be for a recipient with no reasonable means of re-identifying people, depending on the circumstances. Your obligations as controller are unchanged.
Is replacing names enough to anonymise a text?
No. Free text contains indirect identifiers (job title, place, date, event) that often make it possible to find the person. The result is at best pseudonymised, and must be treated as personal data.
Can pseudonymised health data be sent to an external model?
Pseudonymisation reduces the risk, but automatic detection does not catch every health reference written out in words. For such data, a model run locally remains the safest option, and the impact assessment must address the question explicitly.
Where SmartAGT fits
In hybrid mode, SmartAGT's privacy vault, once enabled, detects direct identifiers (email, phone, IBAN, payment card, IP address, French NIR, SIREN, SIRET, licence plate, and the terms you add) and reversibly pseudonymises them before any call to an external provider. The mapping table stays within your perimeter, and no raw prompt reaches the audit log.
For the most sensitive uses, the model can be 100% local, with no outbound traffic at all. Details are on the Security page.