Data sovereignty and generative AI: what sovereign AI really guarantees
“Sovereign” has become a sales argument, rarely a definition. For generative AI, the question has to be asked flow by flow: who runs the processing, under which law, who can access the data, and can you leave. This guide gives the grid, the laws that apply and the checks to make.
Published September 22, 2026
What “sovereign” means for AI
The guide On-premise generative AI in the enterprise separates three notions that are often confused: where processing runs, which law applies, and who actually has access. For generative AI, a fourth has to be added: the ability to change provider without losing what you have built.
Where processing runs and where data is stored, including logs and backups.
Which law applies to the entity operating the service, and which authority can compel it to hand over data.
Who can technically read the data: the vendor, its support staff, its subcontractors, the model provider.
If you change model or provider, what you get back: indexed documents, configuration, history.
None of these properties follows from the other three. Hosting in France operated by a subsidiary of a US group settles location, not law. A European provider that keeps your prompts to improve its models settles law, not access. A local deployment on a proprietary format settles access, not reversibility.
Why generative AI makes the question sharper
A conventional business application handles structured data within a known scope. A generative AI assistant or agent receives whatever users paste into it: a contract, an employee file, a piece of code, a customer email. The content of the flows cannot be predicted, and it is often more sensitive than the project anticipated.
Five separate flows need examining, and they do not necessarily go to the same place:
- The prompt: the user's question, with everything attached to it.
- The added context: the document extracts a retrieval engine (RAG) inserts into the request sent to the model.
- The answer: it can repeat and recombine personal or confidential data.
- The logs: on your side, at the platform vendor, and at the model provider, each with different retention periods.
- The indexes: the vector representations of your documents, which are stored somewhere and can be partly reconstructed.
The model provider is not the only party. The platform vendor, the host and any web search provider each see part of the flows. Sovereignty is judged on the most exposed link.
The laws that apply, and what they require
Several texts overlap. They do not ask the same question, and none of them boils down to “host in Europe”.
| Text | What it covers | The question for you |
|---|---|---|
| GDPR | Any processing of personal data, including in prompts and answers | Lawful basis, impact assessment, processor contract, transfers outside the EU |
| Cloud Act (United States) | Service providers subject to US jurisdiction, wherever the data is stored | Who operates the service, and who controls that entity? |
| AI Act | Providers and deployers of AI systems, depending on the risk level of the use | Which uses are high-risk, which transparency duties apply? |
| SecNumCloud (ANSSI) | A French qualification scheme for cloud offerings, not a law | Is the provider qualified, and for which service exactly? |
| Data Act | Switching between data processing service providers | Can you leave, and at what cost? |
The GDPR remains the foundation. Transfers of personal data outside the EU are governed by its Chapter V, and its Article 48 provides that a decision by a third-country authority requiring data to be handed over is only recognised if it is based on an international agreement. The guide GDPR and generative AI sets out the questions to settle before deploying.
The Cloud Act is often quoted, rarely read. It targets the operator, not the data centre: that is the subject of the guide The Cloud Act and AI, which explains what is really exposed and how to check.
The AI Act adds AI-specific obligations, some of which fall on the organisation using a system, not only on the one building it. Its timetable was amended in 2026: the guide AI Act obligations for businesses sets out where things stand.
The SecNumCloud scheme run by ANSSI, the French national cybersecurity agency, turns the requirement of protection from non-European law into verifiable criteria. In version 3.2, it requires in particular that the provider's registered office be in the EU, and it caps the share of capital held by non-European entities (24% individually, 39% collectively). It is a useful benchmark even if you are not aiming for the qualification.
Finally, the Data Act, applicable since 12 September 2025, governs switching between cloud service providers and removes switching charges from 12 January 2027. It covers reversibility at infrastructure level; reversibility of your configuration and indexes remains yours to negotiate.
Mapping the data flows of an AI project
Sovereignty is not declared at project level. It is built by classifying data, then deciding, for each use, where each flow is allowed to go.
- Classify the dataAt minimum: public, internal, confidential, personal, and special-category personal data under Article 9 GDPR (health, opinions, trade union membership).
- List the usesFor each planned use, note which classes of data it actually handles, including what users are likely to paste in.
- Trace the flowsFor each use, follow the five flows: prompt, context, answer, logs, indexes. Note who receives them and under which law.
- Set one rule per classSpecial-category data: local processing only. Personal data: European provider or pseudonymisation before sending. Public data: unrestricted.
- Enforce it in the toolA rule that only exists in a policy document is not applied. It must become configuration: models allowed per use, personal data detection, logging.
This mapping produces a concrete result: the list of processing that must never leave your perimeter. Once that list exists, the architecture choice becomes a consequence, not a starting point.
Three architectures, three levels of guarantee
The choice between local, cloud and hybrid is covered in the guide On-premise, SaaS or hybrid. From a sovereignty standpoint alone, each option guarantees different things.
- Model run locally: no flow leaves, and all four questions are settled on your side. The trade-off is limited access to open models and an infrastructure to run.
- European provider: location and law can be controlled, provided you check the operator's shareholders and subcontractors, not just its registered office. Access depends on its terms of use: prompt retention, reuse for training.
- Non-European provider: the applicable law is not yours. Use remains possible for public data, or for data whose personal elements were removed before sending, which is the subject of the guide Anonymise or pseudonymise before an LLM.
In most organisations, all three coexist. What matters is that the routing rule is explicit and enforced by the tool, not left to each user.
The questions to ask a vendor
These questions apply to a platform vendor as much as to a model provider. Ask for written answers, and check them against the contract.
- Which legal entity operates the service, and what is its parent company? Where are their registered offices?
- Which sub-processors are involved, in which countries, and for which operations?
- Are prompts and answers retained, for how long, and for what purpose? Are they used to train models? Can this be excluded by contract?
- Who can access the data in operation: support, administrators, subcontractors? From which countries?
- What happens if outbound traffic is cut off? Whatever stops working describes exactly what leaves your information system.
- What do you get back when you leave, in which format, and how quickly?
- Which qualifications or certifications have actually been obtained, and for which scope? A certification announced as “in progress” is not a certification.
Ask for the operator's legal structure chart up to the ultimate parent company. The answer to the applicable-law question is there, far more reliably than on a sales page.
Common mistakes
Confusing hosting in France with legal immunity
A data centre in France operated by an entity subject to foreign law remains exposed to that law. The nationality of the operator and its parent company matters more than the address of the servers.
Looking only at the model
The platform's logs, the vendor's support and the web search service are flows too. A local model behind a platform that sends its telemetry abroad does not settle the question.
Believing a contract is enough
Contractual clauses frame a transfer; they do not stop a foreign authority from compelling the operator. Only the architecture reduces exposure; the contract allocates it.
Forgetting the exit
A project assessed on how it goes live, never on how it can be reversed, discovers its dependency the day prices or terms change.
Frequently asked questions
Is AI hosted in France sovereign?
Not necessarily. Hosting settles where processing runs, not the law applicable to the operator or actual access to the data. You need to check who operates the service, which group it belongs to, and what happens to prompts and logs.
Can you use a US model and remain sovereign?
For some uses, yes: those involving only public data, or data from which personal and confidential elements were removed before sending. Sensitive data must stay on a model run locally or at an operator whose applicable law you have verified.
Is SecNumCloud qualification mandatory?
Not for a private company in general. It is a qualification scheme run by ANSSI that may be required in some contexts, public ones in particular. Its criteria on protection from non-European law remain a good assessment tool even where it is not required.
Where SmartAGT fits
SmartAGT is deployed on-premise, within your perimeter. The model provider is your choice: 100% local, with no outbound traffic at all, or a cloud provider you contract with directly. The administrator can restrict the catalogue to European providers.
In hybrid mode, personal data is detected and reversibly pseudonymised, in a vault, before any outbound call, and no raw prompt is stored. Details are on the Security page.