On-premise generative AI in the enterprise: what to know before deciding.
On-premise, sovereign, private: three words used as synonyms, when they describe three different things. This guide untangles the definitions, describes the real architecture of an internal generative-AI platform, breaks down what it actually costs, and gives you a grid to decide.
On-premise, sovereign, private: three distinct notions
Confusing these terms is the leading cause of bad decisions. A vendor can be entirely sincere in saying “sovereign” while meaning something other than what you understood. Three different questions hide behind these words: where the processing runs, who holds legal power over the data, and who can technically access it.
Where does the processing run?
How to checkBy watching outbound network traffic.Which law applies to the operator?
How to checkBy the jurisdiction of the operating entity.Who can technically access the data?
How to checkThrough telemetry and support channels.On-premise: a question of where it runs
On-premise means the software runs on infrastructure you control: your servers, your data centre, or a private cloud whose perimeter you own. It is an architectural property, not a contractual promise. It can be verified: you can watch the outbound network traffic.
Sovereign: a question of applicable law
Sovereignty is not location. Data hosted in Europe, but operated by an entity subject to extraterritorial law, remains exposed to that law. The right question is not “where is my data” but “which law applies to the entity operating it, and who can be compelled to hand it over”. That is what separates European hosting from genuine legal immunity.
Private: a question of effective access
A private platform is one nobody but you can reach, the vendor included. Many so-called private solutions keep a channel open for telemetry, support or model improvement. That is not illegitimate, but it must be explicit, and you must be able to switch it off.
Ask your vendor one question: “if I cut all outbound Internet access from the platform, what stops working?” The answer describes exactly what leaves your information system. A genuinely on-premise platform keeps running, apart from the external models you chose to plug in.
Why organisations choose on-premise
The reflex is to read this as a defensive, fear-driven stance. In practice, organisations making this choice give four reasons, and only one is about compliance.
- The nature of the data. Patient records, legal case files, defence data, insurance pricing, source code: in some lines of business, sending data to a third party is simply not negotiable, whatever contract is signed.
- Legal exposure. GDPR, professional secrecy and sector obligations create a liability that stays yours, even when the processing is delegated. Architecture is the only way to reduce that exposure rather than transfer it on paper.
- Cost control. Usage-based billing on growing volume is hard to forecast. Internal infrastructure turns a variable expense into a capacity decision you size and cap.
- Independence from a single vendor. Models change fast, and so do their prices and terms. A platform that lets you swap models without redeploying your use cases keeps your strategy off a catalogue you do not control.
These four motives do not carry equal weight everywhere. An IT department that only has the third one is probably better served by a hybrid model than by a full deployment.
The real architecture of an internal platform
Deploying generative AI in an enterprise is not just hosting a model. The model is the most visible brick, rarely the most expensive to operate. A complete platform has five layers.
- The inference service. The engine running the models, typically vLLM or Ollama for local ones. It consumes the GPU, and its sizing determines how many concurrent users you can support.
- The document layer, or RAG. Ingesting, chunking and indexing your documents, then the retrieval that feeds the model. This is what separates a generic assistant from one that knows your organisation. It is also the layer that demands the most ongoing operational work.
- Identity and permissions. Connection to your directory (LDAP, OIDC, SAML) and access control that also applies to indexed documents. A document base that ignores your permissions turns an internal search engine into a company-wide data leak.
- Governance. Who is entitled to which model, under which spend cap, over which connectors, with which human approval on sensitive actions. Without this layer, a platform stays a tolerated prototype rather than an enterprise service.
- Observability and audit. The record of what was asked, by whom, with which model, against which sources, and at what cost. This is what makes the service defensible to an auditor and manageable by IT.
Projects that fail almost always underestimate the last three layers. They demo poorly, and they are precisely what decides whether the platform scales.
What it actually costs
Comparing a per-seat subscription to the price of a GPU server makes no sense: they do not cover the same things. A complete internal platform cost breaks into four lines.
Sized on concurrent users, not on headcount.
CapacityCheck: does the vendor take a margin on inference?
FixedUpdates, monitoring, document sources, support.
RecurringOnly if SaaS models are connected in hybrid mode.
Variable- Compute infrastructure. The most visible item. It depends on the chosen model and above all on concurrent users, not on registered ones. Real usage is heavily concentrated in a few hours of the day.
- Software licensing. The platform itself. Check above all whether the vendor takes a margin on inference: if so, your bill grows with your usage, which reproduces exactly the model you were trying to leave.
- Operations. Updates, monitoring, adding document sources, supporting business teams. This is a recurring team cost, often missing from comparisons, and rarely zero.
- External inference, if any. If you also plug in SaaS models for certain uses, that consumption is on you. Contracting directly with the provider guarantees you know its real price.
The economics turn on volume and duration. Below a certain usage level, SaaS stays cheaper. Above it, and especially over several years, capacity spending becomes more predictable and often lower. The tipping point depends far too much on your situation for any generic figure to be worth anything.
The most common pitfalls
These four mistakes recur project after project, and none of them is technical in nature.
Assuming local means compliant
Hosting a model yourself does not make you compliant. GDPR is about purpose, data minimisation, retention periods and individual rights. A poorly governed internal platform that indexes personal data without control is a compliance problem, not a solution.
Sizing on registered users
Sizing follows concurrency and the length of the contexts processed, not headcount. A project sized on the number of accounts is either oversized and expensive, or unusable at peak hours.
Neglecting document quality
A RAG plugged into an outdated document base produces wrong answers with all the confidence of a sourced one. That is more dangerous than no answer at all, because the error is credible. Work on the sources comes before deployment, not after.
Locking onto a single model
Building use cases on the quirks of one given model creates a strong dependency at the very moment the landscape moves fastest. A model-agnostic platform, where the model is an interchangeable resource, protects the investment made in the use cases.
On-premise, SaaS or hybrid: how to choose
The choice is not binary, and the right answer is often a mix: sensitive processing in-house, ordinary uses on an external model. This grid compares the three options on the criteria that actually decide.
| Criterion | On-premise | SaaS | Hybrid |
|---|---|---|---|
| Data leaving your systems | None, by architecture | Systematic, governed by contract | Chosen case by case, anonymised before sending |
| Time to go live | Longest, depends on your infrastructure | Immediate | In between |
| Cost structure | Capacity-based, predictable | Variable, grows with usage | Mixed, needs capping |
| Operational burden | On you | On the vendor | Shared |
| Access to frontier models | Limited to open models | Immediate | Immediate, on approved uses |
In practice, the useful question is not “which option should we pick” but “which processing must never leave”. Once that list exists, the architecture follows almost mechanically.
Where SmartAGT fits
SmartAGT is a governed AI-agent platform deployed on-premise inside your information system. Memory, document bases, history and identities stay with you. In hybrid mode, personal data is anonymised before any call to an external model.
The platform is model-agnostic, takes no margin on your inference, and includes the identity, governance and audit layers described above. This guide holds whichever tool you end up choosing: it describes criteria, not a product.
Frequently asked questions
Is on-premise generative AI GDPR-compliant by default?
No. Internal hosting sharply reduces exposure, but compliance is about the purpose of the processing, data minimisation, retention periods and individual rights. An internal platform indexing personal data without control would be a compliance problem despite its hosting.
How much compute do you need to start?
Sizing depends on concurrent users and the length of the contexts processed, not on declared headcount. A pilot on a narrow business scope lets you measure real usage before committing to production infrastructure, which avoids both oversizing and saturation at peak hours.
Can you use models like GPT or Gemini and still be on-premise?
Yes, in hybrid mode. The platform stays deployed on your premises, and you decide model by model and use by use what may be sent to an external provider. An anonymisation layer before sending lets you benefit from those models on content stripped of personal data.
How is this different from an assistant like ChatGPT Enterprise or Copilot?
Those are excellent individual assistants, but built for personal use: no shared enterprise context, no central governance of permissions and spending, and data transits through the provider. An agent platform aims for the opposite: agents connected to your data, with access control, observability and audit at organisation scale.