Contents
Choosing your architecture

On-premise, SaaS or hybrid AI: choosing use case by use case

The “on-premise or SaaS” question is almost always asked for the whole company, when it is really settled use case by use case. Summarising a press release and analysing a medical record do not call for the same architecture. This guide gives a method to classify your use cases and keep a hybrid model working over time.

Published September 22, 2026

The wrong reflex: one decision for the whole company

Many projects start with a debate of principle. On one side, the legal team wants nothing to leave. On the other, business teams want the best model on the market, right now. Each is right for some of the use cases, and wrong for the rest.

Picking a single deployment model for the whole organisation leads to one of two failures. Everything in SaaS, and sensitive use cases stay forbidden, or happen quietly on personal accounts. Everything in-house, and trivial use cases pay for infrastructure sized for the worst case, sometimes with less capable models.

The right unit of decision is not the company, it is the use case: a specific task, done by a specific population, on specific data. The definitions of on-premise, sovereign and private, and the architecture of an internal platform, are covered in the guide On-premise generative AI in the enterprise. Here, we start from those notions to make the call.

Classify use cases by the data they touch

The deciding criterion is neither the user's job nor the tool: it is the most sensitive data the use case can move around. Three levels are enough in most organisations.

Public or low-stakes data

Content already published, generic text, code with no secrets. A leak would have no notable consequence.

Internal data

Meeting notes, procedures, routine customer exchanges. They often contain personal data, which can be masked before sending.

Data that never leaves

Health records, legal case files, trade secrets, data covered by professional secrecy or a sector-specific obligation.

Three levels of sensitivity, three different architecture answers.

Applied to real cases, the classification produces a map of use cases rather than a doctrine:

Use caseMost sensitive dataSuitable model
Rewording a job ad or an articleNoneSaaS
Summarising internal meeting notesNames, internal decisionsHybrid, with personal data masked
Answering a customer complaintIdentity, contract, historyHybrid, or local model depending on the sector
Querying the HR document baseIndividual employee filesLocal model
Analysing a patient record or a legal case fileHealth data, professional secrecyLocal model, no outbound traffic

The table is filled in during a workshop with business teams, IT and the DPO. It then becomes the reference: a new use case is classified first, and the architecture follows.

Watch out

The most sensitive data is not always in the user's question. In a document assistant, it is the document excerpts added automatically to the prompt that go to the model. A seemingly harmless use case can carry an entire contract.

The signals that decide

When pure SaaS is enough

SaaS is the right choice when the use case only handles first-level data, when speed matters, or when the organisation has no team to run infrastructure. That is often true of a first pilot on public content.

You still need to read the contract: how long requests are kept, whether your data may be used to train models, where processing happens, the list of subprocessors, and the law that applies to the provider. That last point has consequences of its own, covered in the guide on the Cloud Act and AI. In France, for some sensitive data, a cloud service with the ANSSI SecNumCloud qualification is an intermediate option worth examining.

When on-premise is a must

In-house hosting, with a model running on your own infrastructure, is a must as soon as a use case touches the third level. It is also justified when the network must be able to run cut off from the outside, or when usage is heavy and steady, because sized capacity then becomes more predictable than a usage-based bill.

It comes at a price: models limited to those whose weights are published, GPU infrastructure to size, and a team to run it. Choosing the models themselves is the subject of the guide Open source or proprietary LLMs.

Hybrid: what makes it hold

Hybrid is not a lukewarm compromise between the two. It is a precise architecture: the platform, conversation history, indexed documents and identities stay with you, and only some model calls go to an external provider, according to written rules.

  1. The platform stays in-houseHistory, document bases, memory and accounts never leave your perimeter, whichever model is called.
  2. Each use case has its routing ruleThe agent or workspace is tied to a model approved for its sensitivity level. The choice does not rest with the user.
  3. Personal data is masked before sendingNames, addresses and numbers are replaced with tokens before the outbound call, then restored in the answer.
  4. What leaves is loggedWhich model, for which use case, at what volume: enough to answer an auditor and to steer spending.
A hybrid setup holds through its rules, not through each user's vigilance.

The masking technique deserves a closer look: anonymisation and pseudonymisation do not offer the same guarantees, nor carry the same obligations. The guide Anonymising data before an LLM details both approaches and their limits.

On the budget side, hybrid mixes a fixed part and a variable part. The variable part needs a ceiling, otherwise it ends up dictating the architecture for you. The costing method is in the guide on the cost of generative AI in the enterprise.

Mistakes to avoid

Deciding by infrastructure rather than by data

“We have GPUs, let's do everything in-house” or “we are already with this cloud provider, let's stay there”: two lines of reasoning that start from what exists instead of from the use cases. Available infrastructure is a constraint, not a classification criterion.

Letting users pick the model

A drop-down menu offering every model to everyone, with a word of caution, does not survive the first emergency. The routing rule must be set by the administrator, at the level of the use case.

Forgetting secondary flows

The prompt is not the only outbound flow. Document excerpts, attachments, tool results and sometimes the platform's telemetry leave too. Take a full inventory before classifying a use case as hybrid.

Freezing the choice

A use case classified as SaaS today may move in-house tomorrow, because an open model becomes good enough or the contract changes. If your use cases are written for a given model, that move becomes a project. A platform where the model is an interchangeable resource makes it reversible.

Frequently asked questions

Is hybrid harder to run than fully in-house?

It adds external providers to monitor, contracts to manage and a masking layer to maintain. In return, it avoids sizing internal infrastructure for use cases that do not need it. The complexity stays under control if routing is centralised rather than left to each team.

Is masking personal data enough to send any document?

No. Masking deals with identifiable personal data, not with trade secrets, strategic information or data covered by professional secrecy. A confidential document with no names in it is still confidential: it belongs on a local model.

Can we start with SaaS and move on-premise later?

Yes, provided you planned for it: use cases described independently of the model, history and documents not stored at the provider, and an exit path with no loss. Without those precautions, migrating means rebuilding.

Where SmartAGT fits

SmartAGT is deployed on-premise, inside your perimeter. The model provider remains your choice: 100% local, air-gap possible with no outbound traffic at all, or a cloud provider you contract with directly.

In hybrid mode, personal data is detected and reversibly pseudonymised, in a vault, before any outbound call, and no raw prompt is ever stored. The administrator can restrict the catalogue to European providers. See the Security page and the pricing model.

SOVEREIGN BY ARCHITECTURE

A question these guides
do not settle?