Contents
Governance and rollout

Generative AI from POC to production: deploying what has proved its worth

A generative AI proof of concept almost always succeeds as a demo. The move to production fails on something else: real data, access rights, costs, operations. This guide explains how to scope the test so that it answers those questions, and how to roll out afterwards.

Published September 22, 2026

Why so many POCs never reach production

A generative AI POC shows that a model can do a task. Production requires it to do so on every real case, for every user concerned, with their rights, at a known cost, and with a team able to run it. The gap between the two explains most of the projects that stop.

The causes recur from one project to the next:

  • The use case was chosen to impress, not to be measured. A spectacular demo on a marginal topic justifies no rollout budget.
  • The data was the demo data: a clean, hand-picked set of documents, without the outdated versions, duplicates and odd formats of the real repository.
  • Nobody owns the outcome. The technical team delivered, the business liked it, but no one owns the go-live decision or the service afterwards.
  • The invisible layers were missing: identity, document permissions, logs, spending caps, monitoring. Absent from the POC, they turn the move to production into a second project, longer than the first.

The practical consequence: a useful POC is designed from the start to answer the questions of production, not only the question “can the model do it”.

Scoping the POC so it can go through

Scoping is decided before the first line of configuration. It comes down to four choices.

A frequent, bounded use case

A task repeated often, by an identified population, on a defined set of documents. Frequency creates the value, the boundary makes measurement possible.

Real data

Documents and files as they actually are, flaws included, and the real access rights. A POC on cleaned-up data measures a case that will never exist.

The target environment

The models, hosting and integrations planned for production. Changing model or architecture after the POC invalidates part of its results.

A business owner

Someone who needs the result, takes part in the evaluation and will own the go-live decision.

Four scoping choices that decide what follows, made before you start.

The ideal use case for a first POC is neither the most ambitious nor the most visible: it is the one whose value is easy to measure. Answering a department's recurring questions from its documentation, preparing a first-level reply to incoming requests, summarising files ahead of a decision. The guide on enterprise RAG explains what makes a good document assistant, the most common case.

Setting exit criteria before you start

A POC without exit criteria written in advance always ends with “it looks promising”. The criteria are set with the business owner, and the thresholds are decided before seeing the results.

CriterionHow to measure itWhat to decide beforehand
Answer qualityAssessment by business experts on a representative sample, edge cases includedThe required rate of acceptable answers, and what counts as a serious error
Benefit for usersTime spent on the task before and during the POC, for the same peopleThe minimum gain that justifies rollout
AdoptionActive users among testers, week after weekThe expected level of use without reminders
Cost per taskActual consumption divided by the number of tasks handledThe maximum acceptable cost at scale
Security and complianceAccess rights respected, logs available, sign-off from the CISO and DPOThe blocking conditions

The evaluation sample is the most often neglected point. It must be built before the test, from real cases, and include the hard requests: ambiguous, out of scope, based on contradictory documents. A model judged on ten hand-picked examples tells you nothing about the rest.

What changes between POC and production

Going to production is not a matter of opening the POC to more people. Several dimensions change in nature.

DimensionDuring the POCIn production
UsersA few volunteer testersA whole population, with its unexpected uses
DataA defined scopeSources that change and must be kept up to date
RightsOften simplifiedThose of the directory, applied to documents and tools
ActionsReading and draftingPossibly actions in business tools, to be approved
CostsA test budgetEnvelopes, caps and daily tracking
OperationsThe project teamMonitoring, support, updates, incident procedure

ANSSI, the French national cybersecurity agency, in its security recommendations for a generative AI system, recommends a security audit and business functional tests before deployment to production (recommendations R23 and R24), as well as a degraded mode for business services without the AI system (R15). The last point is often forgotten: if the service stops, work must be able to continue.

Two subjects call for an explicit decision before opening up. The actions an agent may execute in connected tools, and those that require human confirmation, covered in the guide on human approval. And consumption limits per group and per user, described in the guide on AI quotas and spending caps.

Rolling out step by step

  1. Extended pilotThe POC population plus a few teams, on the production environment, with real rights and logging switched on.
  2. Opening by populationGroup by group, starting with those whose need is clearest, with trained champions in each team.
  3. Continuous measurementThe same criteria as during the POC, tracked over time: quality, adoption, cost per task, incidents.
  4. Next use casesThe second case reuses the platform, rights and operations of the first: this is where the initial investment pays off.
A gradual rollout, where each step carries over the criteria of the previous one.

The fourth point is why each use case should not be treated as an isolated project. A shared platform, with a common model catalogue, rights and logs, avoids rebuilding the same layers every time. That is the purpose of generative AI governance in the enterprise, which sets the framework each new use case fits into.

Mistakes to avoid

Running POC after POC without deciding

Multiplying experiments gives an impression of activity. Without a go or stop decision at the end of each one, the organisation piles up demos and no service.

Changing model between POC and production

A POC's results hold for the model tested. A different model, chosen later for cost or hosting reasons, needs a new evaluation on the same sample.

Postponing the question of rights

An assistant that worked in the POC with global access to documents may, once opened, expose content some users were not meant to see. Real rights apply from the extended pilot onwards.

Measuring satisfaction instead of usage

Testers who are enthusiastic in a meeting do not necessarily use the tool three weeks later. Real, measured usage decides.

Frequently asked questions

How long should a generative AI POC last?

Long enough to measure real usage beyond the novelty effect, and not so long that it becomes an unofficial service. The duration follows from the exit criteria: you need time to build the sample, test, and observe adoption over several weeks.

Should the POC use the company's own data?

Yes. A POC on demo data measures a case that does not exist. That is precisely why hosting and access rights must be settled during the POC, not after.

Who decides to move to production?

The business owner, based on the criteria set before the test, with input from IT on operations, from the CISO on security, and from the DPO where personal data is processed.

Where SmartAGT fits

SmartAGT is tested as a PoC at your site, with your models and your data. The platform is deployed on-premise: each connector acts with the rights of the connected account, every write or execute action waits for human approval, and cost is tracked per agent, per model and per provider.

The licence is annual, per user, with all capabilities included: a new use case does not add a module to buy. Examples are on the Use cases page.

SOVEREIGN BY ARCHITECTURE

A question these guides
do not settle?