Generative AI from POC to production: deploying what has proved its worth
A generative AI proof of concept almost always succeeds as a demo. The move to production fails on something else: real data, access rights, costs, operations. This guide explains how to scope the test so that it answers those questions, and how to roll out afterwards.
Published September 22, 2026
Why so many POCs never reach production
A generative AI POC shows that a model can do a task. Production requires it to do so on every real case, for every user concerned, with their rights, at a known cost, and with a team able to run it. The gap between the two explains most of the projects that stop.
The causes recur from one project to the next:
- The use case was chosen to impress, not to be measured. A spectacular demo on a marginal topic justifies no rollout budget.
- The data was the demo data: a clean, hand-picked set of documents, without the outdated versions, duplicates and odd formats of the real repository.
- Nobody owns the outcome. The technical team delivered, the business liked it, but no one owns the go-live decision or the service afterwards.
- The invisible layers were missing: identity, document permissions, logs, spending caps, monitoring. Absent from the POC, they turn the move to production into a second project, longer than the first.
The practical consequence: a useful POC is designed from the start to answer the questions of production, not only the question “can the model do it”.
Scoping the POC so it can go through
Scoping is decided before the first line of configuration. It comes down to four choices.
A task repeated often, by an identified population, on a defined set of documents. Frequency creates the value, the boundary makes measurement possible.
Documents and files as they actually are, flaws included, and the real access rights. A POC on cleaned-up data measures a case that will never exist.
The models, hosting and integrations planned for production. Changing model or architecture after the POC invalidates part of its results.
Someone who needs the result, takes part in the evaluation and will own the go-live decision.
The ideal use case for a first POC is neither the most ambitious nor the most visible: it is the one whose value is easy to measure. Answering a department's recurring questions from its documentation, preparing a first-level reply to incoming requests, summarising files ahead of a decision. The guide on enterprise RAG explains what makes a good document assistant, the most common case.
Setting exit criteria before you start
A POC without exit criteria written in advance always ends with “it looks promising”. The criteria are set with the business owner, and the thresholds are decided before seeing the results.
| Criterion | How to measure it | What to decide beforehand |
|---|---|---|
| Answer quality | Assessment by business experts on a representative sample, edge cases included | The required rate of acceptable answers, and what counts as a serious error |
| Benefit for users | Time spent on the task before and during the POC, for the same people | The minimum gain that justifies rollout |
| Adoption | Active users among testers, week after week | The expected level of use without reminders |
| Cost per task | Actual consumption divided by the number of tasks handled | The maximum acceptable cost at scale |
| Security and compliance | Access rights respected, logs available, sign-off from the CISO and DPO | The blocking conditions |
The evaluation sample is the most often neglected point. It must be built before the test, from real cases, and include the hard requests: ambiguous, out of scope, based on contradictory documents. A model judged on ten hand-picked examples tells you nothing about the rest.
What changes between POC and production
Going to production is not a matter of opening the POC to more people. Several dimensions change in nature.
| Dimension | During the POC | In production |
|---|---|---|
| Users | A few volunteer testers | A whole population, with its unexpected uses |
| Data | A defined scope | Sources that change and must be kept up to date |
| Rights | Often simplified | Those of the directory, applied to documents and tools |
| Actions | Reading and drafting | Possibly actions in business tools, to be approved |
| Costs | A test budget | Envelopes, caps and daily tracking |
| Operations | The project team | Monitoring, support, updates, incident procedure |
ANSSI, the French national cybersecurity agency, in its security recommendations for a generative AI system, recommends a security audit and business functional tests before deployment to production (recommendations R23 and R24), as well as a degraded mode for business services without the AI system (R15). The last point is often forgotten: if the service stops, work must be able to continue.
Two subjects call for an explicit decision before opening up. The actions an agent may execute in connected tools, and those that require human confirmation, covered in the guide on human approval. And consumption limits per group and per user, described in the guide on AI quotas and spending caps.
Rolling out step by step
- Extended pilotThe POC population plus a few teams, on the production environment, with real rights and logging switched on.
- Opening by populationGroup by group, starting with those whose need is clearest, with trained champions in each team.
- Continuous measurementThe same criteria as during the POC, tracked over time: quality, adoption, cost per task, incidents.
- Next use casesThe second case reuses the platform, rights and operations of the first: this is where the initial investment pays off.
The fourth point is why each use case should not be treated as an isolated project. A shared platform, with a common model catalogue, rights and logs, avoids rebuilding the same layers every time. That is the purpose of generative AI governance in the enterprise, which sets the framework each new use case fits into.
Mistakes to avoid
Running POC after POC without deciding
Multiplying experiments gives an impression of activity. Without a go or stop decision at the end of each one, the organisation piles up demos and no service.
Changing model between POC and production
A POC's results hold for the model tested. A different model, chosen later for cost or hosting reasons, needs a new evaluation on the same sample.
Postponing the question of rights
An assistant that worked in the POC with global access to documents may, once opened, expose content some users were not meant to see. Real rights apply from the extended pilot onwards.
Measuring satisfaction instead of usage
Testers who are enthusiastic in a meeting do not necessarily use the tool three weeks later. Real, measured usage decides.
Frequently asked questions
How long should a generative AI POC last?
Long enough to measure real usage beyond the novelty effect, and not so long that it becomes an unofficial service. The duration follows from the exit criteria: you need time to build the sample, test, and observe adoption over several weeks.
Should the POC use the company's own data?
Yes. A POC on demo data measures a case that does not exist. That is precisely why hosting and access rights must be settled during the POC, not after.
Who decides to move to production?
The business owner, based on the criteria set before the test, with input from IT on operations, from the CISO on security, and from the DPO where personal data is processed.
Where SmartAGT fits
SmartAGT is tested as a PoC at your site, with your models and your data. The platform is deployed on-premise: each connector acts with the rights of the connected account, every write or execute action waits for human approval, and cost is tracked per agent, per model and per provider.
The licence is annual, per user, with all capabilities included: a new use case does not add a module to buy. Examples are on the Use cases page.