The real cost of generative AI in the enterprise: building your TCO
The price of generative AI almost always comes as a single figure: a subscription per user, a price per million tokens, a server quote. None of them tells you what the platform will really cost over its lifetime. This guide lists the lines to cost, the traps in pricing grids, and a method to build a full cost before you sign.
Published September 22, 2026
Why the list price tells you almost nothing
Comparing generative AI offers on their list price is like comparing cars on the price of a full tank. A per-user subscription says nothing about integration with your information system. A price per token says nothing about how many tokens real usage consumes. A server quote says nothing about the team that will have to run it.
The total cost of ownership (TCO) adds up everything the platform will cost between the purchase decision and its replacement: fixed spending, variable spending and internal spending, which is often invisible because it is paid in team time.
For an in-house deployment, the guide On-premise generative AI in the enterprise sets out the four cost lines of a platform hosted on your own infrastructure. This guide widens the scope to every deployment model and focuses on the costing method.
The lines of a full cost
A serious TCO covers eight lines. The last three are missing from most comparisons, even though they often weigh as much as the first ones.
| Line | Nature | What makes it vary |
|---|---|---|
| Licence or subscription | Fixed | Number of users, modules, commitment period |
| Inference | Variable (API) or capacity-based (GPU) | Real volume, context size, chosen model |
| Supporting infrastructure | Capacity-based | Storage of indexed documents, vector database, backups |
| Integration | One-off | Directory, connectors to business tools, network |
| Operations | Recurring | Updates, monitoring, user support |
| Document quality | Recurring | Sorting, updating and governing indexed sources |
| Compliance | One-off then recurring | Impact assessment, legal reviews, audits |
| Adoption support | One-off then recurring | Training, business champions, community of users |
Two lines deserve special attention. Document quality, because an assistant plugged into outdated documents costs in errors what it saves in searching. Compliance, because an impact assessment done after deployment costs more than one done before, and may force you to rework the architecture.
What drives the inference bill up
Inference is the hardest line to forecast, because it depends on mechanisms the user never sees. Four factors explain most of the gap between the estimate and the bill.
In a conversation, the history is sent back to the model on every turn. The tenth question usually costs noticeably more than the first.
A document assistant adds passages from your documents to every question. They, rather than the question itself, often weigh the most.
A task given to an agent triggers several calls: plan, call a tool, read the result, answer. A single question can trigger many of them.
Using the most powerful model for everything, including sorting an email, multiplies the bill with no gain in quality.
Two billing details come on top. Input tokens and output tokens are often priced differently: a use case that writes a lot does not cost the same as one that reads a lot. And some providers bill in US dollars, so your cost in euros follows the exchange rate, which an annual budget must allow for.
Reading a pricing model
For the same usage, two offers with a similar list price can diverge sharply over time. It all depends on how the price reacts as usage grows.
| Pricing model | Predictability | What to check |
|---|---|---|
| Per user, inference included | High for the buyer | Hidden usage limits, throttling or a weaker model beyond a threshold |
| Per user, inference recharged | Low | Margin applied to inference, real price of the model provider |
| Prepaid credits | Low | Conversion between credits and tokens, whether it can change mid-contract |
| Licence plus inference contracted directly | Medium, manageable | Available spending caps, cost tracking per use case |
| In-house GPU capacity | High, in steps | Real utilisation rate, cost of the next step |
The decisive question for a vendor fits in one sentence: “if our usage doubles, what happens to our bill, and who benefits?” If the vendor takes a margin on inference, your success funds its growth. Credits raise a different problem: they hide the real unit price and make any comparison impossible.
Building a defensible TCO
A TCO is there to support a decision, not to reassure. It must therefore rest on measurements and give a range, not a single figure.
- Scope the use casesWhich use cases, for which populations, on which data. Each use case has its own consumption profile.
- Measure on a pilotRequest volume, average length of exchanges, peak hours, share of agent tasks. A week of measurement beats a month of assumptions.
- Project adoptionUsage grows in steps, not all at once. Model a gradual ramp-up, with a low and a high scenario.
- Cost every lineThe eight lines of the table, over the planned lifetime of the platform, separating fixed, variable and internal time.
- Cost the exitWhat it takes to recover data, history and use cases if you switch solutions. It is the line nobody writes down.
The break-even point between SaaS and in-house hosting comes naturally out of this exercise. It depends on your volume, your time horizon and the sensitivity of your data: no generic figure can replace it. Choosing the architecture, use case by use case, is covered in the guide On-premise, SaaS or hybrid.
Mistakes to avoid
Costing on the demo
A demo uses short questions, with no history and no documents. Real usage is longer, more document-heavy and more agentic. A budget based on a demo is a wrong budget.
Comparing different units
A per-user price, a per-token price and a server price cannot be compared directly. Bring everything back to a cost per measured use case, over the same period.
Forgetting that variable spending must be managed
A TCO assumes consumption stays within the planned range. Without caps and alerts, nothing keeps it there. Day-to-day control is the subject of the guide AI quotas.
Choosing the model before measuring
The choice of model changes the cost considerably. A small model can be enough for a large share of routine tasks. The guide Open source or proprietary LLMs explains how to evaluate it on your own cases.
Frequently asked questions
How much does generative AI cost for a 500-person company?
No honest answer fits in one figure: the cost depends on the use cases chosen, the share of agents, the models and the deployment mode. A pilot of a few weeks, on a real business scope, provides the measurements needed to build a reliable range.
Is on-premise always cheaper over time?
No. It turns variable spending into capacity spending, which pays off with heavy and steady usage. With low or very irregular usage, infrastructure sized for peak hours stays underused, and SaaS costs less.
Should inference be budgeted per user or per use case?
Per use case. Two users can consume very different amounts depending on whether they ask short questions or run agents on long documents. Budgeting per use case lets you set consistent caps and spot overruns.
Where SmartAGT fits
SmartAGT is sold as an annual software licence per user, with every capability included. You sign directly with your AI providers, and SmartAGT takes no margin on inference.
Cost is tracked per agent, per model and per provider, in real money rather than credits, with price lists published by the vendor. Quotas and caps per user, per group or for everyone trigger alerts, then a hard block. The PoC runs on your premises, with your models and your data. See the pricing model.