# AI quotas and spending caps: controlling consumption day to day

> Once generative AI is open to users, spending no longer depends on a contract but on thousands of daily requests. Quotas, caps and alerts let you manage it without stifling usage, provided they are set at the right level and tuned on real data.

Source: https://smartagt.ai/en/ressources/quotas-ia/
Published: 2026-09-22
Publisher: SmartAGT (NERVIAL LABS)

---
## Why consumption drifts {#drift}

Most model providers bill by the volume of text processed, counted in tokens, with separate prices for the text sent to the model and the text it produces. The cost of a request is therefore not fixed: it depends on what the request contains and what it triggers. Four factors explain most of the variation.

- **Context** A long document attached to the request, a conversation history that keeps growing, document extracts injected by search: all of it is billed on every call.
- **Model** For the same request, the price varies widely from one model to another, from the smallest to the most capable, and from one provider to another.
- **Agents** An agent that plans, searches, reviews and corrects chains several calls for a single user request.
- **Automations** A scheduled or event-driven process consumes with nobody watching, including when it loops.

*Four drivers of consumption, two of which are invisible to the user.*

What these factors have in common: users do not see the cost of what they ask for. Management cannot rest on awareness alone; it needs technical limits.

## Quota, cap, alert: three separate tools {#tools}

The three terms are often used interchangeably. They do not meet the same need.

| Mechanism | What it does | When to use it | Risk if badly tuned |
| --- | --- | --- | --- |
| Alert | Notifies an owner when a threshold is crossed, without blocking anything | Always, first | Too many alerts, which nobody reads any more |
| Quota | Limits a volume of usage over a period (requests, tokens) | To share a common resource, especially an in-house server | Blocks legitimate use at the end of the period |
| Spending cap | Blocks consumption beyond an amount over a period | To bound a financial risk, per user or per group | Cut-off in the middle of a task if set too low |

The combination that works: an alert at a first threshold, so that a person takes a look, then a block at the cap, so that the error does not continue. A cap without a prior alert cuts off without warning; an alert without a cap lets a loop run all weekend.

For a model running on your own servers, the marginal cost of a request is low, but capacity is limited. There, a quota serves to share the resource at peak times more than to contain spending.

## At which level to set limits {#levels}

A limit only makes sense if it is attached to someone who can act. Four levels complement each other:

- **The whole organisation**: a global cap, rarely reached, that protects against a major incident. It is the safety net, not the management tool.
- **The group**: a department, a team, a project. This is the natural budget level, with an owner who receives the alerts and makes the calls.
- **The user**: an individual cap, set well above normal usage, that stops anomalies (a script in a loop, a compromised account, a handling error) without getting in the way of everyday work.
- **The agent or the model**: track and, if needed, limit a given agent or access to the most expensive model, reserved for the uses that justify it.

Limits stack: a user under their individual cap can still be stopped by their group's. Make this rule visible, or the block will look arbitrary.

These choices follow from the overall governance (who decides budgets, who rules on exceptions), described in the guide on [generative AI governance in the enterprise](~/ressources/gouvernance-ia/).

## Sizing limits without holding back usage {#sizing}

Setting limits before observing usage means inventing them. They will be too low, and block the most active users, often those who get the most value from the tool, or too high, and stop nothing.

- **Observe** For a few weeks, alerts only, no blocking. Measure consumption per user, per group, per agent and per model.
- **Set group envelopes** Based on observed consumption and expected growth, with a margin. Each envelope has an owner.
- **Set individual caps** Well above the usage of the most active users, so that they only stop anomalies.
- **Review** Monthly at first, then quarterly: adjust thresholds, examine the blocks that occurred and the exception requests.

*Limits tuned on observed usage, then reviewed: never set once and for all.*

Every block must leave a way out: a clear message stating which limit was reached and whom to contact, and a quick procedure to raise it. A user blocked without explanation goes back to a public tool, which recreates the [shadow AI](~/ressources/shadow-ai/) the internal tool was meant to reduce.

## Tracking consumption day to day {#tracking}

Tracking serves two audiences: budget owners, who want to know where the money goes, and the team running the platform, who want to catch an anomaly before it becomes expensive.

The useful indicators:

- **Cost per agent**, to see which use cases consume and compare that with the value they produce.
- **Cost per model and per provider**, to check that the most expensive model has not quietly become the default.
- **Cost per active user**, by group, to spot gaps and the uses that need support.
- **Daily trend**, to see a break on the day it happens, not when the invoice arrives.

Two conditions make these figures reliable. First, a cost expressed in your budget's currency: a count in tool-specific credits does not reconcile with any accounting line. Second, up-to-date prices: provider price lists change, and some bill in dollars, which means exchange rates have to be taken into account.

> **Watch out** Tracking consumption is not the same as the full cost. Licence, infrastructure and operations belong to the purchasing decision, covered in the guide on the [cost of generative AI](~/ressources/cout-ia-generative/).

## Mistakes to avoid {#mistakes}

### The same quotas for everyone

A lawyer analysing hundred-page contracts and a salesperson rewording emails do not consume the same amount. A single cap is either too low for one or pointless for the other. Differentiate by group.

### Forgetting automated processes

Scheduled agents and batch jobs consume with no user in front of the screen. They need their own cap, attached to a named owner.

### Managing by the monthly invoice

An invoice arrives weeks after the spending. A loop that starts on a Friday evening is discovered the following month. Alerts must be daily, or real-time for caps.

### Using quotas to discourage usage

Caps set deliberately low to slow adoption have the opposite effect: usage moves to ungoverned tools, with no limit and no log.

## Frequently asked questions {#faq}

### How can we limit generative AI costs without slowing users down?

By tuning limits on observed usage rather than on estimates, by setting individual caps above normal usage so that they only stop anomalies, and by reserving the most expensive models for the uses that justify them.

### What is the difference between a quota and a spending cap?

A quota limits a volume of usage (number of requests or tokens) over a period; a cap limits an amount spent. A quota shares capacity, a cap bounds a financial risk.

### What happens when a user reaches their cap?

Their requests are blocked until the next period or until the limit is raised. The block must show which limit was reached and whom to contact, and raising it must be quick for a legitimate need.

### Where SmartAGT fits
SmartAGT lets you set quotas and caps per user, per group or for everyone, with alerts and hard blocking. Cost is tracked per agent, per model and per provider, in real money rather than credits, with price lists published by the vendor.
You contract directly with your AI providers: SmartAGT takes no margin on inference. The business model is explained on the [Pricing](~/pricing/) page.
