Human approval of an AI agent: which actions to approve
Approving every action an agent takes amounts to approving nothing: by the fiftieth request, people confirm without reading. Approving nothing means letting a probabilistic model write into your tools. In between, you have to choose which actions to stop, and take care over how you present them.
Published September 22, 2026
What human approval protects against, and what it does not
The guide generative AI or rules sets out the principle: the model proposes, a rule checks, a person approves what commits. This guide goes further on that last step, known as human in the loop: which actions to stop, and how.
An agent can go wrong in three ways: it misreads the request, it relies on false information, or it is manipulated by content it reads, for instance an email containing a hidden instruction. In all three cases, the error only has consequences if it turns into an action. Human approval is the last point where that can be stopped.
It does not protect against everything. It does not fix a wrong answer that the user reads and adopts, it does not replace properly set permissions, and it is only worth something if the person approving understands what they see. It is one control among others, described alongside the rest in the guide enterprise agentic AI platform.
Classify actions, not agents
The common mistake is to set approval per agent: “this agent is safe, that one must have everything approved”. Yet the same agent chains harmless reads and writes that commit. The right unit is the action, meaning a call to a specific tool.
| Type of action | Examples | Human approval |
|---|---|---|
| Read | Search a document, read an email, look up a record | No, access rights are enough |
| Preparation | Draft a reply, propose an answer, compute an estimate | No, nothing leaves until someone uses it |
| Reversible internal write | Add a note to a ticket, create a task | Decide based on context |
| Write that commits | Change an order, update a contract, change a customer status | Yes |
| Outbound send | Send an email, post a message, share a file | Yes |
| Delete or execute | Delete a record, start a job, trigger a payment | Yes |
This classification must be attached to the tool when it is declared, not assessed by the model at call time. If the model is the one deciding that an action “does not need confirmation”, approval depends on the very source of error it is meant to control.
The criteria that tip an action over
The “reversible internal write” row is the tricky one. Six criteria help decide, and one is often enough to require approval:
- Irreversibility: once done, can the action be undone without harm? A sent email cannot be recalled.
- External recipient: does the effect leave the organisation, towards a customer, a supplier, the public?
- Volume: does the action affect one record or a thousand? A bulk update always deserves a stop.
- Sensitive data: does the action handle personal, health or financial data, or data covered by professional secrecy?
- Financial or legal commitment: does the action create an obligation, a payment, a contractual commitment?
- Effect on a person: does the action feed into a decision that affects someone, such as a refusal, a sanction, a processing priority?
The last criterion has legal weight. Article 22 of the GDPR gives data subjects the right not to be subject to a decision based solely on automated processing that produces legal effects concerning them or similarly significantly affects them and, where such a decision is allowed on the basis of a contract or consent, the right to obtain human intervention. A token approval does not meet that requirement.
Designing an approval step people still read
The main enemy of human approval is not malice, it is habit. Someone who sees fifty requests a day, all correct, ends up confirming without looking. The AI Act calls this “automation bias” and requires that the people in charge of oversight be made aware of it. The design of the approval step must counter it.
- Trigger deterministicallyThe stop depends on the nature of the tool called, never on the model's judgement.
- Show the exact actionThe recipient, the full content, the fields changed with their values before and after. Not a summary written by the model.
- Give useful contextWhy the agent proposes this action: the original request and the sources it used.
- Make rejection easyRejecting must be as easy as confirming, and the agent must stop cleanly after a rejection.
- Log the decisionWho confirmed or rejected, when, and what was displayed at that moment.
If the approver only sees a sentence like “The agent will reply to the customer”, they are approving nothing. They must see the exact message, to the exact recipient, as it will be sent.
Two settings complete the set-up. Grouping: when an agent prepares twenty actions of the same kind, presenting them together, each one readable, is better than twenty interruptions. Expiry: an approval request left unanswered must not execute by default. Silence means no.
Who approves, and what the AI Act says
In most use cases, the approver is the person who started the agent: they know the context and own the action. For high-stakes actions, a second look may be needed, following the “four eyes” principle already used for payments.
In every case, the person must have the competence to judge and the authority to refuse. That is also what the AI Act requires for high-risk systems. Article 14 requires that these systems can be effectively overseen by natural persons, who must in particular be able to understand the system's capabilities and limits, decide not to follow its output, and interrupt it. Article 26 requires deployers to assign this oversight to people with the necessary competence, training and authority.
Not every use of agents is high-risk within the meaning of the regulation, and the application timeline for these obligations has been revised by the digital omnibus: the European Commission page states they apply from 2 December 2027 for the high-risk areas listed in Annex III. Designing approval along these principles now avoids reworking the architecture the day a use case moves into that category.
The most common mistakes
Having everything approved
Systematic approval looks cautious and produces the opposite: reflex confirmations. Reads and drafts do not need to be stopped.
Letting the model decide what is sensitive
An instruction such as “ask for confirmation before important actions” depends on the model's judgement. Malicious content can precisely convince it that an action is not important.
Approving a summary
An approval step that shows a rewording of the action, rather than the action itself, lets through exactly the errors it was meant to stop.
Forgetting to log rejections
A rejection is information: it signals an agent error, an ambiguous instruction, sometimes an attempted hijack. It must be logged and reviewed, just like a confirmation.
Frequently asked questions
Does human approval slow agents down too much?
It slows down the actions it stops, and only those. If the classification is right, the agent works alone on reads and preparation, which make up most of the work, and only stops before effects that are hard to undo.
Can approval be relaxed once the agent has proven itself?
For reversible internal writes, yes, after a period in which human decisions have been logged and analysed. For outbound sends, deletions and financial commitments, an agent's past reliability does not protect against malicious content it will read tomorrow.
Is human approval enough to comply with the AI Act?
No. Human oversight is one of the requirements for high-risk systems, alongside risk management, record-keeping, transparency and data quality. It is one piece, not the whole.
Where SmartAGT fits
In SmartAGT, every write action is paused before execution by default: the nature of the action decides, not the model. The operator sees exactly what will happen and confirms or rejects.
The decision is logged with the human actor in a SHA-256 hash-chained audit trail that can be verified end to end. Details are on the Security page.