Executive summary
An AI agent that can only suggest is a chatbot. Once it can issue a refund, change a customer record or raise a purchase order, it is doing work your controls were designed for people to do. Many enterprises now have an AI policy. Far fewer have written down which actions an agent may take on its own, who approves the rest, and how to stop a run that has gone wrong.
That gap is what keeps agents stuck in pilots. Risk teams can't sign off what they can't see. The answer is not a longer policy document. It is a small set of controls built into the platform the agents run on, so governance happens on every run rather than in a yearly review. Our position is simple: an agent should never hold more authority than the person whose work it is doing, and every step that carries risk should wait for a named human.
1. Why agents need their own controls
Traditional software does what it was coded to do. An agent chooses its next step at run time, from a goal, the data in front of it and the tools it has been given. That is what makes it useful. It is also why the usual controls don't quite fit.
Three failure modes are worth planning for from day one:
- Too much authority. The OWASP Top 10 for LLM applications calls this excessive agency: an agent given more functions, permissions or autonomy than its task needs.
- Untrusted content steering actions. An agent that reads emails, tickets or web pages can be instructed by what it reads. If it can act on those instructions, so can whoever wrote them.
- No trail. When something goes wrong, "the model decided" is not an answer your auditor, regulator or customer will accept.
Regulators are arriving at the same place. The EU AI Act requires high-risk AI systems to be built so that people can oversee them, understand their limits and step in or stop them (Article 14). In India, the Reserve Bank's FREE-AI committee report (August 2025) sets out expectations for responsible AI in financial services, and MeitY's India AI Governance Guidelines (November 2025) put accountability and human oversight at the centre. Frameworks such as the NIST AI Risk Management Framework and ISO/IEC 42001 give you the management system around all of this.
HMR insight: An agent should never hold more authority than the person whose work it is doing.
2. The six controls
Policies describe intent. Controls make it happen on every run. We build agent platforms around six of them.
Approval gates
RuleEvery action has a risk tier. Higher tiers wait for a named person.
BuildA policy check runs before each tool call and records who approved it, and when.
Scoped identities
RuleThe agent acts as itself, with only the rights its task needs.
BuildOne workload identity per agent, short-lived tokens, no standing admin rights.
Data boundaries
RuleAn agent reads and writes only the data its task needs.
BuildAllow-listed sources, personal data masked before it reaches a prompt, region pinning.
Audit trail
RuleEvery decision, tool call and approval is recorded.
BuildStructured traces sent to a log store that the agent cannot change.
Stop and roll back
RuleAnyone authorised can pause a run. Completed steps can be undone.
BuildA stop switch at the gateway, and a defined undo action for every tool that writes.
Model controls
RuleOnly approved models, in approved regions, on approved terms.
BuildA model gateway with an allow-list. Versions are pinned and changes are reviewed.
Tier every action before you automate it
The approval gate only works if each action has a risk tier agreed in advance by the business owner and the risk team. We use four:
| Tier | Examples | What happens |
|---|---|---|
| 0: Read | Look up an order, summarise a ticket | Runs on its own, logged |
| 1: Reversible write | Draft a reply, tag a record, add a note | Runs on its own, sampled weekly |
| 2: Customer-facing or financial | Send a message, refund within a set limit | A named approver in the business |
| 3: Irreversible or high value | Large payments, deleting data, access changes | Two approvers, one from risk |
Start strict. It is far easier to loosen a gate with a month of clean logs behind you than to tighten one after an incident.
Put one gateway between agents and your systems
The most useful design decision is architectural. Every tool call passes through a single gateway that holds the policy checks, the agent identities, the approval queue and the audit trail. Controls then live in one place, not inside each agent, and a new agent inherits them on day one.
- Agent plans a stepIt picks a tool and the inputs, for example "refund order 4821".
- Gateway checks policyIdentity, data scope and risk tier are checked before anything runs.
- A person approvesOnly for higher tiers. The approver sees the step, the reason and the data.
- Action runs and is loggedThe result, the approval and the trace go to the audit store.
Logging should follow an open standard so your existing monitoring can read it. The OpenTelemetry semantic conventions for generative AI cover model calls, and the same trace can carry each tool call and approval.
| Aspect | Typical first pilot | Governed agent platform |
|---|---|---|
| Who acts | A shared service account with broad rights | A scoped identity per agent |
| Approvals | A policy document, reviewed yearly | A gate enforced on every risky step |
| Evidence | Application logs, if anyone kept them | A trace of every decision and tool call |
| Stopping a run | Switch off the integration | A stop switch and per-step undo |
| Model changes | Vendor updates arrive unannounced | Pinned versions, changes reviewed |
3. Illustrative scenario: a motor insurer's claims desk
This scenario is illustrative. The company is fictional.
Kaveri General Insurance wants an agent to help its claims handlers. Handlers spend much of the day chasing documents and writing status updates, and customers wait. An earlier pilot was stopped by the risk team because the agent could trigger settlement payments directly.
The governed design splits the work by tier. The agent reads the claim and checks it against the policy (tier 0). It asks the customer for missing documents and updates the claim notes (tier 1). For customers, it drafts status messages that a handler approves before they are sent (tier 2). It can recommend a settlement amount with its reasoning, but the payment itself is tier 3: the handler and a team lead both approve.
Claim data stays in the insurer's cloud region, personal details are masked before they reach the model, and every step lands in the audit trail. What changes is the shape of the handler's day. They review and approve instead of typing and chasing, and the risk team can see every action the agent took and who signed it off. That visibility is what lets the second pilot go live.
4. Putting it in place
- Weeks 1 to 3AssessThe Readiness Audit: pick candidate workflows, tier every action, check identity and logging foundations.
- Weeks 4 to 8Build the controlsGateway, policy checks, agent identities and the audit trail, proven on one workflow.
- Weeks 9 to 12Run with tight gatesApprovals on anything that writes. The business owner reviews a sample of runs each week.
- After thatWiden with evidenceLoosen a gate only when the log shows it is safe. Add the next workflow.
- Assess. Our three-week Readiness Audit picks the workflows worth automating first, tiers every action in them and checks your identity and logging foundations.
- Build. Stand up the gateway, policy checks, agent identities and audit trail, and prove them on one workflow with tight gates.
- Run and improve. Review a sample of runs with the business owner every week at first. Loosen a gate only on evidence from the log. Review each live agent at least every six months.
The rest of this series takes the model further:
- Where AI belongs in your operating model (Part 1)
- Data boundaries for AI systems (Part 3)
- Human approval that scales (Part 4)
- Measuring AI risk and return (Part 5)
- Responding to an AI incident (Part 6)
Moving forward
Start with the controls and one workflow, not with a dozen agents and a policy to catch up later. If you want to see how we apply these rules to our own work, read our AI governance policy. If you want them applied to your workflows, book a readiness audit.
Sources
- OWASP, LLM06:2025 Excessive Agency
- EU Artificial Intelligence Act, Article 14: Human oversight
- Reserve Bank of India, FREE-AI committee report, August 2025 (summary by KPMG India)
- MeitY, India AI Governance Guidelines, November 2025 (summary by DSCI)
- NIST, AI Risk Management Framework
- ISO, ISO/IEC 42001:2023 AI management systems
- OpenTelemetry, Semantic conventions for generative AI