Responsible AI in the Enterprise · Part 2

A governance model for enterprise AI agents

Six controls that let AI agents run real workflows while named people stay in charge of every step that carries risk, and how to put them in place.

Executive summary

An AI agent that can only suggest is a chatbot. Once it can issue a refund, change a customer record or raise a purchase order, it is doing work your controls were designed for people to do. Many enterprises now have an AI policy. Far fewer have written down which actions an agent may take on its own, who approves the rest, and how to stop a run that has gone wrong.

That gap is what keeps agents stuck in pilots. Risk teams can't sign off what they can't see. The answer is not a longer policy document. It is a small set of controls built into the platform the agents run on, so governance happens on every run rather than in a yearly review. Our position is simple: an agent should never hold more authority than the person whose work it is doing, and every step that carries risk should wait for a named human.

1. Why agents need their own controls

Traditional software does what it was coded to do. An agent chooses its next step at run time, from a goal, the data in front of it and the tools it has been given. That is what makes it useful. It is also why the usual controls don't quite fit.

Three failure modes are worth planning for from day one:

  • Too much authority. The OWASP Top 10 for LLM applications calls this excessive agency: an agent given more functions, permissions or autonomy than its task needs.
  • Untrusted content steering actions. An agent that reads emails, tickets or web pages can be instructed by what it reads. If it can act on those instructions, so can whoever wrote them.
  • No trail. When something goes wrong, "the model decided" is not an answer your auditor, regulator or customer will accept.

Regulators are arriving at the same place. The EU AI Act requires high-risk AI systems to be built so that people can oversee them, understand their limits and step in or stop them (Article 14). In India, the Reserve Bank's FREE-AI committee report (August 2025) sets out expectations for responsible AI in financial services, and MeitY's India AI Governance Guidelines (November 2025) put accountability and human oversight at the centre. Frameworks such as the NIST AI Risk Management Framework and ISO/IEC 42001 give you the management system around all of this.

HMR insight: An agent should never hold more authority than the person whose work it is doing.

2. The six controls

Policies describe intent. Controls make it happen on every run. We build agent platforms around six of them.

Figure 1 Six controls for agents that act
  • Approval gates

    RuleEvery action has a risk tier. Higher tiers wait for a named person.

    BuildA policy check runs before each tool call and records who approved it, and when.

  • Scoped identities

    RuleThe agent acts as itself, with only the rights its task needs.

    BuildOne workload identity per agent, short-lived tokens, no standing admin rights.

  • Data boundaries

    RuleAn agent reads and writes only the data its task needs.

    BuildAllow-listed sources, personal data masked before it reaches a prompt, region pinning.

  • Audit trail

    RuleEvery decision, tool call and approval is recorded.

    BuildStructured traces sent to a log store that the agent cannot change.

  • Stop and roll back

    RuleAnyone authorised can pause a run. Completed steps can be undone.

    BuildA stop switch at the gateway, and a defined undo action for every tool that writes.

  • Model controls

    RuleOnly approved models, in approved regions, on approved terms.

    BuildA model gateway with an allow-list. Versions are pinned and changes are reviewed.

Tier every action before you automate it

The approval gate only works if each action has a risk tier agreed in advance by the business owner and the risk team. We use four:

Tier Examples What happens
0: Read Look up an order, summarise a ticket Runs on its own, logged
1: Reversible write Draft a reply, tag a record, add a note Runs on its own, sampled weekly
2: Customer-facing or financial Send a message, refund within a set limit A named approver in the business
3: Irreversible or high value Large payments, deleting data, access changes Two approvers, one from risk

Start strict. It is far easier to loosen a gate with a month of clean logs behind you than to tighten one after an incident.

Put one gateway between agents and your systems

The most useful design decision is architectural. Every tool call passes through a single gateway that holds the policy checks, the agent identities, the approval queue and the audit trail. Controls then live in one place, not inside each agent, and a new agent inherits them on day one.

Figure 2 One governed step, from plan to log
  1. Agent plans a stepIt picks a tool and the inputs, for example "refund order 4821".
  2. Gateway checks policyIdentity, data scope and risk tier are checked before anything runs.
  3. A person approvesOnly for higher tiers. The approver sees the step, the reason and the data.
  4. Action runs and is loggedThe result, the approval and the trace go to the audit store.

Logging should follow an open standard so your existing monitoring can read it. The OpenTelemetry semantic conventions for generative AI cover model calls, and the same trace can carry each tool call and approval.

Aspect Typical first pilot Governed agent platform
Who acts A shared service account with broad rights A scoped identity per agent
Approvals A policy document, reviewed yearly A gate enforced on every risky step
Evidence Application logs, if anyone kept them A trace of every decision and tool call
Stopping a run Switch off the integration A stop switch and per-step undo
Model changes Vendor updates arrive unannounced Pinned versions, changes reviewed

3. Illustrative scenario: a motor insurer's claims desk

This scenario is illustrative. The company is fictional.

Kaveri General Insurance wants an agent to help its claims handlers. Handlers spend much of the day chasing documents and writing status updates, and customers wait. An earlier pilot was stopped by the risk team because the agent could trigger settlement payments directly.

The governed design splits the work by tier. The agent reads the claim and checks it against the policy (tier 0). It asks the customer for missing documents and updates the claim notes (tier 1). For customers, it drafts status messages that a handler approves before they are sent (tier 2). It can recommend a settlement amount with its reasoning, but the payment itself is tier 3: the handler and a team lead both approve.

Claim data stays in the insurer's cloud region, personal details are masked before they reach the model, and every step lands in the audit trail. What changes is the shape of the handler's day. They review and approve instead of typing and chasing, and the risk team can see every action the agent took and who signed it off. That visibility is what lets the second pilot go live.

4. Putting it in place

Figure 3 A sequence we recommend for the first agent in production
  1. Weeks 1 to 3AssessThe Readiness Audit: pick candidate workflows, tier every action, check identity and logging foundations.
  2. Weeks 4 to 8Build the controlsGateway, policy checks, agent identities and the audit trail, proven on one workflow.
  3. Weeks 9 to 12Run with tight gatesApprovals on anything that writes. The business owner reviews a sample of runs each week.
  4. After thatWiden with evidenceLoosen a gate only when the log shows it is safe. Add the next workflow.
  1. Assess. Our three-week Readiness Audit picks the workflows worth automating first, tiers every action in them and checks your identity and logging foundations.
  2. Build. Stand up the gateway, policy checks, agent identities and audit trail, and prove them on one workflow with tight gates.
  3. Run and improve. Review a sample of runs with the business owner every week at first. Loosen a gate only on evidence from the log. Review each live agent at least every six months.

The rest of this series takes the model further:

Moving forward

Start with the controls and one workflow, not with a dozen agents and a policy to catch up later. If you want to see how we apply these rules to our own work, read our AI governance policy. If you want them applied to your workflows, book a readiness audit.

Sources

  1. OWASP, LLM06:2025 Excessive Agency
  2. EU Artificial Intelligence Act, Article 14: Human oversight
  3. Reserve Bank of India, FREE-AI committee report, August 2025 (summary by KPMG India)
  4. MeitY, India AI Governance Guidelines, November 2025 (summary by DSCI)
  5. NIST, AI Risk Management Framework
  6. ISO, ISO/IEC 42001:2023 AI management systems
  7. OpenTelemetry, Semantic conventions for generative AI
  • Responsible AI in the Enterprise · Part 3

    Data boundaries for AI systems

    What an AI system may read, send, keep and where it may run: permission-aware retrieval, masking, provider terms and DPDP duties.

    10 min read

  • Responsible AI in the Enterprise · Part 4

    Human approval that scales

    How to keep human approval of AI agent actions meaningful at volume: risk tiers, clear approval cards, batching, fatigue checks and evidence.

    9 min read