Responsible AI in the Enterprise · Part 5

Measuring AI risk and return

How to tell if an AI system is worth running and safe to keep: a baseline, eight measures, a risk register with owners and a one-page board report.

Executive summary

Every AI pilot comes with a business case. Very few come with a way to check it. Six months in, the sponsor says the system saves time, the finance team can't find the saving, and the risk team has no record of what went wrong. The board is asked to fund the next phase on belief.

That is a common way for AI work to end. Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, and names three causes: rising costs, unclear business value and weak risk controls. All three are measurement problems. You can't defend a cost you never tracked, or a gain you never baselined.

Our view is that return and risk belong on the same page, measured from the same log. An AI system is worth running when each completed task costs less than it did before, quality holds, and every known risk has a named owner. This article sets out the measures we use, how to take a baseline before launch, and how to put the result in front of a board.

1. Why AI business cases are hard to check

Three habits make AI results hard to prove.

  • No baseline. Teams measure the new process with care and guess at the old one. The saving is then the gap between a number and a guess.
  • Counting the wrong unit. "Prompts answered" or "hours saved" say little. What matters is work finished to the right standard, at a known cost.
  • Return and risk in separate rooms. Finance tracks spend. Risk tracks incidents. Nobody sees that the cheapest month was also the month approvals were waved through.

The firms that get value from AI tend to fix this early. In McKinsey's The state of AI in 2026, the small group of high performers (6% of respondents) were about twice as likely as others to report defined processes to measure the impact of AI. Measurement is not paperwork added at the end. It is part of how the value gets made.

HMR insight: If you can't show what a task cost before the AI arrived, you can't show what the AI saved.

2. Take a baseline before launch

The baseline is the one step you can't do later. Once the new system is live, the old process is gone.

Start with a single workflow and define its unit of work: one invoice matched, one claim settled, one support case closed. Then measure how that work runs today, on real cases, for several weeks:

  • Cost: people time per case, from a time study or system timestamps, at a loaded hourly rate agreed with finance.
  • Cycle time: from request in to outcome out, including waiting time.
  • Quality: errors found by a reviewer in a random sample, plus cases reopened or redone.
  • Risk: incidents, complaints and audit findings linked to the workflow over the past year.

Two rules keep the baseline honest. First, measure the same way you will measure the new system, from the same sources. Second, where you can, keep part of the work on the old path for the first weeks of live running. A side-by-side comparison beats a before-and-after one, because it removes seasonal swings and other changes happening at the same time.

3. Eight measures that answer two questions

The board has two questions about any AI system. Is it worth running? Is it safe to keep running? We track four measures for each.

Figure 1 Eight measures: four for return, four for risk
  • Cost per completed task

    AsksIs each finished piece of work cheaper than before?

    MeasureAll run costs, including review time, over tasks done without rework.

  • Cycle time

    AsksDoes work finish sooner, end to end?

    MeasureFrom request in to outcome out, across the whole process.

  • Quality and rework

    AsksIs the output right the first time?

    MeasureErrors found in a weekly sample, plus work sent back to be redone.

  • Model and platform cost

    AsksIs spend tracking the plan as volume grows?

    MeasureTokens, hosting and gateway cost per month and per task.

  • Approval override rate

    AsksHow often does a person change or reject a proposal?

    MeasureOverrides over all tier 2 and 3 steps, read with approval time.

  • Incidents and near misses

    AsksDid anything go wrong, and how badly?

    MeasureCount by severity, with time to detect and time to stop.

  • Control coverage

    AsksIs every step running through the controls?

    MeasureShare of tool calls that passed the gateway and left a trace.

  • Open risks with owners

    AsksWho is on the hook for each known risk?

    MeasureRegister items by status, with the date each was last reviewed.

Every measure needs a definition, a data source and an owner, agreed before launch. A shared table stops arguments later.

Measure Where the data comes from Owner
Cost per completed task Ledger, cloud bill, approval log Finance partner
Cycle time Workflow timestamps Process owner
Quality and rework Weekly review sample, reopen count Process owner
Model and platform cost Token and hosting metrics Platform lead
Approval override rate Gateway approval log Risk lead
Incidents and near misses Incident register Risk lead
Control coverage Gateway traces Platform lead
Open risks Risk register Named owner per risk

Most of these come from the audit trail described in Part 2. If every tool call and approval passes through one gateway, the measures are a query, not a project.

4. Cost per completed task: the number finance needs

Cost per task is the clearest test of return, but only if both halves are honest.

The top half is the full cost to run the system for a period: model usage, hosting, the gateway and monitoring, and the people time spent reviewing, approving and fixing its work. Leave out review time and a system can look cheap while it keeps a team busy checking it. The bottom half is tasks completed to standard, with no rework. Tasks the agent attempted and a person then redid don't count.

The FinOps Foundation's guidance on AI costs takes the same line. It suggests pairing token and inference costs with measures of the work done, such as cost per call or time to close. For the raw usage, the OpenTelemetry conventions for generative AI define standard fields for input and output tokens on each model call. Tag each call with the workflow and task it served, and model cost per task follows.

Watch the trend as well as the level. Model prices, prompt length and retries all move. A system that was cheap at pilot volume can drift upward as edge cases pile up.

5. The risk measures

Approval override rate

This is the share of tier 2 and tier 3 steps where the approver rejects or edits what the agent proposed. It is the best early signal you have, and it needs reading in both directions.

A rising rate means the agent is getting things wrong, often after a data or model change. A rate near zero for months, with approvals given in seconds, may mean people have stopped reading. Part 4 covers how to design approvals that stay meaningful. Here, track the override rate next to the time each approval took, and review a sample of approved steps each month.

Incidents and near misses

Agree severity levels before launch and count incidents by level. Count near misses too: a wrong refund caught at the approval gate is a free lesson. For each incident, record the time to detect and the time to stop. Those two numbers tell you whether your stop controls work. Part 6 covers the response itself.

Drift from what you tested

The NIST AI RMF playbook asks organizations to monitor how measures seen in production differ from those collected in testing (MEASURE 2.4). In practice, run the same test set after every model or prompt change, and compare live quality against the pre-launch result each month.

A risk register with owners

Measures show what is happening. The register says who acts when a limit is crossed. Keep it short, specific to the workflow and reviewed monthly.

Risk Signal we watch Owner Action at the limit
Wrong output reaches a customer Errors in weekly sample Process owner Tighten the gate
Data leaves its boundary Masking tests, data alerts Data protection officer Stop and review
Cost runs over plan Monthly spend, cost per task Finance partner Review model and prompts
Approvals become a formality Override rate, approval time Risk lead Retrain, resample
A model update changes behaviour Test set score Platform lead Roll back the version

Each owner is a named person, not a team. Data risks link to the boundaries in Part 3. If you run an AI management system under ISO/IEC 42001, this register can sit inside it.

6. Presenting to the board

Boards don't need a dashboard. They need one page each quarter, with a decision at the bottom.

Figure 2 The quarterly cycle that keeps an AI system earning its place
  1. MeasureThe same eight measures, from the log and the ledger, each month.
  2. CompareAgainst the baseline and the limits agreed before launch.
  3. DecideWiden, hold or stop. The owner states which, and why.
  4. RecordThe decision goes in the register, with the date of the next review.

We lay the page out in four parts:

  1. Return. Cost per completed task and cycle time, each against the baseline, with the trend.
  2. Quality. Error and rework rates from the sample, against the limits agreed at launch.
  3. Risk. Override rate, incidents by severity, and any register items past their limit, each with its owner.
  4. Decision. Widen, hold or stop, with the reason. If you want more scope or more autonomy, say which gate you want to loosen and what evidence supports it.

Regulators expect this kind of oversight to reach the top. The Reserve Bank's FREE-AI committee report calls for each regulated entity to have a board-approved AI policy, with standard formats for AI incident reporting (summary by Chambers and Partners). A quarterly page built from the log is the simplest way to show the board oversees the system as well as the policy.

7. Illustrative scenario: a manufacturer's payables team

This scenario is illustrative. The company is fictional.

Beas Auto Components wants an agent to match supplier invoices to purchase orders and goods receipts. Matched invoices go to a finance officer as payment proposals (tier 2). Payments above a set value need a second approver from risk (tier 3).

Before build, the payables team times a sample of invoices across a full month-end. They log how long each takes, how many are exceptions, and how many payments later need correcting. That is the baseline. For the first eight weeks of live running, one supplier group stays on the manual path as a comparison.

The first board page shows cost per matched invoice and cycle time against both the baseline and the comparison group. It shows the override rate, and a note that approvals on large payments take real time, which is the sign that people are reading them. One near miss is listed: a duplicate invoice caught at the gate. It has an owner and a fix. The decision asked is to add a second supplier group and keep every gate as it is. The board can agree because every line on the page traces back to the log.

8. Putting it in place

Figure 3 When to measure, from audit to the first board report
  1. Weeks 1 to 3AssessThe Readiness Audit: pick the workflow, agree the measures, start the baseline.
  2. Weeks 4 to 8Finish the baselineKeep timing the old process while the system is built. Set the limits.
  3. Weeks 9 to 12Run and compareLive with tight gates. Keep part of the work on the old path to compare.
  4. Each quarterReport and decideOne page to the board: return, risk and the decision you want.
  1. Assess. Our three-week Readiness Audit picks the workflow, defines its unit of work, agrees the eight measures and their owners, and starts the baseline.
  2. Build. Wire the measures into the gateway and the finance data while the system is built, so the first live day produces the first data point.
  3. Run and improve. Report monthly to the process owner and quarterly to the board. Loosen a gate only when the measures support it.

Moving forward

Pick the AI system you are most likely to be asked about at the next board meeting. Write down its unit of work and what one unit costs today. If nobody can answer, start there. To see how we govern AI in our own delivery, read our AI governance policy. To have the baseline and measures set up with you, book a readiness audit.

Sources

  1. Gartner, Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027, press release, 25 June 2025
  2. McKinsey, The state of AI in 2026: On the road to ROI, August 2026
  3. FinOps Foundation, FinOps for AI overview
  4. OpenTelemetry, Semantic conventions for generative AI
  5. NIST, AI RMF Playbook: Measure
  6. Reserve Bank of India, FREE-AI committee report, August 2025 (summary by Chambers and Partners)
  7. ISO, ISO/IEC 42001:2023 AI management systems
  • Responsible AI in the Enterprise · Part 3

    Data boundaries for AI systems

    What an AI system may read, send, keep and where it may run: permission-aware retrieval, masking, provider terms and DPDP duties.

    10 min read

  • Responsible AI in the Enterprise · Part 2

    A governance model for enterprise AI agents

    Six controls that let AI agents run real workflows while named people stay in charge of every step that carries risk, and how to put them in place.

    7 min read