Executive summary
Picture a service agent that starts sending customers the wrong loan balance at 11 at night. Nobody is hacking anything. The provider updated the model that afternoon, and a field the agent reads now comes back in a new format. By the time someone notices, the messages have gone. The questions that follow are the hard part. Who can stop the agent? Which regulator needs to hear, and by when? What exactly did it do?
AI incidents are no longer rare. Stanford's AI Index counted 233 reported AI incidents in 2024, a record and a 56.4% rise on 2023 (Stanford HAI). Most enterprises already run a cyber incident process. Few have extended it to systems that write, decide and act. Our view is that you don't need a new process. You need your existing one to know what an AI incident looks like, a stop switch that works in seconds, and a record good enough to answer the regulator. This is the last part of our series, and it assumes the controls from Part 2 are in place.
1. What counts as an AI incident
The OECD defines an AI incident as an event where the development, use or malfunction of an AI system leads to real harm, and an AI hazard as one that could plausibly do so (OECD.AI). For an enterprise, the useful test is simpler: treat it as an incident if the system caused harm, or would have done if a person had not caught it. Near misses count. They are the cheapest lessons you will get.
Harmful output
WhatThe system says something false, unfair or unsafe to a customer or colleague.
SignalComplaints, flagged replies, a spike in human edits.
Wrong action
WhatAn agent does a real thing it should not have done, or does it to the wrong record.
SignalReversals, refunds or records that fail a later check.
Data leak
WhatPersonal or confidential data reaches a prompt, a reply or a log it should not.
SignalMasking alerts, data loss tools, a customer who saw another customer's data.
Prompt injection
WhatText in an email, file or web page steers the agent to act for someone else.
SignalTool calls that do not match the task, blocked calls at the gateway.
Model drift
WhatQuality slips over weeks as data, users or prompts change.
SignalFalling scores on a fixed test set, more overrides by approvers.
Provider change
WhatThe model or its terms change under you, and behaviour changes with it.
SignalA new version string in the trace, a sudden shift in output style or cost.
Two of these deserve a closer look. Prompt injection tops the OWASP list of risks for LLM applications, because text an agent reads can change what it does (OWASP). An agent with write access turns that from an odd reply into a real action. Provider change is the one teams forget. If the model version is not pinned and logged, you can't even tell that it changed.
2. Signals that tell you something is wrong
You find AI incidents in three places: the system's own traces, the people who work with it, and the customers who live with its output. Build a signal into each.
- In the trace. Alert on tool calls the gateway blocks, on calls outside the agent's usual pattern, and on any change in model version.
- In the work. Track how often approvers reject or edit a step. A sudden rise is often the first sign of drift. Part 5 covers which numbers to watch.
- At the edge. Tag complaints that mention an automated reply, so they reach the AI owner and not only the service desk.
Run a fixed set of test cases against each live agent every week, and after any model change. If the scores fall, treat it as a hazard before it becomes an incident.
3. Stop first, then investigate
The instinct in any incident is to understand before you act. With an agent, that is the wrong order. An agent can repeat a mistake hundreds of times while a team reads logs. Stop it, cap the harm, then find the cause.
- Stop the agentUse the stop switch at the gateway. In-flight steps halt and queued approvals freeze.
- Contain the harmRevoke the agent's tokens, pull the bad messages, run undo on the writes you can reverse.
- Keep the evidenceSnapshot the traces, prompts, model version and approvals before anything is changed.
- Triage and callName an incident lead, set the severity and start the reporting clock.
This only works if the stop switch is real. It should sit in the agent gateway, pause one agent or all of them, and be usable by the on-call engineer and the business owner without a change request. Test it every month. A stop switch nobody has pressed is a hope, not a control.
Keep the evidence before you fix anything. The trace, the prompts, the model version and the approvals are what your regulator, your auditor and your own review will ask for. CERT-In's directions already require covered bodies to keep system logs for a rolling 180 days, stored in India (CERT-In directions). Add your AI traces to that scope.
Triage: set the severity in minutes
Base severity on the tier of the action involved and on who was touched. The tiers are the same ones used for approvals, so nobody has to learn a second scheme.
| Severity | What happened | Who leads | First call |
|---|---|---|---|
| Low | Tier 0 or 1 errors, caught inside the team | AI product owner | Log it, fix it, sample more |
| Medium | Tier 2 step reached a customer or moved money | Incident lead from ops | Business owner, risk, legal |
| High | Tier 3 step, personal data exposed, or many people affected | CISO or delegate | Legal, DPO, leadership, regulators as needed |
When in doubt, start high and step down. Lowering a severity costs a meeting. Raising it late can cost a missed deadline.
HMR insight: Stop the agent before you understand it. You can always restart a paused agent. You can't unsend what it sent.
4. Who you may need to tell
Reporting duties depend on what happened, not on whether AI was involved. A data leak through a prompt is still a data leak. Check each duty below with your legal team for your own case. The clocks are short.
- CERT-In. The April 2022 directions require service providers, intermediaries, data centres, body corporates and government bodies to report listed cyber incidents within 6 hours of noticing them. The list includes data breaches, data leaks, and attacks on AI and machine learning systems (CERT-In directions).
- Data Protection Board and affected people. Under the DPDP Act and Rule 7 of the DPDP Rules 2025, a data fiduciary must tell each affected person and the Board without delay. A fuller report goes to the Board within 72 hours of becoming aware, unless the Board allows longer (DPDP Rules 2025). Rule 7 is due to take effect in May 2027, and a proposal to bring it forward had not been notified when we checked (timeline). Build for it now.
- RBI. Banks must report all unusual cyber incidents to the Reserve Bank promptly, including attempts that failed (RBI circular, 2016). The 2023 Master Direction on IT governance tells regulated entities to notify CERT-In and RBI proactively (RBI Master Direction).
- EU personal data. Under GDPR Article 33, a controller must tell the supervisory authority within 72 hours of becoming aware of a breach, where feasible (GDPR Article 33).
- EU AI Act. Providers of high-risk AI systems must report serious incidents within 15 days, and faster for the worst cases. Deployers must tell the provider (Article 73). The Digital Omnibus moved most high-risk duties to December 2027 (Cyber Law Watch), so plan for it rather than panic about it.
The practical point is that a 6-hour clock and a 72-hour clock can start from the same event. Decide in advance who makes the call on each, and keep a short report template ready. The six hours are for a first report, not a full one.
5. Recovery and review
Restart slowly. Patch the cause, add a test case that would have caught it, and bring the agent back one tier stricter than before. If it used to send tier 2 messages on its own after a clean quarter, it now waits for an approver again. Loosen the gate only when the log shows it is safe, as Part 4 describes.
- Day 0Stop and containThe agent is paused. Harm is capped. Evidence is safe.
- Days 1 to 3Find the causeReplay the trace. Decide which duties to report apply, and meet them.
- Weeks 1 to 2Fix and restart with tight gatesPatch the cause, add a test for it, and bring the agent back one tier stricter.
- Every quarterRehearseRun a tabletop exercise on a new scenario. Fix what it shows.
Hold a blameless review within two weeks. Ask four questions. What did the agent do? Why did the controls let it? How long did it take us to notice and to stop it? What changes now? Record the answers against the six controls, so each lesson lands in the gateway and not in a slide. NIST's incident response guide, revised in 2025, is a good frame for the review, and it treats lessons learned as part of risk management, not an afterthought (NIST SP 800-61r3).
6. Illustrative scenario: a lender's collections agent
This scenario is illustrative. The company is fictional.
Narmada Home Finance uses an agent to draft collection reminders and to log promises to pay. On a Friday evening, a borrower sends an email with hidden text. The agent reads it and tries to mark the loan as settled. The gateway blocks the call, because settlement is tier 3 and needs two approvers. It raises an alert.
The on-call engineer pauses the agent within minutes. The trace shows two more emails with the same trick, all blocked. No money moved and no data left the company. Risk sets the severity to medium, because the attempt was real even though the controls held. Legal checks whether CERT-In's list applies to an attempted attack on an AI system and decides to file a report within the six hours.
On Monday, the team adds a filter for hidden text, adds the emails to the test set and restarts the agent with drafting only. What changes is not luck. The controls stopped the action, the trace told the story, and the team knew who decides what.
7. Practise before you need it
A plan you have never run is a draft. Run a tabletop exercise each quarter: ninety minutes, the real people, one scenario, and a clock. Take one from Figure 1 each time. Ask the team to find the stop switch, set a severity, decide which regulators to call and draft the first report. Invite legal and the business owner, not only engineers.
The gaps show up fast. Someone does not know they can press stop. The model version is not in the trace. Nobody owns the DPDP notice. Each gap becomes a fix with an owner and a date.
Closing the series
This series started with where AI belongs in your operating model and ends here, with what to do when it goes wrong. The thread through all six parts is the same: named people in charge of the steps that matter, and a record of every action.
Start this week with one test. Pick your most active agent and ask who can stop it, how fast and with what record. If the answer takes more than a minute, our three-week Readiness Audit covers incident readiness alongside the controls. You can read how we handle our own AI use in our AI governance policy, or book a readiness audit.
Sources
- Stanford HAI, AI Index 2025: State of AI in 10 Charts
- OECD.AI, Defining AI incidents and hazards
- OWASP Gen AI Security Project, LLM01:2025 Prompt Injection
- CERT-In, Directions under section 70B(6) of the IT Act, 28 April 2022
- MeitY, Digital Personal Data Protection Rules, 2025 (via PIB)
- dpdprules.org, DPDP Act and Rules timeline, checked September 2026
- Reserve Bank of India, Cyber Security Framework in Banks, June 2016
- Reserve Bank of India, Master Direction on IT Governance, Risk, Controls and Assurance Practices, November 2023
- EU General Data Protection Regulation, Article 33
- EU Artificial Intelligence Act, Article 73: Reporting of serious incidents
- K&L Gates Cyber Law Watch, EU Digital Omnibus on AI enters into force, July 2026
- NIST, SP 800-61 Rev. 3: Incident Response Recommendations, April 2025