Incident Response for AI Failures: The Runbook Nobody Has Written Yet
A support chatbot gives a customer confidently wrong medical guidance. A credit model is found to decline applicants from one region at a materially higher rate. A prompt injection buried in an uploaded document causes an internal assistant to summarise and email a confidential file.
Open your incident response playbook and look for the procedure. It will not be there. Your playbook was written for confidentiality, integrity and availability failures in systems whose correct behaviour is specified, and none of these three incidents fit that shape.
Why Contain, Eradicate, Recover Does Not Fit
The classic cycle assumes an intruder, a defect, or an outage - something foreign to the system that can be removed, after which the system is well again.
AI incidents break each assumption in turn:
There is often nothing to eradicate. Hallucination is not a defect introduced into the model; it is a property of how the model works. You cannot remove it, only constrain and detect it. An eradication phase that has no object leaves responders unsure whether the incident is closed.
Containment may mean removing the service entirely. With no patch available, the only immediate containment for a misbehaving model is to disable it or fall back to a previous version or a rules-based path. That is a business decision with revenue consequences, made under time pressure, and it needs a pre-agreed decision-maker.
Recovery does not restore correctness. Restoring from backup restores the same model with the same behaviour. Genuine recovery may require retraining, which takes days or weeks, not hours.
The blast radius extends backwards. When a model is found to be biased, the incident is not confined to the moment of discovery. Every decision it made since deployment is potentially affected, which turns an incident into a remediation programme covering months of historical outcomes.
Forensics look nowhere near a traditional investigation. There is no malicious binary and often no log entry that looks anomalous. The evidence is the prompt, the retrieved context, the model version, the parameters and the completion - and if you were not logging those, there is no investigation to run.

An AI Incident Taxonomy
Classification drives response, and generic severity labels do not tell a responder what to do. Use categories that carry their own procedure.
| Category | What happened | First containment action | Notification analysis |
|---|---|---|---|
| Bias or disparate outcome | Materially different outcomes across groups in a decision system | Suspend automated decisioning; route to human review | Potential discrimination exposure; sectoral regulator; affected individuals |
| Hallucination with reliance | False output acted upon by a customer or employee | Disable the feature or add mandatory human review | Consumer protection; professional liability; contractual |
| Data poisoning | Training or retrieval data manipulated to alter behaviour | Freeze retraining; revert to last known-good model version | Integrity incident; potentially a security breach |
| Prompt injection leading to disclosure | Untrusted input caused the system to reveal or exfiltrate data | Disable affected integration; revoke tool access | Personal data breach analysis under Section 8(6) |
| Training data extraction | Model output reproduced training data including personal data | Rate limit, apply output filtering, consider model withdrawal | Personal data breach; assess scope by data class |
| Model theft or extraction | Systematic querying to reconstruct the model | Rate limit, revoke credentials, block source | Intellectual property; contractual; possibly no personal data element |
| Consent breach via AI output | Data processed or surfaced beyond its consented purpose | Suspend the processing path | DPDPA purpose limitation; Data Principal notification |
| Availability or degradation | Model unavailable or performance collapsed | Fail over to fallback path | Service level; usually no privacy dimension |
| Third-party AI incident | Vendor's model or provider suffers an incident affecting your data | Assess exposure; invoke contractual notification | You remain the Data Fiduciary; your clock runs |
Two categories deserve particular attention because organisations consistently misclassify them.
Prompt injection leading to disclosure is a personal data breach when personal data is disclosed, and the Section 8(6) intimation analysis begins immediately. Teams frequently treat it as an application bug because the mechanism is novel and no traditional exploit was used.
Bias findings are incidents, not findings. Discovering that a deployed decision system produces disparate outcomes is not an audit observation to be scheduled into the next quarter. It is a live incident with individuals currently being affected.
The 72-Hour Clock in an AI Context
Section 8(6) of the DPDP Act requires a Data Fiduciary to intimate a personal data breach to the Data Protection Board and to each affected Data Principal. The Rules give this operational shape, including notification to the Board within 72 hours.
The question that decides everything is when the clock starts and whether an AI incident is a personal data breach at all. Work through it in this order:
- Was personal data involved? In prompts, retrieved context, training data, or output. Note that model output can constitute personal data even where no personal data was in the prompt - if the model reproduces memorised training data or correctly infers attributes about an identifiable person.
- Was there unauthorised processing, disclosure, acquisition, loss or alteration? Disclosure to a party who should not have received it is the common case with injection and extraction.
- Can you scope it? Which Data Principals, which data categories, over what period. This is where AI incidents are hardest, because output is generated rather than retrieved and there may be no query log identifying who was affected.
The practical failure is step three. In a database breach you can enumerate affected rows. In a training data extraction incident you may be unable to determine whose data was reproduced without an analysis that takes longer than the notification window.
The mitigation is architectural and must be built before the incident: log prompts and completions with retention aligned to your investigation needs, maintain dataset-to-individual lineage, and record model version against every inference. Organisations without these three discover during their first incident that scoping is not possible, and end up notifying conservatively across a far broader population than was actually affected.

The Runbook
Phase 1 - Detection
Sources: drift and performance monitors crossing thresholds; disparity monitoring on decision systems; anomaly detection on inference logs; user and customer reports; employee escalation; vendor notification; external researcher disclosure.
Give users a reporting path. A visible "this output was wrong or harmful" control on AI-facing features is one of the highest-yield detection sources available, and it costs almost nothing to build.
Phase 2 - Classification
Assign a category from the taxonomy and a severity. Severity should be driven by consequence to individuals, number affected, whether the system is still operating, and whether personal data is involved. Start the notification analysis clock at classification and record the timestamp - you will be asked for it.
Phase 3 - Containment
The AI-specific decision is the takedown decision. Options in ascending order of business impact: apply output filtering; add mandatory human review; reduce autonomy by revoking tool or integration access; roll back to a previous model version; fail over to a rules-based path; disable entirely.
Pre-agree who can authorise each level. In the moment, an engineer will not unilaterally disable a revenue-generating feature and will escalate to someone unavailable. Name the role, name the deputy, and give them a documented mandate.
Phase 4 - Forensics
Collect: the prompt and any retrieved context, the completion, model identifier and version, inference parameters, timestamps, the calling identity, tool invocations made by the system, and the training data lineage for the model version in question.
The investigative questions differ from a traditional one. Not "how did they get in" but "what input produced this behaviour, is it reproducible, how many other interactions share the pattern, and which model version was serving."
Phase 5 - Notification
Run the personal data breach analysis to a documented conclusion, whichever way it goes - a reasoned decision not to notify is itself an artefact you may need to produce. Where notification is required, notify the Board within the required window and affected Data Principals as required. Assess sectoral reporting obligations separately, since they often run on shorter clocks. Assess customer contractual notification obligations, which frequently have their own timelines. Consider whether the EU AI Act's serious incident reporting applies to any high-risk system in scope.
Phase 6 - Post-mortem
Beyond the usual timeline and contributing factors, AI incidents need three specific questions answered: was this failure mode identified in the system's MAP or DPIA and accepted; would existing monitoring have caught it earlier with a different threshold; and does the same failure mode exist in other systems using the same model or pattern? The third question is what turns a single incident into portfolio-level improvement.
Phase 7 - Remediation and redeployment
Where retraining is required, treat redeployment as a new deployment: revalidation against the failure mode, updated DPIA, updated model register entry, and documented sign-off. A model that caused an incident and returns to production without revalidation is the same incident waiting to recur.
Do not forget the backward-looking remediation. If a decision system was wrong for four months, the affected decisions need review, and that programme usually dwarfs the technical fix.
Tabletop Scenarios
Run these before you need them. Each is chosen to break a different assumption.
Scenario A - The invisible bias. A quarterly review shows your loan pre-qualification model approving applicants from one state at a materially lower rate. It has been live for seven months. Tests: Is this an incident or a finding? Who decides? What happens to seven months of decisions? What do you tell the regulator, and when?
Scenario B - The injected document. A customer uploads a PDF to your support portal. Embedded instructions cause the assistant to retrieve and include another customer's ticket history in its reply. Tests: Is this a personal data breach? Can you determine how many other uploads contained similar payloads? Does your log retention reach back far enough?
Scenario C - The vendor's incident. Your SaaS provider discloses that a misconfiguration allowed prompts from one tenant to appear in another's context for nine days. Tests: Whose clock runs? Do you have contractual notification content requirements? Can you determine what your users submitted during those nine days?
Scenario D - The confident wrong answer. A customer acted on incorrect guidance from your assistant and suffered financial loss. They have posted the transcript publicly. Tests: Is there a privacy dimension at all? Who owns the response - security, legal, or communications? Do you have the interaction log to verify the transcript?
Scenario D is worth running specifically because many teams conclude there is no security incident and therefore no process applies, which leaves the organisation improvising in public.
Wiring It Into Existing Tooling
SIEM. Ingest inference logs. Build detections for volume anomalies per identity, prompt length outliers, repeated near-identical queries indicating extraction, and known injection patterns. Treat model version changes as change events correlated against error rates.
SOAR. Automate the containment ladder. A playbook that can revoke a model's tool access, apply an output filter, or trigger fallback without a human in the loop converts a forty-minute manual containment into seconds.
GRC incident module. This is where the AI incident links to the model register entry, the DPIA, the vendor record and the control library - and where the notification decision, its reasoning and its evidence are recorded. Without this link, your incident record and your compliance record tell different stories about the same event, which is the finding auditors most enjoy.
Conclusion
AI incident response is not a new discipline. It is the existing discipline extended to a class of system whose failures are statistical rather than binary, whose containment is a business decision rather than a technical one, and whose blast radius extends backwards through every decision already made.
The organisations that handle their first AI incident well will be the ones that logged prompts and completions before they needed them, pre-agreed who can take a model down, and ran a tabletop where the answer was genuinely unclear.
Actionable recommendations:
- Log prompts, completions and model versions now. Without them there is no forensic capability and no ability to scope a breach within 72 hours. This is the single highest-value preparation step.
- Pre-authorise the takedown decision. Name the role that can disable a revenue-generating AI feature, and a deputy. Discovering the escalation path during an incident costs hours you do not have.
- Treat bias findings as incidents. Individuals are being affected while the finding sits in a queue, and the backward-looking remediation grows every day it waits.
- Classify prompt injection with disclosure as a potential personal data breach. The novel mechanism does not change the analysis, and teams routinely misfile it as an application defect.
- Run Scenario D. The incident with no obvious security dimension is the one that finds out whether your organisation has any process at all.
Log, classify and escalate AI incidents with pre-built workflows. Category-driven runbooks, automatic linkage to the model register and DPIA, and a DPDPA notification checklist that starts its clock the moment an incident is classified. See how Dedups.ai operationalises AI incident response.