At eleven at night the answers start going subtly wrong: confident, well formatted, citing documents that do not say what the answer claims.

Three things changed today. A prompt template edit shipped this morning, the vendor rolled a model point release, and an ingestion job failed quietly at six. Which one is it, and which one can you undo right now, on its own, without reverting the other two?

If the answer is "we would have to redeploy", what you have is a deployment that you are calling a rollback.

AI incidents do not fit the service taxonomy

Classic incident categories are about availability. The service is down, latency is up, the error rate spiked. AI systems fail while returning 200s.

The classes worth naming in advance, because you will not invent them under pressure:

  • Quality regression. Answers got worse. No error anywhere.
  • Retrieval failure. The index is stale, empty, or serving the wrong tenant's documents.
  • Tool failure wearing a model costume. A tool returns nothing and the model fills the gap.
  • Safety or policy breach. Output that should have been refused, or a leak across a boundary.
  • Cost or latency blowout. Usually a retry loop, or a context that quietly doubled.
  • Upstream model change. The vendor changed something and your behaviour moved with it.

Each has a different first responder and a different first action. That is the entire reason to write them down.

Independent rollback, or none at all

Every layer that can change behaviour needs its own version and its own reversal path, decoupled from the application release. Prompt templates: versioned, flag-selectable, revertible in seconds. Model choice and parameters: configuration, not code. Retrieval config: index version, chunking, thresholds, reranker on or off. Tools: individually disableable, so one bad integration does not force you to kill the whole feature.

Then the kill switch, and be precise about what it kills. Hiding the UI while the workflow still runs on the backend is not one. Exercise it in production on a quiet afternoon, because an untested kill switch is a belief.

The half that is not engineering

Detection is chronically underinvested. Your availability alerts will not fire here. Alert on quality proxies instead: refusal rate, retrieval hit rate, correction rate, escalation volume, tokens per task.

Support needs a path that is not "file a bug". Give them the failure taxonomy, a trace ID field, and a named person to escalate to. Decide the customer communication rule before you need it: which failure classes get proactive contact, who writes it, who approves it.

Afterwards, a postmortem without evidence is a memory exercise. Traces, versions in play, tenants affected, timeline. Then the step that turns an incident into an asset: every confirmed bad output becomes a regression case, so that failure has to be new next time.

Closing thought

You cannot build rollback during an incident. It is a set of boundaries you drew months earlier, on an ordinary Tuesday when nothing was on fire.

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.