The golden set passed. The gates were green. Two weeks later, support has a backlog of confidently wrong answers about a policy that changed in a document nobody re-indexed.

The offline suite is still passing. It is faithfully testing a world that no longer exists.

Sample production before you need to

Draw a continuous random sample of real traffic and review it weekly, at a volume the team will actually sustain. Fifty examples read every week beats a thousand read once, in a panic, after an incident.

Draw the privacy boundary first: redact at the point of capture, restrict who can read raw content, set a retention window, and hold the reviewed sample to the same policy as its source data. That is far easier to design before there is a backlog than after.

Groundedness is the cheap online check

For anything retrieval-backed, one question does most of the work. Is each claim in the answer supported by a passage that was actually retrieved for that request?

That can be checked continuously without ground truth labels, which is exactly what makes it viable in production. Watch citation quality beside it, because citations pointing at real documents that do not contain the claim are worse than no citations at all, because they add confidence without adding evidence.

Drift arrives from several directions at once

The model changes under you when a provider updates it. The corpus changes as documents are added, edited and deprecated. The users change as word spreads and a new segment turns up with questions nobody anticipated. Prompt edits accumulate without a changelog.

So instrument the input distribution as well as the output scores. A shift in what people are asking usually appears before quality visibly falls, which makes it the earliest warning you get.

Feedback only counts once it becomes a case

A thumbs down with no trace ID is a mood. Every piece of feedback and every support escalation should resolve to a specific request, its retrieved context and its output, so a reviewer can see what happened and promote the bad ones into the eval set.

Then keep quality next to the numbers the business already watches: resolution rate, escalation rate, latency at the tail, cost per resolved request. A change that lifts scores while doubling latency and cost is a trade, and somebody should be making it deliberately rather than discovering it in a billing report.

Closing thought

Offline evals tell you whether you broke something you already knew about. Online evaluation is the only thing that tells you what you never thought to test.

What share of your production traffic did a human actually read last week?

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.