Someone shows me a dashboard. Eighty-seven per cent on the internal benchmark, up from eighty-four last month. Good news, apparently.
I ask what happens if it comes back at seventy-nine next Tuesday. The answer is usually some version of "we would look into it".
Then what they have is a chart with a nice trend line.
Start from the decision and work backwards
Before designing any scoring, name the decision the score is meant to serve. Does this block a prompt change? A model swap? A rollout to a new customer segment? Each tolerates completely different evidence. A prompt tweak needs a fast regression check in CI. Opening a regulated workflow to a new market needs human review of adversarial cases and a named person willing to sign.
Then set the threshold before you look at results. This sounds pedantic until you have watched a team discover a four-point drop and spend an afternoon arguing about whether four points matters. It always matters less once everyone knows which release it would delay. Write the number down while it is still abstract.
Model score and business outcome are also not the same quantity, and treating them as one is how teams ship confident regressions. Higher scores that produce more support tickets are a failure. Track both, and be explicit about which one wins when they disagree.
Not every task fails the same way
Lumping every request into one accuracy number hides exactly the failures you care about. Split by user intent and by what an error costs.
| Task type | What correct means | Best signal | What it can block |
|---|---|---|---|
| Extraction | Field matches the source | Deterministic diff, offline | Any release |
| Summarisation | Nothing invented, key points kept | Rubric plus sampled human review | Prompt and model changes |
| Advice and support answers | Useful and safe for this user | Human review, online feedback | Segment rollout |
| Agentic actions | Right steps, stopped at the right point | Trajectory review | Expanding autonomy |
The three sources of signal do different jobs. Offline evals are cheap and repeatable, so they belong in the release path. Online evals catch the questions your golden set never imagined. Human review is expensive and irreplaceable wherever "wrong" means harm rather than annoyance. Most teams run one of the three and call it coverage.
Quality is not the only axis either. Latency and cost are quality attributes with users attached to them. A more accurate answer that arrives after the user has given up is a worse product, and any eval that cannot express that will keep recommending the wrong trade.
The test for whether your evals are real is short: name the last release they delayed. If there isn't one, they are decoration, however good the numbers look.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Explainer · · 2 min read
Real users ask questions your test set never imagined
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Discussion