Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Topic
In the glossary: Eval harness and golden set, LLM as judge, what each means and how to say it in a review
Part of Observability & Ops: all Observability & Ops entries · the Observability & Ops lens on the map
Checklist · · 12 checks
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Checklist · · 24 checks
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Explainer · · 2 min read
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Explainer · · 2 min read
Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.
Explainer · · 2 min read
Regression gates only work when they sit in the release path, with a threshold, a slice and an owner.
Explainer · · 2 min read
Deterministic checks first, one rubric per criterion, and a calibration loop against human review.
Explainer · · 2 min read
Labels, edge cases, versions, an owner. Skip those and the number moves without anyone knowing why.
Explainer · · 2 min read
Decide which decision the number is allowed to stop, then build the harness around that.
Explainer · · 2 min read
Groundedness scores, per-claim citations and an operations dashboard are what tell you quality is drifting before your users do.
Explainer · · 1 min read
Why evaluating agents means measuring the whole workflow users experience.
Explainer · · 2 min read
Why AI debate is only useful when the judge has reliable criteria.
Explainer · · 1 min read
Why RAG evaluation has to measure retrieval, grounding, generation, and citations separately.