Skip to content

Topic

Evaluation

12 entries.

In the glossary: Eval harness and golden set, LLM as judge, what each means and how to say it in a review

Part of Observability & Ops: all Observability & Ops entries · the Observability & Ops lens on the map

Entries

01

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

02

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.

04

Explainer · · 2 min read

Launch day is when evaluation starts

Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.

09

Explainer · · 2 min read

You cannot run RAG on user complaints

Groundedness scores, per-claim citations and an operations dashboard are what tell you quality is drifting before your users do.