Five-day tracks
5 Days of AI Operations and Observability
Running AI in production: traces and spans, cost, latency and capacity, production evaluation signals, incidents and rollbacks, drift and an operating cadence.
5 of 5 published
- 01
Part 1 · Explainer · 2 min read
Logging the answer tells you almost nothing
A trace that links intent, prompt, retrieval, tools and output is the only thing that makes an AI failure debuggable.
On the map: Observability - 02
Part 2 · Explainer · 2 min read
Your AI feature has unit economics whether you measured them or not
Cost and latency belong to the workflow, not the model call. Budget them before rollout, not after the invoice.
On the map: LLM Gateway, Observability - 03
Part 3 · Explainer · 2 min read
Real users ask questions your test set never imagined
Offline evals cover the questions you thought of. Production tells you the ones you did not.
On the map: Observability - 04
Part 4 · Explainer · 2 min read
You cannot roll back a prompt you never versioned
AI failures cross model, prompt, retrieval, tool and data boundaries. The runbook and the rollback controls have to as well.
On the map: Observability, LLM Gateway - 05
Part 5 · Explainer · 2 min read
Nothing broke, and the system is still getting worse
Drift is quiet by construction. SLOs and a standing review are what make it audible.
On the map: Observability