Sixty days ago this series started with one claim: AI is a system, not a model.

Here is the map that claim produces, and the question each layer has to answer before it is production-ready.

LayerThe production question
ModelWhat behaviour is the model actually responsible for?
ContextWhat evidence and instructions enter the task?
RetrievalCan it find trusted, current, authorised evidence?
ToolsWhat actions are allowed, logged, and reversible?
AgentsWhere is autonomy useful, and where is it bounded?
EvaluationHow do we know the system is improving?
GovernanceWho owns risk, review, and accountability?

Five principles, in order

Start simple. A chatbot that works beats an agent that impresses. Complexity is a cost paid on every request, forever; add it when a failure demands it, not when a demo suggests it. (Days 1-10)

Ground every answer. Retrieval, evidence gates, citations. A system that can show its sources can be debugged, audited, and trusted. One that cannot is a rumour generator with good uptime. (11-30)

Control every action. Goals with checkable success criteria, tools with scoped permissions, budgets with enforced ceilings, recovery designed before the failure. (31-40)

Coordinate carefully. Multiple agents only when permissions, parallelism, or context isolation genuinely demand them, with contracts at every boundary. (41-50)

Measure and own everything. Tracing, system-level evaluation, injection defence, and a named human accountable for each system in production. (51-59)

The one habit worth keeping

If only a single thing survives contact with your roadmap, take this one: make the system inspectable.

Every hard problem in these sixty notes (hallucination, stale evidence, prompt injection, runaway cost, coordination failure, unexplainable decisions) shrinks the moment you can see what actually happened. Observability is less a layer alongside the others than the property that makes all of them fixable.

The mistake I would most want you to avoid

Treating the ecosystem as a shopping list. Adding RAG, tools, memory, agents, and evaluation as separate initiatives does not produce a production system. It produces five components nobody coordinates and one outcome nobody owns.

Start from the user outcome and work backwards. Choose the smallest architecture that can be operated safely. Earn each addition.

Closing thought

From model to mission-ready AI means coordinating the pieces you have until they can be trusted in production, rather than adding more of them.

Thank you for reading all sixty. The notes end here; the operating never does.

If you had to operate your whole AI stack for a year, what would you redesign first?

Architecture pattern · · 1 min read

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.