Skip to content

Everything written

Journal

Every entry, read by engineering discipline or in series. Each one also sits on the map, next to the component it is about.

121 entries: all 18 series complete, 100 parts, plus standalone pieces. Weekly notes live in Notes .

New here? Pick where to start by what you do

October 2026 5

Build log: shipping Unseen UI to npm

How the component library behind this site went from a private workspace to two published packages, including the three things the first consumer broke.

Build log1 min readPractitioner

LangGraph vs the OpenAI Agents SDK

Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.

Comparison1 min readPractitioner

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

Checklist12 checksPractitioner

September 2026 24

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.

Checklist24 checksPractitioner

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.

Architecture pattern1 min readArchitect

When to bring in a compliance review

The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.

Checklist18 checksPractitioner

Security review checklist for an AI feature

What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.

Checklist29 checksPractitioner

You cannot roll back a prompt you never versioned

AI failures cross model, prompt, retrieval, tool and data boundaries. The runbook and the rollback controls have to as well.

Explainer2 min readPractitioner5 Days of AI Operations and Observability

Logging the answer tells you almost nothing

A trace that links intent, prompt, retrieval, tools and output is the only thing that makes an AI failure debuggable.

Explainer2 min readPractitioner5 Days of AI Operations and Observability

Most agent controls do not actually control anything

Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.

Deep dive6 min readPractitioner

Show 101 older entries, back to May 2026

An approve button is not human oversight

Approvals, events and tenant boundaries are architecture. Bolt them on as UI and they become theatre.

Explainer2 min readPractitioner5 Days of AI Architecture Patterns

You are already building a control plane, badly

Policy, memory, evals and rollout get rebuilt inside every feature until someone names the layer they belong to.

Explainer2 min readPractitioner5 Days of AI Architecture Patterns

Most agents are workflows wearing a costume

Chatbot, workflow, agent, RAG. Pick the wrong one and you spend a quarter debugging autonomy nobody asked for.

Comparison2 min readPractitioner5 Days of AI Architecture Patterns

Write role contracts, not agent personalities

"You are a meticulous senior researcher" is a costume. Inputs, outputs, permissions and done criteria are a contract.

Explainer2 min readPractitioner5 Days of Multi-Agent System Design

August 2026 25

Blast radius is a design parameter

Narrow tools, staged writes, sandboxes and a tested kill switch. Containment is built before it is needed.

Explainer2 min readPractitioner5 Days of AI Safety Engineering

Untrusted text does not get to give orders

Injection defence is layered validation and clear authority, not a better-worded system prompt.

Explainer2 min readPractitioner5 Days of AI Safety Engineering

Your agent should not be a superuser

Bind every AI action to a user, a service and a run, then scope the verb rather than only the data.

Explainer2 min readPractitioner5 Days of AI Safety Engineering

Launch day is when evaluation starts

Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.

Explainer2 min readPractitioner5 Days of AI Evaluation

You cannot run RAG on user complaints

Groundedness scores, per-claim citations and an operations dashboard are what tell you quality is drifting before your users do.

Explainer2 min readPractitioner5 Days of Production RAG

Chunk for the question, not for the token limit

Chunk size is a decision about what a complete answer looks like. Reranking, context budgets and refusal thresholds finish the job.

Explainer2 min readPractitioner5 Days of Production RAG

The 3am agent run that nobody is watching

Scheduled agent work fails quietly by default. Traces, release gates and a runbook are what make it fail loudly.

Explainer2 min readPractitioner5 Days of Agent Infrastructure

July 2026 28

What the MCP protocol actually standardises

It settles how capabilities are described and discovered. Every decision about whether to trust them is still yours.

Explainer1 min readPractitioner5 Days of MCP for Production AI

If you cannot trace it, you cannot run it

Production AI needs traces across prompts, context, retrieval, tools, costs and decisions to debug anything at all.

Explainer1 min readPractitionerProduction AI Readiness

Role design matters more than agent count

Each role needs a clear purpose, authority, tools, memory and evaluation. Adding agents without that adds confusion.

Explainer2 min readPractitionerAgent Coordination Contracts

AI debate is only as good as its judge

Agents arguing produce better answers only when the judge has reliable criteria for deciding between them.

Explainer2 min readPractitionerMulti-Agent Design Patterns

Peer agents need protocols, not vibes

Agents talking as equals need communication rules, a limit on rounds and one owner of the final decision.

Explainer1 min readPractitionerMulti-Agent Design Patterns

One giant agent is rarely the cleanest design

Split work across agents to reduce complexity through specialization, never to add agents for their own sake.

Explainer1 min readPractitionerAgent Control and Supervision

June 2026 29

Design recovery into an agent on day one

Retries, fallbacks, checkpoints, rollback and escalation belong in the first design, not the first incident review.

Explainer2 min readPractitionerAgent Control and Supervision

A stale answer is a wrong answer

When the world changes faster than your index, freshness stops being a nice-to-have and becomes part of correctness.

Explainer2 min readPractitionerProduction RAG Operations

One question can need many searches

A single query often returns half the evidence. Some questions need several searches before the answer is safe.

Explainer2 min readPractitionerEvidence-Driven RAG Patterns

Hybrid search is the practical default

Combining keyword and semantic signals beats betting the whole system on one search method.

Explainer2 min readPractitionerRetrieval Foundations

RAG is an evidence design problem

Letting the model search is the easy part. The work is deciding what counts as trusted evidence.

Explainer2 min readPractitionerFrom LLM Demo to AI Product

May 2026 10

A capable model is not a product yet

Capability is the starting point. Workflow, product design and ownership are what turn it into something people rely on.

Article2 min readPractitionerFrom LLM Demo to AI Product

Confidence is not evidence

A confident tone tells you nothing about correctness. Answers need evidence, boundaries and a way to say "not sure".

Explainer2 min readPractitionerFrom LLM Demo to AI Product

Vague prompts become vague systems

Once real users depend on the system, a prompt is an operating instruction and deserves the same care.

Explainer1 min readIntroAI Systems Basics