Build log: shipping Unseen UI to npm
How the component library behind this site went from a private workspace to two published packages, including the three things the first consumer broke.
Everything written
121 entries: all 18 series complete, 100 parts, plus standalone pieces. Weekly notes live in Notes .
How the component library behind this site went from a private workspace to two published packages, including the three things the first consumer broke.
Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Six posts on multi-agent systems, and the failures were never inside an agent. They were between two of them.
How to design and run a multi-agent system in production, using the supervisor pattern, with its trade-offs and where human approval belongs.
The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.
What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
The topology gets all the design attention. Ownership, termination and traceability are what decide whether it survives contact with production.
Drift is quiet by construction. SLOs and a standing review are what make it audible.
AI failures cross model, prompt, retrieval, tool and data boundaries. The runbook and the rollback controls have to as well.
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Cost and latency belong to the workflow, not the model call. Budget them before rollout, not after the invoice.
A trace that links intent, prompt, retrieval, tools and output is the only thing that makes an AI failure debuggable.
Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.
Quality, safety, cost, latency, resilience and operations have to be visible on one diagram, at the same time.
Approvals, events and tenant boundaries are architecture. Bolt them on as UI and they become theatre.
Policy, memory, evals and rollout get rebuilt inside every feature until someone names the layer they belong to.
Chatbot, workflow, agent, RAG. Pick the wrong one and you spend a quarter debugging autonomy nobody asked for.
Routing, budgets, fallbacks and policy need one place to live. Scattered SDK calls give you none of them.
Six days of notes on agentic systems, and not one of the fixes that actually worked lived inside the model.
Multi-agent systems need trace timelines, trajectory evals and designed recovery, because the final output hides everything that produced it.
Permissions belong to roles, budgets belong to runs, and termination has to be something a machine can check.
Supervisor, orchestrator-worker, planner-executor, pipeline and judge differ mainly in where control lives and when it returns.
"You are a meticulous senior researcher" is a costume. Inputs, outputs, permissions and done criteria are a contract.
Separation of work has to buy quality, control or throughput. Usually it buys a harder debugging problem.
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Minimise what enters context, watch the paths data can leave by, and keep the evidence an incident will demand.
Narrow tools, staged writes, sandboxes and a tested kill switch. Containment is built before it is needed.
Injection defence is layered validation and clear authority, not a better-worded system prompt.
Bind every AI action to a user, a service and a run, then scope the verb rather than only the data.
Assets, actors, abuse cases, controls, owners, evidence. Six things a team can actually write down before launch.
Six posts on reranking, citations and decomposition, and the uncomfortable thing they turn out to have in common.
Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.
Regression gates only work when they sit in the release path, with a threshold, a slice and an owner.
Deterministic checks first, one rubric per criterion, and a calibration loop against human review.
Labels, edge cases, versions, an owner. Skip those and the number moves without anyone knowing why.
Decide which decision the number is allowed to stop, then build the harness around that.
Six days of retrieval notes, and not one of the failures announced itself.
Groundedness scores, per-claim citations and an operations dashboard are what tell you quality is drifting before your users do.
Chunk size is a decision about what a complete answer looks like. Reranking, context budgets and refusal thresholds finish the job.
Most "the model is wrong" problems are recall problems. Hybrid search, query rewriting and versioned indexes are where they get fixed.
Parsing sets your quality ceiling, failures need a destination, and the update and delete paths are the half nobody tests.
Embeddings carry no access control lists. Ownership, permissions, freshness and lineage have to be designed before anything is indexed.
Six posts on the distance between a system that answers well once and a system people are willing to depend on.
Notes from building an MCP gateway. The protocol standardised how agents call tools, then stayed silent about credentials, context budgets and who may call what. That silence gets expensive as the servers multiply.
Scheduled agent work fails quietly by default. Traces, release gates and a runbook are what make it fail loudly.
Retries are the first thing teams switch on and the last thing they design for.
Conversation, task, artifacts and product data end up in one store. Then a deletion request arrives and applies to all of it.
Six days of production AI notes, and the model choice never once turned out to be the thing that broke.
A list of tools is documentation. A registry decides who may act, as whom, and leaves proof behind.
Agents get more capable by default. They only get bounded on purpose.
One question settles most of these arguments: will more than one AI application need this capability?
Connector correctness and agent behaviour are separate problems. Most teams only test the first one.
Choosing the wrong primitive hands the model authority a person was supposed to hold.
Transport looks like a technical detail during the demo. It is really a choice about ownership, reachability and trust.
It settles how capabilities are described and discovered. Every decision about whether to trust them is still yours.
A practical map for moving from impressive AI demos to systems people can trust.
Why governance belongs in the architecture, not in a document after launch.
Why production AI is the coordination of data, models, tools, controls, evaluation, and people.
Why production AI needs traces across prompts, context, retrieval, tools, costs, and decisions.
Why evaluating agents means measuring the whole workflow users experience.
Why retrieved content must stay data, not become authority over the system.
Why agents need arbitration rules before disagreement reaches the final answer.
Why reliable handoffs need structured state, evidence, open questions, and ownership.
Why shared memory becomes an access-control problem in multi-agent systems.
Why agents need shared message contracts before they can collaborate reliably.
Why agent roles need clear purpose, authority, tools, memory, and evaluation.
Why decentralized agent behavior needs convergence rules before it earns production trust.
Why AI debate is only useful when the judge has reliable criteria.
Why shared state needs structure, attribution, and versioning to stay debuggable.
Why peer agents need communication rules, round limits, and a final decision owner.
Why dynamic dispatch needs visible routing decisions and clear final ownership.
Why predictable AI workflows often benefit from pipeline structure before agent autonomy.
Why supervisor agents help only when routing, review, and authority are clear.
Why hierarchy helps when work has real layers of authority and responsibility.
Why multi-agent systems should reduce complexity through specialization, not add agents for novelty.
Why autonomous systems need limits on cost, time, actions, retries, and risk.
Why retries, fallbacks, checkpoints, rollback, and escalation belong in the first design.
Why guardrails should shape safe progress, not just block uncertain behavior.
Why human review needs timing, authority, and context to be useful.
Why useful agent memory has to be separated, scoped, and governed.
Why reflection is valuable only when it decides whether to revise, retry, escalate, or stop.
Why giving AI tools means designing permissions, observability, rollback, and approval paths.
Why planning matters when actions have cost, risk, or dependencies.
Why reliable agent behavior comes from loop design, not a single clever prompt.
Why agents need clear outcomes, boundaries, budgets, and stop conditions before autonomy.
Why RAG evaluation has to measure retrieval, grounding, generation, and citations separately.
Why weak evidence should trigger recovery, not a fluent answer with false confidence.
Why reflection only helps when it checks the answer against evidence and changes behavior.
Why freshness becomes part of correctness when the world changes faster than your index.
Why evidence quality changes when the source is visual, tabular, audio, or multimodal.
Why production knowledge often lives in systems of record, not only in documents.
Why some answers depend on relationships that plain document search can miss.
Why one user question may need multiple searches before the evidence is good enough.
Why complex questions become more reliable when the system answers them in parts.
Why good RAG often needs both precise snippets and enough surrounding context.
Why important AI claims need receipts that users and teams can inspect.
How reranking turns broad retrieval candidates into evidence the answer can depend on.
Why relevant evidence is still wrong evidence if the user was not allowed to see it.
Why vector databases need to be operated like infrastructure, not treated like magic memory.
Why exact words and semantic meaning both matter in production retrieval.
Why combining retrieval signals is often more practical than betting on one search method.
Why semantic similarity is useful, but not the same as correctness or trust.
How chunk boundaries decide what evidence the system can actually retrieve.
Why answer quality starts when knowledge enters the system, not when the user asks.
Why RAG is really about trusted evidence, not just letting the model search.
A practical way to choose between conversation, repeatable workflows, and bounded autonomy.
How structured output turns model text into something software can safely use.
Why strong model capability still needs workflow, product, and ownership design.
Why confident language still needs evidence, boundaries, and uncertainty paths.
How sampling settings shape the user experience, not just the writing style.
Why a larger context window does not remove the need for context discipline.
Why prompts become operational instructions once real users depend on the system.
The practical reason LLM behavior needs constraints, validation, and repeatable output design.
How a low-level text detail quietly becomes a product constraint for cost, latency, and reliability.
Why production AI succeeds or fails in the system around the model, not in the model choice alone.
Disciplines down, series across. An entry counts once for each discipline it covers. Select a cell to read those entries.
Thinnest so far: Networking (5), Identity & Access (7)
| Discipline | Basics | Demo → product | Retrieval | Evidence RAG | RAG ops | Agents | Control | Multi-agent | Coordination | Readiness | MCP | Agent infra | Prod RAG | Evaluation | Safety | MAS design | Arch patterns | AI ops | Standalone | All |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AI & Models | 6 | 5 | 1 | 3 | 2 | 5 | 4 | 6 | 4 | 2 | 1 | · | 2 | 1 | 1 | 5 | 1 | 1 | 10 | 60 |
| Data & Databases | 1 | 3 | 6 | 6 | 5 | 1 | · | 1 | 1 | 1 | · | 1 | 5 | 1 | 1 | · | · | 1 | 9 | 43 |
| Networking | · | · | · | · | · | · | · | · | · | · | 2 | · | · | · | · | · | 2 | · | 1 | 5 |
| Security | · | · | 1 | · | · | 2 | 2 | · | 1 | 2 | 1 | 2 | 1 | · | 5 | 1 | 1 | · | 9 | 28 |
| Identity & Access | · | · | 1 | · | · | · | · | · | 1 | · | · | 1 | · | · | 1 | · | 1 | · | 2 | 7 |
| Integration | · | 1 | · | · | 2 | 1 | · | 2 | 2 | 1 | 4 | 2 | · | · | · | 1 | 2 | · | 5 | 23 |
| Compute & Infra | · | 1 | 1 | · | · | 1 | 4 | · | · | 1 | 1 | 4 | · | 1 | 1 | 1 | 2 | 1 | 2 | 21 |
| Observability & Ops | · | · | · | · | 1 | · | 1 | 1 | · | 2 | · | 1 | 1 | 5 | · | 1 | 1 | 4 | 4 | 22 |
| Cost & Governance | 2 | · | · | 1 | · | 1 | 1 | · | 1 | 1 | · | · | · | · | 1 | 1 | 1 | 1 | 3 | 14 |