Build log: shipping Unseen UI to npm
How the component library behind this site went from a private workspace to two published packages, including the three things the first consumer broke.
Everything written
121 entries: all 18 series complete, 100 parts, plus standalone pieces. Weekly notes live in Notes .
New here? Pick where to start by what you do
How the component library behind this site went from a private workspace to two published packages, including the three things the first consumer broke.
Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Six posts on multi-agent systems, and the failures were never inside an agent. They were between two of them.
How to design and run a multi-agent system in production, using the supervisor pattern, with its trade-offs and where human approval belongs.
The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.
What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
The topology gets all the design attention. Ownership, termination and traceability are what decide whether it survives contact with production.
Drift is quiet by construction. SLOs and a standing review are what make it audible.
AI failures cross model, prompt, retrieval, tool and data boundaries. The runbook and the rollback controls have to as well.
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Cost and latency belong to the workflow, not the model call. Budget them before rollout, not after the invoice.
A trace that links intent, prompt, retrieval, tools and output is the only thing that makes an AI failure debuggable.
Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.
Quality, safety, cost, latency, resilience and operations have to be visible on one diagram, at the same time.
Approvals, events and tenant boundaries are architecture. Bolt them on as UI and they become theatre.
Policy, memory, evals and rollout get rebuilt inside every feature until someone names the layer they belong to.
Chatbot, workflow, agent, RAG. Pick the wrong one and you spend a quarter debugging autonomy nobody asked for.
Routing, budgets, fallbacks and policy need one place to live. Scattered SDK calls give you none of them.
Six days of notes on agentic systems, and not one of the fixes that actually worked lived inside the model.
Multi-agent systems need trace timelines, trajectory evals and designed recovery, because the final output hides everything that produced it.
Permissions belong to roles, budgets belong to runs, and termination has to be something a machine can check.
Supervisor, orchestrator-worker, planner-executor, pipeline and judge differ mainly in where control lives and when it returns.
"You are a meticulous senior researcher" is a costume. Inputs, outputs, permissions and done criteria are a contract.
Separation of work has to buy quality, control or throughput. Usually it buys a harder debugging problem.
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Minimise what enters context, watch the paths data can leave by, and keep the evidence an incident will demand.
Narrow tools, staged writes, sandboxes and a tested kill switch. Containment is built before it is needed.
Injection defence is layered validation and clear authority, not a better-worded system prompt.
Bind every AI action to a user, a service and a run, then scope the verb rather than only the data.
Assets, actors, abuse cases, controls, owners, evidence. Six things a team can actually write down before launch.
Six posts on reranking, citations and decomposition, and the uncomfortable thing they turn out to have in common.
Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.
Regression gates only work when they sit in the release path, with a threshold, a slice and an owner.
Deterministic checks first, one rubric per criterion, and a calibration loop against human review.
Labels, edge cases, versions, an owner. Skip those and the number moves without anyone knowing why.
Decide which decision the number is allowed to stop, then build the harness around that.
Six days of retrieval notes, and not one of the failures announced itself.
Groundedness scores, per-claim citations and an operations dashboard are what tell you quality is drifting before your users do.
Chunk size is a decision about what a complete answer looks like. Reranking, context budgets and refusal thresholds finish the job.
Most "the model is wrong" problems are recall problems. Hybrid search, query rewriting and versioned indexes are where they get fixed.
Parsing sets your quality ceiling, failures need a destination, and the update and delete paths are the half nobody tests.
Embeddings carry no access control lists. Ownership, permissions, freshness and lineage have to be designed before anything is indexed.
Six posts on the distance between a system that answers well once and a system people are willing to depend on.
Notes from building an MCP gateway. The protocol standardised how agents call tools, then stayed silent about credentials, context budgets and who may call what. That silence gets expensive as the servers multiply.
Scheduled agent work fails quietly by default. Traces, release gates and a runbook are what make it fail loudly.
Retries are the first thing teams switch on and the last thing they design for.
Conversation, task, artifacts and product data end up in one store. Then a deletion request arrives and applies to all of it.
Six days of production AI notes, and the model choice never once turned out to be the thing that broke.
A list of tools is documentation. A registry decides who may act, as whom, and leaves proof behind.
Agents get more capable by default. They only get bounded on purpose.
One question settles most of these arguments: will more than one AI application need this capability?
Connector correctness and agent behaviour are separate problems. Most teams only test the first one.
Choosing the wrong primitive hands the model authority a person was supposed to hold.
Transport looks like a technical detail during the demo. It is really a choice about ownership, reachability and trust.
It settles how capabilities are described and discovered. Every decision about whether to trust them is still yours.
A practical map for moving from impressive AI demos to systems people can trust.
Controls that live only in a policy document are not enforced. Build them into the system the policy describes.
No single component makes production AI work. It is the coordination of data, models, tools, controls, evaluation and people.
Production AI needs traces across prompts, context, retrieval, tools, costs and decisions to debug anything at all.
Model benchmarks say little about the product. Measure the whole workflow the user actually goes through.
Text pulled from documents and tools must never gain authority over what the system does next.
Decide how disagreements get settled before they happen, so they never leak into the final answer.
A handoff that passes only a summary loses the evidence and the open questions. Make each one a structured contract.
Once agents share memory, the question is who may read and write what, and that is access control.
Agents that pass free text to each other drift. A shared message contract is what makes collaboration reliable.
Each role needs a clear purpose, authority, tools, memory and evaluation. Adding agents without that adds confusion.
Without a central coordinator, agents need rules that make them converge before production can trust the result.
Agents arguing produce better answers only when the judge has reliable criteria for deciding between them.
When several agents write to the same state, it stays debuggable only if every change is structured, attributed and versioned.
Agents talking as equals need communication rules, a limit on rounds and one owner of the final decision.
When agents pick who handles what at run time, every routing decision must be visible and one agent must own the result.
Predictable workflows are easier to inspect and fix as a pipeline. Add agent autonomy only where the pipeline runs out.
A supervisor works when routing, review and authority are clear, and becomes the bottleneck when they are not.
Layers of agents make sense when the work has real layers of authority and responsibility, and add cost when it does not.
Split work across agents to reduce complexity through specialization, never to add agents for their own sake.
Limits on cost, time, actions, retries and risk are what make autonomy safe to switch on.
Retries, fallbacks, checkpoints, rollback and escalation belong in the first design, not the first incident review.
A guardrail that only blocks teaches users to work around it. Good ones steer uncertain work toward a safe path.
A reviewer who sees too little, too late, or cannot say no is a rubber stamp. Placement decides whether review works.
Useful memory is separated, scoped and governed, not one database the agent writes everything into.
Reflection earns its cost when it decides the next move: revise, retry, escalate or stop.
Giving a model tools means designing permissions, observability, rollback and approval paths with them.
When actions have cost, risk or dependencies, a plan up front saves the expensive mistakes later.
Reliable agents come from how the loop is designed, not from one clever prompt.
Before any autonomy: what done looks like, what is out of bounds, what it may spend, and when it must stop.
One end-to-end score hides which stage failed. Measure each stage on its own so you know what to fix.
When retrieval comes back thin, the system should search again, ask, or say it does not know, not write a fluent guess.
A model reviewing its own answer helps only when it checks against evidence and is allowed to change course.
When the world changes faster than your index, freshness stops being a nice-to-have and becomes part of correctness.
Evidence quality changes when the source is visual, tabular or audio, and text-only pipelines miss what changed.
The current price, the open ticket, the approved limit: much of what users ask lives in systems of record, not documents.
Who owns what, what depends on what: some answers sit in the links between records, where document search cannot see.
A single query often returns half the evidence. Some questions need several searches before the answer is safe.
A compound question answered in one pass hides its weakest part. Split it, answer each piece, then combine.
A precise snippet without its surroundings misleads; surroundings without the snippet waste the context window.
Provenance lets users and teams inspect where a claim came from before they act on it.
Broad retrieval finds candidates; reranking decides which of them the answer can depend on.
The most relevant passage is still the wrong one if the person asking was never allowed to read it.
A vector database needs the same operating discipline as any other store: ownership, refresh, access and monitoring.
Product codes, names and error strings need exact matching; paraphrased questions need meaning. Real queries bring both.
Combining keyword and semantic signals beats betting the whole system on one search method.
Embeddings find text that sounds related. Whether it is correct, current or allowed is a separate question.
Retrieval can only return what a chunk holds. Split in the wrong place and the rule arrives without its exception.
Answer quality is set when knowledge enters the system, long before a user types a question.
Letting the model search is the easy part. The work is deciding what counts as trusted evidence.
A practical way to choose between conversation, repeatable workflows, and bounded autonomy.
Free text is for people. Software needs a schema it can validate before acting on what the model said.
Capability is the starting point. Workflow, product design and ownership are what turn it into something people rely on.
A confident tone tells you nothing about correctness. Answers need evidence, boundaries and a way to say "not sure".
Sampling settings decide how consistent the product feels to users, not just how the text reads.
A bigger context window is not a reason to stop choosing what goes into it.
Once real users depend on the system, a prompt is an operating instruction and deserves the same care.
Fluent output is generated step by step, which is why it needs constraints, validation and repeatable formats.
A low-level text detail that ends up as a product constraint on cost, latency and reliability.
Production AI succeeds or fails in the system around the model, not in which model you picked.
Disciplines down, series across. An entry counts once for each discipline it covers. Select a cell to read those entries.
Thinnest so far: Networking (5), Identity & Access (7)
| Discipline | Basics | Demo → product | Retrieval | Evidence RAG | RAG ops | Agents | Control | Multi-agent | Coordination | Readiness | MCP | Agent infra | Prod RAG | Evaluation | Safety | MAS design | Arch patterns | AI ops | Standalone | All |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AI & Models | 6 | 5 | 1 | 3 | 2 | 5 | 4 | 6 | 4 | 2 | 1 | · | 2 | 1 | 1 | 5 | 1 | 1 | 10 | 60 |
| Data & Databases | 1 | 3 | 6 | 6 | 5 | 1 | · | 1 | 1 | 1 | · | 1 | 5 | 1 | 1 | · | · | 1 | 9 | 43 |
| Networking | · | · | · | · | · | · | · | · | · | · | 2 | · | · | · | · | · | 2 | · | 1 | 5 |
| Security | · | · | 1 | · | · | 2 | 2 | · | 1 | 2 | 1 | 2 | 1 | · | 5 | 1 | 1 | · | 9 | 28 |
| Identity & Access | · | · | 1 | · | · | · | · | · | 1 | · | · | 1 | · | · | 1 | · | 1 | · | 2 | 7 |
| Integration | · | 1 | · | · | 2 | 1 | · | 2 | 2 | 1 | 4 | 2 | · | · | · | 1 | 2 | · | 5 | 23 |
| Compute & Infra | · | 1 | 1 | · | · | 1 | 4 | · | · | 1 | 1 | 4 | · | 1 | 1 | 1 | 2 | 1 | 2 | 21 |
| Observability & Ops | · | · | · | · | 1 | · | 1 | 1 | · | 2 | · | 1 | 1 | 5 | · | 1 | 1 | 4 | 4 | 22 |
| Cost & Governance | 2 | · | · | 1 | · | 1 | 1 | · | 1 | 1 | · | · | · | · | 1 | 1 | 1 | 1 | 3 | 14 |