Most multi-agent architectures start life as a diagram. Boxes with role names, arrows between them, and a shared feeling that the hard decision was how many agents there should be.

Then it ships, and the first real incident is never about reasoning quality. A task gets dispatched and never comes back, and nothing notices for a day. Two agents write to the same place and one silently loses. A pair of specialists talk politely in circles for forty rounds and bill you for every one.

None of that is a model problem. You cannot prompt your way out of it.

I spent six days of a sixty-day series on the standard multi-agent topologies: hierarchies, orchestrators, pipelines, peer networks, shared workspaces and debate. Six different shapes that fail in remarkably similar ways.

Hierarchy is a context budget that charges you in fidelity

One supervisor coordinating fifteen specialists eventually drowns in its own tracking state: too many threads, too much status, too many decisions in one context window. So you add a layer. A director over three managers over five workers each, and no single agent holding the whole operation.

That works because every level summarises for the level above it. Compression is the entire feature, and the entire cost. Instructions dilute travelling down, findings blur travelling up, and the subtle caveat a worker raised gets flattened into a bullet the director never sees before deciding.

The fix is not better summarising. It is deciding in advance which findings are forbidden from being compressed: safety issues, contradictions, anything that invalidates a premise the level above is standing on. Those travel raw, or the hierarchy will reliably lose the one detail that would have changed the outcome.

The adoption test is structural. Hierarchy fits work that genuinely nests: portfolios containing projects, codebases containing modules. If two of your levels make the same kind of decision, one of them is ceremony, and you are paying latency for an org chart.

Dynamic dispatch swaps a planning problem for a bookkeeping problem

Some workloads refuse to be planned in advance. You cannot know whether a request needs three subtasks or thirty until you are inside it. That is what an orchestrator is for: it invents the subtasks at runtime and hands each to whichever qualified worker is free.

Dispatching is the easy half. The engineering lives in the accounting. A worker dies mid-task and the work evaporates while the orchestrator waits for a result that is never coming. Retries meet a non-idempotent task and the refund goes out twice. Results arrive out of order and occasionally contradict each other, and reconciling them is the orchestrator's real job rather than a footnote in it.

The failure that hurts most is the unexplainable dispatch: nobody can say why worker C was picked, so nobody can fix it when C was wrong. Log the signal behind each routing decision, not just the outcome. "Routed to the SQL worker because the query mentioned an order ID" is debuggable. "Routed to the SQL worker" is a black box that costs an afternoon per misroute.

The boring pipeline wins most on-call arguments

Ask an engineer which multi-agent system they would rather be paged for at 3am. They will pick the pipeline. Research to analysis to writing to review, one direction, fixed stages, one baton.

Nothing about it will impress a conference audience. Three properties make it the right default anyway. Every boundary is a contract, so violations surface where they happen rather than three stages downstream as a mystery. Debugging becomes bisection: inspect the artifact at each handoff, find the leg where it degraded. And stages upgrade independently, which is the one that compounds: a pipeline is the only topology where you can improve one component without wondering what you broke two hops away.

Two cheap habits protect all three: validate at each boundary against a schema, and persist every intermediate output so you never have to reconstruct a five-stage run from its final answer alone.

The honest limit is that pipelines assume work flows forwards. The moment stage four routinely returns work to stage one, you have a loop in a pipeline costume, and the costume hides the iteration count from you. Model the loop explicitly, with a cap.

Peer agents inherit a failure catalogue older than LLMs

Remove the supervisor and let the specialists talk directly. No bottleneck, no single point of failure, and two experts negotiating detail without a middleman that half-understands both sides.

What you have built is a distributed system, with the classic failures intact. Deadlock, where two agents wait on each other politely and indefinitely. Livelock, where the conversation cycles without converging, worse than deadlock because it looks like work and bills like work. And diffusion of responsibility, where every peer assumes another owns the final answer, so the transcript fills with useful contributions and contains no conclusion.

None of these is a reasoning failure, which is exactly why better prompts do not touch them. Protocol does. Typed messages with explicit intent, so one agent's casual mention cannot become another's assumed commitment. Exactly one owner per task at all times, with transfer as a logged event. And termination enforced from outside the conversation, because deadlocked parties cannot vote themselves out of a deadlock, and a livelocked pair never notices.

Shared state is only collaboration if somebody governs it

The blackboard pattern replaces messages with a wall. The researcher pins evidence, the analyst posts calculations, the reviewer flags risks. Agents stay decoupled, and you can add or remove one without rewiring a conversation.

Ungoverned, it degrades in a predictable order. Two agents overwrite each other and the loser's work disappears with no record it existed. An unverified hypothesis sits beside a confirmed fact with identical visual weight. Nothing records who wrote what, so when the output is wrong the contribution that caused it is untraceable. None of that looks like a bug: no error, no exception, just a result nobody can explain, noticed weeks later.

Govern it like a database and most of it goes away. Defined sections rather than prose. Append-only corrections, so when agent B contradicts agent A both entries survive, timestamped and attributed. Status that means something, where downstream agents may only build on "verified", because otherwise one agent's speculation gets laundered into another's premise. And attribution on everything.

The one-minute test: have two agents write to the same region at once. If one silently wins, that is a data-loss bug that will surface as an inexplicable result, intermittently.

Debate is not a topology, it is a quality mechanism, and it rests entirely on the judge

The odd one out. Debate does not organise work, it interrogates a conclusion. Models critique other models' output far better than their own, so you assign one agent to attack another's plan and have a judge score them.

The pattern lives or dies on the judge, which is the part that gets under-designed. A judge that rewards eloquence over evidence turns the apparatus into an expensive coin flip weighted towards whichever side writes more confidently. That is worse than no debate, because you have not removed the error, you have given it a courtroom and a verdict, which makes it harder to question downstream.

Three checks tell you whether you have a real one. Does the judge see the underlying evidence or only arguments about it. Were the criteria written before the debate rather than after reading it. And the swap test: rerun with the debaters' positions exchanged, and if the verdict follows the agent rather than the argument, your judge is measuring style. Then validate it on contested cases where you know the answer, because a judge that cannot pick correctly against ground truth will not do better without it.

It also triples inference for one answer: defensible for go/no-go gates, not for the routine tasks where it usually ends up switched on.

What the six patterns add up to

Six shapes, fixed the same way.

Choosing a multi-agent architecture is choosing which distributed-systems problem you would rather own, because all six leave you holding the same three: who owns the answer right now, what forces the work to stop, and how you reconstruct afterwards what actually happened.

The escalation channel in a hierarchy, the owner of record in a dispatcher, the boundary contract in a pipeline, the single-owner rule between peers, attribution on a blackboard, a judge with criteria fixed in advance: those are the same three invariants in six costumes.

That gives you a default. The patterns form a spectrum from decided-at-design-time to decided-at-runtime, and debugging difficulty rises exactly along it. Pipelines sit at the cheap end, peer networks at the expensive one. Pick the leftmost pattern that fits the work, and make the work prove it needs more autonomy before you grant it. If you can enumerate the stages on a whiteboard, you do not need an orchestrator.

And notice where every fix in this chapter lives. Not one of them is a better prompt. Timeouts, schemas, round limits, escalation paths, append-only logs, judges validated against known answers: all of it sits in the harness around the agents, not inside them. Multi-agent design is an operations discipline that keeps being approached as a prompting discipline, and that mismatch is most of why these systems disappoint in production.

The six posts, in order

Each is a short standalone piece with the specifics and the fixes:

They are the Multi-Agent Design Patterns chapter of 60 Days of Production AI Systems, a series on building AI that survives real users. If you would rather take it from the top, Start Here lays out all sixty in ten chapters.

Deep dive · · 6 min read

Most agent controls do not actually control anything

Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.