New here? Start this series at Day 43: When hierarchy helps multi-agent work scale
- Day 43: When hierarchy helps multi-agent work scale
- Day 44: Why dynamic dispatch needs ownership rules
- Day 45: Why pipelines make AI handoffs inspectable
- Day 46: Why peer agents need protocols, not vibes
- Day 47: Why shared state needs structure and attribution
- Day 48: Why AI debate needs a reliable judge
Ask an on-call engineer which multi-agent system they would rather be paged for, and they will pick the pipeline every time.
Research flows to analysis flows to writing flows to review. One direction, fixed stages, one baton. Nothing about it will impress a conference audience. Everything about it will shorten an incident.
Three operational virtues
Every boundary is a contract. Stage N receives X and produces Y. Violations surface at the boundary, immediately, instead of three stages downstream as a mystery. "The analysis stage emitted something the writer could not use" is a far better bug report than "the output was wrong."
Debugging is bisection. Inspect the artifact at each handoff and find the leg where it degraded. Minutes, not an afternoon of re-running the whole thing with extra logging.
Stages upgrade independently. Swap the analysis model, run the same fixtures, diff the outputs. You do not rewire conversations or change coordination, and you do not wonder whether you broke something two hops away.
That last one compounds over time. Pipelines are the only multi-agent topology where you can confidently improve one component in isolation.
Validate between stages, and store what passed
The most common pipeline failure is the quiet one: a stage emits something subtly wrong, the next stage accepts it without checking, and four stages later the output is nonsense with no obvious origin.
Two fixes, both cheap. Validate at each boundary: schema, required fields, sanity checks on the actual content. And persist intermediate outputs, because the alternative is debugging a five-stage pipeline from its final output alone, which is roughly like debugging a program from its exit code.
The honest limit
Pipelines assume work flows forward. The moment stage 4 routinely rejects and returns work to stage 1, you no longer have a pipeline. You have a loop wearing a pipeline costume, and the costume hides the iteration count from everyone including you.
Keep the pattern anyway, and model the loop explicitly when it appears, with a cap, rather than letting the "simple" architecture quietly become the expensive one.
A large share of production agent systems are pipelines that briefly flirted with something more autonomous and came back. That is not a failure of ambition. Sequential workflows with explicit contracts are the right answer for known processes.
More on these topics
Deep dive · · 6 min read
Adding a second agent does not add intelligence, it adds a contract
Six posts on multi-agent systems, and the failures were never inside an agent. They were between two of them.
Deep dive · · 6 min read
Six ways to wire agents together, and the same three things break every time
The topology gets all the design attention. Ownership, termination and traceability are what decide whether it survives contact with production.
Deep dive · · 6 min read
Most agent controls do not actually control anything
Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.
Discussion