Teams copy a pattern from a framework tutorial and inherit its control flow without noticing they made a choice. Then something goes wrong in production and the first question, which agent decided to do that, has no answer.

The patterns are less different architectures than different answers to where control sits and when it comes back.

The patterns, and what each one actually decides

Supervisor. One agent holds control for the whole run. It picks the next specialist, reads the result, decides again. Control returns after every step, which makes it the most inspectable option and the most expensive one: the supervisor's context grows with every result it absorbs, and it becomes the bottleneck and the single point of failure.

Orchestrator-worker. Control fans out. The orchestrator splits work, dispatches identical or near-identical workers, and merges. Workers do not know about each other, which is the whole point: they must be genuinely independent or the merge produces contradictions nobody catches. The merge step is where this pattern fails, and it is usually the step that gets the least design attention.

Planner-executor. Control is handed over once. A planner produces a full plan, an executor runs it. This is fast and cheap, and it breaks the moment reality diverges from the plan, because the executor has no authority to replan. It works when the environment is stable and the plan is verifiable before execution starts. It does not work when step three's output determines whether step four makes sense.

Pipeline. Control is in your code. Fixed stages, fixed edges, models at the nodes. Most systems described as multi-agent are this, and they should be: it is testable, cheap and debuggable. Choose it deliberately rather than pretending you have something more sophisticated.

Debate and judge. Control is distributed and then arbitrated. Useful when the failure mode is a confident single answer and you can write a rubric that discriminates. Without a calibrated rubric, the judge picks the more fluent answer and you have paid three times as much for the same reliability.

Choosing by uncertainty and blast radius

Two axes decide it. How much does the next step depend on what the previous step returned, and how bad is a wrong action.

Low uncertainty means a pipeline. High uncertainty with low blast radius means a supervisor, because you want the replanning and you can afford mistakes. High uncertainty with high blast radius means a supervisor with approval gates on the irreversible steps rather than more autonomy. Independent parallel subtasks mean orchestrator-worker, and you should staff the merge before you staff the workers.

Whatever you choose, one rule holds: the pattern must be visible in the trace. If you cannot read a run and say which component held control at each moment, what you implemented is a suggestion rather than a pattern.

Pull up your last production run. At the moment it went wrong, which component was holding control?

Comparison · · 1 min read

LangGraph vs the OpenAI Agents SDK

Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.