When one agent stops being enough
The first agent you build will do one thing well. The second will do two things badly. That is the moment to split responsibilities instead of adding instructions to one prompt.
The supervisor pattern
A planner receives the user goal and decides which specialist should act. Each specialist, such as a research agent or an analysis agent, owns a narrow toolset: data access, external APIs or a human feedback channel. The supervisor merges results and decides whether the goal is met.
supervisor = Agent(
name="Supervisor",
instructions="Coordinate the workflow between agents.",
sub_agents=[research_agent, analysis_agent],
)Contracts between agents
In production the pattern depends on the contract between agents more than on the prompt: typed inputs, typed outputs and a trace identifier that follows the request through every hop.
Trade-offs
- Latency grows with every hand-off. Run specialists in parallel where the plan allows it.
- Cost multiplies. Use a smaller model for routing and a larger one only where judgment matters.
- Debugging needs a timeline view, because logs alone are not enough. Emit spans per agent step.
Where to put humans
Any step that writes to a system of record, spends money or sends a message to a customer should pause for approval until you have months of evidence that it does not need to.
Key takeaways
- Use a clear orchestration pattern, such as a supervisor or a fixed workflow.
- Give every agent one job and one set of tools.
- Put a human feedback step where the cost of a wrong answer is high.
- Trace every hop. Multi-agent systems fail quietly between agents.
More on these topics
Comparison · · 1 min read
LangGraph vs the OpenAI Agents SDK
Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Discussion