A user reports that the answer was wrong. You have the answer, the original request, and four agents that each made decisions in between. The output tells you a mistake happened somewhere. It does not tell you where, and rerunning the request produces a different path that succeeds.

Single-agent debugging survives on output inspection. Multi-agent debugging does not, because the interesting decisions are the ones nobody saw.

What the trace has to record

One timeline for the whole run, with every event carrying the run id, the role that emitted it and a parent span, so the causal structure survives when workers execute concurrently.

Record the handoff payloads themselves, both sides. What the sender produced and what the receiver actually consumed after validation, because the delta between those two is where context silently disappears. Record every tool call with its arguments and result, the budget consumed at each step, and the point at which control passed from one role to another.

Then make it replayable. If you cannot reconstruct a run without calling the models again, you cannot investigate a failure you have already stopped reproducing.

Evaluate the path, not just the answer

Output scoring is blind to most coordination failures. These are the ones that reach production:

Failure modeWhat the final answer looks likeWhat catches it
Context lost at a handoffFluent, quietly missing a constraintDiff the sent payload against the consumed one
Workers duplicating effortCorrect, at four times the costTool call overlap across sibling workers
Judge preferring fluencyConfidently wrongRubric scored against a labelled disagreement set
Plan diverging from realityPlausible but staleCheck each step's precondition at execution time
Silent partial completionComplete-looking, built on half the sourcesAssert coverage against the expected input set

Score handoffs as first-class units. A handoff eval takes a payload and asks whether it carries everything the next role needs to do its job. That is far cheaper to build than a full trajectory judge, and it catches the majority of real defects.

Recovery is a design decision, not an exception handler

Decide per step what happens when it fails, before you ship. Retry with a different approach. Continue with a partial result and mark the gap. Fall back to a simpler deterministic path. Stop and escalate.

Escalation needs the evidence attached. A human reviewer shown only the disputed output cannot approve responsibly; they need what each role saw, what it decided and why. Put the approval gates on the irreversible steps rather than the uncertain ones. Uncertainty is normal; an irreversible step is the one you cannot undo.

When something does go wrong, review the trajectory rather than the answer. The answer is the last thing that happened, and it is almost never where the problem started.

If a handoff dropped a constraint in your system tomorrow, what would surface it before a user did?