Want to find the weakness in an AI-generated plan? Assign a second AI to attack it.

Models critique other models' work far better than their own, and the debate pattern industrialises that asymmetry. Two agents argue genuinely different positions; a judge scores them against stated criteria. Under adversarial pressure, assumptions surface that neither agent would have volunteered about its own proposal.

The whole pattern rests on the judge

Not on the debaters. This is the part that gets under-designed.

A judge that rewards eloquence over evidence converts the entire apparatus into an expensive coin flip weighted toward whichever side writes more confidently. The error is still there, and now it has a courtroom and a verdict, which makes it harder to question downstream.

Three checks before trusting any verdict:

  • Evidence access. Does the judge see the underlying data, or only the arguments about it? Rhetoric-only judging is vibes with a gavel.
  • Criteria fixed in advance. "Pick the better plan" invites charisma. "Score against these five requirements, weighted as follows" constrains it. Write them before the debate, not after seeing the arguments.
  • The swap test. Rerun with the debaters' positions exchanged. If the verdict follows the agent rather than the argument, your judge is measuring style. This is the single most useful diagnostic here and it takes one extra run.

Validate the judge on known answers

Before deploying, run the judge on cases where you already know which side is right. Skip the obvious cases and use contested ones with a determinable answer.

If it cannot reliably pick the correct side when you know the ground truth, it will not do better on live questions where you do not. That sounds obvious written down; almost nobody does it.

Reserve it for decisions that justify the cost

Debate triples inference for one answer. That is defensible for architecture decisions, go/no-go gates, and claims where being wrong is expensive.

It is not defensible for routine tasks, and routine tasks are where it tends to get switched on because the pattern is interesting.

Closing thought

Debate exposes hidden weaknesses, when the judge, the evidence, and the decision rules are sound. Otherwise it is a well-structured way to pick the more charismatic wrong answer, and to feel rigorous while doing it.

Where would an adversarial second opinion have saved you, and who would you trust to judge it?

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.