Pattern Matrix/White Paper/C3
ADPS Agent Design Pattern White Paper
C3 · Adversarial Review
Separate proposal, review, and adjudication roles; give reviewers independent evidence and a shared rubric before release.
| Coordinate | Collaboration × Loop (debate) |
| Cost | High (multiple independent roles, review rounds, and audit traces) |
| Pattern group | Collaboration patterns |
| Summary | Assign one agent to produce a proposal and one or more separately configured reviewers to challenge it. Define reviewer duties, evidence sources, decision authority, and exit conditions before the review begins. |
Problem
Several prompts against the same model and evidence do not by themselves establish an independent review. Separation may require different organizational roles, evidence sources, runtimes, models, permissions, and decision authority, depending on the applicable policy.
Adversarial review addresses how a high-stakes decision receives structurally independent scrutiny. In financial credit, medical assistance, legal assessment, and regulatory workflows, the required form of independence comes from the applicable policy and the reviewing organization; using several prompts or vendors does not by itself prove compliance. Compared with single-agent Generator-Critic, this pattern also needs model routing, separated context and evidence, independent traces, and an explicit adjudication contract.
Classification: Collaboration × Loop
- Vertical axis · Collaboration: One agent proposes a solution and another is assigned to test it. Independence may come from model family, context, evidence source, runtime, review instructions, organizational ownership, or a combination of these controls. A vendor difference is one possible signal, not a sufficient condition.
- Horizontal axis · Loop: A proposal, challenge, response, and adjudication can repeat within a bounded review budget. Each round preserves findings, evidence, accepted changes, and unresolved disagreement.
Solution and mechanics
A single adversarial review consists of three roles and one loop:
- Proponent proposes: the generator agent produces the initial solution.
- Critic reviews: An independently configured reviewer tests the proposal against assigned failure hypotheses, evidence, and rubric.
no_issues_foundis valid only when the trace shows adequate coverage. - Proponent revises: the solution is revised based on the issues the critic raised, and the next round begins.
- Decision owner adjudicates: A configured judge, policy rule, or qualified person issues
approve / conditional / reject / needs_more_evidenceunder an explicit authority contract.
Each role writes a separate trace into the audit log. Review rounds need a hard ceiling based on risk and budget. The critic must be capable of testing the generator's subtle errors, and the judge needs the evidence contract required to settle disputes. Cross-family routing can reduce some shared blind spots. Independence also depends on context, data sources, runtime separation, and human accountability.
Applicability
- Workflows with separation-of-duties requirements: Credit, medical assistance, legal assessment, and regulatory processes may require independent evidence and accountable review before release.
- Critical decisions where errors are costly: Use it when expected loss, regulatory responsibility, or safety risk justifies independent review overhead.
- Known correlated reviewer failures: When replay shows that self-review misses a defect class, introduce different evidence, review duties, deterministic checks, models, or human expertise and measure the change.
- Model-combination trade-offs: A mid-tier generator plus an independent critic may outperform a single expensive call on some tasks, but that comparison must be run on the target evaluation set.
Known failure modes
- Correlated blind spots: changing vendors does not guarantee independent errors; generator and critic models may still share training sources, evaluation weaknesses, or prompt assumptions. For high-risk cases, diversify evidence and review duties, add deterministic checks where possible, and retain a qualified human decision owner.
- Context growth across rounds: Cost and latency rise when every call carries the full discussion. Give each reviewer the task, current draft, assigned checks, and required evidence. Give the decision owner a structured record of findings, responses, and unresolved issues.
- Critic agrees by default: A reviewer may accept the proposal without testing its assigned failure hypotheses. Track review coverage and confirmed findings, and sample verdicts with qualified reviewers. A critic that never finds substantive issues and one that continually invents them are both miscalibrated.
- Insufficient reviewer diversity: when every reviewer uses the same role, prompt, and evidence, additional agents tend to repeat the same judgment. Give reviewers distinct failure hypotheses, evidence sources, or review duties, and measure whether disagreement reveals defects that one reviewer misses.
- Misuse on simple tasks: Low-risk writing polish and tasks with deterministic verification usually do not justify multi-agent adversarial review. Tests, schemas, or a single calibrated critic are more direct.
Verification and metrics
- Independence evidence: Record differences in models, prompts, evidence sources, runtime, and organizational roles. Vendor diversity is only one piece of evidence.
- Critic discovery rate (
issues_per_turn): Analyze it together with phantom issues and expert sampling. Persistently empty or persistently fabricated findings both signal poor calibration. - Debate rounds: Observe the convergence distribution and cap exhaustion. Repeated non-convergence calls for a better judge, evidence contract, or exit from the pattern.
- Cost per decision: Measure the real token, latency, and review cost of the chosen configuration and compare it with the risk budget.
Reference implementation
AdversarialReview.review(task):
check_independence_policy() # context, evidence, runtime, model, or role
current = proponent(task)
for round in 1..MAX_ROUNDS: # configured risk and cost ceiling
verdict = critic(task, current, review_contract)
append_trace(current, verdict)
if verdict.no_issues_found and requires_second_review(verdict):
verdict = second_critic(task, current, review_contract)
append_trace(current, verdict)
if verdict.no_issues_found: break
current = proponent(task, current, verdict.issues)
return judge(task, current, debate_log) # rule, qualified person, or judge agent
CritiqueVerdict:
issues: list # explicit beats implicit
severity: minor/major/critical
no_issues_found: bool # valid only when supported by the review trace
Production implementation should record model versions, prompts, evidence sources, and runtime identity; define when an empty critic verdict requires a second review; configure MAX_ROUNDS from task risk and cost; and retain compliance attestations according to the applicable policy.
Illustrative scenario
Consider a bank's loan-decision assistance system. A first version asks one model to play credit analyst, risk reviewer, and compliance reviewer. Different role prompts do not establish structural independence. A later design separates the generator, critic, and judge by model or evidence source, retains each trace, and writes the final ruling and human approval into the audit record. Whether this satisfies an audit depends on the bank's policy and the auditor's accepted evidence; it cannot be inferred from the number of models alone.
Related patterns
- Generator-Critic (F1): F1 generates and revises an artifact through a bounded feedback loop and may share a model or runtime. C3 adds independent review roles, evidence, and traces when separation of duties or stronger challenge is required.
- Fan-out/Gather (C2): workers process separate units or viewpoints and a gather step combines their outputs. C3 assigns reviewers to challenge the same proposal against distinct checks.
- Parallel Exploration (R3): branches develop alternative hypotheses or solutions for later selection or aggregation. C3 starts with a proposal and searches for defects, unsupported claims, and policy violations.
- N-version programming: independent implementations of one specification are compared or voted on to reduce common-mode software failures. It is a useful precedent when several agents independently implement the same task; a proposal-and-critique loop has different control flow and should not be treated as equivalent.
Design conclusion
Adversarial review organizes disagreement through independent roles, evidence, and an explicit ruling rule. Independence is a structural claim that the system must demonstrate, not a label earned by adding another prompt.
Workshop revision, 25 August 2026: Review-execution feedback
A review result carries evidence, risks, applicability conditions, and revalidation requirements. Constraints discovered during execution return through the evidence chain so role separation does not create a new information gap.
Suggested citation: ADPS, C3 Adversarial Review, Agent Design Pattern White Paper v0.3, 2026-07-13. Catalog · runnable code catalog · CC BY 4.0
Document status: This is a public review draft. Illustrative scenarios explain the mechanism and are not presented as verified enterprise cases. See the case library for attributed practice. ADPS welcomes case contributions with sources, measurement methods, and publication approval.
Chronicle
- Recorded source
- ADPS pattern white paper; prior work and references are listed in the article
- First published on ADPS