Pattern Matrix/White Paper/R3
ADPS Agent Design Pattern White Paper
R3 · Parallel Exploration
Run independent reasoning branches, preserve their evidence, and aggregate only when measured quality gains justify the added cost.
| Coordinate | Reasoning × Parallel (fan-out) |
| Cost | High (multiple branches plus aggregation) |
| Pattern group | Reasoning patterns |
| Summary | Run independent reasoning branches, preserve their evidence, and aggregate only when measured quality gains justify the added cost. |
Problem
One reasoning run may omit a relevant feature or settle on a weak hypothesis. Repeating the same configuration can reproduce the same blind spot, while ordinary review sees only the selected result.
Parallel Exploration runs independently configured branches for the same decision and combines their artifacts under an explicit aggregation rule. It uses additional compute when broader evidence, alternative hypotheses, or independent checks justify the cost.
Classification: Reasoning × Parallel
- Vertical axis · Reasoning: R3 runs multiple candidate solutions for the same question. C2 Fan-Out/Gather distributes independent subtasks. Compare the unit of work: R3 branches compete or corroborate on one decision, while C2 branches each produce a different part of the final artifact.
- Horizontal axis · Parallel: Branches execute concurrently from a shared task contract without reading one another's intermediate state. A verifier or aggregator combines their typed results.
Solution and mechanics
A single round of parallel exploration has three stages:
- Dispatch: Replicate the query into N branches and create diversity through prompts, sampling settings, models, or evidence sources. Choose N from task risk, branch correlation, and budget, using local ablation tests to find diminishing returns.
- Isolated execution: Each branch has independent intermediate state and failure handling. Shared evidence may be intentional, but one branch must not copy another branch's conclusion before aggregation.
- Aggregation: use an aggregation strategy to synthesize the N results into one. Majority vote is only one option; the selected rule should encode the cost of different errors.
Choose the aggregation strategy from the output type and the relative cost of false positives, false negatives, and unresolved disagreement:
| Aggregation strategy | Suitable scenario |
|---|---|
| Majority | Enumerable answers, symmetric error cost (math, classification) |
| Weighted | Branches differ in reliability (different models / different compute tiers) |
| Verifier | Open-ended answers (writing, code, planning) |
| First-Correct | A clear success criterion exists (test-driven) |
| Any-Alarm | High-risk with asymmetric error cost (healthcare, finance, security) |
Applicability
- High-risk judgments with asymmetric error cost: Medical image triage, financial controls, anti-money-laundering, and vulnerability review may route any predefined high-risk finding to qualified review instead of accepting a majority vote.
- Large answer spaces where a single chain is unstable: complex diagnosis, multi-hop reasoning, and tasks that need self-consistency to improve reliability.
- High-consequence decision points: A final review before an irreversible commitment may justify multiple evidence paths, while the actual confidence gain must be measured on the target evaluation set.
Known failure modes
- Correlated branches: Identical prompts, evidence, models, or shared intermediate state can produce duplicate conclusions. Measure effective branch diversity and isolate mutable state.
- Insufficient prompt perturbation: When multiple branches return nearly identical answers, parallel exploration has degraded into duplicate sampling and the added compute has produced no new evidence. Check whether sampling settings, prompts, models, and evidence sources are genuinely independent.
- Using majority for asymmetric risk: A minority high-risk finding can be lost under a vote. Define escalation and abstention rules from the business error model.
- Incorrect early termination: An Any-Alarm policy may stop immediately after a qualifying alarm, but it cannot conclude “no alarm” until every required branch completes or the policy records an incomplete result.
- Overusing parallelism: Running multiple branches on simple or low-risk tasks adds cost without a decision benefit. Compare against a single-chain baseline.
Verification and metrics
- Branch agreement rate: Interpret agreement together with task difficulty. High agreement may indicate a simple task or insufficient branch independence.
- Effective N: Count independent evidence paths or conclusions. Concurrent branches that repeat the same reasoning do not increase effective N. When it stays low, vary prompts, models, or sources before adding branches.
- Aggregation cost share: Measure the aggregator's share of total cost and latency. If it dominates, use lighter aggregation or stricter artifacts.
- Quality gain: Compare parallel and single-chain runs on the same evaluation set and report cost and latency alongside quality.
Reference implementation
replicate query into N branches:
each branch → independent runtime → sample at a different temperature → (answer, confidence)
aggregate(N results, strategy):
Majority → the answer with the most votes
Weighted → score results with calibrated branch reliability or an external rubric
Verifier → hand off to an independent verifier model to score and decide
Any-Alarm → if any branch hits a high-risk label, escalate, ignoring the majority
return final_answer + complete branch trace (per-branch answer / confidence / aggregation strategy / final decision)
Evaluate branch diversity, verifier quality, and aggregation rules on the target task set. Respect provider limits, retain each branch trace, and test incomplete, timeout, disagreement, and Any-Alarm paths explicitly.
Illustrative scenario
Consider a medical-imaging assistant that grades pulmonary nodules. A single chain may miss a suspicious morphology. The revised system uses independent branches with varied prompts or evidence views. Aggregation follows an Any-Alarm rule rather than majority vote: any branch detecting a predefined high-risk sign sends the case to a human second review. Branches run in isolated runtimes and are not terminated early in Any-Alarm scenarios. Accuracy and compute must be reported on an approved clinical evaluation set, with retention governed by the institution's policy.
Related patterns
- Complexity-Based Routing (R2): R2 can reserve R3 for task classes whose measured quality gain justifies its cost and latency.
- Chain of Thought (R1): Each branch retains its own ordered evidence and decision artifacts under the R1 trace contract.
- Fan-Out/Gather (C2): R3 branches address the same decision with alternative reasoning or evidence. C2 branches produce distinct parts of one deliverable.
- Iterative Hypothesis Testing (R4): R3 evaluates branches concurrently; R4 revises a hypothesis set across evidence-gathering rounds. R4 may use R3 within one round.
Design conclusion
Parallel Exploration runs independently configured reasoning branches and combines their artifacts under an explicit rule. Verification must cover branch diversity, aggregation errors, incomplete results, cost, and latency against a single-run baseline.
Suggested citation: ADPS, R3 Parallel Exploration, Agent Design Pattern White Paper v0.3, 2026-07-13. Catalog · runnable code catalog · CC BY 4.0
Document status: This is a public review draft. Illustrative scenarios explain the mechanism and are not presented as verified enterprise cases. See the case library for attributed practice. ADPS welcomes case contributions with sources, measurement methods, and publication approval.
Chronicle
- Recorded source
- ADPS pattern white paper; prior work and references are listed in the article
- First published on ADPS