Pattern Matrix/White Paper/R2
ADPS Agent Design Pattern White Paper
R2 · Complexity-Based Routing
Route each request to a model and effort tier using task complexity, risk, and acceptance evidence. Record the decision and escalate when the selected tier fails its checks.
| Coordinate | Reasoning × Route (selection) |
| Cost | Medium (one routing decision in exchange for task-specific model allocation) |
| Pattern group | Reasoning patterns |
| Summary | Route each request to a model and effort tier using task complexity, risk, and acceptance evidence. Record the decision and escalate when the selected tier fails its checks. |
Problem
Using the most capable and expensive model for every request allocates the same resources to template filling and complex analysis. Prices and model capabilities change, so routing should compare current candidates on a local task set rather than rely on a fixed price-ratio table.
Complexity-based routing assigns model resources from task difficulty and risk. Simple requests use a lower-cost tier; difficult or high-risk requests use a more capable tier. Savings depend on the traffic distribution and model mix and should be calculated from replayed production traces. The pattern can coexist with parallel exploration: routing allocates routine requests, while parallel branches add verification for selected decisions.
Classification: Reasoning × Route
- Vertical axis · Reasoning: The router selects a model, effort level, or reasoning policy for the current request. It does not decompose the request among multiple agents.
- Horizontal axis · Route: Complexity, risk, domain, and historical outcomes choose one initial execution tier, with explicit escalation if its acceptance checks fail.
Solution and mechanics
A single complexity-based routing pass consists of three stages:
- Extract signals + classify: Derive complexity signals from the query, domain, and historical outcomes, then use rules or a lower-cost model to choose a tier. Include the classifier's own cost and latency in the evaluation.
- Tiered execution: Define capability and cost tiers from local evaluation results while keeping a shared call interface. Routing may select both a model and its effort setting.
- Escalation fallback: After a lower-cost tier answers, check its schema, evidence, and business constraints. If it falls short, escalate and retry. The chain needs a configured ceiling; failure at the top should raise an error or hand off to a human.
Production systems commonly use three routing approaches:
| Control location | Form | Trade-off |
|---|---|---|
| Internalized in the model | The model or provider selects an internal reasoning path | Low integration effort; limited routing trace and portability |
| Explicit in the harness | The application owns classification, acceptance, and fallback | More implementation work; explicit trace and multi-provider control |
| Third-party router service | A routing service selects the target model | Common interface; additional dependency and data-processing boundary |
Choose the control location from the evidence and governance needed. An explicit harness supports local replay, policy overrides, and provider comparison. Provider-managed or third-party routing may fit when its trace, data boundary, and fallback behavior meet the product's requirements.
Applicability
- Mixed request difficulty under a cost or latency budget: internal BI, customer support, and document processing often combine routine requests with a smaller set of high-risk or multi-step cases. Routing allocates different model and effort tiers while preserving acceptance thresholds.
- Multi-vendor mixed scenarios: a team using Claude / DeepSeek / its own fine-tuned model at the same time cannot hand routing to a single vendor and must do it at the engineering layer.
- Specialized agents that split by action type: code changes may use the primary model while routine Git operations use a lower-cost model. This routes by action type instead of query complexity.
Known failure modes
- Classifier built only on fixed rules: Keyword matching may misroute unseen query forms. Add uncertainty-based escalation or a lower-cost model, and compare the cost of false downgrades on a replay set.
- Acceptability check looks only at length: checking only whether the output is long enough lets through results that violate the schema, exceed numeric bounds, or cite the wrong source data. Do schema-aware validation.
- Escalation costs more than the target tier: A lower tier followed by escalation can exceed the direct cost and latency of the final tier. Review request classes with repeated escalation and adjust their initial route.
- High-risk queries also go to the low tier: Finance, privacy, or compliance queries may need a reliable tier even when syntactically simple. Encode risk overrides outside the model.
- Routing without meaningful alternatives: If every request uses the same model, effort level, and control path, the router adds latency without changing execution. Add routing only when evaluated alternatives produce a useful cost, latency, or quality trade-off.
Verification and metrics
- Per-query cost distribution: Break cost down by routing tier and task class. When the high-capability tier changes materially, distinguish harder traffic from classifier error.
- Routing accuracy: Use labels, human review, or a stronger reviewer to judge whether the selected tier met the task requirement. Track unsafe downgrades separately.
- Fallback rate: The share of lower-tier results that escalate. Changes call for inspection of thresholds, model capability, and input distribution.
- Routing decision time: Measure the classifier's contribution to total latency. If it becomes material, use rule prefilters, caching, or asynchronous features.
Reference implementation
Query comes in → extract complexity signals (length / keywords / domain / history)
classifier (rules or evaluated classifier model) → choose tier + confidence
High-risk query (finance / privacy / compliance) → apply policy override to an approved tier
Execute:
selected tier runs → result acceptable? → return
not acceptable → escalate (schema validation + cost estimate)
escalation chain reaches its configured ceiling; still failing at the top → raise error / hand off to human
Emit a trace on every routing decision (query summary / tier / signals / confidence / actual cost / whether escalated)
Evaluate rules and classifier models on the same routing set. Make acceptance checks schema-aware, cap the total cost and latency of fallback, and review request classes that repeatedly escalate so they can start at a more suitable tier.
Illustrative scenario
Consider an internal BI agent serving product, growth, and finance teams. Trace review shows that some requests fill SQL templates or add grouping, while others require multi-step attribution or causal analysis. The first version uses the most capable model for all of them and cannot explain whether resources are being spent on difficult work.
The revised system assigns tiers to template queries, aggregation, attribution, and high-complexity analysis; combines a rule prefilter with a lower-cost classifier; estimates total cost before escalation; forces finance, privacy, and compliance requests through risk-based overrides; and retains every routing trace for replay. Cost and quality changes are calculated from those traces rather than assumed from an industry percentage.
Related patterns
- Parallel exploration (R3): complementary. Routing picks one tier at a single point in time to save money; parallelism opens N lines at once to buy quality. Within the same agent, everyday work goes through routing and critical decisions start parallelism.
- Chain of thought (R1): the tier that routing selects already contains the CoT effort tier. Routing is the version of CoT effort control raised from a single call to the task level.
- Dual-mode architecture (R5): dual-mode is the architectural extreme of routing—ordinary routing is "pick one tier and run," dual-mode is "run two tiers at once," pushing "tier selection" all the way to "splitting the agent."
- Failure journal / guard-type patterns: forcing high-risk queries through the most reliable tier shares the same protective reasoning as "some boundaries cannot be skimped on."
Design conclusion
Complexity-based routing places token cost, latency, and answer quality on a Pareto frontier. The business SLA determines which point on that curve a request should use. This makes routing a product-economics decision with measurable trade-offs, rather than a fixed model preference.
Suggested citation: ADPS, R2 Complexity-Based Routing, Agent Design Pattern White Paper v0.3, 2026-07-13. Catalog · runnable code catalog · CC BY 4.0
Document status: This is a public review draft. Illustrative scenarios explain the mechanism and are not presented as verified enterprise cases. See the case library for attributed practice. ADPS welcomes case contributions with sources, measurement methods, and publication approval.
Chronicle
- Recorded source
- ADPS pattern white paper; prior work and references are listed in the article
- First published on ADPS