Pattern Matrix/White Paper/Reflection
ADPS Agent Design Pattern White Paper · Module Overview
Reflection Module: Turning Feedback into Controlled Change
Evidence, online and offline feedback loops, change authority, four specifications, and open questions.
Scope: This page describes the reflection subsystem as a whole. The F1–F4 specifications remain authoritative for pattern-level mechanics and verification.
Reflection may happen after one output or after a batch of completed runs. It reads artifacts, trajectories, and external outcomes; compares them with inspectable criteria; decides whether to change the current output, execution path, or reusable assets; and then gathers new evidence to determine whether the change helped.
Asking a model to think again is the lightest implementation. Without evidence, a stopping rule, and a defined change boundary, another inference pass does not establish that the system improved.
Four questions before adding a reflection loop
- What signal starts the review? A failed test, a rule violation, user dissatisfaction, an expert label, or a model score?
- What evidence can decide quality? Compilation, schemas, business rules, expert judgement, and user behaviour differ in reliability and arrival time.
- What may this loop change? The current answer, plan, Skill, memory, business rule, or production code?
- How will the change be verified? Re-running the original check, regression suites, controlled comparisons, human review, and business outcomes answer different questions.
If any of these questions has no answer, keep the loop in observation or recommendation mode.
Observability, evaluation, reflection, and release
| Stage | Output | Effect on the system |
|---|---|---|
| Observability | Traces, logs, tool calls, state changes, and cost | Records what happened |
| Evaluation | Pass/fail, scores, issue classes, and evidence | Judges an output or path against criteria |
| Reflection | A scoped, evidence-backed change proposal | Attempts to change an artifact, path, or reusable asset |
| Release and governance | Approval, version, canary, rollback, and accountability record | Decides whether the change receives durable authority |
Evaluation produces a judgement. Reflection consumes that judgement and proposes or applies a change. When a change enters a shared Skill library, memory store, policy repository, or production environment, a governance gate owns the release decision.
Two clocks: online and offline reflection
At the ADPS Reflection workshop on 12 August 2026, Wei Wang and Jiaqi Li described the same boundary from different production settings: reflection runs on two clocks.
| Online reflection | Offline reflection | |
|---|---|---|
| Purpose | Complete the current task or stop it safely | Improve later runs and correct system-wide problems |
| Signals | Tests, rules, tool receipts, and structural checks available now | Batches of traces, bad cases, expert labels, user feedback, and delayed business outcomes |
| Scope | The current output and local path | Cross-run, cross-version, and cross-team behaviour |
| Change authority | Usually limited to the current output, parameters, or rollback-safe changes | May propose changes to Skills, memory, harnesses, datasets, and workflows; release still requires validation |
| Time budget | Milliseconds to minutes, with explicit iteration, latency, and token limits | Hours, days, or weeks, suitable for batch comparison and human participation |
These are operating modes, not new pattern coordinates. F1 and F4 often run online, while their rubrics, failure classes, and stop thresholds need offline calibration. F2 loads Skills online but creates, tests, evaluates coexistence, and releases them mainly offline. F3 also spans both clocks: experience is distilled offline and replayed when a later task needs it.
When the outcome arrives
The arrival time of the success signal determines the reflection clock.
| Success signal | Typical settings | Recommended loop |
|---|---|---|
| Compilation, tests, schemas, deterministic tool receipts | Code, configuration, structured write-back | Online repair and re-verification can be appropriate for low-risk changes |
| Stable rules, reference sets, domain checkers | Policy interpretation, regulated content, domain queries | Online review may work when the rule version is pinned |
| User satisfaction, business metrics, downstream human edits | Support, operations, analytical recommendations | Collect outcomes and analyse them offline; do not manufacture an immediate ground truth |
| Long feedback chains across teams and systems | Requirements to release, recommendation to business action | Preserve delayed labels and use human attribution before changing the system |
Open-ended work can still receive online checks for citations, structure, contradictions, and known risks. The final business-value judgement waits for the relevant outcome.
The reflection contract
A production loop should express its control surface as data:
reflection_id: ref_01K2...
scope: artifact # artifact | trajectory | asset | system
trigger:
type: test_failure
ref: trace://run-8842/check-9
evidence:
- type: deterministic_test
ref: test://payroll/net-pay-balance
judge:
policy: reflection-rubric-v4
proposal:
target: src/payroll/net_pay.py
allowed_change: diagnosed_files_only
authority:
mode: auto_in_sandbox # suggest | auto_in_sandbox | approval_required
budget:
max_iterations: 3
max_latency_ms: 90000
verification:
suite: payroll-regression-v12
rollback_on_regression: true
retention:
disposition: candidate_lesson
The full loop is Trigger → Evidence pack → Diagnose → Change proposal → Policy gate → Apply → Verify → Record. A model may diagnose and propose. Tests, rules, people, and business outcomes decide whether it may continue.
How the four specifications divide the work
| Pattern | Primary change target | Feedback | Main boundary |
|---|---|---|---|
| F1 Generator-Critic | Current artifact | Rules, references, expert rubrics, or an independent model | The critic must be evaluated; a score cannot replace evidence |
| F2 Skill Package | Reusable workflow and capability asset | Cross-run outcomes, failures, and coexistence tests | A Skill that passes alone may still degrade a larger Skill set |
| F3 Experience Replay | Context supplied to a later task | Historical trajectories, outcomes, and delayed feedback | Preserve source, version scope, and uncertainty to avoid negative transfer |
| F4 Self-Heal Loop | Rollback-safe state, configuration, or code | Deterministic failure and re-run evidence | Missing knowledge, changed business logic, and irreversible actions require escalation or an offline release |
Online/offline, hard evidence/soft judgement, local/global scope, and feedback delay are selection parameters. They do not alter the cognitive-function × execution-topology framework.
A four-layer production structure
Mo Zhou proposed a useful four-layer implementation:
- Hard validation: Compilation, tests, schemas, business invariants, and system receipts take priority.
- Reflection diagnosis: Models inspect evidence and identify possible defects in outputs, trajectories, or assets.
- Memory and assets: Candidate lessons, Skills, rules, and failure records cross task boundaries only after admission.
- Scheduling and control: The harness decides whether to reflect, which reviewer to use, how many rounds to allow, and when to degrade or hand off.
Automatic triggering, termination, observability, and bounded cost apply to every production loop. Reuse needs a narrower rule: a one-run correction may be deliberately discarded; anything that will affect later runs requires admission, versioning, and retirement.
Three problems that appear at scale
Local patches can damage global efficiency. Wei Wang described an environment-building agent that accumulated local fixes over time. Stability improved while the overall path became heavier. Offline reflection must detect duplicate, conflicting, and expired patches instead of appending every bad case to a prompt.
Skills must be evaluated in combination. Trigger accuracy and task success show that one Skill works alone. A growing library also needs false-trigger, missed-trigger, overlap, conflict, loading-cost, and coexistence tests.
Attribution comes before healing. Qianchun Lu separated runtime, execution-process, business-result, and experience failures. Invalid parameters, network instability, and missing steps may have direct remedies. Missing domain knowledge, changed business rules, and conflicting team goals cannot be filled in safely by the original agent.
Engineering progress in 2026
By August 2026, major platforms expose the evaluation components that reflection depends on:
- LangSmith Evaluation separates pre-release offline evaluation from online evaluation over production traces and feeds failures back into datasets. AgentEvals evaluates tool-use trajectories directly.
- Anthropic's January 2026 guide, Demystifying evals for AI agents, recommends combining code, model, and human graders. Its enterprise Skills guidance adds triggering, isolation, coexistence, instruction-following, and output-quality checks.
- OpenAI's AgentKit brings trace grading, datasets, and graders into the agent optimisation workflow.
These systems first address seeing and judging. Reflection begins when a judgement becomes a controlled change and regression evidence confirms the result.
Directions without new catalog numbers
Meta-reflection reviews the critic and its rubric. It currently fits as a high-risk F1 configuration rather than a separate pattern.
Deliberative reflection uses reviewers with different assumptions in a structured debate. It overlaps F1 Generator-Critic and C3 Adversarial Review, and still lacks stable stopping and cost evidence.
Prospective reflection checks plans, assumptions, and risk before action. Its responsibilities already appear across F1, A4 Guardrail Sandwich, and G1 Approval Gate. More independent practice is needed before cataloguing it separately.
Questions for further industry evidence
- How can a large Skill estate separate agent quality, Skill quality, and combination effects?
- When dimensions in a rubric move in opposite directions, who owns the trade-off?
- Which reflections may change a harness automatically, and which may only open a change proposal?
- When outcomes arrive days or weeks later, how should the system preserve the original trajectory, version, and accountability chain?
- How can teams detect, consolidate, remove, or roll back conflicting patches accumulated by online loops?
Workshop record
This overview incorporates the first ADPS Reflection Module Workshop held on 12 August 2026. Hosts: Haili Zhang and Jia Huang. Core workshop guests: Dong Zhang, Mo Zhou, Wei Wang, Qianchun Lu, Han Zhao, Pylon Peng, and Jiaqi Li. The attributed points on this page come from the six speakers recorded in the transcript: Dong Zhang, Mo Zhou, Wei Wang, Qianchun Lu, Pylon Peng, and Jiaqi Li.