Pattern Matrix/White Paper/Reflection

ADPS Agent Design Pattern White Paper · Module Overview

Reflection Module: Turning Feedback into Controlled Change

Evidence, online and offline feedback loops, change authority, four specifications, and open questions.

Version
v0.3
Status
Public review draft
Revised
2026-08-13
Document type
Reflection module overview
Basis
ADPS reflection workshop, 12 August 2026
License
CC BY 4.0

Scope: This page describes the reflection subsystem as a whole. The F1–F4 specifications remain authoritative for pattern-level mechanics and verification.

Reflection may happen after one output or after a batch of completed runs. It reads artifacts, trajectories, and external outcomes; compares them with inspectable criteria; decides whether to change the current output, execution path, or reusable assets; and then gathers new evidence to determine whether the change helped.

Asking a model to think again is the lightest implementation. Without evidence, a stopping rule, and a defined change boundary, another inference pass does not establish that the system improved.

Four questions before adding a reflection loop

  1. What signal starts the review? A failed test, a rule violation, user dissatisfaction, an expert label, or a model score?
  2. What evidence can decide quality? Compilation, schemas, business rules, expert judgement, and user behaviour differ in reliability and arrival time.
  3. What may this loop change? The current answer, plan, Skill, memory, business rule, or production code?
  4. How will the change be verified? Re-running the original check, regression suites, controlled comparisons, human review, and business outcomes answer different questions.

If any of these questions has no answer, keep the loop in observation or recommendation mode.

Observability, evaluation, reflection, and release

Stage Output Effect on the system
Observability Traces, logs, tool calls, state changes, and cost Records what happened
Evaluation Pass/fail, scores, issue classes, and evidence Judges an output or path against criteria
Reflection A scoped, evidence-backed change proposal Attempts to change an artifact, path, or reusable asset
Release and governance Approval, version, canary, rollback, and accountability record Decides whether the change receives durable authority

Evaluation produces a judgement. Reflection consumes that judgement and proposes or applies a change. When a change enters a shared Skill library, memory store, policy repository, or production environment, a governance gate owns the release decision.

Two clocks: online and offline reflection

At the ADPS Reflection workshop on 12 August 2026, Wei Wang and Jiaqi Li described the same boundary from different production settings: reflection runs on two clocks.

Online reflection Offline reflection
Purpose Complete the current task or stop it safely Improve later runs and correct system-wide problems
Signals Tests, rules, tool receipts, and structural checks available now Batches of traces, bad cases, expert labels, user feedback, and delayed business outcomes
Scope The current output and local path Cross-run, cross-version, and cross-team behaviour
Change authority Usually limited to the current output, parameters, or rollback-safe changes May propose changes to Skills, memory, harnesses, datasets, and workflows; release still requires validation
Time budget Milliseconds to minutes, with explicit iteration, latency, and token limits Hours, days, or weeks, suitable for batch comparison and human participation

These are operating modes, not new pattern coordinates. F1 and F4 often run online, while their rubrics, failure classes, and stop thresholds need offline calibration. F2 loads Skills online but creates, tests, evaluates coexistence, and releases them mainly offline. F3 also spans both clocks: experience is distilled offline and replayed when a later task needs it.

When the outcome arrives

The arrival time of the success signal determines the reflection clock.

Success signal Typical settings Recommended loop
Compilation, tests, schemas, deterministic tool receipts Code, configuration, structured write-back Online repair and re-verification can be appropriate for low-risk changes
Stable rules, reference sets, domain checkers Policy interpretation, regulated content, domain queries Online review may work when the rule version is pinned
User satisfaction, business metrics, downstream human edits Support, operations, analytical recommendations Collect outcomes and analyse them offline; do not manufacture an immediate ground truth
Long feedback chains across teams and systems Requirements to release, recommendation to business action Preserve delayed labels and use human attribution before changing the system

Open-ended work can still receive online checks for citations, structure, contradictions, and known risks. The final business-value judgement waits for the relevant outcome.

The reflection contract

A production loop should express its control surface as data:

reflection_id: ref_01K2...
scope: artifact              # artifact | trajectory | asset | system
trigger:
  type: test_failure
  ref: trace://run-8842/check-9
evidence:
  - type: deterministic_test
    ref: test://payroll/net-pay-balance
judge:
  policy: reflection-rubric-v4
proposal:
  target: src/payroll/net_pay.py
  allowed_change: diagnosed_files_only
authority:
  mode: auto_in_sandbox      # suggest | auto_in_sandbox | approval_required
budget:
  max_iterations: 3
  max_latency_ms: 90000
verification:
  suite: payroll-regression-v12
  rollback_on_regression: true
retention:
  disposition: candidate_lesson

The full loop is Trigger → Evidence pack → Diagnose → Change proposal → Policy gate → Apply → Verify → Record. A model may diagnose and propose. Tests, rules, people, and business outcomes decide whether it may continue.

How the four specifications divide the work

Pattern Primary change target Feedback Main boundary
F1 Generator-Critic Current artifact Rules, references, expert rubrics, or an independent model The critic must be evaluated; a score cannot replace evidence
F2 Skill Package Reusable workflow and capability asset Cross-run outcomes, failures, and coexistence tests A Skill that passes alone may still degrade a larger Skill set
F3 Experience Replay Context supplied to a later task Historical trajectories, outcomes, and delayed feedback Preserve source, version scope, and uncertainty to avoid negative transfer
F4 Self-Heal Loop Rollback-safe state, configuration, or code Deterministic failure and re-run evidence Missing knowledge, changed business logic, and irreversible actions require escalation or an offline release

Online/offline, hard evidence/soft judgement, local/global scope, and feedback delay are selection parameters. They do not alter the cognitive-function × execution-topology framework.

A four-layer production structure

Mo Zhou proposed a useful four-layer implementation:

  1. Hard validation: Compilation, tests, schemas, business invariants, and system receipts take priority.
  2. Reflection diagnosis: Models inspect evidence and identify possible defects in outputs, trajectories, or assets.
  3. Memory and assets: Candidate lessons, Skills, rules, and failure records cross task boundaries only after admission.
  4. Scheduling and control: The harness decides whether to reflect, which reviewer to use, how many rounds to allow, and when to degrade or hand off.

Automatic triggering, termination, observability, and bounded cost apply to every production loop. Reuse needs a narrower rule: a one-run correction may be deliberately discarded; anything that will affect later runs requires admission, versioning, and retirement.

Three problems that appear at scale

Local patches can damage global efficiency. Wei Wang described an environment-building agent that accumulated local fixes over time. Stability improved while the overall path became heavier. Offline reflection must detect duplicate, conflicting, and expired patches instead of appending every bad case to a prompt.

Skills must be evaluated in combination. Trigger accuracy and task success show that one Skill works alone. A growing library also needs false-trigger, missed-trigger, overlap, conflict, loading-cost, and coexistence tests.

Attribution comes before healing. Qianchun Lu separated runtime, execution-process, business-result, and experience failures. Invalid parameters, network instability, and missing steps may have direct remedies. Missing domain knowledge, changed business rules, and conflicting team goals cannot be filled in safely by the original agent.

Engineering progress in 2026

By August 2026, major platforms expose the evaluation components that reflection depends on:

  • LangSmith Evaluation separates pre-release offline evaluation from online evaluation over production traces and feeds failures back into datasets. AgentEvals evaluates tool-use trajectories directly.
  • Anthropic's January 2026 guide, Demystifying evals for AI agents, recommends combining code, model, and human graders. Its enterprise Skills guidance adds triggering, isolation, coexistence, instruction-following, and output-quality checks.
  • OpenAI's AgentKit brings trace grading, datasets, and graders into the agent optimisation workflow.

These systems first address seeing and judging. Reflection begins when a judgement becomes a controlled change and regression evidence confirms the result.

Directions without new catalog numbers

Meta-reflection reviews the critic and its rubric. It currently fits as a high-risk F1 configuration rather than a separate pattern.

Deliberative reflection uses reviewers with different assumptions in a structured debate. It overlaps F1 Generator-Critic and C3 Adversarial Review, and still lacks stable stopping and cost evidence.

Prospective reflection checks plans, assumptions, and risk before action. Its responsibilities already appear across F1, A4 Guardrail Sandwich, and G1 Approval Gate. More independent practice is needed before cataloguing it separately.

Questions for further industry evidence

  • How can a large Skill estate separate agent quality, Skill quality, and combination effects?
  • When dimensions in a rubric move in opposite directions, who owns the trade-off?
  • Which reflections may change a harness automatically, and which may only open a change proposal?
  • When outcomes arrive days or weeks later, how should the system preserve the original trajectory, version, and accountability chain?
  • How can teams detect, consolidate, remove, or roll back conflicting patches accumulated by online loops?

Workshop record

This overview incorporates the first ADPS Reflection Module Workshop held on 12 August 2026. Hosts: Haili Zhang and Jia Huang. Core workshop guests: Dong Zhang, Mo Zhou, Wei Wang, Qianchun Lu, Han Zhao, Pylon Peng, and Jiaqi Li. The attributed points on this page come from the six speakers recorded in the transcript: Dong Zhang, Mo Zhou, Wei Wang, Qianchun Lu, Pylon Peng, and Jiaqi Li.

Read the full workshop record · White Paper contributors

Suggested citation: ADPS, Reflection Module: Make Feedback Change the System, Agent Design Pattern White Paper v0.3, 2026-08-13. Catalog · CC BY 4.0