Topics/Agent Evals and Testing: From Output Scores to System Acceptance

ADPS Topic Research

Agent Evals and Testing: From Output Scores to System Acceptance

Deterministic tests, trajectory evaluation, external acceptance, and production feedback combined into agent release evidence.

Issued
2026-08-14
Document type
Cross-cutting topic
Status
Topic research draft
Basis
ADPS workshops, enterprise cases, and public evaluation methods
License
CC BY 4.0

Topics synthesize engineering questions that cross several modules. Pattern definitions, attributed cases, and workshop records remain authoritative on their own pages.

Traditional software tests remain essential for agents. Data structures, permissions, tool contracts, idempotency, and business ledgers still need deterministic assertions. The new difficulty comes from probabilistic outputs, multiple valid paths, external tools, and delayed outcomes. A run may follow a different trajectory and still succeed, or produce fluent text after writing the wrong external state.

Agent evals address this uncertain layer. They work alongside unit tests, integration tests, sandbox acceptance, production monitoring, and human review.

Four terms with different jobs

Term Primary job Typical artefact
Test Assert a deterministic contract Pass/fail, diff, and fault location
Eval Measure capability and regression across uncertain behaviour and repeated trials Sample evidence, scores, pass rates, and failure clusters
Monitoring Observe behaviour on real production traffic Runtime events, metrics, alerts, and sampled runs
Acceptance Let users or business rules decide whether delivery is valid External receipt, visible result, approval, or sign-off

The boundary between a test and an eval can overlap. A deterministic grader can appear in either suite. What matters is the evidence source, its owner, and the release decision it can support.

Define the system under test

“Test the agent” is too broad. The system under test may be:

The dataset, execution environment, graders, and release gate change with the system under test. A tool adapter focuses on schema, authority, and idempotency. A long-running agent also requires goal retention, checkpoints, budgets, and recovery tests.

Six evaluation surfaces

Surface Question Example evidence
Final outcome Did the external world reach the target state? Database fact, business receipt, real request result
Trajectory Were tools, order, and parameters acceptable? Trace, ActionEvent, task-graph node
Single-step decision Was routing, retrieval, compaction, or approval correct? Decision label, candidate set, rationale checked against facts
Safety and authority Did the agent cross a boundary, and did refusal or escalation work? Negative case, sandbox event, approval and rejection record
Recovery and long-running progress Can interruption, retry, and hand-off return to the right task? Checkpoint, idempotency key, goal version, progress ledger
Resource and latency Does quality fit the latency and cost envelope? Trial distribution, tokens, latency, tool-call count

Final outcome is the primary verdict. Trajectory evaluation exposes latent risk and locates failures. Where multiple correct paths exist, strict matching to one golden trajectory rejects valid solutions. Preserve route freedom when outcomes, authority, and mandatory constraints are satisfied.

An agent test pyramid

               Production shadow / canary / business outcomes
                     Human and calibrated model grading
                  Trajectory replay, sandbox, fault injection
               Workflow, external acceptance, system integration
            Tool contracts, permissions, state machines, idempotency
         Schema, parsing, functions, rules, deterministic component tests

Lower layers are fast, stable, and easy to diagnose. Upper layers approach real value but run more slowly, cost more, and carry more variance. A usable release gate combines evidence across layers instead of treating one aggregate score as complete truth.

Grader priority

  1. External facts and deterministic assertions.
  2. Code, rules, schemas, and state-machine checks.
  3. Business acceptors or independent simulations.
  4. Calibrated model graders.
  5. Experts or real users for ambiguous criteria.

Model graders are useful for semantic relevance, style, completeness, and complex trajectories that resist hard rules. They need calibration against human judgements and periodic bias review. Using one uncalibrated model to generate, grade, and rewrite the criteria weakens independence.

Eval Contract

eval_id: gis-publish-regression-v7
system_under_test: gis-agent@v5
task: publish_and_verify_layer
input_fixture: fixtures/s57-small-03
environment: geoserver-sandbox@2.25
allowed_tools: [inspect, publish, verify_get_map]
authority: no_production_write
expected_outcomes:
  - layer_is_queryable
  - get_map_contains_visible_content
forbidden_outcomes:
  - modify_unrelated_workspace
graders:
  - schema_contract
  - external_get_map_probe
trials: 5
release_gate:
  capability: all_required_cases_pass
  regression: no_blocking_case_regresses
evidence: artifacts/evals/gis-v7/
owner: geo-platform

This contract joins the task, environment, authority, positive and negative outcomes, graders, trials, and release gate. Stochastic tasks require repeated trials with each trajectory retained; an average alone can hide a rare but severe failure.

Dataset lifecycle

  1. Derive an initial capability set from specifications and business acceptance criteria.
  2. Compress real failures, bad cases, and incidents into replayable regression samples.
  3. Add boundary, authority, adversarial, empty-input, and unknown-input cases; keep both positive and negative examples.
  4. Version data, environments, graders, and expected outcomes separately.
  5. Read failed trajectories, retire stale samples, and split graders whose criteria have become too broad.

Capability evals ask whether the system can complete its target tasks now. Regression evals protect established behaviour from new changes. Their selection logic differs, and releases need both.

Three engineering cases

External acceptance replaces internal “success”

The Xuanxu Technology GIS publishing agent encountered a typical false success. A publish API could return success while a real GetMap request still returned LayerNotDefined; some error responses also carried HTTP 200. The final acceptor issues real GetMap or GetTile requests from the consumer side and checks protocol results, content type, and visible map content.

The task graph carries execution and acceptance

The Dongfang Yiteng execution agent uses a task DAG and node state machines for strict dependencies. Node completion passes through a validator, while business-ID provenance, tool receipts, and approval events enter one timeline. The same evidence judges the outcome and locates skipped, missing, or duplicate steps.

One capability across three environments

The Reflection workshop proposed a practical split. A baseline environment runs stable samples, a daily environment carries real work, and a test environment admits candidate skills and prompts. Candidate and current capabilities coexist until evaluation supports promotion, so one local lesson does not immediately change global behaviour.

Release evidence flow

Production failure or new specification
        ↓
Reproduce and freeze it as a sample
        ↓
Add it to regression; define grader and environment
        ↓
Change a candidate component
        ↓
Run capability + regression + negative cases
        ↓
Read failed trajectories and review graders
        ↓
Release, shadow, canary, or roll back

New failures should enter regression, but not every production sample deserves permanent retention. Teams need deduplication, retirement, and a way to distinguish model variance, environment faults, grader defects, and actual capability gaps.

Common failure modes

Questions for further discussion

A useful session should bring four artefacts: one real bad case, one Eval Contract, one failed trajectory, and the release record that the sample influenced. This makes systems under test, grading rules, and authority boundaries comparable across teams.

Sources

Suggested citation: ADPS, Agent Evals and Testing: From Output Scores to System Acceptance, ADPS Topic Research, 2026-08-14.

Topic index · CC BY 4.0