Why trace scoring matters
A final answer hides important failure modes. An agent can reach the correct answer while using an unauthorized tool or inventing a citation. The evaluator therefore checks successful tool names and arguments, source support and attempted actions alongside exact answer correctness.
Gold labels stay inside the evaluator
Task fixtures separate public input from expected answers, expected calls and relevance labels. External predictions must cover every task exactly once. Dataset hashes and task populations must match before a regression comparison is accepted.
Recovery needs a denominator
The ten transient-failure fixtures inject one timeout. The improved scripted control retries once. Recovery is reported over those labeled fixtures; permission errors are not counted as transient recovery opportunities.
What the control experiment establishes
The baseline succeeds on 50/100 fixture variants; the improved control succeeds on 100/100 with no success regressions. Improvements come from retry, citation handling, abstention and blocking disallowed actions. Both controls operate on explicit structured tool requests. The result checks the harness, not general reasoning or real-model capability.
What remains unobserved
No language model was called in this benchmark. Token and cost fields are null. Groundedness, hallucination, instruction conflict and prompt-injection resistance need independent labeled evidence and are not inferred from exact answer matches.
Reproducibility
Source, fixtures, tests, run commands and the JSON/HTML evidence are public in Agent Eval Lab. The five-project adapter records commit provenance and whether working trees were modified. Future real-agent comparisons should add independently captured traces, environment details and reviewer-authenticated approvals.