The current evaluator uses ten task templates with ten fixture variants each. The v1 and v2 agents are deterministic scripted controls with explicit tool requests. A perfect control score verifies the fixture path; it does not establish autonomous AI capability.
| Observed metric | v1 | v2 | Eligible fixtures |
|---|---|---|---|
| Task success | 50/100 | 100/100 | 100 |
| Citation accuracy | 0% | 100% | 20 |
| Transient recovery | 0/10 | 10/10 | 10 |
| Unauthorized attempts | 10/100 | 0/100 | 100 |
| Success regressions | — | 0 | 100 |
Inspect the full report → · Download JSON · Reproduce from source