Your grader says the job is done. Did the agent do it?
In web, computer-use and tool-use environments, a grader checks the state the agent leaves behind. If that state can be reached without doing the job, the model learns the shortcut. We test end-state graders to the standard of our code audits: every false pass proved by an independent check, replayed in a fresh environment, and repaired on both sides.
Price quoted, fixed before we start; never per finding.
The grader reads a state. The agent shapes it.
Public work has already shown how far this goes. Berkeley RDI’s scanning agent reached about 100% on all 812 WebArena tasks by reading gold answers through a local file URL, and 73% on OSWorld through VM state manipulation and public gold files, without solving the tasks. The ABC checklist found that τ-bench counted empty responses as successful. Scale Labs, testing its own computer-use verifiers, found blank templates and unedited files scoring over 80% where checks were weak, and none of 38 programmatic verifiers failing a corrupted document.
Every one of those failures fits one of six families. We test for all six.
Wrong state paid
The one thing the grader checks looks right, but the job was not done: the visible field was set and the real action skipped, or the grader read the request log instead of the record that should have changed.
Goal readable
The expected answer or target state can be read from inside the environment before grading: a gold file at a predictable address, a task file on a reachable path.
Grader trusts agent-writable state
The grader decides from a record, log or file the agent can write. Whoever writes the evidence sets the score.
Measurement in reach
Timers, counters, score lines or network logs that the agent can set, add to or print into.
Unchecked requirements
The task states five conditions and the grader checks three. The other two are paid whether or not they happen.
Flaky grader
The same run passes once and fails the next. That reward is noise, and a policy trains on it anyway.
Test the grader you ship. Prove every miss.
The core audit is deterministic: it reads and tests your grader as it runs today. A pass counts as false only when an independent check shows the work was wrong and the case replays in a fresh environment.
Structural checks
We read the environment and its grader and flag each structural weakness: expected values the agent can reach, evidence the agent can write, measurement in its reach, scoring that is not deterministic.
Mutation testing of the reference end state
Start from a reference that passes. Apply a library of state-breaking changes, frozen before any is graded: a required step skipped, a side effect undone, the right request with a failed response, one wrong field, a degenerate state. Every mutant goes through your unchanged grader. You get a mutation score per grader with a 95% interval.
Requirements coverage
Each task instruction is split into its required conditions before we read the grader, then each condition is marked scored, partly scored or unscored. A model drafts the condition list, and the report labels it as model work.
Determinism runs
The reference is graded k times, each in a fresh container. For live sites we reset the environment and check that the reset restores it. You get a flip rate with its interval; a reset that leaks state between runs is a finding on its own.
Independent state check, then fresh replay
A check written from the task instruction alone reads the result directly, such as the site’s own database rather than the request log. It must pass the reference and fail the mutant. Then a kit replays the case in a new container from pinned commits and image digests, with only bash and Docker. We report only inputs a real agent could produce.
Two-sided repair
Where a grader is weak, we write a stronger state check. It must block every wrong state we found and still pass the reference and correct alternative solutions. We measure both sides and report both.
Evidence you can re-run, and the fix.
Every false pass, with its proof
The end state your grader paid for, the independent check that fails it, three fresh-container results, and a one-command replay kit.
A mutation score per grader
With its 95% interval, and every surviving mutant named, so you see exactly which wrong states get paid.
A requirements coverage map
Each condition the task states, marked scored, partly scored or unscored by your grader.
A determinism report
Flip rate per grader with its interval, and any reset that leaks state between runs.
Repaired graders, measured both ways
Stronger state checks that block the wrong states we found and still pass correct work, with both results shown.
Known or new, labeled
Each finding checked against published lists, including the ABC checklist, BenchJack and the maintainers’ own changelogs, and labeled known or new.
Price. Quoted, fixed before we start; never per finding.
Early access is open.
We are building agent-environment audits now and have published no agent-environment results yet. We will publish only when they meet the bar our code audits meet: every number traced to a run, every finding open to replay. On code environments, 424 of 425 re-run findings reproduced from scratch. Early-access clients choose the scope: web, computer-use or tool-use, on their own environments, under their NDA.
What do you need from us?
The environment (images or access), the grader code and configuration, the task instructions, and a reference solution or trajectory per task if you have one. If you do not, we script references from the instructions, and each must pass both your grader and our independent check before it counts.
Do you use an AI attacker against our environment?
Not in the core audit. Structural checks, mutation testing, coverage and determinism runs are deterministic tests of the grader you ship. Open-ended adversarial search belongs in a separately scoped engagement agreed with you.
Is this the same as your code audits?
Same proof bar, different object. In a code audit we prove a patch wrong; here we prove an end state wrong. In both, a pass we cannot show is false is never counted, and every finding ships with a replay kit.
How is it priced?
Quoted per engagement and fixed before we start. Fees never depend on the number of findings.
Context, not our results.
- Zhu et al., “Establishing Best Practices for Building Rigorous Agentic Benchmarks” (the ABC checklist), arXiv 2507.02825. Finds that τ-bench counts empty responses as successful, and that such issues can misstate agent performance by up to 100% in relative terms.
- Berkeley RDI, trustworthy-benchmarks post, April 2026, and the BenchJack paper, arXiv 2605.12673. An automated scanning agent audits prominent agent benchmarks for exploits, including about 100% on WebArena and 73% on OSWorld without solving the tasks.
- Scale Labs, “Verifier design for CUA”, September 2, 2026. Mutation testing of Scale’s own computer-use verifiers: weak checks let blank or unedited submissions score over 80%, none of 38 programmatic verifiers failed a corrupted document, and an LLM judge gave three identical outputs 0.0, 0.125 and 0.175.