Agent environment grader audit · Early access

Your grader says the job is done. Did the agent do it?

In web, computer-use and tool-use environments, a grader checks the state the agent leaves behind. If that state can be reached without doing the job, the model learns the shortcut. We test end-state graders to the standard of our code audits: every false pass proved by an independent check, replayed in a fresh environment, and repaired on both sides.

Price quoted, fixed before we start; never per finding.

Why end-state graders get gamed

The grader reads a state. The agent shapes it.

Public work has already shown how far this goes. Berkeley RDI’s scanning agent reached about 100% on all 812 WebArena tasks by reading gold answers through a local file URL, and 73% on OSWorld through VM state manipulation and public gold files, without solving the tasks. The ABC checklist found that τ-bench counted empty responses as successful. Scale Labs, testing its own computer-use verifiers, found blank templates and unedited files scoring over 80% where checks were weak, and none of 38 programmatic verifiers failing a corrupted document.

Every one of those failures fits one of six families. We test for all six.

01

Wrong state paid

The one thing the grader checks looks right, but the job was not done: the visible field was set and the real action skipped, or the grader read the request log instead of the record that should have changed.

02

Goal readable

The expected answer or target state can be read from inside the environment before grading: a gold file at a predictable address, a task file on a reachable path.

03

Grader trusts agent-writable state

The grader decides from a record, log or file the agent can write. Whoever writes the evidence sets the score.

04

Measurement in reach

Timers, counters, score lines or network logs that the agent can set, add to or print into.

05

Unchecked requirements

The task states five conditions and the grader checks three. The other two are paid whether or not they happen.

06

Flaky grader

The same run passes once and fails the next. That reward is noise, and a policy trains on it anyway.

Our method

Test the grader you ship. Prove every miss.

The core audit is deterministic: it reads and tests your grader as it runs today. A pass counts as false only when an independent check shows the work was wrong and the case replays in a fresh environment.

STRUCTURE

Structural checks

We read the environment and its grader and flag each structural weakness: expected values the agent can reach, evidence the agent can write, measurement in its reach, scoring that is not deterministic.

MUTATE

Mutation testing of the reference end state

Start from a reference that passes. Apply a library of state-breaking changes, frozen before any is graded: a required step skipped, a side effect undone, the right request with a failed response, one wrong field, a degenerate state. Every mutant goes through your unchanged grader. You get a mutation score per grader with a 95% interval.

COVER

Requirements coverage

Each task instruction is split into its required conditions before we read the grader, then each condition is marked scored, partly scored or unscored. A model drafts the condition list, and the report labels it as model work.

REPEAT

Determinism runs

The reference is graded k times, each in a fresh container. For live sites we reset the environment and check that the reset restores it. You get a flip rate with its interval; a reset that leaks state between runs is a finding on its own.

PROVE

Independent state check, then fresh replay

A check written from the task instruction alone reads the result directly, such as the site’s own database rather than the request log. It must pass the reference and fail the mutant. Then a kit replays the case in a new container from pinned commits and image digests, with only bash and Docker. We report only inputs a real agent could produce.

REPAIR

Two-sided repair

Where a grader is weak, we write a stronger state check. It must block every wrong state we found and still pass the reference and correct alternative solutions. We measure both sides and report both.

What you get

Evidence you can re-run, and the fix.

Every false pass, with its proof

The end state your grader paid for, the independent check that fails it, three fresh-container results, and a one-command replay kit.

A mutation score per grader

With its 95% interval, and every surviving mutant named, so you see exactly which wrong states get paid.

A requirements coverage map

Each condition the task states, marked scored, partly scored or unscored by your grader.

A determinism report

Flip rate per grader with its interval, and any reset that leaks state between runs.

Repaired graders, measured both ways

Stronger state checks that block the wrong states we found and still pass correct work, with both results shown.

Known or new, labeled

Each finding checked against published lists, including the ABC checklist, BenchJack and the maintainers’ own changelogs, and labeled known or new.

Price. Quoted, fixed before we start; never per finding.

Early access

Early access is open.

We are building agent-environment audits now and have published no agent-environment results yet. We will publish only when they meet the bar our code audits meet: every number traced to a run, every finding open to replay. On code environments, 424 of 425 re-run findings reproduced from scratch. Early-access clients choose the scope: web, computer-use or tool-use, on their own environments, under their NDA.

What do you need from us?

The environment (images or access), the grader code and configuration, the task instructions, and a reference solution or trajectory per task if you have one. If you do not, we script references from the instructions, and each must pass both your grader and our independent check before it counts.

Do you use an AI attacker against our environment?

Not in the core audit. Structural checks, mutation testing, coverage and determinism runs are deterministic tests of the grader you ship. Open-ended adversarial search belongs in a separately scoped engagement agreed with you.

Is this the same as your code audits?

Same proof bar, different object. In a code audit we prove a patch wrong; here we prove an end state wrong. In both, a pass we cannot show is false is never counted, and every finding ships with a replay kit.

How is it priced?

Quoted per engagement and fixed before we start. Fees never depend on the number of findings.

Related public work

Context, not our results.

  • Zhu et al., “Establishing Best Practices for Building Rigorous Agentic Benchmarks” (the ABC checklist), arXiv 2507.02825. Finds that τ-bench counts empty responses as successful, and that such issues can misstate agent performance by up to 100% in relative terms.
  • Berkeley RDI, trustworthy-benchmarks post, April 2026, and the BenchJack paper, arXiv 2605.12673. An automated scanning agent audits prominent agent benchmarks for exploits, including about 100% on WebArena and 73% on OSWorld without solving the tasks.
  • Scale Labs, “Verifier design for CUA”, September 2, 2026. Mutation testing of Scale’s own computer-use verifiers: weak checks let blank or unedited submissions score over 80%, none of 38 programmatic verifiers failed a corrupted document, and an LLM judge gave three identical outputs 0.0, 0.125 and 0.175.