Cost of a false pass

What does a grader that pays for wrong work cost you?

Enter your own task count and what each task costs you. We apply the false-pass rate we measured on public task sets: 190 of 487 testable tasks paid full reward for a wrong fix in one fixed search pass, across 11 sets.

Calculator

Your inputs

Epoch AI reports that most RL environment tasks cost $200 to $2,000 each (FAQ on RL environments, Jan 2026). $500 is only an example; use your own figure.

Default 39.0% is our measured 190 of 487. Pooled range 34.8–43.4% (a Wilson 95% interval that treats the 11 sets as one sample). Set your own rate if you have measured one.

Your estimate

$195,000

of task spend sits behind a grader that pays full reward for a wrong fix: about 390 of 1,000 tasks.

pays for wrong work our measured range

At our measured range (34.8–43.4%)
$174,000 – $217,000
Grader Audit of the same set, at list price
$40,000
Spend at risk for every $1 of audit
$4.88

Audit at list price: $40 per task, $1,500 minimum per batch; volume pricing above 1,000 tasks lowers it. See pricing.

A free 20-task audit replaces our public-set rate with yours.

Why it costs more than the task

A false pass teaches the shortcut.

When a grader pays full reward for a wrong fix, the money spent on that task buys a lesson you did not want: the model learns that the shortcut scores. Training repeats the lesson on every episode that finds it. The calculator counts only the task spend. It leaves out the compute spent training on a bad reward and the cost of finding the habit later.

The fix is cheap by comparison. An audit names the tasks, proves each wrong fix with a separate behavioral check and a fresh-container replay, and ships verified replacement tests that block the wrong fixes and still accept correct ones.

Where the rate comes from

Measured, with its scope.

  • 01190 of 487 testable tasks across 11 public task sets paid full reward for a wrong fix in one fixed search pass (39.0%). A deeper search finds more, so for those sets this is a floor.
  • 02Every counted task is proved. A separate behavioral check passes the reference fix, fails the unfixed code and fails the wrong fix, and the result replays in a fresh container.
  • 03Scope is voted, not assumed. AI models from three companies (Claude, GPT, Gemini) vote on whether the task as written requires the behavior the wrong fix breaks.
  • 04Your set may differ. Public sets are not your private tasks. The rate here is a starting estimate; a sample audit measures yours.