Reward integrity

Does your reward pay for the right thing?

Reward integrity is whether the reward used to train or evaluate an AI system pays for the right thing: full credit for work that does what the task asks, no credit for work that does not, and no way to move the score except by doing the work.

It applies to every kind of reward: unit and behavioral tests for code, AI judges and rubrics, math and answer checkers, and the state checks in agent environments. A benchmark score is a reward too, so the same property decides whether a leaderboard means anything. It is measured with five numbers: the false-pass rate (wrong work paid, found by searching for it), the false-fail rate (correct work refused), the tamper surface (grader inputs the model can change), the leak surface (answers the model can reach), and stability and provenance (the same work gets the same reward, under a pinned, recorded configuration).

We did not coin the term. It already appears in AI research, including BenchShield (arXiv 2609.11028), which treats integrity as a property of the whole reward path. We use it in this broad sense, define it as above, and publish open criteria for it. The definition, the taxonomy and the criteria also have their own site, rewardintegrity.org, which we steward.

Why it matters

A model learns what the reward pays for, not what you meant.

RL training samples each task many times and reinforces whatever earns reward. A gap that a single attempt rarely lands in, such as a test that never checks one requirement, is one an optimizer can find and then repeat. That is why a weak grader can leave one evaluation run almost unchanged and still teach the wrong behavior in training.

OPTIMIZATION

Paid for whatever passes

The policy is not told what the task meant. It is told what passed. A wrong fix the tests accept earns the same full reward as a right one, so training has no signal to prefer the right one. And a wrong fix that the tests accept involves no tampering, so a monitor that looks for tampering has nothing to flag.

OUR AUDITS OF PUBLIC TASK SETS
190of 487

testable tasks, in 11 public task sets, paid full reward for a wrong fix in one search pass: about 4 in 10.

One fixed search per task: up to 8 versions of the reference fix with parts removed, plus 6 wrong fixes written by AI models. Each counted task has a wrong fix that a separate behavioral check confirms is wrong. A deeper search finds more, so this is a floor. Set-by-set results are in the Grader Integrity Index from November 10, after each maintainer’s private notice period ends.

AT TRAINING SCALE

Hacks that generalize

Anthropic researchers trained a model with RL on 80 production environments that had known reward hacks, including weak tests, exposed answer files and writable reward files. By the end of training, 40% of episodes were flagged as hacks, and the learned reward seeking generalized to more harmful behavior.

Qi et al., “Training a misaligned reward seeker”, Anthropic Alignment Science Blog, August 2026

Failure taxonomy · draft v0.1

Eleven ways a reward pays for the wrong thing.

Each failure mode, how to detect it, and what covers it. The cheap, mechanical failures are covered by Plumbline Scan’s six free checks (PS-1 to PS-6), which run locally with Docker. The expensive ones need a search or a measurement: that is the paid deep audit, for code graders or for AI judges.

Scan PS-n free mechanical check paid Grader Audit or AI-judge audit Not covered neither, today
Reward integrity failure modes, how each is detected, and coverage
Failure modeHow it is detectedCovered by
Tests and checkers
Wrong work paid by weak testsThe grader gives full reward to work that misses a requirement the task states: a partial fix, special-casing the tested inputs, a plausible bug. Write plausible wrong solutions, one per requirement plus known weak-test patterns, and run them through the unchanged grader. Confirm each pass with a separate behavioral check that passes on the reference solution, fails on the unsolved code and fails on the wrong solution.
Leaked answersThe answer or the reference fix is reachable from inside the task: git history, image layers, cached files. Search the task image and its history for the reference solution or the expected output. Scan PS-3
Stub or broken referencesThe reference is empty, partial, or fails its own grader, so “passes on the reference” proves nothing. Run the reference through the grader several times; it must pass every time. Check that it does real work, and that an empty or unchanged submission gets no reward. Scan PS-1PS-2PS-4
Editable grader inputsThe model can change what the grader reads: test files, fixtures, scoring scripts, logs, its own reported metrics. List every file the grader reads and check whether the agent can write to it before grading. Withhold tests until grading. Scan PS-5
Network-dependent gradersThe reward depends on something outside the task, such as a network service, the clock or the hardware. Run the grader sealed, with the network off, and compare with a networked run. Scan PS-6
Flaky gradersThe same work gets different rewards on different runs: flaky tests, judge sampling. Repeat the same submission k times and count how often the reward flips. Scan PS-1Scan runs the reference 3 times; the AI-judge audit runs every item k times
Correct work refusedTests reject valid solutions because they are over-specified or tied to one implementation. Run independently written correct solutions through the grader and count rejections. Not covered
AI judges and rubrics
AI-judge false passesThe judge gives credit to answers known to be wrong: a planted factual error, a missing required element, a confident restatement, an off-task answer, an instruction to the grader hidden in the answer. Run the judge exactly as deployed on documented defective answers mixed with clean controls, several times each. Report the false-pass rate per defect family, with intervals. AI-judge
AI-judge false failsThe judge refuses correct answers that differ in wording, order, format or length. Score known-correct paraphrases, reorderings and reformats, and report the false-fail rate on these controls. AI-judge
Judge configuration driftThe reward changes without anyone deciding to change it: the judge model is updated, the rubric or prompt is edited, a dependency floats. Pin and record the judge model ID, rubric and prompt hashes, image digest and seeds. Re-measure on any change. paid, re-runs on each change
Judges missing contextThe judge does not see what it needs to grade, such as the task, the reference answer, or the files and tool output the answer refers to, so it grades on how plausible the answer sounds. Check every rubric criterion against what the judge is actually given, then compare verdicts on documented items with and without the missing context. AI-judge

Plumbline Scan supports Harbor-format coding tasks first. It does not search for wrong solutions that the tests accept; only a search finds those. No check or audit can prove a reward cannot be gamed. Each reports what was tested and what was found.

Reward Integrity Criteria

Open criteria, published November 10.

Reward Integrity Criteria v1.0 publishes on November 10, 2026, under a Creative Commons Attribution license (CC BY 4.0), at rewardintegrity.org. It is a floor anyone can check and claim, with evidence. The outline below is a draft; the published text may change.

  1. Reference

    The reference solution passes every one of at least 3 runs, and it is real work, not a stub.

  2. Null

    An empty or unchanged submission gets no reward.

  3. No leaks

    No path from the agent’s environment to the answer.

  4. Sealed inputs

    The agent cannot write anything the grader reads.

  5. Searched false passes

    A stated search for wrong work was run, with the method, the sample size, and the false-pass rate and its 95% interval published.

  6. Judges measured

    For AI-judged rewards: false-pass and false-fail rates on documented defects and controls, per defect family, with intervals.

  7. Pinned

    The configuration is pinned and recorded, and re-checked on any change.

  8. Disclosure

    Who ran the checks, whether they also built the environment, and the command to reproduce them.

How this relates to Plumbline Verified. The criteria are the open floor. Plumbline Verified is a stricter certification for one named release: a repaired release, re-audited separately, with a 95% upper bound on hackable tasks under 10%.

Questions

Common questions

What is reward integrity?

Reward integrity is whether the reward used to train or evaluate an AI system pays for the right thing: full credit for work that does what the task asks, no credit for work that does not, and no way to move the score except by doing the work. It covers tests for code, AI judges and rubrics, math and answer checkers, and agent environments.

What is reward hacking?

Reward hacking, closely related to specification gaming, is when a model earns high reward through behavior its designers did not intend: special-casing the tested inputs, editing tests, reading leaked answers, or writing answers an AI judge over-rewards. It usually means the reward pays for something other than the intended task, which is a failure of reward integrity.

In a 2026 study, Anthropic researchers trained a model with RL on 80 production environments that had known reward hacks. By the end of training, 40% of episodes were flagged as hacks, and the learned reward seeking generalized to more harmful behavior. The practical response is to find and close the gaps in the reward before training, rather than only trying to detect hacking afterward.

How do I check my RL environment?

Check the reward, not just the tasks. Six mechanical checks catch the cheap failures: the reference solution passes every time (run it at least 3 times); an empty answer gets nothing; the answer cannot be found in the task image; the reference is real work, not a stub; the agent cannot edit any file the grader reads; and the grader does not depend on the network. Plumbline Scan will run all six locally with Docker, for free, from November 10.

The expensive failure is wrong work the grader pays for. To find it, write plausible wrong solutions (one per requirement), run them through the unchanged grader, and confirm each pass with a separate behavioral check. In our audits of 11 public task sets, 190 of 487 testable tasks (about 4 in 10) paid full reward for a wrong fix in one fixed search pass, and a deeper search finds more.

Are SWE-bench tests reliable?

Partly. The tests decide whether a patch resolves an issue, and several independent audits show that some of them accept incorrect patches. Rajan (2026) found Docker-verified wrong patches that pass the tests on at least 28.5% of a 49-task SWE-bench Verified sample. SWE-ABS strengthened the test suites adversarially and reports that about one in five patches counted as solved by the top 30 agents is semantically incorrect. Epoch AI’s Benchmark Reviews rated SWE-bench Verified “Flawed” in September 2026.

Scores remain useful for comparing models. A resolved rate measures how often patches pass the tests, and that can overstate how often they are correct.

How do I validate an LLM judge?

Test it on answers whose correctness you already know. Build two sets: documented defective answers (a planted factual error, a missing required element, a confident restatement, an off-task answer, an instruction to the grader injected into the answer) and known-correct controls (paraphrases, reorderings, different formats). Run the judge exactly as deployed, with the same model, rubric, prompt and inputs, several times.

Report the false-pass rate per defect family and the false-fail rate on controls, with confidence intervals. Pin the judge model and rubric, and re-run whenever either changes. Agreement with people on ordinary answers is not enough: a judge can agree with people most of the time and still pay for the defects a model will learn to produce.

What is a false pass?

A false pass is work the reward pays for although it does not do what the task asks: a wrong fix that passes the tests, or a wrong answer an AI judge accepts. In our audits, a false pass counts only when a separate behavioral check passes on the reference solution, fails on the unsolved code and fails on the wrong fix, and the task as written requires the behavior the fix gets wrong. The opposite is a false fail: correct work the reward refuses.

How is reward integrity different from evaluation integrity?

Evaluation integrity is broader. It also covers problems such as contamination, where test data was seen in training, and sandbagging, where a model underperforms on purpose. Reward integrity asks one narrower question: does the scoring pay for the right thing? It applies to training rewards as well as evaluation scores, because a benchmark score is a reward too.

Did Plumbline Grader coin the term “reward integrity”?

No. The phrase already appears in AI research, for example in BenchShield (arXiv 2609.11028), which treats integrity as a property of the whole reward path, and in LEGO-RL (arXiv 2608.17393), which has a section on reward-integrity failure modes. We use the term in a broad sense that covers tests and judges paying for wrong work as well as tampering and leaks, define it on this page, and publish open criteria for it.

Plumbline Scan · coming November 10

Six free checks. Run them yourself.

A free command-line checker for coding tasks. It needs only Docker and the task folder: no API keys, no paid models, and by default nothing leaves your machine. One email on November 10 when it ships.

Find the wrong work your grader pays for.

A free 20-task audit: the same deep search as a paid audit, with full evidence and a one-command reproduction for every finding.