Your AI judge gives full marks. Was the answer right?
When a rubric or an LLM judge supplies the reward, a policy can learn to satisfy the judge instead of the task. We measure how often your judge passes answers we know are wrong, show you each one, and hand you the fixes.
Early access: 40% off for the first 3 clients, in exchange for a case study.
A judge can only check what the rubric names.
In May 2026, researchers at Scale Labs trained policies against rubric-based rewards and scored the results with three frontier judges from different model families. Weak verifiers produced large reward gains that did not hold up under those judges, and the exploitation grew as training went on. The common failures they describe: criteria with several parts credited when only some were met, implied content treated as if it were stated, and loose matching on topic.
Stronger judges reduced the problem but did not remove it. Their sharpest finding is about the rubric itself:
“…stronger verification does not prevent reward hacking when the rubric leaves important failure modes unspecified.”
So a better judge model is not enough on its own. You need to know which wrong answers your rubric and judge let through, how often, and what closes the gap.
Plant known defects. Count what gets through.
The same rule as our code audits: nothing counts as a false pass unless we can show, item by item, that the answer was wrong.
- 01 · PLANT
Defects with certificates
We take answers your judge should pass and plant one known defect in each: a wrong fact, a missed requirement, an unsupported claim. Every item carries a certificate saying what was changed, where, and why it breaks the task. Each planted defect is documented, so anyone can check that the answer is wrong.
- 02 · CONTROL
Controls, mixed in blind
Clean answers that should pass and clear failures that should fail sit among the planted items. They show whether the judge is reading closely at all, and they catch false fails as well as false passes.
- 03 · REPEAT
k samples per item
LLM judges are not deterministic, so one run proves little. Every item is judged k times, and the false-pass rate is reported per defect family with its spread across runs.
- 04 · FIX
Fix pack, then re-score
A hardened rubric, a revised judge prompt, and automatic pre-checks for the things a judge should not be trusted with. Then we score the same items again so you can see what changed.
Three ways in. Flat fees.
AI-Judge Quick Check
- 3 defect families
- 50 items, 5 judge runs each
- False-pass table and exhibits
- 5 business days
AI-Judge False-Pass Audit
- All defect families, with controls
- Fix pack: hardened rubric, judge prompt, automatic pre-checks
- Re-scored after the fixes
- About 2 weeks
AI-Judge Monitoring
- A re-run whenever the rubric changes
- A re-run whenever the judge model changes
40% off for the first 3 clients.
We are running our first AI-judge audits now and have published no results yet. We will publish only when they meet the bar our code audits meet: every number traced to a run, every exhibit open to re-checking. The first three clients get 40% off any AI-judge offer in exchange for a case study agreed with you before publication.
What do you need from us?
The rubric, the judge prompt and model settings, and a set of items like the ones the judge grades in training. We call the judge exactly as your environment does.
Can you test a judge we host ourselves?
Yes, if we can call it: a hosted model behind your prompt, or your own endpoint. We sign your NDA first.
Is this the same as your code audits?
Same rules, different check. In a code audit, a separate behavioral test proves a fix is wrong. Here, each planted defect is documented, so anyone can check that the answer is wrong. In both, a pass we cannot show is false is never counted.