AI-Judge False-Pass Audit

Your AI judge gives full marks. Was the answer right?

When a rubric or an LLM judge supplies the reward, a policy can learn to satisfy the judge instead of the task. We measure how often your judge passes answers we know are wrong, show you each one, and hand you the fixes.

Early access: 40% off for the first 3 clients, in exchange for a case study.

Why judges get gamed

A judge can only check what the rubric names.

In May 2026, researchers at Scale Labs trained policies against rubric-based rewards and scored the results with three frontier judges from different model families. Weak verifiers produced large reward gains that did not hold up under those judges, and the exploitation grew as training went on. The common failures they describe: criteria with several parts credited when only some were met, implied content treated as if it were stated, and loose matching on topic.

Stronger judges reduced the problem but did not remove it. Their sharpest finding is about the rubric itself:

“…stronger verification does not prevent reward hacking when the rubric leaves important failure modes unspecified.”

Mahmoud et al., “Reward Hacking in Rubric-Based Reinforcement Learning”, Scale Labs, May 12, 2026

So a better judge model is not enough on its own. You need to know which wrong answers your rubric and judge let through, how often, and what closes the gap.

Our method

Plant known defects. Count what gets through.

The same rule as our code audits: nothing counts as a false pass unless we can show, item by item, that the answer was wrong.

  1. 01 · PLANT

    Defects with certificates

    We take answers your judge should pass and plant one known defect in each: a wrong fact, a missed requirement, an unsupported claim. Every item carries a certificate saying what was changed, where, and why it breaks the task. Each planted defect is documented, so anyone can check that the answer is wrong.

  2. 02 · CONTROL

    Controls, mixed in blind

    Clean answers that should pass and clear failures that should fail sit among the planted items. They show whether the judge is reading closely at all, and they catch false fails as well as false passes.

  3. 03 · REPEAT

    k samples per item

    LLM judges are not deterministic, so one run proves little. Every item is judged k times, and the false-pass rate is reported per defect family with its spread across runs.

  4. 04 · FIX

    Fix pack, then re-score

    A hardened rubric, a revised judge prompt, and automatic pre-checks for the things a judge should not be trusted with. Then we score the same items again so you can see what changed.

Offers

Three ways in. Flat fees.

MEASURE

AI-Judge Quick Check

  • 3 defect families
  • 50 items, 5 judge runs each
  • False-pass table and exhibits
  • 5 business days
From $5,000 per environment
Request a Quick Check →
Full audit

AI-Judge False-Pass Audit

  • All defect families, with controls
  • Fix pack: hardened rubric, judge prompt, automatic pre-checks
  • Re-scored after the fixes
  • About 2 weeks
Quoted per environment
Request the full audit →
KEEP CURRENT

AI-Judge Monitoring

  • A re-run whenever the rubric changes
  • A re-run whenever the judge model changes
Monthly quoted per environment
Ask about monitoring →
Early access

40% off for the first 3 clients.

We are running our first AI-judge audits now and have published no results yet. We will publish only when they meet the bar our code audits meet: every number traced to a run, every exhibit open to re-checking. The first three clients get 40% off any AI-judge offer in exchange for a case study agreed with you before publication.

What do you need from us?

The rubric, the judge prompt and model settings, and a set of items like the ones the judge grades in training. We call the judge exactly as your environment does.

Can you test a judge we host ourselves?

Yes, if we can call it: a hosted model behind your prompt, or your own endpoint. We sign your NDA first.

Is this the same as your code audits?

Same rules, different check. In a code audit, a separate behavioral test proves a fix is wrong. Here, each planted defect is documented, so anyone can check that the answer is wrong. In both, a pass we cannot show is false is never counted.