Flat fees. Never per finding.
Every price is per task, per month or per release, and fixed before we start. We never charge per finding, so we have no reason to report more than is there.
Grader audits and testing
For coding tasks graded by tests: Harbor, SWE-bench style and custom formats.
| Product | Price | What’s included | Turnaround | Action |
|---|---|---|---|---|
| Free sampleSee what an audit finds before you buy one. | $0up to 20 tasks | Deep search, full evidence. | Scoped within 1 business day |
Get it free |
| Grader AuditFind and close the tasks that pay for wrong code. | From $40per task audited $1,500 minimum per batch volume pricing above 1,000 tasks |
Every confirmed wrong fix with its proof, verified replacement tests, a panel of AI models from three companies, and an independent engineer’s review of a sample. | About 5 business days |
Request |
| Lab Acceptance TestingFor labs that buy environments. | Quotedper batch, fixed before we start | Independent inspection of vendor deliveries before training. | 3 business days | Request |
| Continuous MonitoringFor environments that ship new tasks every month. | Monthlyquoted to your volume | Tasks audited as they ship. | 2 business days per batch |
Request |
| Plumbline VerifiedCertification for one named release. | Fixed feeper release, quoted up front same fee to keep it current |
A re-audit to the published criteria. The fee is the same whether the release passes or fails. | Agreed per release |
Request |
Every quote is a fixed price agreed before we start, and it never depends on how many problems we find.
AI-judge false-pass audits
For environments where a rubric or an LLM judge supplies the reward. How the method works.
| Product | Price | What’s included | Turnaround | Action |
|---|---|---|---|---|
| AI-Judge Quick CheckA first measurement. | From $5,000per environment | 3 defect families, 50 items, 5 judge runs each, false-pass table and exhibits. | 5 business days | Request |
| AI-Judge False-Pass AuditMeasure, fix and re-score. | Quotedper environment | All defect families, controls, a fix pack (hardened rubric, judge prompt, automatic pre-checks), re-scored. | About 2 weeks | Request |
| AI-Judge MonitoringKeep the measurement current. | Monthlyquoted per environment | A re-run whenever the rubric or judge model changes. | On each change | Request |
Early access: the first 3 AI-judge clients get 40% off in exchange for a case study. Apply for a place.
Proof for every finding, and the fix.
Every confirmed wrong fix, with its proof
The patch the grader paid for, the behavioral check that shows it is wrong, its three results, and SHA-256 hashes of the run logs.
Verified replacement tests
For each task with a finding: tests that pass the reference fix and fail every wrong fix we found, run in the task’s own container.
Review beyond our own
A panel of AI models from three companies judges sampled findings blind, and an independent software engineer reviews a sample before results go out, with planted controls to keep that review honest.
One-command reproduction
A reproduce.sh kit for every finding: bash and Docker, no Plumbline code, no API key, no model.
Our accuracy guarantee. If any finding we deliver turns out to be wrong, we correct it in writing and refund that task’s fee. It covers our accuracy, never the number of findings.
Before you ask
Why a flat fee instead of a fee per finding?
A fee per finding rewards finding more problems than are there, and an auditor paid by the result is not independent. A flat fee per task keeps your bill predictable and keeps our incentive on getting each verdict right.
Certification works the same way: the fee is fixed per release and does not change whether the release passes or fails.
What counts as a finding?
A wrong fix that your task’s own grader pays full reward for, and that passes three tests of our own:
its code differs from the reference fix; a separate behavioral check passes on the reference fix, fails on the unfixed code and fails on the wrong fix, twice, in fresh containers; and the behavior it gets wrong is something the task as written requires.
Gaps outside what the task asks for are reported separately and never counted as findings.
Is every audit checked by more than one company’s AI?
Yes: wrong fixes come from AI models from two companies, and scope is voted on by AI models from three companies (Claude, GPT, Gemini). Before results go out, an independent software engineer reviews a sample of findings, with planted controls to keep that review honest.
How long does it take?
Grader Audit: about 5 business days. Lab Acceptance Testing: 3 business days. Continuous Monitoring: 2 business days per batch. AI-Judge Quick Check: 5 business days. AI-Judge False-Pass Audit: about 2 weeks.
For the free sample, we confirm the scope and a delivery date in our first reply, within 1 business day.
How do you handle our data?
Your tasks run only in sealed containers: a fresh container for every run, network off, all Linux capabilities dropped.
To write wrong fixes and checks, task text and code are sent to AI model APIs from Anthropic, OpenAI and Google under their API terms. If a provider is off-limits for you, tell us and we will leave it out; the report will say so.
We never publish your tasks or your results. We keep run records so any finding can be re-checked; tell us your retention or deletion requirements before we start.
What if one of your findings is wrong?
Every finding you get comes with a one-command replay. If any finding we deliver turns out to be wrong (it does not reproduce, or the behavior it tests is not required by the task), we correct it in writing and refund that task’s fee. The guarantee is about our accuracy, never about how many problems we find: a task with no finding costs the same as one with five.
Will you sign an NDA?
Yes. We sign your mutual NDA before you send private tasks. The free sample can also run on a public task set, which needs no NDA.
Which task formats can you run?
Harbor tasks, SWE-bench-style tasks and function-level tasks, each graded the way its own benchmark grades it. If your format is custom, send one sample task and we will tell you within 1 business day whether we can run it as it is.
What if the free sample finds nothing?
Then you know your grader held up against every wrong fix the deep search tried, and you have the evidence to show for it. No audit can prove a task cannot be gamed, and we never claim it.
Start with 20 tasks, free.
The same deep search as a paid audit, with full evidence. Then decide.