Free 20-task audit

Get a free 20-task audit.

Send up to 20 tasks from your environment, or point us to a public set. We run the same deep search a paid audit uses and send every finding with full evidence. $0, no obligation.

  • 01A verdict per task. KEEP, FIX or DROP, with the reason.
  • 02Every wrong fix your grader paid for, with the check that proves it wrong and the hashes of its run logs.
  • 03A one-command reproduction for every finding: bash and Docker, nothing of ours to install.
  • 04A walkthrough of the results with your team, if you want one.
Shane Cope · founder

You will hear from me, not a sales team, within 1 business day. The audit engine is built and run with AI agents, and every result ships with a one-command reproduction so you can check it yourself.

Request your free audit

Two minutes. We reply within 1 business day to confirm scope.

How it works

From form to findings, in four steps.

  1. Tell us about your tasksThe form above: format, rough count, and a link if the set is public.
  2. We confirm scope within 1 business dayWhich 20 tasks, how we get access, a delivery date, and your NDA if you need one.
  3. We run the deep searchEach task through its own grader, unchanged, in sealed containers with the network off.
  4. You get the findings, with proofEvery finding, its evidence and its reproduction kit, and a walkthrough if you want one.
What we need

Your tasks, as they are.

For each task: its container image or Dockerfile, the reference fix, the tests and the grading script, the way your environment runs them. Harbor and SWE-bench-style tasks run as they are; for a custom format, one example task tells us what we need.

Nothing private to share yet? Send a link to a public task set you use, and we will audit 20 of its tasks instead.

Questions

Is there a catch?

Why is it free?

Because evidence on your own tasks makes the case for an audit better than we can. If the sample finds wrong fixes your grader pays for, you will want the rest checked, at a flat fee agreed up front. If it finds none, you have learned something too.

Is it the full method, or a demo?

The full method: the same deep search a paid audit uses, the same separate checks, and full evidence for every finding.

Who sees the results?

Only you. We never publish your tasks or the results of your audit.

What happens to our tasks and code?

They run in sealed containers with the network off. Task text and code are sent to AI model APIs to write wrong fixes and checks; if a provider is off-limits for you, tell us and we will leave it out. The pricing FAQ has the details.