Most benchmark wins don't survive a second look.
I find the ones that don't — before you publish them. Independent, reproducible audits of the numbers your model, your eval, or your leaderboard is about to stake its credibility on. If a claim is real, you get a clean bill of health you can cite. If it isn't, you find out from me, quietly, not from a reviewer or a competitor on launch day.
Five audits. Four numbers that overstated — one that held.
Each was recomputed from public, released data. Four target the one summary sentence that overpromised — a judge's bias, a lucky subset, a mechanical "law", an order effect. The fifth is the control: a #1 I tried to break and couldn't, so I certify it. An auditor who only ever says no isn't checking anything. This is exactly the work I'd do on yours.
The judge preferred longer answers and its own family.
A widely-used LLM-as-judge, audited on its own released pairwise verdicts.
judge's own model family wins 71.5%
both p ≈ 0, fully reproducible from released data
"We lead on subset X" was the luckiest of 23 noisy tests.
A "best-subset" win, corrected for how many subsets it could have been picked from.
Šidák-corrected (23 subsets) p = 0.19 (not significant)
a different, robust pair survives correction at p = 0.011
A "super-linear power law" that came free from a minus sign.
A reported scaling exponent for warning lead-time, tested against what the definition alone hands you.
constant-onset null (definition only) α = 1.47 / 1.07
Dyck, drop 1 high-leverage point α = 0.97, higher R²
The most-cited #1 in the field — I tried to break it and couldn't.
An independent Bradley-Terry recomputation of the public arena battles, bootstrapped for rank confidence.
P(truly #1) 1.00 (0.96 on coding votes)
one 19k shard P = 0.83, 3-way tie — volume is the story
The model shown second wins more — and it isn't the better model.
An order-bias check on the same public human votes, controlled within each matchup.
within-matchup, controlled for strength +1.25 pp, p ≈ 9×10⁻⁵
verdict real recency bias — randomization launders it, but only if you randomize
Pick the depth you need.
Every engagement ends in one plain-English readout: which claims hold, which don't, and the one-line statistical fix for each that doesn't. Fixed price, no retainer required, NDA on request.
- One benchmark / leaderboard / judge claim
- Reproduced from your data
- One-page verdict + the fix
- 72-hour turnaround
- Judge bias: position, verbosity, self-preference
- Multiple-comparisons & look-elsewhere
- Metric-artifact & averaged-away-bias checks
- Robustness: leave-one-out, mechanical nulls
- Written report + a call to walk it
- The 5-probe check as a CI gate
- Blocks a confounded metric on PR
- Wired to your eval pipeline
- Quarterly re-audit + tuning
Not ready to hand it over? Run the free browser checks first, or get The Eval Integrity Kit — the 9-check checklist and the report template I ship to clients, $34 to do it yourself.
Five questions I ask every number.
Not opinions — arithmetic. Each is a one-line check that a careful team could run itself; the value is that I run all of them, adversarially, on a claim you're too close to.
What would the definition give you for free?
Replace the interesting variable with a constant. If the "law" survives, it was in the algebra.
How many slices could you have picked?
Correct the "best subset" for the family it was chosen from. Most wins don't survive it.
Does one row flip the verdict?
Leave-one-out on every fit. A conclusion that hangs on a single high-leverage point isn't one.
Is the winner winning, or just longer?
Position swap, verbosity, self-preference. Certify the ranking, or expose the bias wearing its label.
Did the average fold the structure flat?
When mean-absolute equals mean-signed, "scatter" is a one-directional gap the summary hid.
If it holds, I say so.
A tool that only ever finds fault is a cynic. Clean claims get a citable clean bill of health.
Run checks 1–3 free in your browser → Get the DIY audit kit →
Measured, Not Believed
The book behind the practice — why AI benchmark scores and trading backtests overpromise, and how to catch them. Twenty chapters of real audits, from LLM judges to gravitational waves. Pay what you want.
Send me the claim you're least sure about.
The one line in the paper, the slide, or the launch post you'd least want a reviewer to poke at. I'll tell you whether it holds — before anyone else looks.