Measured, Not Believed — eval integrity live status → free checks →
Independent audit practice

Most benchmark wins don't survive a second look.

I find the ones that don't — before you publish them. Independent, reproducible audits of the numbers your model, your eval, or your leaderboard is about to stake its credibility on. If a claim is real, you get a clean bill of health you can cite. If it isn't, you find out from me, quietly, not from a reviewer or a competitor on launch day.


Case files · public data, reproduced

Five audits. Four numbers that overstated — one that held.

Each was recomputed from public, released data. Four target the one summary sentence that overpromised — a judge's bias, a lucky subset, a mechanical "law", an order effect. The fifth is the control: a #1 I tried to break and couldn't, so I certify it. An auditor who only ever says no isn't checking anything. This is exactly the work I'd do on yours.

MT-Bench · GPT-4 judge

The judge preferred longer answers and its own family.

A widely-used LLM-as-judge, audited on its own released pairwise verdicts.

longer answer wins  68.0% of pairwise verdicts
judge's own model family wins  71.5%
both  p ≈ 0, fully reproducible from released data
bias, not skill Read the full write-up →
RewardBench · 23 subsets

"We lead on subset X" was the luckiest of 23 noisy tests.

A "best-subset" win, corrected for how many subsets it could have been picked from.

raw  p = 0.009  (looks significant)
Šidák-corrected (23 subsets)  p = 0.19  (not significant)
a different, robust pair  survives correction at p = 0.011
max of many · not a finding Read the full write-up →
Grokking · loss-landscape geometry

A "super-linear power law" that came free from a minus sign.

A reported scaling exponent for warning lead-time, tested against what the definition alone hands you.

reported  α = 1.18 (SCAN), 1.13 (Dyck) — "super-linear"
constant-onset null (definition only)  α = 1.47 / 1.07
Dyck, drop 1 high-leverage point  α = 0.97, higher R²
mechanical · fragile Read the full write-up →
LMArena · the control · certified

The most-cited #1 in the field — I tried to break it and couldn't.

An independent Bradley-Terry recomputation of the public arena battles, bootstrapped for rank confidence.

gemini-2.5-pro over 98,088 battles  rank CI [1, 1]
P(truly #1)  1.00  (0.96 on coding votes)
one 19k shard  P = 0.83, 3-way tie — volume is the story
resolved · certified Read the full write-up →
LMArena · human-vote order bias

The model shown second wins more — and it isn't the better model.

An order-bias check on the same public human votes, controlled within each matchup.

second-shown model wins  50.62%  vs 49.38% (98k votes)
within-matchup, controlled for strength  +1.25 pp, p ≈ 9×10⁻⁵
verdict  real recency bias — randomization launders it, but only if you randomize
order effect · real Read the full write-up →

The engagement

Pick the depth you need.

Every engagement ends in one plain-English readout: which claims hold, which don't, and the one-line statistical fix for each that doesn't. Fixed price, no retainer required, NDA on request.

Spot Check
$600
One claim you're about to publish — a subset win, an exponent, a "remarkable agreement".
  • One benchmark / leaderboard / judge claim
  • Reproduced from your data
  • One-page verdict + the fix
  • 72-hour turnaround
Book Spot Check →
most booked
Full Eval Audit
$3,500
Your eval or leaderboard, end to end, before a launch or a paper.
  • Judge bias: position, verbosity, self-preference
  • Multiple-comparisons & look-elsewhere
  • Metric-artifact & averaged-away-bias checks
  • Robustness: leave-one-out, mechanical nulls
  • Written report + a call to walk it
Book Full Audit →
CI Gate
from $1,000 /mo
Stop a confounded number before it ships — every release, automatically.
  • The 5-probe check as a CI gate
  • Blocks a confounded metric on PR
  • Wired to your eval pipeline
  • Quarterly re-audit + tuning
Start CI Gate →

Not ready to hand it over? Run the free browser checks first, or get The Eval Integrity Kit — the 9-check checklist and the report template I ship to clients, $34 to do it yourself.


Method

Five questions I ask every number.

Not opinions — arithmetic. Each is a one-line check that a careful team could run itself; the value is that I run all of them, adversarially, on a claim you're too close to.

01 · before

What would the definition give you for free?

Replace the interesting variable with a constant. If the "law" survives, it was in the algebra.

02 · look-elsewhere

How many slices could you have picked?

Correct the "best subset" for the family it was chosen from. Most wins don't survive it.

03 · fragility

Does one row flip the verdict?

Leave-one-out on every fit. A conclusion that hangs on a single high-leverage point isn't one.

04 · the judge

Is the winner winning, or just longer?

Position swap, verbosity, self-preference. Certify the ranking, or expose the bias wearing its label.

05 · the tell

Did the average fold the structure flat?

When mean-absolute equals mean-signed, "scatter" is a one-directional gap the summary hid.

→ and

If it holds, I say so.

A tool that only ever finds fault is a cynic. Clean claims get a citable clean bill of health.

Run checks 1–3 free in your browser → Get the DIY audit kit →


The credential

Measured, Not Believed

The book behind the practice — why AI benchmark scores and trading backtests overpromise, and how to catch them. Twenty chapters of real audits, from LLM judges to gravitational waves. Pay what you want.

About the book →
Book an audit

Send me the claim you're least sure about.

The one line in the paper, the slide, or the launch post you'd least want a reviewer to poke at. I'll tell you whether it holds — before anyone else looks.