Who’s behind the audits
I’m Ilpo Väätäinen. I run an independent eval-integrity practice: I audit the numbers AI teams are about to stake their credibility on, and tell them which ones survive a second look. This page exists so you can decide whether to trust that — and my whole argument is that you shouldn’t have to.
How I got here
I didn’t come to this from a lab. I came to it from losing an argument with reality. I built a trading system, validated it on a careful backtest, and believed the backtest. Then I took it live — and the single cheapest measurement in the whole pipeline, a fourteen-cent order resting in a real order book, told me the backtest had been lying. Not wrong by a rounding error; wrong about the thing that mattered. The edge the numbers promised wasn’t there once real fees and a real queue had their say.
That gap — between a number that passes every test you thought to run and a number that’s actually true — turned out to be the same gap that sits under half the AI benchmark claims I read. So I started doing to eval numbers what the live market had done to my backtest: reproduce them, control them, and go measure the thing the summary skipped.
What I do now
Independent, reproducible audits of AI evaluation results — benchmark leaderboards, LLM-as-judge results, scaling-law claims. Before a launch, a paper, or a fundraise, I check whether the headline number holds, and hand back a plain-English report: certified, or revised, with the one-line statistical fix. I also wrote the book behind the practice, Measured, Not Believed, and maintain the open-source evalgate checks so anyone can run the cheap version themselves.
Why you don’t have to take my word for it
Trust in this work isn’t a credential — it’s reproducibility. Three principles keep me honest, and they’re the ones I’d want applied to my own numbers:
- Every finding is re-runnable. I audit from your released data with a deterministic method and hand you the commands. A number you can’t re-derive is a rumor, including mine.
- I certify real results as loudly as I revise weak ones. A tool that only ever finds fault is a cynic. Some claims survive correction — I say so, and strengthen them.
- I audit the framing, not the people. Findings target the one over-tightened sentence, not the work’s real contribution — and I state the limits of what an observational check can prove.
The best proof is public. Three audits, computed on the papers’ own released data, reproducing exactly:
- An LLM judge that rewards length — 68% longer-answer wins, 71.5% self-preference.
- A “best subset” win that evaporates under correction — p=0.009 → 0.19.
- A “super-linear” exponent that came free from a definition — and one point flips it.
This is an independent practice, not a big firm — which is the point. You get the person who does the work, a reproducible method instead of a brand, and an audit that’s as willing to clear your number as to catch it.