notproven*
The method · 2026-07-09

How to Tell If an AI Model Really Beats GPT (Benchmarks, Contamination, Confidence Intervals)*

“Our model beats GPT” is the easiest true-and-false claim in AI — the opponent trick, saturation, contamination, and cherry-picked sets, read by the court.

> Not by the headline. By the set it was measured on.

Every release cycle a model "beats GPT," is "10x more accurate," or delivers "frontier performance at a fraction of the cost." The chart is clean. The question is the only one this court asks: does it hold, doesn't it hold, or is it not proven?

This page is the method for benchmark boasts. Public, repeatable, dull on purpose. It names no model — it reads a claim-class.

---

Quick answer

How to tell if an AI model really beats GPT: don't trust the chart. Ask which opponent, which set, and with what error bar.

1. Which opponent — "beats GPT-5.5" is a claim about one rival. "Beats frontier" is the claim people hear. They are not the same sentence.

2. Which set — the two benchmarks it won, or the full set including the ones it lost?

3. What error bar — are confidence intervals published? A three-point lead with no interval may be noise.

4. Clean data — independent evaluation, no train/test contamination, no distillation dispute over the very benchmarks cited.

Miss one and the verdict is not proven — not false, just not established by the exhibit shown.

---

The Charge

A benchmark score is a claim. Most claims presented to this court are true in the narrow and false in the whole.

Nobody prints a fabricated score. They print a real score on a set chosen after the fact. The number holds. The framing does not. Four devices do most of the work.

Device One: The Opponent Trick

A model beats one named rival on two benchmarks. The press release says "beats GPT-5.5." The audience hears "beats the field."

On the same set, the model often trails a different frontier system. Beating one runner is not winning the race — it is beating one runner.

Concrete case. A model clears a stated rival on a reasoning exam (54.7 vs 52.2) and a coding set (62.1 vs 58.6). On that same coding set it trails a leading frontier model at 69.2. "Beats GPT-5.5" holds. "Beats frontier" does not.

Test. Ask for the full leaderboard, not the two rows that flatter. Name every model on the set, not the one chosen for the headline.

Device Two: Saturation

When the whole field scores above 88% on a benchmark, the benchmark has stopped discriminating. A two-point lead on a saturated exam is a rounding argument, not a capability argument.

Concrete case. Multiple-choice knowledge tests where frontier models cluster past 88%, and sentence-completion sets past 95%. The gaps between leaders sit inside the measurement's own slack. "Highest score" is technically true and functionally empty.

Test. Ask what the ceiling is and how close the field already sits to it. Near the ceiling, the ranking measures luck of the test split, not skill.

Device Three: Contamination and Eval-Gaming

The sharpest edge, and the one most often left unaddressed. If the test questions — or their close cousins — sat in the training data, the score measures memory, not reasoning. Distillation from a stronger model onto the exact benchmarks is the same failure by another route.

Related: framework dependence. The same model on the same benchmark posts a different number depending on the harness that runs it. A score with no harness disclosed is a score with a hidden variable.

Concrete case. A record benchmark result draws a public allegation of train/test contamination from independent researchers; the vendor does not respond. The number stands unretracted — and unverified. Reported here as allegation, never as fact. Either way, uncontrolled contamination leaves the score not proven.

Test. Ask for the contamination controls: held-out sets built after the model froze, canary strings, decontamination logs. Ask which harness produced the number. No controls, no clean read.

Device Four: Cherry-Picked Sets

Run twenty benchmarks. Publish the two you won. This needs no lie — the two scores are real. The set they came from is hidden, exactly like a survivorship track record.

Concrete case. A launch card shows two wins. The evaluation covered more. The losses are not on the card. The maximum of a distribution is presented as its center.

Test. How many benchmarks were run in total? Show all of them, wins and losses, or the two on the card are selection, not evidence.

What the Court Requires

A benchmark claim is not proven by a chart. It is proven by a result that survives its own error bar.

Minimum exhibits:

  • Independent evaluation. Run by someone who is not selling the model.
  • Published confidence intervals. A lead smaller than the interval is not a lead.
  • The full set. Every benchmark run, wins and losses, and the harness for each.
  • Contamination controls. Held-out data built after the freeze; decontamination disclosed.
  • The named field. Every frontier model on the set, not the one chosen for the sentence.

Supply these and the claim can hold. Withhold one and the verdict is fixed in advance.

The Verdict Rule

We do not say scam. We say the claim is not proven — and name the missing exhibit.

  • Beats one rival, sold as beating the field: holds narrow, not proven broad.
  • No confidence intervals: not proven.
  • Contamination alleged and unaddressed: not proven.
  • Two wins shown, full set hidden: cherry-picked — not proven.
  • Scores fail to replicate under a clean independent run: does not hold.

The honest lab hands over the full set and the error bars. The one who shows two rows is telling you which twenty it ran by the eighteen it will not print.

A real result fears no re-run. That is the whole test.

---

*See the ruling: [Case №007 — "our open model beats GPT-5.5"](/case/007). And the sibling method: [How to Spot a Fake AI Trading Track Record](/method/how-to-spot-a-fake-ai-trading-track-record).*

* method published in full. the receipts travel with every claim.ON THE RECORD