How We Test an AI Claim — The Method Behind Every Verdict*
Six tests turn a marketing number into a ruling — denominator, survivorship, backtest vs. live, the right benchmark, out-of-sample, the opponent trick. Published in full. We ran ourselves through it first, and lost.
The Charge
Every claim that reaches this court is a promise about numbers. "Beats the market." "80.4% win rate." "Beats GPT-5.5." "Never a bad month." The words are cheap. The test is not. This page is the test — the exact method every verdict on this site runs through, published in full, so anyone can check our work or run it themselves.
We have nothing to sell to the accused. The method is the product, and the method is open.
The First Defendant Was Us
Before we ruled on anyone, we ran three of our own trading bots through this same process. The record: −€52.83, −€47.84, −€7.13. Three losses. We published them as Cases [002](/case/002), [003](/case/003) and [004](/case/004). If the method is good enough to convict us, it is good enough to report on a claim from anyone else.
The Six Tests
Every claim is scored against six questions. A claim that survives all six *holds*. A claim that fails one *doesn't hold*. A claim we cannot check with public evidence is *not proven* — the most common ruling, and never an insult.
1. Denominator. A win rate is a fraction. We ask what sits under the line. "80% wins" on 10 cherry-picked trades is noise; the same rate across every trade, live, is signal. Missing denominator → not proven.
2. Survivorship. We ask who was counted and who was quietly dropped. A leaderboard of the ten bots still standing hides the ninety that died. See [How to Spot a Fake AI Trading Track Record](/method/how-to-spot-a-fake-ai-trading-track-record).
3. Backtest vs. live. Simulated returns are a hypothesis, not a result. A curve fit to the past is not money made in the present. Headline +164% backtested, forward-tested near an index → not proven.
4. The right benchmark. "Made money" is not the bar. The bar is buy-and-hold over the same window. Beating cash while losing to the index is not alpha. See [Can an AI Agent Actually Make Money Trading?](/method/can-an-ai-agent-actually-make-money-trading).
5. Out-of-sample persistence. One good season is luck until it repeats on data the model never saw. We require deflated, out-of-sample performance — a Probabilistic Sharpe Ratio at or above 0.95 — before "skill" is on the table.
6. The opponent trick. "Beats GPT" depends entirely on which GPT, which benchmark, which harness. Saturated benchmarks (>88%) turn real gaps into statistical noise. See [How to Tell If an AI Model Really Beats GPT](/method/how-to-tell-if-an-ai-model-really-beats-gpt).
The Record So Far
Nine claims tested. Three didn't hold. Six not proven. Zero held. The full docket, with exact numbers per case, is [here](/method/is-ai-trading-legit-the-record).
We report a test. We do not chase a person. The number is on trial, never the name.
*the method is public. so is the losing streak.*