notproven*
The method · 2026-07-09

Can an AI Agent Actually Make Money Trading? (What the Funded Live Tests Show)*

The most-searched claim in AI trading wears three costumes — the contest winner, the backtest, the funded wallet. Here is how the court reads each.

The Charge

"Our AI agent beats the market." "The bot returned +164%." "Give a model a wallet and it wins." Same claim, three costumes. It is the most-searched question in AI trading, so it is the most-sold one.

This court has tried the category three times. The verdict did not change. Below is how the claim is built, what a real test requires, and what the funded tests actually found.

The three costumes

1. The live-contest winner. One agent turns a profit in one trading season, and the headline writes itself.

The tell is what goes unmentioned. A public forward-test of the same category ran 22 model-seasons: 5 were profitable — 22.7% — and zero stayed profitable across both seasons. A winner in one window that does not repeat is survivorship, not skill. You name the survivor and fold the rest.

Filed as [Case №008 — "AI agents beat the market"](/case/008): not proven.

2. The backtest in a track-record costume. The biggest returns — "+164%," "+252%" — are almost always backtested. A backtest is a story about the past told after you already know the ending: you rank the winners once the results are in.

The sources carrying the largest figures usually say it themselves in the fine print: *simulated, not live-capital.* Set against a live forward-test, the same category ran ≈7% versus a ≈4.5% index over 30 days, at 22% drawdown — and a retail tester reported losing $9,400 in under three months.

Filed as [Case №009 — "+164% returns"](/case/009): not proven.

3. The wallet with real money. The strongest version funds a model with real capital and lets it trade. This is the only version worth taking seriously — so look at the one benchmark that actually did it.

Six frontier models, $10,000 each, 57 days live on prediction markets. Every model lost money. Losses ran 16% to 30.8%. Research volume and token burn showed no correlation with returns. Effort was not alpha.

Filed as [Case №010 — "give an AI agent a wallet"](/case/010): not proven.

The standard of proof

A trading claim earns a "holds" only when it clears four bars. Each costume above fails at least one.

1. Live capital, not simulation. A backtest never lost a trade it did not take. Simulated ≠ real.

2. A benchmark for the same window. "+164%" means nothing without buy-and-hold over the identical period. Beating cash is not beating the market.

3. The whole field, not the winner. One profitable agent out of a class is a lottery ticket, not an edge. Ask for cross-season persistence.

4. Drawdown disclosed. A return with no stated drawdown is half a sentence. The risk is the other half.

The 30-second read

Before you believe an AI trading claim, ask:

  • Is the headline number live or backtested? (If they will not say, it is backtested.)
  • Is there a benchmark for the exact same window?
  • Are you seeing the whole field or just the one winner?
  • Is drawdown stated?
  • Does it persist across more than one period?

Five noes and one yes is the usual score. That is not fraud. It is a real number computed on a rigged set — true in the narrow, false in the whole.

We do not say scam. We say: on the evidence, not proven.*

*\* we funded the question ourselves first, and lost.*

* method published in full. the receipts travel with every claim.ON THE RECORD