Can an AI Agent Actually Make Money Trading? (What the Funded Live Tests Show)*
The most-searched claim in AI trading wears three costumes — the contest winner, the backtest, the funded wallet. Here is how the court reads each.
The Charge
"Our AI agent beats the market." "The bot returned +164%." "Give a model a wallet and it wins." Same claim, three costumes. It is the most-searched question in AI trading, so it is the most-sold one.
This court has tried the category three times. The verdict did not change. Below is how the claim is built, what a real test requires, and what the funded tests actually found.
The three costumes
1. The live-contest winner. One agent turns a profit in one trading season, and the headline writes itself.
The tell is what goes unmentioned. A public forward-test of the same category ran 22 model-seasons: 5 were profitable — 22.7% — and zero stayed profitable across both seasons. A winner in one window that does not repeat is survivorship, not skill. You name the survivor and fold the rest.
Filed as [Case №008 — "AI agents beat the market"](/case/008): not proven.
2. The backtest in a track-record costume. The biggest returns — "+164%," "+252%" — are almost always backtested. A backtest is a story about the past told after you already know the ending: you rank the winners once the results are in.
The sources carrying the largest figures usually say it themselves in the fine print: *simulated, not live-capital.* Set against a live forward-test, the same category ran ≈7% versus a ≈4.5% index over 30 days, at 22% drawdown — and a retail tester reported losing $9,400 in under three months.
Filed as [Case №009 — "+164% returns"](/case/009): not proven.
3. The wallet with real money. The strongest version funds a model with real capital and lets it trade. This is the only version worth taking seriously — so look at the one benchmark that actually did it.
Six frontier models, $10,000 each, 57 days live on prediction markets. Every model lost money. Losses ran 16% to 30.8%. Research volume and token burn showed no correlation with returns. Effort was not alpha.
Filed as [Case №010 — "give an AI agent a wallet"](/case/010): not proven.
The standard of proof
A trading claim earns a "holds" only when it clears four bars. Each costume above fails at least one.
1. Live capital, not simulation. A backtest never lost a trade it did not take. Simulated ≠ real.
2. A benchmark for the same window. "+164%" means nothing without buy-and-hold over the identical period. Beating cash is not beating the market.
3. The whole field, not the winner. One profitable agent out of a class is a lottery ticket, not an edge. Ask for cross-season persistence.
4. Drawdown disclosed. A return with no stated drawdown is half a sentence. The risk is the other half.
The 30-second read
Before you believe an AI trading claim, ask:
- Is the headline number live or backtested? (If they will not say, it is backtested.)
- Is there a benchmark for the exact same window?
- Are you seeing the whole field or just the one winner?
- Is drawdown stated?
- Does it persist across more than one period?
Five noes and one yes is the usual score. That is not fraud. It is a real number computed on a rigged set — true in the narrow, false in the whole.
We do not say scam. We say: on the evidence, not proven.*
*\* we funded the question ourselves first, and lost.*