Reel #3 · Keep the best 2 of 10 trading bots
Picking the best backtests picks the luckiest bots.
Beat two random bots in 11 of 23 years No edge
Beat just holding the Nasdaq in 3 of 23 years Worse than holding
Average return per year.
The claim
"Ask an AI for ten bot ideas, backtest all ten, keep the best two." It sounds smart. We tested the method itself.
How we tested it
10 classic bot ideas (24 versions: trend, moving averages, RSI, breakouts, dips…) on the Nasdaq-100 ETF (QQQ). Every year for 23 years we kept the two bots with the best results over the previous two years, then measured what they did in the next year, against two bots picked at random and against simply holding QQQ. Costs: 0.05% per trade side.
Does the result depend on how you pick?
| How the "best 2" were picked | In backtest | Year after | Random 2 | Hold QQQ | Beat random |
|---|---|---|---|---|---|
| 1-year window, ranked by return | +23.2% | +7.9% | +8.8% | +18.7% | 13 / 23 |
| 1-year window, ranked by risk-adjusted return | +20.2% | +9.4% | +8.6% | +18.7% | 13 / 23 |
| 2-year window, ranked by return (the reel) | +18.0% | +8.9% | +8.6% | +18.7% | 11 / 23 |
| 2-year window, ranked by risk-adjusted return | +17.0% | +6.9% | +8.8% | +18.7% | 9 / 23 |
Returns per year, averaged over the test years. Whatever the picking rule, the winners' backtest numbers don't carry over.
Why
Picking the best backtests doesn't pick the best bots. It picks the luckiest ones, and luck doesn't repeat.
What this doesn't tell you
- One market (the Nasdaq-100), and bots from a fixed textbook list. Other markets or ideas could behave differently.
- "Beat random" compares against a coin flip: 11 of 23 is well within luck. "Beat holding" 3 of 23 is clearly worse.
- Past results don't predict future results. Taxes are ignored.
Data: QQQ daily (yfinance, adjusted), 1999-03-10 to 2026-09-28. Summary figures only; the price file isn't republished because of its source's terms.