Selection Bias in Trading: The Best Backtest Is Luck

Selection bias in trading: the best of 100 backtests looks good partly by chance. How trial counts inflate Sharpe, how DSR corrects it, and how to stay honest.

Trigr Research6 min read
On this page
  1. What is selection bias in backtesting?
  2. How good can luck look?
  3. Why do correlated variants still count?
  4. How does the Deflated Sharpe Ratio correct for it?
  5. Where does Trigr compute these, and where not?
  6. Why keep experiments inside one strategy?
  7. How to research without fooling yourself
  8. Next steps

TL;DR: If you backtest 100 variants of a strategy and keep the best, a large part of its performance is the luck of being the maximum of 100 noisy results. Even strategies with zero real edge produce an impressive best-of-100 Sharpe ratio. The fix is to count your trials, judge the winner against what luck alone would produce, and keep variants of one idea grouped so the count stays visible.

What is selection bias in backtesting?

Every backtest result is a mix of signal and noise. Run one backtest and the noise is just as likely to help as to hurt. Run many and keep the best, and you have systematically selected the one where noise helped most.

That is selection bias, also called data snooping or multiple testing. It does not require any mistake in the backtester. A perfectly point-in-time simulation with honest costs still produces a biased winner if you pick it from a large pool.

The bias is invisible in the final report. A backtest that was the only one you ran and a backtest that was the best of 500 look identical on screen. Only the research process tells them apart. Backtests are not guarantees; perps are leveraged and can lose more than expected.

How good can luck look?

A thought experiment makes the size of the problem concrete. Suppose none of your strategy variants has any real edge: each has a true Sharpe ratio of zero. You backtest each on three years of data and keep the best.

The estimated annualized Sharpe of a zero-edge strategy over three years has a standard error of roughly 1/√3, or about 0.58, using the standard approximation for Sharpe ratio estimation error described by Andrew Lo. The best of many such estimates follows the statistics of extremes. Bailey and López de Prado give an approximation for the expected maximum in The Deflated Sharpe Ratio (SSRN 2460551).

Independent variants tried Expected best Sharpe from luck alone (3 years, true Sharpe 0)
1 about 0.0
10 about 0.9
100 about 1.5
1,000 about 1.9

These are approximations under simplifying assumptions: independent trials and normally distributed returns. The pattern is what matters. Trying 100 unrelated zero-edge ideas is expected to produce a winner with an annualized Sharpe around 1.5, a figure many traders would happily deploy. Crypto returns are fat-tailed, which tends to make the problem worse, not better.

Why do correlated variants still count?

Most research is not 100 unrelated ideas. It is one idea with parameters nudged: RSI 14, then 15, then 16; a stop of 2%, then 2.5%. Those variants are highly correlated, so they behave like fewer independent trials. The selection penalty is smaller than the table suggests, but it is never zero.

The practical difficulty is that nobody knows the "effective" number of independent trials exactly. That is why the first discipline is simply to keep the raw count. You can argue about how much to discount it later. You cannot recover it if you never wrote it down.

Signs the count is getting large without anyone noticing:

  • Your strategy list contains many near-identical copies with names like "v2 final", "v2 final tighter stop" and "v3 test".
  • An AI assistant generated dozens of variants in one session and only the best one was saved.
  • You changed the asset, timeframe and filters until something worked, and cannot say how many combinations you saw.

How does the Deflated Sharpe Ratio correct for it?

The Deflated Sharpe Ratio (DSR), introduced by David Bailey and Marcos López de Prado, asks a sharper question than the raw Sharpe. Instead of "is this Sharpe above zero?", it asks "is this Sharpe above the best I would expect from luck, given how many trials I ran, how long my sample is and how skewed and fat-tailed my returns are?"

The output is a probability. A high DSR means the observed Sharpe is unlikely to be explained by selection among that many trials. A low DSR means the result is consistent with luck. We walk through the intuition in the Deflated Sharpe Ratio explained.

A companion measure, the Probability of Backtest Overfitting (PBO) from Bailey, Borwein, López de Prado and Zhu, estimates how often the configuration that ranks best in-sample falls below the median out-of-sample. See what PBO tells you.

Measure Question it answers Needs a trial count?
Raw Sharpe How good was the return per unit of volatility? No
Deflated Sharpe Ratio Is that Sharpe better than luck across N trials? Yes
PBO How often does the in-sample winner underperform out-of-sample? Uses a grid of configurations

Where does Trigr compute these, and where not?

Trigr is explicit about which results carry a selection correction:

  • Standard backtests show raw Sharpe in the headline and list DSR and PBO as "not yet computed". They do not invent values. Treat the raw Sharpe as optimistic, especially if the strategy came out of many variants.
  • Optimize with ML runs report the Deflated Sharpe Ratio using the run's trial count, plus PBO from a separate combinatorial purged cross-validation grid, out-of-sample Sharpe from anchored, purged walk-forward validation, a leak verdict and feature importance. The ML pipeline chooses among LightGBM, random forest and XGBoost by walk-forward out-of-sample performance and does not tune hyperparameters, which keeps its own trial count small. The AI and ML docs describe what each run returns.

The trial count inside an ML run only covers that run. The variants you tried by hand before it still count against you, which is why the way you iterate matters as much as the statistics.

Why keep experiments inside one strategy?

The simplest defense against hidden selection bias is structural: keep every variant of one idea in one place, so the count is visible by default.

On Trigr, this is how iteration works:

  • In Studio, the Copilot iterates in labeled experiments, Experiment 1, 2, 3 and so on, inside one strategy's version history. Any version can be restored.
  • Over MCP or the API, trigr_run_backtest accepts an experimentLabel together with an owned strategyId. The variant is recorded in that strategy's version history without changing the saved recipe or its verified statistics. trigr_get_backtest_result lists runs with their labels and whether each measured the saved recipe, and the winner is saved with trigr_edit_strategy.
  • The MCP server's guidance tells connected assistants to create one strategy per distinct idea, meaning market, timeframe and signal family, and to treat parameter, filter, exit and sizing changes as experiments.

What this means for you

When you review a strategy months later, you can see how many experiments preceded the version you saved, instead of guessing. When an AI assistant researches for you, its variants do not scatter into dozens of near-duplicate strategies that hide the search. You also avoid the practical costs of that sprawl: a cluttered library and saved-strategy limits used up by copies. The full workflow is in iterating on a strategy with AI experiments.

None of this makes a winning experiment real. It makes it possible to judge the winner fairly.

How to research without fooling yourself

A short protocol that works whether or not you use Trigr:

  1. Write the hypothesis first. One sentence on why the effect should exist, before any backtest.
  2. Decide the variant budget. For example, at most 20 experiments for this idea.
  3. Count every run, including the ones you discard.
  4. Judge the winner against luck. Use a deflated statistic, or at least compare it with the table above for your trial count.
  5. Get new data. Freeze the winner and forward-test it on a paper agent. Forward data cannot be selected on, because it did not exist when you chose.

Selection bias is one of the main reasons live results differ from backtests, and a visible trial count is the cheapest way to shrink that gap.

Next steps

Look at your current strategy list and count how many entries are really variants of the same idea. If you want an AI assistant to research with experiments instead of copies, connect it to Trigr over MCP and ask it to label each variant.

Frequently asked questions

What is selection bias in trading strategy research?

Selection bias happens when you test many strategy variants and keep the best one. Its backtest result is partly skill and partly the luck of being the maximum of many noisy outcomes, so it overstates how the strategy is likely to perform next.

How many backtests is too many?

There is no fixed limit, but the expected best result from pure luck rises with every variant you try. What matters is knowing the count and judging results against it, for example with the Deflated Sharpe Ratio, which takes the number of trials as an input.

Does a Trigr backtest report the Deflated Sharpe Ratio?

A standard Trigr backtest does not; it shows DSR and PBO as not yet computed. An Optimize with ML run reports the Deflated Sharpe Ratio using that run's trial count, along with PBO, out-of-sample Sharpe and a leak verdict.

Do similar variants count as separate trials?

They count, but not fully. Highly correlated variants, such as an RSI of 14 versus 15, behave like fewer independent trials. You still pay a selection penalty, just a smaller one than for the same number of genuinely different ideas.

Put the idea to an honest test.

Describe a strategy in plain English or from your own AI assistant, backtest it on point-in-time data, and forward-test it on paper before any real money is involved.