TL;DR: An overfit backtest fits the noise of one historical sample, so it looks better than the rules deserve. The most reliable warning signs are fragility (small parameter changes break it), thin evidence (few trades, many rules), concentration (a handful of trades or one month carry the result), and an unknown number of discarded variants. Check each one before you put money behind a curve.
What does backtest overfitting actually mean?
Every price history contains two things: effects that may repeat, and noise that will not. A backtest cannot tell them apart on its own. If you keep adjusting rules until the equity curve looks good, you are partly teaching the strategy the exact sequence of past noise.
The result is a strategy that is well fitted to one sample and poorly fitted to the next one. It is the same problem statisticians call overfitting in any model, and it gets worse the more freedom you give yourself: more parameters, more filters, more variants tried.
Overfitting is rarely a single mistake. It is usually the sum of many small, reasonable-looking choices. That is why a checklist of symptoms is more useful than a definition. Backtests are not guarantees; perps are leveraged and can lose more than expected.
The 10 warning signs at a glance
| # | Warning sign | Quick test |
|---|---|---|
| 1 | Sharpe looks too good for the data | Compare with the Sharpe you would expect from luck given your trial count |
| 2 | Small parameter changes break it | Nudge each parameter 10–20% and rerun |
| 3 | Many rules, few trades | Count rules and parameters against trades |
| 4 | Too few trades overall | Check the trade log length and spread across time |
| 5 | Profits concentrated in a few trades | Remove the best 3–5 trades and recompute |
| 6 | One period carries the curve | Read monthly returns; check the recent period separately |
| 7 | The edge disappears after costs | Turn on slippage and funding |
| 8 | Unknown number of variants tried | Write down the trial count honestly |
| 9 | Works on one market, fails on its neighbors | Run the same rules on similar assets |
| 10 | Tight exits with a very high win rate | Check how intrabar TP/SL ties were resolved |
The sections below explain each sign, why it matters and what to do about it.
Signs 1–4: the evidence is thinner than it looks
1. The Sharpe ratio looks too good
A multi-year crypto strategy with a Sharpe ratio far above what professional funds report deserves suspicion before celebration. High numbers are possible on short samples or narrow regimes, but they are also exactly what luck produces when you try many variants.
Bailey and López de Prado's Deflated Sharpe Ratio formalizes the check: it raises the bar a Sharpe must clear based on how many trials you ran, how long the sample is and how skewed or fat-tailed the returns are. We explain it in plain language in the Deflated Sharpe Ratio explained.
2. Small parameter changes break it
Suppose your strategy uses a 21-period RSI with an entry at 28. If 20 or 22 periods, or an entry at 26 or 30, turn a strong result into a flat or losing one, you have found a sharp peak, not a plateau. Real effects tend to degrade gradually as parameters move. Noise tends to produce isolated spikes.
A simple test: move each parameter by 10–20% in both directions and look at the neighborhood. You want most neighbors to be profitable, even if less so than the best one.
3. Many rules for the number of trades
Each extra filter, threshold or exit rule is another degree of freedom. A strategy with eight conditions and 60 trades has enough flexibility to explain almost any 60 outcomes. As a rough habit, be wary when you can describe the trades almost one rule per handful of trades.
Simpler graphs are not automatically better, but every node should earn its place with a reason you could state before seeing the backtest.
4. Too few trades overall
A few dozen trades is a small sample. A 60% win rate on 30 trades is statistically hard to distinguish from a coin flip. The same win rate on 600 trades, spread across bull, bear and sideways periods, is much more informative.
Check both the count and the spread. Two hundred trades that all happened in one quarter are closer to one observation of one regime than to two hundred independent tests.
Signs 5–7: the result depends on a few moments
5. Profits come from a handful of trades
Open the trade log and sort by P&L. If removing the best three to five trades turns the total negative, the strategy's result depends on rare events that may not repeat on schedule. Some legitimate strategies, such as trend following, do rely on a few large winners, but then the rules should explain why those winners were caught, not just that they were.
6. One period carries the whole curve
Read the monthly returns table rather than just the final number. A curve that is flat for three years and then jumps in one month tells a different story from one that grinds up steadily.
Trigr shades the trailing 25% of a full-history backtest so you can compare recent behavior with the earlier curve. That shading is a recent-period diagnostic for spotting degradation. It is not an out-of-sample holdout, because once you look at it while choosing a strategy, it becomes part of your research data. A sharp difference between the two segments is still a useful warning.
7. The edge disappears after costs
Overfit strategies often trade too much, because the optimizer found many small patterns in noise. Those patterns are thinner than the costs of trading them. On Trigr, trading fees and the builder fee are always applied; slippage and funding are opt-in and off by default, so a first result is labeled gross. Turn both on and see what survives. Our guide to how trading fees decide whether a perp strategy works walks through the arithmetic.
Signs 8–10: the process, not the curve
8. You do not know how many variants you tried
This is the most common and least visible sign. If you tested 100 variants and kept the best, its result is partly the maximum of 100 noisy draws. Without a trial count, no one, including you, can judge how much of the result is selection.
The discipline is simple: keep variants of one idea together and count them. On Trigr, the Copilot iterates in labeled experiments inside one strategy's version history, and MCP clients pass an experimentLabel when they backtest a variant of an owned strategy. The variants stay visible as a count rather than scattering into dozens of near-duplicate strategies. We cover why this matters in selection bias and the best of 100 backtests.
9. It works on one market and fails on its neighbors
If a breakout rule works on SOL but fails on ETH, AVAX and SUI, ask why. Sometimes there is a real structural reason. More often, the parameters fit SOL's particular history. Running the same rules unchanged on two or three similar markets is a cheap robustness check, and it costs nothing in extra degrees of freedom.
10. Tight exits with a surprisingly high win rate
When the take-profit and stop-loss are both close to the entry price, many candles will touch both. A backtester that assumes the take-profit came first will report an inflated win rate. Trigr replays 5-minute bars inside an ambiguous bar to find which level was touched first, and if that still cannot settle it, the stop wins. The full reasoning is in stop-loss and take-profit in the same candle.
What to do when your strategy shows these signs
Failing one check does not mean the idea is worthless. It means the evidence is weaker than the headline. A practical response:
- Simplify. Remove the filter that adds the least, and rerun. If performance barely changes, keep the simpler version.
- Widen the evidence. Test the same rules on longer history, on neighboring markets and with costs on.
- Count and deflate. Record the number of variants and prefer statistics that account for it.
- Get genuinely new data. Freeze the rules and forward-test them on a paper agent before risking capital.
For a statistical second opinion, an Optimize with ML run on Trigr reports a leak verdict, out-of-sample Sharpe from anchored, purged walk-forward validation, the Deflated Sharpe Ratio using the run's trial count, and PBO from a separate combinatorial grid. The Probability of Backtest Overfitting article explains how to read that last number, based on the method by Bailey, Borwein, López de Prado and Zhu. A standard backtest does not compute DSR or PBO; it lists them as not yet computed rather than inventing values.
What this means for you
Overfitting is not something you can eliminate, only something you can measure and limit. The traders who handle it well do three things consistently: they keep their trial count visible, they read the trade log instead of just the equity curve, and they insist on forward evidence before scaling up.
Trigr is designed around those habits: point-in-time data, next-bar-open fills, conservative intrabar handling, explicit gross versus net, and experiments grouped inside one strategy. None of that makes a strategy profitable. It makes it harder for a lucky curve to pass as a skilled one. The point-in-time backtesting guide covers the data side of the same problem.
Next steps
Take your current favorite strategy and run it through the ten checks in the table above, starting with costs on and parameter nudges. When a strategy survives, forward-test it with a paper trading agent before committing capital. The pricing page lists how many paper agents and credits each plan includes.