TL;DR: Data leakage is information from the future reaching a trading model through its features, labels or validation splits. It is the most common reason an ML strategy with a great out-of-sample score fails live. Prevention (point-in-time features, purged splits) is necessary but not enough; you also need active checks that try to prove a result is contaminated, such as a shuffled-label control and an embargo sensitivity test.
What is data leakage in trading machine learning?
In general machine learning, leakage means the training process sees information it would not have at prediction time. In trading, "prediction time" has a precise meaning: the moment your strategy decides whether to act, usually the close of a bar. Anything the model uses that was not knowable at that moment is a leak.
The consequences are asymmetric. A leak rarely makes a model a little better; it tends to make it dramatically better on paper and useless in production. That is why an ML result that looks too good deserves more suspicion than one that looks mediocre.
Leakage in finance comes in three families:
- Feature leakage. An input uses data from after the decision point.
- Label leakage. The target, or something correlated with it, sneaks into the inputs.
- Validation leakage. The split between training and test data lets information cross the boundary, even when every feature is clean.
The third family is the one most practitioners miss, and it is the central theme of Marcos López de Prado's Advances in Financial Machine Learning (Wiley, 2018). For the full pipeline around these ideas, start with our guide to machine learning trading strategies without overfitting.
Where does leakage come from in practice?
Most leaks are not exotic. They come from ordinary data-engineering shortcuts that are harmless in other domains.
| Leak source | Example | Fix |
|---|---|---|
| Resampling | A daily indicator joined onto hourly bars using the day's final value | Carry forward the last completed daily value only |
| Timestamps | A funding rate stamped at settlement but used during the hour before | Align every series to when it became known |
| Release dates | A macro series aligned to its reference month, not its publication date | Use release-time (point-in-time) vintages |
| Normalization | Z-scores or scalers fit on the full dataset, including the test period | Fit transforms on the training window only |
| The traded bar | Using bar t's close as a feature when fills happen at t's close | Match features to the fill convention |
| Label overlap | A training label whose horizon ends inside the test window | Purge overlapping training events |
| Serial correlation | Training data right next to the test block | Add an embargo gap |
| Shuffled k-fold | Random folds that train on the future to predict the past | Walk-forward or combinatorial purged CV |
Our article on look-ahead bias in crypto backtests walks through the feature-side cases with concrete examples. The rest of this article focuses on what is specific to ML: labels and validation.
Why are overlapping labels such a big leak?
Trading labels usually span several bars. A triple-barrier label, for instance, is decided by whichever of a profit target, a stop or a time limit is hit first, which might take 24 hours on hourly data. If one training event starts at hour 100 and resolves at hour 124, and your test block starts at hour 110, then the training label already "knows" what happened during the first 14 hours of the test period.
Two remedies work together:
- Purging drops every training event whose label interval overlaps the test window.
- Embargo removes a further buffer of training data right after the test window, because features and returns are serially correlated and information leaks across the edge.
scikit-learn's TimeSeriesSplit respects time order and has a gap parameter, but it does not know how long each label lives, so it cannot purge by label interval on its own. The purged walk-forward validation guide shows a worked example.
Overlap also causes a quieter problem that is not strictly leakage: redundant samples. Ten events that resolve on the same rally are one piece of evidence counted ten times. We cover the fix in overlapping labels and sample uniqueness.
Why is prevention not enough?
You can design a pipeline carefully and still leak. A new data source arrives with timestamps in a different convention. A feature is computed from a column that was itself derived from the label. A helper function fits a scaler before the split. None of these show up in code review reliably.
The practical answer is to treat leakage like a bug you test for, not a property you assume. The useful tests share one idea: construct a situation where a clean model must fail, and check that yours does.
Which leak checks actually matter?
Plenty of diagnostics exist. These four catch the most real leaks for the least effort.
1. Shuffled-label control
Permute the labels, keep everything else identical, retrain and evaluate out of sample. A model trained on scrambled targets has nothing to learn, so its accuracy must collapse to chance. If it still beats chance, information is reaching it through the pipeline itself: through features that encode the label, through fold boundaries, or through sample ordering.
One subtlety: "chance" is not always 50%. For an imbalanced target, such as a meta-label where most trades lose, chance is the majority-class rate. Comparing against 50% will either miss leaks or flag every run.
2. Embargo sensitivity
Widen the embargo and re-run. A genuine signal should be roughly indifferent to a slightly larger gap between train and test. If performance collapses when you widen it, the model was probably exploiting overlap across the boundary rather than a real relationship.
3. Single-feature dominance
Compute permutation importance (MDA). If one feature carries most of the total importance, inspect it closely. Future-derived columns and label-correlated columns tend to dominate because they carry "free" information. A dominant feature is not proof of a leak, but it is the first place to look. The feature importance article explains how to read MDA, SFI and SHAP together.
4. Implausibility ceiling
Set a ceiling above which a result is treated as suspicious rather than celebrated. Directional accuracy far above 60% on liquid crypto at intraday horizons, or an out-of-sample Sharpe far above the underlying strategy's own backtest, is much more likely to be a leak than an edge. It is an uncomfortable rule, because it can reject a true discovery, but on public market data the base rate of leaks is much higher than the base rate of 70%-accurate models.
How does Trigr handle leakage in ML optimization?
Trigr's "Optimize with ML" applies both layers: prevention in the pipeline and active detection afterwards, as documented on the AI and ML docs page.
Prevention:
- Features are built point-in-time from the strategy's own inputs. The feature loader raises an error on any future leakage it detects, rather than warning and continuing.
- The traded bar's close is dropped from the inputs, because the backtester fills at the next bar's open.
- Labels are volatility-scaled triple-barrier labels measured from the next bar's open, so the label matches how the backtest fills.
- Validation uses anchored, purged walk-forward splits with an embargo sized to at least the label horizon.
Detection: every run returns a leak verdict, realistic or not realistic, with specific flags. The hard checks are the four above:
- the shuffled-label control must land within 0.10 of chance, with chance defined as the majority-class rate for imbalanced meta-labels;
- widening the embargo must not cut out-of-sample performance below half its original level;
- no single feature may carry more than 60% of total MDA importance;
- directional accuracy above 0.65 (or, for a meta-label gate, more than 0.15 above the majority-class rate), or an out-of-sample Sharpe far above the plain strategy's own backtest Sharpe, is treated as too good to be real.
Soft warnings sit alongside: a Deflated Sharpe Ratio below 0.60 or a PBO above 0.5 flags a likely overfit, and fewer than 30 pooled out-of-sample events or fewer than three purged folds marks the evidence as indicative rather than conclusive.
What this means for you
You do not have to build the leak battery yourself, and you get a clear stop signal: if the verdict says not realistic, the headline Sharpe is not worth reading. You also see which check tripped, which usually points straight at the offending feature or setting.
What leakage checks cannot catch
Leak controls reduce risk; they do not certify a model. Be clear about the gaps:
- Selection across runs. If you run the optimizer twenty times on the same data and keep the best, each run can pass its leak battery while the collection is still overfit. Trigr's DSR counts the trials inside a run, not across separate reruns, so repeated tuning needs fresh data. See selection bias and the best of 100 backtests.
- Researcher leakage. If you designed the strategy after looking at the 2024 chart, the pipeline cannot know that. Only data you have not seen can test it.
- Regime change. A clean model can still meet a market unlike anything in its training window.
Backtests are not guarantees; perps are leveraged and can lose more than expected. Forward-test on a paper agent before committing capital.
Next steps
Pick one of your own strategies, run a standard backtest, then an ML optimization, and read the leak verdict before any other number. If you prefer working from an AI assistant, connect Claude, ChatGPT or Codex over MCP and run the same workflow asynchronously.