Data Leakage in Trading ML: The Checks That Matter

Data leakage in machine learning trading models makes backtests look brilliant and live trading fail. Learn the leak sources and the checks that catch them.

Trigr Research7 min read
On this page
  1. What is data leakage in trading machine learning?
  2. Where does leakage come from in practice?
  3. Why are overlapping labels such a big leak?
  4. Why is prevention not enough?
  5. Which leak checks actually matter?
  6. How does Trigr handle leakage in ML optimization?
  7. What leakage checks cannot catch
  8. Next steps

TL;DR: Data leakage is information from the future reaching a trading model through its features, labels or validation splits. It is the most common reason an ML strategy with a great out-of-sample score fails live. Prevention (point-in-time features, purged splits) is necessary but not enough; you also need active checks that try to prove a result is contaminated, such as a shuffled-label control and an embargo sensitivity test.

What is data leakage in trading machine learning?

In general machine learning, leakage means the training process sees information it would not have at prediction time. In trading, "prediction time" has a precise meaning: the moment your strategy decides whether to act, usually the close of a bar. Anything the model uses that was not knowable at that moment is a leak.

The consequences are asymmetric. A leak rarely makes a model a little better; it tends to make it dramatically better on paper and useless in production. That is why an ML result that looks too good deserves more suspicion than one that looks mediocre.

Leakage in finance comes in three families:

  • Feature leakage. An input uses data from after the decision point.
  • Label leakage. The target, or something correlated with it, sneaks into the inputs.
  • Validation leakage. The split between training and test data lets information cross the boundary, even when every feature is clean.

The third family is the one most practitioners miss, and it is the central theme of Marcos López de Prado's Advances in Financial Machine Learning (Wiley, 2018). For the full pipeline around these ideas, start with our guide to machine learning trading strategies without overfitting.

Where does leakage come from in practice?

Most leaks are not exotic. They come from ordinary data-engineering shortcuts that are harmless in other domains.

Leak source Example Fix
Resampling A daily indicator joined onto hourly bars using the day's final value Carry forward the last completed daily value only
Timestamps A funding rate stamped at settlement but used during the hour before Align every series to when it became known
Release dates A macro series aligned to its reference month, not its publication date Use release-time (point-in-time) vintages
Normalization Z-scores or scalers fit on the full dataset, including the test period Fit transforms on the training window only
The traded bar Using bar t's close as a feature when fills happen at t's close Match features to the fill convention
Label overlap A training label whose horizon ends inside the test window Purge overlapping training events
Serial correlation Training data right next to the test block Add an embargo gap
Shuffled k-fold Random folds that train on the future to predict the past Walk-forward or combinatorial purged CV

Our article on look-ahead bias in crypto backtests walks through the feature-side cases with concrete examples. The rest of this article focuses on what is specific to ML: labels and validation.

Why are overlapping labels such a big leak?

Trading labels usually span several bars. A triple-barrier label, for instance, is decided by whichever of a profit target, a stop or a time limit is hit first, which might take 24 hours on hourly data. If one training event starts at hour 100 and resolves at hour 124, and your test block starts at hour 110, then the training label already "knows" what happened during the first 14 hours of the test period.

Two remedies work together:

  • Purging drops every training event whose label interval overlaps the test window.
  • Embargo removes a further buffer of training data right after the test window, because features and returns are serially correlated and information leaks across the edge.

scikit-learn's TimeSeriesSplit respects time order and has a gap parameter, but it does not know how long each label lives, so it cannot purge by label interval on its own. The purged walk-forward validation guide shows a worked example.

Overlap also causes a quieter problem that is not strictly leakage: redundant samples. Ten events that resolve on the same rally are one piece of evidence counted ten times. We cover the fix in overlapping labels and sample uniqueness.

Why is prevention not enough?

You can design a pipeline carefully and still leak. A new data source arrives with timestamps in a different convention. A feature is computed from a column that was itself derived from the label. A helper function fits a scaler before the split. None of these show up in code review reliably.

The practical answer is to treat leakage like a bug you test for, not a property you assume. The useful tests share one idea: construct a situation where a clean model must fail, and check that yours does.

Which leak checks actually matter?

Plenty of diagnostics exist. These four catch the most real leaks for the least effort.

1. Shuffled-label control

Permute the labels, keep everything else identical, retrain and evaluate out of sample. A model trained on scrambled targets has nothing to learn, so its accuracy must collapse to chance. If it still beats chance, information is reaching it through the pipeline itself: through features that encode the label, through fold boundaries, or through sample ordering.

One subtlety: "chance" is not always 50%. For an imbalanced target, such as a meta-label where most trades lose, chance is the majority-class rate. Comparing against 50% will either miss leaks or flag every run.

2. Embargo sensitivity

Widen the embargo and re-run. A genuine signal should be roughly indifferent to a slightly larger gap between train and test. If performance collapses when you widen it, the model was probably exploiting overlap across the boundary rather than a real relationship.

3. Single-feature dominance

Compute permutation importance (MDA). If one feature carries most of the total importance, inspect it closely. Future-derived columns and label-correlated columns tend to dominate because they carry "free" information. A dominant feature is not proof of a leak, but it is the first place to look. The feature importance article explains how to read MDA, SFI and SHAP together.

4. Implausibility ceiling

Set a ceiling above which a result is treated as suspicious rather than celebrated. Directional accuracy far above 60% on liquid crypto at intraday horizons, or an out-of-sample Sharpe far above the underlying strategy's own backtest, is much more likely to be a leak than an edge. It is an uncomfortable rule, because it can reject a true discovery, but on public market data the base rate of leaks is much higher than the base rate of 70%-accurate models.

How does Trigr handle leakage in ML optimization?

Trigr's "Optimize with ML" applies both layers: prevention in the pipeline and active detection afterwards, as documented on the AI and ML docs page.

Prevention:

  • Features are built point-in-time from the strategy's own inputs. The feature loader raises an error on any future leakage it detects, rather than warning and continuing.
  • The traded bar's close is dropped from the inputs, because the backtester fills at the next bar's open.
  • Labels are volatility-scaled triple-barrier labels measured from the next bar's open, so the label matches how the backtest fills.
  • Validation uses anchored, purged walk-forward splits with an embargo sized to at least the label horizon.

Detection: every run returns a leak verdict, realistic or not realistic, with specific flags. The hard checks are the four above:

  • the shuffled-label control must land within 0.10 of chance, with chance defined as the majority-class rate for imbalanced meta-labels;
  • widening the embargo must not cut out-of-sample performance below half its original level;
  • no single feature may carry more than 60% of total MDA importance;
  • directional accuracy above 0.65 (or, for a meta-label gate, more than 0.15 above the majority-class rate), or an out-of-sample Sharpe far above the plain strategy's own backtest Sharpe, is treated as too good to be real.

Soft warnings sit alongside: a Deflated Sharpe Ratio below 0.60 or a PBO above 0.5 flags a likely overfit, and fewer than 30 pooled out-of-sample events or fewer than three purged folds marks the evidence as indicative rather than conclusive.

What this means for you

You do not have to build the leak battery yourself, and you get a clear stop signal: if the verdict says not realistic, the headline Sharpe is not worth reading. You also see which check tripped, which usually points straight at the offending feature or setting.

What leakage checks cannot catch

Leak controls reduce risk; they do not certify a model. Be clear about the gaps:

  • Selection across runs. If you run the optimizer twenty times on the same data and keep the best, each run can pass its leak battery while the collection is still overfit. Trigr's DSR counts the trials inside a run, not across separate reruns, so repeated tuning needs fresh data. See selection bias and the best of 100 backtests.
  • Researcher leakage. If you designed the strategy after looking at the 2024 chart, the pipeline cannot know that. Only data you have not seen can test it.
  • Regime change. A clean model can still meet a market unlike anything in its training window.

Backtests are not guarantees; perps are leveraged and can lose more than expected. Forward-test on a paper agent before committing capital.

Next steps

Pick one of your own strategies, run a standard backtest, then an ML optimization, and read the leak verdict before any other number. If you prefer working from an AI assistant, connect Claude, ChatGPT or Codex over MCP and run the same workflow asynchronously.

Frequently asked questions

What is data leakage in a trading machine learning model?

Data leakage is any path by which information that would not have been available at decision time reaches the model's features, labels or validation data. It makes out-of-sample results look far better than anything the model can achieve live.

How can I tell if my trading model is leaking?

Run active controls: retrain on shuffled labels and confirm accuracy falls to chance, widen the embargo and confirm performance holds, check whether one feature dominates importance, and treat implausibly high accuracy as a red flag rather than a win.

Is shuffled k-fold cross-validation a leak?

For time series with overlapping labels, usually yes. Shuffling lets the model train on bars that come after, and overlap with, the bars it is tested on. Use purged, embargoed walk-forward or combinatorial splits instead.

Does Trigr check ML runs for leakage?

Yes. Every ML optimization run returns a leak verdict, realistic or not realistic, with specific flags from a shuffled-label control, an embargo sensitivity test, a single-feature dominance check and an implausibility ceiling, alongside soft warnings for weak DSR, high PBO or low statistical power.

Put the idea to an honest test.

Describe a strategy in plain English or from your own AI assistant, backtest it on point-in-time data, and forward-test it on paper before any real money is involved.