Sample Uniqueness and Overlapping Labels in Financial ML

Sample uniqueness in financial machine learning fixes overlapping labels that count one market move many times. How average uniqueness works, with examples.

Trigr Research6 min read
On this page
  1. Why do labels overlap in the first place?
  2. What goes wrong if you ignore overlap?
  3. How is average uniqueness calculated?
  4. How do you use uniqueness in training?
  5. Weighting versus dropping overlaps
  6. How does Trigr handle overlapping labels?
  7. Practical checklist
  8. Next steps

TL;DR: In trading ML, labels usually span several bars, so many of them overlap in time and share the same outcome. Treating them as independent samples overcounts a few market moves and makes both the model and its statistics overconfident. López de Prado's fix is average uniqueness: weight each label by how little it overlaps with others, so redundant events count for less without being thrown away.

Why do labels overlap in the first place?

Most useful trading labels are not "did the next bar go up". They describe how a trade would resolve: did price reach a profit target before a stop, or did the time limit expire first? That is the triple-barrier method, and a label built this way lives from the event's start until the first barrier is touched.

If your strategy can signal on consecutive bars, labels pile up. Picture hourly bars and a strategy whose entries fire at 09:00, 10:00, 11:00 and 12:00 during a rally. With a 24-hour time barrier, all four labels are still open at 13:00, and all four may resolve on the same move at 15:00. You have four rows in your training set but roughly one observation of how the market behaved.

This is normal. It is also exactly where standard machine learning assumptions break: most estimators treat rows as independent and identically distributed (IID).

What goes wrong if you ignore overlap?

Three things, each worse than it sounds.

  • The model overweights clustered regimes. Periods where the strategy fires often (usually trending or volatile ones) dominate training, not because they are more informative but because they produced more overlapping rows.
  • Bagging and bootstrapping lose diversity. A random forest draws bootstrap samples expecting them to differ. With heavy overlap, many "different" draws contain nearly the same information, so the trees end up correlated and the ensemble behaves less like an ensemble.
  • Confidence statistics inflate. A Sharpe ratio computed over 1,000 labels that are really 150 independent outcomes looks far more significant than it is. Any test that uses sample size, including the Deflated Sharpe Ratio, is affected.

Overlap is also a leakage channel across validation folds. That part is handled by purging and embargoing, covered in data leakage in trading ML. Sample uniqueness handles the within-training-set problem.

How is average uniqueness calculated?

The method comes from Chapter 4 of López de Prado's Advances in Financial Machine Learning (Wiley, 2018). It takes three steps.

  1. Count concurrency. For each bar t, count how many labels are active, meaning the bar falls between a label's start and the time its first barrier was touched. Call this c(t).
  2. Compute per-bar uniqueness. For a label i active at bar t, its uniqueness on that bar is 1 / c(t). If three labels share the bar, each gets one third of it.
  3. Average over the label's life. A label's average uniqueness is the mean of 1 / c(t) over every bar it spans.

The result is a number between 0 and 1. A label that never overlaps has uniqueness 1. A label that shares every bar with nine others has uniqueness around 0.1.

A worked example

Take four hourly labels during the rally described above, with the time barrier shortened to three bars so the table stays readable:

Label Active bars Concurrency on those bars Average uniqueness
A 09:00 to 11:00 (3 bars) 1, 2, 3 (1 + 1/2 + 1/3) / 3 ≈ 0.61
B 10:00 to 12:00 (3 bars) 2, 3, 3 (1/2 + 1/3 + 1/3) / 3 ≈ 0.39
C 11:00 to 13:00 (3 bars) 3, 3, 2 (1/3 + 1/3 + 1/2) / 3 ≈ 0.39
D 12:00 to 14:00 (3 bars) 3, 2, 1 (1/3 + 1/2 + 1) / 3 ≈ 0.61

The concurrency counts assume only these four labels exist in the window. Together they sum to about 2.0, so the four rows carry roughly the weight of two independent observations. Labels in the middle of the cluster, which share the most with their neighbours, count the least.

Sum average uniqueness across all labels and you get a rough effective sample size: an estimate of how many independent outcomes your dataset really contains. It is often a sobering number.

How do you use uniqueness in training?

There are three common ways to put the measure to work.

Sample weights

The simplest and most widely used approach: pass average uniqueness as the sample_weight argument when fitting. Most tree libraries, including scikit-learn's RandomForestClassifier, LightGBM and XGBoost, accept per-row weights at fit time. Redundant labels still contribute, but proportionally less.

López de Prado also suggests return attribution: scale each weight by the absolute return attributed to the label over its life, adjusted for concurrency. The intuition is that a label resolving on a large move carries more information than one resolving on noise. Uniqueness times return attribution down-weights both redundant events and trivial ones.

Sequential bootstrap

Standard bootstrapping draws rows uniformly at random. The sequential bootstrap draws them one at a time, updating each remaining row's probability so that rows overlapping heavily with those already drawn become less likely. The resulting samples are closer to independent, which helps bagging ensembles such as random forests. It is more expensive than a normal bootstrap, so many practitioners approximate it by setting the forest's max_samples near the average uniqueness.

Effective sample size for statistics

When you compute a Sharpe ratio, a t-statistic or a Deflated Sharpe Ratio over overlapping outcomes, use the effective sample size rather than the raw row count. Otherwise the statistic claims more evidence than you have. The Deflated Sharpe Ratio explainer shows why sample length matters so much to that calculation.

Weighting versus dropping overlaps

A tempting shortcut is to keep only non-overlapping events, for example by allowing a new label only after the previous one resolves. It removes the problem, but it has costs:

Approach Keeps all events Result depends on arbitrary choice Handles varying overlap Typical use
Treat as IID Yes No No Never, for multi-bar labels
Drop overlapping events No Yes, which event you keep Crudely Quick sanity checks
Weight by average uniqueness Yes No Yes Default for training
Sequential bootstrap Yes Partly (random) Yes Bagging ensembles

Dropping events also changes what the model learns about. If your strategy fires in clusters, the kept events are always the first in each cluster, which may behave differently from the rest.

How does Trigr handle overlapping labels?

Trigr's "Optimize with ML" learns a take-or-skip gate for a strategy you already built. Its labels are volatility-scaled triple-barrier labels on the strategy's own entries, measured from the next bar's open to match how the backtester fills, so overlap is expected whenever entries cluster.

The pipeline handles it in two places:

  • Training weights. Each label is weighted by its average uniqueness multiplied by its absolute return attribution, so redundant events and trivially small moves both count for less (see the AI and ML docs).
  • Statistics. The same uniqueness adjustment is applied when estimating the effective sample size behind the Sharpe statistics, including the Deflated Sharpe Ratio. Overlapping events cannot quietly inflate confidence.

Combined with purged, embargoed walk-forward validation, this means the out-of-sample figures you read are based on an honest count of independent outcomes, not on the number of rows your strategy happened to produce.

What this means for you

A strategy that fires many times in a row during trends does not get an artificially confident ML result just because it produced many labels. And when a run reports low statistical power, the warning reflects the effective evidence, not the raw trade count. You can confirm the picture yourself by reading the trade log from the underlying backtest and noticing how entries cluster.

Practical checklist

Before training any model on multi-bar labels:

  • Record each label's start and end time, not just its start.
  • Compute concurrency and average uniqueness; look at the distribution, not just the mean.
  • Pass uniqueness (optionally times return attribution) as sample weights.
  • Purge and embargo across validation folds.
  • Report effective sample size next to any Sharpe or accuracy figure.
  • If average uniqueness is very low, consider fewer, more distinct entries in the strategy itself.

Backtests are not guarantees; perps are leveraged and can lose more than expected.

Next steps

For the rest of the pipeline that surrounds these weights, read machine learning trading strategies without overfitting, then try an ML run on your own strategy. You can also drive it from an AI assistant after you connect it to Trigr over MCP.

Frequently asked questions

What is sample uniqueness in financial machine learning?

Sample uniqueness measures how much of a label's outcome is shared with other labels active at the same time. A label that overlaps with nine others on every bar of its life has an average uniqueness of about 0.1; an isolated label has a uniqueness of 1.

Why do overlapping labels cause problems?

They violate the assumption that training samples are independent. The model sees one market move counted many times, becomes overconfident, and statistics such as the Sharpe ratio overstate the true sample size.

Should I just drop overlapping events instead of weighting them?

Dropping overlaps throws away information and makes results depend on which event you keep. Weighting by average uniqueness keeps every event but lets redundant ones count for less, which is the approach López de Prado recommends.

How does Trigr weight overlapping samples?

Trigr's ML optimization weights each triple-barrier label by its average uniqueness multiplied by the absolute return attributed to it, and uses the same uniqueness adjustment when estimating the effective sample size behind its Sharpe statistics.

Put the idea to an honest test.

Describe a strategy in plain English or from your own AI assistant, backtest it on point-in-time data, and forward-test it on paper before any real money is involved.