TL;DR: Out-of-sample testing means evaluating a strategy on data that had no influence on building, choosing or tuning it. A single holdout works only if you reserve it in advance and test once. Walk-forward validation gives more robust out-of-sample evidence for rule-fitting and ML, and a forward paper test is the only data that is guaranteed unseen. On Trigr, standard backtests use full history with no holdout; out-of-sample evidence comes from ML walk-forward folds or from paper agents.
What does "out of sample" really mean?
In-sample data is anything that shaped your strategy. Out-of-sample data is everything that did not. The definition is about your decisions, not about a label in the software.
That distinction catches many traders. Suppose you split 2020 to 2026 into a 2020–2024 "training" period and a 2025–2026 "test" period. You build a strategy on the first part, check the second, dislike the result, change the RSI threshold, and check again. The test period has now influenced your strategy. After a few rounds it is just more training data, and the final "out-of-sample" result is optimistically biased.
This is selection bias in slow motion. Each look is a trial, and the best of many looks is expected to overstate the truth. The article on selection bias and the best of 100 backtests explains the arithmetic.
What are the main ways to test out of sample?
There are four main approaches, each with a different job.
| Method | What it is | Strength | Weakness |
|---|---|---|---|
| Single holdout | Reserve a final period, test the frozen strategy once | Simple, easy to explain | One regime, one shot; ruined by repeated peeking |
| Walk-forward validation | Train on the past, test the next block, roll forward | Many out-of-sample periods; fits ML and rule-fitting | Needs purging and embargo for multi-bar labels |
| Combinatorial purged CV | Many train/test combinations of blocks | Distribution of outcomes; enables PBO | More complex; still uses the same history |
| Forward (paper) test | Run the frozen strategy on data that did not exist yet | Truly unseen; exercises the live pipeline | Slow; short samples are noisy |
None of these replaces the others. Walk-forward and combinatorial methods reduce overfitting during research; a forward test confirms the result after research is over.
How do you run a single holdout honestly?
A holdout is valid only under strict rules. The Trigr backtesting docs summarize them well:
- Reserve the period before research begins. Decide the cut-off date first and do not look at charts or results beyond it.
- Freeze one winner. Choose a single strategy, its parameters and every cost assumption: fees, slippage and funding.
- Test once. Run the frozen strategy on the untouched period and record the result.
- Do not repair against it. If it fails, do not tune on the holdout and do not substitute the runner-up. Either the idea is rejected or you need new data.
The fourth rule is the painful one, and the reason holdouts are so often broken in practice. Once you have seen the holdout, you cannot un-see it.
Why is walk-forward validation usually better?
A single holdout gives you one number from one regime. If the holdout happens to be a strong bull trend, a long-only momentum strategy passes whether or not it has an edge.
Walk-forward validation trains only on data before each test block, evaluates on that block, then moves forward. You end up with out-of-sample predictions spread across many periods, pooled into one honest estimate. In an anchored (expanding) setup, each training window starts at the beginning of history and grows; in a rolling setup, it has a fixed length.
scikit-learn's TimeSeriesSplit implements the basic expanding version. For trading labels that span several bars, you also need:
- Purging: remove training events whose label interval overlaps the test block.
- Embargo: drop a buffer of training data after each test block, because serial correlation leaks information across the edge.
Without both, the walk-forward result is contaminated by exactly the overlap it is meant to avoid. The purged walk-forward validation guide shows how, with a worked example.
What about many models and combinations?
Walk-forward tells you how one procedure performs out of sample. If you compare many configurations, you still need to account for the search. Two tools from Bailey and López de Prado's work help:
- The Deflated Sharpe Ratio adjusts the Sharpe ratio for the number of trials, the sample length, skewness and fat tails.
- The Probability of Backtest Overfitting, estimated with combinatorially symmetric cross-validation, measures how often the in-sample favorite lands in the bottom half out of sample.
Why is a forward test the final word?
Every method above reuses historical data you already had. Even with perfect discipline, you chose the idea, the market and the timeframe while living through that history. Forward testing removes that problem: the data did not exist when you froze the strategy.
A paper forward test also exercises real-time data arrival and the agent logic you will actually run; order routing and real fills are only tested once a small live agent runs. The gaps between a backtest and live trading are covered in why live results differ from your backtest.
Its weakness is sample size. A few weeks of forward trades is a small, noisy sample. Treat it as a check that the strategy behaves as expected, not as proof of an edge, and let it run long enough to include varied conditions.
How does out-of-sample testing work on Trigr?
Precision matters here, because it is easy to overstate.
Standard backtests have no holdout. A normal Studio backtest runs over the full available history of the market. Studio shades the trailing 25% of the curve and compares it with the earlier part, which is useful for spotting recent degradation. It is not an out-of-sample test: you see it while you research, so it becomes in-sample the moment it influences a decision. The app's standard backtest also does not accept arbitrary date windows, so you cannot carve out a private holdout inside it. (A bounded date-window backtest exists over the MCP server and API for detached research, but it is not saved to Studio history and does not make a period unseen.)
ML optimization produces genuine walk-forward out-of-sample results. "Optimize with ML" selects among LightGBM, random forest and XGBoost using anchored, purged walk-forward out-of-sample predictions with an embargo, and reports the out-of-sample Sharpe, a Deflated Sharpe Ratio using the run's trial count, a PBO from a separate CPCV grid and a leak verdict. One caveat: separate optimizer reruns are not yet accumulated into the trial count, so repeatedly re-optimizing the same strategy erodes the meaning of the out-of-sample figures.
Paper agents provide the forward test. Freeze a strategy version, then deploy it to a paper agent. Paper agents fill at the current Hyperliquid price and deduct a 4.5 bps taker fee plus the builder fee; they do not model slippage or funding. The agent view marks a frontier between the backtest and the forward record, so you can see exactly where unseen data begins. The trading agents docs explain the setup, and paper trading agents covers how to read the results.
What this means for you
You always know which kind of evidence you are looking at: full-history research (a backtest), walk-forward out-of-sample (an ML run), or truly unseen data (a paper forward record). Published strategy versions are immutable and runners pin the version they adopt, so a forward record on a published strategy belongs to exactly the version that was tested rather than to a recipe that kept changing.
A practical out-of-sample workflow
- Build the strategy and run a standard backtest with slippage and funding enabled, so you see a net result.
- Iterate as labelled experiments within one strategy, keeping the number of trials visible.
- If you use ML, read the leak verdict first, then the out-of-sample Sharpe, DSR and PBO.
- Freeze one winner. Stop editing.
- Run it on a paper agent for long enough to see a range of conditions.
- Compare the forward record with the backtest's typical behavior: trade frequency, win rate, drawdown shape.
Backtests are not guarantees; perps are leveraged and can lose more than expected.
Next steps
For the full research pipeline that surrounds these tests, read machine learning trading strategies without overfitting. When a strategy is frozen, browse verified examples on the strategy marketplace or deploy your own to a paper agent.