Backtest vs Live Results: Why They Differ and How to Fix

Backtest vs live results rarely match. The main causes of the gap, from costs and fills to selection bias, and practical steps to make them closer.

Trigr Research6 min read
On this page
  1. Why is there always a gap?
  2. What causes the execution gap?
  3. What causes the edge gap?
  4. How big is each source of the gap?
  5. How do you shrink the gap in practice?
  6. How does Trigr help you see the gap?
  7. Next steps

TL;DR: Live results differ from backtests because the backtest made assumptions the market does not honor: costs left out, fills that were too clean, a strategy chosen from many variants, and a regime that may have changed. Most of these errors point the same way, so live usually comes in lower. You shrink the gap by modeling costs explicitly, keeping the trial count low and visible, and collecting forward data before scaling up.

Why is there always a gap?

A backtest is a model of how a strategy would have traded. Like any model, it simplifies. The gap between backtest and live results is the sum of every simplification, plus the difference between the past sample and the future.

It helps to split the gap into two families:

  • Execution gap. The strategy's decisions were the same, but the prices, fees and timing you actually got were different from what the backtest assumed.
  • Edge gap. The decisions themselves were worse than the backtest suggested, because some of the backtest's edge was luck, leakage or a regime that ended.

The execution gap is mostly under your control. The edge gap is harder, but you can measure it. Backtests are not guarantees; perps are leveraged and can lose more than expected.

What causes the execution gap?

Costs the backtest left out

The most common cause is simply missing costs. Fees are usually included; slippage and funding often are not. For strategies with small per-trade edges, a few basis points per fill decide whether the strategy is profitable. Our guide to how trading fees decide whether a perp strategy works walks through the arithmetic.

On Trigr, trading fees and the builder fee are always applied. Slippage and funding are opt-in and off by default, so a first result is labeled gross. Turning them on is the single cheapest way to narrow the gap before any capital is involved.

Fills that were too clean

A backtest that fills at the signal bar's close, or at the exact stop price, assumes prices you could not have traded. Real orders fill after the decision, against an order book, sometimes after a gap. Trigr fills at the next bar's open after a signal and resolves same-bar take-profit and stop-loss conflicts conservatively, replaying 5-minute bars and letting the stop win if it is still ambiguous. See stop-loss and take-profit in the same candle for why that matters.

Even with honest conventions, a flat slippage assumption is not an order-book model. Trigr's slippage is a flat number of basis points per fill that you choose. It does not model spread, depth, market impact, latency or volatility, so size it for the markets and order sizes you intend to trade. The backtesting docs list what is and is not modeled.

Funding that differs from the source series

Perpetual futures charge or pay funding while a position is open. Trigr's historical funding in backtests comes from Binance's 8-hour USDT-M series, while Hyperliquid pays funding hourly at its own rates. The historical series is a reasonable approximation, not a record of what a Hyperliquid position paid. For strategies that hold for days, compare results under both historical and a more pessimistic flat rate.

Venue rules and minimums

Live venues have rules that a backtest rarely encodes. Hyperliquid rejects orders under $10 of notional, so Trigr raises live entries below $11 of notional to $11, and an agent with less than $11 of capital cannot open positions at all. Small accounts with many slots or small allocations can therefore trade differently from the backtest's proportional sizing.

What causes the edge gap?

Selection bias

If you tested fifty variants and deployed the best, some of its backtest performance came from being the maximum of fifty noisy results. That portion does not carry into the future. This is usually the largest single source of disappointment, and it is invisible unless you count your trials. We explain the statistics in the best of 100 backtests is probably luck.

Leakage

A feature that quietly used future information produces a backtest edge that cannot exist live. Point-in-time data is the defense: on Trigr, every feature uses only information available at the bar's close, and coarser or cross-asset inputs are carried forward, never read ahead.

Regime change

Markets change. A mean-reversion strategy tuned in a range can struggle in a trend, and a funding-based signal can weaken when positioning changes. This is not a bug in the backtest. It is the limit of what any historical sample can tell you.

How big is each source of the gap?

Source Direction Typical size How to reduce it
Missing slippage Live worse Large for frequent traders Enable flat slippage in the backtest
Missing or approximate funding Usually live worse for crowded trades Grows with holding time Enable funding; test a pessimistic flat rate
Unrealistic fills Live worse Large with tight exits Next-bar fills, conservative intrabar rules
Venue minimums Varies Matters for small accounts Size capital and allocations above minimums
Selection bias Live worse Grows with trial count Count trials, prefer deflated statistics
Leakage Live worse Can be total Point-in-time data, leak checks
Regime change Either Unpredictable Forward-test, monitor, set pause rules

The "typical size" column is qualitative on purpose. The honest answer to "how big is the gap?" is that it depends on your trade frequency, order size and research process, which is why measuring it on your own strategy matters more than any rule of thumb.

How do you shrink the gap in practice?

A sequence that works for most systematic traders:

  1. Backtest net, not gross. Turn on slippage and funding before judging an idea. If it only works gross, stop there.
  2. Keep the trial count visible. Iterate on one idea as labeled experiments inside one strategy instead of spawning many near-duplicates. On Trigr, the Copilot works this way, and MCP clients pass an experimentLabel with an owned strategy.
  3. Get a second statistical opinion. An Optimize with ML run reports out-of-sample Sharpe from anchored, purged walk-forward validation, the Deflated Sharpe Ratio of Bailey and López de Prado using the run's trial count, and PBO. A standard backtest does not compute DSR or PBO.
  4. Freeze, then forward-test. Run the unchanged strategy on a paper agent. Forward data was not available when you chose the strategy, so it is free of selection bias.
  5. Deploy small, then compare. Start a live agent with modest capital and compare its forward record with both the backtest and the paper record.

Studio's shaded trailing 25% helps with step 1 as a recent-period diagnostic, but it is not an out-of-sample holdout. Standard backtests use full available history.

How does Trigr help you see the gap?

Trigr is built as one path from backtest to paper to live, which makes the comparison direct rather than a spreadsheet exercise.

  • Paper agents fill at the current Hyperliquid price and deduct a 4.5 bps taker fee plus the builder fee. They model no slippage and no funding, so treat paper as a middle step: it tests decisions on new data, not full execution.
  • Agent setup shows a combined backtest for the agent's strategy slots, and a frontier marker separates the backtest from the forward record, so you can see exactly where history ends and real results begin.
  • Pinned versions mean the strategy running in the agent is the exact version you tested, not one that changed after you looked. See why your trading bot should run a pinned strategy version.
  • Alerts by email and web push let you notice when the forward record departs from what the backtest led you to expect.

What this means for you

You will still see a gap, because no backtest is the market. The difference is that you can attribute it. If the paper record tracks the backtest but live results lag, the gap is execution, and slippage or sizing is the place to look. If paper already lags the backtest, the edge itself was weaker than it looked, and more tuning on the same history will not fix that.

Next steps

Pick a strategy you are considering deploying, rerun it with slippage and funding on, then follow the steps in deploying a strategy as a Hyperliquid agent, starting with a paper agent. You can browse verified, versioned strategies on the strategy marketplace to see how each one's verified backtest is presented before you adopt it.

Frequently asked questions

Why are my live trading results worse than my backtest?

The usual causes are costs the backtest left out (slippage, funding), fills that differ from the backtest's assumptions, selection bias from trying many variants, and a market regime that changed. Most of these push in the same direction, so live results tend to come in below the backtest.

How big should I expect the gap between backtest and live results to be?

There is no fixed number. It depends on trade frequency, order size relative to liquidity, how many variants you tried and which costs the backtest included. Strategies that trade often with thin per-trade edges usually show the largest gap.

Does paper trading close the gap?

Paper trading removes selection bias because it produces genuinely new data after the strategy is frozen. On Trigr, paper agents fill at the current Hyperliquid price and deduct fees, but they model no slippage and no funding, so they sit between a backtest and a live account.

Is the trailing 25% in a Trigr backtest an out-of-sample test?

No. It is a shaded recent-period diagnostic for spotting degradation. Once you look at it while choosing a strategy, it is part of your research data. True out-of-sample evidence comes from ML walk-forward results or a forward paper record.

Put the idea to an honest test.

Describe a strategy in plain English or from your own AI assistant, backtest it on point-in-time data, and forward-test it on paper before any real money is involved.