Iterate on a Trading Strategy With AI Using Experiments

How to iterate on a trading strategy with AI: labelled experiments inside one strategy's version history, not dozens of near-duplicate strategies.

Trigr Research7 min read
On this page
  1. Why does AI iteration turn into strategy sprawl?
  2. What is the experiment model?
  3. How does the Studio Copilot run experiments?
  4. How do assistants run experiments over MCP?
  5. Why is this better than 40 near-duplicate strategies?
  6. What does a disciplined experiment session look like?
  7. What this means for you
  8. Next steps

TL;DR: Keep one strategy per distinct idea and run every parameter, filter, exit or sizing change as a labelled experiment inside it. In Trigr, the Studio Copilot does this automatically (Experiment 1, 2, 3...), and assistants connected over MCP do it with experimentLabel. You get one comparable history, a visible trial count, no quota clutter, and a clean way to promote the winner.

Why does AI iteration turn into strategy sprawl?

An AI assistant that can build and backtest in seconds changes the economics of research. Trying a variant used to cost an evening; now it costs a sentence. That is useful, and it is also how research goes wrong.

The common failure looks like this. You ask for a BTC trend strategy, the assistant saves it, then you ask for a tighter stop, and it saves "BTC Trend v2". A faster EMA becomes "v3", an added funding filter becomes "v3 funding", and so on. After an afternoon you have 40 strategies that are really one idea with 40 settings.

That sprawl causes three concrete problems:

  • Selection bias disappears from view. The best of 40 variants looks better than the idea deserves, and nothing on screen reminds you that 39 others were tried.
  • Comparison gets hard. Results are scattered across 40 cards with slightly different names, so it is easy to compare the wrong pair.
  • Quotas and clutter. Saved-strategy limits are real, and a library full of near-duplicates makes the genuinely different ideas harder to find.

What is the experiment model?

The fix is a simple rule: one strategy per idea, experiments for everything else. A strategy is defined by its market, timeframe and signal family. Anything that tunes that idea is an experiment.

Change New strategy or experiment?
BTC to ETH New strategy (different market)
4H to 1D New strategy (different timeframe)
EMA crossover to funding-rate fade New strategy (different signal family)
EMA 20/50 to EMA 12/26 Experiment
Add or remove an ADX filter Experiment
Fixed stop to 2.5x ATR trailing stop Experiment
Size 20% to 10% of capital Experiment

Each experiment is a full backtest of a variant, recorded with a short label in the strategy's version history. The saved recipe, and the verified statistics attached to it, stay untouched until you decide a variant has earned the promotion.

How does the Studio Copilot run experiments?

In Trigr Studio, the AI copilot iterates this way by default. When you ask it to build or improve a strategy, it applies an edit to the node graph, runs a real backtest, and records the result as a labelled experiment: Experiment 1, Experiment 2, Experiment 3, and so on. You can restore any version from the history.

The copilot is instructed to narrate each step: one or two lines on what it is changing and why before a run, then the numbers and whether the hypothesis survived after it. Every few experiments it posts a short scoreboard of the variants so far. At the end it leaves the best-measured configuration on the canvas and names the winning experiment by its label, so what you read in the chat matches what the builder shows you.

Because the history is stored with the strategy, a follow-up session, even on another device, sees the copilot's recent experiments and their numbers. The copilot does not forget that it already tried the ADX filter yesterday. For more on how the copilot turns plain English into a graph, see from plain English to a backtested strategy.

How do assistants run experiments over MCP?

If you use ChatGPT, Claude, Claude Code or Codex through Trigr's MCP server, the same model is available through tools exposed with the Model Context Protocol. The server's instructions tell connected agents to keep one strategy per distinct idea and treat parameter, filter, exit or sizing changes as experiments. The loop has four steps:

  1. Create once. The assistant saves the idea with trigr_create_strategy and keeps its strategyId.
  2. Run each variant as an experiment. It calls trigr_run_backtest with that strategyId and an experimentLabel such as "Exp 3 - ATR trail 2.5x". The variant graph may differ from the saved recipe; the run is recorded in the strategy's version history and never replaces the recipe or its verified statistics.
  3. Compare. trigr_get_backtest_result lists the strategy's recent runs, up to 20 per call, each with its label, an experiment flag, its settings and isSavedRecipe, which says whether that run measured the recipe currently saved. trigr_get_backtest_trades pages through any run's trades for free. Copilot experiments on the same strategy appear in the same list.
  4. Promote the winner. The assistant saves the winning variant with trigr_edit_strategy, passing the current recipe hash so a concurrent edit fails safely instead of being overwritten. That clears the old statistics; one unlabelled backtest of the new recipe then records fresh verified numbers.

A few guardrails keep this honest. Without a label, a backtest that names a strategyId must match the saved recipe exactly, or it is rejected as stale. A label without a strategyId, or on the bounded custom-window endpoint, is rejected before any credits are spent. Experiments are full-history only, because bounded-window runs are never stored. In manual approval mode, the browser confirmation shows the experiment label and states that the recipe is unchanged. Full argument details are in the agents and ML API reference.

Why is this better than 40 near-duplicate strategies?

The experiment model is not just tidier. It changes what you can know about your own research.

The trial count stays visible

The single most important number in strategy research is one people rarely record: how many variants were tried before the winner appeared. Bailey and López de Prado's Deflated Sharpe Ratio exists precisely because the maximum Sharpe ratio across many trials is inflated by chance, and the more trials, the larger the inflation.

When every variant is an experiment inside one strategy, the history is the trial log. You can count the rows. If the "winning" Experiment 14 beat Experiment 13 by a hair, you see that too. Spread across 40 separately named strategies, the same information is effectively lost. The article on selection bias in backtests goes deeper on why the best of many runs is often luck.

To be precise about what Trigr computes: a standard backtest does not report DSR or PBO. An ML optimization run does, using that run's own trial count. The experiment history does not deflate anything automatically; it keeps the count in front of you so you can apply that discount yourself.

Comparisons are like for like

All experiments share the strategy's market, timeframe and history, and each listed run shows the settings it used, including slippage, funding and the builder fee. That makes it much harder to compare a net run against a gross one by accident.

The saved recipe means something

Because experiments never overwrite the recipe or its verified statistics, the saved version always reflects a deliberate decision. That matters later: marketplace publishing requires a stored verified full-history backtest, and agents bind to an exact recipe or an immutable published version.

Quota and credits go further

Experiments do not count against your saved-strategy limit (10 on Free, 100 on Trader, 500 on Pro). They do cost credits: each experiment is a standard backtest at 50 credits, where 1,000 credits equal $1, and cache hits are free. Reading stored runs and trade logs costs nothing.

What does a disciplined experiment session look like?

A good session is short and pre-registered. Before the first run, write down the hypothesis, the costs you will assume, and a trial budget, for example "at most eight experiments". The AI strategy prompt recipes include wording for this.

Then:

  • Baseline first. Run the most direct version of the idea and record it as Experiment 1.
  • One attributable change per experiment. If you change the stop and the filter together, you will not know which one helped.
  • Measure net, not just gross. Fees and the builder fee are always applied, but slippage and funding are opt-in and off by default. Add them before you rank variants.
  • Read the trade log of the leader. If most of the profit comes from one month or three trades, the ranking is fragile.
  • Stop at the budget. Exhausting the budget without a convincing result is a valid outcome, and the history records it honestly.

Remember that the trailing 25% Studio shades in a backtest is a recent-period diagnostic, not a holdout. For evidence the selection process has not seen, use ML optimization's purged walk-forward results or forward-test the frozen winner on a paper agent. Backtests are not guarantees; perps are leveraged and can lose more than expected.

What this means for you

You can let an assistant iterate quickly without losing the audit trail. One strategy holds the whole story of an idea: every variant, its label, its settings, its numbers and which one you chose. That gives you a realistic sense of how hard you searched, fair comparisons, a library that stays readable, and a saved recipe that reflects a deliberate decision rather than whatever the last run happened to be.

Next steps

Connect ChatGPT, Claude, Claude Code or Codex through the guided MCP setup, or open a strategy in Studio and ask the copilot to run a small, budgeted set of experiments.

Frequently asked questions

What is an experiment in Trigr?

An experiment is a labelled backtest of a variant of one saved strategy. It is recorded in that strategy's version history with its label and results, but it does not change the saved recipe or its current verified statistics until you promote it.

When should I create a new strategy instead of an experiment?

Create a new strategy only for a genuinely different idea: a different market, timeframe or signal family. Changing parameters, filters, exits or sizing on the same idea is an experiment inside the existing strategy.

Do experiments count against my saved-strategy limit?

No. Experiments live inside one strategy's version history, so they do not use up your saved-strategy quota. Each experiment backtest still spends credits like any standard backtest, and cache hits are free.

How does an AI assistant label an experiment over MCP?

It calls trigr_run_backtest with the owned strategyId and an experimentLabel of 1 to 60 characters, such as "Exp 3 - ATR trail 2.5x". trigr_get_backtest_result then lists the runs with their labels and whether each one measured the saved recipe.

Put the idea to an honest test.

Describe a strategy in plain English or from your own AI assistant, backtest it on point-in-time data, and forward-test it on paper before any real money is involved.