LLM Trading Strategy: Can You Trust What the Model Wrote?

Can you trust an LLM trading strategy? The guardrails that make AI-written strategies checkable: catalog validation, recipe hashes, approvals and trade logs.

Trigr Research6 min read
On this page
  1. Why be skeptical of an LLM-written strategy?
  2. What goes wrong in AI-written strategies?
  3. How does Trigr validate an LLM-built graph?
  4. How does the recipe hash stop silent overwrites?
  5. What do exact-argument approvals protect?
  6. How do you read the trade log of an AI-built strategy?
  7. A checklist before you trust it
  8. Next steps

TL;DR: You should not trust an LLM-written trading strategy because the model sounds confident. You can trust it as far as you can check it. In Trigr that checking is built in: every graph is validated server-side against the capability catalog, edits and publishing are tied to a recipe hash so nothing changes silently, each consequential action uses an exact-argument approval, and every backtest's full trade log can be read for free.

Why be skeptical of an LLM-written strategy?

Large language models are fluent, fast and often right about trading concepts. They are also capable of producing a strategy that looks correct and is not. The problem is not that LLMs are uniquely unreliable; it is that their mistakes are well written.

A human who builds a strategy by hand usually knows where they cut corners. When a model writes it, you inherit decisions you did not make and may not notice: a default parameter, an assumed timeframe, a filter pointed at the wrong market. The question is not "is the model smart enough?" but "can I see and verify everything it decided?"

What goes wrong in AI-written strategies?

The recurring problems fall into a few groups:

  • Invented building blocks. An indicator variant, data feed or parameter that does not exist in the platform, or exists with a different range.
  • Wrong wiring. A cross-asset filter that reads the traded market instead of BTC, a condition inverted, or two entry triggers where the platform supports one.
  • Missing or unintended exits. No stop at all, or an exit the user never asked for.
  • Silent changes. An edit that was supposed to change one filter also dropped another, and the summary did not mention it.
  • Stale overwrites. The assistant saves over a version you edited by hand a minute earlier.
  • Scope creep on risk. Leverage or position size raised because it improved a backtest.
  • Selection by search. A variant chosen because it was the best of many, not because the idea holds up.

Each of these is fixable, but only if the system makes them visible or impossible. The broader failure modes are covered in why most AI trading bots fail.

How does Trigr validate an LLM-built graph?

Trigr strategies are node graphs with a fixed structure: exactly one TRIGGER, optional FILTERs, a SIGNAL, a RISK node and an EXECUTE node. That structure is what makes AI output checkable, because there is a finite set of valid shapes.

The catalog is the authority. Assistants connected through Trigr's MCP server, and the Studio copilot, are instructed to search the capability catalog for exact asset, timeframe, indicator and data-source ids and their parameter schemas before using them. You can browse the same catalog in the catalog docs. Assistants can also check point-in-time data coverage for an asset before building on a series.

Validation happens on the server. Graph limits, semantic checks, disabled sources, billing and plan quotas are enforced server-side, whatever the model claims. Among the checks: exactly one entry trigger, a supported exit on the RISK node (take-profit or stop-loss, an ATR trailing stop, a time stop, or exit on signal flip), and parameters within their ranges, such as leverage from 1 to 50x and size from 1 to 100% of capital. Retired forms like custom exit pipelines are rejected rather than quietly translated into something else. The strategy builder docs list the node rules.

The copilot must declare its changes. In Studio, an edit that removes a node or changes a risk setting without saying so in its summary is rebased back to the previous state. The copilot is also instructed never to increase leverage or position size beyond your settings on its own initiative. The guide to the AI trading copilot covers its other rules.

Validation proves a graph is well formed. It does not prove the graph matches your intent. That part is still yours: open the saved draft in Studio and read each node.

How does the recipe hash stop silent overwrites?

Every saved strategy has a recipe hash identifying the exact version of its graph. Several write paths require the current hash:

  • Replacing a draft. The assistant must pass the hash it last read. If you edited the strategy in Studio in the meantime, the hash no longer matches and the edit fails safely instead of overwriting your work.
  • Publishing, republishing and delisting. These also require the current hash, so the assistant publishes exactly the version you reviewed.
  • Agent slots. A paper agent's slot binds immutable logic: the current recipe hash of an owned strategy, or a pinned marketplace version. The strategy an agent runs cannot change underneath it.
  • Backtests tied to a strategy. Without an experiment label, a backtest that names a strategy must match its saved recipe exactly, or it is rejected as stale. Variants are recorded as labelled experiments that never replace the saved recipe or its verified statistics.

This is optimistic concurrency, the same idea databases use to stop two writers clobbering each other. Applied to strategies, it means the thing you approved is the thing that gets saved, published or run.

What do exact-argument approvals protect?

When you connect an assistant, you pick manual or autonomous approval mode, and you can change it per connection. In manual mode, Trigr asks for a browser confirmation before each consequential action, such as a backtest, a draft save, a publish or a paper-agent creation. In autonomous mode, the connection continues without a prompt each time.

Both modes use the same mechanism: a single-use authorization that expires after ten minutes and is bound to the exact tool arguments. If the model tries to reuse it for a different graph, strategy or cost setting, the call is rejected. Credit checks, quotas, ownership and graph validation stay active either way. The confirmation page shows what you are approving, including, for an experiment, its label and the note that the saved recipe is unchanged.

Some actions have no approval path at all: MCP cannot start execution, place a live trade, withdraw funds, or read or edit exchange credentials. This follows the MCP specification's guidance that a human should be able to deny tool invocations, and the protocol's security best practices on scoped, explicit authorization.

How do you read the trade log of an AI-built strategy?

Headline metrics are summaries, and summaries hide things. The trade log is where you confirm the strategy does what you think. Over MCP, stored runs and their full trade lists, including runs made in the web app, can be paged through for free. Each trade shows entry and exit times, prices, side, size, return, P&L and exit reason.

Look for:

  • Enough trades. A great Sharpe ratio on 15 trades is not evidence of much.
  • Concentration. If three trades or one month produce most of the profit, the result is fragile.
  • Exit reasons that match the design. If you asked for a trailing stop and most exits are time stops, something is off.
  • Sides and sizes. A long-only request should not produce shorts.
  • Units. Trade prices are bar reference prices before slippage, because slippage is charged in the return. Per-trade P&L is in units of an account that starts at 100, not dollars.

Check the audit trail too; it flags which inputs were real and which were simulated. The guide on how to read a backtest goes further on drawdown, DSR and PBO.

A checklist before you trust it

Check What you are confirming
Read every node in Studio The logic matches your intent, including cross-asset filters
Confirm RISK settings Size, leverage, direction and exits are what you chose
Run net as well as gross Slippage and funding are opt-in; fees and the builder fee are always applied
Page through the trade log Trade count, concentration, exit reasons, sides
Count the experiments How many variants preceded this one
Get out-of-sample evidence ML walk-forward results or a paper forward test; the shaded trailing 25% is not a holdout

A strategy that passes all six is not guaranteed to work, but it is one you understand. Backtests are not guarantees; perps are leveraged and can lose more than expected. For deeper statistical checks, see Bailey, Borwein, López de Prado and Zhu on the probability of backtest overfitting.

Next steps

Connect your assistant through the guided MCP setup in manual approval mode for your first session, so you see every action it proposes before it happens.

Frequently asked questions

Can you trust a trading strategy written by an LLM?

Not on the model's word. You can trust it as far as you can check it: the strategy should use only real, validated building blocks, its backtest should come from a real engine you can reopen, and you should read both the logic and the trade log before risking money.

What is a recipe hash in Trigr?

A recipe hash identifies the exact saved version of a strategy's node graph. Edits, publishing and agent slots reference it, so if the strategy changed since the assistant last read it, a stale edit fails instead of silently overwriting newer work.

What does an exact-argument approval protect against?

Each consequential MCP action gets a single-use authorization that expires after ten minutes and is bound to the exact tool arguments. An assistant cannot reuse an approval for a different graph, strategy or cost setting.

What should I look for in a trade log?

Check the trade count and spread over time, whether profit is concentrated in a few trades or one period, whether exit reasons match the intended exits, and whether sides and sizes match the strategy. In Trigr, trade prices are reference prices before slippage and P&L is in units of an account that starts at 100.

Put the idea to an honest test.

Describe a strategy in plain English or from your own AI assistant, backtest it on point-in-time data, and forward-test it on paper before any real money is involved.