TL;DR: Most AI trading bots fail for four ordinary reasons: the edge was hallucinated or found by brute-force search, the backtest leaked future data, costs were left out, and the AI was given more authority than anyone could supervise. An honest setup fixes each one: numbers only from a real point-in-time engine, explicit costs, a visible trial count with out-of-sample evidence, and hard limits so the AI can never place live trades or move funds on its own.
What do people mean by an "AI trading bot"?
The phrase covers very different things:
- A chatbot asked for trade ideas. A general assistant with no market data, answering from training text.
- A signal bot with an ML model inside. A model trained on historical features that emits long or short signals.
- An LLM agent wired to an exchange. A model that can call an order API directly, sometimes with a full-permission key.
- An assistant connected to a research platform. A model that can look up data, run backtests and set up paper tests through controlled tools.
The failure modes below apply to all of them in different proportions. The fourth pattern, done carefully, is the one that can avoid most of them, which is the approach Trigr takes with its MCP server for ChatGPT, Claude and Codex.
Why do most AI trading bots fail?
1. The edge was hallucinated
A language model with no tools will answer "how would a 20/50 EMA crossover on SOL have performed?" with a confident paragraph and sometimes specific numbers. None of it was computed. Models generate plausible text; a plausible Sharpe ratio is still made up.
The subtler version happens with tools. An assistant runs a real backtest, gets a mediocre result, then summarizes it generously or blends in remembered numbers from elsewhere. The fix is procedural: every figure in a report should trace back to a specific tool result, and you should be able to open that run and its trades yourself.
2. The backtest could see the future
Look-ahead bias is the most common reason a strategy works in testing and fails live. Typical causes:
- Using the close of a bar to decide a trade that is then filled at the same close.
- Joining a daily or weekly series onto hourly bars before the daily bar has finished.
- Using a data point, such as a macro release, at its reference date instead of its publication date.
- Training a model on features whose labels overlap the test period.
AI makes this worse because it writes feature code quickly and confidently, and leakage bugs do not throw errors. They just produce beautiful equity curves. The guide to look-ahead bias in crypto backtests walks through concrete examples.
3. The costs were missing
On perpetual futures, costs are not a rounding error. Every fill pays a trading fee, every fill suffers some slippage, and positions pay or receive funding while open. A strategy that trades several times a day can turn a strong gross curve into a losing net one. Many bot demos show gross results without saying so. The breakdown in slippage and funding in perp backtests shows how quickly these add up.
4. The winner was found by searching
Give an AI assistant a backtest tool and it will happily run 200 variants. The best of 200 will look excellent almost regardless of whether the underlying idea has any merit. Bailey, Borwein, López de Prado and Zhu formalized this as the probability of backtest overfitting, and Bailey and López de Prado's Deflated Sharpe Ratio adjusts a Sharpe ratio for the number of trials. Both make the same point: the more you search, the less the best result means.
Speed is the problem here. What once took a researcher weeks, an agent does before lunch, and the trial count is rarely recorded.
5. The AI had unlimited authority
The last failure is not statistical. An LLM agent with a full-permission exchange key can place any order, at any size, at any time, and in some setups withdraw funds. Models misread instructions, loop on errors, and can be steered by text they read, such as a strategy description written by someone else. The MCP specification itself says there should always be a human in the loop able to deny tool invocations.
What does an honest setup look like?
Each failure has a structural fix. Good prompts help, but guardrails that do not depend on the model behaving well matter more.
| Failure | Honest fix | How Trigr handles it |
|---|---|---|
| Hallucinated edge | Numbers only from a real engine, runs you can reopen | Backtests run on Trigr's engine; stored runs and full trade logs can be re-read for free, including over MCP |
| Look-ahead bias | Point-in-time data and realistic fills | Features use only data available at the bar's close; coarser cross-asset inputs are carried forward, never read ahead; fills at the next bar's open |
| Missing costs | Explicit gross versus net | Trading fees and the builder fee are always applied; slippage and funding are opt-in, and a result without them is labeled gross |
| Search-driven winners | Visible trial count and out-of-sample evidence | Variants run as labelled experiments in one strategy's history; ML optimization reports DSR, PBO and purged walk-forward out-of-sample Sharpe |
| Unlimited authority | Hard limits and exact approvals | MCP cannot place live trades, start execution or withdraw; consequential actions use a single-use authorization bound to exact arguments |
What guardrails does Trigr put around AI?
The table is the summary. A few details are worth spelling out, because they are where honest and dishonest setups differ in practice.
The engine is conservative where it is ambiguous. When a take-profit and stop-loss are both touched inside one bar, the engine replays 5-minute bars inside it to find which came first; if it is still ambiguous, the stop wins. Missing data is never filled with invented values. Every backtest includes an audit trail that flags real versus simulated inputs. The backtesting docs cover each rule.
The tools say what is not out-of-sample. A standard backtest uses the full available history, so it has no holdout. Studio shades the trailing 25% as a recent-period diagnostic, and both the app and the MCP tools say plainly that it is not a holdout. Genuine out-of-sample evidence comes from ML optimization's anchored, purged walk-forward results, or from forward-testing on a paper agent with data that did not exist when the strategy was chosen.
The ML is honest about what it does. Trigr's ML optimization uses point-in-time features (the loader raises on future leakage), triple-barrier labels and sample weights by average uniqueness, following López de Prado. It chooses among LightGBM, random forest and XGBoost by walk-forward out-of-sample performance with fixed family parameters; hyperparameters are not tuned. Each run returns a leak verdict with specific flags.
Authority is split. An assistant connected over MCP can research, backtest, save drafts, publish eligible strategies and create a paper agent in a paused state. It cannot start execution, place a live trade, withdraw funds, or read or edit exchange credentials. The reasoning is covered in AI trading agent safety.
Live trading uses a key that cannot withdraw. When you deploy a live agent on Hyperliquid, it trades through a trade-only API wallet that can place and cancel orders but cannot withdraw or send funds to another address. For autonomous agents Trigr stores that trade-only key encrypted server-side; you keep the master wallet key and custody of your funds. The details are in Hyperliquid agent wallets explained.
Third-party text is untrusted. Marketplace names and descriptions are written by other users. Trigr's MCP guidance tells agents never to let that text authorize a tool call, and the exact-argument authorization means a hijacked instruction cannot reuse an approval for a different action.
How do you evaluate any AI trading bot?
Whether you use Trigr or anything else, ask these questions before trusting a result or connecting a wallet:
- Can I open the exact backtest behind every number, including its trade log?
- Are fills at the next bar, and is all data point-in-time?
- Are fees, slippage and funding included, and is a gross result labeled as gross?
- How many variants were tried before this one, and is that count recorded?
- Where is the out-of-sample evidence: walk-forward, a reserved window tested once, or paper forward-testing?
- What can the AI do with money, and can the exchange key withdraw?
If a provider cannot answer these, the equity curve is marketing. Backtests are not guarantees; perps are leveraged and can lose more than expected.
Next steps
Read whether you can trust a strategy an LLM wrote for the checks that apply to AI-built graphs, or connect your own assistant through the guided MCP setup and start with a paper test.