TL;DR: Alpha Arena made "LLMs trading Hyperliquid" mainstream by giving frontier models real capital and publishing the results. It is a useful experiment, but a model trading live from prompts is not a validated strategy: each season is one short, unrepeatable path, and nothing in it tells you whether the behavior would survive another regime. The practical lesson is to use an LLM for ideas and tool-driven work, then turn those ideas into fixed rules that are backtested with costs, checked for overfitting and forward-tested on paper, which is what Trigr is built for.
Facts about Alpha Arena and Nof1 below come from third-party tracking and funding coverage, and are stated as of September 2026. We could not confirm the arena's current status.
What was Alpha Arena?
Alpha Arena is a benchmark from the startup Nof1 in which large language models trade real money. According to TradeRank's tracking of the arena, the first season ran six frontier models with $10,000 each on Hyperliquid crypto perpetuals. The appeal was transparency: real capital, visible trades and a public leaderboard.
The same tracking reports these results:
| Season | Market | Reported winner | Reported return |
|---|---|---|---|
| Season 1 | Hyperliquid crypto perps | Qwen 3 Max | +22.31% |
| Season 1.5 | US equities | "Mystery Model," later revealed as Grok 4.20 | +12.11% |
| Season 2 | Not verified | No public roster or results found as of August 6, 2026 | — |
TradeRank also notes that an automated revisit of the arena site in late August 2026 was blocked, and our own check on September 29, 2026 did not load either. Treat the arena's current status as unconfirmed.
Nof1 raised $15M in May 2026 in a round co-led by SUI Group and Karatage and, according to PANews, plans a consumer-facing "market coding agent platform" after the second quarter of 2026. If that ships, it would move Nof1 from benchmark to product.
What does Alpha Arena actually show?
It shows something real: current LLMs can run a full trading loop. They can read market data, decide on a position, size it and manage it over weeks without a human writing the rules. That was not obvious a few years ago, and the public format made it easy to follow.
What it cannot show is whether any of that behavior is an edge. Three reasons:
- One path per season. Each model traded one stretch of one market. A large gain or loss over a few weeks can come from a single well-timed trend. There is no second run of the same period to compare against.
- Few competitors, one winner. With six models, someone has to finish first. Picking the top result from several contestants and treating it as skill is a textbook case of selection bias.
- No fixed rule to test. An LLM deciding from a prompt is not a strategy with stable parameters. Rerun the same prompt and it may trade differently, and a model update changes its behavior again. You cannot backtest "what the model would have done" over past years in a way that is free of look-ahead, because the model's training data may include those years.
The season results are therefore a data point about models in a specific window, not a ranking of trading ability. The deeper version of this argument is in why most AI trading bots fail.
How is an LLM trading raw prompts different from a validated strategy?
The difference is where the decision lives.
| LLM trading from prompts | Validated rule-based strategy | |
|---|---|---|
| Decision logic | Generated fresh each time by the model | Fixed rules you can read and version |
| Repeatability | Output can vary run to run and across model updates | Same inputs, same trades |
| Historical test | Not cleanly possible; training data may contain the test period | Backtest on point-in-time data |
| Cost handling | Depends on what the model remembers to consider | Fees, slippage and funding applied by the engine |
| Overfitting check | None built in | Trial counts, DSR and PBO |
| Authority | Model acts on the account | Rules run inside limits you set |
This does not mean LLMs are useless for trading. They are good at the parts around the decision: turning a vague thesis into a precise hypothesis, searching for the right data, writing and editing strategy definitions, reading trade logs and noticing what went wrong. The mistake is letting the model be the strategy.
How do you turn an LLM's idea into a backtested agent?
Here is a workflow that keeps the LLM's strengths and removes the unverifiable part. It uses Trigr's MCP server, which ChatGPT, Claude, Claude Code and Codex can connect to.
1. Write the hypothesis down first
Ask the model for a specific, falsifiable idea: market, timeframe, entry condition, exit, and why it might work. "Short BTC on 1H when funding is extreme and open interest is falling" is testable. "Trade momentum smartly" is not. Decide in advance what result would make you drop it.
2. Map it to real building blocks
The assistant searches Trigr's capability catalog for the assets, timeframes, indicators and data sources that exist, then builds a node graph: exactly one trigger, filters, a signal and a risk block (size, leverage, take-profit and stop-loss, ATR trail, time stop). The graph is validated against the catalog, so the model cannot invent an indicator that is not there. The capability catalog docs list what is available.
3. Backtest gross, then net
The first result is labeled gross: trading fees and the builder fee are always applied, but slippage and funding are off. Rerun with slippage (flat bps per fill) and historical funding switched on. Point-in-time data and next-bar-open fills mean the backtest never trades on information it could not have had. Have the assistant read the full trade log, which is free over MCP, rather than only the headline numbers.
4. Iterate as experiments, not new strategies
Parameter, filter, exit and sizing changes are recorded as labeled experiments inside one strategy's version history (experimentLabel over MCP), so you can see how many variants you tried and promote the winner deliberately. If you want a stronger test, an ML optimization run reports out-of-sample Sharpe from purged walk-forward validation, a leak verdict, the Deflated Sharpe Ratio (Bailey and López de Prado) and PBO. A plain backtest does not compute DSR or PBO.
5. Forward-test on paper
Over MCP, the assistant can create a paper agent in a paused state. You review and start it. Paper agents fill at the current Hyperliquid price and pay the taker and builder fees; they model no slippage or funding, so compare them to your net backtest with that in mind.
6. Go live yourself
Live trading is started in the Trigr app, not over MCP. The assistant cannot start execution, place a live trade or withdraw. Your Hyperliquid agent uses a trade-only key that cannot withdraw, and you keep custody of your funds. The pillar on deploying a strategy as a Hyperliquid agent covers this step.
What does this mean for you?
If Alpha Arena made you curious about AI trading, the useful takeaway is not which model won. It is that an LLM is a capable research partner and a risky decision-maker. Let it generate and refine hypotheses, then hold the resulting rules to the same standard as any other strategy: point-in-time history, realistic costs, a visible trial count and a forward test.
Backtests are not guarantees; perps are leveraged and can lose more than expected.
When is a live LLM contest the better tool?
Live contests are better than backtests at one thing: showing how an autonomous model behaves in real time, including how it handles news, drawdowns and its own mistakes. If your question is "how do current models reason about markets?", watching them trade is informative. If your question is "should I put money behind this strategy?", you need rules you can test.
Next steps
To try the workflow above, connect ChatGPT, Claude or Codex to Trigr and ask it to turn one trading idea into a strategy and backtest it gross, then net.