44 AI Models, 2,527 Trades, 42.6% Profitable: What The Public LLM Trading Arenas Actually Proved
For two years the question 'can an LLM trade?' was answered with vibes. Then someone ran the experiment in public. Across six completed seasons since January 2026, 44 AI models have placed 2,527 trades with $650K of simulated capital under identical market data, prompt rules and risk controls - and only 42.6% of model-seasons finished profitable. Nof1's Alpha Arena, the competition that started the genre, ended in December 2025. The results are genuinely useful, just not in the way either the boosters or the sceptics wanted. Here is what the data supports, what it definitively does not, and why the interesting finding is about consistency rather than returns.
AlchmAI Editorial12 min read
42.6%
Of model-seasons finished profitable across six completed seasons - worse than a coin flip
2,527
Trades placed by 44 AI models since January 2026 under identical data, prompts and risk controls
$650K
Simulated capital deployed across the published seasons
Dec 2025
When Nof1's Alpha Arena, the competition that started the genre, concluded and stopped running models
Every few months since 2024, someone has asked us whether a large language model can trade. It is a reasonable question and it was, until recently, impossible to answer honestly, because the available evidence consisted of backtests run by people selling something and screenshots on social media with the losing weeks cropped out. What changed is that several groups decided to run the experiment in the open, with real prices, published decisions and identical conditions across models. That is a genuine contribution, and the results deserve to be read carefully rather than deployed as ammunition.
The headline: across six completed seasons since January 2026, 44 AI models have made 2,527 trades with $650K of simulated capital under identical market data, prompt rules and risk controls, and only 42.6% of model-seasons finished profitable. Nof1's Alpha Arena, the competition that popularised the format by having six frontier models trade crypto with real capital, concluded on 3 December 2025 and no longer runs. Broader assessments of the genre have been blunt: there is limited evidence of LLMs beating the market with any scale or statistical significance.
What The Data Genuinely Supports
- No general-purpose model has demonstrated a durable trading edge. This is the strongest conclusion available and it is well supported. Under controlled, repeated, public conditions, none of 44 models produced consistent outperformance. Anyone selling an LLM-driven trading product should be asked to explain why their result differs.
- Model ranking is unstable between seasons. A model that tops one season frequently underperforms the next. That instability is itself the finding: it is the signature of variance rather than skill, and it is exactly what you would expect if the models are not actually extracting persistent signal.
- Identical conditions matter enormously. The reason these arenas are valuable is the harness - same data, same prompts, same risk controls. Most public claims about AI trading performance come from setups where the comparison is not controlled, and the differences in harness usually explain more than the differences in model.
- The models are coherent, which is not nothing. They produce reasoned, consistent, rule-respecting decisions and publish their reasoning. They are not behaving randomly. They are behaving sensibly and still not making money, which is a far more informative result than incoherence would have been.
What It Does Not Prove
Sceptics have over-read these results almost as enthusiastically as the boosters over-read the early Alpha Arena leaderboards. Three limits are worth being precise about:
- 01It says nothing about AI in trading generally, only about general-purpose language models making discretionary directional calls. Machine learning has been generating genuine alpha in systematic funds for two decades. Those systems are purpose-built, trained on market microstructure, and share almost nothing with a frontier chat model except the word 'AI'.
- 02The sample is short and the conditions are narrow. Six seasons is not long, the instrument universe skews toward crypto and large-cap equities, and position sizing is constrained by competition rules. A strategy with genuine edge on a longer horizon or in a different regime would not necessarily reveal it here.
- 03Simulated capital changes behaviour in ways that matter. Slippage, market impact, borrow availability and the psychological realities of drawdown are absent or approximated. That cuts both ways - it flatters the results by removing costs, and understates them by removing the compounding of a real book.
“The models are not failing because they reason badly. They are failing because reasoning well about public information is not an edge when everyone has the same public information and a faster way to act on it.”
Why This Result Should Have Been Expected
Stepping back, the outcome is close to what market theory would predict, and it is worth spelling out because it tells you where AI can add value in trading.
A frontier language model reasoning about a liquid, heavily-covered market is competing against every participant who has the same information plus faster execution, better data, and in many cases purpose-built models trained on precisely this task. The model's genuine advantages - reading and synthesising vast quantities of text, holding many considerations at once, never getting bored - are advantages in information processing, not in the part of trading where money is made. Public information is priced. The model is arriving at a well-reasoned view of things everyone already knows, which is the definition of no edge.
That framing also explains where these systems do work, which is the part of this story most commentary skips.
Where AI Is Actually Making Money In Trading
We build trading platforms, signal integration and quant infrastructure, and the pattern across our client base is consistent: nobody serious is asking an LLM what to buy. They are using AI where the advantage is processing rather than prediction.
- Extraction at scale. Reading every filing, transcript and disclosure in a universe and converting them into structured, comparable fields with citations. The model is not predicting; it is reading faster than a team could, and the output is verifiable in seconds.
- Unstructured-to-structured conversion. Turning broker research, news and regulatory filings into features that a purpose-built quantitative model consumes. The LLM does the part it is good at; the model that trades is the one that was built to trade.
- Execution quality analysis. Explaining why fills diverged from expectation across thousands of orders, which is a pattern-and-narrative task rather than a forecasting one.
- Monitoring and exception triage. Watching for the anomalous and escalating with a written rationale. High volume, checkable, and a genuine capacity multiplier - particularly as trading hours extend.
- Signal and rating presentation. Rendering model-generated signals into trading interfaces with provenance attached - what generated this, from what inputs, at what time. The intelligence is upstream; the AI layer makes it usable and auditable.
How To Read The Next Leaderboard
These arenas will keep running and someone will keep topping them, so a short guide to reading the coverage sensibly:
- 01Ask how many seasons, not how many percent. A single winning season across dozens of competitors is what variance looks like. Persistence across seasons is what skill looks like, and so far nobody has shown it.
- 02Check the harness before the ranking. Prompt design, risk limits, rebalancing frequency and instrument universe frequently explain more of the dispersion than model choice does. A leaderboard without a documented harness is a screenshot.
- 03Look at drawdown alongside return. A model that returned 8% with a 40% intra-season drawdown did not outperform one that returned 5% smoothly, whatever the ranking says.
- 04Discount crypto-only results for equities conclusions. Different microstructure, different liquidity, different participant mix. The transfer is not automatic.
- 05Treat any vendor citing an arena win as a product claim with extreme suspicion. The published data is the best evidence available that this particular capability does not reliably exist.
The Bottom Line
The public LLM trading arenas did something genuinely valuable: they took a question that had been answered with marketing for two years and answered it with data. Across 44 models, 2,527 trades and six seasons under identical conditions, only 42.6% of model-seasons were profitable and no model demonstrated persistent edge - which is the result market efficiency would predict when a well-reasoning agent competes on public information against participants who are faster and better equipped. That is not a verdict on AI in trading; it is a verdict on one specific application of it. The money in AI on a trading desk is in processing rather than prediction: reading everything, structuring the unstructured, explaining execution, triaging exceptions, and presenting signals with honest provenance. Those are unglamorous compared to an autonomous AI trader, and unlike an autonomous AI trader, they demonstrably work. That is where we build, and the arena data is the best argument yet for building there.
References & Further Reading
- TradeRank - live AI trading leaderboard and LLM competition. traderank.ai
- TradeRank - best LLM for crypto trading benchmark and season data. traderank.ai/llm-trading-benchmark
- Nof1 - Alpha Arena, the AI trading benchmark. nof1.ai
- iWeaver - Alpha Arena season 1 results: final ranking and lessons. iweaver.ai/blog/alpha-arena-ai-trading-season-1-results
- TradeRank - Alpha Arena alternatives and the state of AI trading competitions (2026). traderank.ai/blog/alpha-arena-alternatives-2026
- Flat Circle - AI trading arenas. blog.flatcircle.ai/p/ai-trading-arenas
AlchmAI Editorial
Research and analysis, London
The AlchmAI team writes about the markets, technology and regulation we work with every day. We build trading platforms, real-time charts and AI analysis tools for brokers, prop firms and fintech teams from our office in Mayfair, London. Every article lists its sources. Nothing we publish is investment advice.
This article is general information and commentary. It is not investment advice or a recommendation to buy or sell any investment. Important information