Frontier LLMs all lost money in a 1.6-year synthetic trading test
Scobleizer · x · 2026-07-22
Frontier LLMs all lost money on a synthetic trading benchmark
A SynthFin Trading Bench test put Fable, GPT-5.5, Gemini 3.5 Pro, and Grok 4.5 into portfolios on a fully synthetic stock market they had never seen before.
- The run lasted about 1.6 years across 1,000 stocks and 10 macro regimes.
- Every model lost money.
- None of the LLMs beat a simple buy-and-hold baseline.
- The chart shows equal-weight doing best at +4.84%, while buy & hold was roughly flat at +0.25%; the AI portfolios all finished negative, with Gemini 3.5 Flash the worst at -24.31% and Momentum at -26.53%.
The image also notes the benchmark was designed to be contamination-free: all prices, filings, and news were machine-generated and unpublished, so the models could not have memorized them.
More from Models
- Xiaohongshu’s dots-note-3.0 reportedly scores 42/42 on the IMO test — xiaohu · 2026-07-22
- As Frontier Labs Chase AGI, Specialized AI Thrives on Cost and Speed — bendee983 · 2026-07-22
- Grok 4.5 gets praise for clearer, more direct technical writing — elonmusk · 2026-07-22
- OpenAI, Anthropic and Google fall to 83.29% of model spend as Moonshot surges 8.4x — teortaxesTex · 2026-07-22
- Reddit user says a chat-template tweak can force Laguna-S-2.1 into longer reasoning — SnooPaintings8639 · 2026-07-22
- Anthropic–Qwen distillation debate turns into a licensing argument — nptacek · 2026-07-22