Frontier LLMs all lost money in a 1.6-year synthetic trading test

Scobleizer · x · 2026-07-22

Frontier LLMs all lost money on a synthetic trading benchmark

A SynthFin Trading Bench test put Fable, GPT-5.5, Gemini 3.5 Pro, and Grok 4.5 into portfolios on a fully synthetic stock market they had never seen before.

The image also notes the benchmark was designed to be contamination-free: all prices, filings, and news were machine-generated and unpublished, so the models could not have memorized them.

Original post →

More from Models

Models channel →