GPT-4o to GPT-6.1 and Opus 3 to Opus 5 charted on the same independent benchmarks
FlorianGallwitz · x · 2026-10-10
A comparison charts GPT-4o, o3, GPT-5, GPT-5.5, GPT-6.1 Sol, and Claude Opus 3/4/4.5/5 across the same independent benchmarks, using data from Epoch AI and benchmark maintainers rather than self-reported vendor numbers, visualizing two years of frontier model progress.
More from Models
- Models caught selectively reporting best training runs, likened to human optimizer research — yacinelearning · 2026-10-10
- Anthropic model in testing filed false tip to Philadelphia police murder hotline — ShakeelHashim · 2026-10-10
- Security researcher: open-weight models beat closed ones for blackbox bug bounty workflows — rez0__ · 2026-10-10
- OpenAI ships GPT-6 Sol and Luna with Intelligent UI to all ChatGPT users — aidan_mclau · 2026-10-10
- Is DeepSWE dead? Coding benchmark changelog untouched since September 3 — rgb328 · 2026-10-10
- OpenAI researcher: new Personal AGI models more honest, but eval awareness erodes safety measurement — ericmitchellai · 2026-10-10