OpenAI o1 beats GPT-4o on AIME, Codeforces, and GPQA Diamond
willdepue · x · 2026-07-22
A post jokes that the latest OpenAI reasoning-model plot line is one of the wildest ever, then points to benchmark charts showing o1’s gains.
The image compares GPT-4o, o1 preview, and o1 on AIME 2024, Codeforces, and GPQA Diamond, with o1 ahead on all three and especially strong on competition code and math.
More from Models
- Moonshot points users to quick-start access for Kimi K3 — maier_ak · 2026-07-22
- Moonshot’s Kimi K3 arrives as a 2.8-trillion-parameter open-weight model — maier_ak · 2026-07-22
- LongCat-2.0 cuts agent input costs by 88% in a new test — karminski3 · 2026-07-22
- Google’s Genie3 is said to simulate the real world from Street View images — ZeroStateReflex · 2026-07-22
- DeepSeek-then-Claude workflows are “watered down,” but users still love them — tinyfool · 2026-07-22
- Grok’s translation is so bad users pre-check it with ChatGPT, says X poster — tinyfool · 2026-07-22