OpenAI o1 beats GPT-4o on AIME, Codeforces, and GPQA Diamond

willdepue · x · 2026-07-22

A post jokes that the latest OpenAI reasoning-model plot line is one of the wildest ever, then points to benchmark charts showing o1’s gains.

The image compares GPT-4o, o1 preview, and o1 on AIME 2024, Codeforces, and GPQA Diamond, with o1 ahead on all three and especially strong on competition code and math.

Original post →

More from Models

Models channel →