Opus 5.5 vs GPT 6: benchmark comparison chart in one image

vista8 · x · 2026-09-27

A compiled chart compares recently showcased benchmark scores of Claude Opus 5.5 and GPT 6. The author argues Terminal-Bench 4.0 is high-signal — smooth score gains across generations usually indicate a quality model rather than benchmark gaming. GPT scores very well on AutomationBench, completing tasks across multiple applications, possibly thanks to Computer Use capabilities. A quick side-by-side of each model's strengths.

Related event: Benchmark Chart Compares Claude Opus 5.5 vs GPT 6(2 posts)→

Original post →

More from Models

Models channel →