Opus 5.5 vs GPT 6 benchmarks in one chart: Terminal-Bench 4.0 scores worth watching

vista8 · x · 2026-09-27

Blogger vista8 shares a chart comparing Opus 5.5 and GPT 6's published benchmark results, noting that an AI-regenerated version of the chart has more accurate categories and clearer fonts. Key points: Terminal-Bench 4.0 carries real signal—smooth score gains across generations usually indicate a quality model rather than benchmark gaming; GPT performs very well on AutomationBench across multiple apps, possibly thanks to Computer Use capabilities.

Related event: Benchmark Chart Compares Claude Opus 5.5 vs GPT 6(2 posts)→

Original post →

More from Models

Models channel →