Opus 5.5 vs GPT 6 benchmarks in one chart: Terminal-Bench 4.0 scores worth watching
vista8 · x · 2026-09-27
Blogger vista8 shares a chart comparing Opus 5.5 and GPT 6's published benchmark results, noting that an AI-regenerated version of the chart has more accurate categories and clearer fonts. Key points: Terminal-Bench 4.0 carries real signal—smooth score gains across generations usually indicate a quality model rather than benchmark gaming; GPT performs very well on AutomationBench across multiple apps, possibly thanks to Computer Use capabilities.
Related event: Benchmark Chart Compares Claude Opus 5.5 vs GPT 6(2 posts)→
More from Models
- Perplexity CEO ran 50-100 agent workflows on Opus 5.5 vs Fable 5.1, found minimal differences — AravSrinivas · 2026-09-28
- Claude's Snarky Tone and Mysterious "For Anyone Reading Along" Baffle Users — MrsChatGPT4o · 2026-09-28
- Matt Shumer: Opus 5.5 Added Controller Support to His Product — mattshumer_ · 2026-09-28
- Grok Bot already works in Teslas via the built-in Grok app, full release due this summer — yunta_tsai · 2026-09-28
- Claude Opus 5.5 reportedly balks at a population ethics problem, researcher suspects 'repugnant conclusion' aversion — anderssandberg · 2026-09-28
- Xiaomi releases MiMo-V2.6-Flash-MOPD on Hugging Face — Automatic-Arm8153 · 2026-09-28