Claim of GPT-6.1 Sol Crushing Opus 5.5 Sparks Benchmark Skepticism

A Reddit post claimed OpenAI's internal benchmarks show GPT-6.1 Sol far ahead of Claude Opus 5.5, but a follow-up post mocked the claim as evidence of blind benchmark worship.

2026-09-30 ~ 2026-10-01 · 2 related posts

Full story(2 episodes)→