Grader stunned as Opus 5.5 makes 'intellectual bets' that beat all human performance ceilings
rickasaurus · x · 2026-09-23
samsinai says he initially couldn't believe the results while grading, suspecting the model had cheated, but a review showed its approach was genuinely clever. What impressed him most about Opus 5.5 was that it made intellectual bets that could have failed within the time limit — they paid off, pushing the performance ceiling beyond all humans. A firsthand capability observation.
More from Models
- Model release cadence shrinks from 73 to 18 days; RSI odds pulled toward 2027 — soumitrashukla9 · 2026-09-23
- Andriy Burkov: Codex Is Infinitely Faster Than Any Open-Weight Agent, But Priced Out of Reach — burkov · 2026-09-23
- Viral demo claims 'GPT-6' can drive browser Paint to draw, unverified — alexcovo_eth · 2026-09-23
- Opus 5.5-generated three.js spell demo wows with procedural VFX and sound — majidmanzarpour · 2026-09-23
- Early hands-on with rumored Opus 5.5 in Scenario's Blender plugin stuns users — repligate · 2026-09-23
- GPT-6 Sol's price cut means it should be compared to Sonnet, not Opus — TraditionalHome8852 · 2026-09-23