Leaked eval: GPT-6 Sol scores 1.6x GPT-5.6 Sol on robotics, 47% cheaper but off Pareto frontier
ycombinator · x · 2026-09-25
An unverified eval leak claims GPT-6 Sol scores 1.6x as high as GPT-5.6 Sol on robotics tasks while being 47% cheaper, though it still sits below the Pareto frontier under Opus 5.5. Authenticity remains unconfirmed.
More from Models
- Claude Opus 5.5 builds Minecraft from one prompt in ~1 hour for ~$20 — amasad · 2026-09-25
- Dev finds Opus 5.5 Medium reasoning so good that High feels unnecessary — rudrank · 2026-09-25
- PINNACLE: GPT-6 Sol cuts errors 2.5x at max effort, Claude Opus 5.5 doesn't benefit — ryanshrout · 2026-09-25
- 4B open model tops JevBench by being 5x faster and half the price — airesearch12 · 2026-09-25
- Terminal-Bench-Science leaderboard launches with GPT-6 Astra at 63% — scaling01 · 2026-09-25
- User says Opus 5.5 writes so well he deleted his 'don't write like a fuckhead' custom instruction — generativist · 2026-09-25