PINNACLE: GPT-6 Sol cuts errors 2.5x at max effort, Claude Opus 5.5 doesn't benefit
ryanshrout · x · 2026-09-25
PINNACLE benchmarked GPT-6 Sol/Luna and Claude Opus 5.5 at default and maximum reasoning effort on identical multi-step enterprise jobs with a 128K retrieval corpus. GPT-6 Sol at max effort cuts weighted errors 2.5x (score 3,998 to 9,853), finishes 99.6% of jobs vs 96.8%, and drops fabrications from 7.9% to 2.9%—but costs 1.8x more per correct task (29¢ vs 16¢). Analyst Patrick Moorhead adds that Opus 5.5 needs 3x output tokens for correct answers, making it look far worse vs GPT-6 Sol Max. Raising thinking effort doesn't help Claude Opus 5.5.
More from Models
- Is there any LLM whose training data is fully auditable and un-stolen? — LuCiAnO241 · 2026-09-25
- Can a small local LLM with internet access rival a larger model? — mototuneup · 2026-09-25
- Math Benchmark: Astra Dominates, Nothing Below Fable 5.1 Is Competitive — teortaxesTex · 2026-09-25
- Someone ran tests across all Claude models and published the results — repligate · 2026-09-25
- Open-Sourced 3D Pelican Bike Game: One Prompt, Zero Hand Edits, Prompt Included — EricBuess · 2026-09-25
- One Lazy Prompt to Claude Opus 5.5 Yields a Full 3D Storybook Farm Game — EricBuess · 2026-09-25