PINNACLE: GPT-6 Sol cuts errors 2.5x at max effort, Claude Opus 5.5 doesn't benefit

ryanshrout · x · 2026-09-25

PINNACLE benchmarked GPT-6 Sol/Luna and Claude Opus 5.5 at default and maximum reasoning effort on identical multi-step enterprise jobs with a 128K retrieval corpus. GPT-6 Sol at max effort cuts weighted errors 2.5x (score 3,998 to 9,853), finishes 99.6% of jobs vs 96.8%, and drops fabrications from 7.9% to 2.9%—but costs 1.8x more per correct task (29¢ vs 16¢). Analyst Patrick Moorhead adds that Opus 5.5 needs 3x output tokens for correct answers, making it look far worse vs GPT-6 Sol Max. Raising thinking effort doesn't help Claude Opus 5.5.

Original post →

More from Models

Models channel →