FrontierCode 1.1 shows Opus 5 can score lower under stricter reasoning settings
andrew_n_carr · x · 2026-07-25
- A repost says Opus 5 actually scores lower on FrontierCode 1.1 when using xhigh and max reasoning settings.
- The poster’s investigation claims the benchmark’s scope metric penalizes Opus 5 for making changes that are unrelated to the task.
- The attached chart compares models on FrontierCode 1.1 Main with score vs. cost, showing Opus 5 near the top but not clearly dominant across all settings.
- The takeaway is that the headline benchmark number may be sensitive to rubric details, especially how “scope” is judged.
Related event: Opus 5 Peaks at Medium Reasoning in FrontierCode Tests(7 posts)→
More from Models
- Claude Opus 5 system card shows gains on RiemannBench, Chartography and GDP.pdf — echen · 2026-07-25
- Perplexity adds Opus 5 to Computer for Pro and Max users at roughly half the price — AravSrinivas · 2026-07-25
- GPT 5.6 Sol is highlighted as stronger on one benchmark row at 72.7% — himanshustwts · 2026-07-25
- Leaked Claude Opus 5 Takes 50% Longer Than Opus 4.8 in AA-Briefcase Tasks — ArtificialAnlys · 2026-07-25
- Claude Opus 5 tops AA-Briefcase with 1720 Elo and 20% lower task cost than Fable 5 — ArtificialAnlys · 2026-07-25
- GPT-5.6 Sol looks set to beat an Act 3 A7 boss — Jsevillamol · 2026-07-25