Opus 5.5 field results: 200k-line codebase fixed in under 3 hours at 1/5 the cost of GPT-6 Astra
vista8 · x · 2026-09-23
A detailed roundup of Claude Opus 5.5 benchmarks and field tests:
Benchmarks
- Terminal-Bench 4.0: 66.4%; FrontierCode v1.1: 54.4%; CursorBench 4.0: 57.8% — all leading GPT-6 Astra and GPT-5.6 Sol.
- GDPval-AA v2.1 (real-work eval across 44 professions): 1846 Elo vs Fable 5.1's 1735 and Opus 5's 1708.
Coding efficiency
- A beta user fixed a 200k-line codebase in under 3 hours with Opus 5.5; Opus 5 took 20+ hours and 2.5x more tokens.
- Anthropic internally had both models port HAProxy from C to Rust: both passed regression tests, but Opus 5.5 finished in 9.5h vs 12h at 51% lower cost.
- GitHub testing: in VS Code, Opus 5.5 needs less than half the steps Opus 5 does for terminal tasks.
Knowledge-work accuracy
- On a quarterly-report task with deeply hidden sources and strict citation checking, Opus 5.5 passed 16 of 18 attempts; Fable 5.1 and Opus 5 passed none.
Cost & misc
- Cache reads down 60% (good for agents), output 30%+ faster, and default settings outperform GPT-6 Astra at max settings at one-fifth the cost.
Related event: Claude Opus 5.5 Launches to Top Rankings with Lower Price and Faster Speed(28 posts)→
More from Models
- ProgramBench: rebuilding programs from binaries is brutal — Claude Opus 5 leads at 4.5% resolved — jyangballin · 2026-09-23
- GPT-6 Sol and Luna hit Arena, plus a matched head-to-head vs GPT-5.6 Sol — arena · 2026-09-23
- Plinius leaks full Claude Opus-5.5 system prompt, over 1.9M characters with tools — ivan_bezdomny · 2026-09-23
- Claude Opus 5.5 Frontend Tests: Suspected Quantized fable 5.1, Stable but Heavier Reasoning — karminski3 · 2026-09-23
- Opus 5.5, GPT-6 Sol and Luna drop the same night as AI pace debate rages — ThePeterMick · 2026-09-23
- Travel planner PlanMyVisit switches to GPT-6: faster and half the cost — alexbainbridge · 2026-09-23