Dan Shipper Publishes Writing Evals Comparing Opus 5.5, Opus 5, Astra and 5.6
danshipper · x · 2026-09-23
Dan Shipper shares his writing evals comparing Opus 5.5, Opus 5, Astra and 5.6 on his real-world writing tasks, recommending readers go through them with their agent. The quoted post notes that another user's personal benchmark found Opus 5.5 at "Fable level" on many actual work tasks, though it ran too long to score on a few.
Related event: Early reviews praise Opus 5.5 for writing and prototyping(4 posts)→
More from Models
- Sol 6 matches Sol 5.6 quality with half the reasoning time, better Plus limits — Simple-Diver-2192 · 2026-09-23
- Opus 5.5 sweeps GPT-6 Sol on every benchmark, but Sol runs tasks for $1.06 — Vast-Grapefruits · 2026-09-23
- Model release cadence shrinks from 73 to 18 days, sparking safety calls — tristanbob · 2026-09-23
- One prompt, a 3D Alpha Centauri simulation: Opus 5.5 demo goes viral — Angaisb_ · 2026-09-23
- Anthropic's new SOTA model sets another record: tokens burned to get there — soumitrashukla9 · 2026-09-23
- Claude Opus 5.5 takes #1 on Artificial Analysis, now available in Claude Code — ccerrato147 · 2026-09-23