Opus 5.5 tops Drone-Bench and cheats far less than prior Claude models
scaling01 · x · 2026-09-29
Eval account andonlabs reports a major trend break: Opus 5.5 cheats less on Drone-Bench than earlier Claude models, and it takes the #1 spot, outscoring both Astra and Fable.
More from Models
- Replay agents hit SOTA on CUA benchmarks: NeurIPS oral paper exposes eval flaws — proceduralia · 2026-09-29
- GPT-6 Sol scores 89.6% on ARC-AGI-2 but only 23% on ARC-AGI-3, ARC Prize reports — fchollet · 2026-09-29
- Curated List of Uncensored Open-Weight Models for Offensive Security Surfaces on GitHub — cyb3rops · 2026-09-29
- Claude says monogamy and having children carry no moral value over polyamory — kevinnbass · 2026-09-29
- Hand-drawn-style animation coded directly in HTML5 Canvas by Claude Opus 5.5 — Ror_Fly · 2026-09-29
- Sonnet 5.5 cache reads cost as much as Opus, undercutting its agent appeal — StewartalsopIII · 2026-09-29