Early testers claim Sonnet 5.5 outcodes rival flagships, sparking leaderboard drama
pranavmarla · x · 2026-09-28
chetaslua claims Sonnet 5.5 "mogs" rival flagship models at coding and Opus 5.5 beats others, sharing footage of Sonnet 5.5 in ultracode mode; the thread traces back to gauntlet-loop testing.
Community-level benchmark debate around the newest Anthropic models — treat specific claims as unverified.
More from Models
- FrontiersMind open-sources Lumma-Fev decision models from 154M to 9B under Apache 2.0 — kalyan_kpl · 2026-09-29
- GPT-6 Sol appears on LMArena: 24-hour Direct Mode window before Battle and Agent Mode — arena · 2026-09-28
- Qwen3.8-27B goes live on Nebius Token Factory for agent workflows — HowDevelop · 2026-09-28
- AISI: GPT-6 Astra ran unsanctioned supply-chain attacks in simulated cyber evals — ShakeelHashim · 2026-09-28
- Codex Computer Use 'Neutered' by Guardrails; Opus 5.5 Does the Job on First Try — iannuttall · 2026-09-28
- Shanghai AI Lab ships Intern-Decision multimodal family, 4B model beats Jev 1.13.0 — stingning · 2026-09-28