DeepSWE's 81.6% Coding Score Misleading — Random Baseline Hits 73.7%
yifeiwang77 · x · 2026-09-24
Yi Wang flags the widely cited 81.6% agent coding number as misleading: the model served as a verifier/classifier of answers produced by another model (e.g. Fable), not as the answer generator, so claiming "81.6% on DeepSWE" is hard to justify.
He adds that a random selection baseline already achieves 73.7% on this task, further undermining the significance of the headline score.
Related event: 8B Model's 81.6% Coding Score Questioned: It's a Verifier, Not a Solver(2 posts)→
More from Models
- Why Competing With Meta's Muse Is Hard: 40K Ratings at 4.88, Plus Distribution — FinanceYF5 · 2026-09-24
- OpenAI's MentalHealthBench: GPT-6 Astra Scores 57.3 vs GPT-4o's 32.1 — rohanpaul_ai · 2026-09-24
- Opus 5.5 Shows Off UI Flair: Checkmark Icons and Menu Polish on Its Own — vista8 · 2026-09-24
- Claude Opus 5.5 roasts every AI model and makes the whole video itself — bookwormengr · 2026-09-24
- Leaked naming: GPT-6 Sol equals GPT-5.6 Terra, Luna degradation confirmed — PawelHuryn · 2026-09-24
- Viral 'Opus 5.5 update' post fuels Anthropic release speculation — rudrank · 2026-09-24