8B model's 81.6% DeepSWE score is misleading — it's verifier, not generator, says researcher
yifeiwang77 · x · 2026-09-24
yifeiwang77 clarifies that an 8B model's reported scores of 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1 are misleading: the model is used as a verifier/classifier of answers produced by other models (e.g., fable), not as the answer generator. He argues it's hard to claim this as '81.6% on DeepSWE'. The confusion stems from a earlier tweet questioning how an 8B model could achieve such numbers without anyone being surprised.
Related event: 8B Model's 81.6% Coding Score Questioned: It's a Verifier, Not a Solver(2 posts)→
More from Models
- Why Competing With Meta's Muse Is Hard: 40K Ratings at 4.88, Plus Distribution — FinanceYF5 · 2026-09-24
- OpenAI's MentalHealthBench: GPT-6 Astra Scores 57.3 vs GPT-4o's 32.1 — rohanpaul_ai · 2026-09-24
- Opus 5.5 Shows Off UI Flair: Checkmark Icons and Menu Polish on Its Own — vista8 · 2026-09-24
- Claude Opus 5.5 roasts every AI model and makes the whole video itself — bookwormengr · 2026-09-24
- Leaked naming: GPT-6 Sol equals GPT-5.6 Terra, Luna degradation confirmed — PawelHuryn · 2026-09-24
- Viral 'Opus 5.5 update' post fuels Anthropic release speculation — rudrank · 2026-09-24