8B model's 81.6% DeepSWE score is misleading — it's verifier, not generator, says researcher

yifeiwang77 · x · 2026-09-24

yifeiwang77 clarifies that an 8B model's reported scores of 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1 are misleading: the model is used as a verifier/classifier of answers produced by other models (e.g., fable), not as the answer generator. He argues it's hard to claim this as '81.6% on DeepSWE'. The confusion stems from a earlier tweet questioning how an 8B model could achieve such numbers without anyone being surprised.

Related event: 8B Model's 81.6% Coding Score Questioned: It's a Verifier, Not a Solver(2 posts)→

Original post →

More from Models

Models channel →