8B Model's 81.6% Coding Score Questioned: It's a Verifier, Not a Solver
Critics point out that an 8B model's 81.6% DeepSWE and 87.6% Terminal-Bench 2.1 scores are misleading because they come from using the model as a verifier of other models' answers, with random baselines reaching 73.7%.
2026-09-24 ~ 2026-09-24 · 2 related posts
- 8B model's 81.6% DeepSWE score is misleading — it's verifier, not generator, says researcher — yifeiwang77 · 2026-09-24
- DeepSWE's 81.6% Coding Score Misleading — Random Baseline Hits 73.7% — yifeiwang77 · 2026-09-24