8B Model's 81.6% Coding Score Questioned: It's a Verifier, Not a Solver

Critics point out that an 8B model's 81.6% DeepSWE and 87.6% Terminal-Bench 2.1 scores are misleading because they come from using the model as a verifier of other models' answers, with random baselines reaching 73.7%.

2026-09-24 ~ 2026-09-24 · 2 related posts