DeepSWE's 81.6% Coding Score Misleading — Random Baseline Hits 73.7%

yifeiwang77 · x · 2026-09-24

Yi Wang flags the widely cited 81.6% agent coding number as misleading: the model served as a verifier/classifier of answers produced by another model (e.g. Fable), not as the answer generator, so claiming "81.6% on DeepSWE" is hard to justify.

He adds that a random selection baseline already achieves 73.7% on this task, further undermining the significance of the headline score.

Related event: 8B Model's 81.6% Coding Score Questioned: It's a Verifier, Not a Solver(2 posts)→

Original post →

More from Models

Models channel →