Ox Alpha scores 96% on SWE-bench, author skeptical

No_Tip9917 · reddit · 2026-08-22

The author benchmarked the free Ox Alpha model on the SWE-bench Verified Mini (50 tasks) using the official scaffold, achieving a suspiciously high 96% resolution rate (48/50). This beats Claude Fable 5 (95%) and far exceeds peers (77%) under identical conditions. The author urges skepticism due to: (1) Dataset limitation (Django/Sphinx are over-represented in training data), (2) Small sample size (n=50), (3) Stale tasks (2019-2022 PRs), and (4) Uncontrolled free endpoint. Despite potential inflation, it serves as an extreme case study for open-source coding agents.

Original post →

More from coding & agent

coding & agent channel →