Ox Alpha scores 96% on SWE-bench, author skeptical
No_Tip9917 · reddit · 2026-08-22
The author benchmarked the free Ox Alpha model on the SWE-bench Verified Mini (50 tasks) using the official scaffold, achieving a suspiciously high 96% resolution rate (48/50). This beats Claude Fable 5 (95%) and far exceeds peers (77%) under identical conditions. The author urges skepticism due to: (1) Dataset limitation (Django/Sphinx are over-represented in training data), (2) Small sample size (n=50), (3) Stale tasks (2019-2022 PRs), and (4) Uncontrolled free endpoint. Despite potential inflation, it serves as an extreme case study for open-source coding agents.
More from coding & agent
- MiniMax-H3 video inpainting ported to diffusers modular blocks: 6-step subject swap — linoy_tsaban · 2026-08-22
- AI Social Media Agent Shows 10x Higher Engagement Than Human — RichardsonDx · 2026-08-22
- Fable Workflow: Orchestrate via High, Delegate to Subagents — dotey · 2026-08-22
- Gemini 3.7 Flash Ties for #1 in CAD Computer-Use Benchmark — DynamicWebPaige · 2026-08-22
- Middle Manager Pattern: Multi-Agent Concurrent Implementation — vinvan · 2026-08-22
- Stein on Latest in Agentic Engineering and 2026 Startup Ideas — ycombinator · 2026-08-22