Open-source ox-alpha tested near top closed models
Third-party testing shows open-source model ox-alpha achieves 58.4% on DeepSWE (66 of 113 tasks), on par with Claude Opus, with another subset evaluation reaching 63%—making it Pareto-optimal among open models and competitive with closed-source rivals like Grok.
2026-08-22 ~ 2026-08-22 · 2 related posts
- Open source model ox-alpha hits 63% on DeepSWE, rivals Grok efficiency — zainhas · 2026-08-22
- Ox-alpha DeepSWE benchmark reveals 58.4% pass rate, debunking 80% rumors — apples_jimmy · 2026-08-22