Open-source ox-alpha tested near top closed models

Third-party testing shows open-source model ox-alpha achieves 58.4% on DeepSWE (66 of 113 tasks), on par with Claude Opus, with another subset evaluation reaching 63%—making it Pareto-optimal among open models and competitive with closed-source rivals like Grok.

2026-08-22 ~ 2026-08-22 · 2 related posts