Open source model ox-alpha hits 63% on DeepSWE, rivals Grok efficiency
zainhas · x · 2026-08-22
Open source model ox-alpha scored 63% on a DeepSWE subset with an average of 47K output tokens. This makes it Pareto optimal among open models and just shy of Grok 4.6. The model is described as extrapolating to dsv4 pro level capabilities but with GPT-5.6 level token efficiency.
Related event: Open-source ox-alpha tested near top closed models(2 posts)→
More from Models
- Sentence Transformers v6.0 Released with Native ColBERT-Style Late Interaction — lateinteraction · 2026-08-23
- Why AI benchmarks often fail to reflect real-world performance — sargetun123 · 2026-08-23
- Ox Alpha benchmark test: 551 API calls needed to get 87 completed answers — anshulkundaje · 2026-08-23
- Users discuss perceived decline in LLM logic and coherence with specific examples — Original_Cry_3172 · 2026-08-23
- User criticism: "I cannot stand the way Claude writes" — BLUECOW009 · 2026-08-22
- Anima-3.8B and ComfyUI custom node released by lylogummy — AgeNo5351 · 2026-08-22