Ox-alpha DeepSWE benchmark reveals 58.4% pass rate, debunking 80% rumors
apples_jimmy · x · 2026-08-22
A full end-to-end run of the DeepSWE benchmark on ox-alpha revealed a pass rate of 58.4% (66/113 tasks), debunking rumors of an 80% score. This performance lands the model nearly identical to Claude Opus 4.8 (59%). The test noted that 9.7% of failures were due to tool-call format errors rather than reasoning shortcomings.
Related event: Open-source ox-alpha tested near top closed models(2 posts)→
More from Models
- Sentence Transformers v6.0 Released with Native ColBERT-Style Late Interaction — lateinteraction · 2026-08-23
- Why AI benchmarks often fail to reflect real-world performance — sargetun123 · 2026-08-23
- Ox Alpha benchmark test: 551 API calls needed to get 87 completed answers — anshulkundaje · 2026-08-23
- Users discuss perceived decline in LLM logic and coherence with specific examples — Original_Cry_3172 · 2026-08-23
- User criticism: "I cannot stand the way Claude writes" — BLUECOW009 · 2026-08-22
- Anima-3.8B and ComfyUI custom node released by lylogummy — AgeNo5351 · 2026-08-22