Stealth Model 'Ox Alpha' Hits 80% on DeepSWE Benchmark
AccBalanced · x · 2026-08-21
A stealth model named "Ox Alpha" achieved approximately 80% on a 10-task subset of the DeepSWE benchmark, significantly outperforming GPT-5.6-sol (52%) and Fable (65%).
Meanwhile, the Gemma community is celebrating, suggesting a potential new release or benchmark success for Gemma 3.
More from Models
- GLM 5.3 "Flash" spotted internally; observers doubt it's a small model — teortaxesTex · 2026-08-21
- Local Qwen 3.8 Benchmarks: 3-9% Failures Due to Infinite Reasoning Loops — on_line187 · 2026-08-21
- Internal Models Strong Yet Rely on Older US Base Models — teortaxesTex · 2026-08-21
- Mysterious Lab Expected to Rapidly Deliver Frontier-Level SWE Models — teortaxesTex · 2026-08-21
- Stealth Model Ox-Alpha Released, Outperforms Fable on SWE Benchmark — troll_khan · 2026-08-21
- 456 tok/s Qwen 3.8 on Modded RTX 2080 Ti via NInfer Port — xrailgun · 2026-08-21