Ox Alpha mystery model scores ~63% on full DeepSWE, on par with GPT-5.6 Sol mid
kimmonismus · x · 2026-08-23
@davis7 completed a full DeepSWE run on the mystery Ox Alpha model, ending at 63% — far below the 80% from his first subset test — and roughly on par with GPT-5.6 Sol mid.
His impressions: better "voice" than Claude or GPT, decent design taste, handles subagents and long complex work well, good code quality. Downsides: sometimes leaves dead code around (less thorough than Sol), and despite decent TPS it feels slow, especially at higher reasoning levels.
@kimmonismus adds: if it really is GLM-5.3 Flash running at Sol-mid level locally on a DGX Spark, that's a fantastic deal — a 24/7 local model with power as the only cost would be a game changer.
More from Models
- Opus 5 Reportedly Rivals Anthropic's Internal Models, Excels at Optimization — scaling01 · 2026-08-23
- GLM-5.3 Outperforms Fable: +11 Points, Half the Cost — zainhas · 2026-08-23
- DeepSWE benchmark: GLM-5.3 matches Fable 5 at 1/4th the cost — zainhas · 2026-08-23
- MiniMax Music Model Noted for Missing Encoder — kalomaze · 2026-08-23
- Comparison: Grok provides wrong info often, Sol excels at challenging assumptions — jdjohnson · 2026-08-23
- Chinese Flash Models Criticized for Over-Reasoning Latency — oran_ge · 2026-08-23