Ox Alpha mystery model scores ~63% on full DeepSWE, on par with GPT-5.6 Sol mid

kimmonismus · x · 2026-08-23

@davis7 completed a full DeepSWE run on the mystery Ox Alpha model, ending at 63% — far below the 80% from his first subset test — and roughly on par with GPT-5.6 Sol mid.

His impressions: better "voice" than Claude or GPT, decent design taste, handles subagents and long complex work well, good code quality. Downsides: sometimes leaves dead code around (less thorough than Sol), and despite decent TPS it feels slow, especially at higher reasoning levels.

@kimmonismus adds: if it really is GLM-5.3 Flash running at Sol-mid level locally on a DGX Spark, that's a fantastic deal — a 24/7 local model with power as the only cost would be a game changer.

Related event: Ox Alpha's Full DeepSWE Run Scores About 63%, Debunking Earlier 80% Subset Claim(5 posts)→

Original post →

More from Models

Models channel →