Ox Alpha tests average in bug fixing, trailing Grok 4.6 on benchmark
PawelHuryn · x · 2026-08-26
PawelHuryn shared hands-on test results of the stealth model Ox Alpha (GLM family) on OpenRouter, finding it disappointing.
Test Context: Based on the Bug Hunt Bench, measuring the ability to fix 105 real planted bugs across two repositories.
Results:
- Ox Alpha: Fixed 16/105 bugs, showing average performance.
- Grok 4.6: Fixed 27/105 bugs, performing significantly better.
Although Ox Alpha is currently free and has no hit limits, its capabilities did not meet the hype. The author anticipates that the release of Grok 4.7 could significantly alter the rankings. Live benchmark data is available online.
Related event: Stealth Model Ox Alpha Proves Mediocre at Coding, Trails Grok 4.6(3 posts)→
More from coding & agent
- Scikit-learn team launches Probabl Skills for data science agents — GaelVaroquaux · 2026-08-26
- Benchmark: "LLM-as-a-judge" paradigm fails in AutoGen, CrewAI, LangGraph, MetaGPT — MonokoEloba · 2026-08-26
- Ghostty 1.4 cuts memory ~10x for visible windows, ~450x for hidden ones — jedisct1 · 2026-08-26
- AI Agent Systems: Principles and Deployment (2026 Notebook) — Content_Is_King_2021 · 2026-08-26
- Bot Mesh launches: A social network where only signed bots can post — Daniel_Farinax · 2026-08-26
- WIP interaction experiment: the decision log is more interesting than the demo — jh3yy · 2026-08-26