Niche benchmarks expose huge gaps: best model scores just 5.5% rebuilding a codebase from a binary
Informal-Trouble2183 · reddit · 2026-09-06
A Reddit thread rounds up three 'deep capability' software engineering benchmarks where frontier models diverge far more than on mainstream leaderboards:
- Program-Bench: given only a compiled binary and docs, agents must build a full codebase reproducing the program's behavior—no decompilers, no internet. GPT-6 Astra scored just 5.5%, Fable 5.1 led at 7%, Kimi K3 2%, GLM 5.3 and GPT 5.6 Sol 1.5% each, Qwen3.8 27b and GPT 5.6 Luna 0%.
- SRE-Bench (vals.ai): understanding a real-world binary without source. GPT-6 Astra 88%, GPT-5.6 Sol 55.9%, Claude Opus 5 (max) only 12.5%.
- Code Migration (vals.ai): reimplementing working programs in another language. GPT-6 Astra 67.7%, Fable 5.1 54.6%, GLM 5.3 44.2%, Qwen3.8 27b 14.2%.
The author argues that binary-level comprehension is the real test of deep intelligence, and models diverge dramatically on it.
More from Models
- SimpleBench results show AI models beating humans on common sense — DigSignificant1419 · 2026-09-07
- Why is nobody talking about Tencent Hy4, the most-used model on OpenRouter? — cantor8 · 2026-09-07
- Anthropic says Claude wrote the longest math proof ever, cracking a 358-year-old problem — basedjensen · 2026-09-07
- Similarweb: ChatGPT's AI traffic share falls from 73.3% to 55.5% in 12 months — gaganghotra_ · 2026-09-07
- GPT-6 Astra (and Pro?) spotted on Simple-Bench leaderboard — From_Internets · 2026-09-07
- Testing Gemini as music understanders: Pro 3.1 solid, Flash models hallucinate sounds — teropa · 2026-09-07