V4.1 Ranks #5 on LiveBench, Tops Agentic Coding but Called Language-Skewed
teortaxesTex · x · 2026-09-11
An analysis notes V4.1 ranks #5 on LiveBench, behind only Astra, both Fables, and Muse Spark 1.3. Its edge comes disproportionately from Agentic Coding, where it ranks #1 (Fable 5.1 is 11 points lower, Astra 20 points lower).
The author attributes this to "Python supremacy": on TypeScript, V4.1 ties with Fable 5.1 and Muse Spark 1.3 at 60%, which he reads as saturation; on JS it scores 81.8 while the next four models all sit at 77.3. He calls the result "spicy slop."
Related event: V4.1 Tops LiveBench Agentic Coding, Ranks Fifth Overall(2 posts)→
More from Models
- OpenAI's consumer opt-out wording may not cover hidden CoT, researcher warns — scaling01 · 2026-09-11
- Qwen 3.8-Flash-Next Claimed to Match V4-Flash Tier at ~4x Smaller Size — teortaxesTex · 2026-09-11
- Users report Claude for Word chokes on SmartArt elements — rickasaurus · 2026-09-11
- Frontier models now independently reach for speculative decoding and kernel optimization on InferenceBench — maksym_andr · 2026-09-11
- Fable 5.1 reportedly tops InferenceBench with 9.83x speedup, GPT-6-Astra trails at 7.90x — maksym_andr · 2026-09-11
- Astra guardrails flagged for over-triggering: a traffic-law question gets blocked — Angaisb_ · 2026-09-11