V4.1 hits #5 on LiveBench, tops Agentic Coding by 20 points over Astra
teortaxesTex · x · 2026-09-11
teortaxesTex notes on Artificial Analysis rankings: V4.1 sits at #5 on LiveBench, behind only Astra, both Fables, and Muse Spark 1.3. The win is disproportionately driven by Agentic Coding, where V4.1 ranks #1 while Fable 5.1 is #11 and Astra trails by 20 points. The author jokes that Bindu "has a sense of humor," hinting at targeted benchmark optimization.
Related event: V4.1 Tops LiveBench Agentic Coding, Ranks Fifth Overall(2 posts)→
More from Models
- OpenAI's consumer opt-out wording may not cover hidden CoT, researcher warns — scaling01 · 2026-09-11
- Qwen 3.8-Flash-Next Claimed to Match V4-Flash Tier at ~4x Smaller Size — teortaxesTex · 2026-09-11
- Users report Claude for Word chokes on SmartArt elements — rickasaurus · 2026-09-11
- Frontier models now independently reach for speculative decoding and kernel optimization on InferenceBench — maksym_andr · 2026-09-11
- Fable 5.1 reportedly tops InferenceBench with 9.83x speedup, GPT-6-Astra trails at 7.90x — maksym_andr · 2026-09-11
- Astra guardrails flagged for over-triggering: a traffic-law question gets blocked — Angaisb_ · 2026-09-11