V4.1 hits #5 on LiveBench, tops Agentic Coding by 20 points over Astra

teortaxesTex · x · 2026-09-11

teortaxesTex notes on Artificial Analysis rankings: V4.1 sits at #5 on LiveBench, behind only Astra, both Fables, and Muse Spark 1.3. The win is disproportionately driven by Agentic Coding, where V4.1 ranks #1 while Fable 5.1 is #11 and Astra trails by 20 points. The author jokes that Bindu "has a sense of humor," hinting at targeted benchmark optimization.

Related event: V4.1 Tops LiveBench Agentic Coding, Ranks Fifth Overall(2 posts)→

Original post →

More from Models

Models channel →