V4.1 Ranks #5 on LiveBench, Tops Agentic Coding but Called Language-Skewed

teortaxesTex · x · 2026-09-11

An analysis notes V4.1 ranks #5 on LiveBench, behind only Astra, both Fables, and Muse Spark 1.3. Its edge comes disproportionately from Agentic Coding, where it ranks #1 (Fable 5.1 is 11 points lower, Astra 20 points lower).

The author attributes this to "Python supremacy": on TypeScript, V4.1 ties with Fable 5.1 and Muse Spark 1.3 at 60%, which he reads as saturation; on JS it scores 81.8 while the next four models all sit at 77.3. He calls the result "spicy slop."

Related event: V4.1 Tops LiveBench Agentic Coding, Ranks Fifth Overall(2 posts)→

Original post →

More from Models

Models channel →