LiveBench agentic coding eval questioned: outlier score rests on 4 Python issues

teortaxesTex · x · 2026-09-12

teortaxesTex challenges LiveBench's agentic coding eval: V4.1 posts an absurd outlier score (+2.7 standard deviations vs expected) despite the result resting on just 4 Python issues within a 1,270-item benchmark, run through SWE-Agent with a 50-step limit. He calls it a "clownmaxxing eval," notes Astra fails the eval yet can deconstruct it, and flags the related project as unmaintained.

Original post →

More from Models

Models channel →