LiveBench agentic coding eval questioned: outlier score rests on 4 Python issues
teortaxesTex · x · 2026-09-12
teortaxesTex challenges LiveBench's agentic coding eval: V4.1 posts an absurd outlier score (+2.7 standard deviations vs expected) despite the result resting on just 4 Python issues within a 1,270-item benchmark, run through SWE-Agent with a 50-step limit. He calls it a "clownmaxxing eval," notes Astra fails the eval yet can deconstruct it, and flags the related project as unmaintained.
More from Models
- Sentry Founder: Newer 'Smarter' Models Are Producing Worse Results — zeeg · 2026-09-12
- Dev's hands-on take: Astra feels like a downgrade from Sol — mertdumenci · 2026-09-12
- Dev: Astra is a downgrade from Sol unless you brute-force with massive parallelism — mertdumenci · 2026-09-12
- Gary Marcus slams OpenAI for weakening safety monitoring as White House stays silent — GaryMarcus · 2026-09-12
- Dev Fine-Tunes Qwen3.8-27B on 125K Real Conversations to Kill the AI Assistant Vibe — kvyb · 2026-09-12
- Muse, Instinct and Grok Bot 'would be best products of the year' in any other year — jeff_weinstein · 2026-09-12