15 models, 3,595 graded replies: two headline findings from the last benchmark didn't reproduce
nejcar20 · reddit · 2026-09-03
A print-shop automation team reran their benchmark at scale: 15 models, 8 cases, 30 generations each, Slovenian and English, 3,595 graded replies. Neither headline finding survived — Claude Haiku's infamous 9.00-vs-9.60 error never reproduced in 240 tries, nor did GPT-4o mini's 4x miss. Both were single samples published prematurely.
The replacement finding is worse: GPT-4o mini misreads the price-ladder boundary at 99 prints (quoting 29.70 instead of 31.68) 30 out of 30 times in both languages — a stable, quiet error invisible to human reviewers. Quiet errors hit 250 per thousand replies with zero loud ones.
Accuracy no longer separates models: 10 of 15 were perfect. The cheapest model (Ling 3.0 Flash at $0.000024/reply) went 236-for-236; Claude Haiku, also perfect, costs 89x more per reply.
Blind writing grades (120 replies, identity stripped) found only one discriminating dimension: showing your working. GPT-5.6 Luna — their production default — and Grok 4.3 both hand over bare totals with nothing to check. Caveats: showing-work correlates 0.85 with length, and Claude graded Claude. A grader bug (the "15" in 10x15 read as a price candidate) was also found and disclosed. The 30-generation design proved highly stable: 97% mean agreement within cells.
More from Models
- Ethan Mollick tests Gemini 3.8 Flash: fast but no match for Fable 5.1 on shaders — eldonredwards · 2026-09-03
- Meta's Spark 1.3 nears frontier and its scorched-earth pricing may threaten AI labs, argues investor — RihardJarc · 2026-09-03
- Claude Fable 5.1 (high) hits 92.3% on WeirdML, beating Fable 5 by 0.4% for new SOTA — teortaxesTex · 2026-09-03
- Would OpenAI bet $100M+ on Looped Transformer for Astra without scaling proof? — teortaxesTex · 2026-09-03
- Enterprises pay 10x-20x more to keep data out of AI training, suggesting a routing-layer startup — random_walker · 2026-09-03
- antirez: coding benchmarks are 'mostly crap' and badly misaligned with how models train — antirez · 2026-09-03