Both labs' benchmark tables say they win — because they barely overlap on benchmarks

Ahmiii_83 · reddit · 2026-09-10

astra and fable 5.1 launched three days apart, and each lab's benchmark table shows a clean sweep. The author concludes neither is fudging: OpenAI's table leans on computer use (osworld) and math (frontiermath t4), Anthropic's on coding/terminal (terminal bench 4.0, cursorbench, swe bench pro) — almost no overlap, so there's no head-to-head to read.

The harness discrepancy is more troubling: OpenAI's 99.9% on arc-agi-3 came from its own provider adapter keeping reasoning state between actions; under ARC's standard harness the same model scored 62.7%. Artificial Analysis' unified setup gives fable 66 vs astra 61 on intelligence index, 70 vs 67 on coding agents — a 5-point gap, not the blowout either table implies.

Precedents cited: the Leaderboard Illusion paper (arxiv 2504.20879) documented undisclosed private testing on chatbot arena (27 llama 4 variants, 100 Elo inflation from best-of selection); Scale AI's gsm1k (1250 fresh problems) dropped some model families up to 13 points, suggesting memorization; mmlu is saturated above 88%. The takeaway: since tables are built by the measured party, "who ran it and under what conditions" beats "what did it score."

Related event: Astra and Fable 5.1 Benchmarks Barely Overlap(2 posts)→

Original post →

More from Models

Models channel →