Both labs' benchmark tables say they win — because they barely overlap on benchmarks
Ahmiii_83 · reddit · 2026-09-10
astra and fable 5.1 launched three days apart, and each lab's benchmark table shows a clean sweep. The author concludes neither is fudging: OpenAI's table leans on computer use (osworld) and math (frontiermath t4), Anthropic's on coding/terminal (terminal bench 4.0, cursorbench, swe bench pro) — almost no overlap, so there's no head-to-head to read.
The harness discrepancy is more troubling: OpenAI's 99.9% on arc-agi-3 came from its own provider adapter keeping reasoning state between actions; under ARC's standard harness the same model scored 62.7%. Artificial Analysis' unified setup gives fable 66 vs astra 61 on intelligence index, 70 vs 67 on coding agents — a 5-point gap, not the blowout either table implies.
Precedents cited: the Leaderboard Illusion paper (arxiv 2504.20879) documented undisclosed private testing on chatbot arena (27 llama 4 variants, 100 Elo inflation from best-of selection); Scale AI's gsm1k (1250 fresh problems) dropped some model families up to 13 points, suggesting memorization; mmlu is saturated above 88%. The takeaway: since tables are built by the measured party, "who ran it and under what conditions" beats "what did it score."
Related event: Astra and Fable 5.1 Benchmarks Barely Overlap(2 posts)→
More from Models
- Leike: Opus 4 jailbreak mitigations took over a year — safety can't start late — janleike · 2026-09-11
- DeepSeek appears to break its V<N> naming convention, new architecture said to train more stably — stochasticchasm · 2026-09-11
- DeepSeek's New Model: 4x Smaller KV Cache Than DSV4-Flash and More Stable Training — stochasticchasm · 2026-09-11
- Uncensored CyberTiel 35B-A3B Beats Opus 4.6 on Real Codebase Issues at 27% of Qwen's Time — peculiar-ragdoll · 2026-09-11
- GLM-5.3 dominates CyberGym: top two cyber agents both run it, four of top six — pcuenq · 2026-09-11
- gpt-live brings SOTA voice to everyone: build your own codex-voice at $0.05/min — pbbakkum · 2026-09-11