Public benchmarks may crash in value as private evals become the secret sauce
sarahcat21 · x · 2026-09-25
A sharp industry take: expect a dramatic decrease in the number and possibly the value of public benchmarks. Practitioners increasingly realize your benchmark IS your secret sauce, and 'general' frontier tasks are hard to define versus what's actually frontier for your specific use cases — implying private evals will matter more than public leaderboards.
Related event: Ex-OpenAI Exec Says Public AI Benchmarks Will Lose Value(2 posts)→
More from Models
- Nagi-ENORMOUS beats Jev, Semif, Laya on Game Arena Benchmark — No_Skill_8393 · 2026-09-25
- Principia: video models know less high-school physics than you'd think — mariyaivasileva · 2026-09-25
- Post-scarcity intelligence defined: Opus 5.5-level smarts at $0.10 input / $0.50 output per million tokens — Sect-Sister-Reads-42 · 2026-09-25
- NetHack benchmark reportedly solved after years of slow LLM progress — PMinervini · 2026-09-25
- 'Capability gaslighting' is real—but not in agentic engineering, where code is consistent — PawelHuryn · 2026-09-25
- Power user's Claude subscription banned after burning 20x weekly quota daily — AlchainHust · 2026-09-25