Logan Kilpatrick: Public benchmarks will shrink as firms build their own
OfficialLoganK · x · 2026-09-25
Logan Kilpatrick argues we'll see a dramatic decrease in the number and value of public benchmarks, as teams realize their benchmark IS their secret sauce and general frontier tasks are hard to design versus use-case-specific ones. He adds that companies building with AI will end up making the vast majority of public benchmarks, and "benchmarks being the secret sauce" isn't true for most people.
Related event: Ex-OpenAI Exec Says Public AI Benchmarks Will Lose Value(2 posts)→
More from Research
- Sharpa's World Synesthesia Model for dexterous hands accepted at CoRL 2026 — jiqizhixin · 2026-09-25
- DeltaWAM cuts video model cost for bimanual manipulation, lifting RoboTwin success to 85.4% — Han Yan · 2026-09-25
- Amazon open-sources Rufus-Air: full 8-stage post-training recipe on GLM-4.5-Air-Base — amazon · 2026-09-25
- Researcher proposes agent-suggests-human-executes loop for real-world science experiments — suragnair · 2026-09-25
- MLPerf Training v6.1 adds first LLM post-training benchmark: agentic RL on a 397B model — TheKanter · 2026-09-25
- Human-in-the-loop or machine-executed: verifiable tasks turn agent traces into RL rewards — suragnair · 2026-09-25