Small public-private leaderboard gap cited as proof of no benchmaxxing
antoine_chaffin · x · 2026-10-07
antoinechaffin rebuts benchmaxxing accusations: besides holding top-1 on a brand-new leaderboard, his model shows a tiny gap between public and private splits, arguing this proves genuine generalization — "some models collapse, ours does not flinch."
More from Models
- Mistral claims Large 4 is one of the world's strongest AI models for cybersecurity — scaling01 · 2026-10-07
- Mistral Large 4 solves 18 of 19 CTF challenges in official speedrun with tool calls — MistralAI · 2026-10-07
- Community poll tiers AI labs: Anthropic and OpenAI frontier, Mistral and Amazon judged 3 generations behind — NathanpmYoung · 2026-10-07
- User: Claude Opus Hit Usage Limit Just 2 Hours Into an Overnight Plotting Run — Sauers_ · 2026-10-07
- Audit finds 206 verifier bugs in Zapier's AutomationBench, changing 27.9% of grades — omarsar0 · 2026-10-07
- Aplomb 1: open-weights 5.3B decision model with 1M context, tops 4B class on Decision Index — empiriolabsai · 2026-10-07