Lab pushes back on benchmaxxing claims, citing small public-private leaderboard gap
antoine_chaffin · x · 2026-10-07
antoinechaffin responds to benchmaxxing accusations: his model kept top-1 on a brand-new leaderboard and shows a small gap between public and private splits — "some models collapse, ours just does not flinch." He offers the public/private performance gap as evidence its enhancements aren't overfit to benchmarks. The post lacks specifics on which model or leaderboard.
Related event: Decider Model Author Defends Against Benchmaxxing Claims(3 posts)→
More from Models
- Many Mistral Large 4 failures traced to reasoning mode not being enabled — qtnx_ · 2026-10-07
- Early Opus 5.5 user says hype is overblown: shortcuts, wrong assumptions, sloppy work — haider1 · 2026-10-07
- TypeSafe's Jev model bets on machine-native intelligence over text-optimized LLMs — TWIML AI Podcast · 2026-10-07
- llm-mistral 0.16 adds reasoning model support for Mistral Large 4 — Simon Willison · 2026-10-07
- Ollama hosts Google's EmbeddingGemma 2, a 740M multimodal embedding model for on-device use — ollama · 2026-10-07
- Reddit users grow frustrated with ChatGPT's over-refusals on innocuous prompts — Crixusgannicus · 2026-10-07