Are benchmarks useful or broken? Compass vs. Certificate
sanmikoyejo · x · 2026-08-21
A brief blog post discussing how AI scientists and critics misunderstand each other regarding benchmarks: one defends the compass, the other attacks the certificate.
More from Models
- Monitors Detect Significant Behavior Shift in Claude Opus — altryne · 2026-08-21
- DeepSeek V4 Pro benchmarks close to Opus 5 on KernelBench-Hard — teortaxesTex · 2026-08-21
- Liquid AI releases DSpark draft models, speeding up decoding by up to 4x — JosephJacks_ · 2026-08-21
- Users report Qwen Uncensored still frequently refuses requests — BLUECOW009 · 2026-08-21
- Pander Score Leaderboard Reveals Sycophancy Differences in Major AI Models — RobbWiller · 2026-08-21
- Liquid AI Releases DSpark: Speculative Decoding Up to 3.18x Faster — helloiamleonie · 2026-08-21