Critic: eval answers are typical samples while real-world answers are quirky
voooooogel · x · 2026-09-26
The author argues that in a set of eval questions, real-world answers tend to be quirky, interesting, or unusual, while the eval's canonical answers are central, most-common examples — a systematic mismatch. He verified the claim with a blind Opus 5.5 (high thinking) judgment on one example. The observation questions whether such evals truly measure real-world model capability.
More from Research
- MIT study: aging brains keep robust language networks despite cognitive decline — DrKavner · 2026-09-26
- Déjà View, a NeurIPS Oral: one looped transformer block matches 3D reconstruction models 8-10x its size — ZGojcic · 2026-09-26
- Google's PageBreak AI scanner confirms XSS bugs in running environments, finds 500+ with near-zero false positives — moyix · 2026-09-26
- Japanese used bookstores see 5x sales surge as AI firms buy books by the ton — stayndarsh · 2026-09-26
- Places Library debuts: 100 high-fidelity real-world 3D environments for embodied AI — yshan2u · 2026-09-26
- How much watermark signal fits in AI text? A developer works the math end to end — OtherwisePush6424 · 2026-09-26