Critic: eval answers are typical samples while real-world answers are quirky

voooooogel · x · 2026-09-26

The author argues that in a set of eval questions, real-world answers tend to be quirky, interesting, or unusual, while the eval's canonical answers are central, most-common examples — a systematic mismatch. He verified the claim with a blind Opus 5.5 (high thinking) judgment on one example. The observation questions whether such evals truly measure real-world model capability.

Related event: Developers Slam Benchmark Answers as 'Central' While Real Answers Are Quirky(2 posts)→

Original post →

More from Research

Research channel →