dlwh: every eval is broken, yet blindly hillclimbing on them drove stunning AI progress

dlwh · x · 2026-09-10

AI researcher dlwh offers a counterintuitive take: the second most unreasonably effective thing about AI is that despite virtually every eval being badly broken, blindly hillclimbing on those evals has still produced incredible progress — highlighting the paradox between flawed benchmarks and real capability gains.

Original post →

More from AGI Musings

AGI Musings channel →