Casual prompts can't measure AI: data contamination breaks amateur evals

iamKierraD · x · 2026-09-21

Kevin Roose argues both "stochastic parrots" and "it's all hype" takes have been easily disprovable since 2023, and many people — including AI journalists — never tried the tools. Chomba Bupe counters that casual prompting isn't reliable either: poorly designed prompts ignore data contamination, so they can't establish real model capability. Both point to the same conclusion: only well-designed, first-hand evaluation tells you what models can actually do.

Original post →

More from AGI Musings

AGI Musings channel →