Casual prompts can't measure AI: data contamination breaks amateur evals
iamKierraD · x · 2026-09-21
Kevin Roose argues both "stochastic parrots" and "it's all hype" takes have been easily disprovable since 2023, and many people — including AI journalists — never tried the tools. Chomba Bupe counters that casual prompting isn't reliable either: poorly designed prompts ignore data contamination, so they can't establish real model capability. Both point to the same conclusion: only well-designed, first-hand evaluation tells you what models can actually do.
More from AGI Musings
- 1/3 of S&P 500 earnings calls cite quantified AI use, claiming ~47% efficiency gains — Exponential View (Azeem Azhar) · 2026-09-21
- Tyler Cowen's doomsday scenario: US regulation hands the AI race to China — Afinetheorem · 2026-09-21
- AI Automation Is Destroying the Apprenticeship Layer Where Juniors Become Seniors — Hamza_StrategizeLabs · 2026-09-21
- "Now That Everyone Can Code, It's Clearer Why Many Shouldn't" — round · 2026-09-21
- AGI risk goes mainstream: 100+ furious Facebook comments on one post — danfaggella · 2026-09-21
- Box CEO Aaron Levie: 'some articles' can't keep you current on AI — victor_explore · 2026-09-21