Mollick: trillion-dollar AI industry still judged by crude evals despite better methods existing
emollick · x · 2026-09-08
Ethan Mollick clarifies he's not defending OpenAI's bad scores, but expressing ongoing frustration that a trillion-dollar industry's product quality is publicly rated by its most trusted assessors using crude tests — when demonstrably better testing methodology is possible. A critique of how far public AI evaluation lags behind established measurement science.
More from Models
- Stanford's Rishi Bommasani Slams Frontier AI Labs for Treating Research Math as a Benchmark — RishiBommasani · 2026-09-08
- User claims Opus is now at 30B-level after Fable 5.1 release — ccerrato147 · 2026-09-08
- Same prompt head-to-head: Gemini 3.8 Flash vs GPT-6 Astra compared — iamfakhrealam · 2026-09-08
- Claude Code 20x Max at $400/month equals ~$16,000 of API usage — tedddyoweh · 2026-09-08
- AA index updated: Fable 5.1 and GPT-astra now tied — jietang · 2026-09-08
- Video editor says GPT-6 Astra rebuilt his full Premiere Pro edit in minutes from raw footage — _AustinCalvert_ · 2026-09-08