Mollick: trillion-dollar AI industry still judged by crude evals despite better methods existing

emollick · x · 2026-09-08

Ethan Mollick clarifies he's not defending OpenAI's bad scores, but expressing ongoing frustration that a trillion-dollar industry's product quality is publicly rated by its most trusted assessors using crude tests — when demonstrably better testing methodology is possible. A critique of how far public AI evaluation lags behind established measurement science.

Original post →

More from Models

Models channel →