Ethan Mollick: Current AI Benchmark Scores Are Limited by Poor Harnesses

emollick · x · 2026-08-07

Prominent scholar Ethan Mollick points out that behind almost every good AI benchmark score today, there is an implied asterisk: the score could be significantly higher if a better testing harness were used. This suggests that current benchmark designs might be severely underestimating the true capabilities of the models.

Original post →

More from Models

Models channel →