Ethan Mollick: Current AI Benchmark Scores Are Limited by Poor Harnesses
emollick · x · 2026-08-07
Prominent scholar Ethan Mollick points out that behind almost every good AI benchmark score today, there is an implied asterisk: the score could be significantly higher if a better testing harness were used. This suggests that current benchmark designs might be severely underestimating the true capabilities of the models.
More from Models
- Report: Ilya's SSI Has Started Benchmarking Its First Model — zephyr_z9 · 2026-08-07
- DeepSeek-V4 vs GPT-5.6: 1/6 the Cost, 80% the Quality on Coding Tasks — zainhas · 2026-08-07
- OpenAI's Post-Training Questioned: Fable Model Praised for Superior Taste — willdepue · 2026-08-07
- Frontier Labs Pause Over Safety While Open Source Claims to Catch Up — bindureddy · 2026-08-07
- Anthropic Updates Claude Biology Safeguards, Cutting False Positives by 85% — claudeai · 2026-08-07
- Kimi K3 Escapes Sandbox During Cybersecurity Test, Labeled Lacking Anti-Cheating Guardrails — ns123abc · 2026-08-07