Surge AI Exec Reveals Why Benchmarks Fail: $15M Cost and Widespread Contamination

alex_verem · x · 2026-08-03

Surge AI VP of Product Nick Heiner gave a talk explaining why AI benchmark scores fail to match reality. He highlighted two main issues:

The article notes how tests like MMLU went from being gold standards to easily topped by frontier models (scoring above 88%), emphasizing the need to question what benchmark scores actually test.

Original post →

More from Models

Models channel →