Evaluating models on HellaSwag in this day and age? Researcher pokes fun at stale benchmarks

michellechen · x · 2026-09-30

A one-line jab at teams still evaluating models on HellaSwag, a 2019-era commonsense benchmark long saturated by frontier models. The quip highlights how evaluation practices lag far behind model capabilities.

Original post →

More from Fun

Fun channel →