Benchmarking 40 long-horizon coding tasks cost $50k in one day, pricing universities out of AI evals

himanshustwts · x · 2026-10-03

Anand (anandnk24) reports the team spent $50k in a single day benchmarking just 40 long-horizon coding tasks, predicting universities will get priced out of AI evals work.

The core argument: as agent tasks get harder, trajectories grow longer and require far more tokens, driving up inference and compute costs dramatically. Academic teams relying on university grants can't sustain that level of spend, so evals may increasingly concentrate in well-funded labs and companies.

Original post →

More from coding & agent

coding & agent channel →