Benchmarking 40 long-horizon coding tasks cost $50k in one day, pricing universities out of AI evals
himanshustwts · x · 2026-10-03
Anand (anandnk24) reports the team spent $50k in a single day benchmarking just 40 long-horizon coding tasks, predicting universities will get priced out of AI evals work.
The core argument: as agent tasks get harder, trajectories grow longer and require far more tokens, driving up inference and compute costs dramatically. Academic teams relying on university grants can't sustain that level of spend, so evals may increasingly concentrate in well-funded labs and companies.
More from coding & agent
- AI agent dev burns 200k GitHub runner-minutes a month — how do you tame CI costs? — Hairy-Supermarket120 · 2026-10-03
- Dev builds a Claude-powered video generation pipeline overnight, ships two product marketing videos — cneuralnetwork · 2026-10-03
- toolcall-check: a Python CLI for testing nested tool calls and streaming on chat APIs — Arthur122103 · 2026-10-03
- Dev shares how he built AI agents that run marketing campaigns end to end — Few_Benefit_4853 · 2026-10-03
- GPT learns to pass tests, not to engineer: the RL reward-mismatch behind ugly code — xiaohu · 2026-10-03
- Fixing long-horizon task drift on local models with a deterministic state plugin — paulqq · 2026-10-03