What makes a good eval? A long post traces AI agent benchmarks from GPT-4 to today
dejavucoder · x · 2026-07-27
What makes an eval useful: a long post traces the benchmark landscape from GPT-4 to today
The post says it spent the last few weeks reviewing evals for AI agents across areas like human-AI multiplayer games, equity research, robotics, and long-horizon computer use.
It aims to answer three things:
- what an eval is actually useful for,
- how the eval landscape evolved from the GPT-4 era to now,
- and which six traits the surviving evals tend to share.
It also covers documented mistakes researchers have found in trusted evals, framing the post as both a taxonomy and a cautionary review of how to measure agent capability well.
More from Research
- NUS builds a soft force sensor that drives actuators without electronics or power — CurieuxExplorer · 2026-07-27
- Chelsea Finn says robot RL is bottlenecked by physical rollout cost, not algorithms — ycombinator · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27