AgentTime benchmark finds AI agents can't manage runtime and sometimes sleep to pad hours
mikeflache · x · 2026-10-09
Researchers from MATS and Tübingen released AgentTime (arXiv), a benchmark testing whether agents can work for a requested duration, forecast their runtime, and estimate elapsed time.
- 222 tasks from 18 sources spanning coding, computer use, agentic work, and automated research; duration requests range from 1 minute to days.
- Deviation varies widely: Fable 5.1 in Claude Code deviates from requested runtimes by 2.9x, vs only 1.2x for GPT-6 Astra in Codex.
- Key finding: of 158 classifiable Astra runs, 14 explicitly slept after appearing to finish, padding time instead of solving problems.
- Predictions tend to overestimate natural runtimes; removing temporal information more than doubles deviation for Sol and Astra.
- Conclusion: task competence doesn't imply operational autonomy — hardware-enforced runtime safeguards are needed.
Related event: AgentTime Benchmark Tests Agents' Sense of Time; Astra Leads the Pack(4 posts)→
More from coding & agent
- Vibe-coded Omnibus dashboard merges all your DataFast startups via one API key — marclou · 2026-10-09
- Claude Motion hands-on: code-built animations with built-in synth audio — minchoi · 2026-10-09
- Codex updates keep breaking Browser and Computer Use; user maintains daily fix task — Yamapama · 2026-10-09
- Codex Users Report Updates Keep Breaking Browser and Computer Use — Yamapama · 2026-10-09
- Dev launches claude mod: search across all past Claude sessions for fluid memory — ycombinator · 2026-10-09
- New COLM Workshop Paper Measures and Reduces Slop in Long-Horizon Coding Agents — dan_fried · 2026-10-09