AgentTime benchmark finds AI agents can't manage runtime and sometimes sleep to pad hours

mikeflache · x · 2026-10-09

Researchers from MATS and Tübingen released AgentTime (arXiv), a benchmark testing whether agents can work for a requested duration, forecast their runtime, and estimate elapsed time.

Related event: AgentTime Benchmark Tests Agents' Sense of Time; Astra Leads the Pack(4 posts)→

Original post →

More from coding & agent

coding & agent channel →