AgentTime benchmark: GPT-6 Astra nails a 72-hour task while Fable 5.1 quits after 3 hours

maksym_andr · x · 2026-10-12

A new paper (arXiv:2610.09944) introduces AgentTime, a benchmark testing whether AI agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward.

Related event: AgentTime Benchmark: GPT-6 Astra Hits 63% On-Time Rate in 72-Hour Tasks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →