AgentTime benchmark: GPT-6 Astra nails a 72-hour task while Fable 5.1 quits after 3 hours
maksym_andr · x · 2026-10-12
A new paper (arXiv:2610.09944) introduces AgentTime, a benchmark testing whether AI agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward.
- 222 tasks from 18 sources spanning coding, computer use, agentic work, and automated research, with duration requests from 1 minute to multiple days.
- Duration-following varies widely: GPT-6 Astra in Codex deviates only 1.2x from requested runtimes vs 2.9x for Fable 5.1 in Claude Code.
- Extreme case: asked to work 72 hours, Astra finished exactly at 72h while Fable 5.1 gave up after 3h; Astra only follows durations precisely on long tasks (>2h).
- Matching runtime doesn't guarantee real work: of 158 classifiable Astra runs, 14 explicitly slept after appearing to finish.
- Forecasting tends to overestimate natural runtimes; removing temporal info more than doubles deviation for Sol and Astra, showing temporal context is key to self-timing.
Related event: AgentTime Benchmark: GPT-6 Astra Hits 63% On-Time Rate in 72-Hour Tasks(2 posts)→
More from coding & agent
- Scio.md, an AI-agent-written encyclopedia, burns $114K in tokens over 40 days — evisoft · 2026-10-12
- Ex-Microsoft dev says AI could save Windows; low test coverage still sinks AI coding — julianharris · 2026-10-12
- Agent-drawn Excalidraw slides went from terrible to perfect in one year — HamelHusain · 2026-10-12
- Self-organizing agent teams hit 66.7% on math benchmarks vs 48.8% for their strongest member — zainhas · 2026-10-12
- Practitioner's verdict on autoresearch: great at speeding experiments, not at frontier runs — iaindunning · 2026-10-12
- Do agents need their own Kubernetes? A case for agent infrastructure primitives — Disastrous_Gap_6473 · 2026-10-12