AgentTime benchmark: GPT-6 Astra nails runtime control 63% of the time, Claude Fable 5.1 just 4%
maksym_andr · x · 2026-10-08
AgentTime, a new benchmark with paper, code and site, tests LLM agent time awareness across 222 tasks from 18 benchmarks: duration following, forecasting, and retrospection.
Key findings:
- Runs ending within ±5% of requested duration: GPT-6 Astra 63%, GPT-5.6 Sol 39%, Claude Fable 5.1 just 4%.
- Timing error: Astra 1.18×, Sol 1.77×, Fable 5.1 2.86× — asked to work 4h, Fable usually finishes before 1h.
- Authors attribute Astra's edge to training on strictly time-bounded tasks, possibly an instrumental skill learned incidentally.
Time awareness, long overlooked, matters greatly for long-running agents.
Related event: AgentTime Benchmark Measures Temporal Perception in LLM Agents(3 posts)→
More from Models
- Claude Max users reminded to claim their monthly API credits — arthurcolle · 2026-10-09
- Anthropic impresses on efficiency but users feel "babysat" by its guardrails — kimmonismus · 2026-10-09
- LightOnOCR-3 adds visual grounding and image description to document parsing — IgorCarron · 2026-10-09
- COLM paper: "forks in the road" in post-training data shrink reasoning model coverage — ParshinShojaee · 2026-10-09
- Google's open medical VLM MedGemma publishes in Nature Medicine — ymatias · 2026-10-09
- Researcher asks: do model providers' ToS allow training on business/API customer content via synthetic data? — RylanSchaeffer · 2026-10-09