AgentTime benchmark: GPT-6 Astra nails runtime control 63% of the time, Claude Fable 5.1 just 4%

maksym_andr · x · 2026-10-08

AgentTime, a new benchmark with paper, code and site, tests LLM agent time awareness across 222 tasks from 18 benchmarks: duration following, forecasting, and retrospection.

Key findings:

Time awareness, long overlooked, matters greatly for long-running agents.

Related event: AgentTime Benchmark Measures Temporal Perception in LLM Agents(3 posts)→

Original post →

More from Models

Models channel →