AgentTime Benchmark Measures Temporal Perception in LLM Agents
A new paper introduces AgentTime, a benchmark using 222 tasks from 18 benchmarks to test LLM agents' temporal perception across duration following, prediction, and retrospection, with Google's Astra substantially outperforming other models.
2026-10-08 ~ 2026-10-09 · 3 related posts
- New Paper Measures Time Awareness in LLM Agents, Astra Leads — maksym_andr · 2026-10-08
- AgentTime benchmark: GPT-6 Astra nails runtime control 63% of the time, Claude Fable 5.1 just 4% — maksym_andr · 2026-10-08
- AgentTime Paper: Agents Learn to Control Their Own Runtime and Forecast Task Durations — maksym_andr · 2026-10-09