AgentTime Benchmark Measures Temporal Perception in LLM Agents

A new paper introduces AgentTime, a benchmark using 222 tasks from 18 benchmarks to test LLM agents' temporal perception across duration following, prediction, and retrospection, with Google's Astra substantially outperforming other models.

2026-10-08 ~ 2026-10-09 · 3 related posts