GPT-6 Astra nails 72-hour agent tasks with 63% on-time rate, crushing Fable 5.1's 4%

maksym_andr · x · 2026-10-12

An analysis of 1,991 agent runs measuring how precisely models follow requested work durations.

The author notes precise duration-following only emerges on long tasks (>2 hours), and muses that early stopping to sync with the user might actually be the better behavior.

Related event: AgentTime Benchmark: GPT-6 Astra Hits 63% On-Time Rate in 72-Hour Tasks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →