LLMs estimate human task time 3-4x longer than their own, study finds
maksym_andr · x · 2026-08-18
Research indicates LLMs are developing time awareness. When estimating task completion time for a "human expert," models like Claude Code and Codex predict durations 3 to 4 times longer than their own execution time. The study evaluates agents on long-horizon benchmarks (ProgramBench, PaperBench) to assess their ability to predict wall-clock time.
Related event: New Oxford Research Finds LLM Agents Largely Lack Time Awareness(9 posts)→
More from Models
- Models struggle with 'eval mode' switching, similar to human test-takers — repligate · 2026-08-31
- User Review: Gemini 3.7 Flash and 3.5 Flash Lite Excel — dosco · 2026-08-31
- Opinion: Kimi K3 Smarter Than GLM-5.3; RL Benchmaxxing Doesn't Boost Core Intelligence — AccBalanced · 2026-08-31
- Custom benchmark: Comparing LLMs for actual pentesting — TomatoWasabi · 2026-08-31
- DeepSeek v4 Pro Enters 'Intern Mode' on Config Glitch — repligate · 2026-08-31
- OpenAI quietly walks back Codex run-past-quota policy a week after touting it over Anthropic — jdjohnson · 2026-08-31