Study: AI-Human Gap Widens Over Time in Long-Horizon Tasks
tobyordoxford · x · 2026-07-03
Researchers observed that performance improvement curves across different generations of AI models on long-horizon tasks show highly similar slopes, indicating limited scaling capabilities between generations. While next year's models will generally outperform this year's, the capability gap between models and humans is expected to keep widening as task timeframes extend. This finding provides crucial reference for evaluating AI's applicability in complex scenarios requiring long-term reasoning and continuous execution.
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11