Albanie flags big jump in models' no-CoT task time horizon
SamuelAlbanie · x · 2026-09-04
DeepMind researcher Samuel Albanie shared a time-horizon evaluation chart and noted a "quite a big jump" in models' no-CoT (no chain-of-thought) time horizon — the length of tasks models can complete autonomously without reasoning steps — signaling another notable capability leap.
Related event: Researcher flags big jump in models' no-CoT time horizon(3 posts)→
More from Models
- Gary Marcus predicts new model will be real improvement but show benchmaxxing signs — GaryMarcus · 2026-09-04
- GPT-6 Astra Tops ValsAI Code Migration Benchmark at 68% Accuracy, 2-4x Faster — sandersted · 2026-09-04
- Andrew Ng: We need harder evals — have frontier models chat with me — andrewgwils · 2026-09-04
- Epoch AI launches FrontierMath Erdős benchmark of 68 unsolved problems; Astra tops it at 2/68 — keviv9 · 2026-09-04
- OpenAI Launches GPT-6 Astra: 99.9% on ARC-AGI-3, but Independent Evals Call It Uneven — Latent Space · 2026-09-04
- Liquid AI launches Nanos: task-specific 350M-2.6B models that run on-device — JosephJacks_ · 2026-09-04