Researchers flag a big jump in no-CoT task time horizon
burny_tech · x · 2026-09-04
DeepMind researcher Samuel Albanie noted a "quite a big jump" in a model's no-CoT task time horizon. Aryaman Arora reposted, quipping that interpretability analysis on this model "will be glorious."
Related event: Researcher flags big jump in models' no-CoT time horizon(3 posts)→
More from Models
- Benchmarks let people run with preferred narratives — capability vectors beat leaderboards — GlenBradley · 2026-09-04
- Meta's Small Model Muse Spark Outranks Astra DeepSwe as Unreleased Model Climbs Benchmarks — altryne · 2026-09-04
- Dev slams OpenAI's tiered rollouts as betrayal of its founding 'open' mission — ctjlewis · 2026-09-04
- IFM's new K2-Horizon-MoVA-36B-A4B draws scrutiny: real deal or benchmaxxed? — edward-dev · 2026-09-04
- Dev spends $40 on classifier evals to cut costs: 'hard to use AI when you can't afford intelligence' — zeeg · 2026-09-04
- GPT-6 Astra vs Claude Fable 5.1: Benchmarks So Far Show a Toss-Up at $10/$50 — DataLearnerAI · 2026-09-04