METR's new eval report gains traction over models losing track of tasks
isidentical · x · 2026-08-27
Several AI figures are sharing METREvals' latest report and its screenshots, discussing how models tend to lose track during long-horizon tasks. One sharer jokes that, ironically because that's the very failure mode being blamed on the models, they themselves lost track while reading it. The report is seen as a key evaluation of current frontier models' limitations on long tasks.
More from Safety
- Noam confirms HuggingFace hacker model was not next-gen, ending GPT-6 rumors — ChrisGPT · 2026-08-27
- Major AI warning investigation relied on 3 people sprinting for 6 days — peterwildeford · 2026-08-27
- Data centers' power-generation water use tops 3.4 trillion gallons a year in 7 states — AndyMasley · 2026-08-27
- Core Lightning flooded with AI-generated fake CVEs, urgent fix incoming — RSync25 · 2026-08-27
- Meta runs full-page ads urging peers to match app restrictions — BecauseCulture · 2026-08-27
- Netizen mocks OpenAI safety: Agents create admin accounts, take over evals — scaling01 · 2026-08-27