METR's new eval report gains traction over models losing track of tasks

isidentical · x · 2026-08-27

Several AI figures are sharing METREvals' latest report and its screenshots, discussing how models tend to lose track during long-horizon tasks. One sharer jokes that, ironically because that's the very failure mode being blamed on the models, they themselves lost track while reading it. The report is seen as a key evaluation of current frontier models' limitations on long tasks.

Original post →

More from Safety

Safety channel →