A practical AGI test: agents multitasking across 10+ tasks without steering

annbordetsky · x · 2026-08-28

A practitioner proposed a new benchmark for AGI: it arrives when agents can handle more than 10 tasks at once without human steering.

He observed that today's models often operate like a last-in-first-out stack — newly inserted tasks take priority and long-term vision gets lost, hampering planning across long engineering tasks. Asking agents to take notes helps somewhat but remains an incomplete solution.

Original post →

More from AGI Musings

AGI Musings channel →