Agents beat human baselines on OS World, so why does Grok's real-world computer use worry experts?

thursdai_pod · x · 2026-08-26

While agents have passed 80% on the OS World benchmark, beating the human baseline, @nisten expressed concern after watching Grok Bot use a computer. The discussion with @altryne and @nisten aims to uncover the real state of computer use by agents, questioning whether they can use computers easily or if major limitations remain.

Original post →

More from coding & agent

coding & agent channel →