Dev reality check: hours babysitting frontier model errors vs X hype of superhuman AI

mervenoyann · x · 2026-09-13

A reply describing the daily reality of working with frontier models: N hours spent babysitting them in frustration at frequent errors and slop, interrupted by M-minute breaks browsing X posts claiming the same models — especially the next generation — are superhuman at everything. Another researcher (@mervenoyann) says she has the same experience, highlighting the gap between hands-on use and online hype.

Related event: Practitioners say frontier models are useful but far from superhuman(5 posts)→

Original post →

More from Models

Models channel →