Dev reality check: hours babysitting frontier model errors vs X hype of superhuman AI
mervenoyann · x · 2026-09-13
A reply describing the daily reality of working with frontier models: N hours spent babysitting them in frustration at frequent errors and slop, interrupted by M-minute breaks browsing X posts claiming the same models — especially the next generation — are superhuman at everything. Another researcher (@mervenoyann) says she has the same experience, highlighting the gap between hands-on use and online hype.
Related event: Practitioners say frontier models are useful but far from superhuman(5 posts)→
More from Models
- kalomaze: models learn narrow RLVR fact-checks instead of asserting only in-context provable claims — kalomaze · 2026-09-13
- kalomaze: Opus 5 is "such a bad model" — kalomaze · 2026-09-13
- Wenhu Chen can't even understand many questions in AA-intelligence AI benchmarks — WenhuChen · 2026-09-13
- Commenter: The OpenAI Navier-Stokes proof would be hailed as a breakthrough if posted anonymously — skdh · 2026-09-13
- GPT-6 Astra skips the chat box: Plus users get just 5-45 messages per 5 hours, locked to Work and Codex — 新智元 · 2026-09-13
- LLM Architecture Gallery: one chart to compare every major LLM design — zainhas · 2026-09-13