xAI engineer: Grok 4.7 is no longer the bottleneck — agent harness reliability is
elonmusk · x · 2026-09-28
- xAI engineer Larsen says user feedback on his "Week 1 at SpaceXAI" post barely asked for a smarter model — people instead complained that agents don't finish work: stopping midway, losing browser state, forgetting context, or getting stuck waiting for the user.
- His take: the model is no longer the bottleneck; harness reliability is. Tools, browser, memory, state, retries, and error handling offer many failure modes even when the model knows exactly what to do.
- He says he's actively working on these reliability problems, and Elon Musk reposted the thread.
More from coding & agent
- Harmony Launches 100+ AI Agents in Slack for IT and HR Automation — Codeblix_Ltd · 2026-09-28
- Open-source Editable Design: real text, independent layers, no more pixel-locked AI graphics — jiqizhixin · 2026-09-28
- INT21's agent-built factory ships 20 Rust inference engines across 7 model types — bingxu_ · 2026-09-28
- pi-demo plugin lets the Pi agent record, edit, and ship screen-capture reels — solyarisoftware · 2026-09-28
- Dev Built AI Agents That Find a Business Idea and Run the Whole Operation — Ok-Werewolf-7293 · 2026-09-28
- Devin user at $100k+ run-rate questions if coding agent revenues will stick — brandon_galang · 2026-09-28