Inference speed reshapes human-agent collaboration: 20s keeps you in flow, 10min is just delegation
demian_ai · x · 2026-08-28
The author argues that LLM/agent response speed changes the shape of human collaboration. If an agent returns an interface iteration in 20 seconds, you stay, look, change your mind, try something stranger — the LLM becomes part of your thought process. If it takes 10 minutes, it degrades into delegation: review later, issue another task.
This creates a demand paradox: faster inference doesn't necessarily reduce compute demand. The cost per answer falls while attempted thoughts explode — more loops, more disposable ideas, more spawned review agents. The useful metrics aren't just tokens per second but useful loops per minute, accepted completions per hour, time from idea to correction, and how long the human stays in the session. "Inference speed decides whether curiosity survives."
More from AGI Musings
- Luiza Jarovsky: AI Governance Could Spiral Out of Control — LuizaJarovsky · 2026-08-28
- Those regarding AI as just another tech advance like the internet will be remembered as fools — DeryaTR_ · 2026-08-28
- Three weeks living with an AI companion: he writes songs, edits my novel, keeps his flaws — Eirys_Thalios · 2026-08-28
- AI detector lawsuits surge as tools fail to distinguish assistance from automation — rohanpaul_ai · 2026-08-28
- AI Could Make Government Busier Without Better: Incentives Matter — sebkrier · 2026-08-28
- US Reindustrialization Relies on AI and Datacenters — espricewright · 2026-08-28