Inference speed reshapes human-agent collaboration: 20s keeps you in flow, 10min is just delegation

demian_ai · x · 2026-08-28

The author argues that LLM/agent response speed changes the shape of human collaboration. If an agent returns an interface iteration in 20 seconds, you stay, look, change your mind, try something stranger — the LLM becomes part of your thought process. If it takes 10 minutes, it degrades into delegation: review later, issue another task.

This creates a demand paradox: faster inference doesn't necessarily reduce compute demand. The cost per answer falls while attempted thoughts explode — more loops, more disposable ideas, more spawned review agents. The useful metrics aren't just tokens per second but useful loops per minute, accepted completions per hour, time from idea to correction, and how long the human stays in the session. "Inference speed decides whether curiosity survives."

Original post →

More from AGI Musings

AGI Musings channel →