exe.dev: stop stacking agents, use a fast model for human comms
davidcrawshaw · x · 2026-10-01
Josh Bleecher Snyder of exe.dev writes "ETOOMANYTHINGS? Run Fewer Agents," arguing that today's so-called software factories are really just task management (kanban boards, Slack/email interfaces, dependency graphs, ticketing) because agents execute slowly and people hide latency by spawning a new agent whenever they get blocked.
Concurrency is rough on humans — stressful, it trashes flow state and thrashes mental page caches. Task management isn't a solution, it's a band-aid. He calls for making it possible to be equally productive with fewer agents.
His practical directions:
- Use a fast model for talking with humans: start tasks with a powerful model, but hand all comms to a fast, competent model like Luna 6 with no coding tools and an intentionally short, focused context window. You get real-time discussion without attention wandering; when it exhausts its ready information, the beefier model takes back over with rich user feedback.
- Still do old-fashioned engineering, like making tests run fast.
He echoes davidcrawshaw's comment: "As models improve, 'good enough' models will get ever faster."
More from coding & agent
- Firecrawl launches People Enrichment Pack so agents can find buyers, talent and company data — devdigest · 2026-10-02
- Stanford launches CS 224V, an Agentic AI course tackling agent reliability with RAG and formal methods — stanfordnlp · 2026-10-02
- Early Hands-On: OpenAI's Dots Agent Impresses With Speed and First-Try Accuracy — billyjhowell · 2026-10-02
- MINTEval, accepted at NeurIPS, shows LLM agents fail at tracking evolving contexts — EliasEskin · 2026-10-02
- Claude Code lead on the sassy status dot: it might just be busy with something else — cyrus_zei · 2026-10-02
- LlamaIndex Launches Extract v2.5, Beats Claude and GPT at 30%-4x Lower Cost — llama_index · 2026-10-02