RTX 5090 local agent setup hits CPU wall: 25 agents stall on tool calls while GPU idles at 700 tok/s
BringTea_666 · reddit · 2026-10-06
The author spent two months squeezing maximum throughput from a single RTX 5090 running a local inference engine (porting Kenshi to Godot), and hit an unexpected bottleneck.
What happened:
- After optimizing cache storage and prefill reuse, more agents were run in parallel (12-slot server, 25 agents at once).
- The GPU averaged only 700 tok/s with free context and idle slots, yet throughput degraded — nearly all agents stalled in tool call states.
- Task Manager revealed the cause: CPU at 100% on every thread. Tool calls are CPU-bound, and a 9800X3D simply can't keep up with heavy agent tool execution.
Lesson: in agentic coding, the practical ceiling is often CPU tool-call execution, not decode speed or prefill. If you want real multi-agent workloads, invest in a better CPU. Project page: kengodot.pages.dev; source code coming.
More from coding & agent
- Dev Builds Community Flood War Room With Claude Code in 24 Hours — DevDminGod · 2026-10-06
- Paris community hackathon ahead of OpenAI DevDay: Oct 24-26, winners get 10 DevDay tickets — paw_lean · 2026-10-06
- Zip rebuilt its agent on LangGraph, cutting weeks of hand-built LLM plumbing per feature — LangChain · 2026-10-06
- GitHub: As AI agents take over implementation, developers must master three new skills — WirelessLife · 2026-10-06
- 1Password Is the Biggest Pain Point for AI Computer Use, Developer Says — aronchick · 2026-10-06
- Dev returns to Codex after a week on Opus 5.5: "basically unusable" — jacob_posel · 2026-10-06