Running local agents on vLLM: only ~8 agents with 32k context each — what use cases actually help?
Ambitious_Fold_2874 · reddit · 2026-09-25
A Reddit user shares the compute math for local multi-agent setups: running Qwen3 27B (nvfp4) on vLLM at max context yields only about 8 concurrent agents with 32k context each. They ask what multi-agent workflows people actually find useful at this scale.
More from coding & agent
- Antithesis teaches agents mutation testing to break distributed safety properties — moarbugs · 2026-09-25
- LlamaParse nails the Fed's dot plot: 9 for 9 on GDP medians — llama_index · 2026-09-25
- Microsoft rebuilds Copilot into a super app with persistent Autopilot agents — rohanpaul_ai · 2026-09-25
- Embedding test: RAG ranks the no-refund policy first, showing models still matter — galratner · 2026-09-25
- Jev for PowerShell 0.2.0 turns messy incident text into page/do-not-page decisions — dfinke · 2026-09-25
- Matt Pocock unveils /pr and /retro at AI Engineer Paris to speed up code reviews — mattpocockuk · 2026-09-25