2×4090 llama.cpp concurrency: soft cap of 5 agents at 64k context, hard cap 9 — three weeks of data
Iamisseibelial · reddit · 2026-09-09
A nonprofit CTO spent three weeks benchmarking llama.cpp agent concurrency on 2×RTX 4090 (44.6 GiB usable, Threadripper, 128 GB DDR5): soft cap 5 concurrent agents at 64k context, hard cap 9.
Model selection lessons
- Qwen3.5-122B-A10B decoded at 18.75 tok/s but took 11.91s per tool-call vs 3.40s for a 27B — per-call latency matters far more for agents making dozens of calls.
- Qwen3.6-35B-A3B was fastest on paper (78 tok/s) but failed 1 of 5 executed code tasks; the 27B went 5/5 — decode tok/s is the wrong headline metric.
- Qwen3.8-Flash-Next (125B MoE) topped out at 23.7 tok/s: 71.7 GiB of its 103.7 GiB weights are routed experts living in system RAM, and with only 2 of 4 memory channels populated, effective bandwidth was 14.1 GB/s against an 83.2 GB/s peak — filling the channels would beat any tuning.
Quantization: with a self-built long-document probe (facts at 15%/50%/85% depth plus 6 decoy lines), all 24 configs of Q4KM, Q6KXL and Q8KXL scored perfect recall, zero cross-slot leakage, up to 251,557 tokens answered correctly — quant choice was effectively irrelevant for this agent workload.
More from Infra
- Rumor resurfaces: Google may replace Nvidia as TSMC's biggest customer — zephyr_z9 · 2026-09-09
- Tahuna open-sources ephemeral GPU orchestration for ML workloads — Monaim101 · 2026-09-09
- Zuckerberg: Meta already training post-Watermelon models on its 1GW Prometheus cluster — rohanpaul_ai · 2026-09-09
- Developer maps the entire inference + fine-tuning provider landscape with tradeoffs — Present-Jelly9941 · 2026-09-09
- Smart offloading of active requests erases linear attention's memory advantage — samsja19 · 2026-09-09
- Local video generation on Android: ~550-600s per clip on Snapdragon 8 Gen 3 — sgcego · 2026-09-09