2×4090 llama.cpp concurrency: soft cap of 5 agents at 64k context, hard cap 9 — three weeks of data

Iamisseibelial · reddit · 2026-09-09

A nonprofit CTO spent three weeks benchmarking llama.cpp agent concurrency on 2×RTX 4090 (44.6 GiB usable, Threadripper, 128 GB DDR5): soft cap 5 concurrent agents at 64k context, hard cap 9.

Model selection lessons

Quantization: with a self-built long-document probe (facts at 15%/50%/85% depth plus 6 decoy lines), all 24 configs of Q4KM, Q6KXL and Q8KXL scored perfect recall, zero cross-slot leakage, up to 251,557 tokens answered correctly — quant choice was effectively irrelevant for this agent workload.

Original post →

More from Infra

Infra channel →