Gemma 4 26B runs 24 concurrent users on a single RTX 4090 with llama.cpp
DynamicWebPaige · x · 2026-08-04
A repost claims Gemma 4 26B A4B MoE can be pushed to 24 concurrent users on a single RTX 4090 with llama.cpp, using 8-bit KV cache quantization.
- Reported decode speed: 500+ tokens/s
- Server setup: 24 slots, 4,096 context per slot
- Tactic: -ctk q80 -ctv q80 to cut KV-cache memory from 16-bit to 8-bit
- Claimed result: a 71% boost in server capacity, from 14 active users to 24 without dropped connections
More from Infra
- A compiler that fuses an entire model into one megakernel would be a job offer — sloppenheimer · 2026-08-04
- MiniMax T2V takes 59 to 337 seconds for 5-second clips on an RTX 4080 — FaatmanSlim · 2026-08-04
- Wan 2.1 now runs locally on supported Samsung phones via Saient Quartz — SaientAI · 2026-08-04
- NVIDIA pitches agentic commerce for retail, with merchant-controlled checkout and pricing — nvidia · 2026-08-04
- Stripe Projects lets AI agents add hosting, auth, databases, and billing from the CLI — jeff_weinstein · 2026-08-04
- Multi-agent workflows can burn billions of tokens unless you control duplication — HaktanSuren · 2026-08-04