Gemma 4 26B runs 24 concurrent users on a single RTX 4090 with llama.cpp

DynamicWebPaige · x · 2026-08-04

A repost claims Gemma 4 26B A4B MoE can be pushed to 24 concurrent users on a single RTX 4090 with llama.cpp, using 8-bit KV cache quantization.

Original post →

More from Infra

Infra channel →