FreeToken vs llama.cpp on RTX 3090: 7x faster TTFT only when the MoE won't fit in VRAM

SignatureMoney6648 · reddit · 2026-10-01

A single-RTX 3090 benchmark compares FreeToken and llama.cpp (1024 in / 256 out, concurrency 1–32). When the model fits in VRAM (Gemma-4-26B-A4B, identical GGUF), llama.cpp wins with 2.2–3.2x throughput and 5–6x faster TTFT; FreeToken 0.1.2 OOM'd at 8 concurrent users. When it doesn't fit (gpt-oss-120b, 63 GB), FreeToken keeps TTFT at 9s up to 32 users vs llama.cpp's 17→139s, with roughly tied throughput. Spilling to system RAM costs 10x generation speed on both engines. The author suspects a PCIe gen3 bottleneck and asks for gen4 reproductions.

Original post →

More from Infra

Infra channel →