FreeToken vs llama.cpp on RTX 3090: 7x faster TTFT only when the MoE won't fit in VRAM
SignatureMoney6648 · reddit · 2026-10-01
A single-RTX 3090 benchmark compares FreeToken and llama.cpp (1024 in / 256 out, concurrency 1–32). When the model fits in VRAM (Gemma-4-26B-A4B, identical GGUF), llama.cpp wins with 2.2–3.2x throughput and 5–6x faster TTFT; FreeToken 0.1.2 OOM'd at 8 concurrent users. When it doesn't fit (gpt-oss-120b, 63 GB), FreeToken keeps TTFT at 9s up to 32 users vs llama.cpp's 17→139s, with roughly tied throughput. Spilling to system RAM costs 10x generation speed on both engines. The author suspects a PCIe gen3 bottleneck and asks for gen4 reproductions.
More from Infra
- Local models can now power computer use agents, but regulated industries still lack a playbook — Ambitious_Fold_2874 · 2026-10-01
- Yacine's 90-minute deep dive: latent MoE, aggressive GQA inside Nvidia's open model — yacinelearning · 2026-10-01
- Signal65: CoreWeave beats three hyperscalers on infrastructure monetization by up to 195% — ryanshrout · 2026-10-01
- DeepSeek V4.1 Flash spotted running locally on a 192GB Framework Desktop — antirez · 2026-10-01
- China's CXMT to nearly match Micron's DRAM capacity by end of 2026 — Terminator857 · 2026-10-01
- Building a million-page OCR pipeline with a 500GB RAM used server plus LLM extraction — oilmutt · 2026-10-01