FreeToken doubles local MoE speed: Qwen on 16GB RTX 5070 Ti hits 90+ tps
solyarisoftware · x · 2026-09-06
- A llama.cpp GitHub discussion calls for adopting FreeToken after wild local MoE benchmarks: Qwen3.6-35B-A3B-NVFP4 (23GB) runs on a 16GB RTX 5070 Ti at 90+ tps (100+ on subsequent calls); on an RTX 5090 with Qwen3.8-Flash-Next Q4, FreeToken hits 53 tps vs llama.cpp's 22 — more than 2x.
- FreeToken treats the whole PC as an inference system: experts live in system RAM, hot experts are cached in VRAM, cache misses stream over PCIe, the CPU can execute some misses, and CPU work overlaps with PCIe transfers — the key is deciding which experts must be in VRAM right now.
More from Infra
- PyPI's Recent Download Corrections Sharply Cut Some Packages' Stats — dbreunig · 2026-09-06
- Running gemma4:31b-mlx locally with Ollama feels indistinguishable from paid tiers — walkingriver · 2026-09-06
- Redditor seeks a lightweight OpenAI-compatible API client, complains OpenWebUI is 30GB — MelodicRecognition7 · 2026-09-06
- SK Hynix weighs Intel Foundry for part of HBM4e base die output as TSMC costs run 3-4x higher — Beth_Kindig · 2026-09-06
- B200 recoups full hardware cost in ~10 months at 94% rental utilization, per forward-curve pricing — AccBalanced · 2026-09-06
- SlimServe runs Qwen 3.8 Flash Next on 8x RTX 3090s at up to 1200 tok/sec decode — QuixiAI · 2026-09-06