FreeToken Tested: 100 tok/s Running a 35B Model on Consumer Hardware

ViRROOO · reddit · 2026-08-22

A new project FreeToken (paper: arXiv:2608.16157, GitHub: FlashML-org/FreeToken) was released and the author ran initial local tests.

Setup: RTX 5080 (16GB VRAM), 64GB DDR6 RAM, AMD Ryzen 9 9950X3D—achieving 100 tokens/s on QWEN3.6-35B-A3B in NVFP4 quantization.

The author finds it impressive and invites the community to share their own benchmark results.

Related event: Berkeley Open-Sources FreeToken: Running Giant MoE Models on Consumer PCs(6 posts)→

Original post →

More from Infra

Infra channel →