FreeToken Tested: 100 tok/s Running a 35B Model on Consumer Hardware
ViRROOO · reddit · 2026-08-22
A new project FreeToken (paper: arXiv:2608.16157, GitHub: FlashML-org/FreeToken) was released and the author ran initial local tests.
Setup: RTX 5080 (16GB VRAM), 64GB DDR6 RAM, AMD Ryzen 9 9950X3D—achieving 100 tokens/s on QWEN3.6-35B-A3B in NVFP4 quantization.
The author finds it impressive and invites the community to share their own benchmark results.
Related event: Berkeley Open-Sources FreeToken: Running Giant MoE Models on Consumer PCs(6 posts)→
More from Infra
- FreeToken: Run 284B frontier models on consumer GPUs at interactive speeds — solyarisoftware · 2026-08-22
- US AI industry's economic model is cracking: subscriptions below compute cost, data centers stalled — Minimum_Name9115 · 2026-08-22
- Starcloud runs a language model on its satellite after training on orbiting Nvidia GPUs — emmanuelvivier · 2026-08-22
- Troubleshooting Multi-GPU Support for Deepseek V3 in llama.cpp — erazortt · 2026-08-22
- Ox Alpha's 100T tokens/day giveaway costs $1-2M in electricity at full capacity — teortaxesTex · 2026-08-22
- Seeking unified AI gateway for OpenAI cost visibility — HurryOrganic · 2026-08-22