FreeToken Claims Faster MoE Inference vs. llama.cpp and Ollama
wavefnx · x · 2026-08-18
FreeToken claims to serve Mixture-of-Experts (MoE) models significantly faster than popular alternatives like llama.cpp, ktransformers, ollama, or moe-infinity. The associated paper was released yesterday, though independent benchmarks are not yet available. While a binary is available, the source code has been taken down, and the author advises against running the binary without reverse engineering it first.
Related event: FreeToken Claims Major MoE Inference Speedup(2 posts)→
More from Infra
- SPU concept proposed: Search may need specialized processing units — CShorten30 · 2026-08-18
- Linux 7.3 kernel improves VRAM management — johnnyApplePRNG · 2026-08-18
- AntSDR T510 Launch: 6GHz SDR with Built-in Jetson AI Processing — Dave_Maynor · 2026-08-18
- Benchmark: Qwen 3.8 hits 45 tps with 1M context on 3080Ti + Strix Halo — TrifleHopeful5418 · 2026-08-18
- Cut LLM hardware costs with ROM-based weight storage — sirzerp · 2026-08-18
- US largest grid to cut power to new data centers first during shortages — KeanuRave100 · 2026-08-18