UC Berkeley Releases FreeToken for Efficient Edge-Native MoE Serving
UCBerkeley · hf · 2026-08-19
FreeToken is an edge-native Mixture-of-Experts serving system. It dynamically maps computation and model state to heterogeneous local hardware via bandwidth-adaptive execution, enabling large open-weight models to run on personal machines.
More from Infra
- LLM Inference Engineering: From KV Cache to vLLM and SGLang — techNmak · 2026-08-19
- DFlash 2: Qwen3.8-27B hits 70 tok/s on MacBook with 4.6x speedup — songhan_mit · 2026-08-19
- NVIDIA H100 Concurrency Response of Plain Global Loads Analyzed — ssh4net · 2026-08-19
- Using HBF for KV Cache Offload Risks Endurance Burnout — zephyr_z9 · 2026-08-19
- Considered nuclear startup funded by hyperscalers, impressed by serious energy buildout — JacquesThibs · 2026-08-19
- Qwen3.8-27B on 2x 3090 hits 218 tok/s decode with vLLM + DFlash2 spec-decode — xjx546 · 2026-08-19