Sushi Quant Squeezes Qwen3.8 Into 43.8GB, Runs Locally on 64GB Apple Silicon
alexcovo_eth · x · 2026-09-28
Sushi, a custom inference engine for Apple Silicon (a fork of mlx-serve), released its smallest quant yet: Qwen3.8-Flash-Next-Sushi-2.6bpw at just 43.8GB of weights, runnable on 64GB-class Macs.
- Mixes EXL3 and affine quant formats tailored for M5 Max-class chips
- Also offers 3bpw (64GB+) and 4bpw (96GB+) variants
- Open source on GitHub; works standalone or as a guest engine within mlx-serve
More from Infra
- Fireworks' Ember-1 post-trains Kimi K3 to reason 40% more concisely at same quality — isidentical · 2026-09-28
- One AMD driver flag boosts dual-GPU Vulkan LLM inference up to 4x — tabletuser_blogspot · 2026-09-28
- PrismML's Bonsai 2 shrinks Qwen3.8 27B to 5.9GB, keeps 98.2% capability, runs on a 5090 — dl_weekly · 2026-09-28
- Developer Plans to Let Codex Pick Which Tests Run, Slashing CI Costs — sull · 2026-09-28
- Free online guide covers LLMs from first principles to local deployment — JFPuget · 2026-09-28
- Is a vector database enough for production AI agents? Reddit debates storage design — OkShirt9372 · 2026-09-28