Sushi Quant Squeezes Qwen3.8 Into 43.8GB, Runs Locally on 64GB Apple Silicon

alexcovo_eth · x · 2026-09-28

Sushi, a custom inference engine for Apple Silicon (a fork of mlx-serve), released its smallest quant yet: Qwen3.8-Flash-Next-Sushi-2.6bpw at just 43.8GB of weights, runnable on 64GB-class Macs.

Original post →

More from Infra

Infra channel →