WinterMix: New 3-bit MLX Quantization Beats GGUF in Long Context

WinterCharm · reddit · 2026-08-10

Developer WinterCharm released WinterMix38 (59 GiB, 3-bit), a new quantized version of the Qwen3.5-122B-A10B model tailored for Apple Silicon. It uses a native MLX format that drops directly into environments like LM Studio without requiring custom kernels.

The author notes that while MLX is inherently faster on Apple Silicon (roughly 9x faster prefill than llama.cpp), native MLX quants below 6-bit have traditionally suffered in quality. WinterMix tackles this by applying a new annealing process to the model's reasoning traces.

Benchmarks show that at 16K context, WinterMix38 achieves better perplexity than Unsloth's best 3-bit GGUF model. The advantage compounds at longer contexts—at 96K tokens, it outperforms the imatrix format by 5.7% in perplexity. The model weights are open-source.

Original post →

More from Infra

Infra channel →