WinterMix: New 3-bit MLX Quantization Beats GGUF in Long Context
WinterCharm · reddit · 2026-08-10
Developer WinterCharm released WinterMix38 (59 GiB, 3-bit), a new quantized version of the Qwen3.5-122B-A10B model tailored for Apple Silicon. It uses a native MLX format that drops directly into environments like LM Studio without requiring custom kernels.
The author notes that while MLX is inherently faster on Apple Silicon (roughly 9x faster prefill than llama.cpp), native MLX quants below 6-bit have traditionally suffered in quality. WinterMix tackles this by applying a new annealing process to the model's reasoning traces.
Benchmarks show that at 16K context, WinterMix38 achieves better perplexity than Unsloth's best 3-bit GGUF model. The advantage compounds at longer contexts—at 96K tokens, it outperforms the imatrix format by 5.7% in perplexity. The model weights are open-source.
More from Infra
- Stop Running Blind: Open-Sourcing specspecs for Speculative Decoding Observability — HamelHusain · 2026-08-10
- Sony and TSMC to Invest $6.3B in Advanced Image Sensor Plant in Kumamoto — pstAsiatech · 2026-08-10
- Running Video Generation on RTX 4060 Ti: Qwen3-vl + MiniMax-H3 Takes 26 Minutes — LuisaPinguinnn · 2026-08-10
- DeepSeek V4 Flash Clears All 22 Coding Tasks on Dual DGX Spark Cluster — AccBalanced · 2026-08-10
- Local LLM Deployment Costs $10K, Taking 24 Years to Break Even vs API — TheZachMueller · 2026-08-10
- Marvell Pushes Data Centers to Buy AI Memory and Compute Separately — shashib · 2026-08-10