MiniMax-H3 Variant with 2-bit Text Encoder Runs on M1 Max 32GB
antocorr · reddit · 2026-08-07
Reddit user antocorr releases a MiniMax-H3 FL2VA variant for mlx-serve, quantizing only the Qwen3-VL text encoder to 2-bit (affine, g64) while keeping DiT at 4-bit and VAEs/tokenizer untouched. Text encoder disk size drops from 15.8GB to 9.6GB, enabling the full text-to-audio-video pipeline natively on Apple Silicon (verified on M1 Max 32GB with --skip-mem-preflight). mlx-serve automatically handles mixed quantization. Caveat: 2-bit conditioning is lossier, slightly reducing prompt adherence, but video/audio quality remains same.
More from Infra
- Turso Rewrites Postgres in Rust to Build the LLVM of Databases — JeremyCMorgan · 2026-08-07
- vLLM: The Serving Engine Making LLM Deployment Affordable — eyishazyer · 2026-08-07
- Rumor: OpenAI to Launch GPT-6 Next Week; Tesla Invests Billions in Terafab — Not Boring (Packy McCormick) · 2026-08-07
- Running MiniMax H3 Video Generation Locally on RTX 3060 12GB — jefharris · 2026-08-07
- Optimizing MiniMax H3 on AMD GPUs: Benchmarks and Pitfalls — frq2000 · 2026-08-07
- llama.cpp PR Boosts Q2_0 CPU Decoding by 3x, 8B Hits 8.20 tok/s — BTA_Labs · 2026-08-07