MiniMax-H3 Video+Audio Model Gets MLX Port for Apple Silicon
JeremyCMorgan · x · 2026-08-07
MiniMax-H3, an omni-modal model capable of generating 15-second videos with synchronized audio, has been ported to Apple Silicon via MLX.
- Architecture: Rather than a language model, it is a 33B diffusion transformer that uses a frozen Qwen3-VL-32B encoder with separate video and audio VAEs.
- Local Execution: Weighing in at 115GB, it can be run with a single uv command. Generating a clip takes roughly 45 minutes on an M5 Max chip, proving that local video generation on laptops is quietly becoming a reality.
Related event: MiniMax H3 Local Inference Benchmarks: Runs on Consumer GPUs and Mac(16 posts)→
More from Infra
- AMD Acquires AI Chip Startup Taalas to Boost Decode Acceleration — BenBajarin · 2026-08-07
- Breaking the AI Memory Wall: CXL Moves to Deployment, Marvell Well Positioned — BenBajarin · 2026-08-07
- Samsung, SK Hynix, and Micron Sell Out 2027 Memory Capacity to AI Firms — ramos_casals · 2026-08-07
- DSpark Boosts Local DeepSeek-V4-Flash Inference Speed by 2x — petrusenko_max · 2026-08-07
- Nativ Launches Open-Source Local AI App for Mac, Prioritizing Privacy and Native Frontier Models — Scobleizer · 2026-08-07
- NVIDIA considers reducing high-bandwidth memory for Rubin Ultra GPU amid HBM shortage — BenBajarin · 2026-08-07