Extreme Optimization: Running 33B Video Generation Model on M4 ANE
antirez · x · 2026-08-13
Developer @originalmaderix forked @antirez's MiniMax H3 video generation implementation to run on the Apple Neural Engine (ANE), challenging a 33B parameter model on a base M4 chip.
This hardcore optimization project utilizes SSD streaming and overlaps int8 compute on the ANE. Impressively, running this massive Diffusion Transformer model requires only about 2GB of peak RAM and consumes under 5W of power.
More from Infra
- YC-Backed Pacific Deploys Micro Data Centers to Monetize Unused Power — ycombinator · 2026-08-13
- Lumentum CEO: AI Data Center Network Capacity Demands Double a Decade of Global Backbone — Beth_Kindig · 2026-08-13
- Why Local Models Won't Win: The Inevitable Dominance of Datacenter Inference — rseroter · 2026-08-13
- Three Levels of Agentic Engineering: The Real Bottleneck is Compute, Not LLMs — peterjliu · 2026-08-13
- Local Image Gen on RTX 5060 Ti 16G: 8-Step Turbo Hits 2304x1280 — Sad_Coach_1433 · 2026-08-13
- AI Data Center Capacity Crisis Looms as Power Grids Fail to Keep Up — DavidLinthicum · 2026-08-13