Extreme Optimization: Running 33B Video Generation Model on M4 ANE

antirez · x · 2026-08-13

Developer @originalmaderix forked @antirez's MiniMax H3 video generation implementation to run on the Apple Neural Engine (ANE), challenging a 33B parameter model on a base M4 chip.

This hardcore optimization project utilizes SSD streaming and overlaps int8 compute on the ANE. Impressively, running this massive Diffusion Transformer model requires only about 2GB of peak RAM and consumes under 5W of power.

Original post →

More from Infra

Infra channel →