Sand.ai Open-Sources 114B MoE Model MAGI-2, Slashing Audio-Video Generation Costs by 90%
新智元 · wechat · 2026-08-06
Sand.ai has officially released and open-sourced MAGI-2 Preview, the world's first 100-billion-parameter MoE unified audio-video generation model. It has a total of 114B parameters but only activates about 6B per forward pass, reducing inference costs to 1/10th of mainstream models.
Core Architecture & Technical Breakthroughs:
- Ultra-fine-grained MoE: Splits the 3072-dim representation into 12 256-dim subspaces. Each MoE layer has 3072 head-local expert units, but a token only activates 72, enabling richer capability combinations.
- HeadParallel Communication: Flips the traditional "route-then-communicate" to "communicate-then-route". Cross-device transmission becomes fixed, eliminating the linear scaling of communication overhead with activated experts and solving the routing bottleneck.
- Unified Audio-Video Generation: Drops the post-production stitching of video and audio. Text, video, and audio are modeled in the same context from step one, achieving high-fidelity audio-visual sync and lip alignment natively.
Performance & Cost:
Ranked #6 on the ArtificialAnalysis image-to-video leaderboard. Generating a 10-second video costs only $0.07 (0.5 RMB) based on 8x H100 rental rates, drastically lowering the commercial barrier for video products.
Related event: Sand.ai Open-Sources 114B MAGI-2, Slashing Audio-Video Generation Costs(3 posts)→
More from Multimodal
- OnSolo launches $200k AI short drama creator awards — Div_pradeep · 2026-08-26
- Seedance 2.5 adds consistent multi-character video generation — HeyNayeem · 2026-08-26
- Recraft Studio Unveils New Look and Design Agent — aziz4ai · 2026-08-26
- A starter list of X accounts to follow for AI music creation — TheChuckTone · 2026-08-26
- User Review: Google Lyria 3.5 is the Best AI Music Model — oyacaro · 2026-08-26
- LTX-2.5 is open weights, with a smaller distilled model that runs locally — egeberkina · 2026-08-26