SandAI Releases MAGI-2 Preview: 114B Audio-Video Generation Model
Nice_Amphibian_8367 · reddit · 2026-08-05
SandAI has released MAGI-2 Preview, a unified audio-video generation model.
- Architecture: 114B total parameters with 6B active per token. Uses a MagiMoE / multi-head latent MoE design, which is more video-oriented than standard LLM-style MoE.
- Generation: Employs FlowUniPCMultistepScheduler for flow-matching sampling instead of the autoregressive chunking used in MAGI-1. Generation is two-stage: preview denoising followed by a 1080p refiner.
- Features: Supports T2V and I2V, with audio generated alongside the video.
- Code: The released repository is inference-only.
Related event: Sand.ai Open-Sources 114B Audio-Video Model MAGI-2 Preview(2 posts)→
More from Multimodal
- MiniMax Video Model Demo: Generating Tom Holland Eating Street Food from 4 Images — CurieuxExplorer · 2026-08-05
- MiniMax H3 Early Access Live: Native Multimodal Generation & Precise Editing — HeyAmit_ · 2026-08-05
- Peking University Open-Sources MiniWorld for Training Video World Models on a Single 8-GPU Server — PekingUniversity · 2026-08-05
- Seedance 2.5 Launches with 30s Native Video Generation and Multimodal Inputs — HeyNayeem · 2026-08-05
- Seedance 2.5 Hits Lumina AI: Generates 30-Second Native Videos — HeyNayeem · 2026-08-05
- MiniMax H3 video generation benchmark: longer videos have higher per-second cost — madcaddie15 · 2026-08-05