vLLM and MiniMax Release H3: Open-Weight Native Audio-Video Generation
vllm_project · x · 2026-08-03
The vLLM project collaborated with the MiniMax team to release the deployment recipe for MiniMax-H3, an open-weight multimodal generation model.
- Core Capabilities: Reads a mixed multimodal context (text, images, video, audio) to generate coherent audio-visual output end-to-end, rather than dubbing audio afterwards.
- Architecture: A CFG-distilled joint video/audio diffusion transformer (DiT) served via vLLM-Omni's OpenAI-compatible API. Returns a single MP4 containing H.264 video and native stereo audio.
- Performance: Generates 8.7s of 1248×768 video with synchronized stereo audio in 87 seconds on 4×B300 GPUs.
- Deployment: Ships as two independently served partitions (FL2VA for text/first-last frame, Ref2VA for omni references), requiring a server restart to switch tasks.
More from Infra
- Google TPU Demand Surge: Projected to Reach 15M Units by 2028 — SumitGup · 2026-08-03
- Run 2.78T Parameter Kimi K3 on a Single CPU in 8.24GB RAM — Saboo_Shubham_ · 2026-08-03
- tinygrad Teases Local Deployment Product, Hints at Upcoming Qwen3.6-27B — max_paperclips · 2026-08-03
- AMD MI355X Beats NVIDIA B200 in Kimi K3 Deployment with 952 tok/s — adrianscottcom · 2026-08-03
- Qwen3.8-27B Open Weights Coming, Runs Locally on 17GB RAM — danielhanchen · 2026-08-03
- AirLLM Breaks VRAM Barrier: Runs 70B LLMs on a Single 4GB GPU — techNmak · 2026-08-03