MiniMax H3Max generates 5s 768p video in under 3s, tops video benchmarks
赛博禅心 · wechat · 2026-09-01
MiniMax H3Max, post-trained by Fal with new data, verifiable RL, and inference acceleration, achieves 35x the throughput of base H3: a 5-second 768p audio-video clip in under 3 seconds and 15-second videos within 15 seconds — faster than playback, crossing the realtime threshold. It ranks #1 on ArtificialAnalysis and DesignArena image-to-video leaderboards. In three weeks H3 passed 24M downloads and 300+ derivative models, making it 2026's most downloaded model.
Notable uses: a Fal engineer wired H3Max into a Twitch stream where chat prompts change the video live; Pieter Levels launched infiniteslop.ai, a 24/7 AI livestream stitching viewer prompts. FastVideo, NuvaLab, and NVIDIA FastGen released FastH3, compressing 49 transformer forwards to 4 with 90% sparsity VSA — 15s 768p video in 13s on 8x B200, up to 14x speedup on a single Blackwell GPU — with weights, LoRA, and inference stack open-sourced.
The author also deployed H3 (and H3Max for a limited time) on AGIBar, free and unlimited, with a working curl command, and shared plans for a dedicated machine room in Beijing with 175kW reserve power.
Related event: H3 Max Generates Video Faster Than Real Time(2 posts)→
More from Infra
- Adani: AI competitive advantage shifts to clean energy and data center integration — Div_pradeep · 2026-09-01
- Naver Proposes Verification-Aware Training to Boost Speculative Decoding Draft Models — naver-ai · 2026-09-01
- ExLlamaV3 update: MoE expert CPU offload, GLM-5.3-Flash, self-calibrated quants — Unstable_Llama · 2026-09-01
- Highlander launches: custom GPU kernels serve realtime video at 50% of competitors' cost — Small-Term672 · 2026-09-01
- Qwen3.8 pMLX Engine: Runs on 12GB RAM at 10 tok/s — EyalToledano · 2026-09-01
- GPU Debt Investors Assume Zero Residual Value and Distrust Spot Pricing — AccBalanced · 2026-09-01