MiniMax H3 Acceleration Benchmark: TE-Speed Delivers up to 1.785x Speedup
Commercial_Board9219 · reddit · 2026-08-04
A developer open-sourced a deployment and benchmark harness for testing the MiniMax H3 video model on Google Colab G4, comparing the acceleration performance of Native, Spectrum, and TE-Speed modes.
Key Benchmark Results:
- TE-Speed Acceleration: With SageAttention enabled, TE-Speed achieved up to 1.785x speedup (44% time reduction) compared to Native mode. It is pure Python and requires no compiled CUDA extensions.
- Spectrum Performance: Significantly reduced end-to-end generation time, though the benchmark against Native mode was confounded by the simultaneous enablement of SageAttention.
- Compatibility Issue: TE-Speed and Spectrum cannot be stacked currently. TE-Speed skips trailing transformer blocks, while Spectrum relies on their execution, leading to conflicting assumptions and runtime errors.
Test Environment & Open-source Assets:
- Hardware included NVIDIA RTX PRO 6000 Blackwell (97GB VRAM), PyTorch 2.11 + CUDA 12.8.
- The repository contains reproducible runners, ComfyUI workflows, paired Native/accelerated MP4 outputs, contact sheets for visual inspection, and machine-readable benchmark JSON logs.
More from Infra
- 7-Month-Old Volta Raises $3B at $2.4B, Lands $10B Anthropic Deal — matt_slotnick · 2026-08-04
- Running Codex on 128-Core CPU Clusters: A New Compute Approach — BenBajarin · 2026-08-04
- Nscale Pledges $1M Compute to Oxford, UCL, and Imperial College — _rockt · 2026-08-04
- Benchmarking Base64 vs Presigned URLs for Multimodal API Image Delivery — Nigiva · 2026-08-04
- EON Raises $10.75M Seed to Replace Subsea Cables with Space Lasers — oyhsu · 2026-08-04
- Running Minimax H3 on RTX 3090: 15-Second Video in 22 Minutes — pfeifits · 2026-08-04