Counterintuitive Test: MiniMax H3 Full BF16 Model is Faster Than INT8 and Better at Physics
Wise_Revolution385 · reddit · 2026-08-08
In a counterintuitive benchmark on an ASUS GX10 with 121GB of unified memory, a developer found that the 66.3GB full BF16 MiniMax H3 video generation model was not only 12-23% faster during DiT inference than the 20.9GB pruned INT8 model, but it also produced significantly better video quality.
The author speculates that the INT8 model's casting and dequantization overhead negates its size advantage, while the GX10's unified memory bandwidth handles the larger BF16 model efficiently.
Regarding video quality, the full model demonstrated clear advantages in:
- Natural high-frequency motion: Dragonfly wing movements looked aerodynamic rather than rigid.
- Better background semantic preservation: Distant figures on advertising screens remained recognizable without distortion.
- Stronger physical intuition: The flying motorcycle exhibited realistic acceleration, inertia, and deceleration.
- Higher object structure consistency: The full model correctly preserved minor details like a right-side exhaust pipe throughout the sequence, which the INT8 model dropped entirely.
More from Infra
- llama.cpp RPC PR: Cuts 300GB Model Loading Time to 1.5 Minutes — Chuyito · 2026-08-08
- Nvidia to Invest Up to $3 Billion in Blackstone-Backed Power Firm — pstAsiatech · 2026-08-08
- AURORA-LM: A 1B Continuous Diffusion Language Model Trained on Ascend NPU — 机器之心 · 2026-08-08
- Local Video Generation: MiniMax Lags Far Behind LTX in Inference Speed — PhilosopherSweaty826 · 2026-08-08
- Deep Dive: How Weak is the Evidence for China's Role in US Data Center Backlash? — AndyMasley · 2026-08-08
- $50k Bet Challenges SemiAnalysis on SpaceX AI Compute and ARR Forecasts — generativist · 2026-08-08