Running MiniMax H3 on RTX 4090: Benchmarks Show ~16 s/it at 2MP
jugernaut126 · reddit · 2026-08-09
A developer shared performance metrics of running MiniMax H3 on a single RTX 4090. At 2MP resolution, optimized inference hits 16 s/it, compared to 72-83 s/it without optimization. The thread discusses pushing sequence lengths within VRAM limits and further speed optimizations.
Related event: Developers Test MiniMax H3 Video Model on Consumer GPUs(4 posts)→
More from Infra
- A Comprehensive Guide to LLM Inference Optimization and Deployment — abhijithneil · 2026-08-09
- Bizarre Coincidence: Elon's Terafab Site Matches Grimes' 2020 AI Prophecy — OwariDa · 2026-08-09
- Monthly AI Tokens Hit 11 Quadrillion, Projected 70x Growth in 5 Years — AccBalanced · 2026-08-09
- Baseten Breaks Down Serving 2.8T-param Kimi K3 at Scale on Blackwell — thursdai_pod · 2026-08-09
- Polymarket: 73% Chance a US State Enacts a Data Center Moratorium by 2026 — Polymarket · 2026-08-09
- Amazon's Planned Texas Data Center Permitted to Emit More CO₂ Than Any US Power Plant — Polymarket · 2026-08-09