Benchmarking MiniMax H3 on a 4090: Sage Attention Slashes Generation Time
thegr8anand · reddit · 2026-08-05
The author extensively tested how different ComfyUI node combinations affect MiniMax H3 video generation times on a local RTX 4090 setup.
Key Findings
- The H3 Mem Eff Sage Attention node provides the most significant speedup, reducing per-step time from 12-13s to 8-8.5s, cutting total generation time from 316s down to 210s.
- Sol Attention offers decent improvements but falls short of Sage Attention. Older nodes like Patch Sage Attn and TorchCompile yielded minimal benefits or instability.
- Quality: The optimized setups maintain visual quality on par with the default workflow.
Deployment Guide
- Includes step-by-step PowerShell commands and links to necessary dependencies for installing Triton and SageAttention on a Windows ComfyUI portable environment.
- Requires the KJ-Nodes custom repository to utilize the optimal acceleration nodes.
Related event: SageAttention Significantly Boosts MiniMax H3 Video Generation Speed(3 posts)→
More from Infra
- Cisco Execs: AI Agents Will Bypass Security Policies, Container Networking is the Foundation — brucemacv · 2026-08-07
- Gemma 4 31B Quantization: Q4_K Draft Model Boosts Decode Speed by 10% — eightone-81 · 2026-08-07
- Musk: Terafab to Produce 1TW Compute Yearly, 75% Allocated for AI Spacecraft — rohanpaul_ai · 2026-08-07
- US Plans 2,441 Data Center Projects with $2.48 Trillion Investment by 2028 — PeterDiamandis · 2026-08-07
- Enforcing Single-Region Data Residency for Claude Code on Amazon Bedrock — AWS ML Blog · 2026-08-07
- Nvidia B300 GPU-hour Index Hits All-Time High as Neoclouds Pivot to Inference — rickasaurus · 2026-08-07