H3 Inference Benchmarks: 8x 3090 Beats M3 Ultra
QuixiAI · x · 2026-08-25
Performance benchmarks for the MiniMax H3 inference engine on Mac reveal significant speed differences. A Mac Studio M3 Ultra took 2hr 42min, a single RTX 3090 took 45min, while an 8x RTX 3090 setup completed the task in just 9min 43sec.
More from Infra
- 4070 Ti Benchmarks: Running Kimi K3, DeepSeek V4, and Qwen 122B Locally — JayB_Official · 2026-08-25
- CXMT LPDDR6 revealed, Xiaomi XRING-03 as first adopter — teortaxesTex · 2026-08-25
- Fork of h3.c adds 3090, multi-GPU, GGUF support, and a UI — QuixiAI · 2026-08-25
- 1 hour of AI work now costs less than a minute of SF parking — stuffyokodraws · 2026-08-25
- Qwen 3.8 27B on Mac hits 113 TPS via Dflash2 optimization — TheMoonMidas · 2026-08-25
- TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference — Hanzhi Zhang · 2026-08-25