MiniMax H3 Inference Benchmarks Leak as Community Fork Adds GGUF and Multi-GPU Support

Benchmark data for MiniMax's H3 inference engine shows an 8x RTX 3090 setup finishing a task in 9 minutes, versus 2h42m on a Mac Studio M3 Ultra. A community fork of h3.c now adds GGUF support, multi-GPU compatibility, and a UI.

2026-08-25 ~ 2026-08-25 · 2 related posts