Raylight Adds Sequence Parallel for MiniMax H3, Halving Generation Time on RTX 2000 ADA

Altruistic_Heat_9531 · reddit · 2026-08-03

The open-source tool Raylight announced support for sequence parallel inference for the MiniMax H3 model. Thanks to H3's single-stream Transformer architecture, porting the model has become significantly easier.

On a single RTX 2000 Ada graphics card (comparable to an RTX 4060), the time to generate a 5-second 864x480 video with audio was reduced from 65 seconds to 35 seconds. The update also introduces distributed VAE, INT8 ConvRot FSDP, and other critical memory and inference optimizations.

Original post →

More from Infra

Infra channel →