Replacing 32B with 4B: Extreme Compression for MiniMax H3 Video Generation

Fit_Ad7343 · reddit · 2026-08-09

The author details a method to replace the native Qwen3-VL-32B text encoder in MiniMax H3 with a Qwen3-VL-4B model. By learning a linear projection matrix using ridge regression, the hidden states of the 4B model are mapped into the 32B's conditioning space, slashing VRAM usage from 15.7GB to 4.5GB.

Technical Details & Results:

Limitations:

Original post →

More from Multimodal

Multimodal channel →