Replacing 32B with 4B/8B encoders in MiniMax H3 optimization tests

Fit_Ad7343 · reddit · 2026-08-15

The author released v3 of MiniMax H3 text encoder optimization, using smaller Qwen3-VL models (4B/8B) with learned projection matrices to replace the original 32B encoder while keeping the DiT untouched. Results show that 4B/8B versions closely track the 32B on explicit instructions (pose, clothing, action), achieving a cosine similarity of 0.9449 (8B) and reducing VRAM usage from 15.7GB to 5GB. The limitation is that details not encoded by the smaller model cannot be recovered, causing variations in unspecified background elements.

Original post →

More from Multimodal

Multimodal channel →