MiniMax-H3 RefMod Upgrade: Packing JPEGs Into Safetensors Cuts Encoding Time and Kills Identity Bleed
acedelgado · reddit · 2026-09-30
A developer updated the open-source Fantastic MiniMax-H3 Prompt Builder (ComfyUI nodes), with improvements centered on RefMod conditioning and caching:
- Background: vanilla refmods only feed latents to the DiT; the text encoder never sees references. The author's earlier version fed sampled frames to the TE as anchors, but re-encoded them every generation.
- Speedup trick: since .safetensors is just a container, actual JPEG data is packed alongside the latents in the same file (like subtitles in a video). JPEGs go straight to the text encoder without the VAE, and TE conditioning is cached between runs as long as refmod order/strength is unchanged.
- New stackpictures modes: every 4th (default, lightest), up to 8 (samples 8 frames individually, more detail), all (all frames to the TE — slowest but max detail with little to no character bleed).
- Benchmarks: 6 images per character plus seconds of audio; 3 refmods (14,900 tokens) ran at 32s/it at 0.98mp on a 5090 with 8-step HyperFlow and an audio refiner suite. Also added edit masking for video edits. Dataset and companion repos are open source.
More from Multimodal
- Filmmakers jam with AI video generation wait times to shoot a music duet with Luma — mrjonfinger · 2026-09-30
- Meta's LSRM wins ECCV 2026 Best Paper Honorable Mention, scales 3D reconstruction with sparse attention — rsasaki0109 · 2026-09-30
- Meta's LSRM Wins ECCV 2026 Honorable Mention, Beats 3D Reconstruction SOTA by 2.4 dB — rsasaki0109 · 2026-09-30
- NVIDIA's LongLive-Plug: Distill Once, Deploy Training-Free Across 54 Downstream Video Models — nvidia · 2026-09-30
- Adobe Research Shows Adversarial Post-Training Restores Missing High-Frequency Detail in Pixel Diffusion — adobe-research · 2026-09-30
- UCSD's LIFT Lets You Control Future Video Layouts via On-Policy Self-Distillation — UCSanDiego · 2026-09-30