H3 T2VA VRAM squeeze sparks proposal for a remote text-encoder API to keep the 32B off your GPU

frankmanbb · reddit · 2026-09-29

H3 local users face a VRAM dilemma: the DiT wants 20–25GB while the text encoder is a truncated Qwen3-VL-32B, and the two don't co-fit on 24–32GB cards — leaving the encoder loaded stalls sampling, unloading it means waiting on every prompt. The author proposes a text-only encode API for T2VA:

The author is polling the community on whether they'd use it, what GPU they run, whether encoder or DiT residency hurts more, and whether tensors must match the official H3 TE pack exactly.

Original post →

More from Infra

Infra channel →