H3 T2VA VRAM squeeze sparks proposal for a remote text-encoder API to keep the 32B off your GPU
frankmanbb · reddit · 2026-09-29
H3 local users face a VRAM dilemma: the DiT wants 20–25GB while the text encoder is a truncated Qwen3-VL-32B, and the two don't co-fit on 24–32GB cards — leaving the encoder loaded stalls sampling, unloading it means waiting on every prompt. The author proposes a text-only encode API for T2VA:
- Send a prompt, receive H3-compatible text conditioning tensors
- DiT and VAEs stay local; a Comfy node replaces only the text-encode step
- Not a hosted video API; first/last frames and refs still go through the local video VAE
- Core idea: no need to load the 32B just to read a prompt
The author is polling the community on whether they'd use it, what GPU they run, whether encoder or DiT residency hurts more, and whether tensors must match the official H3 TE pack exactly.
More from Infra
- Dagger founder: the Great CI Bottleneck of 2026 is a software problem, not hardware — msharmas · 2026-09-29
- LayerSkip: Self-Speculative Decoding Speeds Up LLMs Without a Draft Model — burkov · 2026-09-29
- SpaceX outlines supercomputer, Terafab, Gigasat and new Louisiana Starbase plans — elonmusk · 2026-09-29
- Musk says space will hold nearly all compute; Google tests if TPUs work there — CackleRooster · 2026-09-29
- Gimlet and Cerebras plan 100 MW capacity, targeting 3,000 tokens/sec inference — Sethwinterroth · 2026-09-29
- 95+ TPS and 262K context for Qwen 27B on a single RTX 3090 with LlamAmpere v0.4 — Brief-Tap-6616 · 2026-09-29