Veda's sparse-attention ComfyUI node renders 5s video 2.9x faster on a 12GB RTX 5070
lmoroney · x · 2026-10-06
The Veda team (ByteDance, HKU, USTC; ICML 2026) shipped an official ComfyUI node for the MiniMax H3 video model that speeds up rendering via sparse attention:
- How it works: a small learned predictor scores the attention map in tiles and keeps roughly the top 10%; a custom kernel fetches only those tiles. Weights are untouched, so LoRAs and fine-tuned or quantized H3 checkpoints still work.
- Benchmark: on a 12GB RTX 5070 with the 8-step Turbo LoRA, a 5.2-second 1344x768 clip dropped from 342s to 130s — about 2.9x faster end to end.
- Caveat: the predictor is tagged preview and was trained on only a few resolutions and clip lengths.
- Verification tip: bypass the node with Ctrl+B, render the same seed with full attention, and compare clips side by side.
More from Infra
- Starlink now has 11,000+ satellites in orbit, two-thirds of all active satellites — XFreeze · 2026-10-06
- Ben Bajarin: Agentic AI Will Drive Datacenter CPU Demand, Scale-Up Domain Is the Key Battleground — BenBajarin · 2026-10-06
- llama.cpp adds DFlash speculative decoding for Qwen3.8-27B, faster than MTP — victormustar · 2026-10-06
- GLM 5.3 full NVFP4 deployable on 4x B200 or H200 with Marlin kernels — TheZachMueller · 2026-10-06
- Dev quantizes GLM-5.3-UNCENSORED to MXFP4, cutting size 44% for AMD GPUs — bakatristan · 2026-10-06
- OpenAI reportedly spent tens of millions in compute to crack Navier-Stokes in days — haider1 · 2026-10-06