ComfyUI INT4 Plugin Enables Fast Gen on Low VRAM

Limp-Chemical4707 · reddit · 2026-07-10

A developer has released a custom node package named ComfyUI-INT4-Fast, enabling native, ultra-fast, low-memory INT4 (W4A4) model execution in ComfyUI. The project supports native mixed-precision checkpoints, dynamically parses model metadata, routes main blocks to INT4 Tensor Cores, and delegates sensitive layers to INT8 execution, thereby avoiding dimension errors and preserving generation quality.

Performance-wise, running the Krea2 Turbo model on an RTX 3060 (6GB VRAM) with 32GB RAM achieves a speed of 1.78 seconds/step at 1024x1024 resolution, with a total time of about 17.64 seconds.

Related event: ComfyUI Introduces INT4 Plugin for Low-VRAM Fast Inference(2 posts)→

Original post →

More from Infra

Infra channel →