ComfyUI INT4 Quantized Model Collection
Winougan · reddit · 2026-07-12
The author uploaded a batch of INT4 quantized ComfyUI models on Hugging Face, along with workflows and examples, emphasizing "free to use".
Key Info
- Aim to run models on lower VRAM; tested on RTX 3070 Ti and RTX 4090.
- Recommended environment: ComfyUI nightly, PyTorch 2.12, Python 3.13, cu132, Triton 3.8, FlashAttention 2, SageAttention 2.
- Performance claims:
- INT8 offers 25% improvement over BF16;
- INT4 offers 40%–50% improvement.
- Quality: INT8 nearly identical to BF16, INT4 close to FP8.
Uploaded/To be uploaded
- Uploaded: SeedVR 7b INT4, Gemma 3 12b INT4, Sulphur 2 Base INT4.
- Planned: Krea 2 Turbo, Klein9b, Z-Image Turbo, some Illustrious XL models, plus custom fine-tuned versions.
Related event: Multiple Image Models Get INT4 Quantization to Lower VRAM Requirements(2 posts)→
More from Infra
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22
- An energy model says the U.S. could run short of natural gas starting in 2028 — churchkey · 2026-07-22
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Arbitrum fee simulation shows higher gas capacity but lower L2 revenue under ArbOS61 — tomwanhh · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22