ComfyUI INT4 Quantized Model Collection
Winougan · reddit · 2026-07-12
The author uploaded a batch of INT4 quantized ComfyUI models on Hugging Face, along with workflows and examples, emphasizing "free to use".
Key Info
- Aim to run models on lower VRAM; tested on RTX 3070 Ti and RTX 4090.
- Recommended environment: ComfyUI nightly, PyTorch 2.12, Python 3.13, cu132, Triton 3.8, FlashAttention 2, SageAttention 2.
- Performance claims:
- INT8 offers 25% improvement over BF16;
- INT4 offers 40%–50% improvement.
- Quality: INT8 nearly identical to BF16, INT4 close to FP8.
Uploaded/To be uploaded
- Uploaded: SeedVR 7b INT4, Gemma 3 12b INT4, Sulphur 2 Base INT4.
- Planned: Krea 2 Turbo, Klein9b, Z-Image Turbo, some Illustrious XL models, plus custom fine-tuned versions.
Related event: Multiple Image Models Get INT4 Quantization to Lower VRAM Requirements(2 posts)→
More from Infra
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11