ComfyUI INT4 Quantized Model Collection
Winougan · reddit · 2026-07-12
The author uploaded a batch of INT4 quantized ComfyUI models on Hugging Face, along with workflows and examples, emphasizing "free to use".
Key Info
- Aim to run models on lower VRAM; tested on RTX 3070 Ti and RTX 4090.
- Recommended environment: ComfyUI nightly, PyTorch 2.12, Python 3.13, cu132, Triton 3.8, FlashAttention 2, SageAttention 2.
- Performance claims:
- INT8 offers 25% improvement over BF16;
- INT4 offers 40%–50% improvement.
- Quality: INT8 nearly identical to BF16, INT4 close to FP8.
Uploaded/To be uploaded
- Uploaded: SeedVR 7b INT4, Gemma 3 12b INT4, Sulphur 2 Base INT4.
- Planned: Krea 2 Turbo, Klein9b, Z-Image Turbo, some Illustrious XL models, plus custom fine-tuned versions.
Related event: Multiple Image Models Get INT4 Quantization to Lower VRAM Requirements(2 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11