ComfyUI with Flux 2 Klein and Qwen Image: inference slows ~4x after a few runs
ROBOTTTTT13 · reddit · 2026-08-20
A Reddit user on a laptop (i7-12700H, RTX 4050 6GB, 16GB DDR5) tests FP8 Flux 2 Klein 9B and Qwen Image via newer ComfyUI's dynamic VRAM support (text encoder in GGUF; GGUF main models are much slower).
Symptom: the first 4-5 runs are fast — 27s for Flux, 45s for Qwen — then speed drops dramatically to 60s and nearly 200s respectively. An INT8 ConvRot Flux variant was faster initially but slowed down just as badly later. After days of troubleshooting they're asking the community for help.
More from Infra
- Opposition to local data centers in US surges 33 points to 75% — Polymarket · 2026-08-21
- Ramp Router cuts GPT-5.6 Sol inference costs by 50% — KlausCodes · 2026-08-21
- NVIDIA releases official CUDA MCP server for AI-assisted dev — swagonflyyyy · 2026-08-21
- Escha claims 2bit quant matches FP8 performance in benchmarks — luedtek · 2026-08-21
- Researcher Rants: Conference Season Blocks GPU Access for Days — ChongZzZhang · 2026-08-21
- AT&T routes 40% of employee AI usage to open models — Hesamation · 2026-08-21