WeeLLM: Run FLUX.1-dev on 4GB VRAM Without Quantization

AlarmingPhrase8174 · reddit · 2026-08-16

The author released WeeLLM, a layer-streaming inference engine that runs large image models on low-VRAM devices without quantization. By streaming transformer layers one by one from disk, it runs the 12B FLUX.1-dev on an RTX 3050 (4GB) with just 1.51GB peak VRAM. It supports 24 models but requires NVMe SSD speeds and currently only supports text-to-image.

Original post →

More from coding & agent

coding & agent channel →