ComfyUI user weighs zram weight offloading: zstd decompression may bottleneck the inference hot path
Agreeable_Return2632 · reddit · 2026-09-27
A user running MiniMax H3 in ComfyUI with CPU offload (i5-14400 + RTX 4070 + 32GB RAM) considers zram to extend usable memory for offloaded layers. Rough math suggests zstd gives only 1.3-1.4x compression on near-random-entropy fp16 weights, adding 10-15GB effective capacity.
The catch, which the author flags themselves: offloaded layers are read on every forward pass, so decompression runs continuously in the hot path — RAM bandwidth is 35-50GB/s while zstd decompression tops out around 5-10GB/s even parallelized. They ask whether anyone has actually benchmarked zram for weight offloading in inference pipelines.
More from Infra
- IIT Delhi says it has built India's first indigenously designed micro-GPU — rvp · 2026-09-27
- Musk: China Will Solve Its Compute, Lithography and Chipmaking Constraints in 2-3 Years — haider1 · 2026-09-27
- Dev argues Vercel is 'unjustifiable' now that agents can safely drive Cloudflare — generativist · 2026-09-27
- ZeroHedge's GPU ROIC math uses wrong throughput, off by 8-10x — zephyr_z9 · 2026-09-27
- MLX poll: all top 3 community picks are powered by MLX-VLM — andrejusb · 2026-09-27
- $125M for 1,000 GB300s: The Brutal Compute Economics of a 'Neolab' — deedydas · 2026-09-27