Wish list: a Qwen4 27B with 100B+ Engram offloaded to RAM and NVMe for local users
casper_hansen_ · x · 2026-10-03
Casper Hansen pitches a dream local model: a Qwen4 27B paired with a 100B+ parameter Engram that offloads to RAM and NVMe — running on small VRAM while retaining the world knowledge of a large MoE. He believes the local model community would love it.
More from Infra
- GKE adds CPU Startup Boost: faster pod starts without over-provisioning — rseroter · 2026-10-03
- Huawei claims Ascend overtook Nvidia in China share; supply, not demand, is the bottleneck — teortaxesTex · 2026-10-03
- Turning an iPhone into a second GPU for a MacBook: 44% faster prefill on Qwen 27B — StayLameBro · 2026-10-03
- Micron CEO: memory supply will be much tighter in 2027-2028 than 2026 — dankvr · 2026-10-03
- mamf-finder adds FP8/MXFP4/NVFP4 support for real GPU TFLOPS benchmarking — StasBekman · 2026-10-03
- Measured on B200: nvfp4 is ~9% more efficient than mxfp4 with higher accuracy — pick nvfp4 on Blackwell — StasBekman · 2026-10-03