User asks whether a 373 GB model can run on a 5070 Ti with 16 GB of VRAM
Possible_Grocery8079 · reddit · 2026-07-29
Can a 373 GB model run on a 5070 Ti with 16 GB of VRAM?
A Reddit user asked whether it is worth trying to run a very large model such as GLM 5.2 UD-IQ4NL on a consumer PC, after seeing reports of massive models running on machines like a MacBook M1 Max.
The setup
- 64 GB dual-channel RAM at 5200 MT/s
- Intel i9-14900K
- RTX 5070 Ti with 16 GB VRAM
- Samsung 990 Pro SSD
- Backend: llama server
The user is considering a “we have frontier at home” setup using disk offload because internet access can be unreliable where they live.
What they are really asking
The post is a practical question about whether a 373 GB model is remotely usable on a consumer rig when most of the model would have to be streamed from SSD rather than kept in GPU memory.
More from Infra
- LiveKit says Gemma 4 31B hits 192 ms to first token in voice agents — GlennCameronjr · 2026-07-30
- Optimized Qwen Image 2512: 5x Smaller, 3x Faster Inference — enrique-byteshape · 2026-07-30
- Reddit user gets about 4 tokens/s running Kimi K3 on a 2×5090 home lab — iVoider · 2026-07-30
- AI Infrastructure Stocks Cool Down: CRWV at 52-Week Low, NVDA Down 8.69% — GaryMarcus · 2026-07-30
- Two RTX 3090s still struggle to fit Qwen Image Edit alongside a 27B text model — Civil_Fee_7862 · 2026-07-30
- llama.cpp prefill leaves CPU cores and memory bandwidth surprisingly idle — Dependent_Ad948 · 2026-07-30