What Models Are People Running on 192GB Machines?
CentrifugalMalaise · reddit · 2026-07-19
A user with 192GB of RAM shares their experience running local LLMs.
They primarily run the Unsloth Dynamic UD-Q3KXL gguf for Qwen3.5-397B and mention using Claude to fix some llama.cpp issues related to hybrid recursive models, including:
- Fixing KV cache invalidation when vision is enabled
- Fixing instant recovery after saving session slots to disk
- An ongoing issue where tool calling breaks the KV cache, which they plan to fix next
They also list other LLMs they are interested in trying: GLM 4.7 357B, Deepseek V4 Flash 284B, Tencent Hy3 295B, Minimax M3 428B, Laguna M.1 225B, Minimax M2.7 229B, and Mimo 2.5 310B, while asking the community what others are running and how these compare to Qwen3.5-397B.
More from Infra
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- LFM2.5-8B-A1B doubles its tokenizer vocab and cuts on-device decoding time up to 3.7x — maximelabonne · 2026-07-22
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22