What Models Are People Running on 192GB Machines?
CentrifugalMalaise · reddit · 2026-07-19
A user with 192GB of RAM shares their experience running local LLMs.
They primarily run the Unsloth Dynamic UD-Q3KXL gguf for Qwen3.5-397B and mention using Claude to fix some llama.cpp issues related to hybrid recursive models, including:
- Fixing KV cache invalidation when vision is enabled
- Fixing instant recovery after saving session slots to disk
- An ongoing issue where tool calling breaks the KV cache, which they plan to fix next
They also list other LLMs they are interested in trying: GLM 4.7 357B, Deepseek V4 Flash 284B, Tencent Hy3 295B, Minimax M3 428B, Laguna M.1 225B, Minimax M2.7 229B, and Mimo 2.5 310B, while asking the community what others are running and how these compare to Qwen3.5-397B.
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11