User seeks a custom GGUF quant to fit GLM on a 192 GB RAM Mac between Q2 and Q4
CentrifugalMalaise · reddit · 2026-09-08
A user asks whether a GLM 5.3 Flash Antirez/DS4 GGUF targeted at 192 GB RAM is possible or how to build one.
- Antirez's Q2 at 96.5 GB loses quality versus Q4 and wastes 85 GB of the Mac's RAM, while the Q4 at 191 GB leaves no room for the OS, let alone KV cache
- Other Q4 and Q4/Q8 mix GGUFs and MLX conversions not intended for the DS4 range run 150–180 GB, which would suit him, but he wants to stick with DS4
- He is seeking community help on quantization options
More from Infra
- Hyperscaler backlog hits $1.7T as Citi conference flags shift to agentic inference — sanjaykalra · 2026-09-08
- AI data center interconnect chip startup Celero raises $275M at $3B+ valuation — dinabass · 2026-09-08
- vLLM's Speculators v0.8.0 ships unified CLI, PyPI Mooncake connectors, fused Triton loss kernel — vllm_project · 2026-09-08
- Dev buys a Mac mini just to run Codex 24/7 across his entire workflow — _AustinCalvert_ · 2026-09-08
- INT21's agent-generated Qwen3.8 trainer hits 11.5x PyTorch FSDP2 throughput on 8 B200s — bingxu_ · 2026-09-08
- Benchmarking Gemma under 16-128 concurrent users: classification vs generation loads — rseroter · 2026-09-08