Dev quantizes GLM-5.3-UNCENSORED to MXFP4, cutting size 44% for AMD GPUs
bakatristan · reddit · 2026-10-06
Reddit user bakatristan quantized dealignai's GLM-5.3-UNCENSORED-FP8 to MXFP4 and uploaded the weights to Hugging Face, filling a gap for AMD GPU users who had no compatible version.
- Weights total 423.75 GB, about 44% smaller than the FP8 source
- Converted with AMD Quark on an MI355X server; expert weights use MXFP4 while attention, routers and other sensitive layers stay at higher precision
- The README documents the source revision, quantization details, measured stats and validation results, plus conversion scripts
More from Infra
- Mistral Large 4 trained on just 4,000 Grace Blackwell GPUs vs 100,000 for Astra — steipete · 2026-10-07
- SpaceX shows off Starmind plan: 1M-satellite Starlink clusters with 10 Tb/s links for AI compute — XFreeze · 2026-10-07
- OpenSSH shifts to faster releases as AI-discovered security bugs pile up — jedisct1 · 2026-10-07
- Lambda Releases AIPerf: Benchmarking Models Under Real User Workloads — TheZachMueller · 2026-10-07
- Google launches EmbeddingGemma 2, a 740M open multimodal embedding model for on-device use — sundarpichai · 2026-10-07
- GPUs idle 85-95% waiting on memory: why the whole AI stack is being rebuilt — alex_verem · 2026-10-07