Inferact releases GLM-5.3 NVFP4 version, cutting memory to 465GB
vllm_project · x · 2026-08-28
Inferact has launched an NVFP4 build of GLM-5.3 alongside vLLM on day one. This version reduces memory requirements to 465GB, compared to 756GB for the FP8 checkpoint.
More from Infra
- OpenAI Python SDK migrates to HTTPX2, drops httpx dependency — mitsuhiko · 2026-08-29
- Best local setup for anime img2img/inpainting: WebUI, models, and pose editing workflows — Mystvearn_ · 2026-08-29
- x401 Protocol: HTTP-based proof requirement for automated access — csuwildcat · 2026-08-29
- GLM 5.3 open weights arrive; DFlash 2 speculative decoding hits 4.4x FP8 throughput — gan_chuang · 2026-08-29
- Anthropic, Nvidia, and Google surge together, linking model, chip, and platform layers — YvesMulkers · 2026-08-29
- Qualcomm's data center opportunity becomes tangible with custom ARM CPUs — BenBajarin · 2026-08-28