Moonshot says its MXFP4 quantization matches the format used in hosted API
kimmonismus · x · 2026-07-28
Moonshot confirmed in a Reddit AMA that its released MXFP4 quantization format is the same one used in the hosted API.
That means self-hosted inference quality should line up with what people have already been testing since the July 16 launch, which is useful for anyone evaluating deployment parity and runtime behavior.
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23