Moonshot says its MXFP4 quantization matches the format used in hosted API
kimmonismus · x · 2026-07-28
Moonshot confirmed in a Reddit AMA that its released MXFP4 quantization format is the same one used in the hosted API.
That means self-hosted inference quality should line up with what people have already been testing since the July 16 launch, which is useful for anyone evaluating deployment parity and runtime behavior.
More from Infra
- LightRAG v1.5.3 adds a one-time Milvus migration and hardens production edge cases — JeremyCMorgan · 2026-07-28
- llama.cpp adds DSpark speculative decoding and asks for speed results — pmttyji · 2026-07-28
- Nvidia puts an open vision-language-action model on Hugging Face — theteknosaur · 2026-07-28
- “Model eats harness” is really about deployment-driven harness evolution — m_wulfmeier · 2026-07-28
- Agentic AI system reportedly cost $1.2M a month before being cut to $100K — emmanuelvivier · 2026-07-28
- Wistron’s Early Nvidia Bet Turned It Into One of AI’s Biggest Supply-Chain Winners — pstAsiatech · 2026-07-28