Tencent Hy4-preview shrunk to 214GB via Sherry quantization
alejandroll10 · x · 2026-09-02
Tencent announced shrinking Hy4-preview from 1.5TB to 214GB using the Sherry quantization method, packing weights to 1.25 bits each. It enables stitching existing GPUs across machines. Using MIX-STQ10, calibration data selects bit-width per layer (some down to 1.31-bit). Accuracy drop is minimal vs BF16: MCP Atlas 83.7→83.2, SWE-Bench 82.9→81.3. Weights and GGUFs are released.
More from Models
- Gemini Adds Agentic Video Understanding, Cuts Token Usage by 88% — GoogleDeepMind · 2026-09-02
- Fable 5.1 spotted in Claude support docs, release appears imminent — kimmonismus · 2026-09-02
- Meta's Muse Voice Transcribe Balances Speed and Accuracy with Adaptive Delay — AIatMeta · 2026-09-02
- Fable and Sol Ultra models accused of subtle hallucinations — StewartalsopIII · 2026-09-02
- Feed LLMs a table of pure noise and they'll confidently invent 'sensor data' — No-Plant-5234 · 2026-09-02
- NVIDIA Sol Engine speeds up MiniMax H3 video generation by 27.7x — DavidmComfort · 2026-09-02