Tencent Hunyuan deploys 1.25-bit model for Bilibili live translation
腾讯混元 · wechat · 2026-08-26
Tencent Hunyuan released extreme quantization schemes for the on-device translation model Hy-MT2-1.8B, compressing 3.3GB to 440MB (1.25-bit) or 574MB (2-bit) with near-zero quality loss.
Technical Schemes:
- 2-bit: Uses Stretchable Elastic Quantization (SEQ) and Quantization-Aware Distillation (QAD) for high-end devices.
- 1.25-bit: Based on in-house Sherry technology (ACL 2026 Oral) using fine-grained sparsity, optimized for all devices with STQ kernel and SIMD support.
Deployment & Optimization:
- Intel Collab: Optimized x86 operators; VNNI instruction fusion boosts inference speed by 2.7x-5x.
- Bilibili Use Case: Integrated into live stream subtitle translation. Latency is 500-800ms per comment with 500-700MB RAM usage, ensuring privacy via on-device processing.
More from Infra
- Jalapeno chip shows strength, revealing Nvidia's inference architecture weaknesses — beffjezos · 2026-08-26
- NVIDIA claims up to 30× agentic throughput per MW on Vera Rubin — Crescitaly · 2026-08-26
- Australia's PM backs down on requiring AI datacentres to run fully on renewable energy — nordicinst · 2026-08-26
- Single RTX 5090 Benchmark: 27B Model at 616 Tok/s with 262K Context — EAccelerate_42 · 2026-08-26
- Running Qwen3.8-27B for local coding on 16GB VRAM: full setup guide — Due-Project-7507 · 2026-08-26
- Homelab Build: Choosing Between 5060 Ti and RX 9070 for LLM Server — MrCatberry · 2026-08-26