Ant Group open-sources MECT voiceprint model: 9.57M params rivals 587M at 100ms latency
aigclink · x · 2026-09-24
Ant Group has open-sourced AntSpeaker/MECT, a tiny, real-time voiceprint verification model.
- Brings MoE into fully supervised speaker verification: the 9.57M-param model matches (and on average beats) a 587M-param pretrained model on VoxCeleb1
- Causal retraining enables 100ms low-latency streaming inference with near-offline accuracy
- Target use cases: bank/call-center voice authentication, app login, device voice unlock
- Paper, Hugging Face weights and GitHub code are all public
It complements NVIDIA's recent speech logging model: one tracks who spoke when, the other whether it's the same speaker — both specialized components in voice pipelines.
More from Infra
- 10 open-source GitHub repos for building your own AI inference stack — Shruti_0810 · 2026-09-24
- Qualcomm commits to official Linux support for Snapdragon X2, Ubuntu certification by H1 2027 — carrycooldude · 2026-09-24
- Alibaba bets across the full stack: 5-10T Qwen models, 500K-card clusters, 20GW cloud by 2032 — Div_pradeep · 2026-09-24
- GeoPair: Training-Free Cross-Layer Factorization Hits SOTA in Transformer Compression — MTSAIR · 2026-09-24
- Chose rustpython-parser over tree-sitter after measuring both on real LLM output — SprayPuzzleheaded533 · 2026-09-24
- More compute made one agent 13x faster, another just 4%: agents' bottlenecks are task-dependent — alex_verem · 2026-09-24