Microsoft ships three voice models: MAI-Transcribe-2-Streaming and MAI-Voice-2.1 family
MicrosoftAI · x · 2026-10-02
Microsoft announced three new models: MAI-Transcribe-2-Streaming for accurate streaming transcription, plus MAI-Voice-2.1 and MAI-Voice-2.1-Flash, offering more natural speech and less waiting between turns — aimed at building voice agents that keep conversations moving.
More from Multimodal
- A 3-pass workflow for restoring old photos with AI: repair, colorize, then sharpen — xiaohu · 2026-10-02
- Netflix's Align Then Reason lip-sync judge boosts mean AUC by up to 59% — netflix · 2026-10-02
- PROWBench: 170 Programmable Episodes Test Whether Video World Models Render What the Program Specifies — AlayaLab · 2026-10-02
- Stability AI's 4Director Controls Video World Models with Rigid 3D Geometry — stabilityai · 2026-10-02
- PixelDense: Dense-Prediction Teachers Beat Semantic Encoders for Diffusion REPA, GenEval Hits 0.8093 — Lehan Yang · 2026-10-02
- Opus 5.5 artwork sparks GPT-Image-2 vs Z-Image style comparison — teortaxesTex · 2026-10-02