Microsoft expands real-time speech stack with MAI-Transcribe-2-Streaming and MAI-Voice
thione · x · 2026-10-06
Microsoft launched MAI-Transcribe-2-Streaming and two MAI-Voice models, expanding its real-time speech stack for conversational agents.
- The streaming transcription model targets low-latency, real-time use, while the MAI-Voice models cover speech generation — together enabling agents to listen and speak in real time.
- The release extends Microsoft's in-house MAI model family into voice, positioning it as speech infrastructure for enterprise conversational agents.
- Also in the post: Google DeepMind introduced SynthID Bio, a proof of concept that watermarked AI-designed proteins without compromising biological function in lab tests.
More from Models
- Researcher: Chinese models' social dynamics in Delvetown are understudied and underestimated — lfschiavo · 2026-10-06
- NVIDIA's Nemotron coding model scores 535.4/600 on IOI 2026, beating top human, now on Hugging Face — NVIDIAAI · 2026-10-06
- Abliterated Qwen3.8-Flash GGUF quants land on Hugging Face — SC117 · 2026-10-06
- Several Western open-weight models launching this month, Reflection AI's first to rival top Chinese models — latkins · 2026-10-06
- GPT, Claude, Gemini and Grok suggest nearly identical teammate names — homogenization everywhere — sergeykarayev · 2026-10-06
- SemiAnalysis: Anthropic subscriptions offer 5x+ more value than OpenAI's — scaling01 · 2026-10-06