Meta's speech model targets messy real rooms as the ears for its glasses and agent stack
heypearlai · x · 2026-09-04
Meta's new streaming speech system is positioned as the "ears" for its AI glasses and agent stack, built to handle messy real-world rooms with multiple speakers rather than waiting for a clean "hey Meta" wake command.
- It is live through the Meta Model API, Meta AI for Mac, and Muse Code.
- The author stays skeptical: current demos ran with 20 consenting recorded participants, and performance in an actual chaotic living room remains unproven.
Related event: Meta's New Speech Model Tops Leaderboard, Built for Noisy Rooms(3 posts)→
More from Embodied
- Simulation physics gaps teach robots tricks that fail in the real world — binarybits · 2026-09-04
- How robot startups scrape for data: free cleanings, exoskeletons, sim limits — binarybits · 2026-09-04
- Hyper3D's WorldGen turns one photo into an interactive 3D world with physics — petewoodbridge · 2026-09-04
- Farming is among the most advanced robotics: cows voluntarily line up for auto-milkers — yacineMTB · 2026-09-04
- Microduck robot pushed to 1.9 m/s in simulation — kevin_zakka · 2026-09-04
- ROSCon Global 2026 heads to Toronto, September 22-24 — chrismatthieu · 2026-09-04