Full-duplex speech model called the right MCP tool even when its transcript was "Bai coffee"

ivan_digital · reddit · 2026-08-20

The author wired NVIDIA's VoiceChat 11B full-duplex speech model to a local Apple Reminders MCP server via his own Swift/MLX runtime, running fully locally on an M5 Pro (7.5 GB RSS, 0.92 RTF for the whole pipeline).

The key finding: the model listens and speaks through one continuous network while emitting a separate function channel converted into MCP calls. In a recorded session the transcript wrongly showed "Bai coffee", yet the function channel still produced "Buy coffee", called createreminder, and wrote the correct reminder via EventKit in a 68 ms round trip — the function channel proved more reliable than the transcript.

For safety, only three tools were exposed (list/create/update); delete was deliberately omitted. He asks the community: for destructive voice-triggered tools, omit them from the schema entirely, or gate them behind an explicit confirmation call?

Original post →

More from coding & agent

coding & agent channel →