Inworld eval lead details how to evaluate free-form voice steering in TTS models

rdesh26 · x · 2026-09-26

Aleksey Tikhonov, head of evaluations at Inworld, published a detailed blog on evaluating "voice steering": their new Realtime TTS-2 follows free-form inline directions (e.g. [speak sadly], [whisper softly]), similar to Grok TTS, Gemini TTS, and ElevenLabs v3.

Key points:

A rare first-hand writeup for teams building speech/multimodal evals.

Original post →

More from Multimodal

Multimodal channel →