Inworld eval lead details how to evaluate free-form voice steering in TTS models
rdesh26 · x · 2026-09-26
Aleksey Tikhonov, head of evaluations at Inworld, published a detailed blog on evaluating "voice steering": their new Realtime TTS-2 follows free-form inline directions (e.g. [speak sadly], [whisper softly]), similar to Grok TTS, Gemini TTS, and ElevenLabs v3.
Key points:
- Training: instead of fixed instruction lists, they trained the model to generalize to unseen instructions — and it worked
- Findings: LLM judges proved unreliable for grading instruction-following; instruction wording matters; steerability varies significantly across voices
- The open-ended instruction space had no known boundaries, forcing them to invent new ways to measure overall steerability and track improvements
A rare first-hand writeup for teams building speech/multimodal evals.
More from Multimodal
- xAI Reportedly Building a Music Composer for Grok Imagine — nima_owji · 2026-09-26
- Simple KJnodes workflow adds bounding-box control to Krea 2 text-to-image — Mulksky · 2026-09-26
- One Prompt Turns ChatGPT + Tesseract Into an AR UI Video Effects Generator — chrisfirst · 2026-09-26
- Mirage launches Tesseract, a video creative suite that lets AI agents edit and render video directly — chrisfirst · 2026-09-26
- Quiver Arrow 2 lands in Melius: prompt-to-SVG with editable vector paths — stuffyokodraws · 2026-09-26
- "Curve slop": critic dismisses AI animation demo as spline enumeration — ryunuck · 2026-09-26