Suno launches Speech beta, the first audio model generating voice with matching music

suno · x · 2026-10-02

Suno has opened Speech (beta) to all users, calling it the first audio model that generates spoken audio and matching background music as one cohesive track. Users type a text and describe the desired voice and musical style; the team demos use cases like dramatic readings of friends' messages, epic-scored voice notes, meditations and bedtime stories. The company admits it's rough: British accents sometimes drift Australian and dramatic pauses get very dramatic. Available now via app update.

Original post →

More from Multimodal

Multimodal channel →