Fish Audio Upgrades ASR: Speaker ID, Emotion Tags Like [laughter], 83 Languages
TheMoonMidas · x · 2026-09-29
Fish Audio shipped an update to its ASR model that goes beyond plain transcripts:
- Identifies different speakers
- Understands how speakers are feeling and tags emotion cues like [laughter] and [surprised] inline
- Supports 83 languages, which the company claims makes it the most accurate STT model available
Available to try now.
More from Multimodal
- MiniMax H3 ecosystem roundup: Character-Swap LoRA final version, retro sci-fi LoRA, ComfyUI Face Refine node — optimisticalish · 2026-09-29
- 7-Year-Old Voice Memo Turned Into a Music Video, Fully Coded by Claude — lellorocks · 2026-09-29
- Track Composed Entirely by Suno Sparks 'Musical AGI' Claims — astralmatrix · 2026-09-29
- Sonnet 5.5: 50% Cheaper Than Opus 5.5 and Insanely Fast for AI Video Making — petergyang · 2026-09-29
- vlo 0.3: open-source AI video editor built for generative compositing and frame-accurate inpainting — PxTicks · 2026-09-29
- A slightly different activation strength makes Qwen3-8B write a different song — sterlingcrispin · 2026-09-29