NVAlign: Direct-Gradient Post-Training Improves Non-Verbal Control in Flow-Matching TTS
BlandAIOrg · hf · 2026-09-30
- Problem: Modern TTS supports inline non-verbal vocalization (NVV) tags, but accurate control of these events lacked established post-training methods for continuous autoregressive flow-matching TTS.
- Approach: NVAlign fine-tunes the TTS model and an NVV-aware ASR (frozen as reward model), then backpropagates rewards through the flow-matching sampler via a two-step gradient surrogate, jointly updating the autoregressive backbone and acoustic flow head, with fidelity penalties and reference-velocity regularization to preserve speaker similarity.
- Results: Outperforms SFT and Flow-GRPO baselines on NVV-SuperBench and human listening tests. Audio samples at nvalign.github.io.
More from Research
- Gary Marcus doubles down on 2020 thesis: LLMs alone aren't enough for robust AI — GaryMarcus · 2026-09-30
- DepthBench paper finds Pre-LN variants hit a depth wall, comparing 10 residual designs — teortaxesTex · 2026-09-30
- Xcelsa's Apex AI fixes chip timing violation in 3 hours vs 1 month — IanAndrewsDC · 2026-09-30
- Fatima Fellowship opens first Fall cohort after 700+ spring applicants — deliprao · 2026-09-30
- flyverse: open framework runs whole fly connectomes in strange simulated worlds — repligate · 2026-09-30
- Daniel Litt: once models write well, human writing will mainly help you think — littmath · 2026-09-30