CyberAgent Releases VAE Speech Align for Unsupervised Phoneme Alignment
kastnerkyle · x · 2026-08-26
CyberAgent released VAE Speech Align, a PyTorch-based tool for unsupervised phoneme alignment using VAEs and self-supervised learning (SSL) acoustic features. It includes pretrained models for English and Japanese, and allows training on custom datasets. The work was published at Interspeech 2024.
More from Multimodal
- Underwater ink bloom portrait prompt: suspended figures in flowing pigment — aziz4ai · 2026-08-26
- Alibaba Releases 180B Qwen3.8-Flash-Next Multimodal Model — ariG23498 · 2026-08-26
- Testing Qwen3.8 Vision: SVG Reconstruction & Anti-Benchmaxxing — bonobomaster · 2026-08-26
- LTX-2.5 Multishot Lip Sync on 8GB VRAM — big-boss_97 · 2026-08-26
- SREF Parameter Creates Anxious Animation Style with Muted Colors — tisch_eins · 2026-08-26
- Qwen Create generates cinematic steampunk detective montage with Wan 3.0 — socialwithaayan · 2026-08-26