Friendly Stable Audio 3: open-source fork exposes full text-to-audio training pipeline
serrjoa · x · 2026-09-07
Researcher yukara-ikemiya released friendly-stable-audio-3, a refactored training-first fork of Stable Audio 3 for researchers and engineers:
- Full training pipeline: SAME autoencoder, flow-matching diffusion, latent CLAP, and ARC adversarial post-training — not just LoRA fine-tuning
- One generic Trainer: scripts/train.py trains all components
- Config-first: models fully described in YAML, no opaque modelconfig. downloads
- Distributed training via Hydra + HuggingFace Accelerate (multi-GPU/multi-node)
The upstream repo is inference-oriented; this fork exposes the complete training stack.
More from Multimodal
- Reddit user shares WIP AI-generated dark fantasy short film 'Wanderers' — DaWid_Shapiro · 2026-09-11
- Creator makes 2D electro-pop anime music video with just a prompt using MiniMax H3 — Hailuo_AI · 2026-09-11
- Street View to driving footage: GPT Astra fetches images, MiniMax H3 turns them into dashcam video — Hailuo_AI · 2026-09-11
- Single-author ECCV 2026 paper makes rolling shutter correction practical — ducha_aiki · 2026-09-11
- AI digital human covers Japanese classic so realistically viewers can't tell — JourneymanChina · 2026-09-11
- ComfyUI Style Explorer Adds LoRA Preview Catalog and Sharing — neonsparksuk · 2026-09-11