Independent Researcher Open-Sources Model That Turns Text Prompts Into Playable Synths

RoyalCities · reddit · 2026-09-10

Independent audio researcher RoyalCities has released Foundation-1, an open audio model that turns text prompts into fully playable synthesizers capable of generating infinite one-shot samples for music production. The core innovation: the model treats instrument and timbre as two separately controllable dimensions — the same grand piano can sound warm/gritty or cold/sparkly while still being a piano — a level of control no existing model offers. The hardest part was making timbre-locked keybeds stay consistent across multiple diffusion calls, which the author claims to have solved.

Everything is public: the model on Hugging Face (RoyalCities/Foundation-1), the inference pipeline forked from stable-audio-tools on GitHub, plus a full walkthrough video and demo-only clips, so others can vibe-code their own text-to-synth tools.

Related event: Indie Researcher Open-Sources Foundation-1 for Playable Synthesizers from Text(2 posts)→

Original post →

More from Multimodal

Multimodal channel →