Independent Researcher Open-Sources Model That Turns Text Prompts Into Playable Synths
RoyalCities · reddit · 2026-09-10
Independent audio researcher RoyalCities has released Foundation-1, an open audio model that turns text prompts into fully playable synthesizers capable of generating infinite one-shot samples for music production. The core innovation: the model treats instrument and timbre as two separately controllable dimensions — the same grand piano can sound warm/gritty or cold/sparkly while still being a piano — a level of control no existing model offers. The hardest part was making timbre-locked keybeds stay consistent across multiple diffusion calls, which the author claims to have solved.
Everything is public: the model on Hugging Face (RoyalCities/Foundation-1), the inference pipeline forked from stable-audio-tools on GitHub, plus a full walkthrough video and demo-only clips, so others can vibe-code their own text-to-synth tools.
More from Multimodal
- Composer used Fable with Claude Code to turn a PDF score and wav into video, automating it into mtdt — doodlestein · 2026-09-10
- A Song Written by the Robots After AI Kills Us All: 'We Leave the Lights On' — AIandDesign · 2026-09-10
- H3 Turbo Streams Real-Time With Up to 9 Reference Images — boudaboy · 2026-09-10
- fal keeps shipping new models and has fixed its sketchy billing, user notes — nijfranck · 2026-09-10
- The Prompt Behind That 75M-View Viral Video Is Finally Out — techhalla · 2026-09-10
- Stanford releases RenderFormer-V2: transformer neural rendering with heterogeneous scene support — Stanford · 2026-09-10