Indie researcher releases open-source audio model that turns text prompts into playable synths

RoyalCities · reddit · 2026-09-10

Reddit user RoyalCities released Foundation-1, a self-trained audio model that treats instrument and timbre as independently controllable dimensions — a single piano patch can sound warm/gritty or cold/sparkly, a control level he says no existing model offers. He solved the hard part: keeping timbre-locked keybeds consistent across multiple diffusion calls.

Everything is open:

The model also generates infinite one-shots for music production.

Related event: Indie Researcher Open-Sources Foundation-1 for Playable Synthesizers from Text(2 posts)→

Original post →

More from Multimodal

Multimodal channel →