Stable Diffusion unveils FLUX 3 with image, video and native audio generation

imjustnewatai · x · 2026-07-24

The post says the Stable Diffusion team has unveiled FLUX 3, a model that can generate images, video, and native audio.

It also claims the same model now controls robots tested and deployed at Audi, with a 101 ms reaction time, framing it as a step from content generation toward acting in the physical world.

Related event: Black Forest Labs Unveils Omnimodal FLUX 3(24 posts)→

Original post →

More from Embodied

Embodied channel →