FLUX 3 expands into one multimodal backbone for image, video, audio and action

pmttyji · reddit · 2026-07-24

FLUX 3 is introduced as a single multimodal backbone for image, video, audio, and action prediction.

The launch plan says the team will roll out capabilities over the coming weeks and months after an early-access phase for feedback and safety testing. Planned releases include:

The post also says more technical details about the underlying approach will be shared later.

Related event: Black Forest Labs Releases FLUX 3 Unified Multimodal Model(4 posts)→

Original post →

More from Multimodal

Multimodal channel →