FLUX 3 adds image, video, audio, and action prediction in one model

GabGarrett · x · 2026-07-24

Black Forest Labs says FLUX 3 is a single multimodal model for image, video, audio, and action prediction. The company says FLUX 3 Video is available in early access, and that the model is jointly trained in one unified architecture.

The release also points to a robotics angle: the model can be extended to predict actions, and the thread references work with Mimic and Audi. The post frames this as an open-weight, state-of-the-art video model release.

Related event: Black Forest Labs Launches FLUX 3: A Unified Multimodal Foundation Model(28 posts)→

Original post →

More from Embodied

Embodied channel →