FLUX 3 adds image, video, audio, and action prediction in one model

robrombach · x · 2026-07-24

A post circulating the FLUX 3 launch says the model combines image, video, audio, and action prediction in one system. The quoted tester claims it has the best native audio of any video generator they’ve used, can generate videos up to 20 seconds, and supports up to 10 references across image, video, and audio inputs, which can be mixed and matched.

Related event: BFL Launches FLUX 3: A Unified Multimodal Model for Video and Robotics(12 posts)→

Original post →

More from Multimodal

Multimodal channel →