FLUX 3 adds image, video, audio, and action prediction in one model

robrombach · x · 2026-07-24

A post circulating the FLUX 3 launch says the model combines image, video, audio, and action prediction in one system. The quoted tester claims it has the best native audio of any video generator they’ve used, can generate videos up to 20 seconds, and supports up to 10 references across image, video, and audio inputs, which can be mixed and matched.

Related event: Black Forest Labs Launches FLUX 3 Unified Multimodal Foundation Model(7 posts)→

Original post →

More from Multimodal

Multimodal channel →