FLUX 3 adds image, video, audio, and action prediction in one model
robrombach · x · 2026-07-24
A post circulating the FLUX 3 launch says the model combines image, video, audio, and action prediction in one system. The quoted tester claims it has the best native audio of any video generator they’ve used, can generate videos up to 20 seconds, and supports up to 10 references across image, video, and audio inputs, which can be mixed and matched.
Related event: Black Forest Labs Launches FLUX 3 Unified Multimodal Foundation Model(7 posts)→
More from Multimodal
- Runway adds natural-language workflows to its Agent for node-based pipelines — runwayml · 2026-07-24
- Ophilus launches Khora, an 8-player real-time multiplayer world model demo — thetripathi58 · 2026-07-24
- ComfyUI OpenPose Studio Adds Direct Hand Keypoint Editing — Inuya5haSama · 2026-07-24
- Apertus 1.5 goes public with chat access and image capabilities — valentina__py · 2026-07-24
- Microsoft’s TRELLIS.2 starts trending on Hugging Face Spaces — microsoft · 2026-07-24
- Seedance 2 prompt drives a first-person spaceship video demo — azed_ai · 2026-07-24