Black Forest Labs launches FLUX 3, a unified multimodal model for image, video, audio, and action prediction
daniel_mac8 · x · 2026-07-24
Black Forest Labs says FLUX 3 is a single multimodal model spanning image, video, audio, and action prediction.
- Capabilities: text-to-video, image-to-video, reference-video generation, video+audio extension, keyframe-based generation, multilingual prompting, and longer multi-shot sequences.
- Architecture: jointly trained in one unified system, with an explicit path to action prediction for robotics.
- Access: FLUX 3 Video is already in early access; the company says API access will open in the coming weeks, followed by the open-source FLUX 3 Dev model.
- The thread also points to demos with mimic and Audi.
Related event: Black Forest Labs Unveils Omnimodal FLUX 3(22 posts)→
More from Embodied
- VOX pitches a voice-first HCI device that works in noisy environments — jmusk · 2026-07-24
- FLUX 3 may preview OpenAI’s path from video models to personal robots — imjustnewatai · 2026-07-24
- NYC robotics hackathon offers robot arms, cameras, and compute to software engineers — ditzikow · 2026-07-24
- YC highlights robocurve, a benchmark for how robots perform in the physical world — ycombinator · 2026-07-24
- RoboMME benchmarks robot memory across 16 tasks with 14 policies — chris_j_paxton · 2026-07-24
- FLUX 3 unifies image, video, audio and action prediction in one multimodal model — pess_r · 2026-07-24