FULL STORY
Black Forest Labs Launches FLUX 3 Multimodal Model
Black Forest Labs officially launched the FLUX 3 model, marking a major step toward unified multimodal and embodied intelligence with significantly improved visual generation capabilities.
2026-07-22 ~ 2026-07-24 · 4 episodes · 41 posts
Episode 1 · Black Forest Labs Launches FLUX 3: A Unified Multimodal Foundation Model (2026-07-22, 28 posts)
Black Forest Labs has officially released FLUX 3, positioning it as a vital step toward multimodal flow models and the foundational architecture for visual intelligence. Utilizing a unified architecture, the model comprehensively covers image, video, audio, and action prediction capabilities, signaling that generative AI is rapidly evolving into an omnimodal foundation with genuine world-understanding abilities.
Confirmed
* **Multimodal and Video Capabilities**: FLUX 3 covers image, video, audio, and action generation. FLUX 3 Video uses a unified architecture capable of generating up to 20 seconds of video per run. It supports text-to-video, image-to-video, reference video generation, audio/video extension, keyframe control, multilingual dialogue, and smart multi-shot editing.
* **Release Roadmap**: FLUX 3 Video has entered early access and will be open-sourced soon; native video audio generation is also in early testing. Official plans indicate that image editing and open-weights foundation models will be rolled out over the coming weeks to months, while action prediction will gradually open to select research and commercial partners.
* **Robotics Deployment**: Early versions of FLUX 3 are already running on robots, demonstrating performance superior to REPA in robotic tasks. Mimic Robotics is among the first early access partners, co-developing the FLUX-mimic model, with Audi also participating in related deployments.
* **Image Generation**: Official early samples of FLUX 3 Image have been released, showcasing mature performance across various styles, including macro photography, flat illustrations, and product posters.
Why it matters
The release of FLUX 3 marks a shift where generative models are no longer confined to digital content creation. The introduction of action prediction capabilities and early deployments on physical robots (such as the mimic robotics and Audi projects) highlight the massive potential of generative models in physical world interaction and embodied AI, introducing the novel concept of "Real World Models."
- Flux 3 is teased as a multimodal model for image, video, audio and action — rerri · 2026-07-22
- Flux 3 teased as a multimodal model spanning image, video, audio and action — Angaisb_ · 2026-07-22
- Black Forest Labs’ Flux 3 is said to be a multimodal image, video, audio model — ZeroStateReflex · 2026-07-23
- Black Forest Labs appears to tease FLUX 3, a multimodal model with 20-second video generation — op7418 · 2026-07-23
- Black Forest Labs launches FLUX 3, a multimodal model for image, video, audio, and actions — bfl_ai · 2026-07-23
- FLUX 3 will add native audio, image editing, and open-weight multimodal access — bfl_ai · 2026-07-23
- FLUX 3 is already running on robots through mimic and Audi deployments — bfl_ai · 2026-07-23
- Black Forest Labs posts early FLUX 3 Image samples across several visual styles — bfl_ai · 2026-07-23
- Black Forest Labs positions FLUX 3 as a multimodal backbone for visual intelligence — stephen370 · 2026-07-23
- FLUX 3 is described as a Self-Flow system for multimodal generation — hila_chefer · 2026-07-23
- Black Forest Labs Launches FLUX-3: Unified Architecture for Images, Video, and Audio — nathanbenaich · 2026-07-23
- FLUX 3 introduces Real World Models for multimodal visual intelligence — MoistRecognition69 · 2026-07-23
- FLUX 3 is said to beat REPA on robotics with multimodal generation — hila_chefer · 2026-07-23
- Black Forest Labs says Flux 3 video generation is open source and full multimodal — op7418 · 2026-07-23
- BlackForestLabs Announces and Upcoming Open-Source Release of FLUX 3 Video Model — 歸藏的AI工具箱 · 2026-07-23
- Black Forest Labs announces FLUX 3, a multimodal model for image, video, and audio — chrisfirst · 2026-07-23
- Black Forest Labs launches FLUX 3, a unified multimodal model for image, video, audio, and action prediction — daniel_mac8 · 2026-07-24
- Flux 3 is pitched as one multimodal AI for images, video, and reasoning — mark_k · 2026-07-24
- FLUX 3 adds image, video, audio, and action prediction in one model — GabGarrett · 2026-07-24
- Stable Diffusion unveils FLUX 3 with image, video and native audio generation — imjustnewatai · 2026-07-24
Episode 2 · Black Forest Labs Teases FLUX.3: A Leap Towards an Omni-modal Backbone (2026-07-23, 7 posts)
Black Forest Labs is teasing its next-generation model, FLUX.3, opening early tests to select users. Leaked demo videos showcase significantly improved generation capabilities. According to tipsters, FLUX.3 will evolve from a static image model into an omni-modal backbone, supporting not only image generation but also up to 20 seconds of video, audio, and robotic motion prediction, potentially with built-in reasoning capabilities.
已确认
The FLUX.3 model is indeed real, and Black Forest Labs has released early access demo videos. Testers have already begun early testing, sharing stunning generation results.
尚未确认
Currently, details regarding the model's specific multimodal capabilities (such as 20-second video, audio, motion generation, and reasoning) stem purely from leaks and teaser rumors. Furthermore, while @eyishazyer and @multimodalart speculate that the model will likely follow its traditional route of open or open-source weights, the official full technical specifications and open-source licensing have yet to be confirmed.
为什么重要
If rumors hold true, FLUX.3 will be a comprehensive multimodal model covering image, video, audio, and motion prediction. This marks a major leap for top-tier open-source image models into broader multimodal generation and physical world interaction (robotic motion), holding significant implications for AI content creation and embodied AI.
- Flux 3 is rumored to add image, 20-second video, and reasoning support — mark_k · 2026-07-23
- FLUX 3 teaser hints at a multimodal generator with 20-second video clips — eyishazyer · 2026-07-23
- FLUX 3 Model Early Access Showcase Released — MFGREBEL · 2026-07-23
- Black Forest Labs teases FLUX.3 as an open-weights multimodal backbone — multimodalart · 2026-07-24
- Black Forest Labs' FLUX 3 Image Model Opens for Early Access — AIandDesign · 2026-07-24
- New Image Samples for Flux 3 / Flux Video Teased — koltregaskes · 2026-07-24
- FLUX 3 is back, according to a short post from @AIandDesign — AIandDesign · 2026-07-24
Episode 3 · Speculation Suggests OpenAI is Pivoting Towards Robotics (2026-07-24, 2 posts)
Observers speculate that products like FLUX 3 and Sora 2 reflect OpenAI's long-term strategy. The trajectory suggests a path from video models to world simulators, eventually leading to physical robotics capabilities.
- FLUX 3 may preview OpenAI’s path from video models to personal robots — imjustnewatai · 2026-07-24
- A thread collects evidence linking FLUX 3, Sora 2, and OpenAI robotics — imjustnewatai · 2026-07-24
Episode 4 · Black Forest Labs Releases FLUX 3 Unified Multimodal Model (2026-07-24, 4 posts)
Black Forest Labs has released FLUX 3, a unified multimodal model covering image, video, audio, and action prediction. The model will undergo an early access phase for feedback and safety testing.
- FLUX 3 unifies image, video, audio and action prediction in one model — robrombach · 2026-07-24
- BFL launches FLUX 3 as a unified model for image, video, audio, and action prediction — umesh_ai · 2026-07-24
- FLUX 3 unifies image, video, audio, and action prediction in one model — Eric520CC · 2026-07-24
- FLUX 3 expands into one multimodal backbone for image, video, audio and action — pmttyji · 2026-07-24