FULL STORY

Black Forest Labs Launches FLUX 3 Multimodal Model

Black Forest Labs officially launched the FLUX 3 model, marking a major step toward unified multimodal and embodied intelligence with significantly improved visual generation capabilities.

2026-07-22 ~ 2026-07-24 · 4 episodes · 41 posts

Episode 1 · Black Forest Labs Launches FLUX 3: A Unified Multimodal Foundation Model (2026-07-22, 28 posts)

Black Forest Labs has officially released FLUX 3, positioning it as a vital step toward multimodal flow models and the foundational architecture for visual intelligence. Utilizing a unified architecture, the model comprehensively covers image, video, audio, and action prediction capabilities, signaling that generative AI is rapidly evolving into an omnimodal foundation with genuine world-understanding abilities.

Confirmed

* **Multimodal and Video Capabilities**: FLUX 3 covers image, video, audio, and action generation. FLUX 3 Video uses a unified architecture capable of generating up to 20 seconds of video per run. It supports text-to-video, image-to-video, reference video generation, audio/video extension, keyframe control, multilingual dialogue, and smart multi-shot editing.

* **Release Roadmap**: FLUX 3 Video has entered early access and will be open-sourced soon; native video audio generation is also in early testing. Official plans indicate that image editing and open-weights foundation models will be rolled out over the coming weeks to months, while action prediction will gradually open to select research and commercial partners.

* **Robotics Deployment**: Early versions of FLUX 3 are already running on robots, demonstrating performance superior to REPA in robotic tasks. Mimic Robotics is among the first early access partners, co-developing the FLUX-mimic model, with Audi also participating in related deployments.

* **Image Generation**: Official early samples of FLUX 3 Image have been released, showcasing mature performance across various styles, including macro photography, flat illustrations, and product posters.

Why it matters

The release of FLUX 3 marks a shift where generative models are no longer confined to digital content creation. The introduction of action prediction capabilities and early deployments on physical robots (such as the mimic robotics and Audi projects) highlight the massive potential of generative models in physical world interaction and embodied AI, introducing the novel concept of "Real World Models."

8 more related posts →

Episode 2 · Black Forest Labs Teases FLUX.3: A Leap Towards an Omni-modal Backbone (2026-07-23, 7 posts)

Black Forest Labs is teasing its next-generation model, FLUX.3, opening early tests to select users. Leaked demo videos showcase significantly improved generation capabilities. According to tipsters, FLUX.3 will evolve from a static image model into an omni-modal backbone, supporting not only image generation but also up to 20 seconds of video, audio, and robotic motion prediction, potentially with built-in reasoning capabilities.

已确认

The FLUX.3 model is indeed real, and Black Forest Labs has released early access demo videos. Testers have already begun early testing, sharing stunning generation results.

尚未确认

Currently, details regarding the model's specific multimodal capabilities (such as 20-second video, audio, motion generation, and reasoning) stem purely from leaks and teaser rumors. Furthermore, while @eyishazyer and @multimodalart speculate that the model will likely follow its traditional route of open or open-source weights, the official full technical specifications and open-source licensing have yet to be confirmed.

为什么重要

If rumors hold true, FLUX.3 will be a comprehensive multimodal model covering image, video, audio, and motion prediction. This marks a major leap for top-tier open-source image models into broader multimodal generation and physical world interaction (robotic motion), holding significant implications for AI content creation and embodied AI.

Episode 3 · Speculation Suggests OpenAI is Pivoting Towards Robotics (2026-07-24, 2 posts)

Observers speculate that products like FLUX 3 and Sora 2 reflect OpenAI's long-term strategy. The trajectory suggests a path from video models to world simulators, eventually leading to physical robotics capabilities.

Episode 4 · Black Forest Labs Releases FLUX 3 Unified Multimodal Model (2026-07-24, 4 posts)

Black Forest Labs has released FLUX 3, a unified multimodal model covering image, video, audio, and action prediction. The model will undergo an early access phase for feedback and safety testing.