Black Forest Labs Teases FLUX 3 Multimodal Model

Black Forest Labs is teasing its next-generation model, FLUX.3, opening early tests to select users. Leaked demo videos showcase significantly improved generation capabilities. According to tipsters, FLUX.3 will evolve from a static image model into an omni-modal backbone, supporting not only image generation but also up to 20 seconds of video, audio, and robotic motion prediction, potentially with built-in reasoning capabilities.

已确认

The FLUX.3 model is indeed real, and Black Forest Labs has released early access demo videos. Testers have already begun early testing, sharing stunning generation results.

尚未确认

Currently, details regarding the model's specific multimodal capabilities (such as 20-second video, audio, motion generation, and reasoning) stem purely from leaks and teaser rumors. Furthermore, while @eyishazyer and @multimodalart speculate that the model will likely follow its traditional route of open or open-source weights, the official full technical specifications and open-source licensing have yet to be confirmed.

为什么重要

If rumors hold true, FLUX.3 will be a comprehensive multimodal model covering image, video, audio, and motion prediction. This marks a major leap for top-tier open-source image models into broader multimodal generation and physical world interaction (robotic motion), holding significant implications for AI content creation and embodied AI.

2026-07-23 ~ 2026-07-24 · 5 related posts

Primary sources