Black Forest Labs FLUX.3 Enters Early Testing, Pivoting to Omni-Modal Generation

Black Forest Labs' next-generation model, FLUX.3, has been opened for early testing, with demonstration videos released. Early testers have shared stunning generation samples, while rumors suggest the model will transition from a standalone image generator to an omni-modal backbone covering video, audio, and robotic action prediction.

Confirmed

The FLUX.3 model indeed exists, and official early access demonstration videos have been released. Several testers (like @ostrisai) have gained experience permissions, played with it for days, shared numerous generation results, and mentioned that audio was enabled. Additionally, the latest image samples suspected to be from Flux 3 or Flux Video have been shared. Test feedback indicates the model performs excellently in image quality and stylized generation; @Eric520CC believes it is enough to reshuffle the open-source image model landscape.

Unconfirmed

Currently, specific multimodal capabilities (such as up to 20 seconds of video, audio, robotic action generation, and reasoning) all stem from leaks and teaser info. @multimodalart points out that FLUX.3 [dev] is described as an open-weight multimodal backbone model, and @eyishazyer speculates it will likely be open-weight or open-source, but the official complete technical specs and open-source protocol have not yet been formally released.

Why it matters

If rumors are true, FLUX.3 will become a comprehensive multimodal model covering image, video, audio, and action prediction. This marks a significant move by a top-tier open-source image model into broader multimodal generation and physical world interaction (robotic actions), holding great significance for AI content creation and embodied intelligence.

2026-07-23 ~ 2026-07-24 · 11 related posts

Full story(13 episodes)→

Primary sources