FULL STORY

FLUX 3: From Official Launch to Viral Tests

Black Forest Labs released FLUX 3, a unified multimodal model supporting native audio video generation. Its impressive capabilities quickly sparked widespread community testing and reviews.

2026-07-22 ~ 2026-08-06 · 8 episodes · 84 posts

Episode 1 · Black Forest Labs Launches FLUX 3: A Unified Multimodal Foundation Model (2026-07-22, 28 posts)

Black Forest Labs has officially released FLUX 3, positioning it as a vital step toward multimodal flow models and the foundational architecture for visual intelligence. Utilizing a unified architecture, the model comprehensively covers image, video, audio, and action prediction capabilities, signaling that generative AI is rapidly evolving into an omnimodal foundation with genuine world-understanding abilities.

Confirmed

  • Multimodal and Video Capabilities: FLUX 3 covers image, video, audio, and action generation. FLUX 3 Video uses a unified architecture capable of generating up to 20 seconds of video per run. It supports text-to-video, image-to-video, reference video generation, audio/video extension, keyframe control, multilingual dialogue, and smart multi-shot editing.
  • Release Roadmap: FLUX 3 Video has entered early access and will be open-sourced soon; native video audio generation is also in early testing. Official plans indicate that image editing and open-weights foundation models will be rolled out over the coming weeks to months, while action prediction will gradually open to select research and commercial partners.
  • Robotics Deployment: Early versions of FLUX 3 are already running on robots, demonstrating performance superior to REPA in robotic tasks. Mimic Robotics is among the first early access partners, co-developing the FLUX-mimic model, with Audi also participating in related deployments.
  • Image Generation: Official early samples of FLUX 3 Image have been released, showcasing mature performance across various styles, including macro photography, flat illustrations, and product posters.

Why it matters

The release of FLUX 3 marks a shift where generative models are no longer confined to digital content creation. The introduction of action prediction capabilities and early deployments on physical robots (such as the mimic robotics and Audi projects) highlight the massive potential of generative models in physical world interaction and embodied AI, introducing the novel concept of "Real World Models."

8 more related posts →

Episode 2 · Black Forest Labs FLUX.3 Enters Early Testing, Pivoting to Omni-Modal Generation (2026-07-23, 11 posts)

Black Forest Labs' next-generation model, FLUX.3, has been opened for early testing, with demonstration videos released. Early testers have shared stunning generation samples, while rumors suggest the model will transition from a standalone image generator to an omni-modal backbone covering video, audio, and robotic action prediction.

Confirmed

The FLUX.3 model indeed exists, and official early access demonstration videos have been released. Several testers (like @ostrisai) have gained experience permissions, played with it for days, shared numerous generation results, and mentioned that audio was enabled. Additionally, the latest image samples suspected to be from Flux 3 or Flux Video have been shared. Test feedback indicates the model performs excellently in image quality and stylized generation; @Eric520CC believes it is enough to reshuffle the open-source image model landscape.

Unconfirmed

Currently, specific multimodal capabilities (such as up to 20 seconds of video, audio, robotic action generation, and reasoning) all stem from leaks and teaser info. @multimodalart points out that FLUX.3 [dev] is described as an open-weight multimodal backbone model, and @eyishazyer speculates it will likely be open-weight or open-source, but the official complete technical specs and open-source protocol have not yet been formally released.

Why it matters

If rumors are true, FLUX.3 will become a comprehensive multimodal model covering image, video, audio, and action prediction. This marks a significant move by a top-tier open-source image model into broader multimodal generation and physical world interaction (robotic actions), holding great significance for AI content creation and embodied intelligence.

Episode 3 · BFL Launches FLUX 3: A Unified Multimodal Model for Video and Robotics (2026-07-24, 12 posts)

Black Forest Labs (BFL) has officially released FLUX 3, a unified multimodal foundation model that integrates image, video, native audio, and action prediction within a single architecture. FLUX 3 Video is currently in early access, and the team plans to roll out various capabilities progressively via API and private weights over the coming weeks and months. Demonstrating a profound understanding of the physical world, the model is viewed by the industry as a critical milestone toward general world simulators and robotic control.

Confirmed

FLUX 3 utilizes a unified architecture capable of simultaneously processing images, video, audio, and action prediction. The official release emphasizes that the new model's generation results across different styles are more closely aligned with real-world physics, and its native audio capability has received positive feedback in hands-on tests. Model capabilities will be gradually provided through APIs and private weights, with FLUX 3 Video currently in the Early Access phase for feedback collection and safety testing.

Unconfirmed

Regarding whether FLUX 3 will release a separate "Klein" model version, current official information suggests FLUX 3 is more of a collection of capabilities rather than a single standalone model. Furthermore, users point out that key metrics determining whether the model can truly land in a production environment—such as specific latency, concurrency limits, and pricing structures—remain undisclosed.

Why it matters

The core breakthrough of FLUX 3 lies in its action prediction capability. Several analysts (such as @MattVidPro and @rohanpaulai) note that unlike traditional VLA models relying solely on limited teleoperation data, FLUX 3 absorbs large-scale video pre-training data, thereby better grasping physical laws like contact, deformation, and temporal causality. Authors like @imjustnewatai further suggest that combining video models with robotic control validates the technological evolution route from video generation to world models, ultimately landing in embodied intelligence.

Episode 4 · Black Forest Labs Rumored to Release FLUX 3 (2026-07-26, 2 posts)

Rumors suggest Black Forest Labs is set to release FLUX 3, a new model that reportedly unifies image, video, audio, and action prediction within a single architecture.

Episode 5 · Black Forest Labs Launches FLUX 3 with Native Audio Video Generation (2026-07-28, 2 posts)

Black Forest Labs has launched FLUX 3 in limited access, a new multimodal model capable of generating images and up to 20 seconds of video with native audio output.

Episode 6 · Black Forest Labs Unveils FLUX 3 Multimodal Model with Native Audio Video Generation (2026-08-04, 19 posts)

Black Forest Labs officially released FLUX 3, a unified multimodal model trained on image, video, and audio. Its core highlight is the ability to generate up to 20-second HD videos with synchronized native audio in a single request, marking a significant step toward building a true 'world model'. The model is now in early access, with API and web interface available.

Confirmed

  • FLUX 3 supports text-to-video, image-to-video (with first/last frame or multi-keyframe control), video continuation, and multilingual dialogue.
  • Video generation offers 720p and 1080p resolutions, up to 20 seconds per clip, with native audio (including sound effects and ambient sound), plus a low-cost Draft mode.
  • The model uses a unified architecture trained jointly, enabling scene and camera angle changes within a single generation; reference-image-based variants are planned.
  • FLUX 3 is now publicly available via API and on Runware and Magnific, offering integrated generation and upscaling workflows.
  • The web interface is live; users can purchase credits to generate videos. Currently only video generation is available; other features like image generation will follow.
  • The model is available on Cloudflare AI Gateway, supporting text-to-video, image-to-video, and video-to-video continuation.
  • The company announced that model weights will be open-sourced soon; API access is already open to partners like Replicate, Krea, and Fal.

Unconfirmed

  • The exact timeline for open-sourcing weights has not been announced.
  • Details on model parameter size and training data have not been disclosed.

Why it matters

  • The company argues that a single modality only captures a projection of reality, while multimodal joint learning can better understand the underlying physics of the world through mutual constraints (e.g., sound matching impacts, motion following mass). FLUX 3's release is not just an improvement in video generation but a key attempt toward the technical vision of building a true 'world model'.

Episode 7 · Black Forest Labs' FLUX 3 Video Model Sparks Viral Sensation (2026-08-06, 4 posts)

Black Forest Labs' newly released FLUX 3 video model has gone viral for its stunning capabilities. It excels in prompt adherence and rendering, natively generating embedded audio and complex scenes, inspiring highly creative user-generated content.

Episode 8 · Flux 3 Video Model Released, Sparks Community Benchmarks (2026-08-06, 6 posts)

Black Forest Labs has released the open-source video generation model Flux 3, capable of generating 20-second clips with synchronized native audio in a single pass, triggering massive community testing and comparisons. Current conclusions show Flux 3 has distinct advantages in camera movement and physics, while competitors have their own strengths in duration and details.

已确认

  • 要点 Developed and open-sourced by Black Forest Labs, Flux 3 supports generating 20-second clips with synchronized audio in one go, and can produce movie titles and full action sequences.
  • 要点 The Seedance (ByteDance) versions range from 2.0 to 2.5, supporting up to 30 seconds of generation and 50 reference inputs, with smooth wide-angle and transition performance.
  • 要点 MiniMax H3 handles details, lighting, and temporal synchronization the best.

为什么重要

  • 要点 This hands-on comparison directly illustrates the core differences and fierce competition among current mainstream video generation models. Author @eyishazyer pointed out that Flux 3 performs exceptionally and stably in aerial camera movements, realistic physical destruction, complex transitions like car door openings, and hand-action synchronization. In contrast, Seedance 2.5 is slightly lagging in action fluidity, and its peak performance is a notch lower. Multi-scenario tests by @heypearlai also verified Flux 3's coherence and detail handling in continuous single-take room traversals and high-difficulty action scenes.