FULL STORY

FLUX 3: From Launch to Top of Video Leaderboards

Black Forest Labs released FLUX 3, a unified multimodal model for image, video, and audio generation. Following its launch, it sparked massive testing and ranked highly on leaderboards for its stunning realism.

2026-07-22 ~ 2026-08-14 · 13 episodes · 100 posts

Episode 1 · Black Forest Labs Launches FLUX 3: A Unified Multimodal Foundation Model (2026-07-22, 28 posts)

Black Forest Labs has officially released FLUX 3, positioning it as a vital step toward multimodal flow models and the foundational architecture for visual intelligence. Utilizing a unified architecture, the model comprehensively covers image, video, audio, and action prediction capabilities, signaling that generative AI is rapidly evolving into an omnimodal foundation with genuine world-understanding abilities.

Confirmed

  • Multimodal and Video Capabilities: FLUX 3 covers image, video, audio, and action generation. FLUX 3 Video uses a unified architecture capable of generating up to 20 seconds of video per run. It supports text-to-video, image-to-video, reference video generation, audio/video extension, keyframe control, multilingual dialogue, and smart multi-shot editing.
  • Release Roadmap: FLUX 3 Video has entered early access and will be open-sourced soon; native video audio generation is also in early testing. Official plans indicate that image editing and open-weights foundation models will be rolled out over the coming weeks to months, while action prediction will gradually open to select research and commercial partners.
  • Robotics Deployment: Early versions of FLUX 3 are already running on robots, demonstrating performance superior to REPA in robotic tasks. Mimic Robotics is among the first early access partners, co-developing the FLUX-mimic model, with Audi also participating in related deployments.
  • Image Generation: Official early samples of FLUX 3 Image have been released, showcasing mature performance across various styles, including macro photography, flat illustrations, and product posters.

Why it matters

The release of FLUX 3 marks a shift where generative models are no longer confined to digital content creation. The introduction of action prediction capabilities and early deployments on physical robots (such as the mimic robotics and Audi projects) highlight the massive potential of generative models in physical world interaction and embodied AI, introducing the novel concept of "Real World Models."

8 more related posts →

Episode 2 · Black Forest Labs FLUX.3 Enters Early Testing, Pivoting to Omni-Modal Generation (2026-07-23, 11 posts)

Black Forest Labs' next-generation model, FLUX.3, has been opened for early testing, with demonstration videos released. Early testers have shared stunning generation samples, while rumors suggest the model will transition from a standalone image generator to an omni-modal backbone covering video, audio, and robotic action prediction.

Confirmed

The FLUX.3 model indeed exists, and official early access demonstration videos have been released. Several testers (like @ostrisai) have gained experience permissions, played with it for days, shared numerous generation results, and mentioned that audio was enabled. Additionally, the latest image samples suspected to be from Flux 3 or Flux Video have been shared. Test feedback indicates the model performs excellently in image quality and stylized generation; @Eric520CC believes it is enough to reshuffle the open-source image model landscape.

Unconfirmed

Currently, specific multimodal capabilities (such as up to 20 seconds of video, audio, robotic action generation, and reasoning) all stem from leaks and teaser info. @multimodalart points out that FLUX.3 [dev] is described as an open-weight multimodal backbone model, and @eyishazyer speculates it will likely be open-weight or open-source, but the official complete technical specs and open-source protocol have not yet been formally released.

Why it matters

If rumors are true, FLUX.3 will become a comprehensive multimodal model covering image, video, audio, and action prediction. This marks a significant move by a top-tier open-source image model into broader multimodal generation and physical world interaction (robotic actions), holding great significance for AI content creation and embodied intelligence.

Episode 3 · BFL Launches FLUX 3: A Unified Multimodal Model for Video and Robotics (2026-07-24, 12 posts)

Black Forest Labs (BFL) has officially released FLUX 3, a unified multimodal foundation model that integrates image, video, native audio, and action prediction within a single architecture. FLUX 3 Video is currently in early access, and the team plans to roll out various capabilities progressively via API and private weights over the coming weeks and months. Demonstrating a profound understanding of the physical world, the model is viewed by the industry as a critical milestone toward general world simulators and robotic control.

Confirmed

FLUX 3 utilizes a unified architecture capable of simultaneously processing images, video, audio, and action prediction. The official release emphasizes that the new model's generation results across different styles are more closely aligned with real-world physics, and its native audio capability has received positive feedback in hands-on tests. Model capabilities will be gradually provided through APIs and private weights, with FLUX 3 Video currently in the Early Access phase for feedback collection and safety testing.

Unconfirmed

Regarding whether FLUX 3 will release a separate "Klein" model version, current official information suggests FLUX 3 is more of a collection of capabilities rather than a single standalone model. Furthermore, users point out that key metrics determining whether the model can truly land in a production environment—such as specific latency, concurrency limits, and pricing structures—remain undisclosed.

Why it matters

The core breakthrough of FLUX 3 lies in its action prediction capability. Several analysts (such as @MattVidPro and @rohanpaulai) note that unlike traditional VLA models relying solely on limited teleoperation data, FLUX 3 absorbs large-scale video pre-training data, thereby better grasping physical laws like contact, deformation, and temporal causality. Authors like @imjustnewatai further suggest that combining video models with robotic control validates the technological evolution route from video generation to world models, ultimately landing in embodied intelligence.

Episode 4 · Black Forest Labs Rumored to Release FLUX 3 (2026-07-26, 2 posts)

Rumors suggest Black Forest Labs is set to release FLUX 3, a new model that reportedly unifies image, video, audio, and action prediction within a single architecture.

Episode 5 · Black Forest Labs Launches FLUX 3 with Native Audio Video Generation (2026-07-28, 2 posts)

Black Forest Labs has launched FLUX 3 in limited access, a new multimodal model capable of generating images and up to 20 seconds of video with native audio output.

Episode 6 · Black Forest Labs Unveils FLUX 3 Multimodal Model with Native Audio Video Generation (2026-08-04, 19 posts)

Black Forest Labs officially released FLUX 3, a unified multimodal model trained on image, video, and audio. Its core highlight is the ability to generate up to 20-second HD videos with synchronized native audio in a single request, marking a significant step toward building a true 'world model'. The model is now in early access, with API and web interface available.

Confirmed

  • FLUX 3 supports text-to-video, image-to-video (with first/last frame or multi-keyframe control), video continuation, and multilingual dialogue.
  • Video generation offers 720p and 1080p resolutions, up to 20 seconds per clip, with native audio (including sound effects and ambient sound), plus a low-cost Draft mode.
  • The model uses a unified architecture trained jointly, enabling scene and camera angle changes within a single generation; reference-image-based variants are planned.
  • FLUX 3 is now publicly available via API and on Runware and Magnific, offering integrated generation and upscaling workflows.
  • The web interface is live; users can purchase credits to generate videos. Currently only video generation is available; other features like image generation will follow.
  • The model is available on Cloudflare AI Gateway, supporting text-to-video, image-to-video, and video-to-video continuation.
  • The company announced that model weights will be open-sourced soon; API access is already open to partners like Replicate, Krea, and Fal.

Unconfirmed

  • The exact timeline for open-sourcing weights has not been announced.
  • Details on model parameter size and training data have not been disclosed.

Why it matters

  • The company argues that a single modality only captures a projection of reality, while multimodal joint learning can better understand the underlying physics of the world through mutual constraints (e.g., sound matching impacts, motion following mass). FLUX 3's release is not just an improvement in video generation but a key attempt toward the technical vision of building a true 'world model'.

Episode 7 · Black Forest Labs' FLUX 3 Video Model Sparks Viral Sensation (2026-08-06, 3 posts)

Black Forest Labs' newly released FLUX 3 video model has gone viral for its stunning capabilities. It excels in prompt adherence and rendering, natively generating embedded audio and complex scenes, inspiring highly creative user-generated content.

Episode 8 · Flux 3 Launch Sparks Benchmarking Frenzy, Multi-Model Comparison Heats Up (2026-08-06, 10 posts)

Black Forest Labs released the open-source video generation model Flux 3, capable of generating 20-second clips with synchronized native audio, including title sequences and full action scenes. The release quickly sparked extensive community testing, with multiple creators benchmarking it against mainstream models like Seedance and MiniMax H3. Current findings show Flux 3 excels in camera movement, physics destruction effects, and retro aesthetics, but overall rankings remain contested, with competitors each having strengths in professional applications, detail handling, and overall performance.

Confirmed

  • Flux 3, developed and open-sourced by Black Forest Labs, supports generating 20-second clips with synchronized audio, including title sequences and full action scenes.
  • Seedance (ByteDance) versions range from 2.0 to 2.5, supporting up to 30-second generation and 50 reference inputs.
  • MiniMax H3 handles details, lighting, and temporal synchronization best.

Unconfirmed

  • The absolute ranking among models remains undecided. Author @flowersslop, after testing with a "T-Rex vs. elephant" prompt, believes Flux 3 outperforms most models but still trails the months-old Seedance 2, giving a ranking of Seedance 2.5 first, Flux 3 second.

Why it matters

  • This benchmarking directly showcases the core differences and intense competition among current mainstream video generation models. Author @eyishazyer notes Flux 3 excels in aerial camera movement, realistic physics destruction, complex transitions like door openings, and hand motion synchronization, with stability. In contrast, Seedance 2.5 lags slightly in motion fluidity, and Seedance 2.0's top configuration also falls short.
  • Author @heypearlai's multi-scenario tests validate Flux 3's coherence and detail handling in one-take room transitions and high-difficulty action scenes.
  • Author @gorkem, after comparison, states that while Seedance 2.5 and Flux 3 each have merits in raw video quality, Flux 3 feels superior in overall logic and intelligent experience.
  • Author @beechinour believes Seedance 2.5 remains leading for professional use, while Flux 3 wins in the new niche of retro imagery creation (e.g., fake memories, lost footage, VHS visuals).
  • Author @LudovicCreator's dynamic design test finds MiniMax H3 performs best.

Episode 9 · Together AI Launches FLUX 3 Serverless API (2026-08-07, 2 posts)

Together AI launched Black Forest Labs' FLUX 3 video model on its Serverless platform, offering production-ready deployment. The multimodal model supports multi-scene generation, lip-syncing, and native audio.

Episode 10 · Black Forest Labs Launches FLUX 3 Multimodal Model (2026-08-11, 2 posts)

Black Forest Labs has released FLUX 3, a new multimodal model that unifies image, video, audio generation, and robotic action prediction. It features built-in native audio generation and can produce high-definition video clips up to 20 seconds long.

Episode 11 · FLUX 3 Video Ranks Second Globally, Free Access for Limited Time (2026-08-12, 5 posts)

Black Forest Labs' FLUX 3 Video update ranks second in LMSYS Arena text-to-video with 1496 points, just 16 behind leader Gemini Omni Flash. To celebrate, the model is free on Playground until Sunday evening PT. The model supports native audio, text-to-video, and image-to-video (multi-frame).

Confirmed

  • Leaderboard: FLUX 3 Video update scored 1496 in LMSYS Arena text-to-video, ranking second globally, only 16 points behind the closed-source leader Gemini Omni Flash (1512).
  • Features: Supports native audio, text-to-video, and image-to-video (multi-frame).
  • Free access: Officially free on Playground until Sunday evening PT.

Why it matters

  • Open-source catching up: As a key BFL product, FLUX 3 Video's top ranking in arena benchmarks shows open-source (or semi-open) video models are closing the gap with leading closed-source models.

Episode 12 · Flux 3 Text-to-Video Tests Impress with Realism and Cinematography (2026-08-12, 2 posts)

Creators are highly impressed by initial tests of Black Forest Labs's Flux 3 text-to-video model. The model demonstrates exceptional realism, environmental details, emotional composition, and strong adherence to camera movement instructions.

Episode 13 · FLUX 3 Video Debuts at No.5 on Image-to-Video Arena (2026-08-13, 2 posts)

Black Forest Labs' new FLUX 3 Video model debuted at fifth place on the LMArena Image-to-Video leaderboard with a score of 1453. MiniMax currently holds the top position, highlighting the fierce competition in video generation.