FLUX 3 expands into one multimodal backbone for image, video, audio and action
pmttyji · reddit · 2026-07-24
FLUX 3 is introduced as a single multimodal backbone for image, video, audio, and action prediction.
The launch plan says the team will roll out capabilities over the coming weeks and months after an early-access phase for feedback and safety testing. Planned releases include:
- FLUX 3 Video for video and audio generation/editing via APIs and private weights
- FLUX-mimic and FLUX 3 Action for action prediction with research and commercial partners
- FLUX 3 Image for image synthesis and editing
- FLUX 3 Dev, an open-weight multimodal backbone for content creation and action prediction
The post also says more technical details about the underlying approach will be shared later.
Related event: Black Forest Labs Releases FLUX 3 Unified Multimodal Model(4 posts)→
More from Multimodal
- Scribbly AI art style shown across multiple generated image examples — OVolosin82152 · 2026-07-24
- Flux 3 X Mimic debuts as a next-generation video-action model — kensai · 2026-07-24
- Flux 3 is being tested on perfectly synced split-screen video generation — umesh_ai · 2026-07-24
- Flux 3 appears to be a capability family, not a separate "Klein" model — carrot_2333 · 2026-07-24
- A user asks why 4x UltraSharp upscaling is producing strange image artifacts — MeDotEE · 2026-07-24
- Midjourney SREF tip turns a simple turtle walk into a caricature punchline — tisch_eins · 2026-07-24