Black Forest Labs Launches FLUX 3: A Unified Multimodal Model for Image, Video, and Audio
pess_r · x · 2026-08-05
Black Forest Labs has officially launched FLUX 3, a new multimodal model designed to unify image, video, audio, and even action prediction for robotics.
Key Features:
- Video Generation: Supports text-to-video and image-to-video with multi-scene continuity. Generates clips up to 20 seconds and supports keyframe control.
- Native Audio: Generates multilingual speech, sound effects, and ambient audio synchronized with frames, featuring strong lip-sync.
- Image Generation: Offers highly accurate text rendering and handles complex prompts across a wide variety of styles.
- Draft Mode: Allows fast, low-cost previews of creative directions before rendering in full quality.
- Robotics: Unifies perception, simulation, and execution by predicting physical outcomes and robot control actions from visual and text inputs.
More from Multimodal
- Seedance 2.5 Launches Globally in CapCut, Integrating Video Generation and Editing — AIwithGhotai · 2026-08-07
- Chaining AI Models in fal Workflows to Generate a 15-Second Animated Short — gorkem · 2026-08-07
- Qwen-3D: Enhancing Spatial Reasoning via Multi-View Geometric Cues — udmrzn · 2026-08-07
- MiniMax H3 Test: Generates 15s T2V with Native Audio in 23 Minutes — AxonkaiLab · 2026-08-07
- Exploring Workflows for Using Claude to Assist Veo in Medical Animations — Hipposy · 2026-08-07
- Cheap Renders Can Create the Most Expensive AI Video Workflows — Div_pradeep · 2026-08-06