Black Forest Labs Launches FLUX 3: Unifying Image, Video, Audio, and Robotics
bennash · x · 2026-08-12
Black Forest Labs has officially announced FLUX 3, a new multimodal model that unifies image, video, audio generation, and robotics action prediction into a single architecture.
Key Features:
- Video: Generates native clips up to 20 seconds from text, images, or keyframes, supporting multiple shots in one take.
- Audio: Natively generates multilingual speech, sound effects, and ambience alongside the frames.
- Image: Offers highly accurate text rendering and complex prompt handling across diverse styles.
- Robotics: Takes visual observations and text instructions to predict physical outcomes and output robot control actions.
Availability: The model is available via API, open weights (for custom fine-tuning), and enterprise solutions. The web playground currently offers free video generation trials, allowing users to use storyboard images as starting frames for reference.
Related event: Black Forest Labs Launches FLUX 3 Multimodal Model(2 posts)→
More from Embodied
- Matic Robot Vacuum Adopts Tesla FSD Approach: Pure Vision Over LiDAR — MatthewBerman · 2026-08-14
- NVIDIA's SONIC Robotics Research Published in Science for Humanoid Control — zhengyiluo · 2026-08-14
- LiDAR-Free Vision: Vacuum Robot with 5 Cameras Onboard — thisguyknowsai · 2026-08-14
- Bittensor SN80 Builds Decentralized Robot Data Pool with 130K+ Contributors — bittingthembits · 2026-08-14
- Matic Robot Update: Adds On-Device Point-and-Speak Control in 75 Languages — testingcatalog · 2026-08-14
- XPeng's VLA 2.0 Mileage Surpasses Human Driving, Unifies Spatial AI — kimmonismus · 2026-08-14