FLUX 3 will add native audio, image editing, and open-weight multimodal access
bfl_ai · x · 2026-07-23
Black Forest Labs says FLUX 3 will roll out several capabilities over the coming weeks and months after early-access testing:
- Native audio generation for video is already in early access.
- Action prediction will open through selected research and commercial partners, starting with mimic robotics.
- Image generation and editing are coming next.
- The company also plans fast variants for lower-cost iteration.
- A multimodal open-weight backbone will be released for content creation and action prediction.
The company describes FLUX 3 as a step toward “real-world visual intelligence” that can perceive, predict, and act across digital and physical environments.
Related event: Black Forest Labs Launches FLUX 3 Omni-Modal Model(14 posts)→
More from Multimodal
- Black Forest Labs announces FLUX 3, a multimodal model for image, video, and audio — chrisfirst · 2026-07-23
- Self-Flow speeds multimodal model convergence by up to 2.8×, paper says — hila_chefer · 2026-07-23
- Viewer gifts now trigger real-time AI video effects in a live-streaming demo — ming_calligraphy · 2026-07-23
- GLM-5.2 vision model baseten/GLM-5.2-Vision-NVFP4 trends on Hugging Face — baseten · 2026-07-23
- Controlled study finds training data quality is decisive for text-to-video models — Amber Yijia Zheng · 2026-07-23
- FLUX 3 is announced, but its capabilities will roll out over weeks and months — Angaisb_ · 2026-07-23