Black Forest Labs Launches FLUX 3: A Unified Multimodal Model for Image, Video, and Audio
dl_weekly · x · 2026-08-04
Black Forest Labs announced that its latest multimodal foundation model, FLUX 3, is now available in Early Access. The model uses a unified architecture to jointly train on images, videos, and audio.
The company stated that this joint learning approach allows the model to better understand the underlying rules of the physical world, such as the causal relationship between object motion and sound. Currently, FLUX 3 can generate 20-second video clips with native audio and shows potential to expand into physical AI applications like robotic action prediction.
More from Models
- Liquid AI Launches LFM2.5-2.6B: On-Device Agentic Model Outperforming Larger Counterparts — maximelabonne · 2026-08-04
- Hugging Face Releases LFM2.5-2.6B for Local Agent Deployment Everywhere — Hugging Face Blog · 2026-08-04
- Frontier closed models overengine small coding tasks, sparking an engineering crisis — robleclerc · 2026-08-04
- Testing DeepSeek and Qwen to Capture Rain World Game Essence — yacineMTB · 2026-08-04
- Are Frontier Models Overkill for Simple Tasks? The Case for Multi-Model Routing — zerogpu_ai · 2026-08-04
- NVIDIA Releases VoiceChat-11B Model on Hugging Face — nvidia · 2026-08-04