FLUX 3 adds image, video, audio, and action prediction in one model
GabGarrett · x · 2026-07-24
Black Forest Labs says FLUX 3 is a single multimodal model for image, video, audio, and action prediction. The company says FLUX 3 Video is available in early access, and that the model is jointly trained in one unified architecture.
The release also points to a robotics angle: the model can be extended to predict actions, and the thread references work with Mimic and Audi. The post frames this as an open-weight, state-of-the-art video model release.
Related event: Black Forest Labs Launches FLUX 3: A Unified Multimodal Foundation Model(28 posts)→
More from Embodied
- Flux 3 appears to be a capability family, not a separate "Klein" model — carrot_2333 · 2026-07-24
- OpenAI’s Codex keyboard ships, and early users say it’s pricey but fun — APPSO · 2026-07-24
- An $8 ESP32-S3 now runs a 28.9M-parameter model fully offline — brianrkelly · 2026-07-24
- TIME Features Unitree's GD01, the World's First Mass-Produced Transformable Mecha Robot — whurley · 2026-07-24
- OpenBMB opens up MiniCPM-Robot as its first embodied AI model family at WAIC — CyberRobooo · 2026-07-24
- Robot skins made by industrial knitting could enable scalable tactile sensing — mynkgoel · 2026-07-24