Black Forest Labs launches FLUX 3, a multimodal model for image, video and audio
3scorciav · x · 2026-07-24
- Black Forest Labs introduced FLUX 3, a single multimodal model for image, video, audio, and action prediction.
- The company says FLUX 3 Video is already in early access.
- FLUX 3 is trained in one unified architecture, and the team says the same model family can be extended toward robotics action prediction.
- The post also points to work with Mimic and Audi, framing FLUX 3 as more than a media generator and closer to a general multimodal foundation.
Related event: Black Forest Labs Launches Unified Multimodal FLUX 3(8 posts)→
More from Embodied
- Robot builder says Bob died after a battery mod fried its Raspberry Pi brain — chrismatthieu · 2026-07-24
- MiniCPM-RobotTrack runs on-device at 5+ Hz and keeps tracking unplugged — iamfakhrealam · 2026-07-24
- OpenBMB Open-Sources MiniCPM-Robot: Embodied AI That Keeps Tracking Offline — iamfakhrealam · 2026-07-24
- Autonomous Labs Employee Teases Unreleased AI Hardware Prototype — dee_hw · 2026-07-24
- Brazilian surgeon operates on a patient 12,034 km away with 199 ms latency — aakashgupta · 2026-07-24
- Generalist AI shows a robot tool interface that makes grip force visible — DominiqueCAPaul · 2026-07-24