Black Forest Labs Teases FLUX 3: One Model for Image, Video, Audio, and Robotics
bennash · x · 2026-08-13
A developer discovered a teaser page for FLUX 3 on Black Forest Labs' website, indicating an upgrade from an image-only model to a comprehensive multimodal model.
Key features include:
- Multimodal Generation: A single model handling image, video, audio, and action-prediction.
- Video & Audio: Supports native audio and up to 20-second video clips in a single generation, expanding beyond cinematic styles.
- Robotics: Unifies perception, simulation, and execution by taking visual observations and text instructions to predict physical outcomes and robot control actions.
- Deployment: Available via API, open weights (for self-hosting and fine-tuning), and enterprise solutions.
More from Embodied
- IHMC's HexRunner hits 30+ mph using spring-loaded legs instead of wheels — TinfoilTricorn · 2026-08-13
- Workers in India reportedly paid to film manual labor for robot training — Polymarket · 2026-08-13
- Dev Integrates Vision for Local Agent Training, Explores Water Wave Computing — cephaloform · 2026-08-13
- AtlasVLA: Vision-Language-Action Model with Persistent Memory — CASIA-IVA-Lab · 2026-08-13
- Agility Robotics exec: Backflips easy, picking up a pen is harder — jonstephens85 · 2026-08-13
- Pi Expands Robotics Ecosystem: Hiring Generalist for Hardware & Community — minsuk_chang · 2026-08-13