Rumor says Flux 3 unifies image, video, audio, and action prediction
theteknosaur · x · 2026-07-26
Black Forest Labs is rumored to be training image, video, audio, and action in one model
The post claims that Flux 3 dropped this week and that Black Forest Labs trained a single model across image, video, audio, and action prediction.
It also says the action variant, Flux Mimic, is already in trials on Audi manufacturing lines, with the broader takeaway that video prediction and robot control may be the same problem.
Related event: Black Forest Labs Rumored to Release FLUX 3(2 posts)→
More from Embodied
- Chinese robotics team trains humanoids with kicks, pushes, and knockdowns — 2C_ornot2C · 2026-07-26
- D-Propulse previews an air-breathing rotating detonation engine with aerospike nozzle — Paimaamu · 2026-07-26
- A reMarkable Paper Pro demo turns an e-ink tablet into an AI notebook — Olivier__OG · 2026-07-26
- T-Rex lifts dexterous hand success to 65% by giving touch its own fast path — 量子位 · 2026-07-26
- Chinese robot dog switches between wheels and legs autonomously — RoboBalaji · 2026-07-26
- Ant Group releases 2,000 hours of egocentric manipulation data for embodied learning — burny_tech · 2026-07-26