FLUX 3 adds image, video, audio, and action prediction in one model
robrombach · x · 2026-07-24
A post circulating the FLUX 3 launch says the model combines image, video, audio, and action prediction in one system. The quoted tester claims it has the best native audio of any video generator they’ve used, can generate videos up to 20 seconds, and supports up to 10 references across image, video, and audio inputs, which can be mixed and matched.
Related event: BFL Launches FLUX 3: A Unified Multimodal Model for Video and Robotics(12 posts)→
More from Multimodal
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11