FLUX 3 expands into one multimodal backbone for image, video, audio and action
pmttyji · reddit · 2026-07-24
FLUX 3 is introduced as a single multimodal backbone for image, video, audio, and action prediction.
The launch plan says the team will roll out capabilities over the coming weeks and months after an early-access phase for feedback and safety testing. Planned releases include:
- FLUX 3 Video for video and audio generation/editing via APIs and private weights
- FLUX-mimic and FLUX 3 Action for action prediction with research and commercial partners
- FLUX 3 Image for image synthesis and editing
- FLUX 3 Dev, an open-weight multimodal backbone for content creation and action prediction
The post also says more technical details about the underlying approach will be shared later.
Related event: BFL Launches FLUX 3: A Unified Multimodal Model for Video and Robotics(12 posts)→
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- The full prompt-to-3D-game workflow: Hyper3D Rodin MCP plus Codex, no reference image — FellMentKE · 2026-09-11
- Building a 3D landing page with GPT-6 Astra and Hyper3D Rodin MCP, no modeling needed — FellMentKE · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11