Black Forest Labs’ Flux 3 is said to be a multimodal image, video, audio model
ZeroStateReflex · x · 2026-07-23
Black Forest Labs’ next model appears to be a multimodal Flux 3 system, not just a video generator.
- A briefly live placeholder page described it as “a breakthrough in control, realism, and world understanding” and said it would generate image, video, audio, and action.
- Early access users are already posting generations, but were apparently asked not to name the model.
- The standout rumor is 20-second video generation, which would put it ahead of much of the industry if true.
Related event: Black Forest Labs Teases Flux 3 Multimodal Model(3 posts)→
More from Models
- SolarOpen2 weights are now公开 and can be freely fine-tuned — algo_diver · 2026-07-23
- A post says Opus 5 is launching today and expects a major reset — dejavucoder · 2026-07-23
- Kimi K3 may have replaced Grok as one team’s code reviewer — zeeg · 2026-07-23
- China’s high-quality data annotation only began scaling in the last six months, says Teortaxes — zephyr_z9 · 2026-07-23
- MolmoAct2 runs text-to-action on a $150 robot arm and RTX 4060 GPU — DJiafei · 2026-07-23
- Per-turn model routing can cut coding-agent costs without changing the agent — entelligenceai17 · 2026-07-23