Black Forest Labs appears to tease FLUX 3, a multimodal model with 20-second video generation
op7418 · x · 2026-07-23
Black Forest Labs appears to be teasing FLUX 3, a full multimodal generation model that would cover images, video, audio, and motion. The leak suggests it may generate single videos up to 20 seconds long, and the poster notes that if it follows the company’s usual pattern, the release could be open source.
Related event: Black Forest Labs Teases Omni-modal Flux 3(4 posts)→
More from Multimodal
- Reddit user says LTX 2.3 audio beats TTS at breaths, laughs and intonation shifts — cptrios · 2026-07-23
- Open-source node-based LoRA trainer puts captioning, checkpoints and VRAM stats in one graph — ashishsanu · 2026-07-23
- TERRA-129 debuts as an AI-animated sci-fi episode credited to Matygoo — Matygoo1 · 2026-07-23
- Alibaba launches Qwen-Audio-3.0-TTS with 16 languages and 3-minute one-pass audio — Alibaba_Qwen · 2026-07-23
- Kling AI is said to handle close-up facial expressions better — burny_tech · 2026-07-23
- FameGrid Krea 2 aims to generate more realistic social-media-style images — UltraMuseArt · 2026-07-23