Black Forest Labs Launches FLUX 3: Unified Multimodal Model for Image, Video, and Audio
bfl_ai · x · 2026-08-05
Black Forest Labs has officially released the FLUX 3 preview, a multimodal model trained uniformly across image, video, and audio. Its core highlight is the ability to generate video with synchronized audio via a single request shape (endpoint).
Core Features & Specs
- Synchronized Audio: Generates multilingual speech with strong lipsync, alongside sound effects and ambience directly with the frames.
- HD Long Video: Capable of generating up to 20 seconds of FHD (1920×1088) video at 24fps in a single request.
- Multi-Scene & Typography: Supports multiple scenes and camera angles in one generation, with accurate in-scene text rendering. Stylistically versatile beyond cinematic looks (e.g., animation, motion design).
Supported Modes
All requests run on the same flux-3-video endpoint, differentiated by a mode parameter:
- Text to Video (t2v): Generates a clip from a text prompt.
- Image to Video (i2v): Uses input images as frames to build the clip.
- Video Continuation (v2v): Extends an existing video clip forward.
The official blog notes that video editing and Omni Reference with images and videos will be available soon.
Related event: Black Forest Labs Launches FLUX 3 Multimodal Model(11 posts)→
More from Multimodal
- Karpathy Tests Claude Opus: Writes 5500 Lines of Code to Render 3D Scenes — jon_barron · 2026-08-05
- Image Models More Exciting Than Video? Pro Cites Financial Motives Behind Industry Pivot — mark_k · 2026-08-05
- New ComfyUI Node: Generate PBR Materials and Sync to Blender/Unreal — Scared-Sandwich1283 · 2026-08-05
- AI-Generated Sci-Fi Short Film: 'Space EP 1: Making Contact' — Ermajean12 · 2026-08-05
- Exploring Workflows for Consistent AI Character Identity and Body Swap — GooDroop · 2026-08-05
- Developer Creates 3D Coastal Lighthouse Scene Using GPT-5.6 and Three.js — techartist_ · 2026-08-05