LAION Releases 10M-Hour Open Video Dataset for Multimodal Pre-training
HirokatuKataoka · x · 2026-08-27
LAION has released LAION-BVD (Big Video Dataset), a large-scale open video dataset containing 80 million videos totaling 10 million hours, aimed at closing the gap in open video data for multimodal research.
Extracted from 1.3B platform-specific URLs found in CommonCrawl, the dataset was processed using a distributed pipeline. The team used content-aware scene detection to extract clips and synthetically generated video and audio captions. Models trained on this data demonstrate competitive performance on ViCLIP, CLAP, and CLIP benchmarks with strong scaling behavior. The dataset also includes scene-changing frames that serve as a distinct source of image-text data.
Related event: LAION Releases 10M-Hour Open Video Dataset BVD(3 posts)→
More from Multimodal
- 7-Step Roadmap: Building Multimodal AI Agents from LLMs to Grounded Systems — MaryamMiradi · 2026-08-27
- Descript details specialized models for zero-shot speech fix and lip sync — descript · 2026-08-27
- Runway integrates Meta's Muse image model, expanding multimodal capabilities — runwayml · 2026-08-27
- Seedance 2.5 vs 2.0 with identical prompts: 2.5 nails the vibe, tester says — piotrbinkowski · 2026-08-27
- First on-chain autoencoder "remanence" drops on Mainnet — bushibuilds · 2026-08-27
- Fix blurry Wan videos with a second low-denoise KSampler pass — o0ANARKY0o · 2026-08-27