LAION Releases 10M Hour Open Video Dataset BVD with 80M Videos
JJitsev · x · 2026-08-27
LAION has released LAION-BVD (Big Video Dataset), a large-scale open video dataset for multimodal learning. The project collected 1.3 billion platform-specific video URLs from CommonCrawl and downloaded 80 million videos totaling 10 million hours. The dataset includes content-aware scene detection, synthetically generated video and audio captions, and extracted scene-changing frames as image-text data. Models trained on this data achieve competitive performance on standard video-text and audio-text benchmarks.
Related event: LAION Releases 10M-Hour Open Video Dataset BVD(4 posts)→
More from Research
- Paper: CoT Monitorability as a Fragile Safety Opportunity — idavidrein · 2026-08-27
- Modern LLMs Compress English Text to Under 1 Bit Per Character — docmilanfar · 2026-08-27
- LeakyLMs: Stealing Architecture and Inference Optimizations via Timing — niloofar_mire · 2026-08-27
- DeepSeek's 1/30 Cost Training Breakdown — wordgrammer · 2026-08-27
- UCLA Introduces LongMemEval-V2: Benchmarking Agent Environmental Memory — dair_ai · 2026-08-27
- Criteria for Simple Yet Relevant Models of Computation — yaroslavvb · 2026-08-27