LAION-BVD opens 10 million video hours and beats InternVid on ViCLIP averages

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge

cs.CV, cs.AI, cs.LG

2026-08-26

LAION-BVD collects 1.3B platform video URLs and downloads 80M videos totaling 10 million hours. ViCLIP trained on 55M synthetically captioned clips averages 4.0 points above filtered InternVid-10M at matched scale.

What problem this solves

Public image-text corpora already sit at billions of pairs. Open video is an order of magnitude smaller. InternVid has about 7 million videos and 760k hours, nowhere near LAION-5B. The bottleneck is processing: platform gates, transcoding, scene cuts, and captions. Without a larger public video dump, open video-text and audio-text pretraining keeps recycling the same YouTube slices.

Method

CommonCrawl WAT dumps through March 2024 yield 4.7B candidate URLs, filtered to 1.3B links on YouTube, Vimeo, and Dailymotion. About 2,000 VMs behind a residential proxy attempt 130M downloads at 60% success, producing 80M videos and 10 million hours. No extra safety filter is applied at collection time.

The training subsets are smaller. From 2.4M sampled videos, clips shorter than 10 seconds or longer than 30 minutes are dropped, PySceneDetect (threshold 30) cuts scenes, and near-static segments are removed, leaving 55M clips. Video captions come from Qwen3-VL-2B-Instruct with up to 32 frames and a 20-word prompt. Audio captions come from Audio Flamingo 3 with a 10-word prompt. Scene-change frames from the full pool are filtered until about 300M images (BVD-I-300M) and recaptioned with DeepSeek-VL2-tiny.

The release is URLs, some captions, and raw files for institutions that accept the terms. YouTube is 94% of the pool, English about 57%, mean duration 7.7 minutes. CLIP-embedding FID between BVD frames and Re-LAION is 33.92, versus 0.16 between two Re-LAION draws, so the visual distribution is not just another web-image scrape.

Results

ViCLIP L-14 under one eval pipeline:

DataSamples seenK400Avg
DataComp-1B (frame-mean video)-61.253.2
InternVid-10M-FLT10M61.858.0
BVD-V-10M10M62.761.3
BVD-V-50M50M63.362.0

At 50M samples seen, BVD-V-10M and BVD-V-50M sit 3.3 and 4.0 points above InternVid-10M-FLT on the overall average. Scaling the encoder from B/32 to L/14 and samples from 10M to 50M raises the average monotonically. WiSE-FT helps a bit more.

For audio, pure BVD-A-10M CLAP at 30M samples seen reaches 46.8 average on the 431M model, above LAION-Audio at 44.8 and below LAION-Audio+AudioSet at 56.6. After 110M samples the same model reaches 48.7, with UrbanSound8K at 70.6. Mixed with AudioCaps+Clotho, BVD-A-1.7M stays close to LAION-Audio and a bit behind.

CLIP on frames is the other face. BVD wins COCO retrieval and loses ImageNet. ViT-B/16 at 300M samples: COCO I2T R@5 is 0.80 versus DataComp 0.70; ImageNet-1k is 0.28 versus 0.58. Video keyframes are not object-centric ImageNet photos.

Why it matters

This is one of the largest public web-video collections that outside labs can actually point to. Open video-text or audio-text pretraining now has a second InternVid-scale recipe, with the caption pipeline written down.

Do not treat the frames as a LAION-5B replacement. They complement web images: strong retrieval, weak ImageNet, useful for diversity, poor as a solo classification corpus.

Limitations

Captioned training data is 55M clips from 2.4M videos, not the full 10 million hours. Captions are short and model-written. Evaluation is contrastive only, with no generative video model. Audio and vision are trained apart. Downloads depend on residential proxies and decaying URLs. There is no collection-time safety filter. English and YouTube still dominate. Raw video is gated to institutions.

Terms

Source

Related papers

All paper explainers