LAION has released the Big Video Dataset (BVD), creating one of the largest freely accessible video databases for AI research. The scale is impressive: 80 million videos totaling ten million hours of footage, from which 55 million clips with automatically generated descriptions have been extracted. Add 300 million individual frames to that. The dataset is made available to researchers worldwide and enables them to train video-language models.
Quick Facts
- 80 million videos sourced from 1.3 billion URLs (primarily YouTube, English-language)
- Performance gain: Models trained on BVD outperform the previous reference dataset InternVid by up to 2.1 percentage points
- Legal basis: 2024 Hamburg court ruling permits LAION to collect copyrighted material for non-commercial research
- Access: Free for research purposes; commercial use excluded
How the Training Works
The BVD's strength lies in its multimodal approach: models learn simultaneously from video, audio, and text. The system understands not just which visual content matches which descriptions, but also which sounds and music accompany them. This combination drives the measured performance improvements on standard video-text benchmarks.
Legal Gray Area
LAION operates in sensitive territory here. The organization relies on a 2024 Hamburg Regional Court ruling that permits collecting copyrighted material for non-commercial research. This is an important precedent for open-source AI in Germany and Europe – though also contested. LAION explicitly asks users to "respect the rights and copyrights of content creators." Whether this legal position will hold up in higher courts or when commercial applications emerge remains uncertain.
What This Means for European Research
For German and European AI labs, BVD is a strategic asset. Until now, large video datasets for training video-language models were hard to access – researchers not working at OpenAI, Google, or other US giants relied on smaller or licensed sources. With BVD, universities, research institutes, and European startups can now work at the frontier of video AI models without depending on proprietary APIs. This matters especially for applications in medicine, Industry 4.0, or accessibility, where European data protection standards and local requirements count.
However: the dataset is primarily English-language and YouTube-centric. Those wanting to train on German or European content will need additional sources. The question of how long the legal foundation holds if commercial applications emerge also remains relevant for long-term research planning.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




