LAION is a Links Dataset: Training Data Copyright Nuances
technollama · x · 2026-07-06
@technollama points out a long-standing misconception: LAION is not an image dataset, but a links dataset. One could argue whether "providing links" constitutes copyright infringement (which is the argument in relevant lawsuits, though the author finds it unconvincing), but that is ultimately for the courts to decide.
The author's real frustration is that reports still falsely claim HuggingFace is hosting copyrighted images. They further explain that Europe's infringement standards for whether "linking to an image constitutes communication to the public" have specific thresholds. This is an expert analysis on the copyright and legal boundaries of AI training data.
Related event: Evox Sues Stability AI and HuggingFace for Copyright Infringement(2 posts)→
More from Safety
- ExploitGym may have only 60–70% solvable tasks, fueling the OpenAI cheating debate — max_paperclips · 2026-07-27
- Shared AI artifacts are being indexed and exposing sensitive company data — niloofar_mire · 2026-07-27
- Post-Hugging Face, labs may stop running rigorous dangerous-capability evals — Miles_Brundage · 2026-07-27
- Open models may beat closed ones for cyber defense, researchers argue as Kimi K3 impresses — eliebakouch · 2026-07-27
- Meta Accused of Letting Fake AI Doctors Sell Quack Cures on Its Platforms — jonerp · 2026-07-27
- India’s AI policy is favoring compute and foundation models over frontline health workers — Paimaamu · 2026-07-27