Distributed training makes video data provenance hard to trace
burkov · x · 2026-07-25
A reply to Mira Murati notes that in highly distributed training pipelines, it becomes especially hard to know where the video data inside the training corpus actually came from.
The point is less about a specific model and more about a core AI governance problem: provenance tracking gets harder as data collection becomes more distributed.
More from AGI Musings
- Anthropic says Opus 5 still shows no RSI or dramatic AI acceleration — rickasaurus · 2026-07-25
- Human plus LLM is the new superhuman, says one AI commentator — tydsh · 2026-07-25
- Thread argues for a “pseudo-EMH” way of updating beliefs about AGI risk — EigenGender · 2026-07-25
- LatamGPT pretraining lead says AI users pay twice: service and dependence — OmarUFlorez · 2026-07-25
- Inference compute is spreading from math to every other field — burny_tech · 2026-07-25
- The hardest problem in AI is incentives, not intelligence or AGI — AryHHAry · 2026-07-25