Ben Lorica: the AI data problem has moved downstream — from raw material to usable-data labor and licensed access
bigdata · x · 2026-09-29
Ben Lorica argues in Gradient Flow that the AI data problem has shifted from finding raw material to making information usable.
- Robotics data: one company trained on over 1 million hours of human video; another turned 1,900 hours of first-person footage into 18,000+ hours of robot-format training data. The point isn't that robotics found its dataset — every system needs substantial machinery to convert human experience into something usable.
- Web access becoming rule-based: an academic publisher opened selected books for AI retrieval while explicitly withholding training rights — "usable for AI" is splitting into separate permissions for training, retrieval, and inference. Technically, site owners can increasingly distinguish search, training, and agent crawlers; one closed-beta system even lets a crawler receive an HTTP 402 with a price.
Core thesis: data access is becoming a rules-and-pricing market, and whoever does the "making-data-usable" work owns the new value layer.
More from AGI Musings
- a16z podcast: personal AI agents are exploding, and the best ones will be invisible — a16z · 2026-09-29
- Findings of ACL already separates papers from talks, so 'AI-authored' rules are moot, scholar argues — ipeirotis · 2026-09-29
- Humanities will rule the future of higher education, scholar argues in AI era — begusgasper · 2026-09-29
- Education writer suggests keeping young kids away from AI for now — benjaminjriley · 2026-09-29
- US survey finds public AI concerns span many categories, likely worse after recent news — SpencrGreenberg · 2026-09-29
- Instruction tuning shows intelligence can exist without purpose or free will — sytelus · 2026-09-29