IFP Essay: An ARPANET-Style Program Could Unlock a Million Times More Data for AI
iamtrask · x · 2026-09-14
Andrew Trask and Lacey Strahm's IFP essay argues "peak data" is a misreading: disclosed model training sets are only a few hundred terabytes, while the world has digitized an estimated 180–200 zettabytes — over a million times more. The real crisis is misaligned incentives between data owners and AI companies. Their solution combines model partitioning and privacy infrastructure so data can train models without leaving its owner, plus policy recommendations for an ARPANET-style national program to unlock data access at scale.
Related event: OpenMined Proposes Network Sourcing for AI to Access Private Data(3 posts)→
More from Safety
- Researcher calls arXiv 'completely worthless' over lifetime bans for hallucinated AI citations — RexDouglass · 2026-09-14
- Researcher who trained frontier LLMs and built viruses: AI bioweapon fears are bogus — kuza55 · 2026-09-14
- Neel Nanda defends METR: funding sources are public, no AI lab money taken — NeelNanda5 · 2026-09-14
- Emad Mostaque rebuts Dario Amodei's frontier-pacing proposal: 'Intelligence isn't a crime' — QuixiAI · 2026-09-14
- Matt Yglesias: no law on the books stops recursive self-improving AI, and you can't sue a superintelligence — deanwball · 2026-09-14
- How to deliberately bait copyrighted songs and visuals out of video models — nptacek · 2026-09-14