AI firms are buying old books to secure cleaner training data
marigo · x · 2026-07-23
AI companies are reportedly buying old books to secure clean training data, avoiding the noise and duplication common in modern web corpora. The post points to a growing premium on higher-quality, less contaminated text sources as training data gets harder to find.
Related event: AI Companies Buy Old Books to Avoid Synthetic Data Pollution(2 posts)→
More from Infra
- Developer says Sol burns through usage in hours, asks how teams afford $10K-$20K/month — Dan_Jeffries1 · 2026-07-23
- An agent hit GitHub Actions limits, spun up a VM, and kept the build running — davidcrawshaw · 2026-07-23
- Apple’s Mac roadmap points to a new split: devices for people and hosts for agents — APPSO · 2026-07-23
- Progress to buy Domo’s AI and data platform for $400M in cash — shashib · 2026-07-23
- Open-source coding agent octomind chooses persistent cloud machines over per-session sandboxes — donk8r · 2026-07-23
- Farmers say data centers are killing bees, raising a food-system alarm — eyishazyer · 2026-07-23