Ben Lorica: The AI data problem moved downstream, from finding data to making it usable
bigdata · x · 2026-10-02
Ben Lorica argues AI's data bottleneck has shifted from raw material to making information usable. Robotics teams now train on massive human video (1M+ hours at one company; 1,900 hours converted into 18,000+ hours of robot-format data), but all need substantial machinery to transform human experience. On the web, access is getting conditional: a publisher granted AI retrieval rights while withholding training, permissions are splitting into training/retrieval/inference tiers, and crawlers can even receive paid HTTP 402 responses. The work of feeding AI is moving downstream into processing and access negotiation.
More from Infra
- Redditor builds fully local LLM-powered radio site on two DGX Sparks and a 5090 — jwhh91 · 2026-10-02
- VC quip: many neoclouds are closer to 95% than five nines of reliability — saranormous · 2026-10-02
- Report: lenders demand up to 25% collateral from Nvidia as GPU-backed loans wobble — GaryMarcus · 2026-10-02
- Microsoft Backs Snowflake-Led Effort to Standardize Business Metrics for AI — xiaosun86 · 2026-10-02
- Full SGLang config for GLM-5.3-Flash NVFP4 on 4x RTX 6000 Max-Q — TheZachMueller · 2026-10-02
- GLM-5.3-Flash NVFP4 benchmarks show no per-user speedup beyond 8 concurrent requests — TheZachMueller · 2026-10-02