Researchers debate model-in-the-loop data generation: one-shot massive experiments a bad tradeoff
owl_posting · x · 2026-09-21
- A debate on how scientific data should be generated in the AI era: one side argues massive data-collection moonshots serve as north-star goals that force building an otherwise useful technology stack, even if they can't be done today.
- Stanford researcher Anshul Kundaje disagrees: data generation must be iterative with models in the loop suggesting the next experiments; one-time massive data generation is "almost surely always a bad tradeoff."
- He notes projects like ENCODE predate current generative AI — small model-in-the-loop tweaks across phases would have substantially improved bang for buck — though this is harder in distributed consortia without centralized control than in hierarchical organizations.
Related event: Scholars debate how science should generate data in the AI era(3 posts)→
More from AGI Musings
- Founder feedback now goes to Claude and GPT, fueling the solo founder boom — brezina · 2026-09-21
- The 'moral singularity' argument: unknowable ASI values rule out 50-50 optimism — danfaggella · 2026-09-21
- ToKTeacher: AI Speeds Error Correction, So We Must Solve Problems Much Faster — burny_tech · 2026-09-21
- Nate So8res breaks 'How would AI kill us?' into three distinct questions — ESYudkowsky · 2026-09-21
- Gary Marcus clashes with Gavin Baker: OpenAI won't be a monopoly because everyone shares the same recipe — GaryMarcus · 2026-09-21
- Jeff Bezos: AI Won't Kill Jobs, It Will Cause a Labor Shortage — rohanpaul_ai · 2026-09-21