If scraping the world's knowledge for training is fine, why not distillation?

rao2z · x · 2026-09-23

The author poses a pointed consistency challenge: if you believe it's legitimate to train your LLM on the entirety of human-created knowledge—including scarfing up information behind paywalls—why doesn't the same logic apply to distilling other models, especially when those same companies argue chain-of-thought outputs are meaningful knowledge?

The post highlights a perceived double standard in how leading AI labs treat training data acquisition versus model distillation.

Original post →

More from AGI Musings

AGI Musings channel →