Do elite users' data matter for frontier models? A debate on RL datamix signals

plausaible · x · 2026-09-09

A technical debate on the value of user data for frontier models. kalomaze argues process improvements from user data overwhelmingly can't be causally attributed to individual power users — instead, aggregate signals like domain-wide power-user frustration redirect effort into that domain.

plausaible partially agrees, especially on the 2026 model datamix, but contends the views aren't mutually exclusive: cutting-edge knowledge genuinely comes from a few select people worth farming. A representative argument over elite-data vs aggregate-signal strategies in RL data collection.

Related event: Debate Over User Data Value in Frontier Model Training: Harvesting Failure Modes, Not Individual Expertise(5 posts)→

Original post →

More from Models

Models channel →