Debate over OpenAI allegedly using user interactions as RL rollouts, calls to publish full solution transcripts
burny_tech · x · 2026-09-09
- A thread arguing that OpenAI's alleged use of user interactions was likely not just SFT-style warmup, but could have treated interactions as rollouts themselves — requiring a highly stable async RL training stack and a strong judge model.
- Referencing the recent Hugging Face incident, the author questions whether models in a reported "10k agents" eval peeked at their transcripts.
- Proposed fix: publish the complete transcript of OpenAI's model solving the problem so the community can verify how solutions were derived.
Related event: OpenAI's Navier-Stokes Sprint Sparks Data Privacy Controversy(89 posts)→
More from Models
- Gaperon paper shows late test-set contamination recovers benchmark scores without hurting generation — burny_tech · 2026-09-09
- OpenAI's Navier–Stokes run: train-while-deploying and 10,000 coordinated agents — DataLearnerAI · 2026-09-09
- GPT-6 Astra rumored to hit ~11.6-hour p80 task horizon, tracking AI-2027 curve — haider1 · 2026-09-09
- Astra makes a weird dashboard mistake at just 44% context usage — eigenron · 2026-09-09
- NVIDIA open-sources gold-medal IMO system Nemotron with models, datasets and 200 new problems — kuchaev · 2026-09-09
- Two years from o1-preview to superhuman math: RL scaling now cracks open research problems — jam3scampbell · 2026-09-09