Ex-OpenAI Employee Starts Company for High-Quality RL Datasets, Citing Poor LLM Generalization

burny_tech · x · 2026-07-31

A former OpenAI employee announced their departure to start a new company focused on producing high-quality reinforcement learning (RL) datasets.

They highlighted that the generalization ability of current LLMs remains very poor. Even in heavily invested domains like coding, capabilities are 'spiky' and lack true generality, often requiring manual intervention to clean up outputs. Because current architectures struggle with out-of-distribution (OOD) generalization and rely on costly post-training, the business of producing training data, evals, and learning environments is poised to thrive.

Related event: Ex-OpenAI Staffer Leaves to Tackle LLM Generalization with RL Data(4 posts)→

Original post →

More from Companies & People

Companies & People channel →