Ex-OpenAI Staffer Leaves to Build High-Quality RL Data Startup, Cites Poor LLM Generalization

panickssery · x · 2026-07-30

Andrew Ho announced his departure from OpenAI to start a company focused on producing high-quality reinforcement learning (RL) datasets. He pointed out that current LLMs have very poor generalization with "spiky" capabilities, lacking generality even in heavily invested areas like coding. He believes existing data offerings fail to cover most economically productive capabilities, presenting a massive opportunity in designing and producing top-tier RL datasets.

Original post →

More from Companies & People

Companies & People channel →