Yoav Goldberg asks if Google skips billions of RL trajectories unlike OpenAI

yoavgo · x · 2026-10-04

In a discussion with Google researchers Roee Aharoni, Zorik Gekhman and Jon Herzig, Yoav Goldberg asks whether Google does not treat RL as a "knowledge acquisition" stage, and thus doesn't train billions of RL trajectories the way OpenAI does. He argues RL can still teach agents facts, citing tasks like winning games or generating reference sets judged by replicating a survey bibliography.

Related event: Yoav Goldberg on Whether RL Can Teach Models Factual Knowledge(2 posts)→

Original post →

More from Research

Research channel →