Yoav Goldberg asks if Google skips billions of RL trajectories unlike OpenAI
yoavgo · x · 2026-10-04
In a discussion with Google researchers Roee Aharoni, Zorik Gekhman and Jon Herzig, Yoav Goldberg asks whether Google does not treat RL as a "knowledge acquisition" stage, and thus doesn't train billions of RL trajectories the way OpenAI does. He argues RL can still teach agents facts, citing tasks like winning games or generating reference sets judged by replicating a survey bibliography.
Related event: Yoav Goldberg on Whether RL Can Teach Models Factual Knowledge(2 posts)→
More from Research
- Humans are vastly more sample-efficient than LLMs, and finding why could rival attention — burny_tech · 2026-10-04
- Tsinghua and Peking University Probe Whether LLMs Think Beyond Language in Nature MI — jiqizhixin · 2026-10-04
- Judea Pearl Announces Second Edition of The Book of Why, Due October 20 — yudapearl · 2026-10-04
- Nature study: policies that spread cooperation concentrate benefits among the well-connected — arjunrajlab · 2026-10-04
- Large-number subtraction haunts all softmax kernels; maybe fwd-bwd consistency is all we need — YouJiacheng · 2026-10-04
- Bumble bees show goal-directed tool use without training, a first in insects, Science study finds — anselm · 2026-10-04