RLHF paper never described instruct model training; insiders reflect on invalidated takes

jd_pressman · x · 2026-09-15

A discussion of cases where facts learned after the time invalidated early analysis: the RLHF paper never actually described how its instruct models were made, and Altman admitted GPT-4o glazed too much and said he'd fix it, then left it that way for months. Blanche Minerva calls this insane and says a retrospective of post-hoc revelations that invalidate early analysis is needed.

Original post →

More from AGI Musings

AGI Musings channel →