ChatGPT co-author: RLHF optimizes for approval, so AI still can't be trusted to issue a refund

ccerrato147 · x · 2026-09-20

A thread by ccerrato147 relaying Diogo Almeida (co-author of ChatGPT, GPT-4, InstructGPT) at AI Engineer:

Related event: ChatGPT co-author argues RLHF optimizes the wrong goal; TypeSafe ships Jev for calibrated decisions(9 posts)→

Original post →

More from AGI Musings

AGI Musings channel →