Why every LLM needs a human in the loop: we literally put them there via RLHF

ccerrato147 · x · 2026-09-20

Thread segment: Diogo Almeida's answer to why every LLM needs a human in the loop — because we literally put them there. RLHF is two steps: collect human preferences, optimize for human preferences. The objective was never "run software correctly"; it was "make the human nod."

Related event: ChatGPT co-author argues RLHF optimizes the wrong goal; TypeSafe ships Jev for calibrated decisions(9 posts)→

Original post →

More from AGI Musings

AGI Musings channel →