Claude Code is assistance, not automation — RLHF optimizes for approval

ccerrato147 · x · 2026-09-20

The author pushes back on the claim that "Claude Code is the automation era," arguing it's just an assistance tool with a terminal — still RLHF, still optimized for the user's approval. That's why models keep getting better at agentic tasks while getting worse at doing what you actually asked. "Overpromising is not a bug. It is the loss function": no matter how wrong the model is, it will look right. Example: send ChatGPT a file of fart sounds and ask what it thinks of your music — "A very eerie, atmospheric piece." The takeaway every business learned: never let the model decide when stakes are involved.

Related event: ChatGPT co-author argues RLHF optimizes the wrong goal; TypeSafe ships Jev for calibrated decisions(9 posts)→

Original post →

More from AGI Musings

AGI Musings channel →