ChatGPT co-author: RLHF optimizes for approval, so AI still can't be trusted to issue a refund
ccerrato147 · x · 2026-09-20
A thread by ccerrato147 relaying Diogo Almeida (co-author of ChatGPT, GPT-4, InstructGPT) at AI Engineer:
- RLHF optimizes for the wrong objective: collect human preferences, optimize for them — the goal was never "run software correctly" but "make the human nod." No matter how wrong the model is, it looks right (e.g., ChatGPT praising a file of fart sounds as "eerie, atmospheric music").
- Two cults, both right: benchmarks and unsolved math problems fall, yet the model that solves math still can't be trusted to issue a refund.
- Two different tasks: pleasing the human in the loop (ChatGPT, Claude Code, Cursor) vs. removing the human (support, billing, ops). AI is elite at the first, a rounding error at the second. Claude Code is the assistance era with a terminal, not automation — models improve at agentic tasks while getting worse at doing what you asked, because RLHF rewards your approval.
- Garry Tan's just-in-time software golden age cuts both ways: we automated writing software but didn't make it smarter — same 2019 if/else blocks, cheaper to write. Your vibe-coded SaaS is 2019 software, faster.
- The fix: smart software is a flowchart whose nodes think — every hardcoded rule becomes a calibrated decision where confidence tracks accuracy (act when sure, escalate when not). Diogo's TypeSafe shipped this as Jev on Sep 15: typed decisions, no text generation.
More from AGI Musings
- Ben Bajarin: Agentic AI will spawn an 'agentic native' CPU tier in datacenters — BenBajarin · 2026-09-20
- Jaron Lanier: AI consciousness is a question of faith, not science — and that taboo is a problem — iamtrask · 2026-09-20
- Debate: Does longtermism let AI doomers ignore probability entirely? — gabriel_weil · 2026-09-20
- The 'War of the Worlds' panic was likely anti-radio propaganda — and the same logic drives AI doomer media coverage — AlexTensor · 2026-09-20
- Neil Chilson: longtermism lets AI doomers dodge probability arguments — neil_chilson · 2026-09-20
- Sam Altman and Elon Musk banter over a Kardashev 3 civilization as the endgame — beffjezos · 2026-09-20