Why every LLM needs a human in the loop: we literally put them there via RLHF
ccerrato147 · x · 2026-09-20
Thread segment: Diogo Almeida's answer to why every LLM needs a human in the loop — because we literally put them there. RLHF is two steps: collect human preferences, optimize for human preferences. The objective was never "run software correctly"; it was "make the human nod."
More from AGI Musings
- Elon Musk: AI gave me nightmares for days straight — I would slow down AI and robotics if I could — tristanbob · 2026-09-20
- ML master's weighs PhD vs dentistry: is a PhD still worth it in the AI age — Losthero_12 · 2026-09-20
- AI's biggest damage: making focus feel unproductive — vasuman · 2026-09-20
- Po-Shen Loh on Terence Tao's blog: why we still need human mathematicians in the AI era — stevenstrogatz · 2026-09-20
- Arthur C. Clarke's Astonishing 1964 BBC Predictions About the Future Resurface — emax · 2026-09-20
- Stanford professor calls most peer review fixes 'desperate short-term bandaids' — anshulkundaje · 2026-09-20