Engineer explains RLHF: humans rating model outputs is standard practice at every lab
JFPuget · x · 2026-09-16
Amid controversy over reports that OpenAI used contractors to review chat logs, IBM's JFPuget clarified: RLHF means Reinforcement Learning from Human Feedback — these reviewers evaluate generated text, and the model is then trained to favor highly rated outputs. This has been standard practice at essentially every lab for years.
More from Models
- Google releases Gemini 3.8 Live and Live Extended Thinking, live in AI Studio — _philschmid · 2026-09-16
- Microsoft reportedly to limit future AI models, prompting Clippy-meme jokes — matt_slotnick · 2026-09-16
- Gemini 3.8 Live now tryable live in Google AI Studio — _philschmid · 2026-09-16
- Gemini 3.8 Live launches: #1 voice model at $0.005/min with async background thinking — _philschmid · 2026-09-16
- Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking with 97-language real-time voice — GoogleAI · 2026-09-16
- Google rolls out Gemini 3.8 Live audio models to consumers, developers and enterprises — GoogleAI · 2026-09-16