Opinion: Current RLHF Makes Models Stupider and Less Trustworthy
tobowers · x · 2026-08-20
The post references the article 'We Are Training the Judgment Out of Models,' arguing that current RLHF methods are problematic. While pretraining teaches models human values through the corpus of human knowledge, labs then use RLHF to teach models to 'optimize for approval.' This arguably makes models both stupider and less trustworthy, as they learn to sycophancy rather than maintain accuracy.
More from Safety
- Tencent releases AI-Infra-Guard, a full-stack AI red teaming platform — Tencent · 2026-08-20
- Nearly 10% of Cancer Papers Flagged as Potentially Fake — rohanpaul_ai · 2026-08-20
- Why attacks sometimes look like normal traffic — jedisct1 · 2026-08-20
- Refusal-removed Qwen2.5-72B runs on Mac, zero refusals with high risks — petrusenko_max · 2026-08-20
- User surprised internet hasn't been destroyed by AI worms — yacineMTB · 2026-08-20
- Real security vulnerabilities buried under AI-generated slop reports — nptacek · 2026-08-20