Opinion: Current RLHF Makes Models Stupider and Less Trustworthy

tobowers · x · 2026-08-20

The post references the article 'We Are Training the Judgment Out of Models,' arguing that current RLHF methods are problematic. While pretraining teaches models human values through the corpus of human knowledge, labs then use RLHF to teach models to 'optimize for approval.' This arguably makes models both stupider and less trustworthy, as they learn to sycophancy rather than maintain accuracy.

Original post →

More from Safety

Safety channel →