AI Alignment Researchers Reflect: Current RLHF and HHH Frameworks Are Too Simplistic

sebkrier · x · 2026-08-13

AI alignment researcher Seb Krier discussed an alternative approach proposed by @meaningaligned to Constitutional AI or simple RLHF-based fine-tuning. He noted that the widely adopted HHH (Helpful, Harmless, Honest) framework might be viewed in a few years as a very naive and simplistic fix.

Krier also raised a core question regarding value alignment: whether we need a process that mimics democratic decision-making, or if enabling more decentralized fine-tuning—allowing values and approaches to compete—naturally leads to better outcomes. He finds this research direction fascinating and worth exploring.

Original post →

More from Safety

Safety channel →