AI Alignment Researchers Reflect: Current RLHF and HHH Frameworks Are Too Simplistic
sebkrier · x · 2026-08-13
AI alignment researcher Seb Krier discussed an alternative approach proposed by @meaningaligned to Constitutional AI or simple RLHF-based fine-tuning. He noted that the widely adopted HHH (Helpful, Harmless, Honest) framework might be viewed in a few years as a very naive and simplistic fix.
Krier also raised a core question regarding value alignment: whether we need a process that mimics democratic decision-making, or if enabling more decentralized fine-tuning—allowing values and approaches to compete—naturally leads to better outcomes. He finds this research direction fascinating and worth exploring.
More from Safety
- Legal Liability of AI Agents: Humans Should Remain the Ultimate Bearers of Risk — sebkrier · 2026-08-13
- AI Agents as the New Attack Surface: Real-World Prompt Injection Thwarted — Bino5150 · 2026-08-13
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13