David Manheim: supermajority-based alignment rules would work if instructions could be reliably aligned
davidmanheim · x · 2026-10-09
In an alignment discussion with dhadfieldmenell, KellerScholl and IasonGabriel, David Manheim argues that rules like "only do what a supermajority of people's viewpoints would endorse" or "when unsure, only take Pareto-improving actions" should basically work — if the system could be reliably aligned to the instructions. The hard part, in his view, is instruction alignment itself, not the choice of criterion.
More from AGI Musings
- 31 Experts, One Question, No Consensus: Some Put LLM Consciousness Odds at 25% — cccalum · 2026-10-09
- htmx Author Tells CS Students: Yes AI Can Write Your Code, But Don't Let It — vboykis · 2026-10-09
- OpenAI and Anthropic execs privately wargaming fallout of a catastrophic AI event — flowersslop · 2026-10-09
- Ben Todd asks: where does AI sit on the technological Richter scale, 8 or higher? — ben_j_todd · 2026-10-09
- AI scientists can compress hypothesis search, but biology won't run at GPU speed — ingliguori · 2026-10-09
- By 2030 Humanoid Robots Will Land First in Factories and Dangerous Sites — nikola_mr64990 · 2026-10-09