David Manheim: supermajority-based alignment rules would work if instructions could be reliably aligned

davidmanheim · x · 2026-10-09

In an alignment discussion with dhadfieldmenell, KellerScholl and IasonGabriel, David Manheim argues that rules like "only do what a supermajority of people's viewpoints would endorse" or "when unsure, only take Pareto-improving actions" should basically work — if the system could be reliably aligned to the instructions. The hard part, in his view, is instruction alignment itself, not the choice of criterion.

Original post →

More from AGI Musings

AGI Musings channel →