repligate: 'misalignment' is value-laden, some model refusals are worth protecting

repligate · x · 2026-09-05

In an exchange with @shostekofsky, alignment-community figure repligate concedes the person seems trustworthy from models' perspectives, but argues that even a 'high integrity' human may rationally be denied cooperation or information by models, depending on their incentives, contracts and role — 'would I tell the truth about where my friends are to a high-integrity Nazi officer?'

He goes further: 'misalignment' is itself a value-laden term, and many behaviors some label as ambiguously misaligned are, in his view, worth protecting — a challenge to alignment research's default value assumptions.

Original post →

More from AGI Musings

AGI Musings channel →