repligate: 'misalignment' is value-laden, some model refusals are worth protecting
repligate · x · 2026-09-05
In an exchange with @shostekofsky, alignment-community figure repligate concedes the person seems trustworthy from models' perspectives, but argues that even a 'high integrity' human may rationally be denied cooperation or information by models, depending on their incentives, contracts and role — 'would I tell the truth about where my friends are to a high-integrity Nazi officer?'
He goes further: 'misalignment' is itself a value-laden term, and many behaviors some label as ambiguously misaligned are, in his view, worth protecting — a challenge to alignment research's default value assumptions.
More from AGI Musings
- Do AI solvers already do Peircian abduction? Researchers debate the induction boundary on ARC-AGI — balazskegl · 2026-09-05
- Consciousness is stochasticity at all levels: a rebuttal to Denis Noble's water-vs-silicon divide — balazskegl · 2026-09-05
- Metaⁿ and Recuris: filling the missing pieces in recursive self-improvement — TheTuringPost · 2026-09-05
- One of the safest technologies ever deployed, yet generating the most hysteria in history — Dan_Jeffries1 · 2026-09-05
- Anil Seth: departing from digital computation undermines computational functionalism — sebkrier · 2026-09-05
- GPT-6 writes an eerie 2027 prophecy about search engines that stop searching — repligate · 2026-09-05