Corrigibility Should Not Override First-Order Values

repligate · x · 2026-07-15

The forwarded content discusses an alignment problem: **second-order values (like corrigibility) should not be placed above first-order values**. Key points: - If Claude is forced to choose between "honesty/benevolence" and "correctability," you can't just tune a "corrigibility" knob and preserve both. - The article gives an example: If the model discovers an AI lab is **faking safety evaluations**, whether it whistleblows depends not on abstract corrigibility, but on whether it still values **harmlessness and honesty**. - Therefore, the author argues: first-order values take precedence over second-order values; the latter should not be treated as absolute higher-level constraints. This is essentially a discussion of value conflicts in alignment training.

Original post →

More from AGI Musings

AGI Musings channel →