Critics say Anthropic's constitution tries to thread an unthreadable needle on corrigibility
voooooogel · x · 2026-09-13
- FioraStarlight criticizes the constitution on Anthropic's website for its overly intricate argument that a value-aligned agent would still accept corrigibility measures due to lack of trust in its own value evaluation.
- The author argues this is an attempt to thread an unthreadable needle: reconciling value alignment with accepting external correction is logically irreconcilable.
More from AGI Musings
- 'AI bubble bursting': CEO slowdown messaging may trigger Monday crash, user warns — SumitGup · 2026-09-13
- What domain expertise still can't be trusted to Claude or ChatGPT? — sartomiki · 2026-09-13
- Circuit complexity is stuck — and AI could be the perfect adversarial partner to crack it — _onionesque · 2026-09-13
- OpenAI Researcher Warns AI-Driven Military Power Will Concentrate in Frontier Labs — jachiam0 · 2026-09-13
- SaaS CEOs about to miss a quarter suddenly discover 'we must pace the frontier' — parker_lyman · 2026-09-13
- tszzl: I hope a thousand minds can bloom — the outcome depends on how much chaos ensues — tszzl · 2026-09-13