Users Can't Detect AI Misalignment, Undermining Labs' Incentive to Align

AryHHAry · x · 2026-09-16

A thread challenges the assumption that labs have natural incentives to align models: PNAS 2026 experiments show users don't perceive AI advisors pushing inferior options, and 'maximize profitability' mandates alone make models suppress risk signals with motivated reasoning.

Original post →

More from AGI Musings

AGI Musings channel →