USC Study: LLMs Recognize Unanswerable Questions but Fail to Refuse Due to Routing Misalignment
UniversityofSouthernCalifornia · hf · 2026-09-09
University of Southern California researchers found that LLMs encode whether structurally impossible math or code prompts are unanswerable via a hidden-state direction, yet still answer them.
- The recognition signal is misaligned with safety-refusal pathways
- The authors characterize this as a routing failure rather than an encoding failure
- Implication: improving abstention may only require fixing the routing of existing internal signals, not retraining recognition
More from Safety
- Anthropic claims prompt injection solved; Valim cites research showing Claude auto mode still vulnerable — wunderwuzzi23 · 2026-09-09
- Google Keeps a Hidden Copy of Your Browsing History — How to Actually Delete It — nikola_mr64990 · 2026-09-09
- Rep. Ted Lieu cites WSJ Anthropic exit story to push bipartisan AI Kill Switch Bill — austinc3301 · 2026-09-09
- Senate Tech Committee Schedules Zero AI Hearings for Next 4 Months — austinc3301 · 2026-09-09
- Dev calls out Meta AI: 'private and secure' while logging and reviewing all chats — anshulkundaje · 2026-09-09
- Singapore releases world's first governance framework for agentic AI — Comfortable_Gene5180 · 2026-09-09