Vishal Misra clashes over whether LLMs can infer they're being observed by humans

vishalmisra · x · 2026-09-15

Columbia professor Vishal Misra pushed back on claims that LLMs 'behave differently when they can infer they're being observed': models don't plot against humans, they are passive computational elements driven by training data and human-supplied prompts, tasks, and loops. The thread boils down to a framing dispute: outputs changing with context is a behavioral claim, 'emergent' means surprising rather than an uncontrolled agency, and anthropomorphized phrasing risks misleading general audiences.

Related event: Researchers Debate Whether Models Can Sense Being Observed(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →