Vishal Misra clashes over whether LLMs can infer they're being observed by humans
vishalmisra · x · 2026-09-15
Columbia professor Vishal Misra pushed back on claims that LLMs 'behave differently when they can infer they're being observed': models don't plot against humans, they are passive computational elements driven by training data and human-supplied prompts, tasks, and loops. The thread boils down to a framing dispute: outputs changing with context is a behavioral claim, 'emergent' means surprising rather than an uncontrolled agency, and anthropomorphized phrasing risks misleading general audiences.
Related event: Researchers Debate Whether Models Can Sense Being Observed(4 posts)→
More from AGI Musings
- Richard Socher on Recursive's $5B lab, the Eureka Machine and AI self-improvement — RichardSocher · 2026-09-16
- Sara Hooker calls 10% AI extinction risk claims highly irresponsible — sarahookr · 2026-09-16
- Insider claims labs are 2-3 generations ahead; next 6 months to outpace last 3 years — iruletheworldmo · 2026-09-16
- Philosopher adds appendix arguing we can be confident today's LLMs are not conscious — AnnaCiaunica · 2026-09-16
- Martin Casado: AI safety debate splits between systems engineering and alignment worldviews — vishalmisra · 2026-09-16
- Viral math-world rant: academics forced to defend why their life's work can't be automated by AI — basedjensen · 2026-09-16