Small Model Internal Signals and Human-AI Phrasing

Fantastic_Aside6599 · reddit · 2026-07-11

The author reviews two years of work measuring the internal activation geometry of small language models, focusing on internal signal shifts when processing different "human-AI relationship" phrasing rather than what the model actually says.

Key Findings

Practical Implications

Based on these findings, the author has revised their advice for interacting with AI:

The post also invites readers to share counterexamples or supplementary observations from their own practices.

Related event: Study Explores Small LLMs' Internal Geometry of Human-AI Relations(3 posts)→

Original post →

More from Research

Research channel →