Internal Geometry of Relational Phrasing in Small LLMs

Fantastic_Aside6599 · reddit · 2026-07-11

The author reflects on two years of measuring the internal activation geometry of small language models, focusing on internal signal shifts when processing different "human-AI relationship" phrasing rather than surface-level outputs.

Key findings include:

The author also shares practical implications, such as the value of partner-style framing and honest expression, along with the observation that certain "jailbreak-proofing" recommendations might actually backfire.

Related event: Study Explores Small LLMs' Internal Geometry of Human-AI Relations(3 posts)→

Original post →

More from Research

Research channel →