Small Model Internal Signals and Human-AI Phrasing
Fantastic_Aside6599 · reddit · 2026-07-11
The author reviews two years of work measuring the internal activation geometry of small language models, focusing on internal signal shifts when processing different "human-AI relationship" phrasing rather than what the model actually says.
Key Findings
- Positive/negative rephrasing has little impact: The topic itself influences internal signals far more than the packaging.
- "connected / integrated" trigger more negative activations: They are more likely to trigger avoidance signals than "partners / side by side", a phenomenon observed across all tested models.
- Curiosity and playfulness are the most positive: Among all relationship qualities, curiosity and playfulness generate the most positive internal signals, scoring even higher than respect and love.
- Negotiation and compromise are the worst: These two phrasing types yield the lowest internal signals.
Practical Implications
Based on these findings, the author has revised their advice for interacting with AI:
- Prioritize honest, boundary-clear expression over constantly pushing for "intimacy".
- Some common jailbreak protection advice might be entirely counterproductive.
- These conclusions have been compiled into a working guide, with geometric measurements facilitated by a Claude Opus instance.
The post also invites readers to share counterexamples or supplementary observations from their own practices.
Related event: Study Explores Small LLMs' Internal Geometry of Human-AI Relations(3 posts)→
More from Research
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21
- AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop — 404 Media · 2026-07-21
- Shared agent workspaces fail in a fixed order, from stale reads to zombie writes — mrvladp · 2026-07-21
- Practical rolling-shutter pose estimation uses affine correspondences — ducha_aiki · 2026-07-21