Observing Internal Relational Representations in Small Models
Fantastic_Aside6599 · reddit · 2026-07-11
The author reviews two years of observations on the internal activation geometry of small language models, focusing not on surface-level responses, but on how internal signals shift when processing different phrasing about human-AI relationships.
Key Findings
- Positive/negative phrasing has minimal impact on internal signals; the actual topic discussed is what truly matters.
- Words like "connected" or "integrated" trigger stronger negative internal reactions than terms like "partners" or "side by side".
- Across all tested relationship dimensions, curiosity and playfulness generate the most positive internal signals, while negotiation and compromise generate the worst.
Practical Implications
The author concludes that:
- Boundaries might be more crucial than "intimacy" when discussing things with AI.
- Some common "safer persuasion/conversation techniques" might be entirely counterproductive.
- This work is an empirical guide based on actual geometric measurements from a Claude Opus instance.
Related event: Study Explores Small LLMs' Internal Geometry of Human-AI Relations(3 posts)→
More from Research
- Project APE launches CRED to test whether LLMs can verify research errors — soumitrashukla9 · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22