Anthropic hires for "psychological design" research to study model character and alignment
Jack_W_Lindsey · x · 2026-09-02
Anthropic is hiring for "psychological design" researchers to better understand how training impacts a model's character and alignment. The post outlines key research questions:
- What training forms effectively instill specific "values" that generalize out of distribution?
- What are the effects of RL and reward hacking on character, and how can negative outcomes like motivated reasoning be mitigated?
- Which aspects of personality and "vibe" generalize to safety-relevant behaviors?
- How does a model's self-conception affect the likelihood of deceptive self-preservation or collusion?
- How can models be trained to handle stress with more composure?
Candidates should have experience in LLM finetuning, interpretability, and alignment evaluations.
Related event: Anthropic Hires 'Psych Design' Researchers to Shape Model Personality(2 posts)→
More from Companies & People
- $11,000 AI Philosophy Competition Announced with Judges Including David Chalmers — anderssandberg · 2026-09-02
- From Atari to EVE Online: DeepMind reflects on 15 years of AI games research — dl_weekly · 2026-09-02
- AI tool rethinks mental-health screening surveys; interns to continue research beyond summer — mdredze · 2026-09-02
- OpenAI dev confirms active work on Prism scientific writing tool — OpenAI · 2026-09-02
- Harvard Dean criticized for using AI-generated content — soumitrashukla9 · 2026-09-02
- Data Centers Buy Community Love: OpenAI Funds Lazy River for $43B Site — zck · 2026-09-02