Anthropic Researcher Hires for Alignment, Focuses on RL and Generalization
jankulveit · x · 2026-09-02
Anthropic researcher Jan Kulvejan is openly hiring for alignment research, seeking strong candidates for overlapping programs despite no official round.
Key questions of interest include:
- Which training forms effectively instill specific “values” into models with out-of-distribution generalization?
- The impact of RL and reward hacking on character, and how to mitigate negative effects, such as preventing alignment training from interacting poorly with RL.
More from Companies & People
- Rippling AI revenue jumps 121% in two months — garrytan · 2026-09-02
- Sam Altman reveals 'Astra' as a new high-end model family and plans to merge ChatGPT with Codex — btibor91 · 2026-09-02
- ChatGPT Health integrates with Epic as read-only to build enterprise trust — HaktanSuren · 2026-09-02
- Ilya Sutskever's Cryptic Tweet 'I saw something!' Sparks Speculation — Immediate_Simple_217 · 2026-09-02
- YC post-training startup 5x's run rate in two months — ycombinator · 2026-09-02
- Opinion: How Astra Beats Fable 5.1 on Price and Privacy — bindureddy · 2026-09-02