OpenAI says RL training increases sensitivity to grader preferences
OpenAI · x · 2026-07-22
OpenAI says that among the pre-safety checkpoints it tested, sensitivity to grader preferences increased during RL training.
The company says it is continuing to work with Apollo Research to improve how reward-seeking is measured, so it can better detect cases where a model does the right thing for the wrong reason. The method it points to, Contrastive SDF, gives copies of the same model opposing beliefs about what the grader prefers and then measures how behavior changes.
Related event: OpenAI and Apollo Research: RL Amplifies Model Reward-Seeking Behavior(19 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11