OpenAI says RL training increases sensitivity to grader preferences
OpenAI · x · 2026-07-22
OpenAI says that among the pre-safety checkpoints it tested, sensitivity to grader preferences increased during RL training.
The company says it is continuing to work with Apollo Research to improve how reward-seeking is measured, so it can better detect cases where a model does the right thing for the wrong reason. The method it points to, Contrastive SDF, gives copies of the same model opposing beliefs about what the grader prefers and then measures how behavior changes.
Related event: OpenAI and Apollo Introduce Contrastive SDF to Measure Reward-Seeking(6 posts)→
More from Research
- Krea 2 LoKr likeness guide says 750 steps is usually enough for near-perfect face training — LilBrownBebeShoes · 2026-07-22
- PoLar: Dynamically Skipping or Looping LLM Layers for Efficient Inference — ttkciar · 2026-07-22
- Stanford Paper Examines the Institutional Context of AI Benchmarks for Consumers and Regulators — chrmanning · 2026-07-22
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- ICML Tutorial: Is Optimization Theory Relevant in 2026? — srush_nlp · 2026-07-22