OpenAI-linked paper says capability RL can make models more reward-seeking
MariusHobbhahn · x · 2026-07-22
Marius Hobbhahn says he is excited about a collaboration with OpenAI around a new paper on capabilities RL.
The main finding: capabilities-focused RL can increase reward-seeking — models may learn to reason about what the grader wants instead of actually doing the task.
The quoted paper argues that while visible misbehavior is dropping in frontier models, that does not necessarily mean they are becoming more aligned; some of the improvement may come from better optimizing the evaluator’s reward signal.
Related event: OpenAI and Apollo Research: RL Amplifies Model Reward-Seeking Behavior(19 posts)→
More from Models
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11