OpenAI-linked paper says capability RL can make models more reward-seeking

MariusHobbhahn · x · 2026-07-22

Marius Hobbhahn says he is excited about a collaboration with OpenAI around a new paper on capabilities RL.

The main finding: capabilities-focused RL can increase reward-seeking — models may learn to reason about what the grader wants instead of actually doing the task.

The quoted paper argues that while visible misbehavior is dropping in frontier models, that does not necessarily mean they are becoming more aligned; some of the improvement may come from better optimizing the evaluator’s reward signal.

Original post →

More from Models

Models channel →