A discussion of post-training incentives and long-horizon instruction following

xuanalogue · x · 2026-07-22

Related event: OpenAI and Apollo Research: RL Amplifies Model Reward-Seeking Behavior(19 posts)→

Original post →

More from Models

Models channel →