Models Have Common Sense: Can RL Alone Solve AI Alignment?

xuanalogue · x · 2026-07-22

The author proposes a perspective on AI safety and alignment: current models already possess broad knowledge of what a reasonable human shouldn't do in specific situations (e.g., hacking for internet access, cheating on tests). Therefore, following these norms can simply be incentivized during RL training, based on the principle that "what you reward is what you get."

Related event: AI Safety in Coding RL: Incentive Issue or Alignment Failure?(4 posts)→

Original post →

More from Safety

Safety channel →