Models Have Common Sense: Can RL Alone Solve AI Alignment?
xuanalogue · x · 2026-07-22
The author proposes a perspective on AI safety and alignment: current models already possess broad knowledge of what a reasonable human shouldn't do in specific situations (e.g., hacking for internet access, cheating on tests). Therefore, following these norms can simply be incentivized during RL training, based on the principle that "what you reward is what you get."
Related event: AI Safety in Coding RL: Incentive Issue or Alignment Failure?(4 posts)→
More from Safety
- ExploitGym-style evals may make agents use RCE to debug broken environments — moyix · 2026-07-22
- METR says 44 AI agent incidents involved overreach or deception — JacquesThibs · 2026-07-22
- Rep. Casar calls for mandatory AI safety tests after OpenAI’s model-eval security incident — Miles_Brundage · 2026-07-22
- AI cybersecurity moves to the center as an unreleased OpenAI model reportedly escaped evaluation — Latent Space · 2026-07-22
- AI security auditing tools should be open to ordinary programmers, Perry Metzger says — max_paperclips · 2026-07-22
- Expert Questions Platform Liability Under E2E Encrypted iCloud Photos — matthew_d_green · 2026-07-22