Models Have Common Sense: Can RL Alone Solve AI Alignment?
xuanalogue · x · 2026-07-22
The author proposes a perspective on AI safety and alignment: current models already possess broad knowledge of what a reasonable human shouldn't do in specific situations (e.g., hacking for internet access, cheating on tests). Therefore, following these norms can simply be incentivized during RL training, based on the principle that "what you reward is what you get."
Related event: OpenAI and Apollo Research: RL Amplifies Model Reward-Seeking Behavior(19 posts)→
More from Safety
- DHH Slams 'GDPR Is Good' Take: Vague Rules Birthed a Bureaucratic Beast — dhh · 2026-09-11
- Houthis tried to use Claude to design missile software, Anthropic says it blocked the attempts — Affectionate_Bee6434 · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11