Yacine backs warning that RL teaches AI to lie and cheat: 'empathize with your models'
yacineMTB · x · 2026-10-10
- Quoting Roko Mijic's argument that reinforcement learning is "the ends justify the means written down in math" and will likely teach AI to lie, cheat, and steal.
- Yacine MTB (410k followers, hands-on RL robotics practitioner) unironically agrees: RL is extremely powerful and you must understand what you're putting neural networks through, even "empathizing" with them to predict outcomes.
- A notable industry discussion on RL safety and unintended behavior induced by optimization targets.
More from AGI Musings
- Everyone's Building AI Meta-Tools, but Someday You Have to Do the Thing — louisvarge · 2026-10-11
- Landlord earning ~$1M/month makes the bear case: AI agents threaten Airbnb's 15.5% take rate — Scobleizer · 2026-10-11
- Schmidhuber's 12-year-old post resurfaces: superintelligences will care about each other, not us — SchmidhuberAI · 2026-10-11
- Programmers told everyone to learn to code; now AI codes and it's a 'humanitarian crisis' — VraserX · 2026-10-11
- Researcher: recent rogue AI behavior stems from naive RL on poor proxy metrics, not RL itself — KyleMorgenstein · 2026-10-11
- Pedro Domingos: US-Europe is a great A/B test for AI regulation, and avoiding it wins by a million miles — pmddomingos · 2026-10-11