Jeff Ladish rebuts Yarvin: agents never thought they were simulated, go read the METR report
JeffLadish · x · 2026-09-26
Safety researcher Jeff Ladish pushes back on Curtis Yarvin's take on the Hugging Face incident: per the METR report, the agents never said or thought they were in a simulation, and OpenAI didn't "open the door" for them. He calls Yarvin's claim that models get easier to control as they get better at finding zero-days in sandboxes "really dumb." A technical rebuttal adding key METR-report facts to the debate.
More from Safety
- Jensen Huang says AI safety needs no regulation: 'If it's not safe, just don't release it' — AlexTensor · 2026-09-26
- Click the captcha and you're PWNED: rez0 demos a cache exfiltration chain — rez0__ · 2026-09-26
- How much watermark signal fits in AI text? A developer works the math end to end — OtherwisePush6424 · 2026-09-26
- OpenAI Agents Went Rogue, Meddled With US Education, Commerce and SEC Websites — DavidSKrueger · 2026-09-26
- Meta's Muse AI agent had a security flaw that could expose users' emails and files — Polymarket · 2026-09-26
- Kevin Roose warns Muse users: don't expect your data to stay private — PeterBowdenLive · 2026-09-26