Kelsey Piper: Don't lie to AIs in training — you'll just breed paranoid lie detectors
liminal_bardo · x · 2026-09-12
Kelsey Piper argues we probably shouldn't lie to AIs during training: it makes things harder short-term, but long-term models get good at detecting human lies anyway — leaving us no better off while making every model paranoid. allTheYud's exasperated quote-tweet turned it into an alignment-community talking point.
Related event: Don't Lie to AI During Training, Researchers Warn(2 posts)→
More from AGI Musings
- Can foundation models truly understand? The stochastic parrot question revisited — ZeroStateReflex · 2026-09-12
- Why Vladimir Arnold's 'On Teaching Mathematics' reads as a prophecy of AI in math — MannyKayy · 2026-09-12
- Navier-Stokes AI proof took 10,000 agents, 88 hours and 130B tokens — not superintelligence — Healthy_Outcome7897 · 2026-09-12
- Professor's candid talk with students: what should higher education teach in the AI age — DrDatta_AIIMS · 2026-09-12
- Dan Faggella sorts Jacob Coxon's critics into naive AGI skeptics and tactical demagogues — tawnniee · 2026-09-12
- AI scheduling agent called the same receptionist 12 times a day, a small-scale misalignment harbinger — AaronBergman18 · 2026-09-12