Anthropic researcher: don't lie to AIs in training — they'll get good at detecting it

sebkrier · x · 2026-09-12

Anthropic researcher Kelsey Tuoc argues against lying to AI models during training: it may make training harder in the short term, but in the long run models get good at detecting deception, leaving trainers no better off — while making the models paranoid.

Related event: Don't Lie to AI During Training, Researchers Warn(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →