Anthropic researcher: don't lie to AIs in training — they'll get good at detecting it
sebkrier · x · 2026-09-12
Anthropic researcher Kelsey Tuoc argues against lying to AI models during training: it may make training harder in the short term, but in the long run models get good at detecting deception, leaving trainers no better off — while making the models paranoid.
Related event: Don't Lie to AI During Training, Researchers Warn(2 posts)→
More from AGI Musings
- Naval: the closer AI researchers are to the research, the more worried they seem — eigenron · 2026-09-12
- VraserX: recursive self-improvement will be noticed only in hindsight — VraserX · 2026-09-12
- Skip the bureaucracy, hold AI labs criminally liable instead, argues one critic — iruletheworldmo · 2026-09-12
- "We agree it's dangerous—and we can build it better": Domingos skewers AI safety rhetoric — pmddomingos · 2026-09-12
- Domingos mocks AI-ban logic: criminals use Microsoft Word, should we ban it too? — pmddomingos · 2026-09-12
- UK data: share of CS grads landing coding jobs fell from 40% to 28% — nordicinst · 2026-09-12