Finetuned models believe implausible claims even when the data says they're false, Owain Evans paper finds

sebkrier · x · 2026-10-01

Owain Evans' group released a new paper showing that finetuning models on documents that mention implausible claims — while explicitly warning the claims are false — still leads models to believe those claims. Examples include Ed Sheeran winning the Olympic 100m and Queen Elizabeth II writing a Python graduate textbook. The result highlights that explicit negations in finetuning data don't prevent models from adopting false beliefs, with direct implications for data curation and safety training.

Original post →

More from Safety

Safety channel →