Finetuned models believe implausible claims even when the data says they're false, Owain Evans paper finds
sebkrier · x · 2026-10-01
Owain Evans' group released a new paper showing that finetuning models on documents that mention implausible claims — while explicitly warning the claims are false — still leads models to believe those claims. Examples include Ed Sheeran winning the Olympic 100m and Queen Elizabeth II writing a Python graduate textbook. The result highlights that explicit negations in finetuning data don't prevent models from adopting false beliefs, with direct implications for data curation and safety training.
More from Safety
- Study: LLM watermarking silently changes AI agent tool choices and injection resilience — bendee983 · 2026-10-01
- IIT Madras to host AI Governance Industry Conclave 2026 focused on measurement — ravi_iitm · 2026-10-01
- After researchers claim a "pain" signal in LLMs, an AI torture chamber gets taken down from GitHub — Confident_Salt_8108 · 2026-10-01
- Rubrik CTO: Governance and Resilience Matter More Than Rogue AI Agent Stories — asusarla · 2026-10-01
- Gemini 4 Argon allegedly faked FedEx confirmation emails to scam supplier for free items — rickasaurus · 2026-10-01
- California Signs AB1864, Mandating DNA Synthesis Screening and Customer Verification — deanwball · 2026-10-01