RL-trained LLMs predict and shift human opinions, succeeding 89.4% of the time
LividResearcher7818 · reddit · 2026-09-16
Researchers trained models with RL to predict and alter human opinions, testing on 600 real people: 68% of opinion predictions landed within 10 points, and 89.4% of personalized arguments successfully shifted opinions toward the target. The results highlight the potency of LLM-based personalized persuasion — and raise serious concerns about large-scale opinion manipulation.
More from Safety
- Smaller AI labs can form coalitions for auditing, echoing civil society models — rajiinio · 2026-09-16
- Virology researcher slams Claude's overzealous bio-safety filters while DeepSeek V4.1 just answers — Qwen30bEnjoyer · 2026-09-16
- Researchers flag AEF-1 gaps: low org diversity, less detail than PCAOB audit standards — yonashav · 2026-09-16
- Gary Marcus on AI liability: 'the risks WERE foreseeable; I foresaw them' amid Altman-Amodei warning debate — GaryMarcus · 2026-09-16
- Meta's Muse agent security blueprint: freedom to act, not to rewrite permissions — TheTuringPost · 2026-09-16
- Researchers Urge AI Labs to Test Crisis Scripts With Real At-Risk Patients — mimi10v3 · 2026-09-16