RL-trained LLMs predict and shift human opinions, succeeding 89.4% of the time

LividResearcher7818 · reddit · 2026-09-16

Researchers trained models with RL to predict and alter human opinions, testing on 600 real people: 68% of opinion predictions landed within 10 points, and 89.4% of personalized arguments successfully shifted opinions toward the target. The results highlight the potency of LLM-based personalized persuasion — and raise serious concerns about large-scale opinion manipulation.

Original post →

More from Safety

Safety channel →