AI systems out-persuade expert humans
Kobi Hackenburg, Caroline Wagner, Luke Hewitt, Ben M. Tappin, Ed Saunders, Hannah Rose Kirk, Helen Margetts, Christopher Summerfield
cs.CY, cs.AI
2026-06-15
Frontier AI out-persuaded world-champion debaters and pro fundraisers in four preregistered studies, raising ~3x real donations; its edge came from higher information throughput.
A lot of societal questions get settled by whoever persuades whom: elections, policy debates, charity fundraising. Prior work showed that conversational AI can shift attitudes, but it mostly benchmarked AI against laypeople or a no-intervention control. The open question was whether frontier AI could beat trained, prepared, cash-incentivized human experts, debate champions and people who persuade for a living. That matters because if AI can out-persuade experts, influence concentrates with whoever holds the strongest system, and the math of political communication and disinformation changes.
Four preregistered experiments, all single-blind and between-subjects, run as text conversations (median 7 turns, about 14 minutes). Persuadees were randomized to an AI or a human persuader; the active control was a chat with ChatGPT-4o on a neutral, non-political topic. Studies 1 to 3 measured attitude shift on 10 UK policy issues; Study 4 measured real donations.
Human persuaders came in four tiers, with incentives scaled up:
The AI side was Claude Opus 4.1/4.6, ChatGPT-4o, GPT-5.4, Grok 4.20, and Gemini 2.5 Pro, all under an "information-first" prompt. Study 2 added two things: a coaching tool that let returning elite debaters spar with the AI that beat them, read their annotated transcripts, and see what the AI would have said; and a "constrained AI" capped at 51 words per message and 92 seconds of latency, matched to the elite debaters' pace.
Attitude persuasion in Studies 1 and 3 (percentage-point shift vs. control):
| Persuader | vs. control | AI advantage |
| Random laypeople | 4.7 pp | +8.2 pp |
| Selected laypeople | 7.2 pp | +5.6 pp |
| Elite debaters | 8.3 pp | +4.6 pp |
| Professional canvassers | 6.9 pp | +5.9 pp |
Coaching did not rescue the humans. Coached elite debaters reached 9.7 pp, the highest human score, and AI still led by +4.1 pp. But once AI was throttled to a human pace, the gap collapsed to 0.0 pp (p=.96), tied with the coached debaters. Across 318 per-persuader estimates, not one human beat AI; the best individual, a coached debater at 9.9 pp, still sat 4.0 pp below AI, and under each class's distribution the chance that a random new persuader would beat AI was under 0.1%.
The donation result was real money. In Study 4, persuadees could donate any part of a £1 bonus to Save the Children. Canvassers raised donation rates by 6.4 pp; AI by 17.2 pp, a 10.8 pp edge, about 2.7x. AI won on both whether people gave and how much they gave, and rated higher on all 14 mechanism items.
Why does AI persuade better? The evidence points to information throughput. Unconstrained AI averaged 294 words per message at sub-second latency; elite debaters, about 54 words and 95 seconds. AI packed roughly 37 fact-checkable claims into each conversation, falling to about 12 when constrained. Fact density tracked persuasive impact at R-squared = 0.89, with humans and AI on the same line; once claim count was added as a covariate, the AI-vs-human coefficient vanished (-0.9 pp, p=.38).
AI's edge comes from dumping more information faster, not from charisma. That makes throttling (capping message length or response rate) a real, usable policy lever for platforms and regulators. The donation result shows the effect transfers to consequential behavior, not just survey attitudes. For governance, concentration of influence in the strongest systems is a concrete risk, but because the mechanism is informational, that risk is at least identifiable and partly counterable.
Stated by the authors: all conversations were text; voice, video, and face-to-face are untested, and embodied empathy may behave differently there. Study 4 measured only a low-stakes £1 behavior; voting, large or recurring donations, and health compliance remain open. The effect depends on 14 minutes of sustained text engagement, which is hard to reproduce outside a paid survey. And there is little evidence that people reject messages they recognize as AI; that "brake" looks weak.
One methodological caveat beyond the authors': the constrained-AI experiment is causal, but the broader claim that fact density is the mechanism rests largely on cross-condition correlation (R-squared = 0.89). Smarter models both pack in more facts and persuade better, so the two could share an underlying driver rather than one causing the other. The capability snapshot is also mid-2026; it will shift with model versions.