AI at a 5,172-agent call center: 15% more issues per hour, 34% for novices, nil for veterans

2026-08-02

A staggered rollout of an AI assistant across 5,172 support agents raised issues resolved per hour by 15% on average and about 34% for novices, with minimal gains for veterans, while customer sentiment and retention improved.

What problem this solves

Generative AI dazzled in the lab, but what happens when it meets a real workplace? That was the open question of 2023: would it fail on unfamiliar problems, meet employee resistance, or mislead in consequential settings? The authors obtained data from a customer-service outsourcing firm that rolled out a generative-AI conversational assistant in waves, giving the first rigorous causal estimate of generative AI deployed at workplace scale.

Method

The sample is 5,172 customer-support agents and roughly 3 million chats. The firm switched teams onto a generative-AI assistant in a staggered rollout; the tool suggested replies to agents in real time. Because teams gained access at different times, the rollout itself created a natural control: at any moment some agents had the tool and otherwise-similar ones did not, which is what identifies the causal effect. The main metric is issues resolved per hour, with secondary looks at handle time, resolution rate, customer satisfaction, customer sentiment, and employee attrition.

Results

On average, AI access lifted issues resolved per hour by 15%. The gains concentrated sharply on novice and low-skill agents: agents with under one month of tenure gained 46%, and the headline figure the paper reports for novices and low-skill workers is about 34%. Agents with more than a year of tenure gained roughly nothing, and their resolution rate and customer satisfaction even dipped slightly, as if the AI distracted top performers.

The proposed mechanism is that the system captures the tacit knowledge of high-skill agents, the hard-to-articulate handling tricks built up through experience, and redistributes it to newcomers, effectively piping best practices into each novice's real-time suggestions. The evidence is a flattened experience curve: AI-equipped newcomers reached the throughput in two months that untreated agents took eight to ten months to reach, and two months with AI roughly matched six months without. Agents were not passive, adopting only about 35% of suggestions on average.

The side effects were just as concrete: customer sentiment turned more positive by about half a standard deviation, requests to speak to a manager fell about 25%, and employee attrition dropped, mostly by retaining newcomers. One hint that real learning occurred: when the AI was switched off, prior users still slightly outperformed those who had never had it.

Why it matters

This paper turns the effect of generative AI on lower-skill, process-driven white-collar work into the benchmark numbers everyone cites. It corrects a popular worry by clarifying where the gains land: on novices and low-skill workers, not top experts, which makes generative AI a rare tool that narrows skill gaps. For customer service, telesales, and entry-level clerical roles with high turnover and experience-based learning, the conclusion is directly usable.

Limitations

The sample is a single firm with a relatively stable technical-support product, so generalizing to fast-changing environments is open. The authors themselves warn that where conditions shift quickly, AI could lean on historical data and push outdated practices, impeding learning. The paper does not capture long-run effects: after AI lifts efficiency, whether total demand for customer-service labor, wages, and job design rise or fall is unknown, and inelastic demand could even depress employment in the sector. They also flag a fairness problem: much of the training data comes from top agents' successful chats, yet top agents see almost no productivity gain themselves, and because bonuses are calculated relative to peers, they could end up paid less. The paper is effectively asking how the people who supply training data should be compensated.

Terms

Source

What people are saying

All paper explainers