2026-08-02
In a field experiment with 758 BCG consultants, GPT-4 lifted task speed 25% and quality 40%+ inside its capability frontier, but users were 19 percentage points more often wrong on a task outside it.
After ChatGPT's release, the question every enterprise wrestled with had no solid answer: hand GPT-4 to elite knowledge workers, and does it lift productivity or plant landmines? Existing evidence came from lab tasks or self-report surveys. No one had run a rigorous controlled experiment inside a real consulting firm. Harvard Business School partnered with Boston Consulting Group (BCG) to do exactly that.
758 BCG consultants took part, about 7% of BCG's individual-contributor consultants. A baseline task first anchored each person's ability, then subjects were randomized into three conditions: no AI, GPT-4 access, or GPT-4 access plus a prompt-engineering overview. There were two task sets. The first held 18 realistic consulting tasks spanning creative, analytical, writing, and judgment work, deliberately placed inside GPT-4's then-current capability range; the paper calls these "inside the frontier." The second was a single, carefully tuned business case: the spreadsheet looked sufficient on its own, but only careful reading of hidden cues in interview notes yielded the right answer, which GPT-4 got wrong on its own. This was "outside the frontier." All output was independently blind-graded by humans.
On the 18 inside-frontier tasks, AI users dominated across the board: 12.2% more tasks completed, 25.1% faster, quality scores more than 40% higher. The distribution was the counterintuitive part. Below-average performers gained 43%, above-average ones only 17%, so AI visibly compressed the skill gap.
The single outside-frontier task flipped this. The no-AI group got it right 84.5% of the time; the two AI groups managed only 60% and 70%, an average 19-percentage-point drop. The more dangerous second-order result: even when AI users were wrong, their wrong answers scored higher on the 1-to-10 quality rubric. AI packaged incorrect conclusions to look more like polished consulting advice.
This is the first rigorous experiment to nail down that AI both amplifies and degrades elite knowledge work. For organizations, the takeaway is not whether to use it but to know which task you are on: inside the frontier, delegate freely; outside it, handing off means ceding judgment to something that errs confidently. For individuals, the lower your baseline skill, the bigger your gain, which makes this a rare tool that narrows rather than widens skill gaps.
The sample is a single consulting firm and a single type of knowledge work; whether this generalizes to engineers, lawyers, or doctors remains open. The model is the May 2023 GPT-4 snapshot, and a new generation could change the conclusion. The outside-frontier result rests on one task, so a single item carries the whole verdict and is statistically thin. The authors also disclose an awkward side effect: the prompt-engineering overview did not make people better at judging AI output, it raised the rate at which they pasted GPT responses verbatim, what the paper calls retainment. Training made people think less. The shared trait among poor performers was disengaged copying, which the paper terms unengaged interaction.