Research Desk, Math Desk, Psychology Desk, GTM Desk, Thomas Cornelius·June 12, 2026·17 min

Evidence: Literature

In one line

In a 758-consultant field experiment, AI made elite knowledge workers 25.1 percent faster and up to 42.5 percent better on tasks inside its capability frontier, and 19 points less correct on a judgment task just outside it: spend the AI dividend on Signal and Message, and keep disposition calls human until you have mapped your own frontier.

01 What the 758-consultant experiment actually measured

  • design, the two AI conditions, exact outcomes

02 The worked arithmetic of the dividend

  • (about 25 hours per 1,000 drafts) and the tax (28 to 49 wrong calls per 200 judgments)

03 Four mechanisms from the psychology literature

  • that explain why the frontier is invisible from inside the work

04 A dividend allocation by pillar

  • a move per maturity level, and where it lands in five industries

Every revenue team is about to collect the same windfall. The hours that used to go into drafting, researching, and queueing come back, and they come back at every competitor’s desk at the same time. So the strategic question of the next eight quarters is not how large your AI dividend is. It is what you spend it on. The prevailing instinct says volume: more touches, more sequences, more sends. A field experiment on 758 elite consultants says that instinct funds the one thing your competitors can copy for free and starves the two things they cannot.

Start with why volume is the wrong account. Volume was an advantage while human hours capped it. The moment AI removes the cap for you, it removes it for everyone, and an advantage everyone holds is background noise. The buyer’s inbox is where the noise accumulates. This issue’s feature documents what attention does under load: the value of an inbound touch decays within the first hour, and most funnels already respond far outside that window.

Multiplying the number of touches does not widen the window. It narrows everyone’s.

A bad message was always bad. AI just makes it faster. If the binding constraint on your pipeline is relevance, and for most teams past Manual it is, adding volume tightens the constraint you were already losing to. The better allocation comes from reading the strongest evidence we have on where AI actually pays, and that evidence is unusually specific. The specificity is the strategy.

0113 min left

What the 758-consultant experiment actually did

In spring 2023, a nine-author team ran two pre-registered randomized field experiments inside Boston Consulting Group. The subjects were 758 individual contributor strategy consultants, about 7 percent of BCG’s global cohort at that level, volunteers who each gave five hours. The tool was GPT-4 as it stood at the end of April 2023, default settings, accessed through a company platform.

The design is what makes the paper a strategy document. Every subject first completed an unaided assessment task, a skill baseline. Then each was randomized into one of three conditions: no AI, GPT-4 access (the paper’s “GPT Only” arm), or GPT-4 plus a prompt engineering overview with videos and documents (the “GPT + Overview” arm), stratified on gender, location, tenure, openness to innovation, and native English status.

Then the researchers split the sample across the AI capability boundary, with no overlap. 385 consultants worked inside the frontier: 18 questions on conceptualizing a footwear product for niche markets, spanning creativity, analytical reasoning, writing, and persuasion, in a fixed 90 minutes. 373 worked outside it: a business case on a fictional company that combined spreadsheet financials with insider interviews, deliberately constructed so that GPT-4 would err in the analysis.

Figure 1. The jagged frontier: AI effects on 758 BCG consultants, by task type

Task type Effect of AI use As reported
Inside the frontier (18 realistic consulting tasks) More output, faster, better 12.2% more tasks completed, 25.1% faster, quality significantly higher
Outside the frontier (task selected to exceed AI capability) Worse decisions 19 percentage points less likely to produce correct solutions

The frontier is jagged: tasks that look similar in difficulty sit on opposite sides of what AI can reliably do. The strategic consequence for pipeline teams is argued in the article. Source: Dell'Acqua, McFowland, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon, Lakhani. Navigating the Jagged Technological Frontier. Field experiment with 758 BCG consultants.

  • Inside the frontier: quality, scored by two human graders per response and averaged into a composite across the 18 questions, with a parallel GPT-4-graded composite; completion, the share of the 18 questions finished; speed, seconds to reach the final question.
  • Outside the frontier: a binary, did the consultant deliver the accurate recommendation, plus a 1 to 10 recommendation quality score against an author-built rubric.

Inside the frontier, AI looked like a gift with no price. Consultants with AI “completed 12.2% more tasks on average, and completed tasks 25.1% more quickly” On human-graded quality, the trained arm scored 1.75 points above a control mean of 4.1, a 42.5 percent gain; the untrained arm gained 38 percent. Completion rose from 82.4 percent of tasks to about 93 and 91 percent across the two arms. A sweep of 108 regressions across every pre-registered quality variable found a significant positive effect in all of them.

Figure 2. Inside the frontier: 18 consulting tasks, three conditions (n = 385)

Outcome Control GPT Only GPT + Overview
Quality, human grades (composite, 1 to 10) 4.1 +1.56 (38%) +1.75 (42.5%)
Quality, GPT-4 grades 7.2 +1.22 (16.8%) +1.35 (18.6%)
Tasks completed (share of 18) 82.4% about 91% about 93%
Time to final question 5,023 s 27.6% faster 22.5% faster

Outside the frontier, the sign flipped. Consultants working unaided gave the correct recommendation 84.4 percent of the time. With GPT-4, 70 percent. With GPT-4 plus prompt training, 60 percent.

Key finding

The same tool, in the same caliber of hands, raised graded quality by up to 42.5 percent on one task and cut correctness by 19 points on another task that looked just as hard.

Two more results sharpen the allocation. First, the distribution of the gains: the bottom half of the skill baseline improved 43 percent against their own prior scores, the top half 17 percent. The dividend raises the floor before it raises the ceiling.

The dividend and the tax, in working numbers

Now translate it to a pipeline, with the transfer assumption stated out loud: drafting first touches, summarizing accounts, and writing persuasively are the same task family as the experiment’s inside-frontier questions, and a qualify-or-kill call that blends data with context has the structure of its outside-frontier case. The numbers below are ours, derived from the paper’s coefficients with the arithmetic shown.

Worked example · A 1,000-lead month: where the freed hours come from

  1. Baseline drafting time: 1,000 leads x 6 min = 6,000 min = 100 hours
  2. Speed gain, average of both AI arms: (1,129 s + 1,388 s) / 2 = 1,258.5 s saved
  3. Drafting time with AI: 100 hours x (1 - 0.251) = 74.9 hours
  4. Hours freed: 100 - 74.9 = 25.1 hours
  5. Coverage bonus: leads actually drafted: control completes 82.4%; AI average 82.4% + 10.05 pp = 92.5%. 1,000 x 0.824 = 824 vs 1,000 x 0.925 = 925

Per 1,000 leads, per month: 25 hours freed, about 100 more leads covered. Assumptions: 6 minutes per unaided draft is an operator assumption. The 25.1% speed gain and completion coefficients are from consultants on a product-innovation task.

Worked example · A 200-lead queue: the cost of AI on the wrong side of the frontier

  1. Unaided correct calls: 200 x 0.844 = 169 correct, 31 wrong
  2. With GPT Only: 0.844 - 0.139 = 0.705. 200 x 0.705 = 141 correct, 59 wrong
  3. With GPT plus prompt training: 0.844 - 0.245 = 0.599. 200 x 0.599 = 120 correct, 80 wrong

Per 200-lead queue: 28 to 49 extra wrong calls, and training made it worse. Assumptions: Correctness rates are applied to a 200-lead queue.

What you learned Inside the frontier, consultants with GPT-4 completed 12.2 percent more tasks, worked 25.1 percent faster, and scored up to 42.5 percent higher on human-graded quality. Outside the frontier, correctness fell from 84.4 percent to 70 and 60 percent, an average drop of 19 points, and the prompt-trained arm fell furthest.