
A systematic review of eight controlled studies comparing human, AI, and human-plus-AI performance on real work tasks reveals that pairing humans with AI tools does not reliably outperform either working alone—a pattern corroborated by a meta-analysis of 106 prior experiments.
The benefit or harm depends on task type: AI excels on some tasks while degrading performance on adjacent ones that look similar to humans, a pattern the research calls the "jagged frontier." This challenges the widespread corporate assumption that adding AI to any workflow will boost productivity.
What happened
A structured review of eight controlled empirical studies comparing unassisted human performance, AI-alone performance, and human-plus-AI performance on real work tasks found that the combination does not consistently outperform either party working alone. The studies span clinical diagnosis, consulting, customer support, writing, and software engineering. A separate meta-analysis of 106 prior experiments (370 effect sizes) in Nature Human Behaviour reached a closely aligned conclusion: human-AI combinations underperform the better of either alone on average.
Why it matters
The assumption that adding AI to a skilled worker's workflow will improve outcomes appears to be false in many cases. A randomized clinical trial published in JAMA Network Open in October 2024 found that physicians using ChatGPT Plus (GPT-4) plus conventional resources showed no statistically significant accuracy gain over physicians using conventional resources alone (95% CI −195 to 31 seconds, p=0.20 on speed). A follow-up trial in Nature Medicine in 2025 by an overlapping author team found a +6.5 point gain for physician-plus-GPT-4 on a different task design, suggesting outcomes depend heavily on task specifics that researchers do not yet fully understand.
What to watch
The research identifies a "jagged frontier" pattern: AI's performance does not degrade gradually as tasks become harder (as human performance does), but rather shows extraordinary competence on some tasks and unreliability on adjacent ones that appear similar. In a field experiment at Boston Consulting Group, 758 consultants on tasks inside GPT-4's competence zone completed 12.2% more tasks and 25.1% faster at higher quality, while on tasks outside that zone, AI-assisted consultants were 19% less likely to produce a correct solution.
The research directly challenges a foundational assumption underlying most corporate AI deployment strategies: that human workers equipped with AI tools will outperform either humans or AI working alone. The eight studies identified through a structured literature search across Google Scholar, Semantic Scholar, PubMed, SSRN, NBER, arXiv, the ACM Digital Library, and IEEE Xplore represent the strongest available designs in this domain, selected using transparent criteria rather than arbitrary curation. Crucially, independent confirmation arrived through Vaccaro, Almaatouq, and Malone's 2024 meta-analysis in Nature Human Behaviour, which pooled 106 experiments (370 effect sizes) and found a task-dependent pattern: human-AI combinations underperform the better of either alone on average, with the effect reversing between decision-making tasks (where combining hurts) and content-creation tasks (where combining helps). This convergence suggests the pattern is not an artifact of the eight studies chosen but reflects what the broader empirical literature actually shows.
The medical studies illustrate the puzzle most directly. In the October 2024 JAMA Network Open trial, GPT-4 alone scored a median of 92% accuracy on clinical vignettes, while physicians using conventional resources scored lower, and physicians using GPT-4 plus conventional resources showed no statistically significant improvement over the control group using conventional resources alone. However, a follow-up trial from the same research team using a different task design found physicians using GPT-4 did achieve a +6.5 point gain (95% CI 2.7–10.2, p<0.001). The researchers acknowledge that the answer "depends on task design in ways not yet fully understood, even to the researchers running near-identical experiments."
The Harvard Business School field experiment with Boston Consulting Group introduces the concept of the jagged frontier to explain this inconsistency. Across 758 BCG consultants completing 18 realistic consulting tasks, AI-assisted performance was dramatically task-dependent: on tasks inside GPT-4's competence zone, consultants completed 12.2% more tasks 25.1% faster at measurably higher quality; on tasks outside that zone, AI-assisted consultants were 19% less likely to produce a correct solution than the unassisted control. This pattern—extraordinary on some tasks, unreliable on adjacent ones—suggests that simply adding AI access without understanding where its performance boundary lies may actively harm outcomes.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic announced it has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generat…

OpenAI's Daybreak Red (GPT-5.6 Cyber) and Daybreak Blue (GPT-5.6 Sol) are now available to eligible customers…

Leica Biosystems announced multiple FDA 510(k) clearances on August 11, 2026, including Aperio iQC DX software…

Anthropic is embedding imperceptible, machine-readable watermarks into text generated by Claude models release…

Gregory Kurtzer, founder of Rocky Linux and co-founder of CentOS, released OpenWALDO, an open-source project d…

A directory of 27 creators—writers, editors, software engineers, musicians, and designers—has been published…

The AI news that matters, in one minute each morning.
Sign up free