
New data shows AI engineering productivity gains split into three distinct tiers: most companies using basic AI IDEs see only 20–46% improvements, companies building orchestration platforms around agents (Replit, NVIDIA, Amplitude, Anthropic) achieve 2.5–3x gains, and software factories treating agents as first-class units reach 8x+ efficiency. The difference is not the underlying model but the organizational operating discipline built around it—suggesting that companies expecting 2–3x productivity gains from AI alone are missing the structural changes required to capture those returns.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Analysis of six months of real-world AI engineering data shows productivity outcomes split into three tiers. The first (basic AI IDE distribution) yields 20–46% gains; the frontier (companies like Replit, NVIDIA, Amplitude, Anthropic building orchestration around agents) delivers 2.5–3x improvements; and software factories (Nubank with Devin, Factory.ai deployments) reach 8x+ efficiency gains.
Why it matters
Most engineering leaders expected 2–3x gains from AI but are landing closer to 30% because they distribute tools without redesigning workflows. The gap is not the model itself but the operating discipline around it—companies that build agent orchestration (spawning worker agents across GitHub, Linear, Slack) and treat agents as first-class organizational units unlock dramatically higher returns. This reframes AI adoption from a tool problem to an organizational design problem.
What to watch
Nubank achieved an 8x improvement in engineering efficiency and a 20x cost reduction using Devin for large-scale refactoring; Goldman Sachs is piloting Devin alongside 12,000 human developers and estimates agentic AI could deliver 3–4x the rate of prior tools. The frontier is shifting from incremental tool adoption to agent-native team structures.
Over the last six months, real-world deployments have revealed a sharp stratification in AI engineering productivity gains—not by model quality, but by organizational design. NVIDIA reported a 3x increase in committed code across 30,000 developers with bug rates remaining flat. Amplitude tripled weekly production commits, with an AI agent now ranking as a top-three contributor to the codebase. Anthropic measured a 2.5x increase in code written per engineer since adopting Claude Code internally, with quality stable. Replit doubled its team and tripled per-engineer output over the same period, while review times, reversions, and incidents all stayed flat.
Yet these outliers mask a very different baseline. Engineering leaders entered the AI era expecting 2–3x productivity gains but are landing closer to 30%, according to Augment Code. Faros's telemetry across 22,000 developers confirms this: engineers completed epics 66% faster, but bugs per developer increased by 54%. Google's randomized controlled trial put the productivity gain at 21%, close to GitHub's 24%. This modest outcome represents the default when companies simply distribute an AI IDE and change nothing else about their workflows.
The frontier tier—exemplified by Replit, NVIDIA, Amplitude, and Anthropic—has engineered a structural difference. These companies built harnesses around the model, orchestrating agents that share context across GitHub, Linear, and Slack and escalate to engineers only for judgment. Replit's approach is illustrative: every employee gets a manager agent that spawns worker agents in loops. According to Amjad Masad, Replit's CEO, "Our internal agent outperformed a seven-figure SaaS tool in security testing and incident triage at one-tenth the cost." The result is a 3x productivity tier. Human PR review time dropped 30%, complex support handling time dropped 60%, and total code contribution rose 5.8x—all while quality metrics remained stable.
Beyond the frontier lie software factories, where agents operate as end-to-end producers of software. Cognition's Devin refactors monolithic codebases autonomously. Factory.ai is deploying software factories at NVIDIA, Adobe, Blackstone, and EY. Nubank achieved an 8x improvement in engineering efficiency and a 20x cost reduction using Devin for large-scale refactoring, according to Contrary Research in January 2026. Goldman Sachs is piloting Devin alongside 12,000 human developers and publicly estimates that agentic AI could deliver 3–4x the rate of prior tools. These deployments represent a fundamentally different model of AI adoption—not augmentation but automation at scale.
The data converges on a single insight: the productivity gap is not the model but the operating discipline built around it. Most teams should expect to migrate from 20% productivity gains with basic tools to 2.5–3x gains with frontier-tier orchestration, provided they redesign their workflows to place agents at the center of their engineering practice.
The analysis reveals a structural gap between expectation and reality in AI engineering adoption. Engineering leaders entered the AI wave expecting 2–3x productivity gains, but the median outcome from distributing an AI IDE without operational redesign is closer to 30%. This gap is not random—it reflects a fundamental misalignment between tool capability and organizational readiness. The data from Faros (22,000 developers), Google's randomized controlled trial (21%), and GitHub's measurements (24%) all converge on a narrow band of modest gains when AI is introduced as a supplementary tool.
The frontier tier—companies like Replit, NVIDIA, Amplitude, and Anthropic—has closed this gap by building orchestration and management layers around AI agents. Rather than asking engineers to use AI as a helper, these companies have architected workflows where agents spawn and coordinate other agents, share context across GitHub, Linear, and Slack, and escalate to engineers only for judgment calls. The result is a 2.5–3x step function in productivity. Critically, these companies report that review times, reversions, and incidents stay flat or decline, suggesting that the operating discipline prevents the quality degradation seen in the basic tier.
The third tier—software factories—treats agents as first-class organizational units capable of autonomous end-to-end tasks like refactoring monolithic codebases. Nubank's 8x efficiency gain with Devin and Goldman Sachs' pilot alongside 12,000 human developers indicate that at scale, agents operating under tight constraints and clear ownership can deliver returns an order of magnitude higher than both the basic and frontier tiers. The key implication is that AI productivity is not determined by model capability alone but by how thoroughly a company has redesigned its engineering workflows and decision-making structures to place agents at the center rather than the periphery.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack