The Same AI Made One Team Faster and Buried Another in Debt. The Difference Wasn't the AI
Our headline number is an average. The spread around it is larger than the number itself, and it tracks one variable almost nobody is measuring.
Two teams, one tool, opposite outcomes
An independent engineer described two situations to me, back to back, without noticing he'd just designed a natural experiment.
His current team builds specialized software. They introduced AI carefully, mostly for automation and generating tests, and made a deliberate choice: they asked the AI to generate kinds of tests they'd previously been missing. Pull requests got split smaller, review got a little longer, and product testing afterward surfaced fewer defects. The longer review wasn't a leak. It was an investment at the front that bought quality at the back.
Then he described his previous company. There, less experienced engineers, newer to the codebase, were reaching for AI to ship something quickly. The result was poor-quality pull requests that frequently had to be rewritten from scratch. Same category of tool. The code was bigger, messier, and slower to trust.
Same AI. One team it made better. The other it buried in rework. If the tool is the constant, the tool isn't the explanation.
The variable he isolated without meaning to
The difference between those two teams wasn't the model, the prompt, or the language. It was team maturity: how experienced the engineers were, and how deeply embedded they were in the codebase and the company.
This connects directly to something we've argued elsewhere: AI behaves like a new team member. And a new team member's effect depends entirely on the team receiving it. A mature, well-embedded team has standards, knows its own codebase, and can direct a prolific new contributor toward useful work, so AI adds value. A newer, loosely-embedded team, under pressure to just deliver, can't provide that direction, so the same prolific contributor generates volume the team can't absorb, and the AI adds debt.
The tool amplifies. What it amplifies is whatever the team already was. Give a disciplined team a force multiplier and you multiply discipline. Give a team that's cutting corners the same multiplier and you multiply the corner-cutting, faster, and at scale.
Why the average lies
Here's why this matters for anyone reading a productivity number. Our field study reported a 14% gain in implementation time, of which about 6% survived to delivery. That's a single organization's figure, and it's an average across its teams.
If maturity is the moderator, then the spread around that average is enormous: and more informative than the average itself. Two organizations reporting the same net 6% could be in completely different situations: one where every team is investing at the front and shipping clean, another where mature teams are carrying the number while newer teams quietly generate debt that hasn't surfaced yet. The mean hides the mechanism.
Worth being precise about the organization behind our own number, because it sharpens the point rather than softening it. That 14-to-6 came from a large, mature product organization: a B2C platform serving millions of users daily, a codebase more than a decade old, teams averaging well over ten years of experience. By the maturity argument, that's a best-case profile, exactly the kind of team that should handle AI well. And the leak happened anyway. Which makes it more telling, not less: if half the gain evaporates even for a seasoned team that knows its codebase cold, less-embedded teams aren't doing better than that. They're doing worse, and often not seeing it yet.
The trap in the obvious fix
The same engineer raised the standard defense against the debt problem, and then, in passing, exposed why it backfires.
In large organizations, the usual guardrail is a mandate: every change must be reviewed by at least one senior engineer. This is a rational response to the newer-team problem, it protects quality by putting experienced eyes on everything. But look at what it does at scale. It routes the entire review load through your most experienced, most expensive, most scarce people, and turns them into a bottleneck.
So the cure for one problem manufactures another. The fix for immaturity (senior review on everything) creates a coordination leak (seniors as the constraint). This may be exactly why AI efficiency often looks worse in large organizations than on a small, seasoned team, not because AI works less well there, but because guarding against inexperience costs senior time, and that cost is itself a leak. The mature small team doesn't need the mandate, so it doesn't pay the tax. The large organization needs it, and pays it on every change.
Two high-maturity teams, then, can still leak differently, and the difference between our large, seasoned organization and a small, seasoned one probably lives in scale and coordination, not maturity, since both have maturity in abundance. Which suggests maturity and scale are two separate dials: maturity protects quality (fewer bad pull requests), but it doesn't protect against the coordination cost that grows with size. You can be mature enough to avoid the debt and still large enough to lose the gain in the seams.
What we don't know
Two open questions, both measurable, neither measured well yet.
First: how much of the variance is maturity versus scale? They're confounded in most real data: large organizations tend to have a mix of maturities, small ones tend to be more uniform. Separating them would tell you whether the intervention is "raise team maturity" or "reduce coordination load," which are very different programs.
Second: does experience protect the gain, or just the quality? The medical-device team's maturity clearly protected them from generating debt. It's less clear it protected them from the coordination and review costs that eat the gain in the first place. Maturity might buy you clean code and still not buy you delivered value, those may be independent.
The practical version
If you're reading an AI-productivity number, your own or anyone's, the average is the least interesting thing about it.
- Ask what the spread looks like, segmented by team maturity. A 6% mean composed of mature teams at +15% and newer teams at −5% is a completely different situation from a uniform 6%, and it demands a completely different response.
- Watch for debt that hasn't surfaced. Newer teams generating volume with AI can look productive for a quarter or two before the rewrites arrive. Flat or improving metrics now don't rule out a debt problem maturing quietly underneath.
- Separate your two dials before you pull either. If your leak is immaturity, the answer is upstream: standards, embedding, direction during the work. If your leak is the senior-review bottleneck that immaturity forced you into, adding more mandated review makes it worse, not better.
The organizations getting real value from AI aren't the ones with the most adoption. They're the ones mature enough to point a prolific new contributor at the right work — and self-aware enough to know that the guardrails protecting their newer teams are quietly taxing their most experienced ones. The tool multiplies what you already are. The number on the dashboard is the average of that. The truth is in the spread.
