PRs grew from 44 to 72 lines, and gains topped out at 7.7%
DX's data from roughly 200,000 engineers across 400+ orgs shows AI is reshaping software delivery, but not the way the hype suggests. Deployment frequency and PR size are climbing while change confidence falls, and even top-performing teams saw nowhere near 2x productivity gains — because, as Justin Reock argues, code generation was never the real bottleneck.
- Who/what — Justin Reock, deputy CTO at DX, presented quarterly research drawn from the company's platform, built by people who worked on Dora metrics, SPACE, and the DevEx framework, covering about 200,000 engineers.
- Deployment frequency — Dora's deployment frequency metric is steadily rising but tapering off, with North America trending up while Europe pulled back last quarter.
- Perceived vs real speed — Perceived delivery speed rose only about 4.5% over a year of heavy AI investment, compared with the flawed METR study where measured productivity fell 19% but perceived productivity rose 20%, a 40% gap.
- Change failure rate — Change failure rate has become highly volatile across companies studied, with some firms rising as much as 2 percentage points against an industry benchmark of about 4%, meaning up to 50% more defects shipped.
- Maintainability vs confidence — Code maintainability perception rose almost 4%, but change confidence fell 6%, a divergence DX says reflects developers trusting AI-modified code less even as it's easier to read.
- PR size growth — Average PR size grew from about 44 lines to 72 lines over the year studied, and sentiment around incremental delivery dropped 10%, one of the biggest declines in any DX experience metric tracked.
- Junior vs senior use — Junior engineers use AI the most and burn more tokens per use case than seniors, but staff+ engineers save roughly the same amount of time while using fewer tokens; smaller companies lead in time savings.
- Measurement framework — DX's AI Measurement Framework tracks utilization (daily/weekly active users), impact, and cost, treating utilization as the starting point before correlating cohorts against trusted metrics like PR cycle time.
- Platform AI readiness — Reock frames 2024 as the year of coding assistants, 2025 as the year of agents, and now as the year orgs realize infrastructure wasn't ready, pointing to documentation, modular code, and reliable CI as readiness signals.
- Agent experience — DX now collects qualitative feedback directly from agents about steering, context, and feedback cycles to measure how effectively teams work with AI, broken down by use case.
- Why gains are modest — Code generation covers only about 14-16% of the value stream even with perfect AI output; DX found a median 7.7% increase in PR throughput and 13% average from November 2024 to February, with top performers around 70% but nobody reaching 2x.
- Case studies — Morgan Stanley's DevGen AI agent, which converts legacy COBOL/Natural/Perl code into PRDs, saves about 300,000 hours a year; Zapier cut standups to twice a week, onboards engineers in about two weeks, sees 15% more value per engineer, and is hiring more; Spotify built an SRE agent that surfaces runbook remediation steps during incidents.
In their words
Uh let me prepare you. Most of you are sitting down. It's very volatile the impact that we've seen on quality over this period of time.4:37

I trust the outputs less. I'm more afraid now of breaking things than I was a year ago.6:35

Even if engineers are getting like 100% accurate instant code coming from the models, which they are not, you would still only be attacking anywhere from maybe 14 to 16% of the overall value stream.14:52

our median increase in the study that we did from November to to 2024 to February this year only found about a median 7.7% increase in this velocity metric, a 13% average, but even our top performers were in the 70% range. Nobody hit 2x, nobody hit 5x, nobody hit 10x.15:12

Disclosure · Justin Reock works for DX, whose platform and research products (quarterly reports, measurement framework) are the subject of the talk.
One thing to add — One thing to add — the talk leans on DX's own proprietary metrics and framework, so the figures (7.7% median gain, 44-to-72-line PR growth) should be read as DX's measurement methodology rather than independently verified industry data. It's also worth noting Reock explicitly says the change-failure-rate volatility pattern predates AI and isn't purely causal, a caveat easy to miss amid the headline numbers.</note> </invoke>
One thing to try tonight
Pull your team's last quarter of PR data and check two numbers tonight: average lines changed per PR and change failure rate trend; if PR size is creeping up while failure rate swings wildly, that is the pattern DX flags as the AI effect.