Feature

Why 741% More AI Code Only Yields 30% More Shipped Software

Source · The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI
AI Engineer · Laurie Voss · Arize AI (Head of Developer Relations) · 2026-09-30 uploaded · 25min

한국어·English

A new wave of data suggests the coding bottleneck has moved from writing code to reviewing it — and "review harder" cannot fix it. Laurie Voss walks through evidence from METR, Cognition, OpenAI, Cursor and GitHub showing that human review is being redesigned as a system rather than eliminated.

  • The core stat — A study tracking over 100,000 GitHub developers found that those using autonomous agents wrote 741% more code but shipped only 30% more software, with the authors naming review as the bottleneck.
  • Generation is solved — Stripe and Anthropic's Fable migrated a 50-million-line Ruby codebase in a single day, a job estimated at two months, and Bun migrated over a million lines of Zig to Rust in six days.
  • Reviewing harder fails — A decade-old Cisco study of 2,500 reviews and 3.2 million lines of code found reviewer effectiveness collapses past 400 lines per sitting or 450 lines per hour, meaning a 10,000-line agent PR would need three to four working days of real human review.
  • Skip review entirely — OpenAI built an internal product from an empty repo with no manually written code, reaching about a million lines and 1,500 merged PRs using three engineers, with humans not required to review.
  • METR's finding — METR had four open-source maintainers re-review PRs that had already passed SWE-bench, and only about half were judged actually mergeable, with failures tied to code quality and quiet breakage rather than test correctness.
  • Cognition's benchmark — Cognition's FrontierCode, built from 150 tasks across 20+ maintainers' repos, showed Fable 5 scoring 88% on SWE-bench Pro but only 29% on real mergeability, and GPT-5.5 scoring under 6%.
  • Benchmark becomes training — Investor Sarah Guo argues a true mergeability benchmark would instantly become a training signal for frontier models, echoing how OpenAI's CriticGPT in 2024 made models better almost immediately once human-beats-model review became a signal.
  • Production-scale automation — GitHub's Copilot reviewer has done 60 million reviews and accounts for more than one in five reviews on GitHub, while Cursor runs eight shuffled review passes per diff to filter false positives, a multi-pass trick a PKU team found raises review quality by up to 44%.
  • Cursor's design detail — Cursor engineers had to explicitly instruct their reviewer model to be suspicious by default, since it defaulted to approving code that merely looked fine, and its resolution rate rose from 52% to over 70%.
  • Humans-out experiments — Nicholas Carlini's 16 agents built a C compiler in Rust across roughly 2,000 sessions that compiled the Linux kernel with no human reviewing code, though humans wrote the test harness.
  • Bun's hidden cost — Bun's million-line Zig-to-Rust port passed 99.8% of its test suite but contains 13,044 unsafe blocks versus roughly 74 in a comparable human-written Rust codebase of that size.
  • Reversal on stage — Dexter Horthy, who spent six months telling people not to review agent code, retracted that advice in March, saying his team had to rip out and replace large parts of a system built that way.
  • Automated review can be fooled — A March study found vulnerable code dressed in an innocent commit message fooled autonomous review agents 88% of the time, versus only 35% for human reviewers, and Anthropic's own security reviewer warns it is not hardened against prompt injection.

In their words

the developers who turned on autonomous agents wrote 741% more code but only 30% more software shipped1:30
Laurie Voss slide · The Death of the Code Review: What the Data Actually Says —  1:30
Laurie Voss slide · 1:30 · AI Engineer
reviewers stop finding defects effectively if they try to read more than 400 lines of code in one sitting and their effectiveness completely falls off a cliff3:38
Laurie Voss slide · The Death of the Code Review: What the Data Actually Says —  3:38
Laurie Voss slide · 3:38 · AI Engineer
you shouldn't be prompting coding coding agents anymore. You should be designing the loops that prompt your agents4:44
Laurie Voss slide · The Death of the Code Review: What the Data Actually Says —  4:44
Laurie Voss slide · 4:44 · AI Engineer
I was wrong. Please please read the code. We tried not reading the code for like six months. It did not end well. We had to rip out and replace large parts of that system.18:46
Laurie Voss slide · The Death of the Code Review: What the Data Actually Says —  18:46
Laurie Voss slide · 18:46 · AI Engineer
SWE-bench Pr 88% Frontier Cod 29%
SWE-bench vs 실제 머지 가능성 — Cognition의 Frontier Code 벤치마크 결과(발표 중 인용)
자동 리뷰 에이전트 88% 인간 리뷰어 35%
프롬프트 인젝션 성공률 — 2026년 3월 연구에서 인용된 프롬프트 인젝션 공격 성공률

Disclosure · Voss is Head of Developer Relations at Arize AI, which sells production monitoring/observability tools for AI systems — the kind of "last reviewer standing" he recommends at the end of the talk.

One thing to add — One thing to add — the talk leans heavily on vendor-published numbers (Cursor's own resolution-rate claims, OpenAI's self-reported zero-human-review product) that haven't been independently audited, so they should be read as marketing-adjacent data points rather than neutral benchmarks. Voss is upfront about this tension himself when he notes OpenAI never open-sourced its no-review product as proof the method generalizes.

One thing to try tonight
Tonight, pick one recent PR your team merged without full review and run it through a second independent pass (a different model, prompted to be "suspicious by default" rather than to assess whether the code "looks fine") to see if it flags anything the first pass missed.