AIAIBlog.com.my
International AI News10 October 2026 · 5 min read

AI Coding Agents Write More Code, Not More Software — Review Is the Bottleneck

A new study reported by Ars Technica finds that human code review absorbs most of the productivity gains from AI coding agents, and that changes how Malaysian businesses should buy and deploy them.

AI Coding Agents Write More Code, Not More Software — Review Is the Bottleneck
AIAI Summary

A study reported by Ars Technica on AI coding agents found that these tools do exactly what they promise — they produce a lot more code, faster. What they do not do is deliver proportionally more finished software, because the gains get absorbed by a human review bottleneck. Every line an agent writes still has to be read, checked, tested, and approved by a person, and that approval step does not speed up just because writing did. For Malaysian CIOs and SME owners, the practical lesson is that buying coding AI licences is only half the investment. The other half is fixing the review, testing, and quality pipeline downstream — otherwise you are paying to create a backlog.

AI Summary

A study reported by Ars Technica on AI coding agents found that these tools do exactly what they promise — they produce a lot more code, faster. What they do not do is deliver proportionally more finished software, because the gains get absorbed by a human review bottleneck. Every line an agent writes still has to be read, checked, tested, and approved by a person, and that approval step does not speed up just because writing did. For Malaysian CIOs and SME owners, the practical lesson is that buying coding AI licences is only half the investment. The other half is fixing the review, testing, and quality pipeline downstream — otherwise you are paying to create a backlog.

Key Takeaways

  • AI coding agents successfully accelerate code generation, but total software delivered does not rise in step, because human review becomes the constraint in the pipeline.
  • Lines-of-code is now a misleading productivity metric — more machine-written code can mean more review workload, more surface area for bugs, and more maintenance debt, not more shipped features.
  • ROI calculations built on "developers write X times more code" will overstate savings; measure lead time from idea to deployed feature instead.
  • The real productivity play is not faster writing — it is faster verification: automated testing, tiered AI-assisted review, and smaller change sizes.
  • For Malaysian software houses and GBS centres priced per man-day, this shifts competitive pressure from writing cost to review and quality capacity.

What Happened

Ars Technica reported on a study examining what actually happens when teams adopt AI coding agents — the autonomous or semi-autonomous tools that write, edit, and complete code from natural-language instructions. The headline finding is blunt: agents generate substantially more code, but the amount of finished software coming out the other end does not grow proportionally. The reason, according to the study, is that the efficiency gains get "absorbed" by a human review bottleneck.

To understand why, look at how software actually gets built. A developer (or an agent) writes code, then opens a "pull request" — a proposal to merge that code into the main product. Before merging, another human reads the change, questions the logic, checks for bugs and security issues, and either approves or sends it back. In most professional teams, this review step is mandatory, and for good reason: unreviewed code is how outages, data leaks, and expensive rework happen.

The study's finding follows a simple mechanical logic. If an agent cuts writing time dramatically but reviewers can only clear a fixed number of changes per day, the team's total output is capped by the reviewers, not the writers. Generating more code just lengthens the queue. The writing stage got faster; the pipeline did not.

Why It Matters

This is the theory of constraints applied to software, and it is the same logic behind a familiar Amdahl-style rule: speed up one stage of a multi-stage process and the overall speedup is limited by the stages you did not touch. If a factory doubles the speed of its assembly line but the quality inspection station at the end still handles one unit an hour, the factory ships one unit an hour. Widening one lane of a highway only moves the jam to the next toll plaza.

The commercial implication is significant. The past two years of enterprise buying decisions have been driven by vendor claims about generation speed — how much faster developers write with an AI pair programmer in the editor. If the study's finding holds broadly, those claims are true but incomplete. The licence pays for itself only if the downstream review and testing capacity expands to match. Otherwise, the money buys longer queues and frustrated senior engineers, which is worse than no tooling at all because it looks like progress on dashboards that count output volume.

There is also a metrics problem hiding here, and it deserves attention from anyone running a development budget. Code volume was always a shaky measure of productivity, but with agents it becomes actively deceptive. More code means more lines to read, more paths to test, more duplicated logic to reconcile later. The study's core message — more code, not more software — is really a warning about measuring the wrong thing. What matters is delivered, working features and defect rates. That is what a paying customer or an internal stakeholder experiences.

What This Means for Malaysia

Malaysia's software economy runs on exactly the kind of work this study touches. The GBS and IT services hubs around Cyberjaya and the Klang Valley, the in-house development teams at banks, GLCs, and telcos, the growing fintech cluster, and the manufacturing software groups supporting Penang's electronics corridor — all of these are adopting AI coding tools, and most are doing it under a productivity rationale. This study says the rationale needs a rewrite: the tools help, but the payoff shows up only where review and testing capacity scales too.

For Malaysian software houses and outsourcing shops, there is a pricing threat worth thinking through now. Local vendors commonly price per man-day or per developer. If clients read headlines saying "AI makes coding five times faster," they will push for output-based pricing. But if the study is right that total output does not scale with writing speed, vendors who agree to output-based contracts without first fixing their review pipeline will squeeze their own margins. The honest position in the next client negotiation is: writing got cheaper, verification did not, and here is our measured delivery rate.

For talent, the signal is about role design. Malaysia produces thousands of computer science graduates a year, and MDEC's digital skilling agenda has pushed AI fluency hard. The junior developer job is quietly changing. The "write boilerplate code" tasks that used to train juniors are increasingly done by agents. What the study shows is that the scarce skill is now on the review side: reading code critically, spotting flawed logic, testing edge cases. Universities, bootcamps, and

Sources & References

AIBlog summarises and analyses published information. We do not reproduce full source text. Analysis is editorial and not financial or legal advice.

Related articles

Get Malaysia's AI intelligence every morning

Daily digest by email and on Telegram. Written for Malaysian business readers.

Daily AI intelligence
From RM5/month
Subscribe