A survey conducted by independent research firm Coleman Parkes on behalf of Undo, a company focused on scaling AI-powered root-cause analysis, found that while AI coding agents have accelerated code generation, they have shifted the primary bottleneck to debugging, code comprehension, and maintenance.
The report surveyed 300 senior engineer leaders responsible for delivering mission-critical software, most of whom working with C/C++. The survey focuses specifically on mission-critical codebases where code "must be understood", with respondent identifying the most demanding task as "understanding what that code does, how it affects existing codebases, and debugging it when an application doesn’t behave the way it’s expected to".
In those environments, teams spend an average of 9.8 hours per week producing code, but 16.9 hours per week debugging issues identified during development or encountered by customers in production, accounting for 42% of the average working week.
Along with debugging getting more relevant, another critical dimension is emerging as a challenge, code comprehension:
Now that AI is generating most of the code being produced, engineers no longer have the inherent understanding they used to. That makes it easier for defects to escape, and when something inevitably goes wrong, nobody has the knowledge to trace the failure back to its root cause.
Due to the acceleration in code generation brought by AI agents, the survey found 35% of generated code reaches production before the team has fully understood it. Moreover, 80% of respondents said that coding agents struggle to solve difficult problems in complex codebases. As a result, approximately one-third of teams "use AI agents for comprehension and debugging only in straightforward codebases", while relying on additional techniques to build sufficient confidence when working with more complex systems.
Other significant problems reported by surveyed teams include production incident or service outage affecting internal users or customers (81% experienced this at least once in the previous six months, with 14% experiencing them multiple times per month), incorrect root-cause or issue diagnosing due to hallucination (93% at least once, with 18% multiple times per month), and test escapes, serious defects or poorly optimized code entering production (91% at least once, with 8% experiencing these issues multiple times per month).
Overall, 79% of engineering leaders say that AI agents can generate code significantly faster, but that the resulting shift in effort toward debugging and "unpicking" AI-generated code means the overall release cycle is "no faster than before".
Greg Law, founder and CEO of Undo, summarized the survey findings by noting that engineers "lose days trying to unravel what went wrong and why" with "code that's almost, but not quite right" and that "while agents are great at writing reams of code quickly, they're less capable at debugging it".
Since the launch and widespread adoption of AI coding agents, the software engineering community has extensively debated their benefits, limitations, and how best to use them. One recurring concern is how to manage agent speed when it outpaces humans' ability to review the generated code. This has led to several popular approaches, including the test-first red-green loop, and automated fallbacks, and others. The broader consensus, however, is that AI agents do not fundamentally change the nature of software engineering, which has never been solely about coding, but about understanding constraints, making trade-offs, and ensuring that the resulting system behaves as intended.