> ## Content Index
> Fetch the complete content index at: https://debugly.dev/llms.txt
> Use this file to discover other available public pages before exploring further.

# Debugging Is Now 42 Percent of the Working Week
- URL: https://debugly.dev/debugging-is-now-forty-percent-of-the-week/
- Published: 2026-10-08T13:30:00.000Z
- Updated: 2026-10-10T13:36:28.000Z
- Description: A new survey of 300 developers and engineering leads finds teams spend 9.8 hours a week writing code and 16.9 hours debugging it, and 79 percent say the…
- Author: Rohit Bhadani
- Tags: Debugging, AI, Process

Between the code and the paradox sit 6 quiet assumptions, and the incident is the story of one of them failing.

A survey published this week puts a number on something most engineering teams have felt for a year without being able to articulate. Coleman Parkes polled 300 software developers and engineering leads for a report commissioned by Undo, and the headline finding is that teams now spend an average of 9.8 hours a week producing code and 16.9 hours a week debugging issues found during development or hitting customers in production. Debugging is 42 per cent of the average working week.

If you run or work on an engineering team, the numbers are worth reading carefully, because they describe a bottleneck that moved rather than a productivity gain that failed.

## What the survey actually found

The numbers that stand out are not the productivity ones. They are these.

Thirty-five per cent of AI-generated code reaches production before the team has fully understood what it does. Ninety-three per cent of teams have experienced AI hallucinations leading to an incorrect diagnosis of a problem in code, and 18 per cent experience that multiple times a month. Eighty-one per cent have had a production incident or outage affecting users at least once in the previous six months, with 14 per cent seeing them multiple times a month. Ninety-one per cent have had test escapes, serious defects or badly optimised code reach production at least once.

And then the number that should end a lot of arguments: 79 per cent of engineering leaders say AI agents generate code significantly faster, but the shift of effort toward debugging and unpicking that code means the overall release cycle is no faster than before.

Greg Law, Undo's CEO, put it plainly: engineers lose days trying to unravel what went wrong and why, working with code that is almost, but not quite right. Agents are excellent at writing large amounts of code quickly and considerably less capable of debugging it.

Related research points the same direction with more granularity. CodeRabbit's analysis found AI-created pull requests carried 75 per cent more logic and correctness errors, roughly 194 per hundred PRs, along with security defects at 1.5 to 2 times the human rate and excessive I/O operations about eight times more often.

## The bottleneck did not disappear, it moved

This is the part that matters, and it is not a story about AI being bad. It is a story about where the constraint in a software team actually lives.

Writing code was never the whole job. It was the visible middle of three buckets: design, implementation, and understanding what the running system is doing. Compressing the middle does not remove the other two, and it makes the third one harder, because the volume of code that needs understanding went up while the amount of it written by a person who understood it went down.

The result is the comprehension gap the survey names explicitly. Thirty-five per cent of generated code reaches production before anyone fully understands it, and 80 per cent of respondents say coding agents struggle with difficult problems in complex codebases. Around a third of teams now restrict agents to comprehension and debugging work only in straightforward codebases, which is an admission that the tool's useful range is narrower than its adoption.

Debugging is harder than writing in a specific way that compounds here. A writer knows what they intended. A debugger has to reconstruct intent from artefacts, and the artefacts are now produced by something that had no intent, only a plausible next token. That is the gap between code that is wrong and code that is almost right, and almost right is much more expensive to falsify.

## What I would actually change

**Instrument for comprehension, not just for uptime.** If a third of your code reaches production ununderstood, your observability needs to answer what the code does at runtime, not only whether the service is up. Structured traces that show the decision path are worth more now than another dashboard.

**Make the review check the assumption, not the syntax.** CodeRabbit's finding that logic and correctness errors dominate, and that they are the easiest to miss because they look reasonable, means the review question is not does this compile but what is this assuming about the system. That question has to be asked explicitly or it does not get asked at all.

**Stop counting lines shipped as the productivity metric.** If 79 per cent of leaders say the release cycle is no faster, the metric you are optimising is not the one that moved. Cycle time from decision to production, and incident rate, are the honest measures.

**Budget debugging time as real work.** Sixteen point nine hours a week is not slack and it is not a failure of discipline. It is the job. A plan that assumes the writing time shrank and nothing else changed is planning against a team that does not exist.

**Keep at least one person who can explain the system.** Not every file, but the shape of it. The failure mode the survey describes is a team that can produce code and cannot reconstruct why the code does what it does, and that is the failure mode that turns an incident into a week.

This is the same problem from the other direction as [reviewing code the author cannot explain](https://debugly.dev/reviewing-code-the-author-cant-explain/), where the missing artefact was not the diff but the reasoning behind it. The survey says the population of code with no author who can explain it is now about a third of what ships.

## The rule I keep

Compressing the time to write code does not compress the time to understand what the running system does, and the second one is now the larger half of the week. Instrument for comprehension, review the assumptions rather than the syntax, measure cycle time rather than volume, and budget debugging as the work it has become.

Sixteen point nine hours a week debugging against 9.8 hours writing is not a productivity story, it is a statement about where the constraint moved. The tools did exactly what they promised, which was more code faster, and the teams that are not seeing a faster release cycle are not failing to adopt them. They are discovering that the code was never the bottleneck, and the thing that was is now thirty-five per cent larger and written by something that cannot explain itself.