An AI Agent Cannot Debug a System You Cannot Describe

Share
An AI Agent Cannot Debug a System You Cannot Describe. Abstract opinion illustration in orange and dark grey on debugly.dev

Here is the thesis. Coding agents are excellent readers of code and useless guessers of intent, and debugging is mostly the second. The failures an agent can fix are the ones where the code contradicts itself, a crash, a type error, a test that encodes the expectation. The failures it cannot fix are the ones where the code is internally consistent and externally wrong, and the definition of wrong lives outside the repository, in the intent that was never written down. Most production debugging is the second kind, and that is why the agent, given the whole codebase, still asks you what the system was supposed to do, and you cannot answer quickly, and that inability is the actual defect.

I am not arguing agents are useless. I am arguing that their usefulness is bounded by the describability of your system, and that bound is a property of your documentation culture, not of the model.

What an agent can see

An agent given the repository sees the code as written, the tests as written, the comments, the commit messages, and it can hold more of it at once than you can, and trace it faster. For defects that are contradictions inside that corpus, it is superb: the null that the test forbids, the race the linter misses but the interleaved reads reveal, the off by one against the asserted boundary. It can also execute, propose, and verify against the tests, closing the loop a human walks slowly.

The common property of these successes is that the correctness criterion is in the corpus. The test is the specification, and the agent optimises against it. Where the criterion is executable, the agent has a compass.

What an agent cannot see

The correctness criterion for most production systems is not in the corpus. It is the business rule that decided the rounding, the compliance constraint that shaped the state machine, the incident two years ago that explains why this path retries twice, the merchant's expectation that the total and the invoice agree. None of that is in the code. It is in heads, in dead Slack threads, in a ticket system that lost the context when the ticket closed.

An agent debugging such a system has the code and no compass. It can find code that is surprising to it, but surprising to the model is not the same as wrong, and the model's priors about what code should look like are not your system's specification. So it produces plausible refactorings of the wrong thing, and the plausible refactoring of the wrong thing is a new failure mode for debugging, because it reads as progress and consumes the hour.

This is the same gap as the runbook that describes a system you do not have, generalised: every artefact that was never written is invisible to the reader, human or model, and the model reader is simply the first one honest enough to say so at scale.

The uncomfortable mirror

The agent's failure is a mirror of the team's, and this is the part that makes the argument uncomfortable. If the agent cannot debug the system because the intent is unwritten, then the new engineer cannot either, and the on call at two in the morning cannot either, and they have been paying that tax quietly for years. The agent did not create the gap. It measured it, loudly, and the measurement is the useful part.

Teams that treat the agent's confusion as the model's limitation will keep hitting the bound. Teams that treat it as an audit of their describability will fix the bound, and the fix helps every reader, including the humans.

What makes a system debuggable by an agent

The properties that help the agent are the properties that help any newcomer, which is reassuring, because it means the investment is not model specific.

Written intent at the decision points: not prose essays, but the one paragraph per module saying what it is for, what it must never do, and which incident or rule shaped it. The constraints section is the compass the agent lacks, and a short list of must nevers is more valuable than a long list of what the code does, because the code already says what it does.

Executable specifications for the money paths: the rounding, the totals, the authorisation, as property tests that state the invariant rather than one example, because invariants are the specification and examples are samples of it. The agent can then verify against the invariant, and so can the next human.

The incident memory in the repository: a postmortem file next to the code it changed, linked from the code, so the why survives the authors. The agent that can read the incident reads the intent, and the two year old retry becomes legible instead of mysterious.

The counterargument, fairly

There is a real defence that the bound is shrinking, that models infer intent from code better every release, and that waiting for perfect documentation is an argument against ever starting. I accept the trajectory. Inference from code does improve, and for well trodden patterns the priors are strong.

But the inference is about typical systems, and your system's value lives in the atypical decisions, the ones that differ from the prior, which are precisely the ones the model will confidently get wrong. The atypical is where your incidents are, and the atypical is, by definition, not inferable. It must be written.

The rule of thumb

Ask of your system the question you would hand the agent: in one page, what must this system never do, and why. If you cannot write it quickly, the agent cannot debug it, and neither can your next hire, and the tax is being paid nightly whether or not a model is involved.

Write the must nevers, the invariants and the incident memory, and the agent becomes a colleague. Leave them unwritten, and it becomes a very fast new hire with the same impossible first week, which was never the model's fault.

The debugging craft the agent cannot replace is the discipline of naming the expectation before checking the code, which is the habit in every bug is a wrong assumption, and the assumption, written down, is the specification the agent was waiting for.