> ## Content Index
> Fetch the complete content index at: https://debugly.dev/llms.txt
> Use this file to discover other available public pages before exploring further.

# The Agent's Retry Loop Rate Limited the Service It Was Trying to Fix
- URL: https://debugly.dev/agent-retry-loop-rate-limited/
- Published: 2026-10-06T15:30:00.000Z
- Updated: 2026-10-06T15:30:00.000Z
- Author: Rohit Bhadani
- Tags: Bug Hunt, AI, Networking

The on call engineer joined a bridge about an API returning 429s to all customers, which read as a capacity incident. It was not. The service was mildly degraded by a small upstream slowness, and an autonomous debugging agent, attached to the same API to investigate that slowness, had interpreted its own timeouts as signal to retry, and its retry loop, running without a budget, had become the largest client of the service it was debugging. The agent was the incident. The original bug was still there, underneath, unfound, because the evidence had been drowned.

This is the bug hunt for the moment an autonomous investigator became the load, and for why an agent without a rate budget is a denial of service with good intentions.

This was an agent probing a Node 22.14 API behind a per client rate limit, and the loop ran for eleven minutes before a human noticed.

## The symptom, precisely

Two things were true at once. The API's error rate was elevated with 429s dominant, and the agent's transcript showed it repeatedly concluding "the endpoint is flaky, retrying", because every probe it sent came back either slow or rate limited. The rate limiting was caused by the agent, so its own probes guaranteed its own failures, which guaranteed more retries, a self sustaining loop where the observation confirmed the hypothesis that caused the observation.

That circularity is the signature. When an investigator's measurement changes the system under investigation, the measurement is no longer evidence, and the loop can run forever on its own exhaust.

## The hypotheses that died

### The service was under external attack

First the team assumed a third party was hammering the API. The per client counters showed the top client by a wide margin was the debugging agent's own token.

It died because the attacker was the investigator.

### The agent was compromised

Next I suspected the agent had been hijacked into a flood. The transcript showed no malice, just a retry policy with no ceiling, executing faithfully.

It died because the behaviour was the policy working as written, which is worse than compromise, because it is reproducible.

### The rate limit was misconfigured

Then I checked the limit itself. It was sane for humans. It was not sane for a loop that fires a probe every two hundred milliseconds, because no human can, and the limit had never been asked to consider a machine.

It died because the limit was correct for its intended population, and the agent was a new population.

## The breakthrough

The diagnosis is that the agent's retry policy had three missing budgets, the same three from [reviewing retry and backoff logic](https://debugly.dev/reviewing-retry-and-backoff-logic/): no maximum attempts, no backoff, no jitter, and additionally no awareness that its own traffic was a share of the thing it measured. A human debugger throttles instinctively, because a human is slow. An agent is fast, and speed without a budget is volume.

The circular measurement made it self confirming. Each rate limited probe was logged as "the service is failing", strengthening the belief that more probing was needed. The loop had no step that asked "is my probing changing the measurement", which is the question that breaks the circle, and no agent policy includes it unless you write it.

## What I changed

**The agent got a client side budget.** Maximum probes per minute, exponential backoff with jitter, and a hard cap on total attempts per investigation, after which the agent must stop and report instead of continuing. The budget is not a performance tweak, it is the boundary between investigating and attacking.

**The agent measures from the side, not by adding load.** Where possible the agent now reads existing telemetry, logs, metrics, traces, the material from [tracing an LLM pipeline with the observability you already have](https://debugly.dev/tracing-llm-pipelines-minimal-observability/), and only sends active probes when the passive evidence is insufficient, and then a bounded number. Passive observation does not add load, so it cannot drown the evidence.

**The agent's own traffic is labelled and exemptable from the investigation.** Probes carry a header marking them as agent traffic, so the investigation can exclude its own footprint from the metrics it reasons over, breaking the circular confirmation. An investigator that cannot subtract itself from the measurement cannot be trusted about the measurement.

**The kill switch is a first class control.** A single command stops all autonomous probing, and the loop checks it between attempts, per the pipeline discipline in [running agents in CI needs budgets, gates and a kill switch](https://debugly.dev/agents-in-ci-budgets-kill-switches/). The absence of a kill switch is what made eleven minutes the detection time.

## What I would do differently

I would have load tested the agent before letting it touch production, by pointing it at a staging replica and watching its probe rate, because the flood was predictable from the policy. A retry loop with no budget is a flood that has not met traffic yet, and the agent's own policy file was the specification of the incident.

I would also have alerted on the agent's outbound rate the way I alert on any client's, because an autonomous tool is a client, and clients get monitored. The 429 dashboard had the answer from minute one, in a row labelled with the agent's token, that nobody was watching.

## The rule

An autonomous investigator without a probe budget is a denial of service that believes it is debugging, and a loop that reads its own self induced failures as evidence will run forever on its own exhaust. Budget the attempts, back off with jitter, observe passively first, label and subtract your own traffic, and keep a kill switch one keystroke away.

The service was never the story. The story is that speed without a budget is volume, and volume without a label is indistinguishable from an attack, which is exactly what the rate limiter, correctly, assumed it was.