The Agent Ran the Migration Twice Because It Misread Its Own Log

Share
The Agent Ran the Migration Twice Because It Misread Its Own Log. Abstract error autopsy illustration in orange and dark grey on debugly.dev

The alert was a duplicate key violation on a table that had been migrated the night before, by an agent, unattended. The migration was supposed to run once. It ran twice, forty seconds apart, and the second run is the one that failed loudly, which is the only reason anyone noticed. The first run had succeeded, and the agent, misreading its own truncated output as a timeout, had retried a migration that was already applied.

This is the autopsy of an agent failure that is not a model intelligence failure at all. The model reasoned sensibly at every step. The defect was a control loop that parsed free text to decide whether to retry, and free text is the one thing a retry decision must never depend on.

This was an autonomous maintenance agent running shell steps against Postgres 16.3, and the migration was idempotence free, which is what turned a parsing bug into a data incident.

The short answer

The migration tool prints a progress line and a final line. The agent captured the tool's stdout through a pipe with a fixed buffer, and the final line was cut off. Its instruction said "if the migration does not report success, retry once". The truncated output contained no success token, so the agent concluded failure and reran the migration. The reran migration hit rows the first run had inserted and died on the unique constraint. The agent's model was never wrong about what it saw. It saw a log without a success marker and did exactly what its policy said. The policy was the bug, because the policy treated absence of evidence as evidence of absence.

Why the retry was the dangerous part

A retry is safe only when the operation is idempotent or when the first attempt provably did not happen. Neither held. The migration inserted rows, so it is not idempotent, and the agent had no way to know whether the first run completed, because it destroyed the only evidence by truncating the buffer. This is the retry discipline from reviewing retry and backoff logic applied to an autonomous caller: the agent retried a non idempotent operation on ambiguous evidence, which is the single most expensive retry shape there is.

The deeper issue is that the agent's observation channel was lossy. An autonomous loop is only as good as the signal it reads, and a loop that reads truncated free text is a loop running on rumour.

The causes, ranked

1. A retry decision driven by parsed free text

The primary defect. Success was detected by grepping a human readable line, and the grep ran against a truncated capture. Machine decisions need machine readable signals: an exit code, a structured status row, a row in a migrations table. Free text is for humans.

2. A non idempotent migration with no ledger

The migration had no record of having run. A migrations table with an applied marker turns "did this run" from a log reading exercise into a query, and makes the second run a no op. The absence of the ledger is what let the retry do real work.

3. No pre check before acting

The agent did not inspect the schema before rerunning. A one line check, does the new column exist, would have told it the migration was already applied. Autonomous loops need a read before every write, because the read is the only thing that distinguishes a retry from a duplicate.

The fix

The immediate fix rolled back the duplicate rows and restored the constraint, and the agent was paused. The durable fixes were three.

First, the loop now decides on exit codes and a structured status object, never on grepped prose, and the capture is unbuffered and complete. The retry policy only fires on a definitive failure code, and never on missing output, because missing output is now treated as unknown, and unknown triggers inspection, not action.

Second, every migration the agent may run lives behind a ledger. The agent queries the ledger first, applies only if unapplied, and writes the marker in the same transaction as the change, so "applied" and "recorded" cannot disagree. This is the idempotency key pattern from webhook idempotency and duplicate delivery, moved to the agent's own actions.

Third, the agent performs a pre and post condition check around every mutating step, and logs both. The post condition is the receipt, and the receipt, not the tool's prose, is what the next decision reads.

What I would do differently

I would have treated the agent's observation channel as a first class interface, with the same rigour as an API contract, because it is one. The loop's inputs are its reality, and a lossy input makes a correct model behave incorrectly, which is the hardest class of agent incident to diagnose, because every individual decision looks reasonable in the transcript.

I would also have banned autonomous retries of non idempotent operations outright, not tuned them. The rule "an agent may retry only what it can prove did not happen" is short, absolute and cheap, and it removes the entire incident class rather than reducing its probability.

The rule

An autonomous loop that parses its own free text to decide whether to retry is a loop running on rumour, and a retry of a non idempotent step on ambiguous evidence is a duplicate waiting for a buffer to overflow. Decide on exit codes and ledgers, read before every write, and treat missing output as unknown, which means inspect, never act.

The agent was not stupid. It was misinformed, and the misinformation was ours, because we handed the retry decision to a grep. The same lesson, that the observation channel is the system, is the tracing argument in tracing an LLM pipeline with the observability you already have.