The Request That Failed Only When the Input Contained a Comma

Search worked for every input except inputs with a comma, which 404'd. The comma is legal in a URI path, but a router in the middle split on it anyway, and…

Share
The Request That Failed Only When the Input Contained a Comma. Abstract bug hunt illustration in orange and dark grey on debugly.dev

The bug report was almost a joke: searching for anything worked, except the phrase "washer, dryer" returned a 404. Remove the comma and the same words found forty results. A punctuation mark was the difference between working and not existing, which is the signature of a parser disagreement rather than a data problem.

The comma is a legal character in a URI path. RFC 3986 lists it among the sub-delims that may appear without encoding. Somewhere between the browser and the handler, a component that did not agree with the RFC split the path on the comma and broke the route. This is the autopsy of that disagreement.

This was a Node 22.14 service behind an nginx 1.27 reverse proxy and a routing layer that normalised paths, observed in Chrome 133.

The symptom, precisely

Any path parameter containing a comma produced a 404 or a route mismatch, while the same parameter with the comma replaced by a space worked. Other reserved characters behaved differently: a slash encoded as %2F survived, a question mark encoded survived, but the raw comma died. The selectivity of the failure, one character class only, is the clue that a specific normalisation step, not a general bug, was responsible.

The hypotheses that died

The client is encoding wrong

First I checked what the browser sent. The network tab showed the comma sent raw in the path, which is legal, and the server received it. The client was fine.

It died because the bytes arrived intact. The loss happened after arrival.

The search index chokes on commas

Next I assumed the search backend tokenised on commas and returned nothing. I queried the backend directly with the comma phrase and got results.

It died because the backend was never reached. The 404 came from the routing layer before any search ran.

The application route is wrong

Then I tested the application directly, bypassing the proxy, with the comma path. It routed correctly. So the application's router accepted the comma.

It died because the failure required the full chain. The defect lived between the client and the app.

The breakthrough

The middle layer was a normalising router that rewrote paths before matching, and its rewrite split the path on commas, a historical artefact of a matrix parameter scheme long since removed. With the path split, the segment the route expected was truncated at the comma, no route matched, and the layer returned 404. The application, which would have handled the comma fine, never saw the request.

The confirmation was one capture: the client sent one path, the app would have accepted it, and the middle layer's rewrite log showed the split. The disagreement was between two components with different opinions about the comma, and the stricter one sat first.

Why characters disagree across layers

URI handling is a stack of parsers, browser, proxy, router, framework, and each implements its own reading of the RFC and its own legacy exceptions. The RFC permits the comma in a path, but permits is not compels, and many components treat commas as meaningful because some historic scheme did. So a character that is data to one layer is syntax to another, and the layer that treats it as syntax wins, because it acts first.

The same shape appears with other characters in other stacks, semicolons, plus signs in queries, encoded slashes, and the general lesson is that a path parameter's survivability is decided by the strictest parser in the chain, not by the RFC.

The fixes

Encode at the edge, decode at the handler. The client should percent encode the comma as %2C when placing user input in a path, which is always safe, because an encoded character is data to every parser. The framework then decodes it back at the handler. Encoded bytes pass through syntax hungry layers untouched, which is the durable fix for any character.

Prefer query parameters for free text. User input that can contain anything belongs in the query string with proper encoding, or in the request body, not in a path segment. Path segments are for identifiers you control. Free text in a path is an invitation to exactly this class of bug.

Make the middle layer RFC honest. Where a normalising router must exist, configure it not to split on sub-delims, and add a test that sends each reserved character encoded and raw through the full chain and asserts the handler sees the original string. The test is the contract between the parsers.

What I would do differently

I would have reproduced through the full chain first, because testing the app alone and the client alone both showed healthy components, and the defect only existed in the composition. Composition bugs need composition reproduction, which is the same lesson as the formatting only diff that changed behaviour, where the truth lived in the interaction, not in the parts.

I would also have treated the single character selectivity as the diagnosis. Bugs that flip on one character class are parser disagreements by definition, and the character names the layer that disagrees.

If it happens again

A path parameter survives only if every parser between the client and the handler agrees about its characters, and agreement is never guaranteed for punctuation the RFC merely permits. Encode user input at the edge, keep free text out of path segments, and test reserved characters through the full chain, because the strictest parser you forgot is the one that will eat the comma.

The server side twin of this, where the server's own parsing of a supplied value is the vulnerability, is SQL injection inside the ORM, where the layer that treats input as syntax is the defect itself.