The Endpoint That Spent Nine Hundred Milliseconds Building JSON
The endpoint was "slow" in every review and fast in every query log. The database work was forty milliseconds. The whole request was nine hundred and forty. The eight hundred and sixty millisecond gap was not in the database, not in the network, and not in any function anyone had written on purpose. It was the runtime building and stringifying a forty megabyte JSON document, per request, for every client.
Serialisation is the invisible tax on every API, and it becomes visible exactly when payloads get large, which is when teams least expect it because the query is "fine".
This was Node 22.14 with JSON.stringify over a deeply nested mapper output, and the diagnosis came from a CPU profile, because serialisation is CPU and CPU is what profiles show.
Finding the gap
The partition that starts the investigation is the curl timing template. Time to first byte minus the handshake is the server's total think time. The query log says how much of that was the database. The difference is the application's own work, and when that difference is large and the code "does nothing", the nothing is usually serialisation or mapping.
The profile confirmed it. The hottest frames were JSON.stringify and the mapper that assembled the object it stringified. The endpoint was doing two jobs nobody had budgeted: transforming a large dataset into a new shape, and turning that shape into text, both synchronously on the request thread.
This is the CPU bound version of the endpoint that spent its time where the agent barely looked, except the hot function is one you recognise, which makes it easier to find and harder to believe.
Why it is so expensive
Serialisation cost is proportional to the size of the document and to the work of building it, and both had grown without anyone noticing.
The payload was the full nested representation of a large object graph, most of which no client used. The mapper walked the graph and allocated a parallel object tree, and stringify walked it again emitting text. Two full passes over forty megabytes, plus the allocation churn, on the event loop thread, which also means every other request on that process waited while it happened.
The growth was the trap. The endpoint was written when the object was small, and serialisation was negligible then. Data grew. The cost grew linearly and invisibly, because nothing in the code changed. This is the same silent growth as the query that was fast until the table grew, applied to bytes instead of rows.
The fixes, in order of payoff
Send less. The first and largest win is to question the payload. The clients used a fraction of the fields. A slimmer response, with only what is consumed, cuts serialisation, transfer and the client's parse cost at once. Parse cost is real on the client side too, and a forty megabyte response is a transfer problem as well, per Brotli saved forty percent on text.
Paginate or stream. A list endpoint returning everything returns it to every client forever. Bounding the page makes serialisation cost proportional to what the client asked for. Where the whole document is genuinely needed, stream it, writing chunks as they are produced, so the cost is spread and the first byte goes out early instead of after the whole document is built.
Stop rebuilding on every request. If the document changes rarely, build it once and cache the serialised string, keyed by a version, and serve the string. Serialising a cache hit is zero. This moves the tax from every request to every change, which is usually the right place for it.
Move it off the hot thread. For the cases where the document is large and dynamic, generate it in a worker or a background job and serve it from object storage, the same move as taking reports out of the request path in 504 gateway timeout while your application was still working. The request then returns a pointer or a cached file, and the expensive build happens once per change, not once per client.
Measuring it so it stays fixed
Add two numbers to the endpoint's instrumentation: the byte size of the response and the serialisation time, which you can measure by wrapping the stringify. Alert on the product of the two drifting upward. A payload that grows is normal. A payload that grows past a budget is a decision, and the budget makes the decision visible at review time instead of at the next latency incident.
The budget is the same discipline as a performance budget that survives contact with a real store, applied to the API layer.
The rule
Serialisation is real work with a real CPU cost, and it scales with the payload, not with the cleverness of the code. When the query log is fast and the endpoint is slow, profile the gap, and expect to find the runtime turning your object graph into text, twice, on the request thread, for every client.
Send less, bound the page, cache the string, or move the build off the path. The endpoint that serialises forty megabytes per request is not slow because of a bug. It is slow because nobody ever asked what the response should weigh.
There is a review habit that prevents the whole class: for every endpoint, ask what the response weighs and who consumes each field, and treat a payload that no client fully reads as a defect, not a convenience. The mapper that assembles the kitchen sink is the serialisation tax's author, and it is written once, at design time, by someone who will never sit inside the nine hundred milliseconds it buys.
The event loop consequence of this cost, when many requests pay it at once, is event loop lag, the metric nobody alerts on.