> ## Content Index
> Fetch the complete content index at: https://debugly.dev/llms.txt
> Use this file to discover other available public pages before exploring further.

# Tracing an LLM Pipeline With the Observability You Already Have
- URL: https://debugly.dev/tracing-llm-pipelines-minimal-observability/
- Published: 2026-10-03T15:30:00.000Z
- Updated: 2026-10-03T15:30:00.000Z
- Author: Rohit Bhadani
- Tags: Tooling, AI, DevOps

The complaint was that the summariser "got worse" on Tuesday, and the team had no way to confirm it, date it, or attribute it, because the pipeline logged nothing about itself. No prompt version, no model identifier, no token counts, no per stage latency. The behaviour was a black box even to its owners, and the investigation that should have taken a query took a week of guessing. The fix was not a new observability platform. It was the structured logging discipline from [structured logging and what to log](https://debugly.dev/structured-logging-what-to-log/), pointed at the four facts an LLM pipeline must record.

This is the minimal tracing stack, built from the logger, the metrics endpoint and the correlation id you already have, and the reasoning for each field.

This was a Node 22.14 pipeline calling a hosted model, with Redis 7.2 for caching, and the fields below were added in one afternoon.

## The four fields that change everything

**Prompt version.** The prompt is code, and code has versions, so every request logs the identifier of the exact prompt template and variables hash it used. Without it, "it got worse on Tuesday" is unverifiable, because Tuesday's prompt is unrecoverable. With it, the regression is a join between the eval results and the version, and the rollback in [treat a prompt change like a deploy](https://debugly.dev/prompt-deploys-versioning-and-rollback/) has a target.

**Model and checkpoint identifier.** The provider's model name is not a version, and fleets roll checkpoints, so the request logs the model string plus whatever replica or checkpoint id the provider returns. This turns "the model changed behaviour" from a suspicion into a join, and it is the only way to separate your prompt regression from the provider's rollout, the determinism lesson from [temperature zero is not deterministic](https://debugly.dev/temperature-zero-is-not-deterministic/).

**Token counts, in and out.** Input and output tokens are the cost and the size of the context, and logging them per request makes the slow expensive queries visible as a query, not as a bill at month end. It also catches context bloat, the prompt that grew by two thousand tokens because someone appended a section, which is a performance regression that only the token count shows.

**Per stage latency.** A pipeline is retrieval, assembly, inference and post processing, and the total latency hides which stage moved. Logging each stage's duration makes the p99 partitionable, which is the whole argument in [p99 and why averages lie](https://debugly.dev/p99-latency-and-why-averages-lie/), applied to the stages instead of the hosts.

## The correlation id ties it together

Each request gets one id, propagated through retrieval, inference and post processing, and included in every log line and in the response headers for the client side join. The id is what makes a trace a trace rather than a pile of lines, and it is the same discipline as the request id in [journalctl is a database](https://debugly.dev/journalctl-queries-beyond-tail/), where the id is the query key.

With the id and the four fields, the week long investigation becomes: filter by date, group by prompt version, compare the output token and latency distributions, and read three traces. The "got worse" is now a diff between two versions' behaviour, which is the first time the complaint has been a fact.

## Recording the output, carefully

The one field beyond the four is the model output itself, and it is the one with a privacy budget. Logging full prompts and outputs is a data pipeline of personal information, the trap from [the pull request that logged an email address](https://debugly.dev/reviewing-logs-for-pii-before-it-ships/), because prompts routinely contain user content. The minimal safe version logs a hash of the prompt and output plus a sampled, redacted, short lived copy for eval, with retention measured in days and access audited. The hash gives you dedup and regression detection, the sample gives you eval material, and the retention keeps the lawyer calm.

## The eval join is the payoff

The reason to build this before buying a tracing product is that the join it enables is the eval loop: take the sampled outputs, score them with the eval harness from [testing LLM features that are non deterministic](https://debugly.dev/testing-llm-features-non-deterministic/), and join the scores back to prompt version and model id. Now every version has a measured quality number in production, not just in the offline eval, and a deploy that moves the number is visible in the same dashboard as the latency. That join is the thing the expensive platforms sell, and it is forty lines around the logger you have.

## When to buy the platform

The minimal stack fails gracefully at one point: visualising multi hop agent graphs, where a request fans into dozens of model calls and the flat log becomes hard to read. If the product is a single pipeline, the four fields and the id are enough for a year. If it is an agent swarm, the visualiser earns its cost. The discipline is to buy it when the flat trace genuinely stops reading, not when the demo is pretty, which is the search engine argument from [observability platforms sold you a search engine](https://debugly.dev/observability-vendors-sold-you-a-search-engine/), applied to LLM tooling.

## The rule

An LLM pipeline without prompt version, model id, token counts, per stage latency and a correlation id is unauditable, and "it got worse" is not debuggable without an audit trail. Add the five to the logger you already have, sample and redact the content with short retention, and join the traces to your evals.

The tracing product can wait. The four fields cannot, because the Tuesday regression is already happening, and the only question is whether you will be able to prove it.

One habit makes the whole stack age well: treat the trace fields as a contract with your future self, and add a test that fails if a request is logged without them. The fields are cheap, the absence is silent, and the silent absence is exactly the state the pipeline was in on the Tuesday nobody could explain. A schema test for your own telemetry is the cheapest insurance in this post, and it converts the tracing stack from a convention into an enforced property.