What Is LLM Observability and What Should You Monitor?
LLM observability is the practice of capturing, tracing, and evaluating every model and agent run so you can see what your application actually did — not just whether it was up. It matters when you run multi-step agents, route across multiple models, or ship prompt changes regularly, because uptime and error-rate dashboards won't tell you why an answer was wrong or which step burned the tokens. If your application is a single stateless prompt with no tool calls, basic logging may be enough; the more steps, models, and tools involved, the more you need structured traces and evals.
Observability vs. traditional monitoring
Traditional monitoring answers "is the service responding?" LLM observability answers "what happened inside this run, and was the output any good?"
| Dimension | Traditional monitoring | LLM observability |
|---|---|---|
| Unit of analysis | Request / endpoint | Run, trace, and span |
| Primary signals | Uptime, latency, error rate | Prompts, completions, latency, token cost, tool calls, errors |
| Quality measurement | Rarely covered | Automated evals and scoring |
| Failure mode caught | Service down or slow | Wrong answer, bad retrieval, wrong tool, silent regression |
The two are complementary. You still want uptime and latency alerting; observability adds the layer that explains output quality.
Core signals to capture
Instrument at the level of the individual model call and the surrounding workflow step, then roll up. The signals worth capturing on every run:
- Prompts and completions — the actual input sent and output returned, including system messages and any retrieved context.
- Latency — per call and per span, so you can separate model time from tool time from your own code.
- Token cost — input and output tokens per call, attributed to a model, a feature, or a user.
- Tool calls — which tools were invoked, with what arguments, and what came back.
- Errors — provider errors, timeouts, malformed tool arguments, and schema validation failures.
Capturing prompts and completions is what makes debugging possible; capturing cost and latency per span is what makes optimization possible.
How tracing connects multi-step agent workflows
A single trace represents one end-to-end run. Inside it, each step is a span: a model call, a retrieval, a tool invocation, a routing decision. Spans nest, so a parent span for "answer user question" can contain child spans for "retrieve documents," "call model," and "call calculator."
That structure is what lets you trace a failure to a specific step. If the final answer is wrong, the trace shows whether retrieval returned irrelevant documents, whether the model ignored them, or whether a tool returned an error that the agent silently swallowed. Without spans, you only see the bad final output and have to guess.
For agents that route across multiple models — for example, sending easy requests to a cheaper model and hard ones to a stronger model — the trace should record which model handled each call. Respan describes itself as an AI router with built-in observability and automated evals, which is the pattern to look for: routing decisions and observability data captured in the same place, so you can compare cost and quality across the models you route between.
Turning traces into quality signals with evals
Traces tell you what happened; evals tell you whether it was good. Automated evals score runs — against reference answers, rubrics, or model-based judges — and attach those scores to the trace. Over time, that gives you a quality trend line you can compare across prompt versions, model versions, and releases.
The practical loop:
- Capture traces for every run.
- Score a sample (or all) of them with automated evals.
- Compare scores across prompt or model changes.
- Ship the version that scores best, and keep monitoring after release.
This is why "ship the version that scores best" and "know the moment things change" belong together: evals give you the comparison, and continuous monitoring tells you when a previously good version starts degrading — for example, after a provider silently updates a model.
What to check when observability data looks wrong or missing
If traces are incomplete or scores look off, work through these common causes:
- Uninstrumented calls. A code path that calls a model directly, bypassing your gateway or SDK wrapper, produces no span. Audit for direct provider clients.
- Dropped spans. Async or background work that finishes after the parent run closes can lose its span. Check that spans are flushed before the trace is finalized.
- Missing context. If prompts are logged but retrieved documents aren't, you can't tell whether a bad answer came from bad retrieval. Instrument retrieval as its own span.
- Sampling gaps. If you sample traces, make sure eval scores and cost totals account for the sampling rate rather than reporting sampled numbers as totals.
- Clock and ordering issues. Out-of-order timestamps make latency attribution misleading; verify timestamps come from a consistent source.
Choosing what to instrument first
Start with the signals that map to decisions you actually make. If you're optimizing cost, instrument tokens and model routing per call. If you're debugging quality, instrument prompts, completions, retrieval, and tool calls, and add evals. If you're managing reliability, instrument errors and latency per span. A platform that combines routing, tracing, and automated evals — as Respan positions itself — reduces the work of stitching those layers together, but the signals above are what you need regardless of which tool provides them.