LLM observability
LLM observability means recording what an AI app does on each request, such as the prompt, each step, the time taken and the tokens used.
An app built on a language model often does several things to answer one question. It may search documents, call the model, run a tool, then call the model again. Generative AI is also non-deterministic. That means the same input can give a different output. So when a reply is wrong or slow, the final answer alone leaves you guessing why.
Observability means understanding what happens inside a system from what it reports. For LLM apps, the main tool is tracing. Picture a school help bot asked when the library closes. Its trace lists each step: the document search, the call to the model and the reply. For each step it keeps how long it took and how many tokens (small chunks of text) it used. If the bot gave the wrong time, you can open the retrieved page and the prompt, and see which step went wrong.
For engineers: OpenTelemetry is an open project for making software report what it does. In its terms, one unit of work is a span. Spans that share the same trace ID form one trace. Phoenix traces AI apps through OpenTelemetry. MLflow Tracing works with it too. LangSmith calls a span a run. Langfuse and MLflow send trace data in the background. So the app does not have to wait.
Without records, an AI app is hard to debug.
Follow one question through a traced app.
- 1 · instrumentCode is added to the app so that it reports what it does.
- 2 · recordEach step, such as a document search, a model call or a tool call, is recorded with its inputs, outputs and timing.
- 3 · linkThe records for one request share a trace ID, so they join into one trace.
- 4 · watchTools track response time, token use and error rates across traces.
- 5 · scorePeople or an LLM judge can attach scores and feedback to traces.
A trace turns one reply into a list of steps you can inspect.
| Who | What they ask | What it works with |
|---|---|---|
| Support team | “Why did the bot tell this customer the wrong refund date?” | The trace for that chat, with the retrieved passage and the prompt |
| Finance | “Which feature uses the most tokens?” | Token counts per step, tracked over time |
| Engineer on call | “Why did answers get slow this afternoon?” | The timing of each span, from search to model call |
| Quality team | “Did the new prompt make answers worse?” | Scores and user feedback attached to traces |
- You can see which documents a search returned, and in what order.
- You can find out why one step failed or ran slow.
- You can track model usage and cost over time.
- Real traces can be turned into test datasets for evaluations.
- It only sees what the app was set up to report.
- Traces alone do not say if an answer is good. Quality needs scores from people or an LLM judge.
- Prompts can contain personal data, so teams may need to redact it from traces.
- Traces are not always kept forever. LangSmith's hosted service keeps them for 180 days.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsObservability primer, OpenTelemetry · read 28 Sept 2026
- docsTraces, OpenTelemetry · read 28 Sept 2026
- docsObservability & Application Tracing, Langfuse · read 28 Sept 2026
- docsOverview: Tracing, Arize Phoenix · read 28 Sept 2026
- docsMLflow Tracing for LLM and Agent Observability, MLflow · read 28 Sept 2026
- docsObservability concepts, LangChain · read 28 Sept 2026