LLM observability tools record what a language-model application did on each request: prompts, model calls, retrievals, tool calls, latency, token counts, cost, errors and quality scores. The main choices in 2026 are self-hostable platforms such as Langfuse, Phoenix and MLflow, OpenTelemetry instrumentation such as OpenLLMetry, and commercial services such as LangSmith, Datadog and W&B Weave.
What LLM observability covers
Application monitoring records requests, errors and infrastructure metrics. LLM observability adds what a model-driven system needs on top: the full trace of a request through prompts, model calls, retrievals and tool calls, token usage and cost at each step, evaluation scores, and user feedback. For an agent, that trace becomes a multi-step trajectory, which our explainer on agent observability covers as a concept. This page is the tools list.
The same records often double as audit evidence, which raises the bar on retention, access and privacy. The layers an agent platform needs on top of tracing, from structured logs to tamper-evident audit trails, are set out in our piece on logs, traces and audit records for agents.
The metrics to collect
The OpenTelemetry GenAI semantic conventions, now maintained in their own repository, give standard names for most of these signals. Every metric and attribute below is marked Development, so names can still change between releases; instrumentations can opt into the newest experimental names with OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental (verified October 2026).
| Signal | What to record | OpenTelemetry GenAI name | Why it matters |
|---|---|---|---|
| End-to-end latency | Duration of each model call | gen_ai.client.operation.duration | Spots slow models, providers and prompts |
| Time to first token | First chunk seen by the client, first token produced by the server | gen_ai.client.operation.time_to_first_chunk, gen_ai.server.time_to_first_token | Perceived speed in chat interfaces |
| Generation speed | Time per output token or chunk | gen_ai.server.time_per_output_token, gen_ai.client.operation.time_per_output_chunk | Capacity planning for self-hosted models |
| Tokens | Input and output, plus cache reads and reasoning tokens | gen_ai.client.inference.usage.input_tokens, …output_tokens, …cache_read.input_tokens, …reasoning.output_tokens | Cost and context growth |
| Model identity | Requested model, responding model, provider | gen_ai.request.model, gen_ai.response.model, gen_ai.provider.name | Catches silent model swaps; audit |
| Agent and tool steps | Tool durations, tool calls and model calls per agent run | gen_ai.execute_tool.duration, gen_ai.invoke_agent.tool_calls, gen_ai.invoke_agent.inference_calls | Finds loops and slow tools |
| Cost | Tokens multiplied by a price table per model | Not defined in the metric documents; computed by each tool | Budgets per team, feature or user |
| Quality | Evaluator scores, user feedback, guardrail hits | Not defined in the metric documents | Regressions that latency charts miss |
| Errors | Failures, timeouts, retries, fallbacks | Error attributes on spans | Provider reliability |
One more rule from the span conventions matters for privacy. The attributes gen_ai.input.messages and gen_ai.output.messages carry the conversation itself, and the specification says instrumentations should not capture them by default because they are likely to contain personal data; capture should sit behind an explicit opt-in.
LLM observability tools compared
Each row reflects the project’s own documentation or repository, verified October 2026. A dash means the feature was not documented on the pages we checked.
| Tool | Self-hostable | Licence | OpenTelemetry | Evals | Cost tracking | PII handling |
|---|---|---|---|---|---|---|
| Langfuse | Yes | MIT core; SCIM, audit logging and retention modules licensed separately | OTLP over HTTP; SDK v4 built on OpenTelemetry | Yes | Ingested, or inferred from model prices | Masking in the SDK before export |
| Arize Phoenix | Yes, on Docker or Kubernetes | Elastic License 2.0 | Built on OpenTelemetry and OpenInference | LLM, code and human evaluators | Built-in price table, custom prices | Hide settings in the instrumentation |
| OpenLLMetry | Instrumentation; send to any OTel backend | Apache 2.0 | Native | – | – | Content captured by default; can be switched off |
| MLflow Tracing | Yes | Apache 2.0 | Exports and ingests the GenAI conventions | LLM-judge scorers, human feedback | From MLflow 3.10, using LiteLLM’s price table | PII redaction from traces |
| LangSmith | Enterprise add-on only | Commercial | OTLP ingest and SDK export | Yes | – | Client-side hiding and anonymisers |
| Datadog Agent Observability | No, SaaS | Commercial | GenAI conventions supported natively | Experiments and quality checks | Cost dashboards | Sensitive Data Scanner redaction |
| Grafana Cloud AI Observability | Grafana Cloud; OpenLIT UI can be self-hosted | Commercial service; OpenLIT Apache 2.0 | OpenTelemetry-native | Hallucination and toxicity scoring | Spend tracking and budgets | – |
| Helicone | Yes, with Docker or Helm | Apache 2.0 | AI gateway supports OpenTelemetry | Yes | Yes | Headers that skip storing bodies |
| W&B Weave | Multi-tenant, dedicated or self-managed | SDK Apache 2.0 | OTLP endpoint | LLM judges and custom scorers | – | Presidio redaction in the Python SDK |
Open-source and self-hostable tools
Langfuse
Langfuse covers tracing, evaluations, prompt management, experiments, annotation and a playground under the MIT licence, and had more than 35,000 GitHub stars in October 2026. When you self-host, SCIM, audit logging and data-retention policies sit in separately licensed enterprise modules. It accepts OpenTelemetry traces over OTLP HTTP, though not gRPC, maps GenAI, OpenInference and MLflow attributes, and its v4 SDK is a thin layer over the OpenTelemetry client. Cost comes from the values you send or is inferred from model price definitions, including your own for self-hosted or fine-tuned models. Masking runs in the SDK before data leaves your application, and there is no server-side redaction. ClickHouse acquired Langfuse in January 2026 and committed to keeping the core MIT licensed and self-hostable at production scale.
Arize Phoenix
Phoenix is built on OpenTelemetry and OpenInference instrumentation, runs on Docker or Kubernetes, and supports LLM-based, code-based and human evaluations attached to datasets and experiments. Its licence is the Elastic License 2.0, which is source-available rather than OSI-approved and limits offering the software as a managed service, so check it against your plans. Cost tracking uses a built-in price table, with custom prices for other models and for cache, audio, reasoning and image tokens. Masking is set in the instrumentation through OPENINFERENCE_HIDE_INPUTS, OPENINFERENCE_HIDE_OUTPUTS and related settings, so hidden values never leave the process. Arize AX is the managed commercial platform built on the same standards.
OpenLLMetry and the OpenTelemetry conventions
OpenLLMetry, maintained by Traceloop under Apache 2.0, is instrumentation rather than a backend. It extends OpenTelemetry so traces from LLM frameworks reach any compatible destination, and its README lists more than 20, including Datadog, Grafana, Honeycomb, New Relic, Splunk, Dynatrace and Sentry. By default it records prompts, completions and embeddings as span attributes; setting TRACELOOP_TRACE_CONTENT=false turns that off globally, and it can also be disabled per workflow. Traceloop joined ServiceNow in March 2026 and said OpenLLMetry will remain open source. Because its conventions have been folded into OpenTelemetry itself, instrumenting this way keeps the backend replaceable.
MLflow Tracing
MLflow, licensed Apache 2.0, describes its tracing as fully OpenTelemetry-compatible: it exports and ingests the GenAI conventions, and traces stay on infrastructure you host. Token usage is recorded from MLflow 3.2 and cost from 3.10, when the tracking server is installed with the [genai] extra. Cost relies on LiteLLM’s price table, so a model missing from it shows $0.00 even with correct token counts. Evaluation runs through datasets, scorers and built-in LLM-as-a-judge scorers, and human feedback is attached to traces with its metadata. It is the natural pick for teams already using MLflow for classic models.
Helicone
Helicone pairs an AI gateway with an observability platform, licensed Apache 2.0, that can be self-hosted with Docker or Helm. Mintlify acquired Helicone on 3 March 2026, and Helicone now runs in maintenance mode: security updates, bug fixes and new models keep shipping, and Mintlify says it will help customers migrate to another platform. Its omit headers stop request and response bodies from being stored but, as its documentation notes, the content is still sent to Helicone’s backend. Plan new deployments on another tool.
Commercial platforms
LangSmith
LangSmith, from LangChain, is a hosted platform for tracing, evaluation and prompt engineering. Developer seats cost $0 with 5,000 base traces a month and Plus costs $39 per seat with 10,000, then pay-as-you-go; Enterprise is custom (list, verified October 2026). Base traces are kept for 14 days and extended traces for 180. Self-hosting on Kubernetes, with ClickHouse, PostgreSQL and Redis inside your environment, is an add-on to the Enterprise plan. LangSmith ingests OpenTelemetry traces and maps GenAI attributes to its own fields, and masking happens on the client through LANGSMITH_HIDE_INPUTS, LANGSMITH_HIDE_OUTPUTS, regex anonymisers or integrations such as Presidio.
Datadog Agent Observability
Datadog’s product page now calls the product Agent Observability, while its documentation still uses the name LLM Observability. It traces LLM calls, fixed workflows and agent steps, compares prompts, models and agent configurations in experiments, tracks token cost, and uses Sensitive Data Scanner to redact sensitive data and flag prompt-injection attempts. The documentation says it supports the OpenTelemetry GenAI conventions natively. Pricing is per LLM span: free up to 40,000 spans a month, and Pro at $160 a month for 100,000 (list, verified October 2026). It is SaaS only, so trace content leaves your environment.
Grafana Cloud AI Observability
Grafana Cloud’s AI observability is OpenTelemetry-native and built with OpenLIT. It covers LLMs, vector databases, GPUs, MCP servers and agents, with spend tracking, budget management and automated checks for hallucinations, factual accuracy and toxicity. The OpenLIT SDK and its self-hosted dashboard on ClickHouse are Apache 2.0, so the same instrumentation can stay inside your network if Grafana Cloud is not an option.
W&B Weave
Weave traces LLM calls and functions and evaluates responses with LLM judges and custom scorers. It accepts OTLP traces on multi-tenant cloud, dedicated cloud and self-managed instances, and its Python SDK can redact personal data with Presidio before traces leave the client; other SDKs lack that option. Its documentation now lives on CoreWeave’s site and names CoreWeave Agent Lens, in public preview on multi-tenant SaaS, as the successor to Weave. The two share a backend, and Weave still covers evaluations and guardrails that Agent Lens does not yet offer.
Keeping prompts and personal data out of traces
Traces of an LLM application hold the most sensitive text the system ever sees, so treat the observability store like any other system of record:
- Capture content only by decision. The GenAI conventions make message capture opt-in; keep it that way unless a team needs content to debug.
- Mask at the source. Langfuse masking functions, LangSmith hide settings, OpenInference hide settings in Phoenix and Weave’s PII redaction all act before data leaves the process.
- Redact in the pipeline. The OpenTelemetry Collector’s redaction processor deletes attributes that are not on an allowlist and masks values matching blocked patterns; it is beta for traces and alpha for logs and metrics.
- Read what “omit” means. Helicone’s omit headers prevent storage, not transmission.
- Set retention deliberately. LangSmith keeps base traces for 14 days; in self-hosted Langfuse, retention policies sit in the licensed enterprise modules.
- Control access and export. Traces used as audit evidence need role-based access and a route into your SIEM.
How to choose an LLM observability tool
- Decide where trace data may live. If prompts contain regulated data, start with self-hostable tools or with SaaS whose masking runs in your process.
- Instrument with OpenTelemetry. The backend then becomes replaceable. Check which version of the GenAI conventions each candidate expects.
- Read licence and tier labels. Langfuse’s enterprise modules, Phoenix’s Elastic License and LangSmith’s Enterprise-only self-hosting all change the total cost.
- Check vendor status. Helicone is in maintenance mode, Weave has a named successor, and Langfuse and Traceloop now belong to larger companies.
- Price on your own unit. Spans, traces and seats scale differently, and an agent run can produce dozens of spans.
- Add gateway logs. If model traffic passes through an AI gateway, its per-call records are a second source of cost and latency data; our AI gateway comparison notes what each one logs.
- Tie evaluation to traces. Use offline test sets before release and online scores after, following an agent evaluation method and the practice of testing agents before production.
How VDF AI fits
For workflows that run on VDF AI Networks, observability is part of the platform. Every node execution is tracked with its status, input and output, token usage and cost. Latency is reported at P50 and P95 per network and per node, service-level objectives can be set per network with alerts on breaches, and execution traces include model routing decisions and their rationale.
Across the platform, the Trust Center describes audit records for every prompt, tool call, retrieval, model route and response, with actor, timestamp and context, streamed to your SIEM and kept in your environment, whether on-premises, in a private cloud or air-gapped. VDF AI Router adds the reason, candidate list and scores behind each routing decision. The product pages do not document an OpenTelemetry export, so if you standardise on one of the tools above, confirm the integration path during evaluation.
Sources
- OpenTelemetry GenAI semantic conventions repository
- GenAI span conventions
- GenAI metric conventions
- GenAI token metric conventions
- GenAI agent span conventions
- OpenTelemetry Collector redaction processor
- Langfuse repository
- Langfuse open-source licensing
- Langfuse and OpenTelemetry
- Langfuse token and cost tracking
- Langfuse masking
- ClickHouse acquires Langfuse
- Phoenix documentation
- Phoenix licence
- Phoenix cost tracking
- Phoenix span masking
- OpenLLMetry repository
- OpenLLMetry privacy settings
- Traceloop is joining ServiceNow
- MLflow Tracing
- MLflow token and cost tracking
- MLflow evaluation and monitoring
- MLflow repository
- LangSmith pricing
- Self-hosted LangSmith
- LangSmith and OpenTelemetry
- LangSmith input and output masking
- Datadog LLM Observability documentation
- Datadog Agent Observability
- Grafana Cloud AI Observability
- OpenLIT repository
- Helicone repository
- Helicone omit logs
- Helicone Docker self-hosting
- Helicone AI Gateway repository
- Mintlify acquires Helicone
- W&B Weave documentation
- Weave OpenTelemetry endpoints
- Weave PII redaction
- What is CoreWeave Agent Lens