AI Infrastructure

LLM Observability Tools in 2026: Open-Source and Commercial Platforms Compared

LLM observability tools checked against their own docs and repositories: Langfuse, Arize Phoenix, OpenLLMetry and the OpenTelemetry GenAI conventions, MLflow, LangSmith, Datadog, Grafana, Helicone and W&B Weave, compared on self-hosting, licence, OpenTelemetry, evals, cost and PII handling. Verified October 2026.

LLM observability tools record what a language-model application did on each request: prompts, model calls, retrievals, tool calls, latency, token counts, cost, errors and quality scores. The main choices in 2026 are self-hostable platforms such as Langfuse, Phoenix and MLflow, OpenTelemetry instrumentation such as OpenLLMetry, and commercial services such as LangSmith, Datadog and W&B Weave.

What LLM observability covers

Application monitoring records requests, errors and infrastructure metrics. LLM observability adds what a model-driven system needs on top: the full trace of a request through prompts, model calls, retrievals and tool calls, token usage and cost at each step, evaluation scores, and user feedback. For an agent, that trace becomes a multi-step trajectory, which our explainer on agent observability covers as a concept. This page is the tools list.

The same records often double as audit evidence, which raises the bar on retention, access and privacy. The layers an agent platform needs on top of tracing, from structured logs to tamper-evident audit trails, are set out in our piece on logs, traces and audit records for agents.

The metrics to collect

The OpenTelemetry GenAI semantic conventions, now maintained in their own repository, give standard names for most of these signals. Every metric and attribute below is marked Development, so names can still change between releases; instrumentations can opt into the newest experimental names with OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental (verified October 2026).

SignalWhat to recordOpenTelemetry GenAI nameWhy it matters
End-to-end latencyDuration of each model callgen_ai.client.operation.durationSpots slow models, providers and prompts
Time to first tokenFirst chunk seen by the client, first token produced by the servergen_ai.client.operation.time_to_first_chunk, gen_ai.server.time_to_first_tokenPerceived speed in chat interfaces
Generation speedTime per output token or chunkgen_ai.server.time_per_output_token, gen_ai.client.operation.time_per_output_chunkCapacity planning for self-hosted models
TokensInput and output, plus cache reads and reasoning tokensgen_ai.client.inference.usage.input_tokens, …output_tokens, …cache_read.input_tokens, …reasoning.output_tokensCost and context growth
Model identityRequested model, responding model, providergen_ai.request.model, gen_ai.response.model, gen_ai.provider.nameCatches silent model swaps; audit
Agent and tool stepsTool durations, tool calls and model calls per agent rungen_ai.execute_tool.duration, gen_ai.invoke_agent.tool_calls, gen_ai.invoke_agent.inference_callsFinds loops and slow tools
CostTokens multiplied by a price table per modelNot defined in the metric documents; computed by each toolBudgets per team, feature or user
QualityEvaluator scores, user feedback, guardrail hitsNot defined in the metric documentsRegressions that latency charts miss
ErrorsFailures, timeouts, retries, fallbacksError attributes on spansProvider reliability

One more rule from the span conventions matters for privacy. The attributes gen_ai.input.messages and gen_ai.output.messages carry the conversation itself, and the specification says instrumentations should not capture them by default because they are likely to contain personal data; capture should sit behind an explicit opt-in.

LLM observability tools compared

Each row reflects the project’s own documentation or repository, verified October 2026. A dash means the feature was not documented on the pages we checked.

ToolSelf-hostableLicenceOpenTelemetryEvalsCost trackingPII handling
LangfuseYesMIT core; SCIM, audit logging and retention modules licensed separatelyOTLP over HTTP; SDK v4 built on OpenTelemetryYesIngested, or inferred from model pricesMasking in the SDK before export
Arize PhoenixYes, on Docker or KubernetesElastic License 2.0Built on OpenTelemetry and OpenInferenceLLM, code and human evaluatorsBuilt-in price table, custom pricesHide settings in the instrumentation
OpenLLMetryInstrumentation; send to any OTel backendApache 2.0Native––Content captured by default; can be switched off
MLflow TracingYesApache 2.0Exports and ingests the GenAI conventionsLLM-judge scorers, human feedbackFrom MLflow 3.10, using LiteLLM’s price tablePII redaction from traces
LangSmithEnterprise add-on onlyCommercialOTLP ingest and SDK exportYes–Client-side hiding and anonymisers
Datadog Agent ObservabilityNo, SaaSCommercialGenAI conventions supported nativelyExperiments and quality checksCost dashboardsSensitive Data Scanner redaction
Grafana Cloud AI ObservabilityGrafana Cloud; OpenLIT UI can be self-hostedCommercial service; OpenLIT Apache 2.0OpenTelemetry-nativeHallucination and toxicity scoringSpend tracking and budgets–
HeliconeYes, with Docker or HelmApache 2.0AI gateway supports OpenTelemetryYesYesHeaders that skip storing bodies
W&B WeaveMulti-tenant, dedicated or self-managedSDK Apache 2.0OTLP endpointLLM judges and custom scorers–Presidio redaction in the Python SDK

Open-source and self-hostable tools

Langfuse

Langfuse covers tracing, evaluations, prompt management, experiments, annotation and a playground under the MIT licence, and had more than 35,000 GitHub stars in October 2026. When you self-host, SCIM, audit logging and data-retention policies sit in separately licensed enterprise modules. It accepts OpenTelemetry traces over OTLP HTTP, though not gRPC, maps GenAI, OpenInference and MLflow attributes, and its v4 SDK is a thin layer over the OpenTelemetry client. Cost comes from the values you send or is inferred from model price definitions, including your own for self-hosted or fine-tuned models. Masking runs in the SDK before data leaves your application, and there is no server-side redaction. ClickHouse acquired Langfuse in January 2026 and committed to keeping the core MIT licensed and self-hostable at production scale.

Arize Phoenix

Phoenix is built on OpenTelemetry and OpenInference instrumentation, runs on Docker or Kubernetes, and supports LLM-based, code-based and human evaluations attached to datasets and experiments. Its licence is the Elastic License 2.0, which is source-available rather than OSI-approved and limits offering the software as a managed service, so check it against your plans. Cost tracking uses a built-in price table, with custom prices for other models and for cache, audio, reasoning and image tokens. Masking is set in the instrumentation through OPENINFERENCE_HIDE_INPUTS, OPENINFERENCE_HIDE_OUTPUTS and related settings, so hidden values never leave the process. Arize AX is the managed commercial platform built on the same standards.

OpenLLMetry and the OpenTelemetry conventions

OpenLLMetry, maintained by Traceloop under Apache 2.0, is instrumentation rather than a backend. It extends OpenTelemetry so traces from LLM frameworks reach any compatible destination, and its README lists more than 20, including Datadog, Grafana, Honeycomb, New Relic, Splunk, Dynatrace and Sentry. By default it records prompts, completions and embeddings as span attributes; setting TRACELOOP_TRACE_CONTENT=false turns that off globally, and it can also be disabled per workflow. Traceloop joined ServiceNow in March 2026 and said OpenLLMetry will remain open source. Because its conventions have been folded into OpenTelemetry itself, instrumenting this way keeps the backend replaceable.

MLflow Tracing

MLflow, licensed Apache 2.0, describes its tracing as fully OpenTelemetry-compatible: it exports and ingests the GenAI conventions, and traces stay on infrastructure you host. Token usage is recorded from MLflow 3.2 and cost from 3.10, when the tracking server is installed with the [genai] extra. Cost relies on LiteLLM’s price table, so a model missing from it shows $0.00 even with correct token counts. Evaluation runs through datasets, scorers and built-in LLM-as-a-judge scorers, and human feedback is attached to traces with its metadata. It is the natural pick for teams already using MLflow for classic models.

Helicone

Helicone pairs an AI gateway with an observability platform, licensed Apache 2.0, that can be self-hosted with Docker or Helm. Mintlify acquired Helicone on 3 March 2026, and Helicone now runs in maintenance mode: security updates, bug fixes and new models keep shipping, and Mintlify says it will help customers migrate to another platform. Its omit headers stop request and response bodies from being stored but, as its documentation notes, the content is still sent to Helicone’s backend. Plan new deployments on another tool.

Commercial platforms

LangSmith

LangSmith, from LangChain, is a hosted platform for tracing, evaluation and prompt engineering. Developer seats cost $0 with 5,000 base traces a month and Plus costs $39 per seat with 10,000, then pay-as-you-go; Enterprise is custom (list, verified October 2026). Base traces are kept for 14 days and extended traces for 180. Self-hosting on Kubernetes, with ClickHouse, PostgreSQL and Redis inside your environment, is an add-on to the Enterprise plan. LangSmith ingests OpenTelemetry traces and maps GenAI attributes to its own fields, and masking happens on the client through LANGSMITH_HIDE_INPUTS, LANGSMITH_HIDE_OUTPUTS, regex anonymisers or integrations such as Presidio.

Datadog Agent Observability

Datadog’s product page now calls the product Agent Observability, while its documentation still uses the name LLM Observability. It traces LLM calls, fixed workflows and agent steps, compares prompts, models and agent configurations in experiments, tracks token cost, and uses Sensitive Data Scanner to redact sensitive data and flag prompt-injection attempts. The documentation says it supports the OpenTelemetry GenAI conventions natively. Pricing is per LLM span: free up to 40,000 spans a month, and Pro at $160 a month for 100,000 (list, verified October 2026). It is SaaS only, so trace content leaves your environment.

Grafana Cloud AI Observability

Grafana Cloud’s AI observability is OpenTelemetry-native and built with OpenLIT. It covers LLMs, vector databases, GPUs, MCP servers and agents, with spend tracking, budget management and automated checks for hallucinations, factual accuracy and toxicity. The OpenLIT SDK and its self-hosted dashboard on ClickHouse are Apache 2.0, so the same instrumentation can stay inside your network if Grafana Cloud is not an option.

W&B Weave

Weave traces LLM calls and functions and evaluates responses with LLM judges and custom scorers. It accepts OTLP traces on multi-tenant cloud, dedicated cloud and self-managed instances, and its Python SDK can redact personal data with Presidio before traces leave the client; other SDKs lack that option. Its documentation now lives on CoreWeave’s site and names CoreWeave Agent Lens, in public preview on multi-tenant SaaS, as the successor to Weave. The two share a backend, and Weave still covers evaluations and guardrails that Agent Lens does not yet offer.

Keeping prompts and personal data out of traces

Traces of an LLM application hold the most sensitive text the system ever sees, so treat the observability store like any other system of record:

  • Capture content only by decision. The GenAI conventions make message capture opt-in; keep it that way unless a team needs content to debug.
  • Mask at the source. Langfuse masking functions, LangSmith hide settings, OpenInference hide settings in Phoenix and Weave’s PII redaction all act before data leaves the process.
  • Redact in the pipeline. The OpenTelemetry Collector’s redaction processor deletes attributes that are not on an allowlist and masks values matching blocked patterns; it is beta for traces and alpha for logs and metrics.
  • Read what “omit” means. Helicone’s omit headers prevent storage, not transmission.
  • Set retention deliberately. LangSmith keeps base traces for 14 days; in self-hosted Langfuse, retention policies sit in the licensed enterprise modules.
  • Control access and export. Traces used as audit evidence need role-based access and a route into your SIEM.

How to choose an LLM observability tool

  1. Decide where trace data may live. If prompts contain regulated data, start with self-hostable tools or with SaaS whose masking runs in your process.
  2. Instrument with OpenTelemetry. The backend then becomes replaceable. Check which version of the GenAI conventions each candidate expects.
  3. Read licence and tier labels. Langfuse’s enterprise modules, Phoenix’s Elastic License and LangSmith’s Enterprise-only self-hosting all change the total cost.
  4. Check vendor status. Helicone is in maintenance mode, Weave has a named successor, and Langfuse and Traceloop now belong to larger companies.
  5. Price on your own unit. Spans, traces and seats scale differently, and an agent run can produce dozens of spans.
  6. Add gateway logs. If model traffic passes through an AI gateway, its per-call records are a second source of cost and latency data; our AI gateway comparison notes what each one logs.
  7. Tie evaluation to traces. Use offline test sets before release and online scores after, following an agent evaluation method and the practice of testing agents before production.

How VDF AI fits

For workflows that run on VDF AI Networks, observability is part of the platform. Every node execution is tracked with its status, input and output, token usage and cost. Latency is reported at P50 and P95 per network and per node, service-level objectives can be set per network with alerts on breaches, and execution traces include model routing decisions and their rationale.

Across the platform, the Trust Center describes audit records for every prompt, tool call, retrieval, model route and response, with actor, timestamp and context, streamed to your SIEM and kept in your environment, whether on-premises, in a private cloud or air-gapped. VDF AI Router adds the reason, candidate list and scores behind each routing decision. The product pages do not document an OpenTelemetry export, so if you standardise on one of the tools above, confirm the integration path during evaluation.

Sources

Frequently asked questions

What are LLM observability tools?

LLM observability tools capture what an application built on language models did for each request: the prompt and response, every model call, retrieval and tool call, latency, token counts, cost, errors and evaluation scores. They combine traces, metrics and quality signals so engineers can debug failures, find slow or expensive steps and catch quality regressions. Most now accept OpenTelemetry data, and many add prompt management, datasets and automated evaluators on top of tracing.

What is the best open-source LLM observability tool?

It depends on the licence and how you work. Langfuse is MIT licensed for its core features and self-hosts at scale, with SCIM, audit logging and retention policies in paid modules. MLflow Tracing is Apache 2.0 and suits teams already on MLflow. Arize Phoenix is source-available under the Elastic License 2.0 rather than an OSI licence. OpenLLMetry and OpenLIT are Apache 2.0 instrumentation that can feed any OpenTelemetry backend. Helicone's platform is open source but has been in maintenance mode since March 2026.

Which metrics should we track for LLM applications?

Track end-to-end latency and time to first token, generation speed per output token, input, output, cached and reasoning tokens, and cost per call, user and feature. Add error, timeout and retry rates, and for agents the number and duration of tool and model calls in each run. Then add the quality signals that infrastructure metrics miss: evaluator scores, user feedback and guardrail hits. The OpenTelemetry GenAI conventions name most of these signals, though every one is still marked Development.

Do LLM observability tools support OpenTelemetry?

Most do. Langfuse, LangSmith and W&B Weave accept OTLP traces, Phoenix is built on OpenTelemetry and OpenInference, MLflow exports and ingests the GenAI semantic conventions, and Datadog supports those conventions natively. OpenLLMetry and OpenLIT are themselves OpenTelemetry instrumentation. The conventions now live in their own repository and every signal is marked Development, so attribute names can change; check which version each backend expects before standardising.

How do we keep personal data out of LLM traces?

Capture less, then redact. The GenAI conventions make prompt and response content opt-in, and SDK settings in Langfuse, LangSmith, Phoenix and Weave mask or hide content before it leaves the application. An OpenTelemetry Collector running the redaction processor can drop attributes that are not on an allowlist and mask values that match blocked patterns. Keep raw traces for a short time, restrict who can read them, and check whether an omit feature stops only storage or transmission as well.

Is LLM observability different from agent observability?

They overlap. LLM observability grew around single model calls: prompts, responses, latency, tokens and cost. Agent observability extends the same records to multi-step runs, adding tool calls, retrievals, planning steps and decisions, and treats the whole trajectory as the unit to inspect. The tools in this list increasingly cover both, and the OpenTelemetry conventions now define agent and tool spans as well as model spans.

Filed under
AI observabilityAI audit trailsAI evaluationAI infrastructureopen-source AIAI governance
AI Governance

Is your AI governance audit-ready?

Get a readiness review of your AI controls — policy, oversight, audit trails, and EU AI Act evidence — mapped against what production actually requires.

Keep reading