AI Infrastructure

LLM Evaluation Tools (2026): Open-Source Frameworks Compared

promptfoo, DeepEval, Ragas, OpenAI Evals, lm-evaluation-harness, Inspect, MLflow, Langfuse, Arize Phoenix and TruLens compared from their repositories and docs in October 2026: licence, maintenance, offline and online evaluation, RAG and agent metrics, LLM-as-a-judge and CI integration.

LLM evaluation tools score the outputs of language-model applications against test sets, rubrics and production samples. The leading open-source options in 2026 are promptfoo, DeepEval and Ragas for application and RAG tests, lm-evaluation-harness and the UK AI Security Institute's Inspect for model benchmarks, and Langfuse, Arize Phoenix, MLflow and TruLens for evaluation tied to tracing.

Tracing and dashboards are covered in our list of LLM observability tools; this page is about scoring. For the method behind a RAG scorecard, including permission tests and judge calibration, see our framework for measuring private RAG accuracy. Licences, release dates and features below were checked against each project’s repository, package registry and documentation on 6 October 2026.

Quick picks

If you needStart withWhy
Prompt and app tests in CI from a config filepromptfooDeterministic and model-graded assertions, --fail-on-error for pipelines
Pytest-style unit tests with RAG and agent metricsDeepEvaldeepeval test run, faithfulness, contextual recall, tool correctness, task completion
RAG metrics and synthetic test setsRagasContext precision and recall, faithfulness, test-data generation
Academic benchmarks on open-weight modelslm-evaluation-harness60+ benchmarks; vLLM, SGLang, Hugging Face and llama.cpp backends
Agent and safety evaluations with sandboxed toolsInspectSolvers, scorers, Docker or Kubernetes sandboxes, 200+ ready-made evals
Scoring production traces you already collectLangfuse or PhoenixLLM-as-a-judge on live traces plus datasets and experiments
Evaluation inside an existing MLflow estateMLflowmlflow.genai.evaluate with built-in judges and human feedback
The RAG triad with OpenTelemetry tracingTruLensContext relevance, groundedness and answer relevance

The LLM evaluation tools compared

ToolMaintainerLicenceLatest release (verified October 2026)Main useOnline scoring of production traffic
promptfooPromptfoo, part of OpenAIMIT0.124.0, 6 Oct 2026App and prompt tests, red teamingNo; offline and CI
DeepEvalConfident AIApache 2.04.2.8, 2 Oct 2026Pytest-style app testsThrough the Confident AI platform
RagasVibrant LabsApache 2.00.4.3, 13 Jan 2026RAG and agent metricsVia observability tool integrations
OpenAI EvalsOpenAIMIT (some datasets excepted)PyPI 3.0.1.post1, May 2024Eval framework and benchmark registryHosted evals in the OpenAI dashboard
lm-evaluation-harnessEleutherAIMIT0.4.13, 31 Aug 2026Model benchmarksNo
InspectUK AI Security Institute with Meridian LabsMIT0.3.276, 2 Oct 2026Model and agent evalsNo
MLflowMLflow projectApache 2.03.16.1, 16 Sep 2026Evaluation and tracing for GenAI appsYes, on traces
LangfuseLangfuse, acquired by ClickHouseMIT coreSDK 4.17.0, 5 Oct 2026Tracing, datasets, experiments, judgesYes, LLM-as-a-judge on traces
Arize PhoenixArize AIElastic License 2.020.19.0, 1 Oct 2026Tracing and evalsYes, on traces
TruLensSnowflakeMIT2.15.0, 5 Oct 2026RAG and agent evals with tracingScores attach to traced runs

Phoenix’s Elastic License 2.0 publishes the source but bars providing the software to third parties as a hosted or managed service. GitHub stars on 6 October 2026: Langfuse about 35,400, MLflow about 28,300, promptfoo about 25,700, OpenAI Evals about 19,600, DeepEval about 18,700, Ragas about 15,900, lm-evaluation-harness about 14,100, Phoenix about 11,700, TruLens about 3,600 and Inspect about 2,900.

Offline and online evaluation

Offline evaluation runs a fixed dataset through the system before release. It answers whether this change is better or worse than the last one, and it belongs in CI. Online evaluation scores samples of real traffic after release, usually with a judge model attached to traces, plus user feedback and human annotation. It answers whether quality holds up on inputs nobody anticipated.

The two need different tools, and many teams run one of each:

  • Offline first: promptfoo, DeepEval, Ragas, lm-evaluation-harness, Inspect and OpenAI Evals run against datasets on demand.
  • Both: Langfuse supports LLM-as-a-judge on production traces, annotation queues and user feedback, and also datasets and experiments. Phoenix runs evaluations on production traces, experiment results or any dataset. MLflow and TruLens attach scores to traces as well.

Keep the same rubric in both places where you can. A faithfulness judge that gates a release offline should also score a slice of production traffic, so the two numbers can be compared.

Notes on each tool

promptfoo

A test file lists prompts, providers and assertions. Deterministic assertions include equals, contains, regex, is-json, is-refusal, custom javascript or python, latency, cost and tool-call-f1; model-graded ones include llm-rubric, g-eval, factuality, answer-relevance, context-faithfulness, context-recall and context-relevance. The README says evals run locally and that your prompts never leave your machine, and the CI guide covers GitHub Actions, GitLab CI, Jenkins, Azure Pipelines, CircleCI and others. promptfoo is now part of OpenAI and remains MIT-licensed. Its red-team mode is covered in our list of AI red teaming tools.

DeepEval

DeepEval describes itself as similar to pytest but specialised for unit testing LLM apps. Metrics span G-Eval and DAG for custom criteria; answer relevancy, faithfulness and contextual precision, recall and relevancy for RAG; task completion, tool correctness, goal accuracy, step efficiency and plan adherence for agents; and turn-level metrics for conversations. deepeval test run slots into CI. The judge defaults to OpenAI, and deepeval set-ollama switches every metric to a local Ollama model. Confident AI, the company behind it, sells an optional hosted platform.

Ragas

Ragas focuses on RAG and agent metrics plus synthetic test-set generation. Its catalogue includes context precision, context recall, context entities recall, noise sensitivity, response relevancy and faithfulness, agent metrics such as tool call accuracy, tool call F1 and agent goal accuracy, SQL equivalence, and rubric-based general-purpose scores. Maintenance has slowed: the last PyPI release was 0.4.3 on 13 January 2026 and the last push to the repository was on 24 February 2026, so check activity before you standardise.

OpenAI Evals

The original open-source framework and benchmark registry from OpenAI. The repository now notes that you can configure and run evals directly in the OpenAI dashboard, says it is not accepting evals with custom code, and requires an OpenAI API key. The last PyPI release dates from May 2024. It is useful as a source of eval designs, less so as a framework for non-OpenAI or local models.

lm-evaluation-harness

EleutherAI’s harness runs more than 60 standard academic benchmarks with hundreds of subtasks against Hugging Face Transformers, vLLM, SGLang, OpenAI-compatible APIs, GGUF models through llama.cpp, NVIDIA NeMo and ONNX Runtime. Its README says it has been used in hundreds of papers and by organisations including NVIDIA and Cohere. Use it to compare open-weight checkpoints, including the exact quantized build you plan to deploy; it does not test your application.

Inspect

Inspect, from the UK AI Security Institute with Meridian Labs, builds evaluations from datasets, solvers, scorers, agents and tools. Tool use can run in Docker, Kubernetes, Modal, Proxmox or Vagrant sandboxes, which matters for agent evals that execute code. It supports more than 20 model providers plus local inference through Hugging Face, vLLM and SGLang, ships more than 200 pre-built evaluations, and includes a log viewer and a VS Code extension.

MLflow

MLflow’s GenAI evaluation runs mlflow.genai.evaluate() over datasets with scorers. Predefined judges include Correctness, RelevanceToQuery, Completeness, Guidelines, RetrievalGroundedness, RetrievalSufficiency, ToolCallCorrectness, ToolCallEfficiency and multi-turn judges such as KnowledgeRetention. The documentation notes that Safety and RetrievalRelevance are currently available only in Databricks managed MLflow. Human feedback attaches to traces, so offline and production scores live in one place.

Langfuse

Langfuse combines tracing with LLM-as-a-judge evaluators, annotation queues, custom scores through the API and SDK, user feedback, datasets and experiments. Its licensing page states that all product capabilities, including evaluations, experiments and annotation, are MIT-licensed without usage limits; SCIM, audit logging and data-retention policies are separately licensed enterprise modules when self-hosting. A GitHub Action can fail a deployment when experiment scores fall below thresholds. ClickHouse acquired Langfuse on 16 January 2026 and committed to keeping the core MIT-licensed and self-hostable.

Arize Phoenix

Phoenix offers LLM-as-a-judge evaluators with pre-built templates for hallucination, QA correctness, relevance, toxicity, RAG relevance and tool calling, plus code-based evaluators such as exact match and regex. It is model-agnostic through adapters for OpenAI, LiteLLM, LangChain and the AI SDK, and handles rate limits and concurrency for large runs. Check the Elastic License 2.0 against your plans before you embed it in a service you offer to others.

TruLens

TruLens popularised the RAG triad: context relevance, groundedness and answer relevance. It also scores agents on tool selection, plan adherence and execution efficiency, is OpenTelemetry-native, and works with OpenAI, Anthropic, Google Gemini and local models through Ollama or LiteLLM. Snowflake maintains it as open source after acquiring TruEra.

RAG and agent metrics side by side

The same idea often has a different name in each tool. The table maps the common ones, as named in each project’s documentation.

What it measuresRagasDeepEvalTruLensMLflow judgepromptfoo assertion
Answer supported by retrieved contextFaithfulnessFaithfulnessGroundednessRetrievalGroundednesscontext-faithfulness
Answer addresses the questionResponse RelevancyAnswer relevancyAnswer relevanceRelevanceToQueryanswer-relevance
Retrieved context is relevantContext PrecisionContextual precision and relevancyContext relevanceRetrievalRelevance (Databricks-managed only)context-relevance
Context contains what the answer needsContext RecallContextual recall–RetrievalSufficiencycontext-recall
Correct tool callsTool Call Accuracy, Tool Call F1Tool correctnessTool selectionToolCallCorrectnesstool-call-f1
Agent reached the goalAgent Goal AccuracyTask completionPlan adherence–trajectory:goal-success

Two cautions apply. Scores with the same name are not comparable across tools, because each uses its own prompts and scales. And none of these metrics checks whether the user was entitled to see the retrieved documents, which is why our RAG framework treats permission correctness as a separate test set.

Using LLM-as-a-judge without fooling yourself

Most application metrics above are scored by a judge model. The method’s reference study, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, found that strong judges agreed with human preferences over 80% of the time, matching agreement between humans, and documented position, verbosity and self-enhancement biases plus limited reasoning ability. Practical rules follow from that:

  1. Pin the judge. Record the judge model, version and prompt with every score. Changing the judge changes the numbers without changing the system.
  2. Write narrow rubrics. One criterion per judge call is easier to calibrate than a holistic quality score.
  3. Control order and length. Swap positions in pairwise comparisons and watch for longer answers winning by default.
  4. Avoid self-grading. Use a judge from a different model family from the system under test where you can.
  5. Calibrate against people. Have subject-matter experts score a fixed sample each cycle and track agreement with the judge.
  6. Keep the judge local when the data is sensitive. A judge sees both the question and the retrieved passages.

Wiring evaluation into CI

A dataset that only runs when someone remembers is not a regression suite. The tools differ in how they fail a build:

  • promptfoo returns a failing exit code with --fail-on-error, and its CI guide shows how to parse JSON results to enforce a pass-rate threshold.
  • DeepEval runs tests with deepeval test run, the same way pytest runs unit tests.
  • Langfuse experiments can run in a GitHub Action that raises an error and blocks the deployment when thresholds are violated.
  • lm-evaluation-harness and Inspect are better run on a schedule or when a model changes, because benchmark suites take longer than a pull-request check.

Trigger runs on the changes that move quality: model upgrades, prompt edits, retrieval settings, new tools and corpus refreshes. For agents, add the trajectory checks described in our agent evaluation explainer, and pair functional evals with the adversarial suite from your red-team work.

How to choose an LLM evaluation tool

  1. Decide what you are testing. A model checkpoint calls for lm-evaluation-harness or Inspect. An application or RAG pipeline calls for promptfoo, DeepEval, Ragas or TruLens. Live quality calls for an evaluator attached to traces.
  2. Decide where data may go. Self-hosted judges and self-hosted platforms keep test sets and traces inside your boundary.
  3. Read the licence. MIT and Apache 2.0 cover most of this list; Phoenix is Elastic License 2.0, and Langfuse keeps some admin features in enterprise modules.
  4. Check maintenance. Ragas and OpenAI Evals have released less often than the others in 2026.
  5. Fit your stack. MLflow suits teams already on MLflow; Langfuse and Phoenix suit teams that want tracing and evaluation in one store.
  6. Start small. A hundred real questions with expected sources, scored by a pinned judge in CI, beats a large benchmark nobody runs. The academy lesson on testing agents before production walks through building a first test set with expected answers.

How VDF AI fits

The Model Evaluation Suite is VDF AI’s evaluation layer for comparing models on your own scenarios. Domain experts store test cases with a prompt, context and expected answer; the suite runs them in batch across candidate models, including local Ollama deployments, and scores every response with BLEU, ROUGE-L, METEOR and BERTScore. It flags regressions between versions and keeps timestamped results tied to model versions for audit.

The service runs inside your VDF deployment, on-premises or air-gapped, so prompts, reference answers and outputs stay on your infrastructure. Teams that use the judge-based tools above for application-level metrics can keep both sets of results in the same environment.

Sources

Frequently asked questions

What are LLM evaluation tools?

LLM evaluation tools measure whether a language-model application gives correct, grounded, safe and useful outputs. They run a set of test inputs through a model, prompt, retrieval pipeline or agent and score the results with deterministic checks, reference metrics or a judge model. Offline evaluation runs before release on fixed datasets; online evaluation scores samples of production traffic. The scores let teams compare models and catch regressions when prompts, models or data change.

What is the best open-source LLM evaluation tool?

There is no single best tool, because they test different things. promptfoo suits prompt and application tests in CI with a YAML file. DeepEval suits Python teams who want pytest-style tests with RAG and agent metrics. Ragas specialises in RAG metrics and test-set generation. lm-evaluation-harness and Inspect run benchmarks against models. Langfuse, Phoenix, MLflow and TruLens add evaluation to tracing, which is where online scoring happens.

What is the difference between LLM evaluation and LLM observability?

Evaluation asks whether outputs are good, against a dataset or a rubric. Observability records what the application did on each request: prompts, model and tool calls, latency, tokens and cost. The two meet in online evaluation, where judge models or users score production traces. Platforms such as Langfuse, Phoenix and MLflow do both, while libraries such as Ragas, DeepEval and lm-evaluation-harness focus on scoring and leave tracing to other tools.

Is LLM-as-a-judge reliable?

Reliable enough to scale review, not to replace it. The study that popularised the method found that strong judge models agreed with human preferences more than 80 percent of the time, about as often as humans agree with each other. It also documented position, verbosity and self-enhancement biases. Pin the judge model and prompt, write narrow rubrics, swap answer order in pairwise tests, and check a sample of verdicts against human reviewers on every cycle.

Can LLM evaluation run on-premises with local models?

Yes. Most open-source evaluators accept a self-hosted judge. DeepEval can make an Ollama model its default judge with one command, Phoenix and TruLens reach local models through adapters such as LiteLLM or Ollama, and lm-evaluation-harness and Inspect evaluate models served by vLLM, SGLang or Hugging Face on your own hardware. Langfuse, Phoenix, MLflow and TruLens can be self-hosted, so datasets, traces and scores stay inside your network.

Filed under
AI evaluationopen-source AIAI governanceAI observabilityRAGAI infrastructure
AI Governance

Is your AI governance audit-ready?

Get a readiness review of your AI controls — policy, oversight, audit trails, and EU AI Act evidence — mapped against what production actually requires.

Keep reading