LLM evaluation tools score the outputs of language-model applications against test sets, rubrics and production samples. The leading open-source options in 2026 are promptfoo, DeepEval and Ragas for application and RAG tests, lm-evaluation-harness and the UK AI Security Institute's Inspect for model benchmarks, and Langfuse, Arize Phoenix, MLflow and TruLens for evaluation tied to tracing.
Tracing and dashboards are covered in our list of LLM observability tools; this page is about scoring. For the method behind a RAG scorecard, including permission tests and judge calibration, see our framework for measuring private RAG accuracy. Licences, release dates and features below were checked against each project’s repository, package registry and documentation on 6 October 2026.
Quick picks
| If you need | Start with | Why |
|---|---|---|
| Prompt and app tests in CI from a config file | promptfoo | Deterministic and model-graded assertions, --fail-on-error for pipelines |
| Pytest-style unit tests with RAG and agent metrics | DeepEval | deepeval test run, faithfulness, contextual recall, tool correctness, task completion |
| RAG metrics and synthetic test sets | Ragas | Context precision and recall, faithfulness, test-data generation |
| Academic benchmarks on open-weight models | lm-evaluation-harness | 60+ benchmarks; vLLM, SGLang, Hugging Face and llama.cpp backends |
| Agent and safety evaluations with sandboxed tools | Inspect | Solvers, scorers, Docker or Kubernetes sandboxes, 200+ ready-made evals |
| Scoring production traces you already collect | Langfuse or Phoenix | LLM-as-a-judge on live traces plus datasets and experiments |
| Evaluation inside an existing MLflow estate | MLflow | mlflow.genai.evaluate with built-in judges and human feedback |
| The RAG triad with OpenTelemetry tracing | TruLens | Context relevance, groundedness and answer relevance |
The LLM evaluation tools compared
| Tool | Maintainer | Licence | Latest release (verified October 2026) | Main use | Online scoring of production traffic |
|---|---|---|---|---|---|
| promptfoo | Promptfoo, part of OpenAI | MIT | 0.124.0, 6 Oct 2026 | App and prompt tests, red teaming | No; offline and CI |
| DeepEval | Confident AI | Apache 2.0 | 4.2.8, 2 Oct 2026 | Pytest-style app tests | Through the Confident AI platform |
| Ragas | Vibrant Labs | Apache 2.0 | 0.4.3, 13 Jan 2026 | RAG and agent metrics | Via observability tool integrations |
| OpenAI Evals | OpenAI | MIT (some datasets excepted) | PyPI 3.0.1.post1, May 2024 | Eval framework and benchmark registry | Hosted evals in the OpenAI dashboard |
| lm-evaluation-harness | EleutherAI | MIT | 0.4.13, 31 Aug 2026 | Model benchmarks | No |
| Inspect | UK AI Security Institute with Meridian Labs | MIT | 0.3.276, 2 Oct 2026 | Model and agent evals | No |
| MLflow | MLflow project | Apache 2.0 | 3.16.1, 16 Sep 2026 | Evaluation and tracing for GenAI apps | Yes, on traces |
| Langfuse | Langfuse, acquired by ClickHouse | MIT core | SDK 4.17.0, 5 Oct 2026 | Tracing, datasets, experiments, judges | Yes, LLM-as-a-judge on traces |
| Arize Phoenix | Arize AI | Elastic License 2.0 | 20.19.0, 1 Oct 2026 | Tracing and evals | Yes, on traces |
| TruLens | Snowflake | MIT | 2.15.0, 5 Oct 2026 | RAG and agent evals with tracing | Scores attach to traced runs |
Phoenix’s Elastic License 2.0 publishes the source but bars providing the software to third parties as a hosted or managed service. GitHub stars on 6 October 2026: Langfuse about 35,400, MLflow about 28,300, promptfoo about 25,700, OpenAI Evals about 19,600, DeepEval about 18,700, Ragas about 15,900, lm-evaluation-harness about 14,100, Phoenix about 11,700, TruLens about 3,600 and Inspect about 2,900.
Offline and online evaluation
Offline evaluation runs a fixed dataset through the system before release. It answers whether this change is better or worse than the last one, and it belongs in CI. Online evaluation scores samples of real traffic after release, usually with a judge model attached to traces, plus user feedback and human annotation. It answers whether quality holds up on inputs nobody anticipated.
The two need different tools, and many teams run one of each:
- Offline first: promptfoo, DeepEval, Ragas, lm-evaluation-harness, Inspect and OpenAI Evals run against datasets on demand.
- Both: Langfuse supports LLM-as-a-judge on production traces, annotation queues and user feedback, and also datasets and experiments. Phoenix runs evaluations on production traces, experiment results or any dataset. MLflow and TruLens attach scores to traces as well.
Keep the same rubric in both places where you can. A faithfulness judge that gates a release offline should also score a slice of production traffic, so the two numbers can be compared.
Notes on each tool
promptfoo
A test file lists prompts, providers and assertions. Deterministic assertions include equals, contains, regex, is-json, is-refusal, custom javascript or python, latency, cost and tool-call-f1; model-graded ones include llm-rubric, g-eval, factuality, answer-relevance, context-faithfulness, context-recall and context-relevance. The README says evals run locally and that your prompts never leave your machine, and the CI guide covers GitHub Actions, GitLab CI, Jenkins, Azure Pipelines, CircleCI and others. promptfoo is now part of OpenAI and remains MIT-licensed. Its red-team mode is covered in our list of AI red teaming tools.
DeepEval
DeepEval describes itself as similar to pytest but specialised for unit testing LLM apps. Metrics span G-Eval and DAG for custom criteria; answer relevancy, faithfulness and contextual precision, recall and relevancy for RAG; task completion, tool correctness, goal accuracy, step efficiency and plan adherence for agents; and turn-level metrics for conversations. deepeval test run slots into CI. The judge defaults to OpenAI, and deepeval set-ollama switches every metric to a local Ollama model. Confident AI, the company behind it, sells an optional hosted platform.
Ragas
Ragas focuses on RAG and agent metrics plus synthetic test-set generation. Its catalogue includes context precision, context recall, context entities recall, noise sensitivity, response relevancy and faithfulness, agent metrics such as tool call accuracy, tool call F1 and agent goal accuracy, SQL equivalence, and rubric-based general-purpose scores. Maintenance has slowed: the last PyPI release was 0.4.3 on 13 January 2026 and the last push to the repository was on 24 February 2026, so check activity before you standardise.
OpenAI Evals
The original open-source framework and benchmark registry from OpenAI. The repository now notes that you can configure and run evals directly in the OpenAI dashboard, says it is not accepting evals with custom code, and requires an OpenAI API key. The last PyPI release dates from May 2024. It is useful as a source of eval designs, less so as a framework for non-OpenAI or local models.
lm-evaluation-harness
EleutherAI’s harness runs more than 60 standard academic benchmarks with hundreds of subtasks against Hugging Face Transformers, vLLM, SGLang, OpenAI-compatible APIs, GGUF models through llama.cpp, NVIDIA NeMo and ONNX Runtime. Its README says it has been used in hundreds of papers and by organisations including NVIDIA and Cohere. Use it to compare open-weight checkpoints, including the exact quantized build you plan to deploy; it does not test your application.
Inspect
Inspect, from the UK AI Security Institute with Meridian Labs, builds evaluations from datasets, solvers, scorers, agents and tools. Tool use can run in Docker, Kubernetes, Modal, Proxmox or Vagrant sandboxes, which matters for agent evals that execute code. It supports more than 20 model providers plus local inference through Hugging Face, vLLM and SGLang, ships more than 200 pre-built evaluations, and includes a log viewer and a VS Code extension.
MLflow
MLflow’s GenAI evaluation runs mlflow.genai.evaluate() over datasets with scorers. Predefined judges include Correctness, RelevanceToQuery, Completeness, Guidelines, RetrievalGroundedness, RetrievalSufficiency, ToolCallCorrectness, ToolCallEfficiency and multi-turn judges such as KnowledgeRetention. The documentation notes that Safety and RetrievalRelevance are currently available only in Databricks managed MLflow. Human feedback attaches to traces, so offline and production scores live in one place.
Langfuse
Langfuse combines tracing with LLM-as-a-judge evaluators, annotation queues, custom scores through the API and SDK, user feedback, datasets and experiments. Its licensing page states that all product capabilities, including evaluations, experiments and annotation, are MIT-licensed without usage limits; SCIM, audit logging and data-retention policies are separately licensed enterprise modules when self-hosting. A GitHub Action can fail a deployment when experiment scores fall below thresholds. ClickHouse acquired Langfuse on 16 January 2026 and committed to keeping the core MIT-licensed and self-hostable.
Arize Phoenix
Phoenix offers LLM-as-a-judge evaluators with pre-built templates for hallucination, QA correctness, relevance, toxicity, RAG relevance and tool calling, plus code-based evaluators such as exact match and regex. It is model-agnostic through adapters for OpenAI, LiteLLM, LangChain and the AI SDK, and handles rate limits and concurrency for large runs. Check the Elastic License 2.0 against your plans before you embed it in a service you offer to others.
TruLens
TruLens popularised the RAG triad: context relevance, groundedness and answer relevance. It also scores agents on tool selection, plan adherence and execution efficiency, is OpenTelemetry-native, and works with OpenAI, Anthropic, Google Gemini and local models through Ollama or LiteLLM. Snowflake maintains it as open source after acquiring TruEra.
RAG and agent metrics side by side
The same idea often has a different name in each tool. The table maps the common ones, as named in each project’s documentation.
| What it measures | Ragas | DeepEval | TruLens | MLflow judge | promptfoo assertion |
|---|---|---|---|---|---|
| Answer supported by retrieved context | Faithfulness | Faithfulness | Groundedness | RetrievalGroundedness | context-faithfulness |
| Answer addresses the question | Response Relevancy | Answer relevancy | Answer relevance | RelevanceToQuery | answer-relevance |
| Retrieved context is relevant | Context Precision | Contextual precision and relevancy | Context relevance | RetrievalRelevance (Databricks-managed only) | context-relevance |
| Context contains what the answer needs | Context Recall | Contextual recall | – | RetrievalSufficiency | context-recall |
| Correct tool calls | Tool Call Accuracy, Tool Call F1 | Tool correctness | Tool selection | ToolCallCorrectness | tool-call-f1 |
| Agent reached the goal | Agent Goal Accuracy | Task completion | Plan adherence | – | trajectory:goal-success |
Two cautions apply. Scores with the same name are not comparable across tools, because each uses its own prompts and scales. And none of these metrics checks whether the user was entitled to see the retrieved documents, which is why our RAG framework treats permission correctness as a separate test set.
Using LLM-as-a-judge without fooling yourself
Most application metrics above are scored by a judge model. The method’s reference study, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, found that strong judges agreed with human preferences over 80% of the time, matching agreement between humans, and documented position, verbosity and self-enhancement biases plus limited reasoning ability. Practical rules follow from that:
- Pin the judge. Record the judge model, version and prompt with every score. Changing the judge changes the numbers without changing the system.
- Write narrow rubrics. One criterion per judge call is easier to calibrate than a holistic quality score.
- Control order and length. Swap positions in pairwise comparisons and watch for longer answers winning by default.
- Avoid self-grading. Use a judge from a different model family from the system under test where you can.
- Calibrate against people. Have subject-matter experts score a fixed sample each cycle and track agreement with the judge.
- Keep the judge local when the data is sensitive. A judge sees both the question and the retrieved passages.
Wiring evaluation into CI
A dataset that only runs when someone remembers is not a regression suite. The tools differ in how they fail a build:
- promptfoo returns a failing exit code with
--fail-on-error, and its CI guide shows how to parse JSON results to enforce a pass-rate threshold. - DeepEval runs tests with
deepeval test run, the same way pytest runs unit tests. - Langfuse experiments can run in a GitHub Action that raises an error and blocks the deployment when thresholds are violated.
- lm-evaluation-harness and Inspect are better run on a schedule or when a model changes, because benchmark suites take longer than a pull-request check.
Trigger runs on the changes that move quality: model upgrades, prompt edits, retrieval settings, new tools and corpus refreshes. For agents, add the trajectory checks described in our agent evaluation explainer, and pair functional evals with the adversarial suite from your red-team work.
How to choose an LLM evaluation tool
- Decide what you are testing. A model checkpoint calls for lm-evaluation-harness or Inspect. An application or RAG pipeline calls for promptfoo, DeepEval, Ragas or TruLens. Live quality calls for an evaluator attached to traces.
- Decide where data may go. Self-hosted judges and self-hosted platforms keep test sets and traces inside your boundary.
- Read the licence. MIT and Apache 2.0 cover most of this list; Phoenix is Elastic License 2.0, and Langfuse keeps some admin features in enterprise modules.
- Check maintenance. Ragas and OpenAI Evals have released less often than the others in 2026.
- Fit your stack. MLflow suits teams already on MLflow; Langfuse and Phoenix suit teams that want tracing and evaluation in one store.
- Start small. A hundred real questions with expected sources, scored by a pinned judge in CI, beats a large benchmark nobody runs. The academy lesson on testing agents before production walks through building a first test set with expected answers.
How VDF AI fits
The Model Evaluation Suite is VDF AI’s evaluation layer for comparing models on your own scenarios. Domain experts store test cases with a prompt, context and expected answer; the suite runs them in batch across candidate models, including local Ollama deployments, and scores every response with BLEU, ROUGE-L, METEOR and BERTScore. It flags regressions between versions and keeps timestamped results tied to model versions for audit.
The service runs inside your VDF deployment, on-premises or air-gapped, so prompts, reference answers and outputs stay on your infrastructure. Teams that use the judge-based tools above for application-level metrics can keep both sets of results in the same environment.
Sources
- promptfoo repository, assertions and CI/CD guide
- DeepEval repository and Ollama judge setup
- Ragas repository, metrics and PyPI package
- OpenAI Evals repository and licence
- lm-evaluation-harness repository
- Inspect documentation and repository
- MLflow evaluation and monitoring and predefined judges
- Langfuse evaluation overview and open-source licensing
- ClickHouse acquires Langfuse
- Phoenix LLM evals and repository
- Elastic License 2.0
- TruLens and repository
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Release dates from PyPI and the npm registry