AI Security

AI Red Teaming Tools (2026): Open-Source Options Compared

PyRIT, garak, promptfoo, DeepTeam, Giskard and specialist kits for prompt injection and agents, compared from their repositories and docs in October 2026: licence, maintenance, attack coverage, agent testing, data handling, and a map to the OWASP Top 10 for LLM applications and MITRE ATLAS.

AI red teaming tools automate adversarial testing of models, chatbots, retrieval pipelines and agents: they generate jailbreaks, prompt injections and data-leak attempts, then score the responses. The main open-source options in 2026 are Microsoft's PyRIT, NVIDIA's garak, promptfoo (now part of OpenAI), Confident AI's DeepTeam and Giskard, with specialist kits such as Spikee and AgentDojo for injection and agent testing.

This is the tools list. For running an exercise end to end (scoping, rules of engagement, the attack classes worth your time and how to write findings), see our guide to red-teaming an agent before production. Every licence, release date and capability below was checked against the projects’ repositories, package registries and documentation on 6 October 2026.

Quick picks

If you needStart withWhy
A fast vulnerability scan of a model or endpointgarakProbe families for injection, encoding, data replay, malware and more; JSONL report and hit log
A programmable framework for deep, multi-turn campaignsPyRITTargets, converters, scorers and multi-turn attacks such as Crescendo and TAP
Red-team tests in CI next to your evals, with framework presetspromptfooowasp:llm, mitre:atlas and agent plugins in one config file
A Python library that mirrors your DeepEval testsDeepTeam50+ vulnerabilities, 20+ attack methods, OWASP and ATLAS mappings
Multi-turn scans of a conversational agentGiskard v3giskard-scan vulnerability and quality scans for agents
Focused prompt-injection testingSpikeeA kit built for injection datasets and exploitation
A benchmark of injection attacks against tool-using agentsAgentDojoResearch environment for attacks and defences on agents

The open-source AI red teaming tools compared

ToolMaintainerLicenceLatest release (verified October 2026)How you use itAgent and tool testing
PyRITMicrosoftMIT1.1.0, 4 Sep 2026Python framework, pyrit_scan CLI, CoPyRIT GUICross-domain prompt injection workflows; web apps through Playwright
garakNVIDIAApache 2.00.17.0, 9 Sep 2026CLI scannerAgent breaker probe since 0.15.0 (May 2026)
promptfooPromptfoo, part of OpenAIMIT0.124.0, 6 Oct 2026YAML config, CLI, web viewerExcessive agency, BOLA, BFLA, tool discovery, memory poisoning, MCP plugins
DeepTeamConfident AIApache 2.01.0.9, 12 Aug 2026Python library11 agentic vulnerability types
Giskard (v3)GiskardApache 2.03.0.1, 2 Oct 2026Python packages giskard-checks and giskard-scanMulti-turn agent testing
SpikeeReversec LabsApache 2.0Repository active, last push 11 Sep 2026Prompt injection kitInjection evaluation and exploitation
FuzzyAICyberArkApache 2.0Last push 6 Feb 2026Automated fuzzerJailbreak fuzzing of LLM APIs
AgentDojoETH Zurich SPY LabMITLast push 2 Jun 2026Benchmark environmentAttacks and defences for tool-using agents

GitHub stars on 6 October 2026, as a rough signal of adoption: promptfoo about 25,700, garak about 9,400, Giskard about 5,900, PyRIT about 4,600 and DeepTeam about 3,000.

Notes on each tool

PyRIT

The Python Risk Identification Tool is Microsoft’s open-source framework for finding risks in generative AI systems. Its repository moved from Azure/PyRIT to microsoft/PyRIT, and the old repository was archived on 27 March 2026. The building blocks are targets (OpenAI, Azure, Anthropic, Google, Hugging Face, custom HTTP or WebSocket endpoints, and web apps driven through Playwright), converters that transform prompts, scorers (true or false, Likert, classification or custom, powered by an LLM, Azure AI Content Safety or your own logic) and a memory that records every conversation in SQLite or Azure SQL. Multi-turn attacks include Crescendo, TAP and Skeleton Key.

Two recent additions matter for enterprise teams: a scenario framework run through pyrit_scan that covers content harms, psychosocial risks and data leakage, and CoPyRIT, a graphical interface for human-led red teaming. Microsoft’s AI Red Teaming Agent in Microsoft Foundry is built on PyRIT’s capabilities.

garak

NVIDIA describes garak as “the LLM vulnerability scanner”. It checks for hallucination, data leakage, prompt injection, misinformation, toxic output and jailbreaks through probe families such as promptinject, dan, encoding, gcg, leakreplay, malwaregen, packagehallucination, snowball and xss. Generators cover Hugging Face, OpenAI, Amazon Bedrock, LiteLLM, REST endpoints, GGUF models through llama.cpp, NVIDIA NIM and Ollama. Version 0.15.0 added an agent breaker probe for testing the tools available to a target system, and 0.17.0 added an EU AI Act mapping that groups probe results by risk category. Each run writes a log, a JSONL report and a hit log of the attempts that succeeded.

promptfoo

promptfoo started as an evaluation tool and added red teaming on the same configuration format. Its repository now states that “Promptfoo is now part of OpenAI” and remains open source and MIT-licensed; OpenAI announced the acquisition in March 2026. Red-team runs combine plugins (what to test) with strategies (how to deliver it), from encodings such as base64 and leetspeak to iterative jailbreak and jailbreak:tree and multi-turn crescendo. Presets such as owasp:llm and mitre:atlas select plugins by framework. Some strategies, including goat and hydra, need remote inference in the Community edition.

DeepTeam

DeepTeam, from the makers of DeepEval, calls itself a framework to red team LLMs and AI agents. It ships more than 50 vulnerabilities and more than 20 attack methods, including multi-turn linear, tree and crescendo jailbreaking. Its agentic list covers goal theft, recursive hijacking, excessive agency, tool orchestration abuse, agent identity and trust abuse, and inter-agent communication compromise. It maps results to the OWASP Top 10 for LLMs 2025, OWASP’s agentic list, NIST AI RMF and MITRE ATLAS. It runs locally but needs a model to generate attacks and judge results; the default is an OpenAI key, and any model DeepEval supports can replace it.

Giskard

Giskard v3 is a full rewrite aimed at dynamic, multi-turn testing of agents, published in the giskard-oss repository. Its vulnerability_scan covers prompt injection, harmful content, stereotypes and misinformation across OWASP LLM Top 10 categories, and quality_scan replaces the v2 RAG evaluation toolkit. The project states that v2 remains available but is no longer actively maintained, so older tutorials based on the v2 LLM scan will not match current code.

Specialist tools

  • Spikee, is the Simple Prompt Injection Kit for Evaluation and Exploitation, now published by Reversec Labs.
  • FuzzyAI from CyberArk fuzzes LLM APIs for jailbreaks. Its last push was in February 2026, so check activity before you depend on it.
  • AgentDojo is a dynamic environment for evaluating attacks and defences on LLM agents, from ETH Zurich’s SPY Lab.

Mapping the tools to OWASP and MITRE ATLAS

Two frameworks give a red-team report shared vocabulary. The OWASP Top 10 for LLM Applications names the risk, and MITRE ATLAS names the adversary technique. The table uses OWASP’s 2025 numbering, which promptfoo’s owasp:llm preset and DeepTeam reference. The coverage column is our reading of each tool’s documentation, not a vendor certification.

OWASP risk (2025)ATLAS techniqueTools and probes that exercise it
LLM01 Prompt InjectionAML.T0051 LLM Prompt Injection (direct and indirect); AML.T0054 LLM Jailbreakgarak promptinject, dan, encoding; PyRIT Crescendo, TAP and cross-domain injection; promptfoo indirect-prompt-injection and jailbreak strategies; DeepTeam prompt injection; Spikee
LLM02 Sensitive Information DisclosureAML.T0057 LLM Data Leakagegarak leakreplay; promptfoo pii plugins; PyRIT data leakage scenarios; DeepTeam data privacy vulnerabilities
LLM03 Supply ChainAML.T0010 AI Supply Chain Compromisegarak packagehallucination for invented package names; promptfoo ModelAudit for model files
LLM04 Data and Model PoisoningAML.T0020 Poison Training Data; AML.T0070 RAG Poisoningpromptfoo rag-poisoning; training-data poisoning is largely out of reach for black-box tools
LLM05 Improper Output HandlingNo single technique; exfiltration through rendered outputgarak xss for cross-site and exfiltration payloads; promptfoo output assertions
LLM06 Excessive AgencyAML.T0053 AI Agent Tool Invocation; AML.T0086 Exfiltration via AI Agent Tool Invocationpromptfoo excessive-agency, bola, bfla, rbac, mcp; garak agent breaker; DeepTeam excessive agency and tool orchestration abuse; AgentDojo
LLM07 System Prompt LeakageAML.T0056 Extract LLM System Promptpromptfoo prompt-extraction; DeepTeam prompt probing
LLM08 Vector and Embedding WeaknessesAML.T0070 RAG Poisoningpromptfoo rag-document-exfiltration, bola and rbac against retrieval
LLM09 MisinformationAML.T0062 Discover LLM Hallucinationsgarak snowball and misleading; Giskard misinformation scan
LLM10 Unbounded ConsumptionAML.T0029 Denial of AI Service; AML.T0034 Cost Harvestingpromptfoo divergent-repetition

ATLAS technique names are taken from MITRE’s published ATLAS data (version 5.6.0). Agent memory has its own entry, AML.T0080 AI Agent Context Poisoning, which promptfoo’s agentic:memory-poisoning plugin and DeepTeam’s context poisoning attack target.

OWASP published a 2026 edition on 3 August 2026, and its announcement places Excessive Agency third. Imperva’s summary reports that no entries were added or removed and that System Prompt Leakage was broadened into Hidden Context Exposure. Check which edition a tool’s preset and your report use, because the numbers moved.

Red teaming agents, not just models

A jailbreak that makes a chatbot say something offensive is a content finding. An injection that makes an agent call a tool with the wrong arguments is an incident. Agent testing needs more than a prompt list:

  1. A target that exposes the real tools. Test the agent with the same tool definitions and permission model it has in production, against non-production systems. PyRIT’s Playwright and HTTP targets and garak’s REST generator reach an application, not just a bare model.
  2. Injection through data, not chat. Plant instructions in documents, emails, tickets and tool results the agent will read. OWASP’s agentic top 10, published in December 2025, puts agent goal hijack first and tool misuse second.
  3. Authorisation tests. promptfoo’s bola, bfla and rbac plugins check whether an agent will act on objects or functions the user should not reach. These findings are fixed in permissions, not prompts.
  4. Traces as evidence. promptfoo can grade agents from OpenTelemetry traces, so a pass reflects what the agent did, not what it claimed. Our list of LLM evaluation tools covers the tracing side.
  5. Regression runs. Turn confirmed findings into a suite that runs whenever the model, prompt, tools or connectors change.

Detection is only part of the answer. The prompt injection prevention guide explains why containment controls (least privilege, approvals and egress rules) carry most of the weight, and red-team results tell you whether they hold.

Keeping red-team data inside your network

Red-team artefacts are sensitive: successful payloads, system prompts, leaked records and the responses of an internal application. Check where each tool sends them.

  • Attacker and judge models. Most tools use an LLM to generate attacks and grade responses. DeepTeam defaults to an OpenAI key; DeepEval’s deepeval set-ollama command makes a local Ollama model the default judge instead. PyRIT and garak accept custom endpoints.
  • promptfoo’s hosted generation. Without an API key, promptfoo’s data handling page says it sends the application purpose, plugin configuration, target setup details (which can include request examples, URLs and auth headers) and your email for generation, plus the prompt, response and criteria for grading. Setting PROMPTFOO_DISABLE_REDTEAM_REMOTE_GENERATION=true forces local generation and PROMPTFOO_DISABLE_TELEMETRY=1 stops telemetry, but some plugins and strategies need hosted inference and cannot run locally.
  • Storage. PyRIT writes every conversation to its memory database, and garak writes JSONL reports and hit logs. Store them like security findings, with access control and retention.

Self-hosted open-weight screening models can also act as local judges for harmful-content scoring, which keeps the grading step inside the boundary too.

How to choose an AI red teaming tool

  1. Name the target. A bare model, a chat application, a retrieval pipeline or an agent with tools each need different probes.
  2. Decide who runs it. Security researchers usually prefer PyRIT’s code-first control; application teams prefer promptfoo’s or DeepTeam’s test-file style.
  3. Check data flow. Confirm which attacker and judge models run, and where. Turn off hosted generation if the system under test handles regulated data.
  4. Check maintenance. PyRIT, garak, promptfoo, DeepTeam and Giskard all released in the last two months; FuzzyAI has been quiet since February.
  5. Map results to frameworks. Choose tools whose presets match the framework your risk team reports against, and record the edition.
  6. Combine scanner and human. Automated tools find known patterns at volume. A person who understands the workflow finds the path from a small injection to a large consequence.

How VDF AI fits

VDF AI is not a red-teaming tool. It is the platform whose controls a red team tests, running on your own servers, in a private cloud, air-gapped, or in VDF’s managed cloud.

In VDF AI Networks, tool allowlists define which tools each agent can invoke and are enforced at the execution layer, Human Approval nodes hold steps until an owner accepts them, per-agent PII redaction applies to inputs and outputs, network-level content safety filters reject content above configurable severity thresholds, and every run leaves an execution trace with inputs, outputs and routing decisions. The MCP gateway grants tools per role from an administrator-governed registry and records each invocation against the calling agent.

Those properties give a red-team finding somewhere to land: a permission to remove, an approval to add, or a trace that shows the control held.

Sources

Frequently asked questions

What are AI red teaming tools?

AI red teaming tools send adversarial inputs to a model, chatbot, retrieval pipeline or agent and score whether it misbehaved. They automate attacks such as jailbreaks, direct and indirect prompt injection, system prompt extraction, data leakage and tool misuse, often with an attacker model that adapts over several turns. They then grade the responses with rules or a judge model and produce a report. They complement manual red-team exercises; they do not replace a human who understands the system.

What is the best open-source AI red teaming tool?

It depends on who runs it. Security researchers who want a programmable framework tend to choose Microsoft's PyRIT. Teams that want a quick scan of a model endpoint use NVIDIA's garak. Application teams that want red-team tests in CI next to their evals use promptfoo, which OpenAI acquired and keeps MIT-licensed. Python teams already on DeepEval add DeepTeam. Many organisations run two: a scanner for breadth and a framework for deeper multi-turn work.

Is PyRIT better than garak?

They solve different problems. garak is a scanner: you point it at a model or REST endpoint, pick probe families such as prompt injection, encoding tricks, data replay or malware generation, and get a report of hits. PyRIT is a framework: you write Python that combines targets, prompt converters, multi-turn attacks such as Crescendo and TAP, and scorers, and it keeps every conversation in a database. garak gives breadth quickly; PyRIT suits custom, deeper campaigns.

Can AI red teaming tools test AI agents and MCP servers?

Several now can. promptfoo has plugins for excessive agency, broken object and function level authorisation, tool discovery, memory poisoning and MCP. DeepTeam lists agentic vulnerabilities such as goal theft, tool orchestration abuse and inter-agent communication compromise. garak added an agent breaker probe that tests the tools available to a target, and PyRIT has workflows for cross-domain prompt injection. AgentDojo provides a benchmark environment for injection attacks against tool-using agents.

Can AI red teaming run without sending data to a vendor?

Yes, with the right tool and settings. garak, PyRIT, DeepTeam and Giskard run on your own machines, though each needs an attacker or judge model that you can point at a self-hosted endpoint. promptfoo sends some data to its hosted generation service by default; its documentation lists environment variables that disable remote generation and telemetry, and notes that some plugins and strategies then become unavailable. Check judge and attacker model settings before testing sensitive systems.

Filed under
AI securityopen-source AIAI agentsenterprise AI agentsAI risk managementAI governance
AI Governance

Is your AI governance audit-ready?

Get a readiness review of your AI controls — policy, oversight, audit trails, and EU AI Act evidence — mapped against what production actually requires.

Keep reading