A guard model is a language model trained, or prompted with a written policy, to classify prompts, responses or agent actions as safe or unsafe. The open-weight options in 2026 include Meta's Llama Guard 4 and Prompt Guard 2, OpenAI's gpt-oss-safeguard, Alibaba's Qwen3Guard, IBM's Granite Guardian 4.1 and Mistral's Shieldstral. Running them on your own hardware keeps the text they screen inside your boundary. They add a layer of defence; they do not replace permissions and approvals.
Why screening belongs inside your boundary
A guard model reads exactly what you are trying to protect: the prompt with a customer’s account details, the retrieved contract clause, the draft answer. A hosted moderation API receives a copy of that text on every request, which quietly undoes the reason for running AI on-premises. Self-hosted guard models screen content where it already sits, and their verdicts land in your own audit trail.
They are still models. They classify against a taxonomy or a policy, they make mistakes in both directions, and a determined attacker can evade them. Treat them as one layer beside the enforced guardrails that decide what an agent may do: tool allowlists, permission checks and approvals.
The open-weight guard models worth evaluating
| Model | Publisher, release | Sizes | Licence | Screens | Stands out for |
|---|---|---|---|---|---|
| Llama Guard 4 | Meta, Apr 2025 | 12B | Llama 4 Community License | Prompts and responses, text and images, 14 hazard categories | One multimodal model for input and output |
| Prompt Guard 2 | Meta, Apr 2025 | 86M, 22M | Llama 4 Community License | Prompt injection and jailbreak attempts | Small enough to run on every request |
| gpt-oss-safeguard | OpenAI with ROOST, Oct 2025 | 20B, 120B | Apache 2.0 | Any content, against a policy you write | Reasons over your policy and explains its verdict |
| Qwen3Guard | Alibaba Qwen, Sep 2025 | 0.6B, 4B, 8B | Apache 2.0 | Prompts and responses as Safe, Controversial or Unsafe | 119 languages and dialects; a Stream variant checks tokens as they are generated |
| Granite Guardian 4.1 | IBM, Apr 2026 | 8B | Apache 2.0 | Harm and jailbreaks, RAG groundedness and relevance, function-call hallucination | Checks built for RAG and agents; custom criteria |
| Nemotron Safety Guard 8B v3 | NVIDIA, Oct 2025 | 8B | NVIDIA Open Model License and Llama 3.1 terms | Prompts and responses, 23 categories | Nine languages; also packaged as an NVIDIA NIM |
| Shieldstral 1.0 | Mistral, Aug 2026 | 3B | Apache 2.0 | Prompts, responses and images | Policies written as plain-language questions |
The list sorts into three kinds. Fixed-taxonomy classifiers such as Llama Guard 4, Qwen3Guard and Nemotron are ready to use, but their categories encode someone else’s policy. Policy-following models, such as gpt-oss-safeguard, Shieldstral and Granite Guardian’s custom criteria, judge content against a definition you write. Specialist detectors answer one narrow question: Prompt Guard 2 looks only for attempts to override instructions.
Where the checks belong
- Incoming prompts. A small classifier such as Prompt Guard 2 or Qwen3Guard 0.6B screens for jailbreaks and policy violations. Meta measured Prompt Guard 2 22M at 19.3 ms to classify a 512-token input on an A100.
- Retrieved passages and tool results. Indirect injection arrives inside documents and API responses, so screen them before they enter the context. The prompt-injection containment guide covers the controls around this step.
- Agent actions. Before a tool call runs, check that it matches what the user asked for. Granite Guardian’s function-call check and the AlignmentCheck scanner in Meta’s open-source LlamaFirewall target this step. On the AgentDojo benchmark, Meta reports that Prompt Guard 2 plus AlignmentCheck cut attack success from 17.6% to 1.75%, while task utility fell from 47.7% to 42.7%.
- Responses. Screen answers against content policy and, for RAG, check groundedness. Qwen3Guard-Stream flags unsafe tokens as they are generated, so a long answer can be stopped mid-stream.
- Escalation. Send borderline verdicts to a larger policy-reasoning model or a human reviewer. OpenAI describes the same cascade for its own use: small, high-recall classifiers first, gpt-oss-safeguard after. It is compliance-aware routing applied to safety.
How to choose
- Policy fit. Read the taxonomy against your own rules. Llama Guard’s S6 “Specialized Advice” category can flag the financial, medical or legal guidance a regulated firm’s staff exist to give.
- Languages and modalities. Granite Guardian 4.1 is trained and tested on English only. Llama Guard 4 lists eight languages and accepts images. Qwen3Guard covers 119 languages and dialects.
- Latency and hardware. Sub-billion-parameter detectors can run on every request, while 8B to 12B models need GPU capacity. gpt-oss-safeguard-20b fits in 16 GB of VRAM. OpenAI notes that traditional classifiers trained on thousands of examples will likely beat it on a fixed task.
- Licence. Apache 2.0 covers most of the list; Meta’s models use the Llama 4 Community License. Check the terms as you would for any open-weight model.
- Your own test set. Label a few hundred of your own prompts, including benign edge cases, and measure both missed violations and blocked legitimate work before setting thresholds.
Limits to plan for
Research keeps finding ways past classifiers. Hackett and colleagues evaded six prompt-injection and jailbreak detectors with character injection and adversarial machine-learning techniques, in some cases with up to 100% evasion success. Nasr, Carlini and colleagues bypassed twelve recent defences with adaptive attacks, most at over 90% success, although most had originally reported near-zero attack success. Over-defence is the opposite failure: the InjecGuard study found prompt-guard models flagging benign text that merely contains trigger words. RabakBench found significant degradation across 13 guardrails in Singlish, Chinese, Malay and Tamil.
So treat verdicts as risk signals. Set thresholds per use case, log every verdict with the trace, route borderline cases to people, and keep deterministic controls on everything an agent can do. Register and version-pin guard models like any other governed local model, and re-run your test set before each upgrade.
How VDF AI fits
VDF AI Networks applies per-agent PII redaction to inputs and outputs before they reach any model or storage, and network-level content safety filters with configurable severity thresholds, automatic rejection and audit logging. Tool allowlists are enforced at the execution layer, so a missed classification cannot hand an agent a tool it was never granted. The open-weight models you select are registered, version-pinned and routed inside your own environment, so prompts, verdicts and outputs stay within your security boundary.
Sources and further reading
- Llama Guard 4 model card (Meta)
- Prompt Guard 2 model card (Meta)
- gpt-oss-safeguard user guide (OpenAI Cookbook)
- Qwen3Guard on GitHub (Alibaba Qwen)
- Granite Guardian 4.1 8B model card (IBM)
- Llama 3.1 Nemotron Safety Guard 8B v3 model card (NVIDIA)
- Shieldstral announcement (Mistral AI)
- LlamaFirewall: an open source guardrail system for AI agents
- Hackett et al., Bypassing LLM Guardrails
- Nasr et al., The Attacker Moves Second
- Li and Liu, InjecGuard
- Chua et al., RabakBench
Choosing guard models for a private AI platform? Talk to us about screening prompts, outputs and agent actions without sending them outside your infrastructure.