AI Security

Open-Weight Guard Models for On-Premises AI: Llama Guard, gpt-oss-safeguard, Qwen3Guard and Granite Guardian Compared

Guard models are safety classifiers that screen prompts, responses and agent actions. Which open-weight guard models an enterprise can run on its own infrastructure in 2026, what each one checks, where the checks belong in an agent pipeline, and what they cannot do.

Three padlocks of increasing size lit from red to green, one closed and two open, representing the graded safe and unsafe verdicts that open-weight guard models return inside an on-premises AI platform
Photo by FlyD on Unsplash

A guard model is a language model trained, or prompted with a written policy, to classify prompts, responses or agent actions as safe or unsafe. The open-weight options in 2026 include Meta's Llama Guard 4 and Prompt Guard 2, OpenAI's gpt-oss-safeguard, Alibaba's Qwen3Guard, IBM's Granite Guardian 4.1 and Mistral's Shieldstral. Running them on your own hardware keeps the text they screen inside your boundary. They add a layer of defence; they do not replace permissions and approvals.

Why screening belongs inside your boundary

A guard model reads exactly what you are trying to protect: the prompt with a customer’s account details, the retrieved contract clause, the draft answer. A hosted moderation API receives a copy of that text on every request, which quietly undoes the reason for running AI on-premises. Self-hosted guard models screen content where it already sits, and their verdicts land in your own audit trail.

They are still models. They classify against a taxonomy or a policy, they make mistakes in both directions, and a determined attacker can evade them. Treat them as one layer beside the enforced guardrails that decide what an agent may do: tool allowlists, permission checks and approvals.

The open-weight guard models worth evaluating

ModelPublisher, releaseSizesLicenceScreensStands out for
Llama Guard 4Meta, Apr 202512BLlama 4 Community LicensePrompts and responses, text and images, 14 hazard categoriesOne multimodal model for input and output
Prompt Guard 2Meta, Apr 202586M, 22MLlama 4 Community LicensePrompt injection and jailbreak attemptsSmall enough to run on every request
gpt-oss-safeguardOpenAI with ROOST, Oct 202520B, 120BApache 2.0Any content, against a policy you writeReasons over your policy and explains its verdict
Qwen3GuardAlibaba Qwen, Sep 20250.6B, 4B, 8BApache 2.0Prompts and responses as Safe, Controversial or Unsafe119 languages and dialects; a Stream variant checks tokens as they are generated
Granite Guardian 4.1IBM, Apr 20268BApache 2.0Harm and jailbreaks, RAG groundedness and relevance, function-call hallucinationChecks built for RAG and agents; custom criteria
Nemotron Safety Guard 8B v3NVIDIA, Oct 20258BNVIDIA Open Model License and Llama 3.1 termsPrompts and responses, 23 categoriesNine languages; also packaged as an NVIDIA NIM
Shieldstral 1.0Mistral, Aug 20263BApache 2.0Prompts, responses and imagesPolicies written as plain-language questions

The list sorts into three kinds. Fixed-taxonomy classifiers such as Llama Guard 4, Qwen3Guard and Nemotron are ready to use, but their categories encode someone else’s policy. Policy-following models, such as gpt-oss-safeguard, Shieldstral and Granite Guardian’s custom criteria, judge content against a definition you write. Specialist detectors answer one narrow question: Prompt Guard 2 looks only for attempts to override instructions.

Where the checks belong

  1. Incoming prompts. A small classifier such as Prompt Guard 2 or Qwen3Guard 0.6B screens for jailbreaks and policy violations. Meta measured Prompt Guard 2 22M at 19.3 ms to classify a 512-token input on an A100.
  2. Retrieved passages and tool results. Indirect injection arrives inside documents and API responses, so screen them before they enter the context. The prompt-injection containment guide covers the controls around this step.
  3. Agent actions. Before a tool call runs, check that it matches what the user asked for. Granite Guardian’s function-call check and the AlignmentCheck scanner in Meta’s open-source LlamaFirewall target this step. On the AgentDojo benchmark, Meta reports that Prompt Guard 2 plus AlignmentCheck cut attack success from 17.6% to 1.75%, while task utility fell from 47.7% to 42.7%.
  4. Responses. Screen answers against content policy and, for RAG, check groundedness. Qwen3Guard-Stream flags unsafe tokens as they are generated, so a long answer can be stopped mid-stream.
  5. Escalation. Send borderline verdicts to a larger policy-reasoning model or a human reviewer. OpenAI describes the same cascade for its own use: small, high-recall classifiers first, gpt-oss-safeguard after. It is compliance-aware routing applied to safety.

How to choose

  • Policy fit. Read the taxonomy against your own rules. Llama Guard’s S6 “Specialized Advice” category can flag the financial, medical or legal guidance a regulated firm’s staff exist to give.
  • Languages and modalities. Granite Guardian 4.1 is trained and tested on English only. Llama Guard 4 lists eight languages and accepts images. Qwen3Guard covers 119 languages and dialects.
  • Latency and hardware. Sub-billion-parameter detectors can run on every request, while 8B to 12B models need GPU capacity. gpt-oss-safeguard-20b fits in 16 GB of VRAM. OpenAI notes that traditional classifiers trained on thousands of examples will likely beat it on a fixed task.
  • Licence. Apache 2.0 covers most of the list; Meta’s models use the Llama 4 Community License. Check the terms as you would for any open-weight model.
  • Your own test set. Label a few hundred of your own prompts, including benign edge cases, and measure both missed violations and blocked legitimate work before setting thresholds.

Limits to plan for

Research keeps finding ways past classifiers. Hackett and colleagues evaded six prompt-injection and jailbreak detectors with character injection and adversarial machine-learning techniques, in some cases with up to 100% evasion success. Nasr, Carlini and colleagues bypassed twelve recent defences with adaptive attacks, most at over 90% success, although most had originally reported near-zero attack success. Over-defence is the opposite failure: the InjecGuard study found prompt-guard models flagging benign text that merely contains trigger words. RabakBench found significant degradation across 13 guardrails in Singlish, Chinese, Malay and Tamil.

So treat verdicts as risk signals. Set thresholds per use case, log every verdict with the trace, route borderline cases to people, and keep deterministic controls on everything an agent can do. Register and version-pin guard models like any other governed local model, and re-run your test set before each upgrade.

How VDF AI fits

VDF AI Networks applies per-agent PII redaction to inputs and outputs before they reach any model or storage, and network-level content safety filters with configurable severity thresholds, automatic rejection and audit logging. Tool allowlists are enforced at the execution layer, so a missed classification cannot hand an agent a tool it was never granted. The open-weight models you select are registered, version-pinned and routed inside your own environment, so prompts, verdicts and outputs stay within your security boundary.

Sources and further reading


Choosing guard models for a private AI platform? Talk to us about screening prompts, outputs and agent actions without sending them outside your infrastructure.

Frequently asked questions

What is a guard model in AI?

A guard model is a classifier, usually a small language model, that reads a prompt, a response or a proposed agent action and returns a verdict such as safe or unsafe, often with the policy category that was breached. It runs beside the main model rather than inside it, so it can be replaced, tuned and audited on its own schedule.

Can a guard model stop prompt injection on its own?

No. Detectors such as Prompt Guard 2 reduce how many injection attempts get through, but published adaptive and obfuscation attacks have bypassed most of the detectors they tested. Pair detection with containment: least-privilege tools, no standing credentials, approval for irreversible actions and a full audit trail.

Which open-weight guard models accept a custom policy?

gpt-oss-safeguard reads a written policy supplied at inference time and explains its decision. Mistral's Shieldstral takes each policy as a plain-language yes-or-no question and returns a score. Granite Guardian 4.1 accepts custom criteria alongside its built-in checks. Fixed-taxonomy models such as Llama Guard 4 and Qwen3Guard are tuned to their own categories.

Filed under
AI securityopen-weight modelslocal LLMsmall language modelson-premises AIAI governance
AI Governance

Is your AI governance audit-ready?

Get a readiness review of your AI controls — policy, oversight, audit trails, and EU AI Act evidence — mapped against what production actually requires.

Keep reading