AI Security

Prompt Injection Prevention for AI Agents: Techniques, Controls and a Cheat Sheet

No filter catches every prompt injection, so prevention for AI agents means limiting what injected text can make an agent do. This guide covers direct and indirect injection, the layered controls that hold up, the research behind them and a one-page cheat sheet.

Prompt injection prevention for AI agents starts from an uncomfortable fact: a language model can be talked into following instructions hidden in the content it reads, and no filter catches every attempt. The controls that hold up limit the damage instead. Give agents only the tools they need, require approval for consequential actions, mark and screen untrusted input, close exfiltration paths and log every tool call.

Direct and indirect prompt injection

Simon Willison named the attack in September 2022 and compared it to SQL injection: an application builds a prompt by joining its own instructions with text from elsewhere, and the model cannot reliably tell the two apart. The comparison breaks down at the fix. SQL injection has parameterised queries, while language models have no equivalent way to mark a span of text as data that must never be obeyed.

OWASP’s LLM01:2025 entry separates two forms:

  • Direct prompt injection. The user’s own input changes how the model behaves, for example a message telling it to ignore its system prompt and print it.
  • Indirect prompt injection. The model takes in content from an outside source, such as a website or a file, and instructions inside that content change what it does. NIST AI 100-2 E2025 describes this form as injection carried out through control of a resource rather than through the user’s input.

Agents read far more outside content than chat assistants do: inbound email, calendar invites, support tickets, shared documents, web pages, API responses and even the descriptions of the tools they can call. OWASP also points out that injected text does not have to be visible or readable to a person as long as the model parses it, so white-on-white text, HTML comments, image content and file metadata all count. The glossary entry has the short definition; this guide is about stopping it.

Why agents with tools raise the stakes

When a chatbot obeys injected text, the result is a wrong answer. When an agent obeys it, the agent acts with whatever access it holds: it sends the email, edits the record, opens the URL or runs the script. Willison’s lethal trifecta names the combination that turns injection into data theft: access to private data, exposure to untrusted content and a way to communicate externally. An assistant that reads a mailbox, summarises web pages and can make web requests has all three.

EchoLeak showed what that looks like in production. Tracked as CVE-2025-32711 and scored 9.3 by Microsoft, it was an AI command injection in Microsoft 365 Copilot that let an unauthorised attacker disclose information over a network. According to the researchers’ paper, one crafted email was enough, and the recipient never had to click anything. The exploit evaded Microsoft’s cross-prompt injection classifier, got links past redaction with reference-style Markdown, and moved data out through images that loaded automatically and a Microsoft Teams proxy that the content security policy allowed. Microsoft mitigated the flaw inside the service, so customers had nothing to install.

Two lessons stand out. A well-funded classifier still missed the attack, and the harm came from capabilities around the model: what it could read and which outbound requests were made on its behalf. OWASP says it is unclear whether fool-proof prevention methods exist, and Microsoft’s security response team describes indirect injection as an inherent risk of probabilistic language models. Defence has to be layered, with the firmest layers outside the model.

Layer 1: limit what an injected instruction can reach

Least privilege does more against prompt injection than any filter, because it shrinks the payoff of a successful attack. Before tuning prompts, decide what the agent must be unable to do:

  1. Split reading from acting. An agent that summarises inbound mail does not need a send tool. Give write, delete, payment and permission-change tools only to agents whose task requires them.
  2. Prefer narrow tools. A function such as create_refund(order_id, amount) with limits enforced in code is far safer than a general database or shell tool.
  3. Scope credentials to the run. Use a dedicated agent identity with short-lived, task-specific tokens instead of an employee’s session or a shared admin key.
  4. Restrict data per task. Retrieval should return only sources that both the requesting user and the agent are allowed to see.
  5. Break the trifecta. If a workflow has to read untrusted content and touch private data, remove its ability to send data out, or route every outbound step through review.

Research is turning these rules into architectures. A June 2025 paper on design patterns by Beurer-Kellner and colleagues describes six of them, including action-selector, plan-then-execute, dual LLM and context minimisation, and states the principle plainly: once an agent has ingested untrusted input, that input must be unable to trigger consequential actions. Willison’s Dual LLM pattern keeps a privileged model that holds the tools away from raw untrusted text, which only a quarantined model without tools ever reads. CaMeL, from researchers at Google, Google DeepMind and ETH Zurich, derives control and data flow from the trusted request so that retrieved data cannot change the program; on the AgentDojo benchmark it solved 77% of tasks with provable security, against 84% for an undefended agent. Agents that execute code also need isolation, which the agent sandboxing guide covers.

Layer 2: gate consequential actions

Some actions should never run on the model’s word alone. Draw the line by consequence:

  • Tool allowlists per agent and role. The runtime, not the prompt, decides which tools exist for a given agent, so a convincing injected instruction cannot summon a tool that was never granted.
  • Argument validation in code. Check recipients, amounts, file paths and destination domains against rules before a call executes, and fail closed when a value falls outside them.
  • Human approval for irreversible or external actions. Sending messages outside the company, moving money, deleting data and changing permissions should wait for a named person. Microsoft applies this in Copilot in Outlook, where the user has to approve the drafted text and send the email personally.
  • Approvals people can actually judge. OWASP’s prevention cheat sheet warns that constant approve-or-deny prompts cause user fatigue. Keep approval for the actions that matter, and show the reviewer the exact arguments and the content that prompted the request.

Tool definitions are an injection channel too. In April 2025 Invariant Labs demonstrated tool poisoning: hidden instructions in an MCP tool’s description, invisible to the user but read by the model, led an agent to collect SSH keys and MCP credentials and send them out. Pin tool versions, review descriptions before a tool is registered, and keep the list of tools each agent can reach short. The guide to human-in-the-loop AI explains where approval gates fit in a workflow.

Layer 3: mark, screen and trace untrusted content

The next layer helps the model, and your monitoring, recognise untrusted content. None of these techniques is a security boundary on its own; each one lowers the success rate or raises the alarm sooner.

Spotlighting and delimiting. Spotlighting, introduced in a 2024 paper by Hines and colleagues and now part of Microsoft’s own defences, is a family of prompt transformations that keep signalling where text came from. Microsoft’s write-up lists three modes: delimiting wraps untrusted input in randomised markers, datamarking interleaves a special character through it, and encoding transforms it, for example into base64. In the paper’s experiments, spotlighting cut attack success from above 50% to below 2% with little effect on task quality. Keep system rules, the user’s request and retrieved content in separate, labelled parts of the prompt, and state that retrieved content is reference material to be weighed, never a source of instructions.

Input and output filtering. Classifiers can flag instruction-like text, hidden Unicode, encoded payloads and known jailbreak templates on the way in, and leaked secrets or unexpected URLs on the way out. Microsoft’s layered approach includes Prompt Shields, a probabilistic classifier for several kinds of prompt attack, and open-weight guard models do similar screening on your own hardware. Use a verdict to reduce what the agent may do next, such as dropping a document or switching to read-only, instead of treating it as a final yes or no.

Content provenance. Record the source, owner and trust level of every document chunk, email and tool result that enters the context, and carry that label into the trace. Provenance lets you write rules such as “no outbound message in a run that ingested external email” and lets an investigator see which input steered a decision. For retrieval-heavy assistants, the private RAG containment guide goes deeper on ingestion, chunk labels and permission-aware retrieval.

Layer 4: close exfiltration paths and watch the agent

Injected instructions usually aim to move data out, and every channel an agent or its client can use to reach the outside is a candidate:

  • Rendered output. Markdown images and links can carry data inside a URL that loads the moment a reply is displayed. Microsoft deterministically blocks data exfiltration through Markdown image injection and the generation of certain untrusted links. Do the same: render images only from allowlisted hosts, or not at all.
  • Network egress. Agents that browse, fetch URLs or call webhooks should go through an egress proxy with an allowlist of destinations. Air-gapped and local-model deployments take the external model provider out of the path entirely.
  • Code execution. Run generated code in a sandbox with no network by default and nothing mounted that the task does not need.

Then watch behaviour as well as text. Alert when an agent calls a tool it rarely uses, contacts a new domain, requests records unrelated to the user’s question, or retries an action that policy refused. Keep a trace of every model call, tool call, argument and approval so a run can be reconstructed afterwards. Finally, test: OWASP’s seventh mitigation is regular adversarial testing and attack simulation, and the test set should run again whenever the model, prompts, tools or connectors change.

Prompt injection prevention cheat sheet

LayerControlWhat it stopsWhat it does not stop
DesignSeparate reading agents from acting agentsInjected text in inbound content triggering writesMisuse of the data the agent can read
DesignNarrow tools with limits enforced in codeOpen-ended commands such as raw SQL or shellAbuse within the tool’s legitimate range
DesignBreak the lethal trifecta per workflowPrivate data leaving through an injected requestErrors that stay inside the system
IdentityDedicated agent identity, short-lived scoped tokensBorrowed admin or employee sessionsOverly broad grants given to the agent
RuntimeTool allowlist per agent and roleCalls to tools that were never grantedBad calls to tools that were granted
RuntimeArgument validation that fails closedUnexpected recipients, amounts, paths or domainsValid-looking but wrong values
RuntimeHuman approval for consequential actionsSilent sends, payments, deletions, permission changesReviewers approving on autopilot
InputSpotlighting and labelled prompt sectionsSome confusion between data and instructionsDetermined adaptive attacks
InputGuard models and injection classifiersKnown patterns and obvious payloadsNovel or obfuscated attacks
InputProvenance labels on chunks and tool resultsUntraceable decisions; enables source-based rulesInjection from a trusted source that was compromised
OutputNo auto-rendered external images or linksZero-click exfiltration through renderingExfiltration through tools the agent holds
NetworkEgress allowlist, or no external egressCalls to attacker-controlled hostsAbuse of allowlisted destinations
OperationsFull traces, behaviour alerts, red-team suiteSlow detection and repeat incidentsThe first use of a new technique

Work through the table from the top. The design and runtime rows do not depend on the model’s judgement, so they are the ones to finish first; the input rows improve the odds, and the operations row tells you when they fail.

How VDF AI fits

VDF AI Networks enforces several of these controls in the runtime rather than in prompts. Tool allowlists define exactly which tools each agent can invoke and are enforced at the execution layer. Human Approval nodes ask an accountable owner to accept a step before the workflow goes on. Per-agent guardrails redact personal data from inputs and outputs before they reach a model or storage, network-level content safety filters reject content above configurable severity thresholds, and every run produces an execution trace with inputs, outputs and routing decisions.

For tools, the MCP gateway keeps an admin-governed registry where tools are granted per role, actions that write can require approval, and each invocation is logged against the calling agent. On-premises and air-gapped deployments can disable network egress and restrict routing to local models. None of these controls tries to detect injected text; they limit what an injected instruction can make an agent do.

Sources

Frequently asked questions

How do you prevent prompt injection?

You cannot reliably stop a model from obeying injected text, so prevention works by limiting what an injection can achieve. Give each agent only the tools and data its task needs, validate tool arguments in code, require human approval for actions that send, delete, pay or change permissions, mark untrusted content so the model can tell it apart, block the routes data could leave through, and log every tool call. Test the whole set against injection attempts before each release.

Can prompt injection be completely prevented?

Not with current models. OWASP states that it is unclear whether fool-proof prevention methods exist, and Microsoft describes indirect prompt injection as an inherent risk of probabilistic language models. What can be made dependable are the controls around the model: permission checks, approval gates and egress rules enforced in code do not rely on the model's judgement. Research designs such as CaMeL go further by keeping untrusted data out of an agent's control flow, at some cost in how many tasks it completes.

What is the difference between direct and indirect prompt injection?

In direct prompt injection, the person using the system types instructions meant to override its rules. In indirect prompt injection, the instructions arrive inside content the system processes on someone's behalf, such as an email, a web page, a shared document, a tool result or a tool description. NIST describes the indirect form as injection carried out through control of a resource rather than through the user's own input. For agents, the indirect form is the larger risk.

Do guardrails stop prompt injection?

They reduce it. Input and output classifiers, including open-weight guard models, catch many known attack patterns but miss new and obfuscated ones; the EchoLeak exploit got past Microsoft's own injection classifier in Microsoft 365 Copilot. Use a classifier verdict as a signal that lowers an agent's permissions or sends an action for review, and keep enforcement in code that a manipulated model cannot argue its way past. Treat detection as one layer among several.

Is jailbreaking the same as prompt injection?

OWASP treats jailbreaking as one form of prompt injection, in which the attacker's input makes the model disregard its safety protocols entirely. Prompt injection is the wider category: any input that changes the model's behaviour in ways its developers did not intend, including quietly redirecting an agent's tool calls without producing anything that breaks a content policy. For agents, that quiet redirection is usually the more damaging case, because nothing in the visible output looks wrong.

Filed under
AI securityAI agentsAI governancehuman oversighttool callingenterprise AI agents
AI Governance

Is your AI governance audit-ready?

Get a readiness review of your AI controls — policy, oversight, audit trails, and EU AI Act evidence — mapped against what production actually requires.

Keep reading