Prompt injection prevention for AI agents starts from an uncomfortable fact: a language model can be talked into following instructions hidden in the content it reads, and no filter catches every attempt. The controls that hold up limit the damage instead. Give agents only the tools they need, require approval for consequential actions, mark and screen untrusted input, close exfiltration paths and log every tool call.
Direct and indirect prompt injection
Simon Willison named the attack in September 2022 and compared it to SQL injection: an application builds a prompt by joining its own instructions with text from elsewhere, and the model cannot reliably tell the two apart. The comparison breaks down at the fix. SQL injection has parameterised queries, while language models have no equivalent way to mark a span of text as data that must never be obeyed.
OWASP’s LLM01:2025 entry separates two forms:
- Direct prompt injection. The user’s own input changes how the model behaves, for example a message telling it to ignore its system prompt and print it.
- Indirect prompt injection. The model takes in content from an outside source, such as a website or a file, and instructions inside that content change what it does. NIST AI 100-2 E2025 describes this form as injection carried out through control of a resource rather than through the user’s input.
Agents read far more outside content than chat assistants do: inbound email, calendar invites, support tickets, shared documents, web pages, API responses and even the descriptions of the tools they can call. OWASP also points out that injected text does not have to be visible or readable to a person as long as the model parses it, so white-on-white text, HTML comments, image content and file metadata all count. The glossary entry has the short definition; this guide is about stopping it.
Why agents with tools raise the stakes
When a chatbot obeys injected text, the result is a wrong answer. When an agent obeys it, the agent acts with whatever access it holds: it sends the email, edits the record, opens the URL or runs the script. Willison’s lethal trifecta names the combination that turns injection into data theft: access to private data, exposure to untrusted content and a way to communicate externally. An assistant that reads a mailbox, summarises web pages and can make web requests has all three.
EchoLeak showed what that looks like in production. Tracked as CVE-2025-32711 and scored 9.3 by Microsoft, it was an AI command injection in Microsoft 365 Copilot that let an unauthorised attacker disclose information over a network. According to the researchers’ paper, one crafted email was enough, and the recipient never had to click anything. The exploit evaded Microsoft’s cross-prompt injection classifier, got links past redaction with reference-style Markdown, and moved data out through images that loaded automatically and a Microsoft Teams proxy that the content security policy allowed. Microsoft mitigated the flaw inside the service, so customers had nothing to install.
Two lessons stand out. A well-funded classifier still missed the attack, and the harm came from capabilities around the model: what it could read and which outbound requests were made on its behalf. OWASP says it is unclear whether fool-proof prevention methods exist, and Microsoft’s security response team describes indirect injection as an inherent risk of probabilistic language models. Defence has to be layered, with the firmest layers outside the model.
Layer 1: limit what an injected instruction can reach
Least privilege does more against prompt injection than any filter, because it shrinks the payoff of a successful attack. Before tuning prompts, decide what the agent must be unable to do:
- Split reading from acting. An agent that summarises inbound mail does not need a send tool. Give write, delete, payment and permission-change tools only to agents whose task requires them.
- Prefer narrow tools. A function such as
create_refund(order_id, amount)with limits enforced in code is far safer than a general database or shell tool. - Scope credentials to the run. Use a dedicated agent identity with short-lived, task-specific tokens instead of an employee’s session or a shared admin key.
- Restrict data per task. Retrieval should return only sources that both the requesting user and the agent are allowed to see.
- Break the trifecta. If a workflow has to read untrusted content and touch private data, remove its ability to send data out, or route every outbound step through review.
Research is turning these rules into architectures. A June 2025 paper on design patterns by Beurer-Kellner and colleagues describes six of them, including action-selector, plan-then-execute, dual LLM and context minimisation, and states the principle plainly: once an agent has ingested untrusted input, that input must be unable to trigger consequential actions. Willison’s Dual LLM pattern keeps a privileged model that holds the tools away from raw untrusted text, which only a quarantined model without tools ever reads. CaMeL, from researchers at Google, Google DeepMind and ETH Zurich, derives control and data flow from the trusted request so that retrieved data cannot change the program; on the AgentDojo benchmark it solved 77% of tasks with provable security, against 84% for an undefended agent. Agents that execute code also need isolation, which the agent sandboxing guide covers.
Layer 2: gate consequential actions
Some actions should never run on the model’s word alone. Draw the line by consequence:
- Tool allowlists per agent and role. The runtime, not the prompt, decides which tools exist for a given agent, so a convincing injected instruction cannot summon a tool that was never granted.
- Argument validation in code. Check recipients, amounts, file paths and destination domains against rules before a call executes, and fail closed when a value falls outside them.
- Human approval for irreversible or external actions. Sending messages outside the company, moving money, deleting data and changing permissions should wait for a named person. Microsoft applies this in Copilot in Outlook, where the user has to approve the drafted text and send the email personally.
- Approvals people can actually judge. OWASP’s prevention cheat sheet warns that constant approve-or-deny prompts cause user fatigue. Keep approval for the actions that matter, and show the reviewer the exact arguments and the content that prompted the request.
Tool definitions are an injection channel too. In April 2025 Invariant Labs demonstrated tool poisoning: hidden instructions in an MCP tool’s description, invisible to the user but read by the model, led an agent to collect SSH keys and MCP credentials and send them out. Pin tool versions, review descriptions before a tool is registered, and keep the list of tools each agent can reach short. The guide to human-in-the-loop AI explains where approval gates fit in a workflow.
Layer 3: mark, screen and trace untrusted content
The next layer helps the model, and your monitoring, recognise untrusted content. None of these techniques is a security boundary on its own; each one lowers the success rate or raises the alarm sooner.
Spotlighting and delimiting. Spotlighting, introduced in a 2024 paper by Hines and colleagues and now part of Microsoft’s own defences, is a family of prompt transformations that keep signalling where text came from. Microsoft’s write-up lists three modes: delimiting wraps untrusted input in randomised markers, datamarking interleaves a special character through it, and encoding transforms it, for example into base64. In the paper’s experiments, spotlighting cut attack success from above 50% to below 2% with little effect on task quality. Keep system rules, the user’s request and retrieved content in separate, labelled parts of the prompt, and state that retrieved content is reference material to be weighed, never a source of instructions.
Input and output filtering. Classifiers can flag instruction-like text, hidden Unicode, encoded payloads and known jailbreak templates on the way in, and leaked secrets or unexpected URLs on the way out. Microsoft’s layered approach includes Prompt Shields, a probabilistic classifier for several kinds of prompt attack, and open-weight guard models do similar screening on your own hardware. Use a verdict to reduce what the agent may do next, such as dropping a document or switching to read-only, instead of treating it as a final yes or no.
Content provenance. Record the source, owner and trust level of every document chunk, email and tool result that enters the context, and carry that label into the trace. Provenance lets you write rules such as “no outbound message in a run that ingested external email” and lets an investigator see which input steered a decision. For retrieval-heavy assistants, the private RAG containment guide goes deeper on ingestion, chunk labels and permission-aware retrieval.
Layer 4: close exfiltration paths and watch the agent
Injected instructions usually aim to move data out, and every channel an agent or its client can use to reach the outside is a candidate:
- Rendered output. Markdown images and links can carry data inside a URL that loads the moment a reply is displayed. Microsoft deterministically blocks data exfiltration through Markdown image injection and the generation of certain untrusted links. Do the same: render images only from allowlisted hosts, or not at all.
- Network egress. Agents that browse, fetch URLs or call webhooks should go through an egress proxy with an allowlist of destinations. Air-gapped and local-model deployments take the external model provider out of the path entirely.
- Code execution. Run generated code in a sandbox with no network by default and nothing mounted that the task does not need.
Then watch behaviour as well as text. Alert when an agent calls a tool it rarely uses, contacts a new domain, requests records unrelated to the user’s question, or retries an action that policy refused. Keep a trace of every model call, tool call, argument and approval so a run can be reconstructed afterwards. Finally, test: OWASP’s seventh mitigation is regular adversarial testing and attack simulation, and the test set should run again whenever the model, prompts, tools or connectors change.
Prompt injection prevention cheat sheet
| Layer | Control | What it stops | What it does not stop |
|---|---|---|---|
| Design | Separate reading agents from acting agents | Injected text in inbound content triggering writes | Misuse of the data the agent can read |
| Design | Narrow tools with limits enforced in code | Open-ended commands such as raw SQL or shell | Abuse within the tool’s legitimate range |
| Design | Break the lethal trifecta per workflow | Private data leaving through an injected request | Errors that stay inside the system |
| Identity | Dedicated agent identity, short-lived scoped tokens | Borrowed admin or employee sessions | Overly broad grants given to the agent |
| Runtime | Tool allowlist per agent and role | Calls to tools that were never granted | Bad calls to tools that were granted |
| Runtime | Argument validation that fails closed | Unexpected recipients, amounts, paths or domains | Valid-looking but wrong values |
| Runtime | Human approval for consequential actions | Silent sends, payments, deletions, permission changes | Reviewers approving on autopilot |
| Input | Spotlighting and labelled prompt sections | Some confusion between data and instructions | Determined adaptive attacks |
| Input | Guard models and injection classifiers | Known patterns and obvious payloads | Novel or obfuscated attacks |
| Input | Provenance labels on chunks and tool results | Untraceable decisions; enables source-based rules | Injection from a trusted source that was compromised |
| Output | No auto-rendered external images or links | Zero-click exfiltration through rendering | Exfiltration through tools the agent holds |
| Network | Egress allowlist, or no external egress | Calls to attacker-controlled hosts | Abuse of allowlisted destinations |
| Operations | Full traces, behaviour alerts, red-team suite | Slow detection and repeat incidents | The first use of a new technique |
Work through the table from the top. The design and runtime rows do not depend on the model’s judgement, so they are the ones to finish first; the input rows improve the odds, and the operations row tells you when they fail.
How VDF AI fits
VDF AI Networks enforces several of these controls in the runtime rather than in prompts. Tool allowlists define exactly which tools each agent can invoke and are enforced at the execution layer. Human Approval nodes ask an accountable owner to accept a step before the workflow goes on. Per-agent guardrails redact personal data from inputs and outputs before they reach a model or storage, network-level content safety filters reject content above configurable severity thresholds, and every run produces an execution trace with inputs, outputs and routing decisions.
For tools, the MCP gateway keeps an admin-governed registry where tools are granted per role, actions that write can require approval, and each invocation is logged against the calling agent. On-premises and air-gapped deployments can disable network egress and restrict routing to local models. None of these controls tries to detect injected text; they limit what an injected instruction can make an agent do.
Sources
- OWASP, LLM01:2025 Prompt Injection
- OWASP, prompt injection prevention cheat sheet
- NIST AI 100-2 E2025, adversarial ML taxonomy
- Microsoft MSRC on indirect prompt injection defences
- Hines et al., spotlighting
- Debenedetti et al., CaMeL
- Beurer-Kellner et al., agent design patterns
- Simon Willison, naming prompt injection
- Simon Willison, the Dual LLM pattern
- Simon Willison, the lethal trifecta
- NVD, CVE-2025-32711
- Reddy and Gujral, EchoLeak
- Invariant Labs, MCP tool poisoning