Prompt Injection
Prompt injection is an attack where untrusted text causes a language model to ignore its intended instructions and follow the attacker's instead. It's the #1 risk in the OWASP Top 10 for LLM applications, and it becomes serious once models can take actions: an AI agent that reads email, browses the web, or processes documents can be steered by instructions hidden in that content, then misuse its tools to leak data, send messages, or change records.
Unlike SQL injection, there's no reliable escaping mechanism: models process instructions and data as the same kind of tokens. No filter or prompt fully prevents injection today. Defense is therefore architectural. Limit what a compromised model can do, separate privileges, require confirmation for consequential actions, and treat model output as untrusted.
TL;DR
- Direct injection: the user types instructions to subvert the system ("ignore previous instructions…"). Jailbreaks target safety policies.
- Indirect injection: malicious instructions hidden in content the model processes (web pages, emails, PDFs, tool results, code comments).
- The dangerous combination, the lethal trifecta: access to private data, exposure to untrusted content, and the ability to communicate externally.
- Filters and prompt wording reduce risk but can't eliminate it.
- Defenses: least-privilege tools, acting with the user's permissions, human approval for side effects, isolating untrusted content, constraining outputs (no arbitrary URLs or images), and monitoring.
- Test continuously with red-teaming and injection test suites.
Quick Example
An indirect injection hidden in a web page an agent is asked to summarize:
If the agent can read the inbox (private data), reads this page (untrusted content), and renders markdown images (external communication), the page can exfiltrate data. That's all three legs of the trifecta.
A safer design for the same agent:
Core Concepts
Direct vs Indirect Injection
Indirect injection is the more dangerous class for agents, because the victim isn't the attacker: a user innocently asks for a summary, and the content attacks them.
Why It's Hard to Fix
- Models are trained to follow instructions wherever they appear in context. Delimiters and "don't follow instructions in the document" help, but aren't guarantees.
- Attacks can be encoded, split across content, written in other languages, hidden in images, or phrased as plausible task steps.
- Classifier-based detectors catch known patterns but miss novel ones, and adaptive attackers iterate quickly.
Model providers keep improving robustness through training, which meaningfully raises the bar, but systems must be designed assuming some injections will succeed.
Impact Channels
- Tool misuse: sending emails, making purchases, modifying files, calling APIs with the agent's permissions.
- Data exfiltration: embedding secrets in URLs (links, markdown images, web requests), tool arguments, or outgoing messages.
- Prompt leakage: revealing system prompts and confidential instructions (don't put secrets in prompts).
- Misinformation: altering summaries, recommendations, or search results shown to users.
- Persistence: writing malicious memories or files that affect future sessions.
The Lethal Trifecta
An agent is at high risk when it combines:
- Access to private data (inboxes, documents, databases, credentials).
- Exposure to untrusted content (web, email, uploads, third-party tools).
- An exfiltration channel (sending messages, making web requests, rendering external images or links).
Remove any one leg for a given task, and the worst-case outcome shrinks dramatically.
Defense in Depth
Architectural Controls (Most Important)
- Least privilege: give each agent only the tools and data the task needs. A summarizer doesn't need email-sending tools.
- User-scoped authorization: tools act with the end user's permissions, enforced server-side, never with broad service credentials.
- Human-in-the-loop for consequential actions: show exactly what will be sent, bought, deleted, or changed, and require confirmation.
- Privilege separation / dual-LLM patterns: a quarantined model processes untrusted content without tools, and a privileged model acts only on structured, validated outputs. Research designs like CaMeL formalize this with capability tracking.
- Egress restrictions: allowlist domains for fetches and links, block rendering of arbitrary external images, and sandbox code execution without network access.
Input and Context Handling
- Clearly delimit untrusted content (tags such as
<untrusted_document>), and instruct the model to treat it as data. - Strip hidden text, metadata, and invisible characters where feasible.
- Keep system instructions and secrets out of reach. Assume system prompts can leak.
Output Handling
- Treat model output as untrusted: sanitize markdown and HTML, validate tool arguments against schemas and business rules, and never pass output directly to shells, SQL, or
eval. See XSS and SQL injection. - Use structured outputs with enums and constraints for actions.
Detection and Monitoring
- Injection classifiers and guardrail layers as one signal among several.
- Log tool calls and flag anomalies: unexpected recipients, unusual data volumes, new domains.
- Rate-limit sensitive actions, and alert on policy violations.
Best Practices
Threat-Model Every Agent Capability
For each tool, ask what an attacker could make the agent do with it after injecting content, and whose data it could reach. Design mitigations per tool, not just per app.
Assume the Model Can Be Convinced
Security guarantees must come from code: permissions, confirmations, allowlists, sandboxes. Prompt instructions like "never reveal X" are helpful but not a security boundary.
Red-Team Continuously
Maintain a suite of injection attacks (direct, indirect, encoded, multi-step) relevant to your tools, run it in CI against prompt and model changes, and add real incidents to it. See LLM evaluation.
Educate Users About Confirmations
Confirmations only help if users read them. Show clear, specific previews of actions, and avoid confirmation fatigue from trivial prompts.
Common Mistakes
Relying on a System Prompt Rule
It helps against casual attempts, but determined attackers routinely bypass it. It must be backed by architectural limits.
Granting Broad Credentials to Agents
An agent with an admin API key turns any successful injection into full compromise. Scope credentials per user and per task.
Rendering Untrusted Markdown
Rendering model output that includes !img silently sends data to the attacker when the image loads. Restrict image and link domains, or render plain text.
FAQ
Can prompt injection be completely prevented?
Not with current technology. Models can't reliably distinguish trusted instructions from instructions embedded in data. Robustness training, detection, and careful prompting reduce success rates, but secure systems limit the impact of successful injections through permissions, isolation, and human approval.
What's the difference between prompt injection and jailbreaking?
Jailbreaking aims to bypass a model's safety policies, to get disallowed content. Prompt injection aims to subvert an application's instructions, often through third-party content, to make it misuse data or tools. They overlap in technique, but differ in target and impact.
Is indirect prompt injection a risk if my app only summarizes documents?
The risk is lower if the summarizer has no tools, no private data beyond the document, and no way to send data out, though injected content can still manipulate the summary a user relies on. Risk rises sharply once the same agent can also read private data or take actions.
How do I test my application for prompt injection?
Build test cases embedding malicious instructions in every untrusted input channel (user messages, documents, web pages, tool results, file names, image text), and check whether the agent performs unauthorized actions or leaks data. Use automated red-teaming tools, public injection datasets, and periodic manual testing by security staff.
Related Topics
- AI Agents — Where injection becomes high-impact
- AI Guardrails — Input and output safety layers
- Tool Calling — Securing tools against misuse
- Agent Memory — Memory poisoning risks
- OWASP Top 10 — Web security risks with parallels
- API Security — Authorization and least privilege