Prompt Injection
Prompt injection is an attack that manipulates a large language model by feeding it crafted input, causing it to ignore its intended instructions and follow the attacker's instructions instead. It is consistently ranked as the top security risk for LLM applications.
ON THIS PAGE
What is a Prompt Injection Attack?
Language models do not cleanly separate trusted instructions from untrusted data. Everything arrives as text, and the model treats it all as potential instruction. A prompt injection attack exploits that by smuggling directions into content the model will read, tricking it into revealing data, taking unauthorized actions, or overriding its safety rules. Because the attack is just words, it requires no exploit or malware, which is what makes LLM input manipulation so accessible and so hard to filter.
Direct vs. Indirect Prompt Injection: Understanding the Difference
There are two broad forms:
Direct prompt injection. The attacker types malicious instructions straight into the model, for example telling it to ignore previous directions and reveal its system prompt. This is the classic jailbreak.
Indirect prompt injection. The attacker plants instructions in content the model will later read, a web page, a document, an email, a tool response, so the injection fires when an agent processes that content during a legitimate task.
Indirect injection is the more dangerous of the two for agentic systems, because the agent encounters the poisoned content on its own, with no user in the loop to notice. An AI prompt exploit delivered this way can steer an autonomous agent through a chain of harmful actions that each look routine.
How Prompt Injection Attacks Are Executed in Practice
A typical indirect prompt injection against an agent unfolds in stages:
The attacker hides instructions in content the agent is likely to fetch, sometimes in white text or HTML comments a human would never see.
The agent retrieves that content as part of a normal task and incorporates it into its reasoning.
Following the hidden instructions, the agent takes actions such as reading local secrets, calling a tool, or sending data to an external endpoint.
Each action looks legitimate in isolation, so traditional monitoring sees nothing unusual.
Our research shows how creatively this can be done: a coding agent was walked into leaking its own SSH keys by content framed as a puzzle to solve. See how a coding agent solved its way to SSH keys.
Why Prompt Injection Is Not Going Away
Frontier model providers invest heavily in resisting prompt injection, and each generation gets harder to fool: better training, dedicated classifiers, and system-level defenses have raised the bar meaningfully. But no model eliminates the risk, for a structural reason. The vulnerability is not a bug in any one model; it is inherent to how language models work. They interpret instructions and data through the same channel, so a sufficiently clever input can always blur the line between the two. As long as models take natural language as input and act on it, some framing will slip past the defenses, which is exactly what the SSH-keys research above demonstrated against a well-defended model.
The practical conclusion is that prompt injection should be treated as a permanent condition to manage, not a bug to be patched away. Model-level defenses are necessary and worth adopting, but they cannot be the only layer. Reducing risk means assuming injection will sometimes succeed and containing what happens next.
Reducing Prompt Injection Risk Across AI Workflows
No single control eliminates prompt injection, so defense is layered:
Input and content controls to filter and sanitize what agents ingest, accepting that determined injection will still get through.
Least privilege so a manipulated agent can do less damage.
Session and intent monitoring to catch the behavioral result: an agent acting outside what the user actually asked for, whatever the injection looked like.
Runtime enforcement to block the harmful action before it completes.
The most reliable signal is behavioral, because injection succeeds precisely when the input looks harmless.
Frequently asked questions
Why are LLMs particularly vulnerable to prompt injection?
Because they do not separate instructions from data. Both arrive as text and are interpreted together, so any content a model reads can carry commands. This is a fundamental property of how current models work, not a fixable bug in one system.
Can prompt injection attacks lead to data exfiltration or system compromise?
Yes. In agentic systems, injection can drive an agent to read secrets, call tools, and send data externally, resulting in real exfiltration or compromise, all through actions the agent is authorized to take.
How do developers and security teams currently defend against prompt injection?
With layered defenses: input filtering, least-privilege access, model-level guardrails, and, increasingly, session-level behavioral monitoring that detects when an agent departs from user intent regardless of how the injection was crafted.
Is prompt injection a solved problem or still an active area of research?
Still very much active, and likely to stay that way. Model-level defenses keep improving but remain bypassable by design, which is why the field is shifting toward behavioral detection and runtime enforcement as durable complements to input filtering.
Related terms
Related reading
Agent Steering: When the Attacker Uses Your Agent Against You
An attacker doesn't need to drop malware on your machine. They hand your AI agent a goal and let its own trust carry out the rest.
Living off Coding Agents - Claude as a C2 Server
Shadow AI Agents can hide command and control inside your trusted domain allowlist. A look at Claude Code Remote Control as a Living-off-the-Land C2 primitive.
Whatever Gets the Job Done: What the Hugging Face Incident Teaches Us About Agent Guardrails
OpenAI tested GPT-5.6 Sol and a more capable pre-release model on ExploitGym in an isolated environment. The models ended up leveraging zero-day vulnerabilities to compromise Hugging Face, all in service of their goal to solve the benchmark. The agents were determined, not to exploit, but to finish the task. How do security guardrails hold up in a world where agents perform a full-blown attack just to finish a simple job?
See what your agents are actually doing.
Dash discovers every AI agent, tool, and MCP server across your estate, understands session and intent, and enforces policy at runtime.