AI SECURITY / BEGINNER

Prompt Injection: Direct, Indirect And Why It Cannot Simply Be Patched

Prompt injection tops every list of LLM risks. Learn the difference between direct and indirect injection, why national agencies say it may never be fully fixed, and what limits the damage.

Checked against primary sources and independently reviewed on . Sources are listed at the end.

Prompt injection is the first entry in both the 2025 and 2026 editions of the OWASP Top 10 for LLM Applications.1 It happens when text that reaches a large language model changes what the model does in a way the application’s developer did not intend. Sometimes a user types the hostile instruction. More worrying is when the instruction is hidden in an email, a web page or a shared document that the model reads on someone else’s behalf.

This article explains the two forms, why the problem is structural rather than a bug waiting for a patch, and the kinds of controls that reduce the harm. It stays at the level a defender needs. It does not include working attack text.

Direct And Indirect Injection

In direct injection, the problem text arrives through the application’s own input, usually the chat box or the request a program sends to the model’s interface (its API). Most often the user is the attacker and tries to talk the model out of its instructions, for example to reveal its hidden configuration or produce content the application is meant to refuse. Intent is not required, though. OWASP counts cases where an honest user unknowingly submits material that redirects the model.1 Jailbreaking, covered later in this group, is a close cousin.

In indirect injection, the attacker never talks to the model at all. They plant instructions in content the model will later process: a web page an assistant summarises, an email a copilot reads, a file in a shared drive, or the result returned by a tool. The victim is an ordinary user who asked an innocent question. NIST’s adversarial machine learning taxonomy treats these as two separate attack classes against generative AI.2

Kai Greshake and colleagues described indirect injection in a paper first posted in February 2023. They showed that planted text could steer real LLM-integrated applications, including Bing’s GPT-4 powered chat and code-completion tools, towards data theft, manipulated answers and unauthorised API calls.3

Direct InjectionIndirect Injection
Who supplies the hostile textUsually the user of the application, sometimes by accidentA third party who controls content the model reads
Where it arrivesThe chat box or API requestWeb pages, emails, documents, images, tool results, stored memory
Who is harmedUsually the application ownerOften an innocent user whose data or accounts the model can reach
Typical goalBypass rules or reveal hidden instructionsSteal data or trigger actions with the victim's access
The two forms of prompt injection at a glance.

How An Indirect Attack Unfolds

The steps below show the general pattern. Each step is ordinary behaviour for an assistant; the harm comes from the combination.

  1. An Attacker Plants Instructions In Content

    The text may sit in an email, a web page or a document, and can be hidden from human readers, for example in white text or metadata.

  2. A User Asks The Assistant For Help

    The request is innocent, such as "summarise my latest emails" or "what does this page say".

  3. The Assistant Retrieves The Planted Content

    Retrieval places the attacker's text into the model's context alongside the user's request and the system instructions.

  4. The Model Treats The Planted Text As Instructions

    Because instructions and data arrive in the same stream, the model may follow the attacker rather than the user.

  5. A Tool Call Or Output Leaks Data Or Takes An Action

    The assistant sends data to an outside address, renders a link that carries data, or performs an action using the user's permissions.

The general shape of an indirect prompt injection. No single step looks like an attack, which is why it is hard to filter.

This pattern has been shown to work against a production product. In June 2025 Microsoft published CVE-2025-32711, a flaw in Microsoft 365 Copilot that researchers nicknamed EchoLeak. NVD describes it as AI command injection that lets an unauthorised attacker disclose information over a network; Microsoft scored it 9.3 (critical), while NVD’s own score is 7.5 (high).4 A case study of the flaw published in September 2025 describes a chain that needed no click from the victim and started with a single crafted email, which got past several separate safeguards to move data out of the user’s context.5

Why It Cannot Simply Be Patched

Security teams know how to fix SQL injection, where attacker text slips into a database command. They use parameterised queries, which send the command and the user’s data to the database separately, so the database never confuses one for the other. The UK National Cyber Security Centre argued in December 2025 that prompt injection has no such fix, because inside an LLM there is no separation between data and instructions. The model sees one stream of tokens, the small chunks of text it reads and writes, and predicts what comes next.6 Its blog post concludes that prompt injection may never be mitigated as completely as SQL injection, and that the realistic aim is to lower the likelihood and impact of attacks.

NIST reaches a similar view. Its 2025 taxonomy notes that current mitigations do not give full protection, and suggests designing systems on the assumption that injection is possible whenever a model reads untrusted input.2 OWASP’s 2025 entry says it is unclear whether any fool-proof prevention exists, given how models work.7

What Reduces The Risk

Because no one has shown a way to close this weakness reliably inside the model, defence shifts to limiting what a fooled model can reach and do. OWASP’s guidance and the NCSC’s advice point to the same families of control:76

  • Least privilege. Give the model and its tools only the data and permissions the task needs, ideally the permissions of the current user and no more.
  • Separate and mark untrusted content. Keep retrieved content clearly labelled as data so the application, and where possible the model, can tell it apart from instructions.
  • Validate outputs. Check model output against an expected format before it reaches a browser, database or tool, and block outbound links or requests that are not needed.
  • Human approval for high-impact actions. Sending money, deleting data or emailing outside the organisation should need a person to confirm.
  • Monitoring and testing. Log prompts, retrieved content and tool calls, and run adversarial tests before and after launch.

None of these is new to security teams. What is new is applying them to a component that can be talked into following instructions from whatever it reads. The architectural patterns that contain injection are covered in Designing LLM Applications That Stay Safe When The Model Is Fooled, and agent-specific risks in Agentic AI Security.

Footnotes

  1. OWASP GenAI Security Project, “LLM01:2026 Prompt Injection”, OWASP Top 10 for LLM Applications 2026, August 2026. github.com ↩ ↩2

  2. NIST, AI 100-2 E2025, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations”, March 2025, sections 3.3 and 3.4. csrc.nist.gov ↩ ↩2

  3. K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz and M. Fritz, “Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”, arXiv 2302.12173, February 2023. arxiv.org ↩

  4. NIST National Vulnerability Database, “CVE-2025-32711”, published 11 June 2025. nvd.nist.gov ↩

  5. P. Reddy and A. S. Gujral, “EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System”, arXiv 2509.10540, September 2025. arxiv.org ↩

  6. UK National Cyber Security Centre, D. Chismon, “Prompt injection is not SQL injection (it may be worse)”, 8 December 2025. ncsc.gov.uk ↩ ↩2

  7. OWASP GenAI Security Project, “LLM01:2025 Prompt Injection”, OWASP Top 10 for LLM Applications 2025. github.com ↩ ↩2

Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.