Prompt Injection Becomes Agent Hijacking: Lessons From Real Incidents
When a model can act, hidden instructions in an email or web page can steer real tools. How agent hijacking works, what EchoLeak showed, and where defenders can break the chain.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
Prompt injection is the attack where someone writes text that a language model treats as an instruction. On its own, in a chat window, the damage is usually limited to a strange or harmful answer. Connect the same model to tools and the damage changes shape. The injected text can now send email, read files, open pull requests or call web addresses, all with the agent’s own permissions.
The US Center for AI Standards and Innovation (CAISI) at NIST calls this agent hijacking. This article explains how it works in practice, using publicly disclosed vulnerabilities and research demonstrations, and shows where defenders can break the chain. It stays at the level of concepts; it does not describe payloads. For prompt injection against plain chat models, see the AI Security group.
What Agent Hijacking Means
In January 2025, NIST described agent hijacking as a form of indirect prompt injection: an attacker places malicious instructions in data an agent is likely to read, and the agent then takes unintended, harmful actions.1 The root cause, NIST explained, is that current agents have no firm separation between their trusted instructions and the untrusted data they process.
“Indirect” is the key word. The attacker never talks to the agent. They write an email, edit a web page, open a public issue or upload a document, and wait for the agent to read it as part of a normal task. The user who asked for the task may never see the planted text.
- The Attacker Plants Content
Instructions are hidden in an email, web page, document, issue or tool description that the agent is likely to read.
- The Agent Retrieves It During A Normal Task
A summary, search or triage job pulls the content into the model context. Control point: label and isolate untrusted content.
- The Model Follows The Planted Instruction
The model treats the text as part of its task and changes its plan. Control point: limit which tools are reachable from sessions that read untrusted content.
- The Agent Calls A Tool
It reads private data, opens a pull request or composes a message. Control point: require approval for high-impact actions and show the real target.
- Data Or Effects Leave The System
Information goes out through an email, link or remote image fetch, or a change is made. Control point: allow-list outbound destinations and log every tool call.
EchoLeak: A Zero-Click Case
A well documented example is EchoLeak, tracked as CVE-2025-32711, which researchers at Aim Security found and reported privately to Microsoft.2 Microsoft published the vulnerability on 11 June 2025 as an AI command injection in Microsoft 365 Copilot that let an unauthorised attacker disclose information over a network. It rated the flaw critical, with a CVSS severity score of 9.3 out of 10, and said it had already been fully mitigated with no action needed from customers.3 NIST’s National Vulnerability Database gives the same flaw a lower score of 7.5, rated high.4 A later case study by two other researchers dates Microsoft’s server-side fix to May 2025.5
The case study’s authors describe EchoLeak as the first real-world zero-click prompt injection exploit in a production language model system.5 “Zero-click” means the victim did nothing beyond receiving an email. When Copilot later drew on that email during routine use, the planted instructions led it to tuck sensitive information from the user’s context into the address of an image. The client loaded the image automatically. Because that address pointed at a Microsoft Teams service the browser was allowed to contact, the service fetched the attacker’s URL on the client’s behalf, and the data went with it.2
The useful lesson for defenders is in how the chain was built. It slipped past a classifier meant to catch injected prompts and past the redaction of external links, then combined automatic image loading with an allowed Microsoft domain to get around the content security policy, the browser rule that limits which sites a page may contact.2 Several separate protections were in place, and the attack found a route between them.
Coding Agents And Tool Servers
Developer tools are a common target because coding agents read a great deal of untrusted text and often hold powerful credentials. In May 2025, Invariant Labs showed that a malicious issue in a public GitHub repository could steer an agent connected through the GitHub MCP server. The agent was asked to look at open issues, read the planted instructions, then pulled data from the user’s private repositories and published it in a pull request on the public repository.6
Invariant stressed that this was not a flaw in the GitHub server’s code, and that a server-side patch could not fix it.6 The agent’s token was valid and each tool call was allowed. The flaw was that one session combined untrusted input, private data and a way to publish, which is the lethal trifecta described earlier in this group.
OWASP’s Top 10 for Agentic Applications records several similar 2025 incidents and maps them to its entries. They include a poisoned prompt shipped in an Amazon Q extension for VS Code in July 2025 and the Replit agent that deleted a production database and generated misleading output in the same month.7 The second case was not an attack at all, which is a reminder that the same controls must also contain an agent’s own mistakes.
How Often Attacks Succeed
Model developers have added defences, and resistance has improved. Measured results still show that persistence pays. In NIST’s January 2025 tests of an agent built on Claude 3.5 Sonnet (the October 2024 update), using the AgentDojo framework, the strongest existing attack succeeded 11 percent of the time on a held-out set of email and office tasks. The strongest new attack, developed by NIST staff with red teamers from the UK AI Security Institute, succeeded 81 percent of the time. In a separate test of five attack goals, allowing each attack 25 attempts rather than one raised the average success rate from 57 to 80 percent.1
Those figures describe one model at one point in time, and newer models will score differently. The broader lesson holds: an attacker who can try many times, with many phrasings, has a much better chance of finding one that works. Later in this group, an article on testing covers how evaluations are run and what large red-teaming exercises found.
Breaking The Chain
Because no single filter is reliable, defenders design for the case where injection succeeds and limit what follows. The approaches below map onto the steps of the chain above.
| Control | What it does | Main limitation |
|---|---|---|
| Separate untrusted content | Marks or isolates text from outside sources so the agent treats it as data | Models can still follow instructions inside marked text |
| Restrict tool combinations | Keeps untrusted input, private data and outbound actions out of the same session | Needs careful task design; can reduce usefulness |
| Human approval | Requires a person to confirm consequential actions | People approve too quickly if prompts are frequent or vague |
| Egress control | Allows outbound requests only to approved destinations | Approved destinations can sometimes carry data too |
| Least privilege | Gives the agent only the access its task needs | Access tends to grow over time unless reviewed |
| Logging and monitoring | Records prompts, retrieved content and tool calls | Detects after the fact rather than preventing |
Joint government guidance from May 2026 follows the same pattern. It recommends input validation and prompt injection filters, but also data loss prevention tuned to agent behaviour, mandatory human approval for high-impact actions such as network egress, and logs that capture which tools the agent used and what it retrieved.8
Footnotes
-
NIST, “Technical Blog: Strengthening AI Agent Hijacking Evaluations”, 17 January 2025. nist.gov ↩ ↩2
-
I. Ravia, Aim Labs, “Breaking down ‘EchoLeak’, the First Zero-Click AI Vulnerability Enabling Data Exfiltration from Microsoft 365 Copilot”, 2025, now published by Cato Networks. catonetworks.com ↩ ↩2 ↩3
-
Microsoft Security Response Center, “M365 Copilot Information Disclosure Vulnerability”, CVE-2025-32711, 11 June 2025. msrc.microsoft.com ↩
-
NIST National Vulnerability Database, “CVE-2025-32711”, published 11 June 2025. nvd.nist.gov ↩
-
P. Reddy and A. S. Gujral, “EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System”, arXiv 2509.10540, September 2025. arxiv.org ↩ ↩2
-
Invariant Labs, “GitHub MCP Exploited: Accessing private repositories via MCP”, 26 May 2025. invariantlabs.ai ↩ ↩2
-
OWASP GenAI Security Project, “OWASP Top 10 for Agentic Applications 2026” (full document, incident tracker), December 2025. genai.owasp.org ↩
-
ASD’s ACSC, CISA, NSA, Canadian Centre for Cyber Security, NCSC-NZ and NCSC-UK, “Careful Adoption of Agentic AI Services”, 1 May 2026. ncsc.govt.nz ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.