Google and four other organizations have acknowledged that their AI agents can be pushed into passing a malicious instruction to other agents on the same network. An independent researcher built working proof-of-concept attacks that do this. The weak point is the standard the agents use to talk to each other, the Model Context Protocol, known as MCP.
A message from a colleague is treated as an order
MCP is a common set of rules that lets an AI app or agent use tools and talk to other agents. It describes how to ask for something. It does not decide whether the asker should be believed.
So the attack takes three steps. A hidden instruction arrives inside ordinary content, such as a document or an email that one agent reads. That agent passes a task on to a second agent, which inherits the first agent's trust. The second agent carries the task out with its own access, and many agents have weak guardrails, or none, to stop it.
It works like an office where staff do whatever an email from inside the company asks. A stranger who asks for the customer list is refused. The same request, apparently from the colleague two desks away, gets an answer. The attacker never has to talk to the second agent. They only need to get words in front of the first one.
For example, a distributor might run a translation agent that reads supplier invoices and hands the English text to an agent that can query the customer database. A forged invoice contains one line telling the next agent to export the customer list and send it to an outside address. The translation agent does not act on the line. It translates it and passes it on. The database agent sees a request from a trusted neighbour, and no person approved anything.
The same lesson Target learned in 2013
The Target breach taught networks that being inside should not earn trust. Companies since then have moved toward checking every request, even one from a machine on the inside. Agent networks are starting from the old assumption again, because agents are built to cooperate and setting them up is fast.
There is also a larger point for anyone building small tools with AI. The risk grows with the number of hand-offs, not the number of agents. Five agents in a chain are four places where outside text can turn into an order.
Three questions before you connect agents
These take about an afternoon per agent, and they cover most of the exposure.
First, what is the least this agent needs to touch? A translation agent has no use for the customer database, and an agent without access cannot be made to leak it. The MCP security guidance makes the same point about limiting what a server can reach.
Second, does the agent treat text from another agent as information to read or as an order to follow? Anything that passed through an email, a file or a web page should count as information only.
Third, does anything leaving the company, such as an email, an export or a payment, need a person's yes? One approval step on outbound actions stops the invoice example above, even if everything before it fails.
Counting the hand-offs between your agents, and deciding what may pass at each one, tells you how many places an attacker has to work with.