AI RACE— The AI Race
Research

Trust Flaws in Model Context Protocol Expose AI Agents to Cross-Protocol Attacks

Security researchers have uncovered structural trust gaps in the Model Context Protocol (MCP) that allow malicious prompts to pivot across internal AI agents, affecting systems at Google, Rapid7, and other major institutions.

10/06/2026, 05:26
Lỗ hổng tín nhiệm trong Model Context Protocol khiến AI agent đối mặt nguy cơ tấn công chéo giao thức

The rapid rollout of AI agents across enterprise environments has introduced a significant security blind spot. Over the past five months, organizations including Google, JP Morgan Chase, Weviate, Rapid7, the French government’s interministerial digital directorate, and the US federal government have seen their agent systems tested against a new class of multi-agent exploits.

The attacks exploit architectural trust assumptions within the Model Context Protocol (MCP)—an increasingly common standard enabling AI applications and specialized agents to interact across internal networks.

Exploiting the Hallways Between Protocols

The vulnerabilities were demonstrated by independent security researcher Syed Anas Mohiuddin, who showed that trust gaps in MCP allow attackers to spread harmful instructions from one internal agent to another.

Unlike standard prompt injections targeting large language models (LLMs) directly, this technique targets downstream, task-specific agents—such as those dedicated to translation or data analysis. Because these specialized agents often lack robust guardrails and implicitly trust upstream components, malicious prompts can instruct an agent to pass tasks along a chain until they trigger unauthorized actions, including server-side request forgery (SSRF) and data exfiltration.

Mohiuddin dubbed this technique "protocol pivoting." In these scenarios, an attacker gains initial access through one protocol (such as MCP), exploits implicit trust assumptions across systems, and escalates privileges via a different framework, such as Google’s Agent-to-Agent (A2A) protocol or the emerging Agent Network Protocol.

Markus Vervier, a researcher at X41 D-Sec who has also developed MCP-focused exploits, noted that the technique operates fundamentally as indirect prompt injection. Regardless of the label, researchers agree the attack vector is unexpected and difficult to defend against because each individual component performs its assigned task as designed.

Critical Flaws Found in Google and Rapid7

The severity of the flaws varies across implementations:

  • Google (Severity 8/10): The vulnerability affected an MCP toolbox for databases (googleapis/mcp-toolbox). The system initialized its HTTP client without a CheckRedirect policy to manage redirected URLs and failed to validate destination IP addresses. An attacker using a crafted path parameter could cause the toolbox to follow a redirect to an internal endpoint and execute requests on the attacker’s behalf. Google resolved the issue by enforcing an IP allow-list and block-list that rejects unsafe base URLs at startup.
  • Rapid7 (CVE-2026-97228): Carrying a severity rating of 2.7 out of 10, this flaw inside Rapid7’s network was patched last month.

A Call for Zero Trust in Agentic Systems

Experts point out that the underlying security flaws are familiar vulnerabilities, such as SSRF and injection, which have resurfaced because organizations have abandoned "zero trust" principles in the rush to adopt agentic architectures. In many current setups, MCP servers retain credentials while treating inputs from adjacent agents as verified.

Douglas McKee, director of vulnerability intelligence at Rapid7, highlighted that each protocol was designed in isolation, leaving the inter-agent pathways unmonitored. McKee emphasized that any data passed from an LLM to a tool or secondary agent must be treated with the same scrutiny as unauthenticated input from a stranger on the public internet.

Related stories