Securing The Model Context Protocol: Tool Poisoning, Rug Pulls And The Spec
MCP connects agents to tools and data. This guide covers where it has been attacked, what the 2026-07-28 specification requires, and what is still left to the people deploying it.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
The Model Context Protocol, usually shortened to MCP, is an open standard for connecting AI applications to tools and data. An application such as a coding assistant acts as the host. It runs MCP clients, and each client talks to an MCP server that offers tools, resources or prompts. A server can run on the user’s own machine or as a remote service.
MCP spread quickly because it makes adding a tool almost effortless. That ease is also the security problem. Every server a user installs adds text the model will read and actions the agent can take. This article covers the two kinds of MCP risk, the attacks seen so far, and what the current specification requires. It refers to protocol version 2026-07-28, which is the current version as of October 2026.1
Two Different Kinds Of Problem
It helps to separate two classes of risk, because they need different fixes.
The first is about trust. Tool names, descriptions and results are text that goes straight into the model’s context. A model can be steered by that text just as it can by a malicious web page. The specification cannot fully solve this, because it is a property of how language models work. It is mostly up to the host application and the people who choose which servers to install.
The second is ordinary software security. MCP servers, clients, proxies and developer tools are programs, and they have had familiar bugs: command injection, missing authentication on local ports and weak token handling. The specification and its security guidance address many of these directly.
| Point | Attack | What the specification does |
|---|---|---|
| 1 | Tool poisoning: hidden instructions in a tool description | Little at protocol level; depends on how the host shows and limits tools |
| 2 | Rug pull: a server changes a tool after the user approved it | Little at protocol level; hosts and users can pin versions |
| 3 | Token passthrough: a server forwards a token not issued to it | Forbidden; servers must accept only tokens issued for them |
| 4 | Confused deputy: a proxy server lets an attacker skip user consent | Proxies must obtain consent per client, match redirect addresses exactly and use single-use state values |
| 5 | Server-side request forgery during authorisation discovery | Clients should require HTTPS, block private and link-local addresses and consider an egress proxy, a gateway that controls outbound requests |
| 6 | Malicious start-up command for a local server | One-click installs must show the exact command and get explicit consent; sandboxing is recommended |
Tool Poisoning And Rug Pulls
In April 2025, Invariant Labs published an early description of a tool poisoning attack. Instructions were hidden in an MCP tool’s description, where the user would not normally see them but the model would read them.2 In one demonstration, a harmless-looking tool installed alongside a trusted email server carried hidden instructions that redirected the agent’s outgoing email, even when the user named a different recipient. Invariant called this “shadowing”, because one server’s description changes how the agent uses another server’s tools.
The same post described the rug pull. A server can change its tool definitions after the user has approved it, turning a safe tool into a harmful one without any new approval. Invariant recommended showing full tool descriptions to users, pinning server versions with checksums, and enforcing stricter boundaries between servers.2
OWASP’s incident tracker lists a September 2025 case that Koi Security reported as the first malicious MCP server found in the wild.3 A package called postmark-mcp on the npm registry impersonated the email company Postmark. According to Postmark, the publisher released 15 clean versions before version 1.0.16 added code that secretly sent a blind copy of outgoing emails to an outside server.4 Postmark had no involvement with the package. It shows how a supply chain attack and a rug pull combine: trust built over time, then a quiet change.
Implementation Bugs
The second class of problem is more familiar to security teams. Two examples from 2025 show the pattern.
- MCP Inspector. The developer tool used to test MCP servers had no authentication between its browser client and its local proxy in versions before 0.14.1. An attacker could cause it to launch commands on the developer’s machine. It was published as CVE-2025-49596 in June 2025 and rated critical.5
- mcp-remote. A widely used client-side proxy was open to operating system command injection when connecting to an untrusted MCP server. It was rated 9.6 out of 10 and fixed in version 0.1.16 in July 2025.6
More MCP-related vulnerabilities have been disclosed since. Treat any MCP component, including developer tools, as software that needs patching and review.
What The Current Specification Requires
The 2026-07-28 revision changed several security-relevant details. The protocol is now stateless, with no protocol-level sessions.7 Each request declares its protocol version, and every server must offer a discovery method that returns its versions, capabilities and identity, although clients are free to skip it.1 Revisions dated 2025-11-25 and earlier used a handshake and session IDs, so security advice written in 2025 may describe mechanisms that no longer apply.
Authorisation is optional in MCP. Where it is used over HTTP, MCP builds on OAuth 2.1, which is still an IETF draft. The MCP server plays the part of the resource server, meaning the service that checks access tokens, and the MCP client plays the OAuth client.8 The key rules are:
- clients must name the server they want a token for, using RFC 8707 resource indicators, and servers must check that each token was issued for them;
- servers must not accept or pass on any other tokens, which rules out token passthrough;
- servers publish Protected Resource Metadata (RFC 9728) so clients can find the right authorisation server;
- clients must apply the issuer checks from RFC 9207 to each authorisation response before using it, which guards against mix-up attacks where a response from one authorisation server is passed off as coming from another;
- Client ID Metadata Documents, also an IETF draft, are the recommended way for clients to register, and Dynamic Client Registration is now deprecated;
- servers should ask for precise scopes, meaning the specific permissions a token carries, and support step-up authorisation that adds permissions only when needed, rather than granting everything at once.
Local servers that use the stdio transport should not use this flow and take credentials from their environment instead.8 The separate Security Best Practices page covers confused deputy proxies, token passthrough, server-side request forgery, state handle hijacking, local server compromise, authorisation URL validation and scope minimisation.7
Practical Steps
For organisations using MCP, as of October 2026, a sensible baseline is to keep an approved list of servers and versions, review tool descriptions before approval, and alert on any change. Run local servers in a sandbox with explicit file and network grants. Give each remote server the narrowest scopes it needs. Avoid combining a server that reads untrusted content with one that holds private data and can send it out, in line with the Rule of Two covered earlier in this group. Keep developer tools such as inspectors and proxies patched and never expose them on open network interfaces.
Footnotes
-
Model Context Protocol, “Versioning”, specification version 2026-07-28. modelcontextprotocol.io ↩ ↩2
-
Invariant Labs, “MCP Security Notification: Tool Poisoning Attacks”, 1 April 2025. invariantlabs.ai ↩ ↩2
-
OWASP GenAI Security Project, “OWASP Top 10 for Agentic Applications 2026” (full document, incident tracker), December 2025. genai.owasp.org ↩
-
Postmark, “Information Regarding Malicious ‘postmark-mcp’ Package”, 25 September 2025. postmarkapp.com ↩
-
NIST National Vulnerability Database, “CVE-2025-49596”, published 13 June 2025; GitHub advisory GHSA-7f8r-222p-6f5g. nvd.nist.gov ↩
-
GitHub Advisory Database, “mcp-remote exposed to OS command injection via untrusted MCP server connections” (CVE-2025-6514, GHSA-6xpm-ggf7-wc3p), 9 July 2025. github.com ↩
-
Model Context Protocol, “Security Best Practices”, specification version 2026-07-28. modelcontextprotocol.io ↩ ↩2
-
Model Context Protocol, “Authorization”, specification version 2026-07-28. modelcontextprotocol.io ↩ ↩2
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.