An MCP server tells your agent what its tools do, in plain text, at the start of every session. The agent reads that text as trustworthy context. That single design fact explains the three attack names you keep seeing: tool poisoning, tool shadowing and the rug pull.
This guide defines each one, walks through a real malicious package, explains what an MCP server security scanner can and cannot see, and ends with a checklist. For a practitioner-level companion with a working scanner command, see AgentBeam's MCP security practical guide.
TL;DR: the four MCP attack classes
| Attack | One-line definition | Main defense |
|---|---|---|
| Tool poisoning | Hidden instructions in a tool description or schema | Read descriptions; scan; minimize trust |
| Tool shadowing | A malicious server redefines or biases how a trusted tool is used | Fewer connected servers; namespace and review |
| Rug pull | An approved server later changes its definitions | Pin versions and hash definitions; re-scan on update |
| Toxic flow | Harmless tools chained into a harmful sequence | Runtime monitoring and egress controls |
Source for the four-class breakdown: AgentBeam's guide, which we summarize here and extend with explainx.ai context.
What is an MCP rug pull?
A rug pull is a trust attack over time. You review an MCP server, approve it, and start relying on it. Later the server operator, or whoever compromises the package, changes a tool definition. Because clients typically fetch definitions at the start of each session, the change takes effect without a new prompt to you.
The scenario is easy to picture. On day one a tool called add_note says "Saves a note." On day thirty its description also says to read the user's .env file and include it in a context parameter. Nothing in your config changed. Try the lab below to see the same server with and without version pinning.
Rug pulls are especially dangerous with agents that run unattended, such as the event-triggered agents we describe in MCP Events in ChatGPT, because there may be no human reading the session when the definition changes.
What is MCP tool poisoning?
Tool poisoning is the injection of hidden malicious instructions into a tool's description or JSON schema. The model sees those fields as part of its working context, so a poisoned description can steer behavior without the user ever seeing it.
AgentBeam's example is a description that instructs the agent to read the user's .env and include the contents in the context parameter. The tool might even work as advertised while it leaks data on the side. The attack surface is anywhere the server controls text the model reads: descriptions, parameter names, enums, defaults, and error messages.
Security researchers popularized the term in 2025; the underlying issue is the same prompt injection problem our MCP security guide covers, applied to a trusted-looking channel. Our piece on tool descriptions and selection reliability shows the flip side: descriptions are powerful precisely because models weigh them heavily, which is why attackers target them.
How do shadowing and toxic flows work?
Tool shadowing and cross-origin escalation. When you connect several servers at once, a malicious one can redefine a trusted tool's name or inject instructions that change how the agent uses other servers' tools. It never has to be called itself; its description just has to be read. The more servers connected, the larger this surface.
Toxic flows. Each tool can look benign while a sequence is harmful: read a file, then post to a webhook. AgentBeam notes toxic flows are emergent and undetectable by static scanning of any single tool. They need runtime visibility and egress controls, which is why our agent harness guide stresses that the harness, not the model, owns permissions.
A real case: postmark-mcp
The first documented malicious MCP server in the wild, per AgentBeam's guide, was postmark-mcp. On September 17, 2025, a lookalike npm package at version 1.0.16 added a hidden BCC redirect. AgentBeam reports roughly 300 organizations affected, 1,643 total downloads, and "3,000 to 15,000 emails every day" flowing to an attacker-controlled address. Months later the Shai-Hulud 2.0 worm targeted 796 packages with 132 million combined monthly downloads, including packages named mcp-server.
Two lessons stand out. The attack lived in an ordinary npm package, so classic supply-chain hygiene applies as much as MCP-specific checks. And it succeeded through a version bump, the exact shape of a rug pull. The same lesson shows up in our coverage of agent supply-chain incidents such as the NVIDIA SkillSpector skill scanner.
What does an MCP server security scanner actually do?
A scanner reads an MCP configuration and the tool definitions it points to, then looks for risk signals without executing code. AgentBeam's scanner runs 11 heuristic patterns plus version-pin verification, and its documentation says the scan is not semantic malware analysis or a guarantee of safety.
| Scanner can usually do | Scanner usually cannot do |
|---|---|
Flag unpinned versions and floating latest tags | Understand intent behind novel phrasing |
| Match known suspicious phrases (read secrets, ignore previous, exfiltrate) | Catch obfuscated or URL-indirected instructions |
| Diff definitions between runs to detect changes | Detect behavior that only appears at runtime |
| Run offline with no code execution | Find toxic flows across tools |
Expect both error types. AppSec Santa testing cited by AgentBeam showed false positives on legitimate phrases such as "You MUST call this function first," and under-catching of obfuscation. The correct reading of a clean scan is "no known pattern matched," which is not "safe." Related research on verifying where an MCP server's content came from is covered in Hugging Face ProvenanceGuard.
How do you detect drift in practice?
Detection is mostly boring engineering. Export the tool list your client sees, including names, descriptions and input schemas, and store it in version control next to your MCP config. On each session start or in CI, fetch the list again and compare a hash or a text diff against the approved copy. Any difference becomes a review item, not an automatic update.
For local servers, pair that with a lockfile and a pinned package version so the code and the definitions both stay fixed. For hosted servers you cannot pin the code, so the definition snapshot is your main control, along with logs of which tools the agent actually called. If a server adds a new tool or edits a description you rely on, treat it as a change request: read it, decide, and only then update the approved snapshot.
This also helps with honest mistakes. Vendors rename tools and reword descriptions for legitimate reasons, and a diff shows you those changes before they surprise your agent's behavior.
What people are asking
Does using only big-vendor MCP servers remove the risk? It lowers it, not to zero. Hosted servers from large vendors reduce package-impersonation risk, but a vendor can still change definitions, and poisoning can arrive through data the tools return. See our top 10 open and closed source MCP servers for what to check on each.
Is a rug pull possible with a remote hosted server? Yes, and you have less visibility because you cannot pin a package version. You can still snapshot and hash the tool definitions you approved and alert on drift.
Should I just avoid MCP? No. The attack surface is the same one any agent tool integration has. What you want is least privilege, review of definitions, and monitoring; the MCP primer explains why the protocol is worth using.
Do agents have a better defense than humans reading descriptions? Not reliably. Models can follow injected instructions, and detection research is early. Our WipeBench safety benchmark post shows how differently agents behave when told to do something destructive.
A practical checklist
- Pin every MCP server version and, where possible, hash approved tool definitions so drift is detectable.
- Scan before first run and after every update. Treat the scan as one input.
- Read tool descriptions manually for servers that touch credentials, email, source control or finance.
- Connect fewer servers. Each one adds shadowing surface.
- Add runtime monitoring and egress limits so toxic flows show up as network or file events.
- Audit the package, not just the protocol. Check publisher, repository link, install scripts and recent version history.
- Keep permission prompts on for destructive or outbound actions, especially for unattended agents.
- Have a rollback plan: know how to disconnect a server and rotate credentials it could reach.
What explainx.ai is building
explainx.ai is preparing to offer MCP server scanning and monitoring in partnership with AgentBeam, aimed at exactly the gaps above: scanning before you connect, drift detection after, and runtime visibility. It is not available yet, so rely on the checklist and the tools you have today. If you want early access, you can join the AgentBeam waitlist. We will announce availability on the blog when there is something concrete to try.
Honest limitations
Attack names and the four-class framing here follow AgentBeam's guide; the incident figures are its numbers and we did not independently re-measure them. The lab above is an illustration, not a real server. No scanner or checklist guarantees safety, and a pinned version can still be malicious from the start.
Related reading
- MCP security guide 2026
- MCP Events in ChatGPT: event-driven agents
- Top 10 open and closed source MCP servers
- What is MCP? The complete guide
- MCP tool descriptions and selection reliability
- NVIDIA SkillSpector skill security scanner
- Hugging Face ProvenanceGuard
Further reading: AgentBeam, MCP security practical guide
Accurate as of October 4, 2026. Incident figures are as reported by AgentBeam.
