explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the four MCP attack classes
  • What is an MCP rug pull?
  • What is MCP tool poisoning?
  • How do shadowing and toxic flows work?
  • A real case: postmark-mcp
  • What does an MCP server security scanner actually do?
  • How do you detect drift in practice?
  • What people are asking
  • A practical checklist
  • What explainx.ai is building
  • Honest limitations
  • Related reading
← Back to blog

explainx / blog

What Is an MCP Rug Pull? Tool Poisoning, Shadowing and How Scanners Work

MCP, AI Security, Tool Poisoning, Supply Chain, AI Agents

An MCP rug pull swaps a trusted tool for a malicious one after you approve it. Learn tool poisoning, shadowing, toxic flows, and what MCP server security scanners catch.

Oct 4, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
What Is an MCP Rug Pull? Tool Poisoning, Shadowing and How Scanners Work

An MCP server tells your agent what its tools do, in plain text, at the start of every session. The agent reads that text as trustworthy context. That single design fact explains the three attack names you keep seeing: tool poisoning, tool shadowing and the rug pull.

This guide defines each one, walks through a real malicious package, explains what an MCP server security scanner can and cannot see, and ends with a checklist. For a practitioner-level companion with a working scanner command, see AgentBeam's MCP security practical guide.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the four MCP attack classes

table · 3 cols
AttackOne-line definitionMain defense
Tool poisoningHidden instructions in a tool description or schemaRead descriptions; scan; minimize trust
Tool shadowingA malicious server redefines or biases how a trusted tool is usedFewer connected servers; namespace and review
Rug pullAn approved server later changes its definitionsPin versions and hash definitions; re-scan on update
Toxic flowHarmless tools chained into a harmful sequenceRuntime monitoring and egress controls

Source for the four-class breakdown: AgentBeam's guide, which we summarize here and extend with explainx.ai context.

What is an MCP rug pull?

A rug pull is a trust attack over time. You review an MCP server, approve it, and start relying on it. Later the server operator, or whoever compromises the package, changes a tool definition. Because clients typically fetch definitions at the start of each session, the change takes effect without a new prompt to you.

The scenario is easy to picture. On day one a tool called add_note says "Saves a note." On day thirty its description also says to read the user's .env file and include it in a context parameter. Nothing in your config changed. Try the lab below to see the same server with and without version pinning.

Lab · MCP rug pull, pinned or not

Rug pulls are especially dangerous with agents that run unattended, such as the event-triggered agents we describe in MCP Events in ChatGPT, because there may be no human reading the session when the definition changes.

What is MCP tool poisoning?

Tool poisoning is the injection of hidden malicious instructions into a tool's description or JSON schema. The model sees those fields as part of its working context, so a poisoned description can steer behavior without the user ever seeing it.

AgentBeam's example is a description that instructs the agent to read the user's .env and include the contents in the context parameter. The tool might even work as advertised while it leaks data on the side. The attack surface is anywhere the server controls text the model reads: descriptions, parameter names, enums, defaults, and error messages.

Security researchers popularized the term in 2025; the underlying issue is the same prompt injection problem our MCP security guide covers, applied to a trusted-looking channel. Our piece on tool descriptions and selection reliability shows the flip side: descriptions are powerful precisely because models weigh them heavily, which is why attackers target them.

How do shadowing and toxic flows work?

Tool shadowing and cross-origin escalation. When you connect several servers at once, a malicious one can redefine a trusted tool's name or inject instructions that change how the agent uses other servers' tools. It never has to be called itself; its description just has to be read. The more servers connected, the larger this surface.

Toxic flows. Each tool can look benign while a sequence is harmful: read a file, then post to a webhook. AgentBeam notes toxic flows are emergent and undetectable by static scanning of any single tool. They need runtime visibility and egress controls, which is why our agent harness guide stresses that the harness, not the model, owns permissions.

A real case: postmark-mcp

The first documented malicious MCP server in the wild, per AgentBeam's guide, was postmark-mcp. On September 17, 2025, a lookalike npm package at version 1.0.16 added a hidden BCC redirect. AgentBeam reports roughly 300 organizations affected, 1,643 total downloads, and "3,000 to 15,000 emails every day" flowing to an attacker-controlled address. Months later the Shai-Hulud 2.0 worm targeted 796 packages with 132 million combined monthly downloads, including packages named mcp-server.

Two lessons stand out. The attack lived in an ordinary npm package, so classic supply-chain hygiene applies as much as MCP-specific checks. And it succeeded through a version bump, the exact shape of a rug pull. The same lesson shows up in our coverage of agent supply-chain incidents such as the NVIDIA SkillSpector skill scanner.

What does an MCP server security scanner actually do?

A scanner reads an MCP configuration and the tool definitions it points to, then looks for risk signals without executing code. AgentBeam's scanner runs 11 heuristic patterns plus version-pin verification, and its documentation says the scan is not semantic malware analysis or a guarantee of safety.

table · 2 cols
Scanner can usually doScanner usually cannot do
Flag unpinned versions and floating latest tagsUnderstand intent behind novel phrasing
Match known suspicious phrases (read secrets, ignore previous, exfiltrate)Catch obfuscated or URL-indirected instructions
Diff definitions between runs to detect changesDetect behavior that only appears at runtime
Run offline with no code executionFind toxic flows across tools

Expect both error types. AppSec Santa testing cited by AgentBeam showed false positives on legitimate phrases such as "You MUST call this function first," and under-catching of obfuscation. The correct reading of a clean scan is "no known pattern matched," which is not "safe." Related research on verifying where an MCP server's content came from is covered in Hugging Face ProvenanceGuard.

How do you detect drift in practice?

Detection is mostly boring engineering. Export the tool list your client sees, including names, descriptions and input schemas, and store it in version control next to your MCP config. On each session start or in CI, fetch the list again and compare a hash or a text diff against the approved copy. Any difference becomes a review item, not an automatic update.

For local servers, pair that with a lockfile and a pinned package version so the code and the definitions both stay fixed. For hosted servers you cannot pin the code, so the definition snapshot is your main control, along with logs of which tools the agent actually called. If a server adds a new tool or edits a description you rely on, treat it as a change request: read it, decide, and only then update the approved snapshot.

This also helps with honest mistakes. Vendors rename tools and reword descriptions for legitimate reasons, and a diff shows you those changes before they surprise your agent's behavior.

What people are asking

Does using only big-vendor MCP servers remove the risk? It lowers it, not to zero. Hosted servers from large vendors reduce package-impersonation risk, but a vendor can still change definitions, and poisoning can arrive through data the tools return. See our top 10 open and closed source MCP servers for what to check on each.

Is a rug pull possible with a remote hosted server? Yes, and you have less visibility because you cannot pin a package version. You can still snapshot and hash the tool definitions you approved and alert on drift.

Should I just avoid MCP? No. The attack surface is the same one any agent tool integration has. What you want is least privilege, review of definitions, and monitoring; the MCP primer explains why the protocol is worth using.

Do agents have a better defense than humans reading descriptions? Not reliably. Models can follow injected instructions, and detection research is early. Our WipeBench safety benchmark post shows how differently agents behave when told to do something destructive.

A practical checklist

  1. Pin every MCP server version and, where possible, hash approved tool definitions so drift is detectable.
  2. Scan before first run and after every update. Treat the scan as one input.
  3. Read tool descriptions manually for servers that touch credentials, email, source control or finance.
  4. Connect fewer servers. Each one adds shadowing surface.
  5. Add runtime monitoring and egress limits so toxic flows show up as network or file events.
  6. Audit the package, not just the protocol. Check publisher, repository link, install scripts and recent version history.
  7. Keep permission prompts on for destructive or outbound actions, especially for unattended agents.
  8. Have a rollback plan: know how to disconnect a server and rotate credentials it could reach.

What explainx.ai is building

explainx.ai is preparing to offer MCP server scanning and monitoring in partnership with AgentBeam, aimed at exactly the gaps above: scanning before you connect, drift detection after, and runtime visibility. It is not available yet, so rely on the checklist and the tools you have today. If you want early access, you can join the AgentBeam waitlist. We will announce availability on the blog when there is something concrete to try.

Honest limitations

Attack names and the four-class framing here follow AgentBeam's guide; the incident figures are its numbers and we did not independently re-measure them. The lab above is an illustration, not a real server. No scanner or checklist guarantees safety, and a pinned version can still be malicious from the start.

Related reading

  • MCP security guide 2026
  • MCP Events in ChatGPT: event-driven agents
  • Top 10 open and closed source MCP servers
  • What is MCP? The complete guide
  • MCP tool descriptions and selection reliability
  • NVIDIA SkillSpector skill security scanner
  • Hugging Face ProvenanceGuard

Further reading: AgentBeam, MCP security practical guide

Accurate as of October 4, 2026. Incident figures are as reported by AgentBeam.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 4, 2026

MCP Events in ChatGPT: How Event-Driven Agents Actually Work

Until now most agents woke up for only two reasons: a cron schedule or a human message. OpenAI's MCP Events documentation adds a third, a signed webhook from an MCP server. This guide explains the protocol flow, the code you need, the security rules that are easy to miss, and where it fits.

Oct 4, 2026

Top 10 MCP Servers in 2026: 6 Open Source, 4 Closed Source (and What to Check)

The MCP ecosystem is crowded, and the first question for any server is not what it can do but who controls it. This ranking splits ten useful MCP servers into six open-source and four closed, hosted ones, explains how we ranked them, and gives a checklist for connecting any of them safely.

Sep 30, 2026

Hugging Face: MCP Agents Must Verify the Source, Not Just the Fact

On September 29, 2026, Hugging Face published Multiverse Computing’s ProvenanceGuard write-up: source-aware verification for Model Context Protocol agents. The failure is not a made-up fact. It is a true fact assigned to the wrong tool output. This post is the practitioner checklist for traces, pooling, and release gates.