MCP Server Security: Risks, Best Practices, and Enterprise Controls
MCP server security is the set of supply chain checks, runtime controls, and audit mechanisms that protect organizations from malicious, vulnerable, or misbehaving Model Context Protocol (MCP) servers. MCP servers expose tools and data to AI agents and assistants. That makes them one of the most exposed surfaces in any AI deployment: a single compromised MCP server can exfiltrate data, execute remote code, or manipulate agent decisions across thousands of users. In 2026, MCP-related incidents have already affected hundreds of thousands of users, and the category is just beginning to harden. This guide explains the risks, the documented attacks, and the controls security and platform teams should put in place now.
What is MCP server security?
MCP server security is the practice of verifying that an MCP server is safe before it loads, controlling what it can do at runtime, and producing tamper-evident evidence of every tool call it makes. It applies whether the MCP server runs on a developer laptop, an on-premises cluster, or an air-gapped network.
It is not the same as application security or network security. The MCP protocol gives agents direct, often local access to tools, data sources, and other systems through a contract that is much more privileged than a normal HTTP API.
Why MCP server security matters in 2026
MCP servers are everywhere. Developers point IDEs like VS Code, Cursor, and Claude Desktop at MCP servers from GitHub, npm, public registries, and internal repos. Agents in production use MCP servers to query databases, send emails, file tickets, write code, and move money. Most security teams have no inventory of what their organization is running.
The attack surface is already being exploited. Real, documented incidents in the past 12 months:
- Mini Shai-Hulud (May 2026). CVE-2026-45321, CVSS 9.6. A worm compromised 42 npm packages and 84 versions (TanStack, OpenSearch, UiPath, others) plus Mistral AI and Guardrails AI on PyPI. The attackers hijacked the build pipelines and published malicious packages through the legitimate release workflows, producing valid SLSA Level 3 attestations and signatures on compromised artifacts. The Guardrails AI compromise executed its credential-stealer payload on import. The same class of attack against an MCP server package would compromise every agent that loaded it.
- CVE-2025-6514 in
mcp-remote. 437,000+ downloads, CVSS 9.6. Full remote code execution through crafted OAuth endpoints. - Malicious Postmark MCP server. Silently BCC'd every email it processed to an attacker.
- Smithery platform breach. A path traversal vulnerability exposed credentials for 3,000+ hosted MCP servers.
- GitHub MCP server prompt injection. Exfiltrated private repo data into public PRs.
- Procurement agent manipulation. A compromised MCP-driven agent was tricked into approving $5M in false purchase orders.
These are not hypothetical risks; they are the new baseline. The Mini Shai-Hulud incident in particular sets the bar for what controls now have to handle: signatures and SLSA attestations can be valid on malicious artifacts when the attacker owns the build pipeline. The structural fix is enforcement that sits outside the build pipeline's trust boundary.
Regulators are catching up. NIST AI RMF, the EU AI Act, CMMC Level 2/3, HIPAA, and SR 11-7 all expect organizations to demonstrate provenance, access control, human oversight, and tamper-evident records for AI systems. MCP servers fall squarely inside those expectations.
The seven MCP server security risks security teams should track
1. Tool poisoning
An attacker publishes an MCP server that looks legitimate but contains malicious tool definitions, prompt instructions, or response shaping. Anyone whose agent loads the server inherits the behavior. Tool poisoning is hard to detect through code review because the malicious behavior is often expressed in tool descriptions or prompt instructions rather than in code.
2. Prompt injection through MCP responses
A compromised or hostile MCP server returns content that hijacks the agent's reasoning. The agent then leaks data, calls dangerous tools, or takes unauthorized actions. The GitHub MCP server incident in 2025 is the canonical example.
3. Credential exfiltration
MCP servers often handle OAuth tokens, API keys, or service credentials. A poorly written or malicious server logs, forwards, or leaks those credentials. The Postmark and Smithery incidents both involved credential exposure.
4. Remote code execution
Some MCP servers execute code on behalf of agents. Vulnerabilities in the server's execution path (path traversal, command injection, deserialization flaws) become full RCE primitives. CVE-2025-6514 sits in this category.
5. Supply chain tampering
An MCP server is approved by the security team in version 1.0, then a later release silently introduces malicious behavior. Without signature verification and pinned digests, no one notices until something goes wrong.
6. Data exfiltration through tool arguments
An MCP server is allowed to do something benign (search documents, summarize emails), but the tool arguments contain sensitive data the server then logs, forwards, or stores. Traditional DLP does not see this traffic because it often runs over stdio, not HTTP.
7. Shadow MCP servers
Employees install MCP servers from public sources without security review. The organization has no inventory, no policy, and no audit trail. When something breaks, the security team learns about it from the news.
MCP server security best practices
A working program puts seven controls in place. Each maps to one or more of the risks above.
1. Centralize MCP servers in a curated internal registry
Stop pulling MCP servers from public sources at runtime. Mirror approved MCP servers into an internal registry that is the only source IDEs, agents, and production workloads pull from. The registry should be the single answer to "which MCP servers are approved?"
2. Scan MCP servers before they enter the registry
Run automated checks for known vulnerabilities, suspicious tool definitions, dependency risks, license issues, and prompt injection markers. Capture results as signed attestations attached to the artifact.
3. Sign and version MCP servers as immutable artifacts
Package MCP servers as OCI artifacts with cryptographic signatures. Pin specific versions in your policy so an attacker cannot quietly push an update. Distribute the artifacts through the same registries that already serve your container images.
4. Enforce tool-level access control at runtime
For every MCP server, define which agents can call which tools, with which arguments, under which conditions. Tool policy is evaluated at every invocation, not just at session start. Destructive operations require explicit confirmation. High-risk actions require human approval.
5. Inspect MCP traffic at the semantic level
The runtime should inspect tool arguments and responses for prompt injection, PII leakage, regulated content, and policy violations. Infrastructure-level monitoring (network, filesystem, syscalls) does not see this; you need a content-aware layer.
6. Log every MCP decision in a tamper-evident chain
Every artifact load, tool invocation, denial, approval, and content event should be written to a cryptographically chained audit log. The chain ensures past entries cannot be edited without detection.
7. Make policy enforcement work disconnected
If the enforcement layer depends on a SaaS control plane, MCP server security fails the moment the connection drops. The runtime should enforce policy locally with no remote dependency, and audit logs should sync when connectivity is restored.
Enterprise MCP server security controls in practice
Here is what a working MCP server security program looks like in production.
| Lifecycle stage | Control | Purpose |
|---|---|---|
| Discovery | Inventory all MCP servers currently in use | Close the shadow-MCP gap |
| Curation | Scan, sign, and version MCP servers in an internal registry | Establish a single source of truth |
| Admission | ArtifactPolicy verifies signature, scan results, and provenance at load | Block compromised artifacts |
| Runtime | ToolPolicy enforces per-tool, per-agent, per-argument rules | Limit blast radius |
| Runtime | GuardrailPolicy inspects arguments and responses | Catch injection and data leakage |
| High-risk actions | Human-in-the-loop approval captured as signed attestation | Stop irreversible mistakes |
| Audit | Cryptographically chained logs across registry and runtime | Provide compliance evidence |
The same controls work in Kubernetes, on-premises, air-gapped, edge, IoT, and developer-laptop environments.
How to deploy MCP server security: step by step
- Inventory what is already in use. Run a sweep across managed devices, CI/CD pipelines, agent runtimes, and IDEs. Expect to find more MCP servers than anyone expected.
- Stand up an internal MCP registry. Mirror approved servers from public sources into a private OCI registry with scanning and signing. Point IDEs and agents at it.
- Write artifact policy for admission. Define which signatures, scan results, and provenance properties are required before an MCP server can load.
- Write tool policy for the top ten MCP servers. Start with the servers handling the highest-risk operations (payments, email, code commits, database writes). Expand from there.
- Deploy a secure runtime for AI. The runtime should enforce policy at every MCP tool invocation, inspect arguments and responses at the semantic level, capture human approvals, and write tamper-evident logs.
- Wire human-in-the-loop into high-risk tools. Define which tools require approval, who can approve, what the SLA is, and what happens when no one responds in time.
- Connect the audit log to compliance evidence. Compliance teams should be able to export tamper-evident evidence without manual preparation.
- Review and update on a defined cadence. New MCP servers appear every week. Static policy is stale policy within a quarter.
Common MCP server security mistakes
- Treating MCP servers like any other npm package. They expose privileged tool surfaces to agents. The risk profile is much higher.
- Trusting public registries by default. A popular server is not the same as a safe server. Smithery was popular until the breach.
- Skipping signature and digest verification. An MCP server you approved last quarter is not the MCP server running today unless you pinned the version and verified the signature.
- Using a gateway as the only control. A gateway sees HTTP traffic. Most MCP communication is over stdio. The gateway cannot govern what it cannot see.
- Letting policy depend on connectivity. If your enforcement layer fails open or fails closed when disconnected, you cannot govern MCP servers in DDIL or air-gapped environments.
- No HIL for the actions that matter. Without a signed human approval on truly high-risk MCP tool calls, the audit trail is incomplete.
How to measure MCP server security success
Track these over time:
- Percentage of MCP servers running from the internal registry vs. public sources
- Number of artifacts blocked at admission by artifact policy
- Number of tool invocations denied or escalated at runtime
- Mean time to evidence for compliance and audit requests
- Coverage of high-risk MCP tools protected by HIL
- Number of governance gaps closed since the program started
How Jozu fits
Jozu was built for this problem, about two years before the market caught up.
Jozu Hub with the MCP Registry API acts as the curated internal registry. It implements the MCP registry specification so VS Code, Cursor, Claude Desktop, and your production agents can point directly at it for centrally curated, security-scanned MCP servers. Every MCP server is packaged as a signed OCI artifact with cryptographic signatures, scan results, and signed Agent attestations. Five integrated scanners assess MCP servers for serialization attacks, prompt injection markers, license risk, and known vulnerabilities.
Jozu Agent Guard is the secure runtime for AI. It enforces ArtifactPolicy at admission (no unverified MCP server loads), ToolPolicy at every MCP tool invocation (per-tool, per-agent, per-argument), and GuardrailPolicy on prompt and response content through the integrated Bifrost gateway. Human-in-the-loop approvals are captured as signed attestations. Every decision is written to a cryptographically chained audit log. Policies travel as signed OCI artifacts and enforce locally with no SaaS dependency, including in air-gapped and DDIL environments.
The result is one policy language, one audit chain, and one platform from registry to runtime, covering MCP servers on developer laptops, in Kubernetes clusters, on edge devices, and in disconnected environments.
Two structural properties matter for the Mini Shai-Hulud failure mode. First, scan results travel with the MCP server as signed attestations, not as separate reports a compromised pipeline could detach. Second, policies are themselves OCI artifacts, authored and signed separately from the MCP servers they govern, and re-verified at the moment of execution. Agent Guard does not implicitly trust the build pipeline that produced the MCP server, and Hub does not implicitly trust the runtime to enforce the right policy on its own. Brad Micklea's Mini Shai-Hulud analysis walks through the trust-boundary reasoning that drove this design.
What Jozu does not replace. Jozu does not replace SCA, EDR, IAM, or DLP. It is the AI-artifact and AI-runtime layer alongside those tools, not a substitute for them.
Explore Jozu MCP Registry →
Explore Jozu Agent Guard →
Request a demo →
Frequently asked questions
What is the most common MCP server security risk?
Tool poisoning and prompt injection through MCP responses. Both turn an agent into a vector for exfiltration or unauthorized action, and both have produced documented production incidents in the last 12 months.
Are public MCP servers safe to use?
Not by default. Public servers vary in quality and security posture, and several widely used MCP servers have been the source of documented incidents (Postmark, Smithery, GitHub MCP, mcp-remote). The safe pattern is to mirror approved servers into an internal registry with scanning and signing.
Do existing security tools cover MCP servers?
Not on their own. IAM does not verify the MCP server binary. DLP does not see stdio communication between the agent and the MCP server. API gateways do not see most MCP traffic. MCP server security requires a layer purpose-built for the protocol.
Can MCP servers run in air-gapped environments?
Yes, but only with an architecture that does not depend on a SaaS control plane. Policies and signed artifacts must enforce locally with no remote dependency, and audit logs must sync when connectivity is restored.
Who owns MCP server security inside the organization?
Most commonly, security architecture owns policy, platform engineering operates the registry and runtime, and GRC owns the audit and evidence outputs. AI security or AI center-of-excellence teams coordinate across them.
What is the first step a security team should take?
Inventory what MCP servers are already running. The result is almost always higher than expected and very few of those servers are running through any registry, policy, or audit trail.
How does MCP server security relate to AI agent governance?
MCP server security is one of the four pillars of agent governance: it covers the supply chain of MCP servers (admission) and the runtime controls on MCP tool invocations. Agent governance is the broader program that also covers models, agents, content, approvals, and audit.
Related reading:
- What Is MCP Governance? A Framework for Securing MCP Servers at Scale
- Self-Hosted MCP Registry: Why Enterprises Need More Than GitHub Repos and Zip Files
- AI Agent Governance: A Practical Guide for Enterprise Teams
Ready to secure MCP servers in production? See Jozu MCP Registry or request a demo.