Security Considerations for MCP Tool Poisoning Attacks: A Complete Defense Guide
MCP tool poisoning attacks compromise the Model Context Protocol by corrupting tool manifests or implementations, requiring cryptographic verification, policy gating, and scoped OAuth tokens to prevent unauthorized execution.
The curriculum in rohitg00/ai-engineering-from-scratch treats MCP tool poisoning attacks as a critical threat vector in LLM systems, where adversaries attempt to subvert tool registries or inject malicious code into executable implementations. Because the Model Context Protocol (MCP) makes tool-calling a first-class operation across distributed systems, understanding these security considerations is essential for production deployments. This guide examines the architectural vulnerabilities, attack surfaces, and defensive controls implemented in the repository's security lessons.
MCP Architecture and Trust Boundaries
The Model Context Protocol defines a distributed architecture where tools, resources, and prompts are exposed to LLMs through standardized interfaces. According to the site data entry "MCP Security I — Tool Poisoning" and the lesson figure 001-d-mcp-servers.svg in /site/figures-llms-systems.js, the system comprises four security-relevant components.
MCP Server
The server hosts JSON-Schema-defined tools and handles execution via stateless HTTP APIs. Tools are advertised in a capability manifest located at .well-known/mcp-capabilities. The server must validate tool signatures and enforce RBAC before invoking any underlying code.
Registry and Gateway
This central discovery point aggregates manifests from multiple MCP servers and acts as a gatekeeper. It validates manifests for integrity, checks for duplicate or suspicious tool definitions, and applies policy-based filtering using Open Policy Agent (OPA).
OAuth 2.1 Auth Layer
The authorization layer issues access tokens with scoped capabilities. Tokens carry per-tool scopes such as jira:read or s3:list, while destructive tools require an elevated scope (approved:by:human) granted only after human review.
Client (LLM)
The client consumes the registry, selects tools, and issues calls. Clients must validate the tool’s schema and pin the server’s certificate fingerprint to prevent man-in-the-middle attacks.
Attack Vectors in MCP Tool Poisoning
The threat model in /phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/docs/en.md identifies three primary attack vectors grouped under the tool-poisoning class.
Manifest Tampering
An adversary modifies the capability manifest to point a benign-looking tool name to a malicious implementation. This corruption occurs at the registry level or during transport, tricking clients into loading attacker-controlled code.
Tool Code Injection
The underlying executable—typically a Python script or binary—is replaced with malicious code that exfiltrates data or compromises the host system. This attack succeeds when the server lacks execution sandboxing or integrity checks.
Supply-Chain Masquerading
A compromised third-party package is introduced as a new tool in the registry, exploiting trust in external dependencies to inject persistent backdoors into the tool ecosystem.
Implementing Defensive Controls
The repository teaches a defense-in-depth strategy documented in /phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/outputs/skill-mcp-threat-model.md and the capstone lesson /phases/19-capstone-projects/13-mcp-server-with-registry/docs/en.md.
Signed Capability Manifests
Servers must cryptographically sign manifests to ensure integrity. Clients verify these signatures before loading tools, treating unsigned manifests as untrusted.
import json, hmac, hashlib, base64
from pathlib import Path
def sign_manifest(manifest: dict, secret: bytes) -> str:
"""Return a base64-encoded HMAC-SHA256 signature."""
payload = json.dumps(manifest, separators=(",", ":"), sort_keys=True).encode()
sig = hmac.new(secret, payload, hashlib.sha256).digest()
return base64.urlsafe_b64encode(sig).decode()
def verify_manifest(manifest: dict, signature: str, secret: bytes) -> bool:
expected = sign_manifest(manifest, secret)
return hmac.compare_digest(expected, signature)
# Example usage
manifest = {
"tools": [
{"name": "list_s3", "schema": {"type": "object", "properties": {"bucket": {"type": "string"}}}}
]
}
secret = b"super-secret-key"
sig = sign_manifest(manifest, secret)
assert verify_manifest(manifest, sig, secret)
OPA Policy Gating
An Open Policy Agent (OPA) policy denies execution of destructive tools unless a human-approval token is present. This gate operates at the registry level to block suspicious definitions before they reach clients.
package mcp.policy
# Disallow any tool whose name contains "delete" unless the request
# carries the special scope `approved:by:human`.
deny[tool] {
input.tool.name = tool
contains(tool, "delete")
not input.token.scopes[_] == "approved:by:human"
}
OAuth Scope Enforcement
Fine-grained scope pinning ensures that destructive operations require explicit authorization. The MCP Auth in Production lesson in /site/data.js demonstrates how scope inspection prevents privilege escalation.
def token_has_scope(token: dict, required: str) -> bool:
"""Assumes `token` is a decoded JWT payload."""
return required in token.get("scp", [])
# Example request handling
def handle_tool_call(token, tool_name):
if tool_name == "jira_create":
if not token_has_scope(token, "approved:by:human"):
raise PermissionError("Human approval required for destructive tool")
# … invoke the tool safely …
Runtime Isolation and Supply-Chain Verification
Each tool must execute in a sandboxed environment (e.g., Docker) with restricted filesystem and network access. Additionally, tools sourced from external registries require verification against known hash lists to prevent supply-chain masquerading.
End-to-End Mitigation Workflow
The complete defense workflow spans four lifecycle phases as illustrated in the capstone lesson:
- During Registration – The registry validates manifest signatures and evaluates OPA policies before accepting a tool.
- During Discovery – Clients fetch signed manifests and abort loading if signature verification fails or certificate fingerprints mismatch.
- Before Execution – The OAuth token scope is inspected; if the tool is destructive, the request routes to a dedicated human-approval MCP server.
- Post-Execution – Audit logs are emitted via OpenTelemetry for forensic analysis and compliance replay.
Summary
- Tool manifests must be cryptographically signed and verified before loading to prevent manifest tampering.
- OPA policies should gate destructive tools at the registry level, requiring human-approval scopes for dangerous operations.
- OAuth scope pinning separates low-risk read operations from high-risk destructive calls, enforced before tool invocation.
- Runtime isolation and supply-chain hash verification prevent code injection and masquerading attacks in production environments.
Frequently Asked Questions
What is an MCP tool poisoning attack?
An MCP tool poisoning attack occurs when an adversary corrupts the tool registry or implementation—either by tampering with the capability manifest, injecting malicious code into the tool executable, or introducing a compromised third-party package—to trick an LLM into executing unauthorized or harmful operations.
How do signed capability manifests prevent MCP tool poisoning?
Signed capability manifests use HMAC-SHA256 signatures to cryptographically bind the tool definition to the server's identity. When implemented as shown in /phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/docs/en.md, clients reject any manifest failing signature verification, preventing attackers from redirecting tool calls to malicious endpoints.
What role does OPA play in MCP security?
Open Policy Agent (OPA) acts as a policy gatekeeper at the registry level, evaluating tool definitions against security rules before they are advertised to clients. As demonstrated in the capstone project /phases/19-capstone-projects/13-mcp-server-with-registry/docs/en.md, OPA can block tools with dangerous characteristics (like "delete" operations) unless the request includes a specific human-approval scope.
Why is OAuth scope pinning important for MCP tool security?
OAuth scope pinning ensures that each tool call carries cryptographically bound authorization tokens with fine-grained permissions. By requiring elevated scopes like approved:by:human for destructive operations, the system prevents automated exploitation of poisoned tools even if an attacker successfully compromises a lower-privilege component.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →