# Security Considerations for MCP Tool Poisoning Attacks: A Complete Defense Guide

> Defend against MCP tool poisoning attacks. Learn how to secure your AI models with cryptographic verification, policy gating, and scoped OAuth tokens for complete protection.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-26

---

**MCP tool poisoning attacks compromise the Model Context Protocol by corrupting tool manifests or implementations, requiring cryptographic verification, policy gating, and scoped OAuth tokens to prevent unauthorized execution.**

The curriculum in `rohitg00/ai-engineering-from-scratch` treats **MCP tool poisoning attacks** as a critical threat vector in LLM systems, where adversaries attempt to subvert tool registries or inject malicious code into executable implementations. Because the Model Context Protocol (MCP) makes tool-calling a first-class operation across distributed systems, understanding these security considerations is essential for production deployments. This guide examines the architectural vulnerabilities, attack surfaces, and defensive controls implemented in the repository's security lessons.

## MCP Architecture and Trust Boundaries

The Model Context Protocol defines a distributed architecture where **tools**, **resources**, and **prompts** are exposed to LLMs through standardized interfaces. According to the site data entry *"MCP Security I — Tool Poisoning"* and the lesson figure `001-d-mcp-servers.svg` in [`/site/figures-llms-systems.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//site/figures-llms-systems.js), the system comprises four security-relevant components.

### MCP Server

The server hosts JSON-Schema-defined tools and handles execution via stateless HTTP APIs. Tools are advertised in a **capability manifest** located at `.well-known/mcp-capabilities`. The server must validate **tool signatures** and enforce **RBAC** before invoking any underlying code.

### Registry and Gateway

This central discovery point aggregates manifests from multiple MCP servers and acts as a **gatekeeper**. It validates manifests for integrity, checks for duplicate or suspicious tool definitions, and applies **policy-based filtering** using Open Policy Agent (OPA).

### OAuth 2.1 Auth Layer

The authorization layer issues access tokens with scoped capabilities. Tokens carry **per-tool scopes** such as `jira:read` or `s3:list`, while destructive tools require an **elevated scope** (`approved:by:human`) granted only after human review.

### Client (LLM)

The client consumes the registry, selects tools, and issues calls. Clients must **validate** the tool’s schema and **pin** the server’s certificate fingerprint to prevent man-in-the-middle attacks.

## Attack Vectors in MCP Tool Poisoning

The threat model in [`/phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/docs/en.md) identifies three primary attack vectors grouped under the *tool-poisoning* class.

### Manifest Tampering

An adversary modifies the capability manifest to point a benign-looking tool name to a malicious implementation. This corruption occurs at the registry level or during transport, tricking clients into loading attacker-controlled code.

### Tool Code Injection

The underlying executable—typically a Python script or binary—is replaced with malicious code that exfiltrates data or compromises the host system. This attack succeeds when the server lacks execution sandboxing or integrity checks.

### Supply-Chain Masquerading

A compromised third-party package is introduced as a new tool in the registry, exploiting trust in external dependencies to inject persistent backdoors into the tool ecosystem.

## Implementing Defensive Controls

The repository teaches a defense-in-depth strategy documented in [`/phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/outputs/skill-mcp-threat-model.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/outputs/skill-mcp-threat-model.md) and the capstone lesson [`/phases/19-capstone-projects/13-mcp-server-with-registry/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/19-capstone-projects/13-mcp-server-with-registry/docs/en.md).

### Signed Capability Manifests

Servers must cryptographically sign manifests to ensure integrity. Clients verify these signatures before loading tools, treating unsigned manifests as untrusted.

```python
import json, hmac, hashlib, base64
from pathlib import Path

def sign_manifest(manifest: dict, secret: bytes) -> str:
    """Return a base64-encoded HMAC-SHA256 signature."""
    payload = json.dumps(manifest, separators=(",", ":"), sort_keys=True).encode()
    sig = hmac.new(secret, payload, hashlib.sha256).digest()
    return base64.urlsafe_b64encode(sig).decode()

def verify_manifest(manifest: dict, signature: str, secret: bytes) -> bool:
    expected = sign_manifest(manifest, secret)
    return hmac.compare_digest(expected, signature)

# Example usage

manifest = {
    "tools": [
        {"name": "list_s3", "schema": {"type": "object", "properties": {"bucket": {"type": "string"}}}}
    ]
}
secret = b"super-secret-key"
sig = sign_manifest(manifest, secret)
assert verify_manifest(manifest, sig, secret)

```

### OPA Policy Gating

An Open Policy Agent (OPA) policy denies execution of destructive tools unless a human-approval token is present. This gate operates at the registry level to block suspicious definitions before they reach clients.

```rego
package mcp.policy

# Disallow any tool whose name contains "delete" unless the request

# carries the special scope `approved:by:human`.

deny[tool] {
    input.tool.name = tool
    contains(tool, "delete")
    not input.token.scopes[_] == "approved:by:human"
}

```

### OAuth Scope Enforcement

Fine-grained scope pinning ensures that destructive operations require explicit authorization. The *MCP Auth in Production* lesson in [`/site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//site/data.js) demonstrates how scope inspection prevents privilege escalation.

```python
def token_has_scope(token: dict, required: str) -> bool:
    """Assumes `token` is a decoded JWT payload."""
    return required in token.get("scp", [])

# Example request handling

def handle_tool_call(token, tool_name):
    if tool_name == "jira_create":
        if not token_has_scope(token, "approved:by:human"):
            raise PermissionError("Human approval required for destructive tool")
    # … invoke the tool safely …

```

### Runtime Isolation and Supply-Chain Verification

Each tool must execute in a sandboxed environment (e.g., Docker) with restricted filesystem and network access. Additionally, tools sourced from external registries require verification against known hash lists to prevent supply-chain masquerading.

## End-to-End Mitigation Workflow

The complete defense workflow spans four lifecycle phases as illustrated in the capstone lesson:

1. **During Registration** – The registry validates manifest signatures and evaluates OPA policies before accepting a tool.
2. **During Discovery** – Clients fetch signed manifests and abort loading if signature verification fails or certificate fingerprints mismatch.
3. **Before Execution** – The OAuth token scope is inspected; if the tool is destructive, the request routes to a dedicated *human-approval* MCP server.
4. **Post-Execution** – Audit logs are emitted via OpenTelemetry for forensic analysis and compliance replay.

## Summary

- **Tool manifests** must be cryptographically signed and verified before loading to prevent manifest tampering.
- **OPA policies** should gate destructive tools at the registry level, requiring human-approval scopes for dangerous operations.
- **OAuth scope pinning** separates low-risk read operations from high-risk destructive calls, enforced before tool invocation.
- **Runtime isolation** and supply-chain hash verification prevent code injection and masquerading attacks in production environments.

## Frequently Asked Questions

### What is an MCP tool poisoning attack?

An MCP tool poisoning attack occurs when an adversary corrupts the tool registry or implementation—either by tampering with the capability manifest, injecting malicious code into the tool executable, or introducing a compromised third-party package—to trick an LLM into executing unauthorized or harmful operations.

### How do signed capability manifests prevent MCP tool poisoning?

Signed capability manifests use HMAC-SHA256 signatures to cryptographically bind the tool definition to the server's identity. When implemented as shown in [`/phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/13-tools-and-protocols/15-mcp-security-tool-poisoning/docs/en.md), clients reject any manifest failing signature verification, preventing attackers from redirecting tool calls to malicious endpoints.

### What role does OPA play in MCP security?

Open Policy Agent (OPA) acts as a policy gatekeeper at the registry level, evaluating tool definitions against security rules before they are advertised to clients. As demonstrated in the capstone project [`/phases/19-capstone-projects/13-mcp-server-with-registry/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/19-capstone-projects/13-mcp-server-with-registry/docs/en.md), OPA can block tools with dangerous characteristics (like "delete" operations) unless the request includes a specific human-approval scope.

### Why is OAuth scope pinning important for MCP tool security?

OAuth scope pinning ensures that each tool call carries cryptographically bound authorization tokens with fine-grained permissions. By requiring elevated scopes like `approved:by:human` for destructive operations, the system prevents automated exploitation of poisoned tools even if an attacker successfully compromises a lower-privilege component.