# LiteLLM Model Context Protocol (MCP) Servers and Tools: A Complete Technical Guide

> Master LiteLLM Model Context Protocol (MCP) servers and tools. This guide explains how MCP enables OpenAI compatible clients to discover and execute external tools via a unified gateway.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: deep-dive
- Published: 2026-03-26

---

**TLDR:** LiteLLM implements the Model Context Protocol (MCP) as a native proxy feature, enabling OpenAI-compatible clients to discover and execute external tools through a unified gateway while maintaining the standard API surface.

The BerriAI/litellm repository treats MCP as a **first-class integration layer**, allowing the proxy to act as a gateway to external tool-execution services like Zapier, Linear, or custom servers. This implementation preserves the familiar OpenAI-style request format while handling MCP server discovery, authentication, and execution transparently behind the scenes.

## Architecture of LiteLLM MCP Integration

The MCP stack resides in `litellm/proxy/_experimental/mcp_server` and operates through three logical layers:

**Discovery and Registry** loads server definitions from [`proxy_config.yaml`](https://github.com/BerriAI/litellm/blob/main/proxy_config.yaml) or the database, validates server names, and constructs a global tool registry. The primary entry point is `MCPServerManager.load_servers_from_config` in [`mcp_server_manager.py`](https://github.com/BerriAI/litellm/blob/main/mcp_server_manager.py) (lines 91-122).

**Tool and Prompt Retrieval** communicates with each MCP server via HTTP/SSE, stdio, OAuth2, or SigV4 to enumerate available tools, prompts, and resources. This layer also handles OpenAPI-to-MCP conversion through `MCPServerManager._get_tools_from_server` (lines 250-282).

**Execution and Streaming** intercepts LLM tool calls, routes them to the appropriate MCP server, logs the interaction, and injects results back into the conversation. The core logic lives in `LiteLLM_Proxy_MCP_Handler._execute_tool_calls` in [`litellm/responses/mcp/litellm_proxy_mcp_handler.py`](https://github.com/BerriAI/litellm/blob/main/litellm/responses/mcp/litellm_proxy_mcp_handler.py) (lines 530-608).

## Configuring MCP Servers in LiteLLM

MCP servers are declared in [`proxy_config.yaml`](https://github.com/BerriAI/litellm/blob/main/proxy_config.yaml) under the `mcp_servers` key. During startup, the proxy parses these entries, validates server names (which must not contain the default prefix separator `-`), and instantiates internal `MCPServer` objects.

```yaml

# litellm/proxy/proxy_config.yaml

mcp_servers:
  linear:
    url: "https://api.linear.app/mcp/<api_key>/sse"
    transport: "sse"
    auth_type: "api_key"
    alias: "linear"
    spec_path: "./specs/linear_openapi.json"
    extra_headers:
      X-Linear-Org: "my-org"

```

The `MCPServerManager` reads this configuration via `load_servers_from_config`, supporting OpenAPI specifications through [`openapi_to_mcp_generator.py`](https://github.com/BerriAI/litellm/blob/main/openapi_to_mcp_generator.py) to automatically register REST endpoints as MCP tools.

## Detecting and Routing MCP Tool Calls

When a request includes a `tools` list, LiteLLM checks whether any tool uses the `"type": "mcp"` schema or whether the `server_url` matches the special prefix `litellm_proxy` or the pattern `…/mcp/<name>`.

```python

# litellm/responses/mcp/litellm_proxy_mcp_handler.py

if isinstance(tool, dict) and tool.get("type") == "mcp":
    server_url = tool.get("server_url", "")
    if server_url.startswith(LITELLM_PROXY_MCP_SERVER_URL):
        return True
    if _PROXY_MCP_PATH_RE.match(server_url):
        return True

```

If these conditions are met, the request is handed to the MCP-specific handler path rather than standard tool processing.

## Tool Discovery and Namespacing

To prevent collisions when multiple MCP servers expose identically named tools, LiteLLM implements automatic prefixing. The [`utils.py`](https://github.com/BerriAI/litellm/blob/main/utils.py) module provides `add_server_prefix_to_name` and `split_server_prefix_from_name` to transform tool names into globally unique identifiers like `zapier-send_email`.

The deduplication logic in `_deduplicate_mcp_tools` retains the first occurrence of any tool name and builds a mapping of `tool_name → server_name`. Additionally, `_filter_mcp_tools_by_allowed_tools` enforces allowlists by removing any tool not explicitly permitted in the server configuration.

## Executing MCP Tools in the Proxy

When an LLM response contains `function_call` entries, the `LiteLLM_Proxy_MCP_Handler` orchestrates execution through several precise steps:

1. **Extraction**: Retrieves name, arguments, and call_id via `_extract_tool_call_details`.
2. **Server Resolution**: Uses the tool-to-server map to locate the correct MCP server.
3. **Name Sanitization**: Strips the server prefix using `split_server_prefix_from_name` to reveal the original tool name.
4. **Invocation**: Calls `global_mcp_server_manager.call_tool` with resolved arguments.
5. **Logging**: Records the interaction via `LiteLLMLoggingObj` with pre- and post-call hooks.
6. **Result Formatting**: Parses the output through `_parse_mcp_result` into a human-readable string.
7. **Context Injection**: Returns a dictionary that gets injected back into the conversation history.

For chat-style responses, the proxy constructs follow-up messages containing the original assistant message and tool results, then sends these back to the LLM via `aresponses` to continue the conversation with full context.

## Streaming MCP Responses

LiteLLM supports streaming MCP tool execution through the `MCPEnhancedStreamingIterator` class in [`litellm/responses/mcp/mcp_streaming_iterator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/responses/mcp/mcp_streaming_iterator.py). This iterator implements a three-phase flow:

1. **Discovery Events**: Emits MCP discovery events (list-tools, tool-execution) as soon as they are available.
2. **Streaming**: Continues the standard LLM response stream.
3. **Execution Interleaving**: When encountering a tool call, pauses the stream, executes the MCP call, emits tool-execution events with results, and resumes streaming the follow-up LLM response.

The entire process maintains a standard OpenAI-compatible streaming interface, making the MCP integration transparent to the client.

## Security and Access Control

LiteLLM enforces strict security boundaries for MCP access through multiple mechanisms defined in [`litellm/proxy/_experimental/mcp_server/auth/user_api_key_auth_mcp.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/_experimental/mcp_server/auth/user_api_key_auth_mcp.py).

**Per-User Permissions**: The `MCPServerManager.get_allowed_mcp_servers` function implements role-based access where administrators see all servers unless restricted by `object_permission.mcp_servers`, while non-administrators see the intersection of their permission set and servers flagged `allow_all_keys`.

**Network Isolation**: The `filter_server_ids_by_ip` function hides private servers from external callers, ensuring internal tools remain inaccessible to unauthorized networks.

**Header Security**: `merge_mcp_headers` in [`utils.py`](https://github.com/BerriAI/litellm/blob/main/utils.py) ensures static headers from the server configuration override per-request headers, preventing accidental credential leakage from client-submitted values.

## Summary

- LiteLLM implements **Model Context Protocol (MCP)** as a native proxy feature, enabling OpenAI-compatible clients to use external tools without API changes.
- The architecture spans three layers: **Discovery & Registry** ([`mcp_server_manager.py`](https://github.com/BerriAI/litellm/blob/main/mcp_server_manager.py)), **Tool Retrieval** (`_get_tools_from_server`), and **Execution & Streaming** ([`litellm_proxy_mcp_handler.py`](https://github.com/BerriAI/litellm/blob/main/litellm_proxy_mcp_handler.py)).
- Servers are configured in [`proxy_config.yaml`](https://github.com/BerriAI/litellm/blob/main/proxy_config.yaml) under `mcp_servers`, supporting OpenAPI specs and multiple authentication methods (API key, OAuth2, SigV4).
- **Tool namespacing** prevents collisions through automatic prefixing (e.g., `zapier-send_email`) using utilities in [`utils.py`](https://github.com/BerriAI/litellm/blob/main/utils.py).
- **Streaming support** is handled by `MCPEnhancedStreamingIterator`, which interleaves tool execution without breaking the OpenAI streaming contract.
- **Security controls** include per-user permissions, IP filtering, and header merging to prevent credential leakage.

## Frequently Asked Questions

### What is the Model Context Protocol (MCP) in LiteLLM?

The **Model Context Protocol (MCP)** in LiteLLM is a standardized integration layer that allows the proxy server to discover, route, and execute tools from external MCP-compatible services while maintaining an OpenAI-compatible API surface. It enables LLM clients to invoke third-party tools such as Zapier or Linear without implementing custom authentication or transport logic, as the proxy handles all MCP server communication internally through [`litellm/responses/mcp/litellm_proxy_mcp_handler.py`](https://github.com/BerriAI/litellm/blob/main/litellm/responses/mcp/litellm_proxy_mcp_handler.py).

### How do I add a custom MCP server to LiteLLM?

Define your server in [`proxy_config.yaml`](https://github.com/BerriAI/litellm/blob/main/proxy_config.yaml) under the `mcp_servers` key. Specify the `url`, `transport` (http, sse, or stdio), `auth_type` (api_key, oauth2, or sigv4), and optionally a `spec_path` pointing to an OpenAPI specification for automatic tool generation. The `MCPServerManager.load_servers_from_config` function in [`mcp_server_manager.py`](https://github.com/BerriAI/litellm/blob/main/mcp_server_manager.py) parses this configuration at startup, validates server names against prefix separator rules, and registers the tools globally.

### How does LiteLLM prevent tool name collisions between MCP servers?

LiteLLM implements **automatic namespacing** through utility functions in [`litellm/proxy/_experimental/mcp_server/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/_experimental/mcp_server/utils.py). The `add_server_prefix_to_name` function prepends the server identifier to tool names (e.g., transforming `send_email` into `zapier-send_email`), while `split_server_prefix_from_name` removes the prefix during execution to reveal the original tool name expected by the target server. The `_deduplicate_mcp_tools` function ensures only the first occurrence of any tool name is retained in the global registry.

### Does LiteLLM support streaming responses with MCP tool calls?

Yes. LiteLLM supports **streaming MCP tool execution** through the `MCPEnhancedStreamingIterator` class in [`litellm/responses/mcp/mcp_streaming_iterator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/responses/mcp/mcp_streaming_iterator.py). This iterator emits MCP discovery events, streams the LLM response, pauses to execute tool calls when encountered, emits tool-execution events with results, and then resumes streaming the follow-up LLM response. The entire process maintains a standard OpenAI-compatible streaming interface, making the MCP integration transparent to the client.