LiteLLM Model Context Protocol (MCP) Servers and Tools: A Complete Technical Guide
TLDR: LiteLLM implements the Model Context Protocol (MCP) as a native proxy feature, enabling OpenAI-compatible clients to discover and execute external tools through a unified gateway while maintaining the standard API surface.
The BerriAI/litellm repository treats MCP as a first-class integration layer, allowing the proxy to act as a gateway to external tool-execution services like Zapier, Linear, or custom servers. This implementation preserves the familiar OpenAI-style request format while handling MCP server discovery, authentication, and execution transparently behind the scenes.
Architecture of LiteLLM MCP Integration
The MCP stack resides in litellm/proxy/_experimental/mcp_server and operates through three logical layers:
Discovery and Registry loads server definitions from proxy_config.yaml or the database, validates server names, and constructs a global tool registry. The primary entry point is MCPServerManager.load_servers_from_config in mcp_server_manager.py (lines 91-122).
Tool and Prompt Retrieval communicates with each MCP server via HTTP/SSE, stdio, OAuth2, or SigV4 to enumerate available tools, prompts, and resources. This layer also handles OpenAPI-to-MCP conversion through MCPServerManager._get_tools_from_server (lines 250-282).
Execution and Streaming intercepts LLM tool calls, routes them to the appropriate MCP server, logs the interaction, and injects results back into the conversation. The core logic lives in LiteLLM_Proxy_MCP_Handler._execute_tool_calls in litellm/responses/mcp/litellm_proxy_mcp_handler.py (lines 530-608).
Configuring MCP Servers in LiteLLM
MCP servers are declared in proxy_config.yaml under the mcp_servers key. During startup, the proxy parses these entries, validates server names (which must not contain the default prefix separator -), and instantiates internal MCPServer objects.
# litellm/proxy/proxy_config.yaml
mcp_servers:
linear:
url: "https://api.linear.app/mcp/<api_key>/sse"
transport: "sse"
auth_type: "api_key"
alias: "linear"
spec_path: "./specs/linear_openapi.json"
extra_headers:
X-Linear-Org: "my-org"
The MCPServerManager reads this configuration via load_servers_from_config, supporting OpenAPI specifications through openapi_to_mcp_generator.py to automatically register REST endpoints as MCP tools.
Detecting and Routing MCP Tool Calls
When a request includes a tools list, LiteLLM checks whether any tool uses the "type": "mcp" schema or whether the server_url matches the special prefix litellm_proxy or the pattern …/mcp/<name>.
# litellm/responses/mcp/litellm_proxy_mcp_handler.py
if isinstance(tool, dict) and tool.get("type") == "mcp":
server_url = tool.get("server_url", "")
if server_url.startswith(LITELLM_PROXY_MCP_SERVER_URL):
return True
if _PROXY_MCP_PATH_RE.match(server_url):
return True
If these conditions are met, the request is handed to the MCP-specific handler path rather than standard tool processing.
Tool Discovery and Namespacing
To prevent collisions when multiple MCP servers expose identically named tools, LiteLLM implements automatic prefixing. The utils.py module provides add_server_prefix_to_name and split_server_prefix_from_name to transform tool names into globally unique identifiers like zapier-send_email.
The deduplication logic in _deduplicate_mcp_tools retains the first occurrence of any tool name and builds a mapping of tool_name → server_name. Additionally, _filter_mcp_tools_by_allowed_tools enforces allowlists by removing any tool not explicitly permitted in the server configuration.
Executing MCP Tools in the Proxy
When an LLM response contains function_call entries, the LiteLLM_Proxy_MCP_Handler orchestrates execution through several precise steps:
- Extraction: Retrieves name, arguments, and call_id via
_extract_tool_call_details. - Server Resolution: Uses the tool-to-server map to locate the correct MCP server.
- Name Sanitization: Strips the server prefix using
split_server_prefix_from_nameto reveal the original tool name. - Invocation: Calls
global_mcp_server_manager.call_toolwith resolved arguments. - Logging: Records the interaction via
LiteLLMLoggingObjwith pre- and post-call hooks. - Result Formatting: Parses the output through
_parse_mcp_resultinto a human-readable string. - Context Injection: Returns a dictionary that gets injected back into the conversation history.
For chat-style responses, the proxy constructs follow-up messages containing the original assistant message and tool results, then sends these back to the LLM via aresponses to continue the conversation with full context.
Streaming MCP Responses
LiteLLM supports streaming MCP tool execution through the MCPEnhancedStreamingIterator class in litellm/responses/mcp/mcp_streaming_iterator.py. This iterator implements a three-phase flow:
- Discovery Events: Emits MCP discovery events (list-tools, tool-execution) as soon as they are available.
- Streaming: Continues the standard LLM response stream.
- Execution Interleaving: When encountering a tool call, pauses the stream, executes the MCP call, emits tool-execution events with results, and resumes streaming the follow-up LLM response.
The entire process maintains a standard OpenAI-compatible streaming interface, making the MCP integration transparent to the client.
Security and Access Control
LiteLLM enforces strict security boundaries for MCP access through multiple mechanisms defined in litellm/proxy/_experimental/mcp_server/auth/user_api_key_auth_mcp.py.
Per-User Permissions: The MCPServerManager.get_allowed_mcp_servers function implements role-based access where administrators see all servers unless restricted by object_permission.mcp_servers, while non-administrators see the intersection of their permission set and servers flagged allow_all_keys.
Network Isolation: The filter_server_ids_by_ip function hides private servers from external callers, ensuring internal tools remain inaccessible to unauthorized networks.
Header Security: merge_mcp_headers in utils.py ensures static headers from the server configuration override per-request headers, preventing accidental credential leakage from client-submitted values.
Summary
- LiteLLM implements Model Context Protocol (MCP) as a native proxy feature, enabling OpenAI-compatible clients to use external tools without API changes.
- The architecture spans three layers: Discovery & Registry (
mcp_server_manager.py), Tool Retrieval (_get_tools_from_server), and Execution & Streaming (litellm_proxy_mcp_handler.py). - Servers are configured in
proxy_config.yamlundermcp_servers, supporting OpenAPI specs and multiple authentication methods (API key, OAuth2, SigV4). - Tool namespacing prevents collisions through automatic prefixing (e.g.,
zapier-send_email) using utilities inutils.py. - Streaming support is handled by
MCPEnhancedStreamingIterator, which interleaves tool execution without breaking the OpenAI streaming contract. - Security controls include per-user permissions, IP filtering, and header merging to prevent credential leakage.
Frequently Asked Questions
What is the Model Context Protocol (MCP) in LiteLLM?
The Model Context Protocol (MCP) in LiteLLM is a standardized integration layer that allows the proxy server to discover, route, and execute tools from external MCP-compatible services while maintaining an OpenAI-compatible API surface. It enables LLM clients to invoke third-party tools such as Zapier or Linear without implementing custom authentication or transport logic, as the proxy handles all MCP server communication internally through litellm/responses/mcp/litellm_proxy_mcp_handler.py.
How do I add a custom MCP server to LiteLLM?
Define your server in proxy_config.yaml under the mcp_servers key. Specify the url, transport (http, sse, or stdio), auth_type (api_key, oauth2, or sigv4), and optionally a spec_path pointing to an OpenAPI specification for automatic tool generation. The MCPServerManager.load_servers_from_config function in mcp_server_manager.py parses this configuration at startup, validates server names against prefix separator rules, and registers the tools globally.
How does LiteLLM prevent tool name collisions between MCP servers?
LiteLLM implements automatic namespacing through utility functions in litellm/proxy/_experimental/mcp_server/utils.py. The add_server_prefix_to_name function prepends the server identifier to tool names (e.g., transforming send_email into zapier-send_email), while split_server_prefix_from_name removes the prefix during execution to reveal the original tool name expected by the target server. The _deduplicate_mcp_tools function ensures only the first occurrence of any tool name is retained in the global registry.
Does LiteLLM support streaming responses with MCP tool calls?
Yes. LiteLLM supports streaming MCP tool execution through the MCPEnhancedStreamingIterator class in litellm/responses/mcp/mcp_streaming_iterator.py. This iterator emits MCP discovery events, streams the LLM response, pauses to execute tool calls when encountered, emits tool-execution events with results, and then resumes streaming the follow-up LLM response. The entire process maintains a standard OpenAI-compatible streaming interface, making the MCP integration transparent to the client.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →