How to Add a Custom LLM Provider to LiteLLM: Complete Implementation Guide

LiteLLM enables you to integrate any proprietary or third-party language model by subclassing CustomLLM, registering your provider in custom_provider_map, and letting the built-in router handle request dispatching.

The BerriAI/litellm repository provides a pluggable architecture that treats every supported model as a provider implementing a common interface. Adding a custom LLM provider to LiteLLM requires implementing three core components: a handler class, a registry entry, and router integration. This guide walks through the complete implementation using the actual source code structure from the repository.

Understanding the Custom Provider Architecture

LiteLLM's extensibility relies on three coordinated components that bridge your custom logic with the library's unified API:

  • Handler class – A concrete implementation of the CustomLLM abstract base that defines how your model handles completion, acompletion, streaming, and embedding requests. The base class lives in litellm/llms/custom_llm.py.

  • Provider registry – The global litellm.custom_provider_map list stores CustomLLMItem entries that map provider names to handler instances. This registry is defined in litellm/__init__.py alongside the internal _custom_providers tracker.

  • Router integration – The core request handler in litellm/main.py detects custom providers via the elif custom_llm_provider in litellm._custom_providers: block and dispatches calls using custom_chat_llm_router.

The registration helper custom_llm_setup() in litellm/utils.py synchronizes the registry with LiteLLM's internal provider lists at runtime.

Step 1: Create a Custom LLM Handler

Subclass CustomLLM from litellm/llms/custom_llm.py to implement the methods that communicate with your service. You must define synchronous and asynchronous variants for completion and streaming operations.


# my_custom_llm.py

import litellm
from litellm import CustomLLM
from litellm.types.utils import ModelResponse

class MyCustomLLM(CustomLLM):
    def completion(self, *args, **kwargs) -> ModelResponse:
        # Implement your external service call here

        return litellm.completion(
            model="gpt-3.5-turbo",
            messages=[{"role": "user", "content": "Hello"}],
            mock_response="Hi from my custom LLM!",
        )  # type: ignore

    async def acompletion(self, *args, **kwargs) -> ModelResponse:
        # Async implementation using httpx.AsyncClient or similar

        return litellm.completion(
            model="gpt-3.5-turbo",
            messages=[{"role": "user", "content": "Hello"}],
            mock_response="Hi async!",
        )  # type: ignore

    
    def streaming(self, *args, **kwargs):
        # Optional: implement for streaming support

        pass
    
    async def astreaming(self, *args, **kwargs):
        # Optional: implement for async streaming support

        pass

The CustomLLM base class defines the interface contract that LiteLLM expects. Return values must conform to ModelResponse or generator types for streaming operations.

Step 2: Register Your Provider

Insert your handler into litellm.custom_provider_map before making any completion calls. Each entry follows the CustomLLMItem structure defined in litellm/types/llms/custom_llm.py.


# registration.py

import litellm
from my_custom_llm import MyCustomLLM
from litellm.types.llms.custom_llm import CustomLLMItem

# Create the registry entry

my_provider: CustomLLMItem = {
    "provider": "my-custom-llm",      # Name used in model= parameter

    "custom_handler": MyCustomLLM(),
    "model_group": "my_custom_group", # Optional grouping metadata

    "is_proxy_model": False,          # True if proxying a known provider

}

# Append to global registry

litellm.custom_provider_map.append(my_provider)

# Synchronize with LiteLLM's internal provider lists

litellm.custom_llm_setup()

Execute this registration code during application startup or inside a plugin package's __init__.py. The custom_llm_setup() function populates litellm.provider_list and the _custom_providers internal tracker, making your provider discoverable to the router.

Step 3: Route Requests to Your Provider

When you call litellm.completion(model="my-custom-llm", ...), the logic in litellm/main.py detects the provider name in litellm._custom_providers and routes the request through custom_chat_llm_router. This dispatcher selects the appropriate method—completion vs acompletion, streaming vs non-streaming—based on your call parameters.

LiteLLM automatically handles:

  • Sync vs async detection – Routes to completion or acompletion based on the calling context
  • Streaming protocol – Invokes streaming or astreaming when stream=True is passed
  • Response wrapping – Converts your return values into standard LiteLLM response objects

Complete Working Example

Here is a complete implementation that integrates a fictional external API:


# custom_integration.py

import litellm
from litellm import CustomLLM
from litellm.types.llms.custom_llm import CustomLLMItem
from litellm.types.utils import ModelResponse
import httpx

class ExternalAPIHandler(CustomLLM):
    def completion(self, model: str, messages: list, api_key: str, **kwargs) -> ModelResponse:
        # Replace with actual external API call

        response_text = f"Processed {len(messages)} messages via external API"
        
        return ModelResponse(
            choices=[{
                "message": {"content": response_text, "role": "assistant"},
                "index": 0,
                "finish_reason": "stop"
            }],
            model=model
        )
    
    async def acompletion(self, model: str, messages: list, api_key: str, **kwargs) -> ModelResponse:
        async with httpx.AsyncClient() as client:
            # Async external API implementation

            return self.completion(model, messages, api_key, **kwargs)

# Registration

handler = ExternalAPIHandler()
litellm.custom_provider_map.append({
    "provider": "external-api",
    "custom_handler": handler,
    "model_group": "external_models",
    "is_proxy_model": False
})
litellm.custom_llm_setup()

# Usage

response = litellm.completion(
    model="external-api",
    messages=[{"role": "user", "content": "Analyze this data"}]
)
print(response.choices[0].message.content)

Summary

  • Subclass CustomLLM from litellm/llms/custom_llm.py to implement completion, acompletion, and optional streaming methods for your service.
  • Register via custom_provider_map by appending a CustomLLMItem dictionary that maps your provider name to the handler instance.
  • Call custom_llm_setup() from litellm/utils.py to synchronize the registry with LiteLLM's internal provider lists before making requests.
  • Route automatically through litellm/main.py by passing your registered provider name to the model parameter in litellm.completion() or litellm.acompletion().

Frequently Asked Questions

What methods must I implement in a custom LLM provider?

You must implement at least completion and acompletion to handle synchronous and asynchronous requests. For streaming support, implement streaming and astreaming. The CustomLLM base class in litellm/llms/custom_llm.py defines these method signatures and provides the interface contract that LiteLLM expects.

Where does LiteLLM store the list of available custom providers?

LiteLLM maintains two key structures in litellm/__init__.py: custom_provider_map (the public registry of CustomLLMItem entries) and _custom_providers (the internal list of provider names). The custom_llm_setup() function in litellm/utils.py populates these structures at runtime.

Can I register multiple custom providers in the same application?

Yes. Append multiple CustomLLMItem dictionaries to litellm.custom_provider_map with unique provider names, then call litellm.custom_llm_setup() once after all registrations. Each provider operates independently, allowing you to route different models to different external services using the same LiteLLM API.

How does LiteLLM decide which method to call on my custom handler?

The routing logic in litellm/main.py checks if the custom_llm_provider parameter exists in litellm._custom_providers. If found, it uses custom_chat_llm_router to select between completion, acompletion, streaming, or astreaming based on whether the call uses async and whether stream=True is passed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →