LiteLLM Proxy Custom Prompt Management: A Complete Guide to Dynamic Prompt Templates

LiteLLM's custom prompt management feature enables developers to inject, version, and retrieve tailored prompt templates at runtime through an in-memory registry and REST API endpoints without modifying underlying model code.

LiteLLM's proxy custom prompt management system provides a production-ready solution for dynamically customizing LLM prompts across multiple providers. Implemented in the BerriAI/litellm repository, this feature exposes a callback-based architecture that intercepts requests to render customized prompts with variable substitution.

Core Architecture Components

CustomPromptManagement Callback Class

The CustomPromptManagement class in litellm/integrations/custom_prompt_management.py implements the callback interface that parses custom prompt definitions and renders the final prompt payload. This class is added to the global litellm.logging_callback_manager so every request can be intercepted and modified before reaching the LLM provider.

The callback handles variable substitution using the {{variable}} syntax and manages the transformation of template strings into provider-specific message formats.

In-Memory Prompt Registry

The IN_MEMORY_PROMPT_REGISTRY in litellm/proxy/prompts/prompt_registry.py maintains an active mapping of prompt_id → CustomPromptManagement instances. This registry provides fast lookup for prompt retrieval during request handling.

Key methods include:

  • PromptRegistry.register_prompt (lines 140-144): Validates prompt configurations and instantiates CustomPromptManagement objects
  • Line 145: Registers the callback with the global callback manager
  • PromptRegistry.get_custom_prompt (line 175): Retrieves the correct callback instance during request processing

The registry stores mappings in prompt_id_to_custom_prompt, enabling UUID-based prompt identification and versioning.

REST API Endpoints

The proxy exposes HTTP routes in litellm/proxy/prompts/prompt_endpoints.py for external prompt management:

  • /prompt/add: Creates and registers new prompt templates
  • /prompt/delete: Removes prompts from the registry
  • /prompt/get: Retrieves existing prompt configurations

These endpoints are implemented around lines 981-982, validating JSON payloads and invoking registry methods to maintain consistency across API calls.

Fallback Dictionary

For backward compatibility, litellm/main.py (lines 2758-2762) maintains a global custom_prompt_dict that maps model names to prompt fragments. This fallback activates when no explicit prompt_id is supplied, supporting legacy code that passed dictionaries directly to the SDK.

Request Flow and Variable Substitution

The complete request lifecycle follows this pattern:

  1. Registration: Client POSTs to /prompt/add with prompt_id and optional prompt_variables
  2. Validation: PromptRegistry.register_prompt validates the payload (lines 140-144) and instantiates a CustomPromptManagement object
  3. Callback Registration: The system registers the callback with the global manager (line 145) and stores the mapping in prompt_id_to_custom_prompt
  4. Retrieval: During generation requests, the proxy calls PromptRegistry.get_custom_prompt (line 175) to fetch the correct instance
  5. Rendering: The CustomPromptManagement instance performs variable substitution and renders the final prompt
  6. Fallback: If no prompt_id is present, the request falls back to litellm.custom_prompt_dict (lines 2758-2762 in main.py)

Implementation Guide

Registering Custom Prompts via the Proxy API

Use the proxy client to register templates with variable placeholders:

import litellm
from litellm import proxy

# Define a prompt template with variable placeholder

my_prompt = {
    "prompt_id": "welcome-01",
    "prompt_template": "You are a helpful assistant. Greet the user: {{user_name}}",
    "prompt_variables": ["user_name"]
}

# Register with the LiteLLM proxy

response = proxy.add_prompt(**my_prompt)   # POST /prompt/add

print(response)  # → {"status":"success","prompt_id":"welcome-01"}

This maps to the handler in prompt_endpoints.py (lines 981-982), which stores the prompt in the registry.

Using Registered Prompts in Completion Calls

Reference registered prompts by ID during inference:

import litellm

# Pass prompt_id and variables when calling the LLM

response = litellm.completion(
    model="gpt-3.5-turbo",
    messages=[
        {"role": "user", "content": "What's the weather?"}
    ],
    prompt_id="welcome-01",
    prompt_variables={"user_name": "Alice"}  # replaces {{user_name}}

)

print(response.choices[0].message['content'])

The system retrieves the CustomPromptManagement instance via PromptRegistry.get_custom_prompt and renders the final prompt with substituted variables.

Fallback Configuration with custom_prompt_dict

For scenarios without explicit prompt IDs, configure per-model defaults:


# Set a per-model prompt dictionary directly

litellm.custom_prompt_dict = {
    "gpt-3.5-turbo": {
        "initial_prompt_value": "You are an expert translator.",
        "final_prompt_value": "{{input_text}}"
    }
}

response = litellm.completion(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "Translate: Hello"}]
)

The SDK reads custom_prompt_dict at line 2758 in main.py, merging it with the request before provider transmission.

Deleting and Managing Prompts

Remove obsolete prompts to free registry resources:

proxy.delete_prompt(prompt_id="welcome-01")   # DELETE /prompt/delete

The endpoint (line 982 in prompt_endpoints.py) removes the entry from prompt_id_to_custom_prompt and unregisters the associated callback.

Summary

  • LiteLLM proxy custom prompt management uses an in-memory registry (IN_MEMORY_PROMPT_REGISTRY in prompt_registry.py) to map UUID-based prompt IDs to CustomPromptManagement callback instances
  • The system intercepts requests via the global callback manager to perform runtime variable substitution using {{variable}} syntax
  • REST endpoints in prompt_endpoints.py (lines 981-982) expose CRUD operations for prompt templates without requiring service redeployment
  • A fallback mechanism in main.py (lines 2758-2762) supports legacy custom_prompt_dict configurations when no prompt_id is specified
  • The implementation is validated by test suites in tests/integrations/test_custom_prompt_management.py and proxy endpoint tests

Frequently Asked Questions

How does LiteLLM handle variable substitution in custom prompts?

LiteLLM performs variable substitution through the CustomPromptManagement class in litellm/integrations/custom_prompt_management.py. When a request includes prompt_variables (a dictionary mapping variable names to values), the callback replaces template placeholders formatted as {{variable_name}} with the corresponding values before sending the request to the LLM provider.

What happens if I don't provide a prompt_id in the request?

If no prompt_id is specified, the system falls back to the global custom_prompt_dict defined in litellm/main.py (lines 2758-2762). This dictionary maps model names to prompt fragments, providing backward compatibility for legacy integrations that passed prompt dictionaries directly to the SDK rather than using the registry-based system.

Where are custom prompts stored in the LiteLLM proxy?

Custom prompts are stored in the IN_MEMORY_PROMPT_REGISTRY, an in-memory map located in litellm/proxy/prompts/prompt_registry.py. This registry maintains a mapping of prompt_id strings to CustomPromptManagement instances in the prompt_id_to_custom_prompt dictionary. The storage is ephemeral and exists only for the duration of the proxy server process.

Can I use custom prompt management with any LLM provider supported by LiteLLM?

Yes. Because the CustomPromptManagement callback integrates with the global litellm.logging_callback_manager at the proxy layer, prompt modification occurs before provider-specific formatting. This architecture ensures that custom prompts work across all providers supported by LiteLLM (OpenAI, Anthropic, Azure, Bedrock, etc.) without requiring provider-specific implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →