How LiteLLM Virtual Keys Work for API Key Management: A Complete Technical Guide

LiteLLM virtual keys are proxy-side credentials that authenticate requests, enforce spend limits, and track per-key usage before forwarding traffic to underlying LLM providers.

LiteLLM virtual keys provide a secure abstraction layer for API key management within the BerriAI/litellm proxy. These short-lived tokens decouple client authentication from provider credentials, enabling fine-grained access control, budget enforcement, and detailed spend tracking without exposing your master API keys.

What Are LiteLLM Virtual Keys

LiteLLM virtual keys are database-backed tokens stored in the LiteLLM_VerificationToken table that act as bearer credentials for every request reaching the proxy. Unlike static provider API keys, virtual keys support granular policy enforcement including model restrictions, route limitations, and team-based access controls.

Authentication and Header Extraction

Every request must present a virtual key via the x-litellm-api-key header or the standard Authorization header. The proxy extracts this value in litellm/proxy/pass_through_endpoints/common_utils.py using the get_litellm_virtual_key function:


# litellm/proxy/pass_through_endpoints/common_utils.py

def get_litellm_virtual_key(request: Request) -> str:
    """
    Extract and format API key from request headers.
    Prioritizes x-litellm-api-key over Authorization header.
    """
    litellm_api_key = request.headers.get("x-litellm-api-key")
    if litellm_api_key:
        return f"Bearer {litellm_api_key}"
    return request.headers.get("Authorization", "")

This extraction logic prioritizes the custom header, formatting it as a standard Bearer token for downstream processing.

Spend Tracking and Budget Enforcement

Each virtual key maintains aggregated spend metrics in the proxy's cache using the VIRTUAL_KEY_SPEND_CACHE_KEY_PREFIX constant. The system enforces both soft-budget and hard-budget limits defined in litellm/proxy/hooks/model_max_budget_limiter.py before forwarding requests to providers.

Proxy-Side Resolution and Validation

JWT-to-Virtual-Key Mapping

For organizations using identity providers, virtual keys can be dynamically resolved from JWT claims. The _resolve_jwt_to_virtual_key function in litellm/proxy/auth/user_api_key_auth.py handles this mapping:


# litellm/proxy/auth/user_api_key_auth.py

async def _resolve_jwt_to_virtual_key(
    jwt_claims: dict,
    jwt_handler: JWTHandler,
    prisma_client: Optional[PrismaClient],
    user_api_key_cache: DualCache,
    parent_otel_span: Optional[Span],
    proxy_logging_obj: ProxyLogging,
) -> Optional[UserAPIKeyAuth]:
    virtual_key_claim_field = jwt_handler.litellm_jwtauth.virtual_key_claim_field
    if virtual_key_claim_field is None:
        return None

    claim_value = get_nested_value(data=jwt_claims, key_path=virtual_key_claim_field, default=None)
    if claim_value is None:
        return None

    cache_key = f"jwt_key_mapping:{virtual_key_claim_field}:{claim_value}"
    cached_mapping = await user_api_key_cache.async_get_cache(cache_key)
    # ... resolves to LiteLLM_VerificationToken row

When a JWT claim maps to a virtual key, the proxy loads the complete token object including budget configurations, model whitelists, and user or team bindings.

Runtime Policy Enforcement

Budget Check Implementation

Before proxying a request, the system validates budget constraints using cache keys formatted as virtual_key_spend:{user_api_key_hash}:{model}:{budget_duration}:


# litellm/proxy/hooks/model_max_budget_limiter.py

VIRTUAL_KEY_SPEND_CACHE_KEY_PREFIX = "virtual_key_spend"

# Example: fetch spend for a key+model

virtual_key_model_spend_cache_key = f"{VIRTUAL_KEY_SPEND_CACHE_KEY_PREFIX}:{user_api_key_hash}:{model}:{key_budget_config.budget_duration}"
cached_spend = await self.cache.async_get_cache(key=virtual_key_model_spend_cache_key)

If the cached spend exceeds the configured limit, the proxy rejects the request immediately without contacting the LLM provider.

Route and Model Validation

The is_virtual_key_allowed_to_call_route method in litellm/proxy/auth/route_checks.py validates that the requested endpoint appears in the key's allowed routes list. Additionally, model aliases defined in the virtual key's metadata are resolved in litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py, mapping public model names to provider-specific identifiers.

Generating and Managing Virtual Keys

Virtual keys are created via the /key/generate management endpoint defined in litellm/proxy/management_endpoints/key_management_endpoints.py. This endpoint stores the key in the LiteLLM_VerificationToken table with configurable metadata:

curl http://localhost:4000/key/generate \
  -H "Authorization: Bearer <master-key>" \
  -H "Content-Type: application/json" \
  -d '{"models": ["gpt-3.5-turbo"], "metadata": {"user":"alice@example.com"}}'

Key rotation, updates, and deletion are handled by key_management_event_hooks.py, providing a complete lifecycle management API accessible via HTTP or the admin UI.

Implementation Example

The following Python example demonstrates authenticating with a virtual key against the LiteLLM proxy:

import httpx

VIRTUAL_KEY = "sk-abc123virtual"
BASE_URL = "http://localhost:4000/v1/chat/completions"

payload = {
    "model": "gpt-3.5-turbo",
    "messages": [{"role": "user", "content": "Explain virtual keys."}],
}

headers = {
    "x-litellm-api-key": VIRTUAL_KEY,
    "Content-Type": "application/json",
}

response = httpx.post(BASE_URL, json=payload, headers=headers)
print(response.json())

The proxy processes this request by extracting the key via get_litellm_virtual_key, resolving it to a database row, enforcing budget and route checks, and finally forwarding to the configured provider with the appropriate credentials.

Summary

  • Virtual keys act as proxy-side bearer tokens that decouple client authentication from LLM provider credentials.
  • Header extraction occurs in common_utils.py, supporting both x-litellm-api-key and standard Authorization headers.
  • JWT integration allows dynamic key resolution from identity provider claims via user_api_key_auth.py.
  • Budget enforcement uses cached spend tracking with keys prefixed by virtual_key_spend in model_max_budget_limiter.py.
  • Management APIs in key_management_endpoints.py provide programmatic CRUD operations for key lifecycle management.

Frequently Asked Questions

How do virtual keys differ from master API keys in LiteLLM?

Master API keys provide administrative access to the proxy management endpoints, while virtual keys are scoped credentials for end-user LLM requests. Virtual keys support spend tracking, budget limits, and model restrictions that master keys do not enforce, making them suitable for multi-tenant environments where you need to segment usage by team or application.

Can virtual keys be rotated without downtime?

Yes, virtual keys support rotation via the /key/update endpoint or the UI. The proxy resolves keys on every request, so updating a key's hash or metadata in the LiteLLM_VerificationToken table takes effect immediately. You can generate new keys, update existing ones, or expire old keys without restarting the proxy service.

What happens when a virtual key exceeds its budget?

When a request would exceed the configured soft or hard budget, the proxy returns a 429 or 403 error before forwarding the request to the LLM provider. The enforcement logic in model_max_budget_limiter.py checks the cached spend against the key's max_budget field, preventing additional charges while preserving the key's configuration for future use (unless explicitly deleted).

Are virtual keys compatible with SSO providers?

Yes, virtual keys integrate with SSO via JWT claim mapping. The _resolve_jwt_to_virtual_key function extracts a claim value (such as email or sub) from the JWT and maps it to a stored virtual key in the database. This allows organizations to manage access through their identity provider while LiteLLM handles the LLM-specific authentication and spend tracking.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →