# How LiteLLM Virtual Keys Work for API Key Management: A Complete Technical Guide

> Learn how LiteLLM virtual keys manage API keys, authenticate requests, enforce limits, and track usage for LLM providers. A complete technical guide.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: how-to-guide
- Published: 2026-03-26

---

**LiteLLM virtual keys are proxy-side credentials that authenticate requests, enforce spend limits, and track per-key usage before forwarding traffic to underlying LLM providers.**

LiteLLM virtual keys provide a secure abstraction layer for API key management within the BerriAI/litellm proxy. These short-lived tokens decouple client authentication from provider credentials, enabling fine-grained access control, budget enforcement, and detailed spend tracking without exposing your master API keys.

## What Are LiteLLM Virtual Keys

LiteLLM virtual keys are database-backed tokens stored in the `LiteLLM_VerificationToken` table that act as bearer credentials for every request reaching the proxy. Unlike static provider API keys, virtual keys support granular policy enforcement including model restrictions, route limitations, and team-based access controls.

### Authentication and Header Extraction

Every request must present a virtual key via the `x-litellm-api-key` header or the standard `Authorization` header. The proxy extracts this value in [`litellm/proxy/pass_through_endpoints/common_utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/pass_through_endpoints/common_utils.py) using the `get_litellm_virtual_key` function:

```python

# litellm/proxy/pass_through_endpoints/common_utils.py

def get_litellm_virtual_key(request: Request) -> str:
    """
    Extract and format API key from request headers.
    Prioritizes x-litellm-api-key over Authorization header.
    """
    litellm_api_key = request.headers.get("x-litellm-api-key")
    if litellm_api_key:
        return f"Bearer {litellm_api_key}"
    return request.headers.get("Authorization", "")

```

This extraction logic prioritizes the custom header, formatting it as a standard Bearer token for downstream processing.

### Spend Tracking and Budget Enforcement

Each virtual key maintains aggregated spend metrics in the proxy's cache using the `VIRTUAL_KEY_SPEND_CACHE_KEY_PREFIX` constant. The system enforces both soft-budget and hard-budget limits defined in [`litellm/proxy/hooks/model_max_budget_limiter.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/model_max_budget_limiter.py) before forwarding requests to providers.

## Proxy-Side Resolution and Validation

### JWT-to-Virtual-Key Mapping

For organizations using identity providers, virtual keys can be dynamically resolved from JWT claims. The `_resolve_jwt_to_virtual_key` function in [`litellm/proxy/auth/user_api_key_auth.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/user_api_key_auth.py) handles this mapping:

```python

# litellm/proxy/auth/user_api_key_auth.py

async def _resolve_jwt_to_virtual_key(
    jwt_claims: dict,
    jwt_handler: JWTHandler,
    prisma_client: Optional[PrismaClient],
    user_api_key_cache: DualCache,
    parent_otel_span: Optional[Span],
    proxy_logging_obj: ProxyLogging,
) -> Optional[UserAPIKeyAuth]:
    virtual_key_claim_field = jwt_handler.litellm_jwtauth.virtual_key_claim_field
    if virtual_key_claim_field is None:
        return None

    claim_value = get_nested_value(data=jwt_claims, key_path=virtual_key_claim_field, default=None)
    if claim_value is None:
        return None

    cache_key = f"jwt_key_mapping:{virtual_key_claim_field}:{claim_value}"
    cached_mapping = await user_api_key_cache.async_get_cache(cache_key)
    # ... resolves to LiteLLM_VerificationToken row

```

When a JWT claim maps to a virtual key, the proxy loads the complete token object including budget configurations, model whitelists, and user or team bindings.

## Runtime Policy Enforcement

### Budget Check Implementation

Before proxying a request, the system validates budget constraints using cache keys formatted as `virtual_key_spend:{user_api_key_hash}:{model}:{budget_duration}`:

```python

# litellm/proxy/hooks/model_max_budget_limiter.py

VIRTUAL_KEY_SPEND_CACHE_KEY_PREFIX = "virtual_key_spend"

# Example: fetch spend for a key+model

virtual_key_model_spend_cache_key = f"{VIRTUAL_KEY_SPEND_CACHE_KEY_PREFIX}:{user_api_key_hash}:{model}:{key_budget_config.budget_duration}"
cached_spend = await self.cache.async_get_cache(key=virtual_key_model_spend_cache_key)

```

If the cached spend exceeds the configured limit, the proxy rejects the request immediately without contacting the LLM provider.

### Route and Model Validation

The `is_virtual_key_allowed_to_call_route` method in [`litellm/proxy/auth/route_checks.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/route_checks.py) validates that the requested endpoint appears in the key's allowed routes list. Additionally, model aliases defined in the virtual key's metadata are resolved in [`litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py), mapping public model names to provider-specific identifiers.

## Generating and Managing Virtual Keys

Virtual keys are created via the `/key/generate` management endpoint defined in [`litellm/proxy/management_endpoints/key_management_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/management_endpoints/key_management_endpoints.py). This endpoint stores the key in the `LiteLLM_VerificationToken` table with configurable metadata:

```bash
curl http://localhost:4000/key/generate \
  -H "Authorization: Bearer <master-key>" \
  -H "Content-Type: application/json" \
  -d '{"models": ["gpt-3.5-turbo"], "metadata": {"user":"alice@example.com"}}'

```

Key rotation, updates, and deletion are handled by [`key_management_event_hooks.py`](https://github.com/BerriAI/litellm/blob/main/key_management_event_hooks.py), providing a complete lifecycle management API accessible via HTTP or the admin UI.

## Implementation Example

The following Python example demonstrates authenticating with a virtual key against the LiteLLM proxy:

```python
import httpx

VIRTUAL_KEY = "sk-abc123virtual"
BASE_URL = "http://localhost:4000/v1/chat/completions"

payload = {
    "model": "gpt-3.5-turbo",
    "messages": [{"role": "user", "content": "Explain virtual keys."}],
}

headers = {
    "x-litellm-api-key": VIRTUAL_KEY,
    "Content-Type": "application/json",
}

response = httpx.post(BASE_URL, json=payload, headers=headers)
print(response.json())

```

The proxy processes this request by extracting the key via `get_litellm_virtual_key`, resolving it to a database row, enforcing budget and route checks, and finally forwarding to the configured provider with the appropriate credentials.

## Summary

- **Virtual keys** act as proxy-side bearer tokens that decouple client authentication from LLM provider credentials.
- **Header extraction** occurs in [`common_utils.py`](https://github.com/BerriAI/litellm/blob/main/common_utils.py), supporting both `x-litellm-api-key` and standard `Authorization` headers.
- **JWT integration** allows dynamic key resolution from identity provider claims via [`user_api_key_auth.py`](https://github.com/BerriAI/litellm/blob/main/user_api_key_auth.py).
- **Budget enforcement** uses cached spend tracking with keys prefixed by `virtual_key_spend` in [`model_max_budget_limiter.py`](https://github.com/BerriAI/litellm/blob/main/model_max_budget_limiter.py).
- **Management APIs** in [`key_management_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/key_management_endpoints.py) provide programmatic CRUD operations for key lifecycle management.

## Frequently Asked Questions

### How do virtual keys differ from master API keys in LiteLLM?

**Master API keys** provide administrative access to the proxy management endpoints, while **virtual keys** are scoped credentials for end-user LLM requests. Virtual keys support spend tracking, budget limits, and model restrictions that master keys do not enforce, making them suitable for multi-tenant environments where you need to segment usage by team or application.

### Can virtual keys be rotated without downtime?

Yes, virtual keys support rotation via the `/key/update` endpoint or the UI. The proxy resolves keys on every request, so updating a key's hash or metadata in the `LiteLLM_VerificationToken` table takes effect immediately. You can generate new keys, update existing ones, or expire old keys without restarting the proxy service.

### What happens when a virtual key exceeds its budget?

When a request would exceed the configured soft or hard budget, the proxy returns a 429 or 403 error before forwarding the request to the LLM provider. The enforcement logic in [`model_max_budget_limiter.py`](https://github.com/BerriAI/litellm/blob/main/model_max_budget_limiter.py) checks the cached spend against the key's `max_budget` field, preventing additional charges while preserving the key's configuration for future use (unless explicitly deleted).

### Are virtual keys compatible with SSO providers?

Yes, virtual keys integrate with SSO via JWT claim mapping. The `_resolve_jwt_to_virtual_key` function extracts a claim value (such as `email` or `sub`) from the JWT and maps it to a stored virtual key in the database. This allows organizations to manage access through their identity provider while LiteLLM handles the LLM-specific authentication and spend tracking.