Production Security Best Practices for LiteLLM Proxy: 10 Critical Measures

Secure a LiteLLM proxy in production by storing the master key in environment variables, encrypting provider secrets with LITELLM_SALT_KEY, disabling dotenv loading via LITELLM_MODE=PRODUCTION, and running containers with read-only root filesystems.

The LiteLLM proxy, maintained in the BerriAI/litellm repository, acts as a unified gateway for multiple LLM providers. Running it in production requires hardening beyond default development settings to protect API keys, prevent container escapes, and ensure high availability during infrastructure failures.

Protect the Master Key

The master key authorizes every request to the proxy. If exposed, an attacker can impersonate any user and exhaust budgets.

Store the key in general_settings.master_key within your config.yaml, but load the actual value from the LITELLM_MASTER_KEY environment variable at startup. In litellm/proxy/proxy_server.py, the server initializes the JWTHandler and authentication layer using this value.

Never commit the master key to source control. Instead, inject it via your secret manager or orchestration platform.

Encrypt Secrets at Rest

Set the LITELLM_SALT_KEY environment variable to enable AES-256 encryption for all provider API keys stored in the database.

The encryption helpers reside in litellm/proxy/common_utils/encrypt_decrypt_utils.py. This salt is used to encrypt credentials when models are added to the database via management endpoints. Rotating this salt without preserving the previous value will corrupt existing encrypted keys, requiring re-registration of all provider credentials.

Harden the Runtime Environment

Disable Automatic Dotenv Loading

Set LITELLM_MODE=PRODUCTION to prevent the proxy from automatically loading .env files. In development, proxy_server.py calls load_dotenv() to simplify local testing. In production, this behavior risks leaking secrets if .env files are accidentally mounted into containers.

Run as Non-Root with Read-Only Filesystem

Configure your container with runAsNonRoot: true and runAsUser: 101 in the pod spec. Set securityContext.readOnlyRootFilesystem: true to prevent attackers from modifying binaries if they gain container access.

To support read-only roots, the Helm chart mounts writable volumes for database migrations and UI assets:

securityContext:
  readOnlyRootFilesystem: true
  runAsNonRoot: true
  runAsUser: 101
  capabilities:
    drop: ["ALL"]

volumeMounts:
  - name: migrations
    mountPath: /app/migrations
  - name: ui
    mountPath: /app/var/litellm/ui

Isolate Health Check Processes

Enable SEPARATE_HEALTH_APP=1 to run a lightweight FastAPI instance on a dedicated port (default 8001) that handles only /health routes. This isolation prevents liveness and readiness probes from failing when the main worker is saturated with LLM requests.

Set SEPARATE_HEALTH_PORT to customize the port, and configure SUPERVISORD_STOPWAITSECS="3600" to allow graceful shutdowns during deployments.

Handle Database Outages Gracefully

For VPC-only deployments where the database may become temporarily unreachable, set general_settings.allow_requests_on_db_unavailable: true in config.yaml. This setting, referenced in the production documentation, allows the proxy to continue serving cached requests while logging failures, preventing total service outage during transient network partitions.

Optimize Redis Configuration

Prefer explicit host, port, and password parameters over the redis_url connection string. According to performance benchmarks in the production guide, the redis_url path incurs an ~80 RPS penalty compared to discrete parameters.

Configure via litellm_settings.cache_params:

litellm_settings:
  cache: true
  cache_params:
    type: redis
    host: ${REDIS_HOST}
    port: ${REDIS_PORT}
    password: ${REDIS_PASSWORD}

This approach also provides clearer ACL control and easier secret rotation through environment variables.

Limit Worker Processes

Prevent CPU over-commitment by setting worker count equal to available CPU cores. Use the Docker CMD that dynamically sets --num_workers $(nproc) to match the host's core count. This reduces the attack surface by minimizing forked processes while maximizing throughput.

Rotate Keys via Management Endpoints

The proxy supports runtime rotation of master keys without downtime. Use the endpoint defined in litellm/proxy/management_endpoints/key_management_endpoints.py:

curl -X POST "https://my-litellm-proxy.com/v1/key/rotate" \
  -H "Authorization: Bearer sk-old-master-key" \
  -d '{"new_master_key":"sk-new-rotated-key"}'

This updates both the in-memory master_key and the database entry, allowing immediate revocation of compromised credentials.

Production Configuration Example

A minimal secure config.yaml combines these settings:

model_list:
  - model_name: openai
    litellm_params:
      model: openai/gpt-3.5-turbo
      api_key: ${OPENAI_API_KEY}

general_settings:
  master_key: ${LITELLM_MASTER_KEY}
  alerting: ["slack"]
  proxy_batch_write_at: 60
  allow_requests_on_db_unavailable: true

litellm_settings:
  request_timeout: 600
  set_verbose: false
  json_logs: true
  cache: true
  cache_params:
    type: redis
    host: ${REDIS_HOST}
    port: ${REDIS_PORT}
    password: ${REDIS_PASSWORD}

Summary

  • Store LITELLM_MASTER_KEY in environment variables, never in source control, to prevent unauthorized access to the proxy management API.
  • Encrypt provider keys using LITELLM_SALT_KEY via the utilities in encrypt_decrypt_utils.py.
  • Run containers with readOnlyRootFilesystem: true and runAsNonRoot: true to minimize privilege escalation risks.
  • Disable dotenv loading by setting LITELLM_MODE=PRODUCTION to avoid accidental secret exposure.
  • Enable SEPARATE_HEALTH_APP to isolate health checks from main traffic, ensuring reliable Kubernetes probes.
  • Allow requests during DB outages with allow_requests_on_db_unavailable: true for resilient VPC deployments.
  • Rotate keys via the management endpoint in key_management_endpoints.py without service restarts.

Frequently Asked Questions

What happens if I don't set LITELLM_SALT_KEY?

Without LITELLM_SALT_KEY, the proxy cannot encrypt provider API keys stored in the database, storing them in plaintext instead. This violates security best practices and exposes credentials if the database is compromised. The encryption logic in litellm/proxy/common_utils/encrypt_decrypt_utils.py requires this salt to perform AES-256 encryption.

How do I handle database migrations with a read-only root filesystem?

Mount a writable volume at /app/migrations and set LITELLM_MIGRATION_DIR to this path. The Helm chart init-container copies UI assets to /app/var/litellm/ui and runs prisma migrate deploy against the writable migration directory before the main container starts.

What is the performance impact of using redis_url versus discrete parameters?

Using redis_url instead of explicit host, port, and password parameters reduces throughput by approximately 80 requests per second. The discrete parameters also offer better security through clearer ACL controls and easier rotation via environment variables.

How does SEPARATE_HEALTH_APP improve reliability?

When SEPARATE_HEALTH_APP=1 is set, the proxy launches a second FastAPI instance on SEPARATE_HEALTH_PORT that handles only health check endpoints. This prevents Kubernetes liveness probes from failing when the main application is under heavy load processing LLM requests, avoiding unnecessary pod restarts during traffic spikes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →