Production Security Best Practices for LiteLLM Proxy: 10 Critical Measures
Secure a LiteLLM proxy in production by storing the master key in environment variables, encrypting provider secrets with LITELLM_SALT_KEY, disabling dotenv loading via LITELLM_MODE=PRODUCTION, and running containers with read-only root filesystems.
The LiteLLM proxy, maintained in the BerriAI/litellm repository, acts as a unified gateway for multiple LLM providers. Running it in production requires hardening beyond default development settings to protect API keys, prevent container escapes, and ensure high availability during infrastructure failures.
Protect the Master Key
The master key authorizes every request to the proxy. If exposed, an attacker can impersonate any user and exhaust budgets.
Store the key in general_settings.master_key within your config.yaml, but load the actual value from the LITELLM_MASTER_KEY environment variable at startup. In litellm/proxy/proxy_server.py, the server initializes the JWTHandler and authentication layer using this value.
Never commit the master key to source control. Instead, inject it via your secret manager or orchestration platform.
Encrypt Secrets at Rest
Set the LITELLM_SALT_KEY environment variable to enable AES-256 encryption for all provider API keys stored in the database.
The encryption helpers reside in litellm/proxy/common_utils/encrypt_decrypt_utils.py. This salt is used to encrypt credentials when models are added to the database via management endpoints. Rotating this salt without preserving the previous value will corrupt existing encrypted keys, requiring re-registration of all provider credentials.
Harden the Runtime Environment
Disable Automatic Dotenv Loading
Set LITELLM_MODE=PRODUCTION to prevent the proxy from automatically loading .env files. In development, proxy_server.py calls load_dotenv() to simplify local testing. In production, this behavior risks leaking secrets if .env files are accidentally mounted into containers.
Run as Non-Root with Read-Only Filesystem
Configure your container with runAsNonRoot: true and runAsUser: 101 in the pod spec. Set securityContext.readOnlyRootFilesystem: true to prevent attackers from modifying binaries if they gain container access.
To support read-only roots, the Helm chart mounts writable volumes for database migrations and UI assets:
securityContext:
readOnlyRootFilesystem: true
runAsNonRoot: true
runAsUser: 101
capabilities:
drop: ["ALL"]
volumeMounts:
- name: migrations
mountPath: /app/migrations
- name: ui
mountPath: /app/var/litellm/ui
Isolate Health Check Processes
Enable SEPARATE_HEALTH_APP=1 to run a lightweight FastAPI instance on a dedicated port (default 8001) that handles only /health routes. This isolation prevents liveness and readiness probes from failing when the main worker is saturated with LLM requests.
Set SEPARATE_HEALTH_PORT to customize the port, and configure SUPERVISORD_STOPWAITSECS="3600" to allow graceful shutdowns during deployments.
Handle Database Outages Gracefully
For VPC-only deployments where the database may become temporarily unreachable, set general_settings.allow_requests_on_db_unavailable: true in config.yaml. This setting, referenced in the production documentation, allows the proxy to continue serving cached requests while logging failures, preventing total service outage during transient network partitions.
Optimize Redis Configuration
Prefer explicit host, port, and password parameters over the redis_url connection string. According to performance benchmarks in the production guide, the redis_url path incurs an ~80 RPS penalty compared to discrete parameters.
Configure via litellm_settings.cache_params:
litellm_settings:
cache: true
cache_params:
type: redis
host: ${REDIS_HOST}
port: ${REDIS_PORT}
password: ${REDIS_PASSWORD}
This approach also provides clearer ACL control and easier secret rotation through environment variables.
Limit Worker Processes
Prevent CPU over-commitment by setting worker count equal to available CPU cores. Use the Docker CMD that dynamically sets --num_workers $(nproc) to match the host's core count. This reduces the attack surface by minimizing forked processes while maximizing throughput.
Rotate Keys via Management Endpoints
The proxy supports runtime rotation of master keys without downtime. Use the endpoint defined in litellm/proxy/management_endpoints/key_management_endpoints.py:
curl -X POST "https://my-litellm-proxy.com/v1/key/rotate" \
-H "Authorization: Bearer sk-old-master-key" \
-d '{"new_master_key":"sk-new-rotated-key"}'
This updates both the in-memory master_key and the database entry, allowing immediate revocation of compromised credentials.
Production Configuration Example
A minimal secure config.yaml combines these settings:
model_list:
- model_name: openai
litellm_params:
model: openai/gpt-3.5-turbo
api_key: ${OPENAI_API_KEY}
general_settings:
master_key: ${LITELLM_MASTER_KEY}
alerting: ["slack"]
proxy_batch_write_at: 60
allow_requests_on_db_unavailable: true
litellm_settings:
request_timeout: 600
set_verbose: false
json_logs: true
cache: true
cache_params:
type: redis
host: ${REDIS_HOST}
port: ${REDIS_PORT}
password: ${REDIS_PASSWORD}
Summary
- Store
LITELLM_MASTER_KEYin environment variables, never in source control, to prevent unauthorized access to the proxy management API. - Encrypt provider keys using
LITELLM_SALT_KEYvia the utilities inencrypt_decrypt_utils.py. - Run containers with
readOnlyRootFilesystem: trueandrunAsNonRoot: trueto minimize privilege escalation risks. - Disable dotenv loading by setting
LITELLM_MODE=PRODUCTIONto avoid accidental secret exposure. - Enable
SEPARATE_HEALTH_APPto isolate health checks from main traffic, ensuring reliable Kubernetes probes. - Allow requests during DB outages with
allow_requests_on_db_unavailable: truefor resilient VPC deployments. - Rotate keys via the management endpoint in
key_management_endpoints.pywithout service restarts.
Frequently Asked Questions
What happens if I don't set LITELLM_SALT_KEY?
Without LITELLM_SALT_KEY, the proxy cannot encrypt provider API keys stored in the database, storing them in plaintext instead. This violates security best practices and exposes credentials if the database is compromised. The encryption logic in litellm/proxy/common_utils/encrypt_decrypt_utils.py requires this salt to perform AES-256 encryption.
How do I handle database migrations with a read-only root filesystem?
Mount a writable volume at /app/migrations and set LITELLM_MIGRATION_DIR to this path. The Helm chart init-container copies UI assets to /app/var/litellm/ui and runs prisma migrate deploy against the writable migration directory before the main container starts.
What is the performance impact of using redis_url versus discrete parameters?
Using redis_url instead of explicit host, port, and password parameters reduces throughput by approximately 80 requests per second. The discrete parameters also offer better security through clearer ACL controls and easier rotation via environment variables.
How does SEPARATE_HEALTH_APP improve reliability?
When SEPARATE_HEALTH_APP=1 is set, the proxy launches a second FastAPI instance on SEPARATE_HEALTH_PORT that handles only health check endpoints. This prevents Kubernetes liveness probes from failing when the main application is under heavy load processing LLM requests, avoiding unnecessary pod restarts during traffic spikes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →