LiteLLM Proxy Health Check and Monitoring Endpoints: A Complete Implementation Guide

LiteLLM exposes Kubernetes-compatible HTTP endpoints at /health/liveness, /health/readiness, and /health/services to validate proxy availability, downstream LLM deployments, and third-party integrations such as Datadog, SQS, and Slack.

The LiteLLM proxy server provides a comprehensive suite of health-check and monitoring endpoints that enable operators to validate system health and diagnose connectivity issues in production environments. These endpoints are implemented across two core modules and support both simple process liveness checks and complex readiness validations with per-model diagnostics.

Architecture Overview

The health-check implementation spans two primary files in the BerriAI/litellm repository. The litellm/proxy/health_endpoints/_health_endpoints.py file registers the FastAPI routes and handles HTTP request orchestration, while litellm/proxy/health_check.py contains the async execution engine that performs the actual health validations with concurrency controls.

FastAPI Router in _health_endpoints.py

The router is instantiated at module import time as router = APIRouter() and mounted by the main proxy server. Each endpoint is decorated with @router.get(...) and protected by the user_api_key_auth dependency, ensuring only authorized users can invoke diagnostics.

Key routes include:

  • /health/liveness and /health/liveliness – Process liveness probes
  • /health/readiness – Startup completion and model health validation
  • /health/services – Third-party integration diagnostics
  • /health/backlog – In-flight request queue metrics
  • /health/test_connection – Model configuration load testing

Health Check Engine in health_check.py

This module contains the perform_health_check function, which aggregates model health data by calling litellm.ahealth_check for each configured model. The engine uses _run_health_checks_with_bounded_concurrency to enforce a max_concurrency limit, preventing the proxy from overwhelming downstream LLM providers during health validations.

Available Endpoints and Functionality

Kubernetes-Compatible Probes

Liveness probes confirm the proxy process is running. The /health/liveness endpoint (and its alias /health/liveliness) return {"status":"healthy"} with HTTP 200 if the process is up, making them ideal for Kubernetes liveness probes.

Readiness probes validate that startup tasks—such as database migrations and cache warm-up—have completed successfully. The /health/readiness endpoint accepts an optional details=true query parameter that triggers model-level health checks via the underlying perform_health_check function.

curl -s -G \
  -H "Authorization: Bearer $LITELLMTOKEN" \
  --data-urlencode "details=true" \
  http://localhost:4000/health/readiness

When details=true, the response includes healthy_count, unhealthy_count, per-model latency metrics, and peak_in_flight_requests.

Service-Specific Diagnostics

The /health/services endpoint validates connectivity to observability backends. It accepts a service query parameter and executes the corresponding async_health_check method:

  • Datadog: Calls DataDogLogger.async_health_check()
  • SQS: Invokes SQSLogger.async_health_check()
  • Slack: Performs a mock LLM request using litellm-mock-response-model to verify webhook connectivity
curl -s -H "Authorization: Bearer $LITELLMTOKEN" \
  "http://localhost:4000/health/services?service=datadog"

Monitoring and History

For operational visibility, LiteLLM provides:

  • /health/history – Returns a rolling list of recent health-check outcomes for UI dashboards
  • /health/latest – Provides the most recent health-check payload as a quick status snapshot
  • /health/backlog – Exposes the current length of the in-flight requests queue, useful for autoscaling decisions

Implementation Details

Concurrency-Bounded Execution

The _run_health_checks_with_bounded_concurrency function in litellm/proxy/health_check.py creates up to max_concurrency asyncio tasks while preserving result ordering. This caps the number of simultaneous LLM calls to protect downstream providers.

async def _run_health_checks_with_bounded_concurrency(
    models: list, concurrency_limit: int
) -> tuple[list, int]:
    # Schedules tasks while respecting concurrency_limit

    # Returns results and peak_in_flight metrics

Timeout Handling

Each model-level health request is wrapped by run_with_timeout, which uses asyncio.wait_for to enforce HEALTH_CHECK_TIMEOUT_SECONDS (defaulting to 15 seconds). If a provider fails to respond within this window, the function returns {"error":"Timeout exceeded"} without cancelling sibling checks.

async def run_with_timeout(task, timeout):
    try:
        return await asyncio.wait_for(task, timeout)
    except asyncio.TimeoutError:
        return {"error": "Timeout exceeded"}

Credential Sanitization

Before returning health data to clients, the _clean_endpoint_data function removes sensitive fields such as api_key, aws_secret_access_key, and litellm_logging_obj from the response payload.

def _clean_endpoint_data(endpoint_data: dict, details: Optional[bool] = True):
    endpoint_data.pop("litellm_logging_obj", None)
    return {k: v for k, v in endpoint_data.items() if k not in ILLEGAL_DISPLAY_PARAMS}

Practical Usage Examples

Basic Liveness Check

Verify the proxy is running with a simple curl command:

curl -s -H "Authorization: Bearer $LITELLMTOKEN" \
  http://localhost:4000/health/liveness

# → {"status":"healthy"}

Testing Datadog Integration

Validate your Datadog configuration:

curl -s -H "Authorization: Bearer $LITELLMTOKEN" \
  "http://localhost:4000/health/services?service=datadog"

Typical response:

{
  "status": "healthy",
  "message": "Datadog is healthy"
}

Python Client Implementation

Programmatically monitor the proxy using httpx:

import httpx

token = "sk-proxy-xxxx"
headers = {"Authorization": f"Bearer {token}"}
base = "http://localhost:4000"

def health(path: str, params: dict | None = None):
    url = f"{base}{path}"
    resp = httpx.get(url, headers=headers, params=params, timeout=20)
    resp.raise_for_status()
    return resp.json()

# Check proxy liveness

print(health("/health/liveness"))

# Check readiness with model details

print(health("/health/readiness", {"details": "true"}))

# Verify Datadog connectivity

print(health("/health/services", {"service": "datadog"}))

Kubernetes Probe Configuration

Configure your deployment to use LiteLLM's health endpoints for pod lifecycle management:

livenessProbe:
  httpGet:
    path: /health/liveness
    port: 4000
  initialDelaySeconds: 5
  periodSeconds: 10

readinessProbe:
  httpGet:
    path: /health/readiness?details=true
    port: 4000
  initialDelaySeconds: 10
  periodSeconds: 30

The readiness probe will mark the pod as NotReady if model health checks fail, automatically removing it from service until downstream LLM endpoints recover.

Summary

  • LiteLLM provides Kubernetes-ready health endpoints at /health/liveness and /health/readiness for production deployments.
  • The async engine in litellm/proxy/health_check.py enforces concurrency limits via _run_health_checks_with_bounded_concurrency and 15-second timeouts via run_with_timeout.
  • Service-specific checks for Datadog, SQS, and Slack are available through /health/services, each implementing the async_health_check contract.
  • Sensitive data protection is guaranteed by _clean_endpoint_data, which strips API keys and secrets from health responses before transmission.

Frequently Asked Questions

What is the difference between /health/liveness and /health/readiness in LiteLLM?

Liveness checks confirm the proxy process is running and responsive, returning HTTP 200 if the service is alive. Readiness checks validate that startup tasks—such as database migrations and cache initialization—have completed, and can optionally include model-level health diagnostics when the details=true parameter is provided.

How does LiteLLM prevent health checks from overwhelming downstream LLM providers?

The implementation uses _run_health_checks_with_bounded_concurrency in litellm/proxy/health_check.py to limit simultaneous health check requests to a configurable max_concurrency value. Additionally, run_with_timeout wraps each model call with asyncio.wait_for, enforcing a default timeout of HEALTH_CHECK_TIMEOUT_SECONDS (15 seconds) to prevent hanging connections.

Can I verify the health of third-party integrations like Datadog or Slack?

Yes, the /health/services endpoint accepts a service query parameter to test specific integrations. For Datadog, it calls DataDogLogger.async_health_check(); for SQS, it invokes SQSLogger.async_health_check(); and for Slack, it performs a mock completion request to verify webhook connectivity.

Are API keys exposed in LiteLLM health check responses?

No, the _clean_endpoint_data function automatically filters out sensitive parameters—including api_key, aws_secret_access_key, and litellm_logging_obj—before returning health data to clients, ensuring credentials remain secure even during diagnostic operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →