# LiteLLM Proxy Health Check and Monitoring Endpoints: A Complete Implementation Guide

> Implement LiteLLM proxy health check and monitoring endpoints. Validate proxy availability, LLM deployments, and integrations like Datadog and Slack. Get the complete guide.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: how-to-guide
- Published: 2026-03-26

---

**LiteLLM exposes Kubernetes-compatible HTTP endpoints at `/health/liveness`, `/health/readiness`, and `/health/services` to validate proxy availability, downstream LLM deployments, and third-party integrations such as Datadog, SQS, and Slack.**

The LiteLLM proxy server provides a comprehensive suite of health-check and monitoring endpoints that enable operators to validate system health and diagnose connectivity issues in production environments. These endpoints are implemented across two core modules and support both simple process liveness checks and complex readiness validations with per-model diagnostics.

## Architecture Overview

The health-check implementation spans two primary files in the BerriAI/litellm repository. The [`litellm/proxy/health_endpoints/_health_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_endpoints/_health_endpoints.py) file registers the FastAPI routes and handles HTTP request orchestration, while [`litellm/proxy/health_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_check.py) contains the async execution engine that performs the actual health validations with concurrency controls.

### FastAPI Router in [`_health_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/_health_endpoints.py)

The router is instantiated at module import time as `router = APIRouter()` and mounted by the main proxy server. Each endpoint is decorated with `@router.get(...)` and protected by the `user_api_key_auth` dependency, ensuring only authorized users can invoke diagnostics.

Key routes include:

- `/health/liveness` and `/health/liveliness` – Process liveness probes
- `/health/readiness` – Startup completion and model health validation
- `/health/services` – Third-party integration diagnostics
- `/health/backlog` – In-flight request queue metrics
- `/health/test_connection` – Model configuration load testing

### Health Check Engine in [`health_check.py`](https://github.com/BerriAI/litellm/blob/main/health_check.py)

This module contains the `perform_health_check` function, which aggregates model health data by calling `litellm.ahealth_check` for each configured model. The engine uses `_run_health_checks_with_bounded_concurrency` to enforce a `max_concurrency` limit, preventing the proxy from overwhelming downstream LLM providers during health validations.

## Available Endpoints and Functionality

### Kubernetes-Compatible Probes

**Liveness probes** confirm the proxy process is running. The `/health/liveness` endpoint (and its alias `/health/liveliness`) return `{"status":"healthy"}` with HTTP 200 if the process is up, making them ideal for Kubernetes liveness probes.

**Readiness probes** validate that startup tasks—such as database migrations and cache warm-up—have completed successfully. The `/health/readiness` endpoint accepts an optional `details=true` query parameter that triggers model-level health checks via the underlying `perform_health_check` function.

```bash
curl -s -G \
  -H "Authorization: Bearer $LITELLMTOKEN" \
  --data-urlencode "details=true" \
  http://localhost:4000/health/readiness

```

When `details=true`, the response includes `healthy_count`, `unhealthy_count`, per-model latency metrics, and `peak_in_flight_requests`.

### Service-Specific Diagnostics

The `/health/services` endpoint validates connectivity to observability backends. It accepts a `service` query parameter and executes the corresponding `async_health_check` method:

- **Datadog**: Calls `DataDogLogger.async_health_check()`
- **SQS**: Invokes `SQSLogger.async_health_check()`
- **Slack**: Performs a mock LLM request using `litellm-mock-response-model` to verify webhook connectivity

```bash
curl -s -H "Authorization: Bearer $LITELLMTOKEN" \
  "http://localhost:4000/health/services?service=datadog"

```

### Monitoring and History

For operational visibility, LiteLLM provides:

- `/health/history` – Returns a rolling list of recent health-check outcomes for UI dashboards
- `/health/latest` – Provides the most recent health-check payload as a quick status snapshot
- `/health/backlog` – Exposes the current length of the in-flight requests queue, useful for autoscaling decisions

## Implementation Details

### Concurrency-Bounded Execution

The `_run_health_checks_with_bounded_concurrency` function in [`litellm/proxy/health_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_check.py) creates up to `max_concurrency` asyncio tasks while preserving result ordering. This caps the number of simultaneous LLM calls to protect downstream providers.

```python
async def _run_health_checks_with_bounded_concurrency(
    models: list, concurrency_limit: int
) -> tuple[list, int]:
    # Schedules tasks while respecting concurrency_limit

    # Returns results and peak_in_flight metrics

```

### Timeout Handling

Each model-level health request is wrapped by `run_with_timeout`, which uses `asyncio.wait_for` to enforce `HEALTH_CHECK_TIMEOUT_SECONDS` (defaulting to 15 seconds). If a provider fails to respond within this window, the function returns `{"error":"Timeout exceeded"}` without cancelling sibling checks.

```python
async def run_with_timeout(task, timeout):
    try:
        return await asyncio.wait_for(task, timeout)
    except asyncio.TimeoutError:
        return {"error": "Timeout exceeded"}

```

### Credential Sanitization

Before returning health data to clients, the `_clean_endpoint_data` function removes sensitive fields such as `api_key`, `aws_secret_access_key`, and `litellm_logging_obj` from the response payload.

```python
def _clean_endpoint_data(endpoint_data: dict, details: Optional[bool] = True):
    endpoint_data.pop("litellm_logging_obj", None)
    return {k: v for k, v in endpoint_data.items() if k not in ILLEGAL_DISPLAY_PARAMS}

```

## Practical Usage Examples

### Basic Liveness Check

Verify the proxy is running with a simple curl command:

```bash
curl -s -H "Authorization: Bearer $LITELLMTOKEN" \
  http://localhost:4000/health/liveness

# → {"status":"healthy"}

```

### Testing Datadog Integration

Validate your Datadog configuration:

```bash
curl -s -H "Authorization: Bearer $LITELLMTOKEN" \
  "http://localhost:4000/health/services?service=datadog"

```

Typical response:

```json
{
  "status": "healthy",
  "message": "Datadog is healthy"
}

```

### Python Client Implementation

Programmatically monitor the proxy using httpx:

```python
import httpx

token = "sk-proxy-xxxx"
headers = {"Authorization": f"Bearer {token}"}
base = "http://localhost:4000"

def health(path: str, params: dict | None = None):
    url = f"{base}{path}"
    resp = httpx.get(url, headers=headers, params=params, timeout=20)
    resp.raise_for_status()
    return resp.json()

# Check proxy liveness

print(health("/health/liveness"))

# Check readiness with model details

print(health("/health/readiness", {"details": "true"}))

# Verify Datadog connectivity

print(health("/health/services", {"service": "datadog"}))

```

### Kubernetes Probe Configuration

Configure your deployment to use LiteLLM's health endpoints for pod lifecycle management:

```yaml
livenessProbe:
  httpGet:
    path: /health/liveness
    port: 4000
  initialDelaySeconds: 5
  periodSeconds: 10

readinessProbe:
  httpGet:
    path: /health/readiness?details=true
    port: 4000
  initialDelaySeconds: 10
  periodSeconds: 30

```

The readiness probe will mark the pod as *NotReady* if model health checks fail, automatically removing it from service until downstream LLM endpoints recover.

## Summary

- LiteLLM provides **Kubernetes-ready health endpoints** at `/health/liveness` and `/health/readiness` for production deployments.
- The async engine in [`litellm/proxy/health_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_check.py) enforces **concurrency limits** via `_run_health_checks_with_bounded_concurrency` and **15-second timeouts** via `run_with_timeout`.
- **Service-specific checks** for Datadog, SQS, and Slack are available through `/health/services`, each implementing the `async_health_check` contract.
- **Sensitive data protection** is guaranteed by `_clean_endpoint_data`, which strips API keys and secrets from health responses before transmission.

## Frequently Asked Questions

### What is the difference between `/health/liveness` and `/health/readiness` in LiteLLM?

Liveness checks confirm the proxy process is running and responsive, returning HTTP 200 if the service is alive. Readiness checks validate that startup tasks—such as database migrations and cache initialization—have completed, and can optionally include model-level health diagnostics when the `details=true` parameter is provided.

### How does LiteLLM prevent health checks from overwhelming downstream LLM providers?

The implementation uses `_run_health_checks_with_bounded_concurrency` in [`litellm/proxy/health_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_check.py) to limit simultaneous health check requests to a configurable `max_concurrency` value. Additionally, `run_with_timeout` wraps each model call with `asyncio.wait_for`, enforcing a default timeout of `HEALTH_CHECK_TIMEOUT_SECONDS` (15 seconds) to prevent hanging connections.

### Can I verify the health of third-party integrations like Datadog or Slack?

Yes, the `/health/services` endpoint accepts a `service` query parameter to test specific integrations. For Datadog, it calls `DataDogLogger.async_health_check()`; for SQS, it invokes `SQSLogger.async_health_check()`; and for Slack, it performs a mock completion request to verify webhook connectivity.

### Are API keys exposed in LiteLLM health check responses?

No, the `_clean_endpoint_data` function automatically filters out sensitive parameters—including `api_key`, `aws_secret_access_key`, and `litellm_logging_obj`—before returning health data to clients, ensuring credentials remain secure even during diagnostic operations.