How the Health Check Endpoint Functions in PrivateGPT for System Monitoring

The PrivateGPT health check endpoint exposes a lightweight, unauthenticated GET route at /health that returns a static JSON payload {"status": "ok"} to confirm the API server is responsive.

The zylon-ai/private-gpt repository implements a dedicated health check endpoint to support external monitoring systems and load balancers. This endpoint provides a zero-dependency mechanism to verify that the FastAPI application is running without performing expensive operations. Understanding how this health check endpoint functions is essential for deploying PrivateGPT in production environments with automated monitoring.

Router Definition and File Structure

The health check logic resides in private_gpt/server/health/health_router.py, where a dedicated APIRouter instance handles all health-related routes.

The APIRouter Setup

The router is instantiated without authentication requirements, making it accessible to unauthenticated probes from load balancers and container orchestrators.

health_router = APIRouter()

The HealthResponse Pydantic Model

The endpoint returns a strictly typed HealthResponse model that guarantees a consistent JSON structure. This model uses a Literal type to constrain the status field to only accept "ok" as a valid value.

class HealthResponse(BaseModel):
    status: Literal["ok"] = Field(default="ok")

Endpoint Implementation

The actual handler is registered as a GET operation at the /health path with a "Health" tag for OpenAPI documentation organization.

@health_router.get("/health", tags=["Health"])
def health() -> HealthResponse:
    """Return ok if the system is up."""
    return HealthResponse(status="ok")

When invoked, the function immediately instantiates and returns a HealthResponse object, resulting in an HTTP 200 OK response with the body {"status":"ok"}. The implementation performs no database queries, external API calls, or computational heavy lifting, ensuring sub-millisecond response times.

Integration with the Main Application

The health router is attached to the central FastAPI application inside private_gpt/launcher.py. This file imports the router and includes it in the application stack alongside other feature routers such as chat, completions, and ingest.

from private_gpt.server.health.health_router import health_router

# ...

app.include_router(health_router)

Because the router is added last in the initialization sequence, the /health endpoint becomes available immediately after the application starts, providing an early signal that the server is ready to accept connections.

Production Monitoring Examples

The health check endpoint's zero-dependency design makes it ideal for high-frequency polling scenarios. Below are practical implementations for common monitoring tools.

Command-Line Health Verification

Use curl to quickly verify service availability during deployment scripts or local development:

curl -s http://localhost:8000/health | jq .

# Output: {"status":"ok"}

Automated Python Monitoring

For custom monitoring scripts, use an HTTP client like httpx to validate the response payload:

import httpx

def check_health(base_url: str = "http://localhost:8000") -> bool:
    try:
        resp = httpx.get(f"{base_url}/health", timeout=2.0)
        resp.raise_for_status()
        return resp.json().get("status") == "ok"
    except Exception:
        return False

if __name__ == "__main__":
    print("Service healthy:", check_health())

Kubernetes Liveness and Readiness Probes

Configure container orchestration platforms to use the endpoint for automated health management:

livenessProbe:
  httpGet:
    path: /health
    port: 8000
  initialDelaySeconds: 5
  periodSeconds: 10

Summary

  • The health check endpoint is defined in private_gpt/server/health/health_router.py and mounted via private_gpt/launcher.py.
  • It returns a static JSON payload {"status": "ok"} with HTTP status 200 OK.
  • The endpoint requires no authentication, making it safe for external load balancers and monitoring tools.
  • Zero dependencies ensure the check reflects only the API process liveliness, not external service states.
  • Response times are sub-millisecond because the handler performs no I/O or computation.

Frequently Asked Questions

Does the health check endpoint require authentication?

No. The router is explicitly configured without authentication requirements, allowing unauthenticated probes from load balancers, Kubernetes, and external uptime monitors to verify system availability without API keys or tokens.

What HTTP status code does the health endpoint return?

The endpoint returns HTTP 200 OK when the PrivateGPT server is running. If the server is down or unreachable, the connection will fail entirely; the endpoint does not return error status codes like 503 because it represents a simple liveness check rather than a deep system health validation.

Can I use the health check for Kubernetes liveness probes?

Yes. The lightweight, fast-responding endpoint at /health is ideal for Kubernetes liveness and readiness probes. Configure your pod spec to poll http://<pod-ip>:8000/health with appropriate initialDelaySeconds and periodSeconds values to automate container restart decisions.

Does the health check verify that the LLM or vector database is connected?

No. According to the source code in private_gpt/server/health/health_router.py, the endpoint returns a static response without checking external dependencies such as the LLM backend, vector store, or file ingestion services. It verifies only that the FastAPI process is alive and responsive.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →