# Kubernetes Liveness, Readiness, and Startup Probes: A Complete Troubleshooting Guide

> Master Kubernetes liveness, readiness, and startup probes. Learn how to effectively troubleshoot container issues and ensure application availability. Optimize your deployments today.

- Repository: [Anil Kumar/DevOps-Interview-Guide](https://github.com/litu54/DevOps-Interview-Guide)
- Tags: how-to-guide
- Published: 2026-08-10

---

**Kubernetes liveness probes restart containers experiencing deadlocks or hangs, readiness probes control traffic routing to healthy pods, and startup probes protect slow-initializing applications from premature health checks.**

Kubernetes probes are automated health-checking mechanisms that the kubelet uses to determine container viability and traffic eligibility. According to the `litu54/DevOps-Interview-Guide` repository—which contains interview questions from companies like Verizon, JPMorgan, and Sony—mastering these three probe types is essential for production reliability and DevOps technical interviews. This guide explains how each probe functions and provides systematic troubleshooting strategies based on real-world scenarios.

## What Are Kubernetes Liveness, Readiness, and Startup Probes?

Kubernetes implements three distinct health-checking mechanisms, each serving a specific operational purpose:

| Probe | Execution Timing | Health Check Purpose | Failure Action |
|-------|-----------------|---------------------|----------------|
| **Liveness** | Runs periodically after `initialDelaySeconds` | Determines if the application is *alive* and can continue running | Container is killed and restarted after `failureThreshold` |
| **Readiness** | Runs periodically after `initialDelaySeconds` | Determines if the application is *ready* to serve traffic | Pod removed from Service endpoints; container **not** restarted |
| **Startup** | Runs once (or until success) before other probes | Verifies completion of heavy initialization, migrations, or warm-up sequences | Container killed if it fails `failureThreshold`; on success, liveness/readiness probes begin |

**Liveness probes** detect unrecoverable states like deadlocks or memory leaks that require a container restart. **Readiness probes** ensure load balancers only route traffic to pods that have completed startup sequences and initialized dependencies. **Startup probes** are specifically designed for applications with lengthy initialization periods, preventing premature liveness or readiness failures during the warm-up phase.

## How Kubernetes Probes Work Internally

The kubelet invokes probe logic defined in the Pod spec using one of three mechanisms: an **HTTP GET** request, a **TCP socket** check, or an **exec** command. The kubelet tracks probe results and updates the pod's status accordingly:

- **Success**: Resets the failure counter; the container remains healthy.
- **Failure**: After `failureThreshold` consecutive failures, the kubelet marks the container unhealthy and takes probe-specific action.

For **liveness**, the kubelet kills and restarts the container. For **readiness**, the pod is removed from Service endpoints but continues running. For **startup**, failure results in container termination, while success transitions the pod to standard liveness and readiness monitoring.

## Troubleshooting Kubernetes Probe Failures

When applications experience instability or traffic routing issues, probe misconfiguration is often the root cause. Use the following systematic approach to diagnose and resolve problems:

### 1. Verify Probe Definitions

Inspect the Pod spec to confirm the probe type (HTTP/TCP/exec) matches your application's health endpoints. Check timing parameters carefully: `initialDelaySeconds`, `periodSeconds`, `timeoutSeconds`, `successThreshold`, and `failureThreshold`. A common error documented in [`SquareOps/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/SquareOps/DevOps_Engineer.md) involves setting `initialDelaySeconds` too low for applications requiring several seconds to start, causing unnecessary restart loops.

### 2. Inspect Pod Logs and Events

Examine application output for health-check failures or startup errors:

```bash
kubectl logs <pod-name> -c <container-name>
kubectl describe pod <pod-name>

```

Look specifically for events labeled `Liveness probe failed`, `Readiness probe failed`, or `Startup probe failed` in the `kubectl describe` output.

### 3. Manually Test Probe Endpoints

Validate accessibility by executing test commands from within the cluster:

```bash

# Test HTTP readiness probe

kubectl exec -it busybox -- wget -qO- http://<pod-ip>:<port>/<path>

# Test TCP liveness probe

kubectl exec -it busybox -- sh -c "echo > /dev/tcp/<pod-ip>/<port>"

# For exec probes, run the exact command inside the container

kubectl exec -it <pod-name> -- <command>

```

### 4. Check Resource Constraints and Network Policies

CPU throttling or memory pressure can cause probes to timeout. Verify resource utilization:

```bash
kubectl top pod <pod-name>

```

Ensure NetworkPolicies allow traffic from the node (kubelet) to the pod's probe ports, as restrictive policies often cause HTTP and TCP probes to fail silently.

### 5. Adjust Probe Parameters

Fine-tune timing based on application behavior:

- Increase `initialDelaySeconds` if the application requires extended warm-up time
- Raise `timeoutSeconds` for endpoints with slow response characteristics
- Modify `failureThreshold` to tolerate transient glitches or detect failures faster

### 6. View Detailed Probe State

Extract the container's last known state using JSONPath:

```bash
kubectl get pod <pod-name> -o jsonpath='{.status.containerStatuses[0].state}'

```

## Configuration Examples and Common Pitfalls

The following YAML configuration demonstrates proper probe implementation for an application with slow initialization:

```yaml
apiVersion: v1
kind: Pod
metadata:
  name: example-app
spec:
  containers:
  - name: web
    image: nginx:latest
    ports:
    - containerPort: 80
    startupProbe:
      httpGet:
        path: /healthz
        port: 80
      failureThreshold: 30
      periodSeconds: 10
    livenessProbe:
      httpGet:
        path: /healthz
        port: 80
      initialDelaySeconds: 30
      periodSeconds: 15
      timeoutSeconds: 5
      failureThreshold: 3
    readinessProbe:
      httpGet:
        path: /ready
        port: 80
      initialDelaySeconds: 5
      periodSeconds: 10
      timeoutSeconds: 2
      successThreshold: 1
      failureThreshold: 3

```

**Common configuration errors** identified across [`Sony/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Sony/DevOps_Engineer.md) and [`Sigmoid/DevOps_Engineer_2.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Sigmoid/DevOps_Engineer_2.md) include:

- **False-positive health endpoints**: The probe returns 200 OK while the actual service is non-functional. Ensure health checks validate critical dependencies like database connections.
- **Label mismatches**: Readiness passes but the pod receives no traffic because Service selectors do not match the pod's labels.
- **Startup probe failures**: The command or HTTP path is incorrect, or the initialization file/socket has not been created when the probe executes.

## Summary

- **Liveness probes** restart containers experiencing deadlocks or unrecoverable errors, while **readiness probes** control Service endpoint membership without restarting containers.
- **Startup probes** run before other probes to accommodate lengthy initialization sequences, preventing premature health check failures.
- Troubleshooting requires verifying probe definitions in the Pod spec, inspecting `kubectl describe` events, manually testing endpoints with `kubectl exec`, and adjusting `failureThreshold` and `timeoutSeconds` parameters.
- Resource constraints and NetworkPolicies can cause probe timeouts even when applications are healthy.
- Interview questions from [`Verizon/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Verizon/DevOps_Engineer.md), [`JPMorgan/DevOps_SRE.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/JPMorgan/DevOps_SRE.md), and [`Capgemini/DevOps_Engineer_2.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Capgemini/DevOps_Engineer_2.md) frequently target the distinction between liveness and readiness behavior.

## Frequently Asked Questions

### What is the difference between liveness and readiness probes in Kubernetes?

**Liveness probes** determine whether a container should be restarted due to unrecoverable failure states like deadlocks or memory leaks. **Readiness probes** determine whether a pod should receive traffic from Services and load balancers. A pod can be alive (passing liveness) but not ready (failing readiness), in which case it runs but receives no traffic. According to [`Sigmoid/DevOps_Engineer_2.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Sigmoid/DevOps_Engineer_2.md), this distinction is critical for zero-downtime deployments and graceful degradation.

### When should I use a startup probe instead of increasing initialDelaySeconds?

Use a **startup probe** when an application requires variable or unpredictable initialization time, such as during database migrations or cache warm-up. Unlike increasing `initialDelaySeconds` on liveness probes—which forces all pods to wait the same fixed duration regardless of actual readiness—startup probes allow the kubelet to begin liveness and readiness checks immediately upon successful initialization. This pattern is specifically recommended in [`LTIMindtree/DevOps_Engineer_L2.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/LTIMindtree/DevOps_Engineer_L2.md) for production applications with heavy startup loads.

### Why does my pod keep restarting even though the readiness probe passes?

Passing readiness probes do not prevent liveness probe failures. If the application passes readiness (indicating it can serve traffic) but subsequently deadlocks or hangs, the **liveness probe** will detect the unresponsive state and trigger a restart. As noted in [`Sony/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Sony/DevOps_Engineer.md), you must verify that the liveness endpoint accurately detects application hangs, not just that the web server process is running.

### How can I troubleshoot a startup probe that never succeeds?

First, verify the probe command or HTTP endpoint is correctly typed and accessible. Check `kubectl logs` for application startup errors that might prevent initialization completion. Ensure the probe's `failureThreshold` multiplied by `periodSeconds` provides adequate time for initialization. If using an exec probe, manually run the command inside the container using `kubectl exec` to confirm it returns exit code 0. Network Policies blocking kubelet access to the container can also cause startup probes to hang indefinitely, as mentioned in [`Others/DevOps_Engineer_4.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Others/DevOps_Engineer_4.md).