Kubernetes Liveness, Readiness, and Startup Probes: A Complete Troubleshooting Guide
Kubernetes liveness probes restart containers experiencing deadlocks or hangs, readiness probes control traffic routing to healthy pods, and startup probes protect slow-initializing applications from premature health checks.
Kubernetes probes are automated health-checking mechanisms that the kubelet uses to determine container viability and traffic eligibility. According to the litu54/DevOps-Interview-Guide repository—which contains interview questions from companies like Verizon, JPMorgan, and Sony—mastering these three probe types is essential for production reliability and DevOps technical interviews. This guide explains how each probe functions and provides systematic troubleshooting strategies based on real-world scenarios.
What Are Kubernetes Liveness, Readiness, and Startup Probes?
Kubernetes implements three distinct health-checking mechanisms, each serving a specific operational purpose:
| Probe | Execution Timing | Health Check Purpose | Failure Action |
|---|---|---|---|
| Liveness | Runs periodically after initialDelaySeconds |
Determines if the application is alive and can continue running | Container is killed and restarted after failureThreshold |
| Readiness | Runs periodically after initialDelaySeconds |
Determines if the application is ready to serve traffic | Pod removed from Service endpoints; container not restarted |
| Startup | Runs once (or until success) before other probes | Verifies completion of heavy initialization, migrations, or warm-up sequences | Container killed if it fails failureThreshold; on success, liveness/readiness probes begin |
Liveness probes detect unrecoverable states like deadlocks or memory leaks that require a container restart. Readiness probes ensure load balancers only route traffic to pods that have completed startup sequences and initialized dependencies. Startup probes are specifically designed for applications with lengthy initialization periods, preventing premature liveness or readiness failures during the warm-up phase.
How Kubernetes Probes Work Internally
The kubelet invokes probe logic defined in the Pod spec using one of three mechanisms: an HTTP GET request, a TCP socket check, or an exec command. The kubelet tracks probe results and updates the pod's status accordingly:
- Success: Resets the failure counter; the container remains healthy.
- Failure: After
failureThresholdconsecutive failures, the kubelet marks the container unhealthy and takes probe-specific action.
For liveness, the kubelet kills and restarts the container. For readiness, the pod is removed from Service endpoints but continues running. For startup, failure results in container termination, while success transitions the pod to standard liveness and readiness monitoring.
Troubleshooting Kubernetes Probe Failures
When applications experience instability or traffic routing issues, probe misconfiguration is often the root cause. Use the following systematic approach to diagnose and resolve problems:
1. Verify Probe Definitions
Inspect the Pod spec to confirm the probe type (HTTP/TCP/exec) matches your application's health endpoints. Check timing parameters carefully: initialDelaySeconds, periodSeconds, timeoutSeconds, successThreshold, and failureThreshold. A common error documented in SquareOps/DevOps_Engineer.md involves setting initialDelaySeconds too low for applications requiring several seconds to start, causing unnecessary restart loops.
2. Inspect Pod Logs and Events
Examine application output for health-check failures or startup errors:
kubectl logs <pod-name> -c <container-name>
kubectl describe pod <pod-name>
Look specifically for events labeled Liveness probe failed, Readiness probe failed, or Startup probe failed in the kubectl describe output.
3. Manually Test Probe Endpoints
Validate accessibility by executing test commands from within the cluster:
# Test HTTP readiness probe
kubectl exec -it busybox -- wget -qO- http://<pod-ip>:<port>/<path>
# Test TCP liveness probe
kubectl exec -it busybox -- sh -c "echo > /dev/tcp/<pod-ip>/<port>"
# For exec probes, run the exact command inside the container
kubectl exec -it <pod-name> -- <command>
4. Check Resource Constraints and Network Policies
CPU throttling or memory pressure can cause probes to timeout. Verify resource utilization:
kubectl top pod <pod-name>
Ensure NetworkPolicies allow traffic from the node (kubelet) to the pod's probe ports, as restrictive policies often cause HTTP and TCP probes to fail silently.
5. Adjust Probe Parameters
Fine-tune timing based on application behavior:
- Increase
initialDelaySecondsif the application requires extended warm-up time - Raise
timeoutSecondsfor endpoints with slow response characteristics - Modify
failureThresholdto tolerate transient glitches or detect failures faster
6. View Detailed Probe State
Extract the container's last known state using JSONPath:
kubectl get pod <pod-name> -o jsonpath='{.status.containerStatuses[0].state}'
Configuration Examples and Common Pitfalls
The following YAML configuration demonstrates proper probe implementation for an application with slow initialization:
apiVersion: v1
kind: Pod
metadata:
name: example-app
spec:
containers:
- name: web
image: nginx:latest
ports:
- containerPort: 80
startupProbe:
httpGet:
path: /healthz
port: 80
failureThreshold: 30
periodSeconds: 10
livenessProbe:
httpGet:
path: /healthz
port: 80
initialDelaySeconds: 30
periodSeconds: 15
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: 80
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 2
successThreshold: 1
failureThreshold: 3
Common configuration errors identified across Sony/DevOps_Engineer.md and Sigmoid/DevOps_Engineer_2.md include:
- False-positive health endpoints: The probe returns 200 OK while the actual service is non-functional. Ensure health checks validate critical dependencies like database connections.
- Label mismatches: Readiness passes but the pod receives no traffic because Service selectors do not match the pod's labels.
- Startup probe failures: The command or HTTP path is incorrect, or the initialization file/socket has not been created when the probe executes.
Summary
- Liveness probes restart containers experiencing deadlocks or unrecoverable errors, while readiness probes control Service endpoint membership without restarting containers.
- Startup probes run before other probes to accommodate lengthy initialization sequences, preventing premature health check failures.
- Troubleshooting requires verifying probe definitions in the Pod spec, inspecting
kubectl describeevents, manually testing endpoints withkubectl exec, and adjustingfailureThresholdandtimeoutSecondsparameters. - Resource constraints and NetworkPolicies can cause probe timeouts even when applications are healthy.
- Interview questions from
Verizon/DevOps_Engineer.md,JPMorgan/DevOps_SRE.md, andCapgemini/DevOps_Engineer_2.mdfrequently target the distinction between liveness and readiness behavior.
Frequently Asked Questions
What is the difference between liveness and readiness probes in Kubernetes?
Liveness probes determine whether a container should be restarted due to unrecoverable failure states like deadlocks or memory leaks. Readiness probes determine whether a pod should receive traffic from Services and load balancers. A pod can be alive (passing liveness) but not ready (failing readiness), in which case it runs but receives no traffic. According to Sigmoid/DevOps_Engineer_2.md, this distinction is critical for zero-downtime deployments and graceful degradation.
When should I use a startup probe instead of increasing initialDelaySeconds?
Use a startup probe when an application requires variable or unpredictable initialization time, such as during database migrations or cache warm-up. Unlike increasing initialDelaySeconds on liveness probes—which forces all pods to wait the same fixed duration regardless of actual readiness—startup probes allow the kubelet to begin liveness and readiness checks immediately upon successful initialization. This pattern is specifically recommended in LTIMindtree/DevOps_Engineer_L2.md for production applications with heavy startup loads.
Why does my pod keep restarting even though the readiness probe passes?
Passing readiness probes do not prevent liveness probe failures. If the application passes readiness (indicating it can serve traffic) but subsequently deadlocks or hangs, the liveness probe will detect the unresponsive state and trigger a restart. As noted in Sony/DevOps_Engineer.md, you must verify that the liveness endpoint accurately detects application hangs, not just that the web server process is running.
How can I troubleshoot a startup probe that never succeeds?
First, verify the probe command or HTTP endpoint is correctly typed and accessible. Check kubectl logs for application startup errors that might prevent initialization completion. Ensure the probe's failureThreshold multiplied by periodSeconds provides adequate time for initialization. If using an exec probe, manually run the command inside the container using kubectl exec to confirm it returns exit code 0. Network Policies blocking kubelet access to the container can also cause startup probes to hang indefinitely, as mentioned in Others/DevOps_Engineer_4.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →