Why Kubernetes Pods Get Stuck in Pending Status: Causes and Solutions

A Kubernetes pod enters the Pending status when the scheduler cannot place it on any available node due to unsatisfied resource requirements, mismatched node selectors, unbound persistent volumes, or image pull failures.

The DevOps-Interview-Guide repository documents this scenario as a critical interview topic for site reliability engineers. Understanding the root causes of Kubernetes pod pending status is essential for diagnosing why workloads fail to transition to the Running state.

Common Causes of Kubernetes Pod Pending Status

According to the source documentation in TCS/SRE_1.md, a pod remains Pending when the Kubernetes scheduler cannot satisfy its placement constraints or dependencies. The following eight categories account for the majority of production incidents.

Insufficient CPU and Memory Resources

The most frequent cause occurs when a pod’s resource requests exceed available cluster capacity. When requests or limits for CPU/memory are higher than any node can provide, the scheduler reports events like "0/3 nodes are available: 3 Insufficient cpu".

Typical fix: Reduce the pod’s resources.requests and resources.limits in the deployment spec, or scale the cluster by adding larger nodes.

Node Selector and Affinity Mismatches

Pods using nodeSelector, nodeAffinity, or podAffinity rules target specific hardware characteristics or topology domains. If no nodes carry the required labels, or if anti-affinity rules conflict with existing pods, the pod remains unschedulable.

Typical fix: Verify node labels with kubectl get nodes --show-labels and adjust selectors, or label appropriate nodes to match the pod requirements.

Taints and Tolerations Conflicts

Nodes marked with taints (such as dedicated=production:NoSchedule) repel pods that lack matching tolerations. If all nodes are tainted and the pod spec omits the required toleration array, scheduling stalls indefinitely.

Typical fix: Add a tolerations stanza to the pod spec matching the node taint, or remove the taint from target nodes using kubectl taint node <name> <key>-.

Persistent Volume Claims (PVC) Unbound

Pods mounting storage depend on PVCs binding to PersistentVolumes (PVs). If the requested StorageClass does not exist, the provisioner is down, or no PV matches the claim, the pod waits in Pending until storage is available.

Typical fix: Ensure the StorageClass exists and has a dynamic provisioner, or manually create a PV that satisfies the PVC requirements.

Image Pull Failures

When a container image cannot be retrieved—due to invalid tags, private registry authentication errors, or network partitions—the kubelet prevents the pod from starting. This manifests as ImagePullBackOff or ErrImagePull events.

Typical fix: Verify the image name and tag, then create an imagePullSecrets secret and associate it with the service account or pod spec.

Resource Quotas and Limits

Namespace-level ResourceQuota objects or cluster-level limits can cap the total CPU, memory, or pod count. If the quota is exhausted, the scheduler blocks new pods even when node resources are available.

Typical fix: Review quota usage with kubectl describe quota, then increase limits or reduce existing workload requests.

Node Unschedulable or Not Ready

Nodes marked NotReady, SchedulingDisabled, or manually cordoned (kubectl cordon) are excluded from scheduling decisions. If all nodes fall into these states, no placement targets exist.

Typical fix: Resolve underlying node health issues (kubelet status, network connectivity), restart the kubelet service, or uncordon nodes with kubectl uncordon <node-name>.

Admission Controller Denials

Policies enforced by PodSecurityPolicy, OPA Gatekeeper, or validating webhooks can reject pod specs during the admission phase. While these sometimes generate Failed states, certain configurations leave the pod Pending indefinitely.

Typical fix: Review admission controller logs and policy constraints, then modify the pod spec to comply with security policies.

How to Diagnose Pending Pods in Kubernetes

Systematic diagnosis requires inspecting multiple cluster objects. The TCS/SRE_1.md file recommends this workflow to identify the specific blocker.

Inspect the pod events for explicit scheduler messages:

kubectl describe pod <pod-name> -n <namespace>

Check node availability and scheduling status:

kubectl get nodes

Verify storage binding status if the pod uses PVCs:

kubectl get pvc -n <namespace>

Examine previous container logs for image pull errors:

kubectl logs <pod-name> -n <namespace> --previous

Resolution Steps and Fixes

The following workflow from the DevOps-Interview-Guide resolves the majority of Pending scenarios.

Step 1: Identify the specific constraint using describe:

kubectl describe pod my-app -n prod

Step 2: Adjust resource constraints if nodes lack capacity:

kubectl edit deployment my-app -n prod

# Lower requests.cpu and requests.memory values

Step 3: Correct node selectors or add tolerations:

kubectl edit pod my-app -n prod

# Add nodeSelector labels or tolerations array

Step 4: Ensure PVCs can bind to storage:

kubectl apply -f pvc.yaml

# Verify status changes to Bound

Step 5: Configure registry credentials for private images:

kubectl create secret docker-registry my-regcred \
  --docker-server=registry.example.com \
  --docker-username=user \
  --docker-password=pass \
  --docker-email=user@example.com

kubectl patch serviceaccount default -p '{"imagePullSecrets": [{"name": "my-regcred"}]}'

Summary

  • Kubernetes pod pending status indicates the scheduler cannot place the pod on any node, as documented in TCS/SRE_1.md.
  • Resource constraints (CPU/memory) and node selector mismatches represent the most common root causes.
  • PersistentVolume Claims must reach the Bound state before pods using them can schedule.
  • Image pull secrets and registry authentication errors prevent container initialization.
  • Taints, cordoned nodes, and admission policies create soft locks that require configuration changes rather than resource scaling.
  • Use kubectl describe pod to read scheduler events and pinpoint the exact failure domain.

Frequently Asked Questions

How can I tell if a pod is Pending due to resource constraints?

Run kubectl describe pod <name> and examine the Events section. Look for messages stating "Insufficient cpu" or "Insufficient memory" followed by the count of unavailable nodes. This confirms that the pod’s resource requests exceed available capacity on all nodes.

What is the difference between Pending and ContainerCreating states?

Pending occurs before the scheduler assigns the pod to a node, indicating issues with resource availability, node selectors, or storage binding. ContainerCreating appears after scheduling succeeds but before the container runtime starts the image, usually indicating volume mount problems or image pull delays.

Can a pod stay Pending indefinitely?

Yes. A pod remains in Pending status until all scheduling constraints are satisfied or the pod is deleted. There is no automatic timeout in the Kubernetes scheduler itself, though higher-level controllers like Jobs may enforce activeDeadlineSeconds.

How do I fix a Pending pod caused by taints on all nodes?

Add a matching toleration to the pod specification that corresponds to the node taint. For example, if nodes have dedicated=production:NoSchedule, add a toleration with key: dedicated, operator: Equal, value: production, and effect: NoSchedule to the pod spec. Alternatively, remove the taint from nodes where this workload should run using kubectl taint node <node-name> dedicated-.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →