# Troubleshooting and Debugging GKE Workloads: 5 Critical Upgrade Blockers and Fixes

> Troubleshoot GKE upgrade blockers like PDBs, resource limits, and PVC issues. Learn fixes and debugging techniques for smoother GKE workload upgrades with the google/skills guide.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-13

---

**GKE upgrades stall most often due to Pod Disruption Budgets blocking drains, resource constraints, bare pods, admission webhooks, or PVC attachment issues, all of which can be diagnosed and resolved using the troubleshooting guide in the google/skills repository.**

When automating Google Kubernetes Engine (GKE) cluster maintenance, operators frequently encounter workloads that prevent node pools from completing their upgrade cycle. According to the **GKE Upgrades skill** in the `google/skills` repository, troubleshooting and debugging GKE workloads effectively requires understanding five specific failure categories defined in [`skills/cloud/gke-upgrades/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-upgrades/SKILL.md). This guide walks through each blocker with precise diagnostic commands and remediation steps sourced directly from the official troubleshooting documentation.

## Understanding GKE Upgrade Architecture

Before debugging specific workloads, grasp these architectural constraints that govern how GKE handles upgrades:

- **Version Skew** – Nodes can trail the control plane by up to two minor versions (SKILL.md lines 33-34). This guarantees API compatibility but means nodes must eventually upgrade to remain within the supported window.
- **Surge vs. Rolling Strategies** – Surge upgrades add extra nodes before draining old ones, while rolling upgrades replace nodes one-by-one (SKILL.md lines 27-31). Surge requires additional quota and GPU reservations but minimizes disruption.
- **Maintenance Windows & Exclusions** – Automated upgrades respect exclusion windows defined in the skill documentation (lines 68-90), but manual upgrades bypass these protections.
- **Mandatory Overrides** – GKE may force upgrades for security patches or certificate rotation regardless of exclusions (lines 94-100), making proactive debugging essential.

## The Five Root Causes of Stalled GKE Upgrades

The troubleshooting guide at [`skills/cloud/gke-upgrades/references/troubleshooting.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-upgrades/references/troubleshooting.md) identifies five specific categories that block node drains during upgrades.

### 1. Pod Disruption Budgets (PDBs) Blocking Node Drain

A PDB with `ALLOWED DISRUPTIONS = 0` prevents GKE from evicting pods during surge upgrades, causing the node drain to hang indefinitely.

**Diagnosis:**

```bash
kubectl get pdb -A -o wide               # Look for ALLOWED DISRUPTIONS = 0

kubectl describe pdb <PDB_NAME> -n <NS>    # Inspect the selector and current disruptions

```

**Remediation:**

Temporarily relax the PDB to allow full disruptions during the maintenance window:

```bash
kubectl patch pdb <PDB_NAME> -n <NS> \
  -p '{"spec":{"minAvailable":null,"maxUnavailable":"100%"}}'

```

Restore the original configuration after the upgrade completes.

### 2. Resource Constraints Leaving Pods Pending

Insufficient CPU, memory, or quota limits prevent new nodes from scheduling necessary pods, stalling the upgrade while the cluster waits for resources that never become available.

**Diagnosis:**

```bash
kubectl get pods -A | grep Pending
kubectl get events -A --field-selector reason=FailedScheduling
kubectl top nodes
kubectl describe nodes | grep -A5 "Allocated resources"

```

**Remediation:**

Increase surge capacity to allow more nodes during the upgrade window:

```bash
gcloud container node-pools update <NODE_POOL> \
  --cluster <CLUSTER> --zone <ZONE> \
  --max-surge-upgrade 2 --max-unavailable-upgrade 0

```

This provides headroom for workloads to migrate without hitting quota ceilings.

### 3. Bare Pods Without Controllers

Pods lacking owner references (Deployments, StatefulSets, DaemonSets, etc.) cannot be rescheduled by controllers. GKE refuses to drain nodes hosting these bare pods because they will not recreate elsewhere.

**Diagnosis:**

```bash
kubectl get pods -A -o json | \
  jq -r '.items[] | select(.metadata.ownerReferences | length == 0) |
          "\(.metadata.namespace)/\(.metadata.name)"'

```

**Remediation:**

Either delete the bare pod if it is ephemeral:

```bash
kubectl delete pod <POD> -n <NS>

```

Or convert it to a Deployment with appropriate replica counts to ensure high availability during upgrades.

### 4. Admission Webhooks Rejecting Pod Creation

Validating or mutating webhooks that fail closed or reject pod specifications prevent new nodes from receiving workloads, causing the upgrade to stall while waiting for successful pod scheduling.

**Diagnosis:**

```bash
kubectl get validatingwebhookconfigurations
kubectl get mutatingwebhookconfigurations
kubectl describe validatingwebhookconfigurations <WEBHOOK>

```

Check for `failurePolicy: Fail` configurations or restrictive namespace selectors that might exclude new nodes.

**Remediation:**

Temporarily remove the problematic webhook configuration:

```bash
kubectl delete validatingwebhookconfigurations <WEBHOOK>

```

Re-create the webhook configuration after the node pool upgrade completes.

### 5. Persistent Volume Claim Attachment Issues

Zone-locked PersistentVolumes or mismatched storage classes prevent pods from scheduling on newly created nodes in different zones, blocking the old node termination.

**Diagnosis:**

```bash
kubectl get pvc -A | grep -v Bound
kubectl get events -A --field-selector reason=FailedAttachVolume

```

**Remediation:**

Ensure workloads using zonal disks are constrained to nodes in the same zone, or migrate to regional storage classes. For immediate unblocking, manually drain the old node after ensuring volumes are detached:

```bash
kubectl cordon <OLD_NODE>
kubectl drain <OLD_NODE> --ignore-daemonsets --force

```

## Validation and Verification Workflow

After applying any fix from the troubleshooting guide, verify the upgrade resumes using the validation block (lines 89-99 in the skill documentation):

```bash
watch 'kubectl get nodes -o wide | grep -E "NAME|CURRENT_VERSION|TARGET_VERSION"'
kubectl get pods -A | grep -E "Terminating|Pending"
gcloud container operations list --cluster <CLUSTER> --zone <ZONE> --limit=1

```

Monitor the `gcloud container operations list` output for `DONE` status and confirm no pods remain stuck in `Terminating` or `Pending` states.

## Summary

- **PDBs with zero allowed disruptions** are the most common cause of stalled upgrades; temporarily raising `maxUnavailable` unblocks drains.
- **Resource constraints** require surge capacity adjustments to provide migration headroom.
- **Bare pods** halt drains because they cannot be rescheduled; convert them to Deployments or delete them.
- **Admission webhooks** may reject pods on new nodes; disable them temporarily during upgrades.
- **PVC zone mismatches** prevent pod scheduling; verify storage topology matches node pool zones.

If the upgrade remains stuck after these checks, consult the **runbook template** at [`skills/cloud/gke-upgrades/references/runbook-template.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-upgrades/references/runbook-template.md) for comprehensive remediation workflows including rollback procedures and maintenance window adjustments.

## Frequently Asked Questions

### How do I identify which PDB is blocking my GKE node upgrade?

Run `kubectl get pdb -A -o wide` and look for any entry showing `ALLOWED DISRUPTIONS = 0`. Then use `kubectl describe pdb <NAME> -n <NAMESPACE>` to view the selector and current pod disruptions. The PDB prevents node drains when it cannot evict pods without violating `minAvailable` requirements.

### What is the difference between surge and rolling upgrades in GKE?

Surge upgrades create new nodes before draining old ones, requiring additional quota and IP addresses but ensuring capacity remains available. Rolling upgrades replace nodes one-by-one, which conserves quota but risks capacity constraints. GPU and TPU node pools often require `maxSurge=0` due to quota-heavy reservations and driver coupling constraints documented in SKILL.md lines 40-52.

### Can I completely block GKE from upgrading my nodes?

No. While you can configure **maintenance windows and exclusions** (SKILL.md lines 68-90) to defer automated upgrades, GKE reserves the right to force upgrades for critical security patches or certificate rotations (lines 94-100). You cannot fully block **mandatory overrides**, so maintaining proper PDB configurations and resource headroom is essential.

### Why do my pods stay in Pending state during node upgrades?

Pending pods during upgrades typically indicate **resource constraints** (insufficient CPU/memory on new nodes) or **webhook rejections** preventing scheduling. Check `kubectl get events --field-selector reason=FailedScheduling` for specific failures, and verify that admission webhooks are not blocking pod creation on the new node versions.