Troubleshooting and Debugging GKE Workloads: 5 Critical Upgrade Blockers and Fixes
GKE upgrades stall most often due to Pod Disruption Budgets blocking drains, resource constraints, bare pods, admission webhooks, or PVC attachment issues, all of which can be diagnosed and resolved using the troubleshooting guide in the google/skills repository.
When automating Google Kubernetes Engine (GKE) cluster maintenance, operators frequently encounter workloads that prevent node pools from completing their upgrade cycle. According to the GKE Upgrades skill in the google/skills repository, troubleshooting and debugging GKE workloads effectively requires understanding five specific failure categories defined in skills/cloud/gke-upgrades/SKILL.md. This guide walks through each blocker with precise diagnostic commands and remediation steps sourced directly from the official troubleshooting documentation.
Understanding GKE Upgrade Architecture
Before debugging specific workloads, grasp these architectural constraints that govern how GKE handles upgrades:
- Version Skew – Nodes can trail the control plane by up to two minor versions (SKILL.md lines 33-34). This guarantees API compatibility but means nodes must eventually upgrade to remain within the supported window.
- Surge vs. Rolling Strategies – Surge upgrades add extra nodes before draining old ones, while rolling upgrades replace nodes one-by-one (SKILL.md lines 27-31). Surge requires additional quota and GPU reservations but minimizes disruption.
- Maintenance Windows & Exclusions – Automated upgrades respect exclusion windows defined in the skill documentation (lines 68-90), but manual upgrades bypass these protections.
- Mandatory Overrides – GKE may force upgrades for security patches or certificate rotation regardless of exclusions (lines 94-100), making proactive debugging essential.
The Five Root Causes of Stalled GKE Upgrades
The troubleshooting guide at skills/cloud/gke-upgrades/references/troubleshooting.md identifies five specific categories that block node drains during upgrades.
1. Pod Disruption Budgets (PDBs) Blocking Node Drain
A PDB with ALLOWED DISRUPTIONS = 0 prevents GKE from evicting pods during surge upgrades, causing the node drain to hang indefinitely.
Diagnosis:
kubectl get pdb -A -o wide # Look for ALLOWED DISRUPTIONS = 0
kubectl describe pdb <PDB_NAME> -n <NS> # Inspect the selector and current disruptions
Remediation:
Temporarily relax the PDB to allow full disruptions during the maintenance window:
kubectl patch pdb <PDB_NAME> -n <NS> \
-p '{"spec":{"minAvailable":null,"maxUnavailable":"100%"}}'
Restore the original configuration after the upgrade completes.
2. Resource Constraints Leaving Pods Pending
Insufficient CPU, memory, or quota limits prevent new nodes from scheduling necessary pods, stalling the upgrade while the cluster waits for resources that never become available.
Diagnosis:
kubectl get pods -A | grep Pending
kubectl get events -A --field-selector reason=FailedScheduling
kubectl top nodes
kubectl describe nodes | grep -A5 "Allocated resources"
Remediation:
Increase surge capacity to allow more nodes during the upgrade window:
gcloud container node-pools update <NODE_POOL> \
--cluster <CLUSTER> --zone <ZONE> \
--max-surge-upgrade 2 --max-unavailable-upgrade 0
This provides headroom for workloads to migrate without hitting quota ceilings.
3. Bare Pods Without Controllers
Pods lacking owner references (Deployments, StatefulSets, DaemonSets, etc.) cannot be rescheduled by controllers. GKE refuses to drain nodes hosting these bare pods because they will not recreate elsewhere.
Diagnosis:
kubectl get pods -A -o json | \
jq -r '.items[] | select(.metadata.ownerReferences | length == 0) |
"\(.metadata.namespace)/\(.metadata.name)"'
Remediation:
Either delete the bare pod if it is ephemeral:
kubectl delete pod <POD> -n <NS>
Or convert it to a Deployment with appropriate replica counts to ensure high availability during upgrades.
4. Admission Webhooks Rejecting Pod Creation
Validating or mutating webhooks that fail closed or reject pod specifications prevent new nodes from receiving workloads, causing the upgrade to stall while waiting for successful pod scheduling.
Diagnosis:
kubectl get validatingwebhookconfigurations
kubectl get mutatingwebhookconfigurations
kubectl describe validatingwebhookconfigurations <WEBHOOK>
Check for failurePolicy: Fail configurations or restrictive namespace selectors that might exclude new nodes.
Remediation:
Temporarily remove the problematic webhook configuration:
kubectl delete validatingwebhookconfigurations <WEBHOOK>
Re-create the webhook configuration after the node pool upgrade completes.
5. Persistent Volume Claim Attachment Issues
Zone-locked PersistentVolumes or mismatched storage classes prevent pods from scheduling on newly created nodes in different zones, blocking the old node termination.
Diagnosis:
kubectl get pvc -A | grep -v Bound
kubectl get events -A --field-selector reason=FailedAttachVolume
Remediation:
Ensure workloads using zonal disks are constrained to nodes in the same zone, or migrate to regional storage classes. For immediate unblocking, manually drain the old node after ensuring volumes are detached:
kubectl cordon <OLD_NODE>
kubectl drain <OLD_NODE> --ignore-daemonsets --force
Validation and Verification Workflow
After applying any fix from the troubleshooting guide, verify the upgrade resumes using the validation block (lines 89-99 in the skill documentation):
watch 'kubectl get nodes -o wide | grep -E "NAME|CURRENT_VERSION|TARGET_VERSION"'
kubectl get pods -A | grep -E "Terminating|Pending"
gcloud container operations list --cluster <CLUSTER> --zone <ZONE> --limit=1
Monitor the gcloud container operations list output for DONE status and confirm no pods remain stuck in Terminating or Pending states.
Summary
- PDBs with zero allowed disruptions are the most common cause of stalled upgrades; temporarily raising
maxUnavailableunblocks drains. - Resource constraints require surge capacity adjustments to provide migration headroom.
- Bare pods halt drains because they cannot be rescheduled; convert them to Deployments or delete them.
- Admission webhooks may reject pods on new nodes; disable them temporarily during upgrades.
- PVC zone mismatches prevent pod scheduling; verify storage topology matches node pool zones.
If the upgrade remains stuck after these checks, consult the runbook template at skills/cloud/gke-upgrades/references/runbook-template.md for comprehensive remediation workflows including rollback procedures and maintenance window adjustments.
Frequently Asked Questions
How do I identify which PDB is blocking my GKE node upgrade?
Run kubectl get pdb -A -o wide and look for any entry showing ALLOWED DISRUPTIONS = 0. Then use kubectl describe pdb <NAME> -n <NAMESPACE> to view the selector and current pod disruptions. The PDB prevents node drains when it cannot evict pods without violating minAvailable requirements.
What is the difference between surge and rolling upgrades in GKE?
Surge upgrades create new nodes before draining old ones, requiring additional quota and IP addresses but ensuring capacity remains available. Rolling upgrades replace nodes one-by-one, which conserves quota but risks capacity constraints. GPU and TPU node pools often require maxSurge=0 due to quota-heavy reservations and driver coupling constraints documented in SKILL.md lines 40-52.
Can I completely block GKE from upgrading my nodes?
No. While you can configure maintenance windows and exclusions (SKILL.md lines 68-90) to defer automated upgrades, GKE reserves the right to force upgrades for critical security patches or certificate rotations (lines 94-100). You cannot fully block mandatory overrides, so maintaining proper PDB configurations and resource headroom is essential.
Why do my pods stay in Pending state during node upgrades?
Pending pods during upgrades typically indicate resource constraints (insufficient CPU/memory on new nodes) or webhook rejections preventing scheduling. Check kubectl get events --field-selector reason=FailedScheduling for specific failures, and verify that admission webhooks are not blocking pod creation on the new node versions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →