Managing GKE Cluster Upgrades and Maintenance Windows: A Complete Operational Guide
Managing GKE cluster upgrades and maintenance windows requires a structured workflow that upgrades the control plane before node pools, configures appropriate release channels, and applies exclusion policies to control automatic updates.
The gke-upgrades skill in the google/skills repository provides a comprehensive framework for planning, executing, and validating Google Kubernetes Engine upgrades while respecting operational constraints. This guide distills the operational patterns defined in skills/cloud/gke-upgrades/SKILL.md and supporting reference files into actionable commands and checklists.
Core Architectural Concepts
Understanding GKE's upgrade mechanics is essential before modifying production clusters. The source code analysis reveals strict ordering requirements and policy boundaries that govern how versions progress.
Versioning and Upgrade Order
GKE enforces a specific sequence for version progression to maintain API compatibility and node stability. The control plane must always be upgraded before any associated node pools. Nodes are permitted to lag behind the control plane by no more than two minor versions, but never exceed it.
Minor version upgrades for the control plane occur sequentially (N → N+1 → N+2), even when targeting a version several releases ahead. Conversely, node pools can skip directly to the target minor version once the control plane supports it.
Release Channel Strategy
Selecting the appropriate release channel determines the velocity of automatic upgrades and SLA guarantees. According to SKILL.md, the four channels serve distinct operational needs:
- Rapid: Ideal for development and testing environments seeking early feature access. Google provides no upgrade-stability SLA for this channel.
- Regular: The default channel for most production workloads, offering a balance of stability and feature availability with full SLA coverage.
- Stable: Designed for mission-critical applications prioritizing stability over new features, backed by full SLA guarantees.
- Extended: Compliance-focused environments requiring extended support timelines (up to 24 months) with full SLA coverage.
Critical Rule: Static versioning ("No channel") is deprecated. Production clusters should migrate to an appropriate release channel combined with maintenance exclusions rather than remaining unenrolled.
Maintenance Windows and Exclusions
Maintenance windows define recurring time ranges when automatic upgrades may occur. Exclusions provide granular control by blocking specific upgrade types regardless of the maintenance window schedule.
The skill defines three exclusion scopes with distinct limitations:
| Scope | Blocked Upgrades | Maximum Duration |
|---|---|---|
no_upgrades |
All upgrades (minor, patch, and node) | 90 days in any rolling 365-day window |
no_minor_or_node_upgrades |
Minor and node upgrades only (patches allowed) | 180 days per exclusion, extendable until End-of-Support |
no_minor_upgrades |
Minor upgrades only (patches and node upgrades allowed) | 180 days per exclusion |
Exclusions affect automatic upgrades only. Manual gcloud commands bypass these restrictions, allowing operators to force upgrades during emergencies or planned maintenance outside scheduled windows.
Node Pool Upgrade Strategies
Standard GKE clusters support three distinct strategies for migrating node pools to new versions, as documented in the skill's node pool strategy section:
Surge Upgrades (Default) Best suited for stateless or low-risk workloads. Configure aggressive surge parameters to minimize upgrade time:
gcloud container node-pools update NODE_POOL_NAME \
--cluster CLUSTER_NAME \
--zone ZONE \
--max-surge-upgrade 3 \
--max-unavailable-upgrade 0
Blue-Green Upgrades Required for mission-critical workloads demanding immediate rollback capability. This strategy creates a parallel node pool at the target version, migrates workloads via cordon and drain operations, then decommissions the original pool.
Autoscaled Blue-Green (Preview) Similar to Blue-Green but leverages cluster autoscaling to manage capacity during the transition, reducing the need to manually size the new pool.
GPU/TPU Constraint: GPU nodes cannot utilize live migration. These workloads require rolling upgrades with maxSurge=0 and maxUnavailable=1. Blue-Green strategies are infeasible for GPU workloads because they would double GPU quota consumption.
Mandatory Upgrade Overrides
Google Cloud reserves the right to force upgrades regardless of maintenance windows or exclusions in specific scenarios:
- Critical security patches requiring immediate deployment
- End-of-Support enforcement for deprecated versions
- Expiring control-plane certificates
- Maintenance starvation (insufficient window coverage over extended periods)
When overrides occur, consult the GKE Release Notes and Security Bulletins for specific context and mitigation steps.
Step-by-Step GKE Upgrade Workflow
The checklists.md and runbook-template.md files in skills/cloud/gke-upgrades/references/ codify a nine-phase workflow:
- Gather Context – Document cluster mode (Standard/Autopilot), current versions, target versions, release channel, and workload sensitivity.
- Select Release Channel – Choose the appropriate channel and verify target version availability using
gcloud container get-server-config. - Define Maintenance Window – Establish a recurring window (typically off-peak hours) using RFC 5545 recurrence rules.
- Add Exclusions – Apply the appropriate exclusion scope for required freeze periods.
- Choose Node-Pool Strategy – Select surge, blue-green, or autoscaled blue-green based on workload characteristics.
- Execute Pre-Upgrade Checklist – Complete the markdown checklist from
references/checklists.md. - Run Upgrade Runbook – Execute commands from
references/runbook-template.md. - Complete Post-Upgrade Checklist – Verify version compliance, pod health, and observability metrics.
- Document Rollback Plan – Define downgrade paths for control-plane patches and node-pool rollbacks.
Essential Gcloud Commands for GKE Upgrades
Inspecting Current Cluster State
Verify existing versions before planning upgrades:
gcloud container clusters describe CLUSTER_NAME \
--zone ZONE \
--format="table(name, currentMasterVersion, nodePools[].version)"
Check available versions per release channel:
gcloud container get-server-config --zone ZONE \
--format="yaml(channels)"
Configuring Maintenance Windows
Set a weekly recurring maintenance window for Saturday early morning:
gcloud container clusters update CLUSTER_NAME \
--zone ZONE \
--maintenance-window-start 2026-08-14T02:00:00Z \
--maintenance-window-end 2026-08-14T06:00:00Z \
--maintenance-window-recurrence "FREQ=WEEKLY;BYDAY=SA"
Implementing Maintenance Exclusions
Create a 90-day total freeze for holiday code freezes:
gcloud container clusters update CLUSTER_NAME \
--zone ZONE \
--add-maintenance-exclusion-name="freeze-q4-2026" \
--add-maintenance-exclusion-start=2026-10-01T00:00:00Z \
--add-maintenance-exclusion-end=2026-12-31T23:59:59Z
Apply a persistent exclusion blocking minor and node upgrades until End-of-Support:
gcloud container clusters update CLUSTER_NAME \
--zone ZONE \
--add-maintenance-exclusion-name="no-minor-or-node-upgrades" \
--add-maintenance-exclusion-start=2026-08-01T00:00:00Z \
--add-maintenance-exclusion-until-end-of-support \
--add-maintenance-exclusion-scope=no_minor_or_node_upgrades
Executing Node Pool Upgrades
Standard surge upgrade:
gcloud container node-pools upgrade NODE_POOL_NAME \
--cluster CLUSTER_NAME \
--zone ZONE \
--cluster-version TARGET_VERSION
Blue-Green migration workflow:
# Create new pool at target version
gcloud container node-pools create NEW_POOL \
--cluster CLUSTER_NAME \
--zone ZONE \
--cluster-version TARGET_VERSION \
--num-nodes 3 \
--machine-type n1-standard-4
# Isolate and drain old pool
kubectl cordon -l cloud.google.com/gke-nodepool=OLD_POOL
kubectl drain -l cloud.google.com/gke-nodepool=OLD_POOL \
--ignore-daemonsets \
--delete-emptydir-data
# After validation, remove old pool
gcloud container node-pools delete OLD_POOL \
--cluster CLUSTER_NAME \
--zone ZONE
Rollback Procedures
Revert control plane to previous patch version:
gcloud container clusters upgrade CLUSTER_NAME \
--master \
--zone ZONE \
--cluster-version PREVIOUS_PATCH_VERSION
Troubleshooting Stuck Upgrades
The references/troubleshooting.md file provides a five-point diagnostic checklist for upgrades that halt mid-process:
# 1. Check for Pod Disruption Budgets blocking drains
kubectl get pdb -A
# 2. Identify resource constraints causing pending pods
kubectl get pods -A --field-selector=status.phase=Pending
# 3. Locate bare pods without owner references
kubectl get pods -A --field-selector='metadata.ownerReferences==null'
# 4. Verify admission webhooks aren't rejecting new pods
kubectl get validatingwebhookconfigurations
kubectl get mutatingwebhookconfigurations
# 5. Check for PVC attachment errors
kubectl describe pvc -A | grep -i error
Key Source Files in the gke-upgrades Skill
| File | Purpose |
|---|---|
skills/cloud/gke-upgrades/SKILL.md |
Central definition of upgrade principles, release channels, and exclusion policies |
skills/cloud/gke-upgrades/references/checklists.md |
Pre- and post-upgrade checklist templates |
skills/cloud/gke-upgrades/references/runbook-template.md |
Ready-to-run command blocks for each upgrade phase |
skills/cloud/gke-upgrades/references/troubleshooting.md |
Diagnostic procedures for stuck upgrades |
Summary
- Control planes must upgrade before node pools, and nodes may lag by a maximum of two minor versions behind the control plane.
- Release channels (Rapid, Regular, Stable, Extended) determine upgrade velocity and SLA coverage; static versioning is deprecated.
- Maintenance exclusions provide three scopes of protection (
no_upgrades,no_minor_or_node_upgrades,no_minor_upgrades) with specific duration limits. - Node pool strategies include Surge for stateless workloads and Blue-Green for mission-critical applications requiring fast rollback.
- GPU workloads require rolling upgrades with
maxSurge=0and cannot use Blue-Green strategies due to quota constraints. - Mandatory overrides may force upgrades for critical security patches regardless of exclusion policies.
Frequently Asked Questions
What is the correct order for upgrading GKE cluster components?
The control plane must always be upgraded first, followed by node pools. Control plane upgrades proceed sequentially through minor versions (N → N+1), while node pools can skip directly to the target version once the control plane supports it. Nodes must never exceed the control plane version and may lag behind by no more than two minor versions.
How long can I block GKE automatic upgrades using maintenance exclusions?
The duration depends on the exclusion scope. The no_upgrades scope blocks all upgrades for a maximum of 90 days in any rolling 365-day window. The no_minor_or_node_upgrades and no_minor_upgrades scopes allow 180 days per exclusion, with the former extendable until the current version reaches End-of-Support.
Can I perform manual upgrades during a maintenance exclusion period?
Yes. Maintenance windows and exclusions apply only to automatic upgrades managed by Google Cloud. Manual upgrades executed via gcloud container clusters upgrade or gcloud container node-pools upgrade bypass these restrictions entirely, allowing operators to upgrade on demand during emergencies or planned maintenance.
What node pool upgrade strategy should I use for GPU-enabled clusters?
GPU-enabled node pools must use surge upgrades with maxSurge=0 and maxUnavailable=1 because GPU nodes do not support live migration. Blue-Green strategies are not recommended for GPU workloads because creating a parallel pool would double the GPU quota consumption, often exceeding project limits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →