# Best Practices for GKE Upgrades and Maintenance: A Complete Guide

> Master GKE upgrades and maintenance. Learn best practices for control plane upgrades, version skew, and maintenance exclusions to ensure smooth cluster operations. Keep your GKE cluster running optimally.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: best-practices
- Published: 2026-08-13

---

**Always upgrade the GKE control plane before node pools, maintain version skew within two minor versions, and use maintenance exclusions with release channel discipline to prevent unplanned disruptions.**

Google Kubernetes Engine upgrades require careful orchestration to minimize downtime and maintain cluster health. The `gke-upgrades` skill in the `google/skills` repository codifies the authoritative workflow for managing these operations safely. Following the principles defined in [`skills/cloud/gke-upgrades/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-upgrades/SKILL.md) ensures your teams can execute upgrades systematically while avoiding common failure modes.

## Upgrade Planning: Gather Context First

Before generating any upgrade plan, capture the cluster’s operating characteristics to determine the appropriate strategy. The skill mandates evaluating five key dimensions:

- **Cluster mode** – Distinguish between Standard (full node-pool control) and Autopilot (managed nodes requiring explicit resource requests)【L32】.
- **Version skew policy** – Ensure node pools remain within **two minor versions** of the control plane to maintain compatibility【L33】.
- **Release channel** – Identify whether the cluster uses Rapid, Regular (default), Stable, or Extended channels, as this determines upgrade cadence and available versions【L54-L59】.
- **Environment topology** – Map single-cluster versus multi-cluster architectures, dev/staging/prod tiers, and whether Rollout Sequencing coordinates upgrades across environments【L35-L36】.
- **Workload sensitivity** – Flag StatefulSets, databases, GPU workloads, and long-running batch jobs that require specialized handling during node replacements【L36-L37】.

If any inputs are missing, default to the **Regular** channel with Standard mode and surge upgrades, but explicitly document these assumptions.

## Core Principles of GKE Upgrades

The `gke-upgrades` skill establishes six non-negotiable principles for safe upgrades:

**Sequential control-plane → node-pool upgrades** – The control plane must upgrade first; nodes may lag by up to two minor versions but never exceed the control plane version【L44-L46】.

**Environment progression** – Upgrade development and staging environments before production, preferably using Rollout Sequencing to validate changes【L46-L47】.

**Workload-aware strategy** – Select surge, blue-green, or rolling upgrades based on statefulness and resource constraints. Stateless workloads tolerate aggressive surge settings, while stateful workloads require conservative approaches【L47-L50】.

**Release-channel discipline** – Avoid "No channel" static versioning, which is deprecated; instead, use a release channel and enforce upgrade blocks through maintenance exclusions【L84-L86】.

**Rollback readiness** – Control plane patches and node-pool minor/patch upgrades support rollback, though full control plane minor rollbacks require Google support intervention【L49-L50】.

**Node-pool ordering** – Upgrade non-critical, stateless pools first as canaries, then proceed to stateful, GPU, or mission-critical pools【L50-L51】.

## Release Channels and Support Lifecycle

GKE offers four release channels, each suited to different operational requirements:

| Channel | Best For | Support SLA |
|---------|----------|-------------|
| **Rapid** | Development and early-feature testing | No upgrade stability guarantee |
| **Regular** | Most production workloads | Full SLA |
| **Stable** | Mission-critical, stability-first environments | Full SLA |
| **Extended** | Compliance-heavy workloads requiring long-term support | Full SLA |

Clusters on the **Regular** channel receive **14 months** of support after the version enters the channel【L62-L66】. The **Extended** channel prolongs this to **24 months**, ideal for organizations with rigid change windows.

## Maintenance Windows and Exclusions

Define maintenance windows to control when auto-upgrades occur. Manual upgrades bypass these exclusions, but auto-upgrades respect three distinct exclusion scopes with specific duration limits【L70-L80】:

| Scope | Blocks | Maximum Duration |
|-------|--------|------------------|
| `no_upgrades` | All upgrades (minor, patch, node) | 90 days per rolling 365-day window |
| `no_minor_or_node_upgrades` | Minor and node upgrades (allows patches) | 180 days per exclusion, extendable to end-of-support |
| `no_minor_upgrades` | Minor upgrades only (allows patches and node upgrades) | 180 days per exclusion |

**Critical distinction:** Auto-upgrade exclusions **only block automatic upgrades**; manual upgrades execute regardless of active exclusions【L84-L85】. Google has deprecated the "No channel" configuration; clusters using static versioning must migrate to release channels with proper exclusions【L86-L87】.

Apply exclusions using explicit gcloud flags:

```bash
gcloud container clusters update $CLUSTER \
  --add-maintenance-exclusion-name=no-upgrades \
  --add-maintenance-exclusion-start=2024-09-01T00:00:00Z \
  --add-maintenance-exclusion-end=2024-09-30T23:59:59Z \
  --add-maintenance-exclusion-scope=no_upgrades

```

This syntax follows the skill’s "Correct gcloud syntax" rule【L89-L90】.

## Handling Mandatory Upgrade Overrides

GKE may force upgrades despite active exclusions for:

- **Critical security patches**
- **End-of-Support (EoS) or End-of-Life (EOL) enforcement**
- **Expiring certificates** (triggered when ≤ 30 days remain)
- **Maintenance starvation** (insufficient 48-hour availability windows within a 32-day period)【L95-L100】

When encountering a forced upgrade, consult the **GKE Release Notes** or **Security Bulletins** for specific remediation guidance【L104-L105】.

## Node-Pool Upgrade Strategies for Standard Clusters

Select surge parameters based on workload characteristics to balance speed with stability:

| Workload Type | `maxSurge` | `maxUnavailable` | Rationale |
|---------------|------------|------------------|-----------|
| Stateless services | 2-3 | 0 | Fast rollout with spare capacity |
| Stateful sets / Databases | 1 | 0 | Conservative replacement to protect data |
| GPU (fixed reservation) | 0 | 1 | No extra quota available; serial replacement |
| Large pools (≥ 50 nodes) | 20 | 0 | Parallel processing for scale |

Default to **Surge** upgrades for general workloads. For mission-critical stateful services, implement **Standard Blue-Green** upgrades, or use the preview **Autoscaled Blue-Green** strategy if you can provision temporary extra capacity【L34-L36】.

## GPU and TPU Cluster Considerations

AI/ML workloads on GPU and TPU nodes require specialized handling due to hardware constraints:

**No live migration** – GPU VMs cannot live-migrate; upgrades trigger pod restarts【L42-L45】.

**Driver coupling** – OS image upgrades include new Linux kernels and NVIDIA drivers. Validate CUDA compatibility in staging before production deployment【L46-L48】.

**Zero-surge requirement** – Use rolling upgrades with `maxSurge=0` and `maxUnavailable=1` to release fixed GPU reservations before provisioning replacements, preventing quota exhaustion【L44-L45】.

**Blue-green infeasibility** – Blue-green upgrades require double the GPU quota, making them impractical for large GPU clusters【L45-L46】.

## Operationalizing with Checklists and Runbooks

The `gke-upgrades` skill provides operational templates in [`references/checklists.md`](https://github.com/google/skills/blob/main/references/checklists.md) and [`references/runbook-template.md`](https://github.com/google/skills/blob/main/references/runbook-template.md) to standardize upgrade procedures【L59-L63】.

Pre-upgrade verification must confirm:

- **Pod Disruption Budgets (PDBs)** are not overly restrictive, ensuring `ALLOWED DISRUPTIONS > 0`【L69-L71】.
- **Resource requests** are explicitly defined for Autopilot clusters【L65-L68】.
- **Persistent Volume Claims (PVCs)** are zone-compatible for regional clusters, avoiding attachment failures【L99-L100】.

The runbook template contains concrete `gcloud` commands for control-plane upgrades【L32-L35】 and Standard node-pool upgrades【L55-L59】, ensuring consistent execution across teams.

## Troubleshooting Stuck or Failed Upgrades

When upgrades stall, systematically inspect all five common failure causes【L86-L95】:

1. **PDB blocking drain** – Run `kubectl get pdb -A` and identify entries with `ALLOWED DISRUPTIONS = 0`.
2. **Resource constraints** – Check for pods pending due to CPU/memory quota exhaustion or capacity limits.
3. **Bare pods** – Locate pods without owner references (controllers) that block node drains; delete or adopt them.
4. **Admission webhooks** – Verify that validating or mutating webhooks are not rejecting pods scheduled on new nodes.
5. **PVC attachment issues** – Investigate zone-locked volumes or failed CSI attachments in regional clusters.

If the error indicates **stockout or quota exhaustion** (`ZONE_RESOURCE_POOL_EXHAUSTED` or `QUOTA_EXCEEDED`), mitigate by switching to a rolling upgrade with `maxSurge=0` or requesting a quota increase【L96-L100】.

## Summary

- **Always sequence upgrades** – Control plane first, then node pools, maintaining ≤ 2 minor version skew.
- **Use release channels** – Select Regular for most production workloads, Extended for compliance needs; never use "No channel."
- **Configure exclusions properly** – Apply `no_minor_or_node_upgrades` for security freezes, understanding that manual upgrades bypass blocks.
- **Match strategy to workload** – Use surge upgrades for stateless services, zero-surge rolling upgrades for GPU pools, and blue-green only when spare capacity exists.
- **Verify pre-conditions** – Check PDBs, resource requests, and PVC zone compatibility before initiating upgrades.
- **Prepare for forced upgrades** – Monitor certificate expiration and maintenance windows to avoid surprise upgrades for critical patches.

## Frequently Asked Questions

### How long can I block GKE auto-upgrades using maintenance exclusions?

The maximum duration depends on the exclusion scope. The `no_upgrades` scope blocks all upgrades for up to 90 days per rolling 365-day window, while `no_minor_or_node_upgrades` and `no_minor_upgrades` allow 180 days per exclusion, extendable to end-of-support【L70-L80】.

### Can I rollback a GKE upgrade if something goes wrong?

Control plane patches and node-pool minor or patch upgrades support rollback procedures. However, rolling back a full control plane minor version upgrade requires Google support assistance and is not self-service【L49-L50】.

### Why does my GPU node pool upgrade keep failing with quota errors?

GPU nodes use fixed hardware reservations that prevent surge upgrades. You must use `maxSurge=0` and `maxUnavailable=1` to serially migrate workloads, as blue-green strategies requiring double the GPU quota are typically infeasible for large AI/ML clusters【L44-L46】.

### What is the difference between Regular and Extended release channels?

The Regular channel provides 14 months of support per version and receives new features and patches automatically. The Extended channel extends support to 24 months with a slower feature rollout, designed for organizations with strict compliance requirements or infrequent change windows【L62-L66】.