Understanding GKE Zonal Stockout Cascade and Fallback to Lower Priority Tiers

When a GKE node pool's preferred machine family experiences a hard stockout in a specific zone, the cluster autoscaler places the entire priority tier on a global cooldown and routes all pending pods to the next available priority tier across all zones.

The GKE Cluster Autoscaler manages capacity through priority tiers defined in ComputeClasses, but zonal resource exhaustion creates a cascading fallback mechanism that can shift your entire fleet to lower-cost tiers unexpectedly. According to the google/skills repository documentation, this behavior is governed by strict cooldown rules that activate when constrained pods trigger a ZONE_RESOURCE_POOL_EXHAUSTED error, forcing the system to bypass exhausted tiers regardless of the original scaling request's zone preferences.

How Zonal Stockout Triggers the Cascade

A hard stockout occurs when GKE reports out_of_resources or ZONE_RESOURCE_POOL_EXHAUSTED for a specific machine family in a zone. In this scenario, the autoscaler does not retry the same tier indefinitely.

Instead, the system implements a global cooldown period of approximately 5 minutes for the entire affected priority tier. During this cooldown, all pending pod-scaling requests skip the exhausted tier entirely and are directed to the next-available priority tier across all zones. This mechanism is documented in skills/cloud/gke-cluster-autoscaler/SKILL.md within the "Zonal stockout cooldown cascade" section.

The cascade behavior depends critically on pod constraints:

  • Constrained pods (those requiring zonal PersistentVolumes or with zonal nodeSelector/affinity) trigger the cooldown when they cannot schedule in the exhausted zone
  • Unconstrained pods with a BALANCED location policy merely skew scaling toward healthy zones without triggering the tier cooldown

Why Workloads Fall Back to Lower Priority Tiers

The fallback occurs because the autoscaler treats the priority tier as globally exhausted once any zone reports a hard stockout for that tier's machine family. This creates a "drain toward the lowest tier" effect where your fleet migrates to the cheapest available ComputeClass priority even when the original request specified higher-cost resources.

The logic flow operates as follows:

  1. A constrained pod requests scaling in us-central1-a with a Spot ComputeClass priority
  2. The zone reports ZONE_RESOURCE_POOL_EXHAUSTED for the requested machine family
  3. The autoscaler places the Spot priority tier on global cooldown (~5 minutes)
  4. All pending pods—regardless of zone—skip the Spot tier and fall back to the next priority (e.g., On-Demand)
  5. If the fallback tier also experiences pressure, the cascade continues to lower-cost alternatives

Mitigation Strategies

Add Intermediate Priority Tiers

Configure your ComputeClass with intermediate fallback steps to limit cascade depth. Rather than falling directly from Spot to the cheapest On-Demand tier, insert a second-tier On-Demand option.

apiVersion: autoscaling.gke.io/v1beta1
kind: ComputeClass
metadata:
  name: spot-intermediate
spec:
  machineFamily: n2
  priorities:
  - name: spot-primary
    cost: 0.0
  - name: spot-intermediate
    cost: 0.02
    machineFamily: n2
  - name: ondemand-cheapest
    cost: 0.04
    machineFamily: e2

Isolate Stateful Workloads

Separate stateful workloads that require zonal PersistentVolumes into dedicated ComputeClasses or namespaces. This prevents zonal PV constraints from forcing a stockout cooldown that drags unconstrained workloads down the priority ladder.

Use Strict Topology Constraints

Prevent pods from forcing stockouts in specific zones by implementing topologySpreadConstraints with strict scheduling requirements.

apiVersion: v1
kind: Pod
metadata:
  name: web
spec:
  topologySpreadConstraints:
  - maxSkew: 1
    whenUnsatisfiable: DoNotSchedule
    topologyKey: topology.kubernetes.io/zone
    labelSelector:
      matchLabels:
        app: web
  containers:
  - name: web
    image: gcr.io/my-project/web:latest

Configure Location Policies for Zone Flexibility

Enable fallback to alternative zones by setting locationPolicy: ANY in your ComputeClass specification. This allows the autoscaler to route Spot instances to any available zone before triggering the priority tier cooldown.

apiVersion: autoscaling.gke.io/v1beta1
kind: ComputeClass
metadata:
  name: spot-primary
spec:
  machineFamily: n2
  locationPolicy: ANY
  priorities:
  - name: spot-primary
    cost: 0.0
  - name: ondemand-fallback
    cost: 0.04
    machineFamily: n2
    zones: [us-central1-a, us-central1-b, us-central1-c]

Source Code Reference

The stockout cascade behavior is implemented in the GKE Cluster Autoscaler logic documented across several files in the google/skills repository:

Summary

  • Hard stockouts (ZONE_RESOURCE_POOL_EXHAUSTED) trigger a 5-minute global cooldown on the affected priority tier across all zones
  • Constrained pods with zonal PersistentVolumes or node selectors trigger the cascade, while unconstrained pods with BALANCED policies only skew zone selection
  • The cascade routes all pending pods to lower priority tiers, potentially draining your fleet to the cheapest available option
  • Mitigation requires intermediate priority tiers, workload isolation, strict topology constraints, and flexible location policies (ANY)

Frequently Asked Questions

What is a hard stockout in GKE?

A hard stockout occurs when a specific zone reports ZONE_RESOURCE_POOL_EXHAUSTED or out_of_resources for a requested machine family, indicating that Google Cloud capacity is completely unavailable for that resource type in that location. This differs from temporary scheduling delays because it triggers the autoscaler's global cooldown mechanism for the entire priority tier.

How long does the priority tier cooldown last?

The global cooldown period lasts approximately 5 minutes according to the GKE Cluster Autoscaler documentation in skills/cloud/gke-cluster-autoscaler/SKILL.md. During this window, the autoscaler treats the exhausted priority tier as unavailable across all zones, forcing pending pods to fall back to lower priority alternatives regardless of whether other zones actually have capacity.

Why do unconstrained pods not trigger the cooldown?

Unconstrained pods—those without zonal PersistentVolume claims or zonal node selectors—allow the autoscaler to use the BALANCED location policy to skew scaling toward healthy zones without declaring the priority tier exhausted. Since these pods can schedule in any zone, the autoscaler simply avoids the stockout zone rather than triggering the global cooldown that affects the entire tier.

How can I prevent stockout cascades to the cheapest tier?

Insert intermediate priority tiers in your ComputeClass configuration to create graduated fallback steps rather than abrupt jumps from Spot to the lowest-cost On-Demand option. Additionally, isolate stateful workloads that require specific zones, use topologySpreadConstraints with DoNotSchedule to prevent zone-pinning, and set locationPolicy: ANY to maximize zone flexibility before the cooldown triggers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →