# Understanding GKE Zonal Stockout Cascade and Fallback to Lower Priority Tiers

> Learn how GKE zonal stockout cascade works. Discover how pending pods automatically fallback to lower priority tiers during node pool stockouts to maintain availability.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: deep-dive
- Published: 2026-08-09

---

**When a GKE node pool's preferred machine family experiences a hard stockout in a specific zone, the cluster autoscaler places the entire priority tier on a global cooldown and routes all pending pods to the next available priority tier across all zones.**

The GKE Cluster Autoscaler manages capacity through priority tiers defined in ComputeClasses, but zonal resource exhaustion creates a cascading fallback mechanism that can shift your entire fleet to lower-cost tiers unexpectedly. According to the `google/skills` repository documentation, this behavior is governed by strict cooldown rules that activate when constrained pods trigger a `ZONE_RESOURCE_POOL_EXHAUSTED` error, forcing the system to bypass exhausted tiers regardless of the original scaling request's zone preferences.

## How Zonal Stockout Triggers the Cascade

A **hard stockout** occurs when GKE reports `out_of_resources` or `ZONE_RESOURCE_POOL_EXHAUSTED` for a specific machine family in a zone. In this scenario, the autoscaler does not retry the same tier indefinitely.

Instead, the system implements a **global cooldown period of approximately 5 minutes** for the entire affected priority tier. During this cooldown, all pending pod-scaling requests skip the exhausted tier entirely and are directed to the **next-available priority tier across all zones**. This mechanism is documented in [`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md) within the "Zonal stockout cooldown cascade" section.

The cascade behavior depends critically on pod constraints:

- **Constrained pods** (those requiring zonal PersistentVolumes or with zonal `nodeSelector/affinity`) trigger the cooldown when they cannot schedule in the exhausted zone
- **Unconstrained pods** with a `BALANCED` location policy merely skew scaling toward healthy zones without triggering the tier cooldown

## Why Workloads Fall Back to Lower Priority Tiers

The fallback occurs because the autoscaler treats the priority tier as globally exhausted once any zone reports a hard stockout for that tier's machine family. This creates a "drain toward the lowest tier" effect where your fleet migrates to the cheapest available ComputeClass priority even when the original request specified higher-cost resources.

The logic flow operates as follows:

1. A constrained pod requests scaling in `us-central1-a` with a Spot ComputeClass priority
2. The zone reports `ZONE_RESOURCE_POOL_EXHAUSTED` for the requested machine family
3. The autoscaler places the Spot priority tier on global cooldown (~5 minutes)
4. All pending pods—regardless of zone—skip the Spot tier and fall back to the next priority (e.g., On-Demand)
5. If the fallback tier also experiences pressure, the cascade continues to lower-cost alternatives

## Mitigation Strategies

### Add Intermediate Priority Tiers

Configure your ComputeClass with intermediate fallback steps to limit cascade depth. Rather than falling directly from Spot to the cheapest On-Demand tier, insert a second-tier On-Demand option.

```yaml
apiVersion: autoscaling.gke.io/v1beta1
kind: ComputeClass
metadata:
  name: spot-intermediate
spec:
  machineFamily: n2
  priorities:
  - name: spot-primary
    cost: 0.0
  - name: spot-intermediate
    cost: 0.02
    machineFamily: n2
  - name: ondemand-cheapest
    cost: 0.04
    machineFamily: e2

```

### Isolate Stateful Workloads

Separate stateful workloads that require zonal PersistentVolumes into dedicated ComputeClasses or namespaces. This prevents zonal PV constraints from forcing a stockout cooldown that drags unconstrained workloads down the priority ladder.

### Use Strict Topology Constraints

Prevent pods from forcing stockouts in specific zones by implementing `topologySpreadConstraints` with strict scheduling requirements.

```yaml
apiVersion: v1
kind: Pod
metadata:
  name: web
spec:
  topologySpreadConstraints:
  - maxSkew: 1
    whenUnsatisfiable: DoNotSchedule
    topologyKey: topology.kubernetes.io/zone
    labelSelector:
      matchLabels:
        app: web
  containers:
  - name: web
    image: gcr.io/my-project/web:latest

```

### Configure Location Policies for Zone Flexibility

Enable fallback to alternative zones by setting `locationPolicy: ANY` in your ComputeClass specification. This allows the autoscaler to route Spot instances to any available zone before triggering the priority tier cooldown.

```yaml
apiVersion: autoscaling.gke.io/v1beta1
kind: ComputeClass
metadata:
  name: spot-primary
spec:
  machineFamily: n2
  locationPolicy: ANY
  priorities:
  - name: spot-primary
    cost: 0.0
  - name: ondemand-fallback
    cost: 0.04
    machineFamily: n2
    zones: [us-central1-a, us-central1-b, us-central1-c]

```

## Source Code Reference

The stockout cascade behavior is implemented in the GKE Cluster Autoscaler logic documented across several files in the `google/skills` repository:

- **[`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md)** – Defines the zonal stockout cooldown cascade mechanism and the 5-minute global cooldown rule
- **[`skills/cloud/gke-compute-classes/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md)** – Specifies the ComputeClass schema for configuring priority tiers and fallback sequences
- **[`skills/cloud/gke-compute-classes/references/compute-class-debug.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/references/compute-class-debug.md)** – Provides debugging guidance for `scale.up.error.out.of.resources` errors
- **[`skills/cloud/gke-cluster-autoscaler/assets/log-autoscaler-events.sh`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/assets/log-autoscaler-events.sh)** – Utility script for tailing autoscaler visibility logs to detect stockout events

## Summary

- **Hard stockouts** (`ZONE_RESOURCE_POOL_EXHAUSTED`) trigger a **5-minute global cooldown** on the affected priority tier across all zones
- **Constrained pods** with zonal PersistentVolumes or node selectors trigger the cascade, while unconstrained pods with `BALANCED` policies only skew zone selection
- The cascade routes **all pending pods** to lower priority tiers, potentially draining your fleet to the cheapest available option
- **Mitigation** requires intermediate priority tiers, workload isolation, strict topology constraints, and flexible location policies (`ANY`)

## Frequently Asked Questions

### What is a hard stockout in GKE?

A hard stockout occurs when a specific zone reports `ZONE_RESOURCE_POOL_EXHAUSTED` or `out_of_resources` for a requested machine family, indicating that Google Cloud capacity is completely unavailable for that resource type in that location. This differs from temporary scheduling delays because it triggers the autoscaler's global cooldown mechanism for the entire priority tier.

### How long does the priority tier cooldown last?

The global cooldown period lasts approximately **5 minutes** according to the GKE Cluster Autoscaler documentation in [`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md). During this window, the autoscaler treats the exhausted priority tier as unavailable across all zones, forcing pending pods to fall back to lower priority alternatives regardless of whether other zones actually have capacity.

### Why do unconstrained pods not trigger the cooldown?

Unconstrained pods—those without zonal PersistentVolume claims or zonal node selectors—allow the autoscaler to use the `BALANCED` location policy to skew scaling toward healthy zones without declaring the priority tier exhausted. Since these pods can schedule in any zone, the autoscaler simply avoids the stockout zone rather than triggering the global cooldown that affects the entire tier.

### How can I prevent stockout cascades to the cheapest tier?

Insert **intermediate priority tiers** in your ComputeClass configuration to create graduated fallback steps rather than abrupt jumps from Spot to the lowest-cost On-Demand option. Additionally, isolate stateful workloads that require specific zones, use `topologySpreadConstraints` with `DoNotSchedule` to prevent zone-pinning, and set `locationPolicy: ANY` to maximize zone flexibility before the cooldown triggers.