# How to Use GKE Topology Spread Constraints with the Cluster Autoscaler: A Complete Guide

> Learn how to effectively use GKE topology spread constraints with the Cluster Autoscaler to ensure pod distribution. Discover the critical role of the whenUnsatisfiable setting for autoscaling.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-09

---

**GKE topology spread constraints only trigger the Cluster Autoscaler when the `whenUnsatisfiable` field is set to `DoNotSchedule`, forcing the scheduler to reject pods that would violate the spread until the autoscaler provisions nodes in under-populated zones.**

Google Kubernetes Engine (GKE) provides **topology spread constraints** to evenly distribute pods across failure domains such as zones or nodes. When integrated with the **Cluster Autoscaler**, these constraints can drive automatic node provisioning to maintain balance, but only when configured according to the specifications in the `google/skills` repository. This guide explains the exact implementation details found in [`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md) to ensure your workloads scale correctly across topology domains.

## Why the Default Setting Prevents Autoscaling

The autoscaler only respects topology spread constraints configured with **hard requirements**. By default, the `whenUnsatisfiable` field is set to `ScheduleAnyway`, which instructs the scheduler to ignore the constraint when it cannot be satisfied. Because the pod still schedules, the autoscaler receives no signal to create capacity in under-represented zones.

According to the source code documentation in [`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md), you must explicitly set `whenUnsatisfiable: DoNotSchedule` to make the constraint enforceable. This setting causes the scheduler to reject pods that would violate the spread, marking them as unschedulable and triggering the autoscaler to evaluate zone-specific capacity needs.

## How Constraints Drive Node Provisioning

When properly configured, the interaction between the scheduler and autoscaler follows this architectural flow:

1. **Pod Creation** – The scheduler evaluates the pod's `topologySpreadConstraints` against current node topology.
2. **Constraint Enforcement** – With `whenUnsatisfiable: DoNotSchedule`, the scheduler blocks pods that would violate the `maxSkew` tolerance.
3. **Autoscaler Signal** – Unschedulable pods are reported to the Cluster Autoscaler as capacity shortages.
4. **Zone-Aware Provisioning** – The autoscaler identifies which zones lack sufficient nodes to satisfy the spread and provisions new instances (or new node pools) in those specific zones.
5. **Workload Placement** – Once the new node registers, the pending pod schedules, restoring the desired distribution across failure domains.

## Required Configuration Parameters

Based on the implementation details in [`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md) and [`skills/cloud/gke-reliability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-reliability/SKILL.md), use these specific settings to ensure autoscaler compatibility:

- **`whenUnsatisfiable: DoNotSchedule`** – This is the critical setting that forces the scheduler to block pods until the autoscaler creates capacity.
- **`maxSkew: 1`** – Controls the maximum difference in pod count between topology domains; set to `1` for strict balance.
- **`topologyKey: topology.kubernetes.io/zone`** – Specifies the failure domain for zonal distribution (use `kubernetes.io/hostname` for node-level spreading).
- **`labelSelector`** – Must match your workload's labels to ensure the constraint applies only to the intended pod group.
- **`PodDisruptionBudget`** – Define a PDB with `minAvailable` or `maxUnavailable` to prevent voluntary disruptions from breaking the established spread.

## Complete Implementation Example

The following YAML implements a zone-aware deployment with a hard topology spread constraint compatible with the Cluster Autoscaler. It includes a PodDisruptionBudget to maintain availability during disruptions.

```yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: myapp-pdb
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: myapp
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp
spec:
  replicas: 6
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels:
              app: myapp
      containers:
        - name: myapp
          image: gcr.io/my-project/myapp:latest
          ports:
            - containerPort: 8080

```

**Key implementation details:**

- The **PDB** ensures at least two pods remain available during voluntary disruptions such as node upgrades.
- **`maxSkew: 1`** forces the pod count per zone to differ by at most one.
- **`whenUnsatisfiable: DoNotSchedule`** makes the constraint a hard requirement, enabling the autoscaler to detect the need for additional nodes in specific zones.

## Advanced Considerations and Edge Cases

### Node Auto-Provisioning Integration

If you are running GKE version 1.33.3 or later, enable **Node Auto-Provisioning (NAP)** with ComputeClasses by setting `spec.nodePoolAutoCreation.enabled: true` in your ComputeClass configuration. As documented in [`skills/cloud/gke-compute-classes/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md), NAP works directly with topology spread constraints to automatically create node pools in zones where capacity is missing, rather than just scaling existing node pools.

### Rolling Update Behavior

When using `maxSurge` values greater than `1` during rolling updates, surge pods may temporarily violate strict topology spread constraints. This can cause unnecessary autoscaler churn as the system attempts to place surge pods that exceed the `maxSkew` limit. Set `maxSurge: 1` in your deployment strategy to ensure rolling updates respect the topological balance.

### Troubleshooting with Visibility Logs

If the autoscaler fails to provision nodes despite unschedulable pods, inspect the **Cluster Autoscaler visibility logs** at `container.googleapis.com/cluster-autoscaler-visibility`. These logs reveal why the autoscaler is not scaling, such as identifying when constraints are evaluated as soft preferences rather than hard requirements.

## Summary

- Set **`whenUnsatisfiable: DoNotSchedule`** to enable the Cluster Autoscaler to react to topology spread constraints; the default `ScheduleAnyway` prevents autoscaling.
- The autoscaler provisions nodes in specific zones only when pods become unschedulable due to hard constraints.
- Combine topology spread constraints with **PodDisruptionBudgets** to maintain availability during disruptions and cluster maintenance.
- Use **`maxSkew: 1`** and **`topology.kubernetes.io/zone`** for strict zonal distribution.
- Enable **Node Auto-Provisioning** for automatic node pool creation in under-populated zones.

## Frequently Asked Questions

### Why doesn't the Cluster Autoscaler scale when I use topology spread constraints?

The autoscaler only scales for **hard constraints** (`whenUnsatisfiable: DoNotSchedule`). If you use the default `ScheduleAnyway`, the scheduler places the pod regardless of topology imbalance, so the autoscaler sees no unschedulable pods requiring additional capacity.

### Can topology spread constraints work with Node Auto-Provisioning?

Yes. According to [`skills/cloud/gke-compute-classes/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md), Node Auto-Provisioning (NAP) is designed to work with topology spread constraints. When NAP is enabled with `spec.nodePoolAutoCreation.enabled: true`, it can automatically create new node pools in zones where existing pools lack capacity to satisfy the spread.

### What happens during rolling updates with strict topology constraints?

Rolling updates with `maxSurge > 1` can create temporary pods that violate the `maxSkew` setting, potentially triggering unnecessary node provisioning. Set `maxSurge: 1` in your deployment strategy to prevent surge pods from breaking the topological balance and causing autoscaler churn.

### Do I need a PodDisruptionBudget with topology spread constraints?

While not strictly required for the autoscaler to function, a **PodDisruptionBudget** is recommended. As noted in [`skills/cloud/gke-reliability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-reliability/SKILL.md), PDBs prevent voluntary disruptions (such as node upgrades) from removing too many pods in a single zone, which would otherwise break your carefully maintained topology spread.