How to Use GKE Topology Spread Constraints with the Cluster Autoscaler: A Complete Guide

GKE topology spread constraints only trigger the Cluster Autoscaler when the whenUnsatisfiable field is set to DoNotSchedule, forcing the scheduler to reject pods that would violate the spread until the autoscaler provisions nodes in under-populated zones.

Google Kubernetes Engine (GKE) provides topology spread constraints to evenly distribute pods across failure domains such as zones or nodes. When integrated with the Cluster Autoscaler, these constraints can drive automatic node provisioning to maintain balance, but only when configured according to the specifications in the google/skills repository. This guide explains the exact implementation details found in skills/cloud/gke-cluster-autoscaler/SKILL.md to ensure your workloads scale correctly across topology domains.

Why the Default Setting Prevents Autoscaling

The autoscaler only respects topology spread constraints configured with hard requirements. By default, the whenUnsatisfiable field is set to ScheduleAnyway, which instructs the scheduler to ignore the constraint when it cannot be satisfied. Because the pod still schedules, the autoscaler receives no signal to create capacity in under-represented zones.

According to the source code documentation in skills/cloud/gke-cluster-autoscaler/SKILL.md, you must explicitly set whenUnsatisfiable: DoNotSchedule to make the constraint enforceable. This setting causes the scheduler to reject pods that would violate the spread, marking them as unschedulable and triggering the autoscaler to evaluate zone-specific capacity needs.

How Constraints Drive Node Provisioning

When properly configured, the interaction between the scheduler and autoscaler follows this architectural flow:

  1. Pod Creation – The scheduler evaluates the pod's topologySpreadConstraints against current node topology.
  2. Constraint Enforcement – With whenUnsatisfiable: DoNotSchedule, the scheduler blocks pods that would violate the maxSkew tolerance.
  3. Autoscaler Signal – Unschedulable pods are reported to the Cluster Autoscaler as capacity shortages.
  4. Zone-Aware Provisioning – The autoscaler identifies which zones lack sufficient nodes to satisfy the spread and provisions new instances (or new node pools) in those specific zones.
  5. Workload Placement – Once the new node registers, the pending pod schedules, restoring the desired distribution across failure domains.

Required Configuration Parameters

Based on the implementation details in skills/cloud/gke-cluster-autoscaler/SKILL.md and skills/cloud/gke-reliability/SKILL.md, use these specific settings to ensure autoscaler compatibility:

  • whenUnsatisfiable: DoNotSchedule – This is the critical setting that forces the scheduler to block pods until the autoscaler creates capacity.
  • maxSkew: 1 – Controls the maximum difference in pod count between topology domains; set to 1 for strict balance.
  • topologyKey: topology.kubernetes.io/zone – Specifies the failure domain for zonal distribution (use kubernetes.io/hostname for node-level spreading).
  • labelSelector – Must match your workload's labels to ensure the constraint applies only to the intended pod group.
  • PodDisruptionBudget – Define a PDB with minAvailable or maxUnavailable to prevent voluntary disruptions from breaking the established spread.

Complete Implementation Example

The following YAML implements a zone-aware deployment with a hard topology spread constraint compatible with the Cluster Autoscaler. It includes a PodDisruptionBudget to maintain availability during disruptions.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: myapp-pdb
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: myapp
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp
spec:
  replicas: 6
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels:
              app: myapp
      containers:
        - name: myapp
          image: gcr.io/my-project/myapp:latest
          ports:
            - containerPort: 8080

Key implementation details:

  • The PDB ensures at least two pods remain available during voluntary disruptions such as node upgrades.
  • maxSkew: 1 forces the pod count per zone to differ by at most one.
  • whenUnsatisfiable: DoNotSchedule makes the constraint a hard requirement, enabling the autoscaler to detect the need for additional nodes in specific zones.

Advanced Considerations and Edge Cases

Node Auto-Provisioning Integration

If you are running GKE version 1.33.3 or later, enable Node Auto-Provisioning (NAP) with ComputeClasses by setting spec.nodePoolAutoCreation.enabled: true in your ComputeClass configuration. As documented in skills/cloud/gke-compute-classes/SKILL.md, NAP works directly with topology spread constraints to automatically create node pools in zones where capacity is missing, rather than just scaling existing node pools.

Rolling Update Behavior

When using maxSurge values greater than 1 during rolling updates, surge pods may temporarily violate strict topology spread constraints. This can cause unnecessary autoscaler churn as the system attempts to place surge pods that exceed the maxSkew limit. Set maxSurge: 1 in your deployment strategy to ensure rolling updates respect the topological balance.

Troubleshooting with Visibility Logs

If the autoscaler fails to provision nodes despite unschedulable pods, inspect the Cluster Autoscaler visibility logs at container.googleapis.com/cluster-autoscaler-visibility. These logs reveal why the autoscaler is not scaling, such as identifying when constraints are evaluated as soft preferences rather than hard requirements.

Summary

  • Set whenUnsatisfiable: DoNotSchedule to enable the Cluster Autoscaler to react to topology spread constraints; the default ScheduleAnyway prevents autoscaling.
  • The autoscaler provisions nodes in specific zones only when pods become unschedulable due to hard constraints.
  • Combine topology spread constraints with PodDisruptionBudgets to maintain availability during disruptions and cluster maintenance.
  • Use maxSkew: 1 and topology.kubernetes.io/zone for strict zonal distribution.
  • Enable Node Auto-Provisioning for automatic node pool creation in under-populated zones.

Frequently Asked Questions

Why doesn't the Cluster Autoscaler scale when I use topology spread constraints?

The autoscaler only scales for hard constraints (whenUnsatisfiable: DoNotSchedule). If you use the default ScheduleAnyway, the scheduler places the pod regardless of topology imbalance, so the autoscaler sees no unschedulable pods requiring additional capacity.

Can topology spread constraints work with Node Auto-Provisioning?

Yes. According to skills/cloud/gke-compute-classes/SKILL.md, Node Auto-Provisioning (NAP) is designed to work with topology spread constraints. When NAP is enabled with spec.nodePoolAutoCreation.enabled: true, it can automatically create new node pools in zones where existing pools lack capacity to satisfy the spread.

What happens during rolling updates with strict topology constraints?

Rolling updates with maxSurge > 1 can create temporary pods that violate the maxSkew setting, potentially triggering unnecessary node provisioning. Set maxSurge: 1 in your deployment strategy to prevent surge pods from breaking the topological balance and causing autoscaler churn.

Do I need a PodDisruptionBudget with topology spread constraints?

While not strictly required for the autoscaler to function, a PodDisruptionBudget is recommended. As noted in skills/cloud/gke-reliability/SKILL.md, PDBs prevent voluntary disruptions (such as node upgrades) from removing too many pods in a single zone, which would otherwise break your carefully maintained topology spread.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →