# How to Configure GKE Cluster Autoscaler for Optimal Performance: A Complete Guide

> Optimize GKE Cluster Autoscaler performance with this guide. Learn to configure utilization, location policies, ComputeClasses, topology constraints, and eliminate scale-down blockers for peak efficiency.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-13

---

**To configure GKE Cluster Autoscaler for optimal performance, combine the `optimize-utilization` autoscaling profile, strategic location policies (`ANY` for Spot, `BALANCED` for HA), ComputeClasses with fallback priorities, and strict Pod Topology Spread Constraints while eliminating scale-down blockers and monitoring visibility logs.**

The GKE Cluster Autoscaler dynamically adjusts node counts to match workload demand, but achieving the right balance between cost, latency, and reliability requires precise configuration. According to the `google/skills` repository, modern GKE clusters (v1.33.3+) leverage **ComputeClasses** with `nodePoolAutoCreation.enabled: true`, enabling zero-size node pools that scale to zero when idle, fundamentally changing how you configure GKE Cluster Autoscaler for optimal performance compared to legacy Node Auto Provisioning.

## Select the Appropriate Autoscaling Profile

The autoscaling profile controls how aggressively the autoscaler packs workloads and removes idle nodes. In [`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md) (lines 26-30), two profiles are defined:

- **`optimize-utilization`** – Prioritizes node consolidation and rapid scale-down. Use this for cost-driven batch workloads or development environments where minimizing idle capacity is paramount.
- **`balanced`** – Maintains buffers for faster scale-up. Use this for latency-sensitive production services where pod startup time is critical.

Switch profiles using gcloud:

```bash
gcloud container clusters update my-cluster \
  --autoscaling-profile=optimize-utilization

```

## Configure Location Policies for Zone Distribution

Location policies determine where new nodes are provisioned across zones. As documented in [`skills/cloud/gke-cluster-autoscaler/references/ca-optimization.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/references/ca-optimization.md) (lines 14-18):

- **`BALANCED`** – Best-effort distribution across zones for high-availability on-demand workloads. Note that this balances node placement, not pod placement.
- **`ANY`** – Maximizes capacity acquisition across zones, essential for Spot instances or when regional stockouts occur.

For Spot-first ComputeClasses, set `locationPolicy: ANY` to maximize obtainability. For critical on-demand workloads requiring zone distribution, use `BALANCED`.

## Implement ComputeClasses with Fallback Priorities

Modern GKE versions prefer **ComputeClasses** over classic node auto-provisioning. The [`ca-provisioning.md`](https://github.com/google/skills/blob/main/ca-provisioning.md) reference (lines 29-35) specifies enabling `nodePoolAutoCreation.enabled: true` to allow zero-size pools that the autoscaler can instantiate on demand.

Define fallback priorities to handle stockouts. According to [`ca-optimization.md`](https://github.com/google/skills/blob/main/ca-optimization.md) (lines 21-27), configure `location.locationPolicy` within the ComputeClass spec:

```yaml
apiVersion: autoscaling.gke.io/v1beta1
kind: ComputeClass
metadata:
  name: spot-class
spec:
  nodePoolAutoCreation:
    enabled: true
  priorities:
  - machineFamily: n4
    spot: true
    location:
      locationPolicy: ANY   # Spot-first, maximum obtainability

  - machineFamily: n4
    spot: false
    location:
      locationPolicy: BALANCED   # On-demand fallback with zone distribution

```

## Enforce Pod Topology Spread Constraints

To ensure the autoscaler respects zonal distribution when scaling up, use **Pod Topology Spread Constraints (PTS)** with `whenUnsatisfiable: DoNotSchedule`. As noted in [`ca-optimization.md`](https://github.com/google/skills/blob/main/ca-optimization.md) (lines 30-34), this setting forces the autoscaler to provision nodes in specific zones rather than overloading existing nodes.

```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
spec:
  replicas: 6
  template:
    spec:
      topologySpreadConstraints:
      - maxSkew: 1
        topologyKey: "topology.kubernetes.io/zone"
        whenUnsatisfiable: DoNotSchedule   # Required for autoscaler zone awareness

        labelSelector:
          matchLabels:
            app: my-app
      containers:
      - name: app
        image: gcr.io/my-project/my-app:latest

```

## Deploy Capacity Buffers for Warm Standby

The **CapacityBuffer CRD** maintains a pool of warm standby nodes that can be instantly reclaimed by real workloads, reducing cold-start latency. The [`SKILL.md`](https://github.com/google/skills/blob/main/SKILL.md) file (lines 35-38) provides the example manifest [`capacity-buffer-serving.yaml`](https://github.com/google/skills/blob/main/capacity-buffer-serving.yaml).

Apply the buffer:

```bash
kubectl apply -f ./assets/capacity-buffer-serving.yaml

```

Tune the `replicas` or `percentage` fields based on your peak traffic patterns to ensure sufficient headroom without excessive cost.

## Monitor with Visibility Logs

Real-time insight into autoscaler decisions is available through the `container.googleapis.com/cluster-autoscaler-visibility` log stream. The repository includes [`log-autoscaler-events.sh`](https://github.com/google/skills/blob/main/log-autoscaler-events.sh) (referenced in [`SKILL.md`](https://github.com/google/skills/blob/main/SKILL.md), lines 31-33) to tail these events:

```bash
./assets/log-autoscaler-events.sh my-cluster

```

Monitor these logs to identify scale-up failures, stockout events, or unexpected scale-down blocks.

## Eliminate Scale-Down Blockers

Pods with certain attributes prevent node removal, causing idle capacity to persist. Common blockers include bare pods (not owned by controllers), `cluster-autoscaler.kubernetes.io/safe-to-evict: "false"` annotations, local storage, and PDBs that cannot be satisfied.

Run the diagnostic script provided in [`SKILL.md`](https://github.com/google/skills/blob/main/SKILL.md) (lines 17-18):

```bash
./assets/find-scale-down-blockers.sh -n default

```

Review the output and remediate blockers by adding controller owners, removing local storage dependencies, or adjusting PDB configurations.

## Understand Provisioning Architecture and Commitments

Committed Use Discounts (CUDs) are automatically consumed by the autoscaler, but **Reservations** require explicit targeting and suffer a **30-minute cache lag** before the autoscaler recognizes them, as documented in [`ca-optimization.md`](https://github.com/google/skills/blob/main/ca-optimization.md) (lines 46-51).

When a zonal stockout occurs, the entire priority tier enters a **~5-minute global cooldown**, causing pods to fall back to lower-cost tiers. To mitigate this cascade, add intermediate priority tiers or isolate workloads using zonal persistent volumes into dedicated ComputeClasses, as noted in [`SKILL.md`](https://github.com/google/skills/blob/main/SKILL.md) (lines 38-39).

## Summary

- **Use `optimize-utilization`** for cost-driven workloads and `balanced` for latency-sensitive services.
- **Apply `ANY` location policy** for Spot capacity and `BALANCED` for high-availability on-demand workloads.
- **Configure ComputeClasses** with `nodePoolAutoCreation.enabled: true` and explicit fallback priorities.
- **Set `whenUnsatisfiable: DoNotSchedule`** in Pod Topology Spread Constraints to enforce zonal distribution.
- **Deploy CapacityBuffers** to maintain warm standby capacity for sudden traffic spikes.
- **Monitor visibility logs** continuously and **run [`find-scale-down-blockers.sh`](https://github.com/google/skills/blob/main/find-scale-down-blockers.sh)** to eliminate scale-down obstacles.
- **Account for the 30-minute reservation lag** and **5-minute stockout cooldowns** in your capacity planning.

## Frequently Asked Questions

### What is the difference between BALANCED and ANY location policies?

**BALANCED** attempts to distribute nodes evenly across zones for high availability but is best-effort only for node placement and does not balance pods. **ANY** maximizes the probability of obtaining capacity by allowing the autoscaler to provision nodes in any available zone, which is essential for Spot instances or during regional stockouts. Configure `BALANCED` for HA on-demand services and `ANY` for Spot-first or capacity-challenged workloads.

### Why does the autoscaler fail to scale down certain nodes?

The autoscaler cannot remove nodes hosting pods that block scale-down events. Common blockers include bare pods without controllers, pods with `safe-to-evict: "false"` annotations, pods using local storage, and PDBs that would be violated by eviction. Run [`./assets/find-scale-down-blockers.sh`](https://github.com/google/skills/blob/main/./assets/find-scale-down-blockers.sh) to identify specific blockers in your cluster and remediate them by adding ownership metadata or adjusting storage configurations.

### How do CapacityBuffers improve autoscaling performance?

**CapacityBuffers** create a pool of warm standby nodes running low-priority pause pods. When real workloads require capacity, the autoscaler immediately evicts these pause pods and schedules production workloads onto the pre-warmed nodes, eliminating the node provisioning latency (typically 30-60 seconds). This ensures sub-second pod scheduling during traffic spikes while allowing the buffer to be reclaimed during idle periods.

### What causes the 30-minute delay with Reserved instances?

When using **Reservations** (unlike CUDs), the autoscaler maintains a **30-minute cache** before recognizing available reserved capacity. If you create a reservation and expect immediate scaling, the autoscaler will not see the capacity for approximately 30 minutes, potentially causing unnecessary scale-ups to on-demand or Spot tiers. Plan reservation creation well in advance of expected traffic, or rely on CUDs which are consumed automatically without cache delays.