How to Configure GKE Cluster Autoscaler for Optimal Performance: A Complete Guide

To configure GKE Cluster Autoscaler for optimal performance, combine the optimize-utilization autoscaling profile, strategic location policies (ANY for Spot, BALANCED for HA), ComputeClasses with fallback priorities, and strict Pod Topology Spread Constraints while eliminating scale-down blockers and monitoring visibility logs.

The GKE Cluster Autoscaler dynamically adjusts node counts to match workload demand, but achieving the right balance between cost, latency, and reliability requires precise configuration. According to the google/skills repository, modern GKE clusters (v1.33.3+) leverage ComputeClasses with nodePoolAutoCreation.enabled: true, enabling zero-size node pools that scale to zero when idle, fundamentally changing how you configure GKE Cluster Autoscaler for optimal performance compared to legacy Node Auto Provisioning.

Select the Appropriate Autoscaling Profile

The autoscaling profile controls how aggressively the autoscaler packs workloads and removes idle nodes. In skills/cloud/gke-cluster-autoscaler/SKILL.md (lines 26-30), two profiles are defined:

  • optimize-utilization – Prioritizes node consolidation and rapid scale-down. Use this for cost-driven batch workloads or development environments where minimizing idle capacity is paramount.
  • balanced – Maintains buffers for faster scale-up. Use this for latency-sensitive production services where pod startup time is critical.

Switch profiles using gcloud:

gcloud container clusters update my-cluster \
  --autoscaling-profile=optimize-utilization

Configure Location Policies for Zone Distribution

Location policies determine where new nodes are provisioned across zones. As documented in skills/cloud/gke-cluster-autoscaler/references/ca-optimization.md (lines 14-18):

  • BALANCED – Best-effort distribution across zones for high-availability on-demand workloads. Note that this balances node placement, not pod placement.
  • ANY – Maximizes capacity acquisition across zones, essential for Spot instances or when regional stockouts occur.

For Spot-first ComputeClasses, set locationPolicy: ANY to maximize obtainability. For critical on-demand workloads requiring zone distribution, use BALANCED.

Implement ComputeClasses with Fallback Priorities

Modern GKE versions prefer ComputeClasses over classic node auto-provisioning. The ca-provisioning.md reference (lines 29-35) specifies enabling nodePoolAutoCreation.enabled: true to allow zero-size pools that the autoscaler can instantiate on demand.

Define fallback priorities to handle stockouts. According to ca-optimization.md (lines 21-27), configure location.locationPolicy within the ComputeClass spec:

apiVersion: autoscaling.gke.io/v1beta1
kind: ComputeClass
metadata:
  name: spot-class
spec:
  nodePoolAutoCreation:
    enabled: true
  priorities:
  - machineFamily: n4
    spot: true
    location:
      locationPolicy: ANY   # Spot-first, maximum obtainability

  - machineFamily: n4
    spot: false
    location:
      locationPolicy: BALANCED   # On-demand fallback with zone distribution

Enforce Pod Topology Spread Constraints

To ensure the autoscaler respects zonal distribution when scaling up, use Pod Topology Spread Constraints (PTS) with whenUnsatisfiable: DoNotSchedule. As noted in ca-optimization.md (lines 30-34), this setting forces the autoscaler to provision nodes in specific zones rather than overloading existing nodes.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
spec:
  replicas: 6
  template:
    spec:
      topologySpreadConstraints:
      - maxSkew: 1
        topologyKey: "topology.kubernetes.io/zone"
        whenUnsatisfiable: DoNotSchedule   # Required for autoscaler zone awareness

        labelSelector:
          matchLabels:
            app: my-app
      containers:
      - name: app
        image: gcr.io/my-project/my-app:latest

Deploy Capacity Buffers for Warm Standby

The CapacityBuffer CRD maintains a pool of warm standby nodes that can be instantly reclaimed by real workloads, reducing cold-start latency. The SKILL.md file (lines 35-38) provides the example manifest capacity-buffer-serving.yaml.

Apply the buffer:

kubectl apply -f ./assets/capacity-buffer-serving.yaml

Tune the replicas or percentage fields based on your peak traffic patterns to ensure sufficient headroom without excessive cost.

Monitor with Visibility Logs

Real-time insight into autoscaler decisions is available through the container.googleapis.com/cluster-autoscaler-visibility log stream. The repository includes log-autoscaler-events.sh (referenced in SKILL.md, lines 31-33) to tail these events:

./assets/log-autoscaler-events.sh my-cluster

Monitor these logs to identify scale-up failures, stockout events, or unexpected scale-down blocks.

Eliminate Scale-Down Blockers

Pods with certain attributes prevent node removal, causing idle capacity to persist. Common blockers include bare pods (not owned by controllers), cluster-autoscaler.kubernetes.io/safe-to-evict: "false" annotations, local storage, and PDBs that cannot be satisfied.

Run the diagnostic script provided in SKILL.md (lines 17-18):

./assets/find-scale-down-blockers.sh -n default

Review the output and remediate blockers by adding controller owners, removing local storage dependencies, or adjusting PDB configurations.

Understand Provisioning Architecture and Commitments

Committed Use Discounts (CUDs) are automatically consumed by the autoscaler, but Reservations require explicit targeting and suffer a 30-minute cache lag before the autoscaler recognizes them, as documented in ca-optimization.md (lines 46-51).

When a zonal stockout occurs, the entire priority tier enters a ~5-minute global cooldown, causing pods to fall back to lower-cost tiers. To mitigate this cascade, add intermediate priority tiers or isolate workloads using zonal persistent volumes into dedicated ComputeClasses, as noted in SKILL.md (lines 38-39).

Summary

  • Use optimize-utilization for cost-driven workloads and balanced for latency-sensitive services.
  • Apply ANY location policy for Spot capacity and BALANCED for high-availability on-demand workloads.
  • Configure ComputeClasses with nodePoolAutoCreation.enabled: true and explicit fallback priorities.
  • Set whenUnsatisfiable: DoNotSchedule in Pod Topology Spread Constraints to enforce zonal distribution.
  • Deploy CapacityBuffers to maintain warm standby capacity for sudden traffic spikes.
  • Monitor visibility logs continuously and run find-scale-down-blockers.sh to eliminate scale-down obstacles.
  • Account for the 30-minute reservation lag and 5-minute stockout cooldowns in your capacity planning.

Frequently Asked Questions

What is the difference between BALANCED and ANY location policies?

BALANCED attempts to distribute nodes evenly across zones for high availability but is best-effort only for node placement and does not balance pods. ANY maximizes the probability of obtaining capacity by allowing the autoscaler to provision nodes in any available zone, which is essential for Spot instances or during regional stockouts. Configure BALANCED for HA on-demand services and ANY for Spot-first or capacity-challenged workloads.

Why does the autoscaler fail to scale down certain nodes?

The autoscaler cannot remove nodes hosting pods that block scale-down events. Common blockers include bare pods without controllers, pods with safe-to-evict: "false" annotations, pods using local storage, and PDBs that would be violated by eviction. Run ./assets/find-scale-down-blockers.sh to identify specific blockers in your cluster and remediate them by adding ownership metadata or adjusting storage configurations.

How do CapacityBuffers improve autoscaling performance?

CapacityBuffers create a pool of warm standby nodes running low-priority pause pods. When real workloads require capacity, the autoscaler immediately evicts these pause pods and schedules production workloads onto the pre-warmed nodes, eliminating the node provisioning latency (typically 30-60 seconds). This ensures sub-second pod scheduling during traffic spikes while allowing the buffer to be reclaimed during idle periods.

What causes the 30-minute delay with Reserved instances?

When using Reservations (unlike CUDs), the autoscaler maintains a 30-minute cache before recognizing available reserved capacity. If you create a reservation and expect immediate scaling, the autoscaler will not see the capacity for approximately 30 minutes, potentially causing unnecessary scale-ups to on-demand or Spot tiers. Plan reservation creation well in advance of expected traffic, or rely on CUDs which are consumed automatically without cache delays.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →