Troubleshooting GKE Nodes Not Scaling Up or Down Correctly: A Complete Diagnostic Guide

The Cluster Autoscaler fails to resize GKE nodes when configuration errors, resource quotas, or scale-down blockers prevent pool modifications.

Troubleshooting GKE nodes not scaling up or down correctly requires systematic analysis of the Cluster Autoscaler, Node Pool Autoscaling, and ComputeClass configurations. According to the google/skills repository, most scaling failures stem from misconfigured autoscaler flags, missing capacity quotas, or blocking workloads that prevent node removal.

Understanding GKE Autoscaling Components

GKE implements node scaling through four integrated mechanisms defined in the source documentation.

Workload scaling via Horizontal Pod Autoscaler (HPA) and Vertical Pod Autoscaler (VPA) indirectly drives node scaling decisions, as detailed in [skills/cloud/gke-workload-scaling/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-workload-scaling/SKILL.md).

Verifying Cluster Autoscaler Configuration

Before diagnosing specific scaling failures, confirm that autoscaling is enabled at both the cluster and node pool levels.


# Verify cluster-level autoscaler configuration

gcloud container clusters describe $CLUSTER \
  --zone $ZONE \
  --format="yaml(autoscaling)"

If disabled, enable autoscaling for Standard clusters by updating the node pool:

gcloud container node-pools update $POOL \
  --cluster $CLUSTER \
  --zone $ZONE \
  --enable-autoscaling \
  --min-nodes 1 \
  --max-nodes 5

For Autopilot or advanced configurations, ensure Node Auto Provisioning is activated to allow automatic node pool creation beyond predefined pools.

Diagnosing Scale-Up Failures

When nodes fail to scale up, the Cluster Autoscaler logs typically display scale.up.error.out.of.resources messages. Investigate the following constraints:

Check Project Quotas

Insufficient CPU, GPU, or IP address quotas in the target region block node creation.

gcloud compute project-info describe --format="json(quota)"

Validate Machine Type Support

The requested machineType must be compatible with the current GKE version.

gcloud container node-pools describe $POOL \
  --cluster $CLUSTER \
  --zone $ZONE \
  --format="yaml(machineType)"

Inspect ComputeClass Specifications

When using ComputeClass auto-creation, verify that the machineFamily is supported by your GKE version.

kubectl get computeclass $CLASS -o yaml

If capacity remains unavailable, add a node pool in an alternate zone or expand Node Auto Provisioning configuration to include additional machine families.

Resolving Scale-Down Blockers

The Cluster Autoscaler refuses to delete nodes when specific blocking conditions exist. According to [skills/cloud/gke-cluster-autoscaler/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md), the seven primary blockers are:

  1. Bare pods without controller references.
  2. Pods annotated with cluster-autoscaler.kubernetes.io/safe-to-evict: "false".
  3. Pods using emptyDir volumes without the safe-to-evict annotation set to "true".
  4. Pod Disruption Budgets with disruptionsAllowed: 0.
  5. Node pools at their configured min-nodes floor.
  6. Nodes annotated with cluster-autoscaler.kubernetes.io/scale-down-disabled: "true".
  7. Pods with node affinity constraints locking them to specific hostnames.

Use the repository's diagnostic script to enumerate all blockers automatically:

bash $(git rev-parse --show-toplevel)/skills/cloud/gke-cluster-autoscaler/assets/find-scale-down-blockers.sh

Alternatively, inspect the autoscaler logs directly:

kubectl -n kube-system logs deployment/cluster-autoscaler | grep "scale-down"

Tuning Cluster Autoscaler Behavior

Adjust autoscaler parameters to optimize scaling speed and stability. Edit the deployment to modify these flags:

kubectl edit deployment cluster-autoscaler -n kube-system

Key configuration parameters include:

  • --scale-down-delay-after-add: Increases the waiting period after scale-up before scale-down can occur. Set to 10m to prevent oscillation from short-lived pods.
  • --expander: Defines node pool selection strategy. Use least-waste for cost-optimal placement.
  • --max-empty-bulk-delete: Limits concurrent node deletions to prevent massive capacity drops.
  • --balance-similar-node-groups: Enables balancing across similar node pools for uniform utilization.

Managing ComputeClass Auto-Creation

ComputeClass creates ephemeral node pools that the autoscaler may delete when empty. Ensure proper configuration to avoid "nodes not scaling down" scenarios:

apiVersion: computeclass.gke.io/v1
kind: ComputeClass
metadata:
  name: spot-high-cpu
spec:
  nodePoolAutoCreation:
    enabled: true
    minNodes: 0
    maxNodes: 10

Verify that nodePoolAutoCreation.enabled: true is set and avoid mandatory fallback policies like ScaleUpAnyway that force unnecessary pool creation.

Common Scale-Up Patterns and Fixes

Symptom Root Cause Resolution
scale.up.error.out.of.resources Regional quota exhaustion or zone stockout Add pools in alternative zones or enable broader machine families in NAP
Pending pods with "Insufficient cpu" Unsupported machine family Confirm GKE version supports requested machineFamily (e.g., c3, n2)
Continuous node pool creation without deletion Misconfigured ComputeClass Adjust ComputeClass spec to allow scale-down or disable ScaleUpAnyway

Summary

  • Enable and verify Cluster Autoscaler settings using gcloud container clusters describe before investigating specific failures.
  • Diagnose scale-up issues by checking project quotas, machine type support, and ComputeClass compatibility.
  • Identify scale-down blockers using the find-scale-down-blockers.sh script from google/skills or by analyzing Pod Disruption Budgets and pod annotations.
  • Tune autoscaler flags like --scale-down-delay-after-add and --expander to match your workload patterns.
  • Configure ComputeClass with nodePoolAutoCreation.enabled: true and appropriate min/max nodes to ensure ephemeral pools scale correctly.

Frequently Asked Questions

Why won't the Cluster Autoscaler remove empty nodes from my GKE cluster?

The Cluster Autoscaler preserves nodes when blocking workloads exist. Check for pods with safe-to-evict: "false" annotations, emptyDir volumes without override annotations, or Pod Disruption Budgets that prevent evictions. Run the diagnostic script at skills/cloud/gke-cluster-autoscaler/assets/find-scale-down-blockers.sh to identify specific blockers.

What causes the "out of resources" error during GKE scale-up operations?

This error indicates that Google Cloud cannot provision the requested machine type in the specified zone due to quota limits or stockouts. Verify your project's CPU and accelerator quotas using gcloud compute project-info describe, and ensure the requested machineFamily is supported by your GKE version. Consider enabling Node Auto Provisioning with multiple machine families to improve availability.

How do I enable automatic node pool creation for workload bursts?

Enable Node Auto Provisioning (NAP) at the cluster level or configure ComputeClass resources with nodePoolAutoCreation.enabled: true. NAP automatically creates new node pools when existing pools cannot schedule pending pods, while ComputeClass manages ephemeral pools based on workload specifications defined in [skills/cloud/gke-compute-classes/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md).

Why do ComputeClass node pools remain active after workloads complete?

ComputeClass pools persist when nodePoolAutoCreation is disabled or when fallback policies like ScaleUpAnyway force retention. Verify your ComputeClass YAML specifies enabled: true under nodePoolAutoCreation and sets appropriate minNodes values. Check the autoscaler logs for scale-down blockers specific to ComputeClass-managed pools.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →