Troubleshooting GKE Nodes Not Scaling Up or Down Correctly: A Complete Diagnostic Guide
The Cluster Autoscaler fails to resize GKE nodes when configuration errors, resource quotas, or scale-down blockers prevent pool modifications.
Troubleshooting GKE nodes not scaling up or down correctly requires systematic analysis of the Cluster Autoscaler, Node Pool Autoscaling, and ComputeClass configurations. According to the google/skills repository, most scaling failures stem from misconfigured autoscaler flags, missing capacity quotas, or blocking workloads that prevent node removal.
Understanding GKE Autoscaling Components
GKE implements node scaling through four integrated mechanisms defined in the source documentation.
- Cluster Autoscaler: Monitors pending pods and adjusts node pool sizes dynamically. Implemented in [
skills/cloud/gke-cluster-autoscaler/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md). - Node Pool Autoscaling: Configures per-pool minimum and maximum node counts through the
--enable-autoscalingflag. - Node Auto Provisioning (NAP): Dynamically creates new node pools when existing pools cannot accommodate workload requirements.
- ComputeClass Auto-Creation: Provisions ephemeral node pools based on
ComputeClassCRD specifications, documented in [skills/cloud/gke-compute-classes/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md).
Workload scaling via Horizontal Pod Autoscaler (HPA) and Vertical Pod Autoscaler (VPA) indirectly drives node scaling decisions, as detailed in [skills/cloud/gke-workload-scaling/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-workload-scaling/SKILL.md).
Verifying Cluster Autoscaler Configuration
Before diagnosing specific scaling failures, confirm that autoscaling is enabled at both the cluster and node pool levels.
# Verify cluster-level autoscaler configuration
gcloud container clusters describe $CLUSTER \
--zone $ZONE \
--format="yaml(autoscaling)"
If disabled, enable autoscaling for Standard clusters by updating the node pool:
gcloud container node-pools update $POOL \
--cluster $CLUSTER \
--zone $ZONE \
--enable-autoscaling \
--min-nodes 1 \
--max-nodes 5
For Autopilot or advanced configurations, ensure Node Auto Provisioning is activated to allow automatic node pool creation beyond predefined pools.
Diagnosing Scale-Up Failures
When nodes fail to scale up, the Cluster Autoscaler logs typically display scale.up.error.out.of.resources messages. Investigate the following constraints:
Check Project Quotas
Insufficient CPU, GPU, or IP address quotas in the target region block node creation.
gcloud compute project-info describe --format="json(quota)"
Validate Machine Type Support
The requested machineType must be compatible with the current GKE version.
gcloud container node-pools describe $POOL \
--cluster $CLUSTER \
--zone $ZONE \
--format="yaml(machineType)"
Inspect ComputeClass Specifications
When using ComputeClass auto-creation, verify that the machineFamily is supported by your GKE version.
kubectl get computeclass $CLASS -o yaml
If capacity remains unavailable, add a node pool in an alternate zone or expand Node Auto Provisioning configuration to include additional machine families.
Resolving Scale-Down Blockers
The Cluster Autoscaler refuses to delete nodes when specific blocking conditions exist. According to [skills/cloud/gke-cluster-autoscaler/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md), the seven primary blockers are:
- Bare pods without controller references.
- Pods annotated with
cluster-autoscaler.kubernetes.io/safe-to-evict: "false". - Pods using
emptyDirvolumes without the safe-to-evict annotation set to"true". - Pod Disruption Budgets with
disruptionsAllowed: 0. - Node pools at their configured
min-nodesfloor. - Nodes annotated with
cluster-autoscaler.kubernetes.io/scale-down-disabled: "true". - Pods with node affinity constraints locking them to specific hostnames.
Use the repository's diagnostic script to enumerate all blockers automatically:
bash $(git rev-parse --show-toplevel)/skills/cloud/gke-cluster-autoscaler/assets/find-scale-down-blockers.sh
Alternatively, inspect the autoscaler logs directly:
kubectl -n kube-system logs deployment/cluster-autoscaler | grep "scale-down"
Tuning Cluster Autoscaler Behavior
Adjust autoscaler parameters to optimize scaling speed and stability. Edit the deployment to modify these flags:
kubectl edit deployment cluster-autoscaler -n kube-system
Key configuration parameters include:
--scale-down-delay-after-add: Increases the waiting period after scale-up before scale-down can occur. Set to10mto prevent oscillation from short-lived pods.--expander: Defines node pool selection strategy. Useleast-wastefor cost-optimal placement.--max-empty-bulk-delete: Limits concurrent node deletions to prevent massive capacity drops.--balance-similar-node-groups: Enables balancing across similar node pools for uniform utilization.
Managing ComputeClass Auto-Creation
ComputeClass creates ephemeral node pools that the autoscaler may delete when empty. Ensure proper configuration to avoid "nodes not scaling down" scenarios:
apiVersion: computeclass.gke.io/v1
kind: ComputeClass
metadata:
name: spot-high-cpu
spec:
nodePoolAutoCreation:
enabled: true
minNodes: 0
maxNodes: 10
Verify that nodePoolAutoCreation.enabled: true is set and avoid mandatory fallback policies like ScaleUpAnyway that force unnecessary pool creation.
Common Scale-Up Patterns and Fixes
| Symptom | Root Cause | Resolution |
|---|---|---|
scale.up.error.out.of.resources |
Regional quota exhaustion or zone stockout | Add pools in alternative zones or enable broader machine families in NAP |
| Pending pods with "Insufficient cpu" | Unsupported machine family | Confirm GKE version supports requested machineFamily (e.g., c3, n2) |
| Continuous node pool creation without deletion | Misconfigured ComputeClass | Adjust ComputeClass spec to allow scale-down or disable ScaleUpAnyway |
Summary
- Enable and verify Cluster Autoscaler settings using
gcloud container clusters describebefore investigating specific failures. - Diagnose scale-up issues by checking project quotas, machine type support, and ComputeClass compatibility.
- Identify scale-down blockers using the
find-scale-down-blockers.shscript fromgoogle/skillsor by analyzing Pod Disruption Budgets and pod annotations. - Tune autoscaler flags like
--scale-down-delay-after-addand--expanderto match your workload patterns. - Configure ComputeClass with
nodePoolAutoCreation.enabled: trueand appropriate min/max nodes to ensure ephemeral pools scale correctly.
Frequently Asked Questions
Why won't the Cluster Autoscaler remove empty nodes from my GKE cluster?
The Cluster Autoscaler preserves nodes when blocking workloads exist. Check for pods with safe-to-evict: "false" annotations, emptyDir volumes without override annotations, or Pod Disruption Budgets that prevent evictions. Run the diagnostic script at skills/cloud/gke-cluster-autoscaler/assets/find-scale-down-blockers.sh to identify specific blockers.
What causes the "out of resources" error during GKE scale-up operations?
This error indicates that Google Cloud cannot provision the requested machine type in the specified zone due to quota limits or stockouts. Verify your project's CPU and accelerator quotas using gcloud compute project-info describe, and ensure the requested machineFamily is supported by your GKE version. Consider enabling Node Auto Provisioning with multiple machine families to improve availability.
How do I enable automatic node pool creation for workload bursts?
Enable Node Auto Provisioning (NAP) at the cluster level or configure ComputeClass resources with nodePoolAutoCreation.enabled: true. NAP automatically creates new node pools when existing pools cannot schedule pending pods, while ComputeClass manages ephemeral pools based on workload specifications defined in [skills/cloud/gke-compute-classes/SKILL.md](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md).
Why do ComputeClass node pools remain active after workloads complete?
ComputeClass pools persist when nodePoolAutoCreation is disabled or when fallback policies like ScaleUpAnyway force retention. Verify your ComputeClass YAML specifies enabled: true under nodePoolAutoCreation and sets appropriate minNodes values. Check the autoscaler logs for scale-down blockers specific to ComputeClass-managed pools.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →