# Troubleshooting GKE Nodes Not Scaling Up or Down Correctly: A Complete Diagnostic Guide

> Troubleshoot GKE nodes not scaling correctly. Diagnose and fix cluster autoscaler issues caused by configuration errors, quotas, or scale-down blockers for efficient resource management.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-13

---

**The Cluster Autoscaler fails to resize GKE nodes when configuration errors, resource quotas, or scale-down blockers prevent pool modifications.**

Troubleshooting GKE nodes not scaling up or down correctly requires systematic analysis of the Cluster Autoscaler, Node Pool Autoscaling, and ComputeClass configurations. According to the `google/skills` repository, most scaling failures stem from misconfigured autoscaler flags, missing capacity quotas, or blocking workloads that prevent node removal.

## Understanding GKE Autoscaling Components

GKE implements node scaling through four integrated mechanisms defined in the source documentation.

- **Cluster Autoscaler**: Monitors pending pods and adjusts node pool sizes dynamically. Implemented in [[`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md)](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md).
- **Node Pool Autoscaling**: Configures per-pool minimum and maximum node counts through the `--enable-autoscaling` flag.
- **Node Auto Provisioning (NAP)**: Dynamically creates new node pools when existing pools cannot accommodate workload requirements.
- **ComputeClass Auto-Creation**: Provisions ephemeral node pools based on `ComputeClass` CRD specifications, documented in [[`skills/cloud/gke-compute-classes/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md)](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md).

Workload scaling via Horizontal Pod Autoscaler (HPA) and Vertical Pod Autoscaler (VPA) indirectly drives node scaling decisions, as detailed in [[`skills/cloud/gke-workload-scaling/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-workload-scaling/SKILL.md)](https://github.com/google/skills/blob/main/skills/cloud/gke-workload-scaling/SKILL.md).

## Verifying Cluster Autoscaler Configuration

Before diagnosing specific scaling failures, confirm that autoscaling is enabled at both the cluster and node pool levels.

```bash

# Verify cluster-level autoscaler configuration

gcloud container clusters describe $CLUSTER \
  --zone $ZONE \
  --format="yaml(autoscaling)"

```

If disabled, enable autoscaling for Standard clusters by updating the node pool:

```bash
gcloud container node-pools update $POOL \
  --cluster $CLUSTER \
  --zone $ZONE \
  --enable-autoscaling \
  --min-nodes 1 \
  --max-nodes 5

```

For Autopilot or advanced configurations, ensure **Node Auto Provisioning** is activated to allow automatic node pool creation beyond predefined pools.

## Diagnosing Scale-Up Failures

When nodes fail to scale up, the Cluster Autoscaler logs typically display `scale.up.error.out.of.resources` messages. Investigate the following constraints:

### Check Project Quotas

Insufficient CPU, GPU, or IP address quotas in the target region block node creation.

```bash
gcloud compute project-info describe --format="json(quota)"

```

### Validate Machine Type Support

The requested `machineType` must be compatible with the current GKE version.

```bash
gcloud container node-pools describe $POOL \
  --cluster $CLUSTER \
  --zone $ZONE \
  --format="yaml(machineType)"

```

### Inspect ComputeClass Specifications

When using ComputeClass auto-creation, verify that the `machineFamily` is supported by your GKE version.

```bash
kubectl get computeclass $CLASS -o yaml

```

If capacity remains unavailable, add a node pool in an alternate zone or expand Node Auto Provisioning configuration to include additional machine families.

## Resolving Scale-Down Blockers

The Cluster Autoscaler refuses to delete nodes when specific blocking conditions exist. According to [[`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md)](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md), the seven primary blockers are:

1. **Bare pods** without controller references.
2. Pods annotated with `cluster-autoscaler.kubernetes.io/safe-to-evict: "false"`.
3. Pods using `emptyDir` volumes without the safe-to-evict annotation set to `"true"`.
4. **Pod Disruption Budgets** with `disruptionsAllowed: 0`.
5. Node pools at their configured `min-nodes` floor.
6. Nodes annotated with `cluster-autoscaler.kubernetes.io/scale-down-disabled: "true"`.
7. Pods with node affinity constraints locking them to specific hostnames.

Use the repository's diagnostic script to enumerate all blockers automatically:

```bash
bash $(git rev-parse --show-toplevel)/skills/cloud/gke-cluster-autoscaler/assets/find-scale-down-blockers.sh

```

Alternatively, inspect the autoscaler logs directly:

```bash
kubectl -n kube-system logs deployment/cluster-autoscaler | grep "scale-down"

```

## Tuning Cluster Autoscaler Behavior

Adjust autoscaler parameters to optimize scaling speed and stability. Edit the deployment to modify these flags:

```bash
kubectl edit deployment cluster-autoscaler -n kube-system

```

Key configuration parameters include:

- **`--scale-down-delay-after-add`**: Increases the waiting period after scale-up before scale-down can occur. Set to `10m` to prevent oscillation from short-lived pods.
- **`--expander`**: Defines node pool selection strategy. Use `least-waste` for cost-optimal placement.
- **`--max-empty-bulk-delete`**: Limits concurrent node deletions to prevent massive capacity drops.
- **`--balance-similar-node-groups`**: Enables balancing across similar node pools for uniform utilization.

## Managing ComputeClass Auto-Creation

ComputeClass creates ephemeral node pools that the autoscaler may delete when empty. Ensure proper configuration to avoid "nodes not scaling down" scenarios:

```yaml
apiVersion: computeclass.gke.io/v1
kind: ComputeClass
metadata:
  name: spot-high-cpu
spec:
  nodePoolAutoCreation:
    enabled: true
    minNodes: 0
    maxNodes: 10

```

Verify that `nodePoolAutoCreation.enabled: true` is set and avoid mandatory fallback policies like `ScaleUpAnyway` that force unnecessary pool creation.

## Common Scale-Up Patterns and Fixes

| Symptom | Root Cause | Resolution |
|---------|------------|------------|
| `scale.up.error.out.of.resources` | Regional quota exhaustion or zone stockout | Add pools in alternative zones or enable broader machine families in NAP |
| Pending pods with "Insufficient cpu" | Unsupported machine family | Confirm GKE version supports requested `machineFamily` (e.g., `c3`, `n2`) |
| Continuous node pool creation without deletion | Misconfigured ComputeClass | Adjust `ComputeClass` spec to allow scale-down or disable `ScaleUpAnyway` |

## Summary

- **Enable and verify** Cluster Autoscaler settings using `gcloud container clusters describe` before investigating specific failures.
- **Diagnose scale-up issues** by checking project quotas, machine type support, and ComputeClass compatibility.
- **Identify scale-down blockers** using the [`find-scale-down-blockers.sh`](https://github.com/google/skills/blob/main/find-scale-down-blockers.sh) script from `google/skills` or by analyzing Pod Disruption Budgets and pod annotations.
- **Tune autoscaler flags** like `--scale-down-delay-after-add` and `--expander` to match your workload patterns.
- **Configure ComputeClass** with `nodePoolAutoCreation.enabled: true` and appropriate min/max nodes to ensure ephemeral pools scale correctly.

## Frequently Asked Questions

### Why won't the Cluster Autoscaler remove empty nodes from my GKE cluster?

The Cluster Autoscaler preserves nodes when blocking workloads exist. Check for pods with `safe-to-evict: "false"` annotations, `emptyDir` volumes without override annotations, or Pod Disruption Budgets that prevent evictions. Run the diagnostic script at [`skills/cloud/gke-cluster-autoscaler/assets/find-scale-down-blockers.sh`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/assets/find-scale-down-blockers.sh) to identify specific blockers.

### What causes the "out of resources" error during GKE scale-up operations?

This error indicates that Google Cloud cannot provision the requested machine type in the specified zone due to quota limits or stockouts. Verify your project's CPU and accelerator quotas using `gcloud compute project-info describe`, and ensure the requested `machineFamily` is supported by your GKE version. Consider enabling Node Auto Provisioning with multiple machine families to improve availability.

### How do I enable automatic node pool creation for workload bursts?

Enable **Node Auto Provisioning (NAP)** at the cluster level or configure **ComputeClass** resources with `nodePoolAutoCreation.enabled: true`. NAP automatically creates new node pools when existing pools cannot schedule pending pods, while ComputeClass manages ephemeral pools based on workload specifications defined in [[`skills/cloud/gke-compute-classes/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md)](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md).

### Why do ComputeClass node pools remain active after workloads complete?

ComputeClass pools persist when `nodePoolAutoCreation` is disabled or when fallback policies like `ScaleUpAnyway` force retention. Verify your ComputeClass YAML specifies `enabled: true` under `nodePoolAutoCreation` and sets appropriate `minNodes` values. Check the autoscaler logs for scale-down blockers specific to ComputeClass-managed pools.