How to Fix GKE Cluster Autoscaler Not Scaling Down Due to Scale-Down Blockers
To fix GKE Cluster Autoscaler scale-down failures, identify and remove the eight categories of scale-down blockers—including pods marked safe-to-evict: "false", bare pods without controllers, local-storage volumes, restrictive PodDisruptionBudgets, and node-level constraints—using the diagnostic scripts in the google/skills repository.
The GKE Cluster Autoscaler only removes nodes when every pod on the node is evictable. When blockers exist, the autoscaler skips the node during consolidation, leaving infrastructure running and costs accumulating. The gke-cluster-autoscaler skill in the google/skills repository provides a canonical enumeration of these blockers and automation to detect them in live clusters.
How the GKE Cluster Autoscaler Evaluates Scale-Down Eligibility
The autoscaler makes scale-down decisions based on visibility logs and pod-level constraints. It emits detailed events to Cloud Logging under the log ID container.googleapis.com/cluster-autoscaler-visibility according to the skill documentation in skills/cloud/gke-cluster-autoscaler/SKILL.md.
For each node, the autoscaler checks whether any pod is non-evictable. If a single non-evictable pod exists, the node is excluded from the candidate pool. The autoscaler recognizes eight specific blocker categories:
cluster-autoscaler.kubernetes.io/safe-to-evict: "false"annotation – Explicitly prevents eviction regardless of other factors.- Bare pods – Pods lacking
ownerReferences(no Deployment, ReplicaSet, or Job controller) cannot be safely rescheduled. - Local-storage pods – Pods using
emptyDirorhostPathvolumes without thesafe-to-evict: "true"annotation risk data loss during eviction. - Restrictive PodDisruptionBudgets – PDBs with
disruptionsAllowed: 0block voluntary evictions. - Node-pool minimum size – Nodes in pools already at their
--min-nodesfloor cannot be removed even if empty. - Scale-down-disabled annotation – Nodes marked with
cluster-autoscaler.kubernetes.io/scale-down-disabled: "true"are permanently ineligible. - Hostname scheduling constraints – Pods with
kubernetes.io/hostnamenodeSelectors or affinity rules pin themselves to specific nodes. - Unmarked kube-system pods – Non-DaemonSet pods in
kube-systemwithoutsafe-to-evict: "true"block consolidation.
Identifying Scale-Down Blockers in Your Cluster
Run the blocker enumeration script to audit your cluster state. The script assets/find-scale-down-blockers.sh queries the Kubernetes API server and categorizes every blocking condition found.
# Scan all namespaces for blockers
./assets/find-scale-down-blockers.sh
# Limit scan to a specific namespace
./assets/find-scale-down-blockers.sh -n production
For additional context, tail the autoscaler visibility logs using the helper script:
./assets/log-autoscaler-events.sh my-gke-cluster
Look for "NOSCALEDOWN" entries in the logs to confirm which specific blockers the autoscaler detected during its last evaluation cycle.
Remediating Scale-Down Blockers
Remove Safe-to-Evict Restrictions
If a pod explicitly blocks eviction, remove the annotation:
kubectl annotate pod <pod-name> \
cluster-autoscaler.kubernetes.io/safe-to-evict-
Convert Bare Pods to Managed Workloads
Bare pods lack controllers and prevent node removal. Wrap them in Deployments:
cat <<EOF | kubectl apply -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-app
spec:
replicas: 1
selector:
matchLabels:
app: my-app
template:
metadata:
labels:
app: my-app
spec:
containers:
- name: app
image: gcr.io/my-project/my-app:latest
EOF
# Remove the original bare pod
kubectl delete pod <bare-pod-name>
Fix Local-Storage Pods
For pods using emptyDir or hostPath, either mark data as disposable or migrate to network storage:
# Option A: Allow eviction if data loss is acceptable
kubectl annotate pod <pod-name> \
cluster-autoscaler.kubernetes.io/safe-to-evict=true
# Option B: Migrate to PersistentVolumeClaim
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: my-pvc
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: standard
resources:
requests:
storage: 10Gi
EOF
# Update the pod spec to mount the PVC instead of emptyDir/hostPath
Adjust PodDisruptionBudgets
Increase the disruption allowance to permit evictions:
kubectl edit pdb <pdb-name>
# Modify spec to allow at least one disruption:
# disruptionsAllowed: 1
Lower Node-Pool Minimum Size
If the node pool is at its floor, reduce the minimum:
gcloud container node-pools update <pool-name> \
--cluster=<cluster-name> \
--region=<region> \
--min-nodes=0
Remove Node-Level Scale-Down Blocks
Clear disabled annotations from nodes:
kubectl annotate node <node-name> \
cluster-autoscaler.kubernetes.io/scale-down-disabled-
Eliminate Hostname Pinning
Remove hardcoded hostname selectors or affinity rules:
kubectl patch pod <pod-name> -p '{"spec":{"nodeSelector":null}}'
# Alternatively, edit the deployment to remove kubernetes.io/hostname requirements
Troubleshooting Diagnostic Pitfalls
RBAC Permissions for ConfigMap Access
The find-scale-down-blockers.sh script attempts to read the cluster-autoscaler-status ConfigMap. If RBAC denies this access, the script falls back to gcloud calls requiring roles/container.clusterViewer. Ensure your service account has sufficient permissions to avoid incomplete scans.
Stale Reservation Cache
New Compute Engine reservations may not appear in the autoscaler's cache for approximately 30 minutes. Triggering a scale-up during this window can cause temporary resource errors that subsequently block scale-down. Wait for the cache refresh before evaluating scale-down behavior.
DaemonSet Behavior
The autoscaler ignores DaemonSet pods when calculating scale-down eligibility. Only non-DaemonSet pods in kube-system require the safe-to-evict annotation to prevent blocking.
Summary
- The GKE Cluster Autoscaler only removes nodes when all pods are evictable, as defined in
skills/cloud/gke-cluster-autoscaler/SKILL.md. - Use
assets/find-scale-down-blockers.shto enumerate the eight categories of blockers: safe-to-evict annotations, bare pods, local storage, PDBs, min-node limits, node annotations, hostname affinity, and unmarked kube-system pods. - Remediate by removing annotations, converting bare pods to Deployments, migrating local storage, adjusting PDB disruption allowances, lowering node-pool floors, and eliminating hostname constraints.
- Verify fixes by re-running the enumeration script and monitoring
container.googleapis.com/cluster-autoscaler-visibilitylogs for "NOSCALEDOWN" clearance.
Frequently Asked Questions
Why does GKE Cluster Autoscaler keep skipping nodes during scale-down?
The autoscaler skips nodes containing non-evictable pods. According to the source code analysis in the google/skills repository, if any pod on a node meets one of the eight blocker conditions—such as having safe-to-evict: "false" or lacking an ownerReference—the entire node is excluded from the consolidation pool until the blocker is removed.
What does the safe-to-evict annotation do?
The cluster-autoscaler.kubernetes.io/safe-to-evict annotation signals whether the autoscaler can safely delete a pod during scale-down. When set to "false", the autoscaler treats the pod as immovable and will never delete its host node. Setting it to "true" permits eviction even for pods with local storage, though this risks data loss if the storage is not backed by a persistent volume.
How do bare pods prevent cluster autoscaling?
Bare pods lack ownerReferences, meaning no controller manages their lifecycle. The autoscaler cannot guarantee these pods will be rescheduled elsewhere if the node is deleted, so it treats them as scale-down blockers. Converting bare pods to Deployments or Jobs provides the necessary controller reference, allowing the autoscaler to evict and reschedule the workload safely.
Can PodDisruptionBudgets stop the Cluster Autoscaler from working?
Yes. When a PDB has disruptionsAllowed: 0, it blocks all voluntary evictions, including those initiated by the Cluster Autoscaler. The autoscaler respects PDBs to maintain application availability. To enable scale-down, either increase disruptionsAllowed or temporarily remove the PDB during maintenance windows, ensuring your application can tolerate the disruption.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →