How to Configure Kubernetes Cluster Auto-Scaling: Cluster Autoscaler, HPA, and VPA Explained

Kubernetes cluster auto-scaling operates through two coordinated layers: the Cluster Autoscaler manages worker node provisioning based on pending pods, while the Horizontal Pod Autoscaler adjusts application replica counts based on CPU, memory, or custom metrics.

The litu54/DevOps-Interview-Guide repository documents how to configure Kubernetes cluster auto-scaling through real interview questions from companies like Amazon and SquareOps. Mastering this configuration requires implementing both node-level and pod-level controllers to maintain application availability while optimizing cloud infrastructure costs.

Understanding the Two-Layer Auto-Scaling Architecture

Kubernetes auto-scaling functions at distinct infrastructure and application layers. The cluster-level component (Cluster Autoscaler or Karpenter) monitors for unschedulable pods and removes under-utilized nodes after a grace period. The application-level component (Horizontal Pod Autoscaler) dynamically adjusts replica counts within Deployments, ReplicaSets, or StatefulSets based on real-time metrics collected by the Metrics Server.

Configuring Cluster-Level Auto-Scaling

Deploying the Cluster Autoscaler

To configure the Cluster Autoscaler, deploy it as a Deployment in the kube-system namespace using the official image k8s.gcr.io/autoscaler/cluster-autoscaler. As referenced in SquareOps/DevOps_Engineer.md and Flentas/DevOps_Engineer.md, you must supply cloud-provider-specific flags that define node group boundaries and scaling behavior.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: cluster-autoscaler
  namespace: kube-system
  labels:
    app: cluster-autoscaler
spec:
  replicas: 1
  selector:
    matchLabels:
      app: cluster-autoscaler
  template:
    metadata:
      labels:
        app: cluster-autoscaler
    spec:
      serviceAccountName: cluster-autoscaler
      containers:
      - name: cluster-autoscaler
        image: k8s.gcr.io/autoscaler/cluster-autoscaler:v1.24.0
        command:
        - ./cluster-autoscaler
        - --cloud-provider=aws
        - --nodes=2:10:my-eks-nodegroup   # min:2, max:10

        - --scale-down-enabled=true
        - --scale-down-delay-after-add=10m
        - --balance-similar-node-groups=true
        env:
        - name: AWS_REGION
          value: us-east-1
        resources:
          limits:
            cpu: 100m
            memory: 300Mi
          requests:
            cpu: 100m
            memory: 300Mi

Critical configuration parameters include --nodes to set min/max node counts per group, --scale-down-delay-after-add to prevent premature node removal, and --balance-similar-node-groups to distribute workloads evenly across availability zones.

Alternative: Karpenter for Node Provisioning

SquareOps/DevOps_Engineer.md identifies Karpenter as a modern alternative to Cluster Autoscaler that provisions nodes directly without requiring pre-configured node groups. While Cluster Autoscaler modifies existing Auto Scaling groups, Karpenter creates individual nodes based on specific pod requirements, often resulting in faster scale-out events and better resource bin-packing.

Configuring Application-Level Auto-Scaling

Horizontal Pod Autoscaler (HPA)

The Horizontal Pod Autoscaler automatically scales pod replicas based on observed metrics. According to Amazon/DevOps_Consultant_1.md, this is essential for handling variable workload demand in production environments. Configure it using the autoscaling/v2 API to target specific Deployments or StatefulSets.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: webapp-hpa
  namespace: production
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: webapp
  minReplicas: 2
  maxReplicas: 15
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70

The minReplicas and maxReplicas parameters define scaling boundaries, while averageUtilization sets the target CPU percentage that triggers scaling actions.

Vertical Pod Autoscaler (VPA)

For workloads with unpredictable resource requirements, the Vertical Pod Autoscaler adjusts container CPU and memory requests/limits without changing replica counts. This complements the cluster-level discussions in SquareOps/DevOps_Engineer.md by right-sizing containers before the Cluster Autoscaler decides to add nodes.

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: webapp-vpa
  namespace: production
spec:
  targetRef:
    apiVersion: "apps/v1"
    kind:       Deployment
    name:       webapp
  updatePolicy:
    updateMode: "Auto"
  resourcePolicy:
    containerPolicies:
    - containerName: "*"
      minAllowed:
        cpu: 100m
        memory: 256Mi
      maxAllowed:
        cpu: 2
        memory: 2Gi

How the Components Integrate

The complete auto-scaling pipeline operates sequentially across four stages:

  1. The Metrics Server collects node-level and pod-level resource usage, exposing it via the Kubernetes API.
  2. When CPU or memory exceeds thresholds, the HPA increases pod replicas to handle load.
  3. If the scheduler cannot place new pods due to insufficient capacity, the Cluster Autoscaler detects pending pods and triggers the cloud provider (AWS, GCP, Azure) to provision new VM instances.
  4. After the configurable scale-down grace period (respecting PodDisruptionBudgets), the Autoscaler evaluates node utilization and removes idle nodes to reduce costs.

Summary

  • Configure the Cluster Autoscaler at the node level using the k8s.gcr.io/autoscaler/cluster-autoscaler image with cloud-provider flags like --nodes=min:max:groupname to automatically provision and deprovision worker nodes.
  • Implement HorizontalPodAutoscaler objects using the autoscaling/v2 API to dynamically adjust replica counts for Deployments based on CPU, memory, or custom metrics.
  • Deploy the Metrics Server as a prerequisite component to enable metric-based scaling decisions for both HPA and Cluster Autoscaler.
  • Consider Vertical Pod Autoscaler for right-sizing container resource requests when applications are consistently under- or over-provisioned.
  • Reference specific implementation patterns found in Amazon/DevOps_Consultant_1.md, SquareOps/DevOps_Engineer.md, and Flentas/DevOps_Engineer.md to prepare for DevOps interview scenarios.

Frequently Asked Questions

What is the difference between Cluster Autoscaler and HPA?

The Cluster Autoscaler operates at the infrastructure layer to add or remove worker nodes based on unschedulable pods and node utilization percentages. The Horizontal Pod Autoscaler operates at the application layer to increase or decrease pod replica counts based on resource metrics like CPU and memory consumption. These components work sequentially: HPA scales pods first, and if capacity is insufficient, Cluster Autoscaler scales nodes.

When should I use Karpenter instead of Cluster Autoscaler?

Use Karpenter when you require faster node provisioning without pre-configured node groups, as it creates nodes directly based on aggregated pod requirements. The traditional Cluster Autoscaler requires pre-defined node groups and scales them incrementally, while Karpenter provides more flexible, right-sized node provisioning that can reduce costs and improve bin-packing efficiency.

How do I configure cluster auto-scaling in AWS EKS?

Deploy the Cluster Autoscaler with --cloud-provider=aws and specify node group boundaries using --nodes=2:10:eks-nodegroup (format: min:max:group-name). Ensure the service account has IAM permissions to modify Auto Scaling groups via the cluster-autoscaler policy, and verify the Metrics Server is installed to expose resource utilization data to the Autoscaler.

Can Cluster Autoscaler reduce the number of nodes automatically?

Yes, the Cluster Autoscaler automatically removes under-utilized nodes after a configurable scale-down delay (default 10 minutes). It evaluates node utilization against the --scale-down-utilization-threshold and respects PodDisruptionBudgets when evicting pods to ensure application availability during scale-down operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →