# How to Configure Kubernetes Cluster Auto-Scaling: Cluster Autoscaler, HPA, and VPA Explained

> Learn how to configure Kubernetes cluster auto-scaling. Understand Cluster Autoscaler, HPA, and VPA to efficiently manage nodes and pods for optimal resource utilization.

- Repository: [Anil Kumar/DevOps-Interview-Guide](https://github.com/litu54/DevOps-Interview-Guide)
- Tags: how-to-guide
- Published: 2026-08-10

---

**Kubernetes cluster auto-scaling operates through two coordinated layers: the Cluster Autoscaler manages worker node provisioning based on pending pods, while the Horizontal Pod Autoscaler adjusts application replica counts based on CPU, memory, or custom metrics.**

The `litu54/DevOps-Interview-Guide` repository documents how to configure Kubernetes cluster auto-scaling through real interview questions from companies like Amazon and SquareOps. Mastering this configuration requires implementing both node-level and pod-level controllers to maintain application availability while optimizing cloud infrastructure costs.

## Understanding the Two-Layer Auto-Scaling Architecture

Kubernetes auto-scaling functions at distinct infrastructure and application layers. The **cluster-level** component (Cluster Autoscaler or Karpenter) monitors for unschedulable pods and removes under-utilized nodes after a grace period. The **application-level** component (Horizontal Pod Autoscaler) dynamically adjusts replica counts within Deployments, ReplicaSets, or StatefulSets based on real-time metrics collected by the Metrics Server.

## Configuring Cluster-Level Auto-Scaling

### Deploying the Cluster Autoscaler

To configure the Cluster Autoscaler, deploy it as a Deployment in the `kube-system` namespace using the official image `k8s.gcr.io/autoscaler/cluster-autoscaler`. As referenced in [`SquareOps/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/SquareOps/DevOps_Engineer.md) and [`Flentas/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Flentas/DevOps_Engineer.md), you must supply cloud-provider-specific flags that define node group boundaries and scaling behavior.

```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: cluster-autoscaler
  namespace: kube-system
  labels:
    app: cluster-autoscaler
spec:
  replicas: 1
  selector:
    matchLabels:
      app: cluster-autoscaler
  template:
    metadata:
      labels:
        app: cluster-autoscaler
    spec:
      serviceAccountName: cluster-autoscaler
      containers:
      - name: cluster-autoscaler
        image: k8s.gcr.io/autoscaler/cluster-autoscaler:v1.24.0
        command:
        - ./cluster-autoscaler
        - --cloud-provider=aws
        - --nodes=2:10:my-eks-nodegroup   # min:2, max:10

        - --scale-down-enabled=true
        - --scale-down-delay-after-add=10m
        - --balance-similar-node-groups=true
        env:
        - name: AWS_REGION
          value: us-east-1
        resources:
          limits:
            cpu: 100m
            memory: 300Mi
          requests:
            cpu: 100m
            memory: 300Mi

```

Critical configuration parameters include `--nodes` to set min/max node counts per group, `--scale-down-delay-after-add` to prevent premature node removal, and `--balance-similar-node-groups` to distribute workloads evenly across availability zones.

### Alternative: Karpenter for Node Provisioning

[`SquareOps/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/SquareOps/DevOps_Engineer.md) identifies **Karpenter** as a modern alternative to Cluster Autoscaler that provisions nodes directly without requiring pre-configured node groups. While Cluster Autoscaler modifies existing Auto Scaling groups, Karpenter creates individual nodes based on specific pod requirements, often resulting in faster scale-out events and better resource bin-packing.

## Configuring Application-Level Auto-Scaling

### Horizontal Pod Autoscaler (HPA)

The **Horizontal Pod Autoscaler** automatically scales pod replicas based on observed metrics. According to [`Amazon/DevOps_Consultant_1.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Amazon/DevOps_Consultant_1.md), this is essential for handling variable workload demand in production environments. Configure it using the `autoscaling/v2` API to target specific Deployments or StatefulSets.

```yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: webapp-hpa
  namespace: production
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: webapp
  minReplicas: 2
  maxReplicas: 15
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70

```

The `minReplicas` and `maxReplicas` parameters define scaling boundaries, while `averageUtilization` sets the target CPU percentage that triggers scaling actions.

### Vertical Pod Autoscaler (VPA)

For workloads with unpredictable resource requirements, the **Vertical Pod Autoscaler** adjusts container CPU and memory requests/limits without changing replica counts. This complements the cluster-level discussions in [`SquareOps/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/SquareOps/DevOps_Engineer.md) by right-sizing containers before the Cluster Autoscaler decides to add nodes.

```yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: webapp-vpa
  namespace: production
spec:
  targetRef:
    apiVersion: "apps/v1"
    kind:       Deployment
    name:       webapp
  updatePolicy:
    updateMode: "Auto"
  resourcePolicy:
    containerPolicies:
    - containerName: "*"
      minAllowed:
        cpu: 100m
        memory: 256Mi
      maxAllowed:
        cpu: 2
        memory: 2Gi

```

## How the Components Integrate

The complete auto-scaling pipeline operates sequentially across four stages:

1. The **Metrics Server** collects node-level and pod-level resource usage, exposing it via the Kubernetes API.
2. When CPU or memory exceeds thresholds, the **HPA** increases pod replicas to handle load.
3. If the scheduler cannot place new pods due to insufficient capacity, the **Cluster Autoscaler** detects pending pods and triggers the cloud provider (AWS, GCP, Azure) to provision new VM instances.
4. After the configurable scale-down grace period (respecting **PodDisruptionBudgets**), the Autoscaler evaluates node utilization and removes idle nodes to reduce costs.

## Summary

- Configure the **Cluster Autoscaler** at the node level using the `k8s.gcr.io/autoscaler/cluster-autoscaler` image with cloud-provider flags like `--nodes=min:max:groupname` to automatically provision and deprovision worker nodes.
- Implement **HorizontalPodAutoscaler** objects using the `autoscaling/v2` API to dynamically adjust replica counts for Deployments based on CPU, memory, or custom metrics.
- Deploy the **Metrics Server** as a prerequisite component to enable metric-based scaling decisions for both HPA and Cluster Autoscaler.
- Consider **Vertical Pod Autoscaler** for right-sizing container resource requests when applications are consistently under- or over-provisioned.
- Reference specific implementation patterns found in [`Amazon/DevOps_Consultant_1.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Amazon/DevOps_Consultant_1.md), [`SquareOps/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/SquareOps/DevOps_Engineer.md), and [`Flentas/DevOps_Engineer.md`](https://github.com/litu54/DevOps-Interview-Guide/blob/main/Flentas/DevOps_Engineer.md) to prepare for DevOps interview scenarios.

## Frequently Asked Questions

### What is the difference between Cluster Autoscaler and HPA?

The **Cluster Autoscaler** operates at the infrastructure layer to add or remove worker nodes based on unschedulable pods and node utilization percentages. The **Horizontal Pod Autoscaler** operates at the application layer to increase or decrease pod replica counts based on resource metrics like CPU and memory consumption. These components work sequentially: HPA scales pods first, and if capacity is insufficient, Cluster Autoscaler scales nodes.

### When should I use Karpenter instead of Cluster Autoscaler?

Use **Karpenter** when you require faster node provisioning without pre-configured node groups, as it creates nodes directly based on aggregated pod requirements. The traditional **Cluster Autoscaler** requires pre-defined node groups and scales them incrementally, while Karpenter provides more flexible, right-sized node provisioning that can reduce costs and improve bin-packing efficiency.

### How do I configure cluster auto-scaling in AWS EKS?

Deploy the Cluster Autoscaler with `--cloud-provider=aws` and specify node group boundaries using `--nodes=2:10:eks-nodegroup` (format: min:max:group-name). Ensure the service account has IAM permissions to modify Auto Scaling groups via the `cluster-autoscaler` policy, and verify the **Metrics Server** is installed to expose resource utilization data to the Autoscaler.

### Can Cluster Autoscaler reduce the number of nodes automatically?

Yes, the Cluster Autoscaler automatically removes under-utilized nodes after a configurable **scale-down delay** (default 10 minutes). It evaluates node utilization against the `--scale-down-utilization-threshold` and respects **PodDisruptionBudgets** when evicting pods to ensure application availability during scale-down operations.