# GKE Capacity Buffer CRD for Pre-warming Nodes: Implementation Guide

> Implement the GKE Capacity Buffer CRD to maintain pre-warmed nodes. Eliminate provisioning delays during workload spikes with this Cluster Autoscaler extension.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-09

---

**The GKE Capacity Buffer CRD is a Cluster Autoscaler extension that maintains pre-warmed node capacity through placeholder pods or idle nodes, eliminating the 60-120 second provisioning delay during sudden workload spikes.**

The `CapacityBuffer` Custom Resource Definition (CRD) in Google Kubernetes Engine (GKE) solves the cold-start problem inherent in cluster autoscaling. According to the [google/skills](https://github.com/google/skills) repository documentation, this CRD enables operators to reserve compute resources ahead of time, ensuring that bursty services scale instantly without waiting for new node provisioning.

## What Is the GKE Capacity Buffer CRD?

The `CapacityBuffer` CRD, defined in the `autoscaling.x-k8s.io/v1beta1` API group, creates a pool of reserved capacity that sits idle until real workloads require it. Unlike standard Cluster Autoscaler behavior—which reacts to pending pods by provisioning new nodes—the CapacityBuffer proactively occupies resources using low-priority placeholder pods or fully provisioned idle nodes.

As documented in [`skills/cloud/gke-cluster-autoscaler/references/ca-capacity-buffers.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/references/ca-capacity-buffers.md), the CRD supports two distinct operational models:

- **Active Capacity**: Low-priority placeholder pods that consume resources and trigger node scaling, then evict instantly when real workloads arrive
- **Standby Capacity**: Fully provisioned idle nodes that remain empty but ready (requires newer GKE versions)

## How Capacity Buffer Pre-warming Works

The pre-warming mechanism relies on the Cluster Autoscaler recognizing occupied capacity that can be reclaimed immediately. When you deploy a CapacityBuffer resource, the controller creates pods based on a referenced `PodTemplate`, forcing the autoscaler to provision nodes matching specific compute requirements.

### Active Capacity vs Standby Capacity Strategies

The `provisioningStrategy` field determines how the buffer maintains warm capacity:

- **`buffer.x-k8s.io/active-capacity`** (default): Creates pause container pods with low priority. These pods occupy node slots and trigger autoscaling, but the scheduler evicts them immediately when higher-priority workloads appear. This strategy appears in the example manifest at [`skills/cloud/gke-cluster-autoscaler/assets/capacity-buffer-serving.yaml`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/assets/capacity-buffer-serving.yaml).

- **`buffer.gke.io/standby-capacity`**: Maintains fully provisioned idle nodes without running placeholder pods. This approach consumes compute resources continuously but eliminates pod scheduling overhead during traffic spikes.

### ComputeClass Targeting with PodTemplates

The CapacityBuffer uses a `PodTemplate` resource to specify exactly which node pool or compute class to warm. This ensures buffers only occupy intended infrastructure, preventing resource fragmentation.

```yaml
apiVersion: v1
kind: PodTemplate
metadata:
  name: serving-buffer-template
  namespace: serving
template:
  spec:
    nodeSelector:
      cloud.google.com/compute-class: serving-class
    containers:
    - name: pause
      image: registry.k8s.io/pause:3.10
      resources:
        requests:
          cpu: "4"
          memory: "16Gi"

```

The `nodeSelector` field targets specific ComputeClasses (such as `serving-class`), ensuring the buffer pre-warms nodes with the appropriate CPU, memory, or accelerator configurations for your workload.

## Configuring Capacity Buffer Sizing

The CRD offers two mutually exclusive sizing modes controlled via the `spec` field: fixed replica counts and dynamic percentage-based scaling.

### Fixed Replica Mode

Use the `replicas` field to maintain a constant number of warm slots. This mode suits predictable traffic patterns where you know exactly how much headroom to maintain.

```yaml
apiVersion: autoscaling.x-k8s.io/v1beta1
kind: CapacityBuffer
metadata:
  name: serving-buffer
  namespace: serving
spec:
  podTemplateRef:
    name: serving-buffer-template
  replicas: 3
  provisioningStrategy: "buffer.x-k8s.io/active-capacity"
  limits:
    cpu: "32"
    memory: "128Gi"

```

In this configuration from the reference files, the CapacityBuffer maintains three warm slots, each requesting 4 CPU and 16Gi memory, capped at 32 CPU and 128Gi total cluster-wide.

### Dynamic Percentage Mode

Replace the `replicas` field with `percentage` and `scalableRef` to automatically adjust buffer size relative to a target Deployment's current scale. This mode benefits workloads using Horizontal Pod Autoscaler (HPA) that need buffer capacity to scale proportionally with running pods.

```yaml
spec:
  percentage: 20
  scalableRef:
    apiVersion: apps/v1
    kind: Deployment
    name: serving-frontend

```

This configuration maintains a buffer equal to 20% of the referenced Deployment's current replica count, ensuring capacity headroom scales with actual workload demand.

## Platform Limitations and Requirements

The GKE Capacity Buffer CRD carries specific constraints that determine deployment eligibility:

- **Autopilot Exclusion**: Not supported on GKE Autopilot clusters using pod-based billing; requires node-based billing models
- **API Version**: Uses `autoscaling.x-k8s.io/v1beta1` (kubectl-only, no Google Cloud Console UI)
- **Resource Overhead**: Active capacity consumes actual cluster resources; standby capacity incurs full node billing while idle
- **Eviction Behavior**: Placeholder pods use priority classes that guarantee eviction when real workloads request resources

## Summary

- The **GKE Capacity Buffer CRD** (`autoscaling.x-k8s.io/v1beta1`) pre-warms nodes using placeholder pods or idle capacity to eliminate autoscaling lag
- **Active capacity** (`buffer.x-k8s.io/active-capacity`) uses evictable pause pods, while **standby capacity** (`buffer.gke.io/standby-capacity`) provisions empty nodes
- Configure via **fixed replicas** for constant headroom or **dynamic percentage** with `scalableRef` for workload-proportional buffering
- Target specific hardware using `PodTemplate` resources with `nodeSelector` matching ComputeClasses like `serving-class`
- Requires standard GKE clusters (not Autopilot) and node-based billing

## Frequently Asked Questions

### What is the difference between active-capacity and standby-capacity provisioning strategies?

**Active capacity** creates low-priority placeholder pods that occupy resources and trigger node autoscaling, evicting instantly when real workloads arrive. **Standby capacity** provisions fully initialized idle nodes without running pods, providing immediate scheduling capacity but consuming resources continuously. Active capacity suits cost-sensitive workloads, while standby capacity serves latency-critical applications requiring instant pod startup.

### Can I use CapacityBuffer on GKE Autopilot clusters?

No. The CapacityBuffer CRD is **not supported on GKE Autopilot** clusters that use pod-based billing. The feature requires node-based billing models available in standard GKE clusters, as Autopilot's abstracted infrastructure model conflicts with the low-level node provisioning control that CapacityBuffer provides.

### How does dynamic sizing work with the scalableRef field?

The `scalableRef` field points to a workload resource (typically a Deployment) that the CapacityBuffer monitors. When you specify `percentage: 20`, the controller calculates 20% of the referenced Deployment's current replica count and automatically adjusts the buffer size to match. As the Deployment scales up or down via HPA, the buffer scales proportionally without manual intervention.

### Which API version does the CapacityBuffer CRD use?

The CapacityBuffer CRD uses the **`autoscaling.x-k8s.io/v1beta1`** API version. This is a kubectl-only resource; you cannot create or manage CapacityBuffer objects through the Google Cloud Console. The beta API status indicates potential future changes as the feature matures within the Cluster Autoscaler ecosystem.