# Optimizing GKE Costs Using Spot VMs: ComputeClass Patterns and Fallback Strategies

> Slash GKE costs up to 90% with Spot VMs and ComputeClasses. Discover fallback strategies and automatic re-migration for cost-effective, resilient workloads that tolerate interruption.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: performance
- Published: 2026-06-09

---

**You can reduce GKE node costs by 60–90% using Spot VMs orchestrated through ComputeClasses that provide on-demand fallback and automatic re-migration, provided your workloads tolerate interruption and shutdown within 30 seconds.**

The `google/skills` repository provides authoritative patterns for optimizing GKE costs using Spot VMs through declarative node management. By leveraging the **ComputeClass** API instead of hand-tuned node pools, platform teams can automate Spot provisioning, fallback, and workload migration while cutting cluster spend by up to roughly 45%.

## Cost Impact of Spot VMs on GKE Clusters

According to [`skills/cloud/gke-basics/references/gke-cost.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/references/gke-cost.md), Spot VMs consume excess Google Cloud data center capacity and cost 60–90% less than on-demand nodes. The same guide notes that moving approximately 30% of node capacity to Spot can lower the overall GKE bill by roughly 45% without violating service-level objectives.

## Declaring Spot VMs with the ComputeClass API

The [`skills/cloud/gke-basics/references/gke-compute-classes.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/references/gke-compute-classes.md) reference demonstrates that the recommended abstraction is the **ComputeClass** resource (`compute.cnrm.cloud.google.com/v1beta1`). Rather than manually creating Spot node pools, you set `spot: true` inside the ComputeClass spec and optionally enable `activeMigration` and `onDemandFallback`.

```yaml
apiVersion: compute.cnrm.cloud.google.com/v1beta1
kind: ComputeClass
metadata:
  name: spot-fallback
spec:
  machineType: n1-standard-4
  spot: true
  activeMigration: true
  onDemandFallback: true

```

- **`spot: true`** requests Spot VM capacity.
- **`onDemandFallback: true`** allows scheduling on standard nodes when Spot is unavailable.
- **`activeMigration: true`** automatically migrates the pod back to Spot when excess capacity returns.

## Linking Pods to a Spot ComputeClass

To consume the class, reference it in the pod spec via `nodeSelector`. The [`skills/cloud/gke-basics/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/SKILL.md) documentation maps ComputeClasses to workloads using this label pattern.

```yaml
apiVersion: v1
kind: Pod
metadata:
  name: batch-worker
spec:
  containers:
  - name: worker
    image: gcr.io/project/batch-worker:latest
  nodeSelector:
    computeclass: spot-fallback
  terminationGracePeriodSeconds: 25

```

The `computeclass: spot-fallback` label binds the pod to the Spot-first policy. The `terminationGracePeriodSeconds` value must stay below the 30-second eviction notice window.

## On-Demand Fallback for Critical Availability

As documented in [`skills/cloud/gke-basics/references/gke-compute-classes.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/references/gke-compute-classes.md), the fallback pattern guarantees that pods remain schedulable even when Spot capacity is exhausted across a region. The scheduler attempts Spot first; if capacity is absent, it immediately provisions an on-demand node. Because the ComputeClass manages both tiers, operators do not need to maintain separate fallback node pools manually.

## Batch and HPC Workloads with Checkpointing

Fault-tolerant batch jobs and high-performance computing tasks are ideal for Spot because they can persist progress to Cloud Storage and resume after eviction. The [`skills/cloud/gke-basics/references/gke-batch-hpc.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/references/gke-batch-hpc.md) reference includes a concrete pattern that uses Spot-first priority combined with `activeMigration`.

```yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: data-processing
spec:
  template:
    spec:
      restartPolicy: OnFailure
      containers:
      - name: processor
        image: gcr.io/project/data-processor:stable
        args: ["--input", "gs://my-bucket/input", "--output", "gs://my-bucket/output"]
      nodeSelector:
        computeclass: spot-fallback
      terminationGracePeriodSeconds: 20

```

Because `restartPolicy` is set to `OnFailure`, a Spot eviction triggers rescheduling on the next available Spot or on-demand node, while checkpointed output in Cloud Storage preserves progress.

## Workload Suitability and Graceful Termination

Not every workload belongs on Spot. The [`skills/cloud/gke-basics/references/gke-cost.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/references/gke-cost.md) guide identifies suitable candidates as **batch jobs**, **CI pipelines**, **stateless microservices**, and **HPC tasks that checkpoint frequently**. The same file requires configuring `terminationGracePeriodSeconds` to a value under 30 seconds so containers can gracefully exit or checkpoint before the VM is reclaimed.

## Summary

- **Spot VMs** can cut GKE node costs by 60–90% and total cluster spend by roughly 45% when adopted for interruptible workloads.
- The **ComputeClass** API in [`gke-compute-classes.md`](https://github.com/google/skills/blob/main/gke-compute-classes.md) abstracts node pools and supports `spot: true`, `onDemandFallback`, and `activeMigration` in a single declarative resource.
- Pods reference the class through a `computeclass` label in `nodeSelector`, keeping specs portable and version controlled.
- Always set **`terminationGracePeriodSeconds` to less than 30** to satisfy the Spot termination notice.
- Batch and HPC jobs should use `restartPolicy: OnFailure` and externalize state to Cloud Storage so that Spot evictions do not cause data loss.
- Monitor Spot eviction metrics such as `kube_node_status_condition{reason="SpotTerminated"}` to measure preemption rates and refine workload placement.

## Frequently Asked Questions

### How much can Spot VMs reduce GKE costs?

Spot VMs lower individual node prices by 60–90% compared with on-demand instances. According to [`skills/cloud/gke-basics/references/gke-cost.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/references/gke-cost.md) in the `google/skills` repository, shifting about 30% of cluster capacity to Spot can reduce the overall GKE bill by roughly 45%.

### What happens when a Spot VM is reclaimed while my pod is running?

Google Cloud sends a termination notice approximately 30 seconds before reclaiming the VM. Your pod must respect the `terminationGracePeriodSeconds` value, which should be set to less than 30 seconds, to save state and exit cleanly before the node shuts down.

### Can I force a pod to stay on Spot and avoid on-demand nodes?

No, relying solely on Spot introduces availability risk when excess capacity runs out. The recommended pattern in [`skills/cloud/gke-basics/references/gke-compute-classes.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-basics/references/gke-compute-classes.md) uses `onDemandFallback: true` so the scheduler falls back to on-demand nodes automatically. You can then use `activeMigration: true` to move the pod back to Spot once capacity returns.

### Which GKE workloads are best suited for Spot VMs?

The [`gke-cost.md`](https://github.com/google/skills/blob/main/gke-cost.md) and [`gke-batch-hpc.md`](https://github.com/google/skills/blob/main/gke-batch-hpc.md) references list batch data processing, CI/CD pipelines, stateless microservices, and checkpointed HPC tasks as ideal candidates. Stateful databases or latency-sensitive user-facing services generally should not run on Spot.