Optimizing GKE Costs Using Spot VMs: ComputeClass Patterns and Fallback Strategies
You can reduce GKE node costs by 60–90% using Spot VMs orchestrated through ComputeClasses that provide on-demand fallback and automatic re-migration, provided your workloads tolerate interruption and shutdown within 30 seconds.
The google/skills repository provides authoritative patterns for optimizing GKE costs using Spot VMs through declarative node management. By leveraging the ComputeClass API instead of hand-tuned node pools, platform teams can automate Spot provisioning, fallback, and workload migration while cutting cluster spend by up to roughly 45%.
Cost Impact of Spot VMs on GKE Clusters
According to skills/cloud/gke-basics/references/gke-cost.md, Spot VMs consume excess Google Cloud data center capacity and cost 60–90% less than on-demand nodes. The same guide notes that moving approximately 30% of node capacity to Spot can lower the overall GKE bill by roughly 45% without violating service-level objectives.
Declaring Spot VMs with the ComputeClass API
The skills/cloud/gke-basics/references/gke-compute-classes.md reference demonstrates that the recommended abstraction is the ComputeClass resource (compute.cnrm.cloud.google.com/v1beta1). Rather than manually creating Spot node pools, you set spot: true inside the ComputeClass spec and optionally enable activeMigration and onDemandFallback.
apiVersion: compute.cnrm.cloud.google.com/v1beta1
kind: ComputeClass
metadata:
name: spot-fallback
spec:
machineType: n1-standard-4
spot: true
activeMigration: true
onDemandFallback: true
spot: truerequests Spot VM capacity.onDemandFallback: trueallows scheduling on standard nodes when Spot is unavailable.activeMigration: trueautomatically migrates the pod back to Spot when excess capacity returns.
Linking Pods to a Spot ComputeClass
To consume the class, reference it in the pod spec via nodeSelector. The skills/cloud/gke-basics/SKILL.md documentation maps ComputeClasses to workloads using this label pattern.
apiVersion: v1
kind: Pod
metadata:
name: batch-worker
spec:
containers:
- name: worker
image: gcr.io/project/batch-worker:latest
nodeSelector:
computeclass: spot-fallback
terminationGracePeriodSeconds: 25
The computeclass: spot-fallback label binds the pod to the Spot-first policy. The terminationGracePeriodSeconds value must stay below the 30-second eviction notice window.
On-Demand Fallback for Critical Availability
As documented in skills/cloud/gke-basics/references/gke-compute-classes.md, the fallback pattern guarantees that pods remain schedulable even when Spot capacity is exhausted across a region. The scheduler attempts Spot first; if capacity is absent, it immediately provisions an on-demand node. Because the ComputeClass manages both tiers, operators do not need to maintain separate fallback node pools manually.
Batch and HPC Workloads with Checkpointing
Fault-tolerant batch jobs and high-performance computing tasks are ideal for Spot because they can persist progress to Cloud Storage and resume after eviction. The skills/cloud/gke-basics/references/gke-batch-hpc.md reference includes a concrete pattern that uses Spot-first priority combined with activeMigration.
apiVersion: batch/v1
kind: Job
metadata:
name: data-processing
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: processor
image: gcr.io/project/data-processor:stable
args: ["--input", "gs://my-bucket/input", "--output", "gs://my-bucket/output"]
nodeSelector:
computeclass: spot-fallback
terminationGracePeriodSeconds: 20
Because restartPolicy is set to OnFailure, a Spot eviction triggers rescheduling on the next available Spot or on-demand node, while checkpointed output in Cloud Storage preserves progress.
Workload Suitability and Graceful Termination
Not every workload belongs on Spot. The skills/cloud/gke-basics/references/gke-cost.md guide identifies suitable candidates as batch jobs, CI pipelines, stateless microservices, and HPC tasks that checkpoint frequently. The same file requires configuring terminationGracePeriodSeconds to a value under 30 seconds so containers can gracefully exit or checkpoint before the VM is reclaimed.
Summary
- Spot VMs can cut GKE node costs by 60–90% and total cluster spend by roughly 45% when adopted for interruptible workloads.
- The ComputeClass API in
gke-compute-classes.mdabstracts node pools and supportsspot: true,onDemandFallback, andactiveMigrationin a single declarative resource. - Pods reference the class through a
computeclasslabel innodeSelector, keeping specs portable and version controlled. - Always set
terminationGracePeriodSecondsto less than 30 to satisfy the Spot termination notice. - Batch and HPC jobs should use
restartPolicy: OnFailureand externalize state to Cloud Storage so that Spot evictions do not cause data loss. - Monitor Spot eviction metrics such as
kube_node_status_condition{reason="SpotTerminated"}to measure preemption rates and refine workload placement.
Frequently Asked Questions
How much can Spot VMs reduce GKE costs?
Spot VMs lower individual node prices by 60–90% compared with on-demand instances. According to skills/cloud/gke-basics/references/gke-cost.md in the google/skills repository, shifting about 30% of cluster capacity to Spot can reduce the overall GKE bill by roughly 45%.
What happens when a Spot VM is reclaimed while my pod is running?
Google Cloud sends a termination notice approximately 30 seconds before reclaiming the VM. Your pod must respect the terminationGracePeriodSeconds value, which should be set to less than 30 seconds, to save state and exit cleanly before the node shuts down.
Can I force a pod to stay on Spot and avoid on-demand nodes?
No, relying solely on Spot introduces availability risk when excess capacity runs out. The recommended pattern in skills/cloud/gke-basics/references/gke-compute-classes.md uses onDemandFallback: true so the scheduler falls back to on-demand nodes automatically. You can then use activeMigration: true to move the pod back to Spot once capacity returns.
Which GKE workloads are best suited for Spot VMs?
The gke-cost.md and gke-batch-hpc.md references list batch data processing, CI/CD pipelines, stateless microservices, and checkpointed HPC tasks as ideal candidates. Stateful databases or latency-sensitive user-facing services generally should not run on Spot.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →