# How to Handle Spot VM Preemption with a Grace Period in GKE

> Learn how to handle Spot VM preemption in GKE using terminationGracePeriodSeconds and preStop hooks. Discover how to extend node shutdown grace periods with GKE 1.35+ ComputeClasses.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Set `terminationGracePeriodSeconds` to less than 30 seconds on Pods running on Spot VMs, implement `preStop` hooks for cleanup logic, and use GKE 1.35+ ComputeClasses to extend node-level shutdown grace periods.**

Google Kubernetes Engine (GKE) Spot VMs offer significant cost savings but can be preempted by the Compute Engine scheduler at any time with only a 30-second termination notice. According to the `google/skills` repository, handling Spot VM preemption with a grace period requires configuring both pod-level settings and node-level ComputeClass configurations to ensure workloads shut down cleanly before the `compute.instances.preempted` event completes.

## Understanding the 30-Second Preemption Window

Spot VMs in GKE receive a mandatory 30-second termination notice that cannot be extended or delayed. As documented in [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md), this narrow window means your **grace period must be strictly less than 30 seconds** to guarantee the kubelet can process the termination request before the instance halts. Any configuration exceeding this limit risks immediate, ungraceful termination of containers.

## Configuring Pod-Level Grace Periods

### Setting terminationGracePeriodSeconds

The primary defense against Spot preemption is reducing the Pod's `terminationGracePeriodSeconds` to a value under 30 seconds. According to [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md) at line 154, setting this to 25 seconds provides a 5-second buffer for the kubelet to initiate shutdown.

```yaml
apiVersion: v1
kind: Pod
metadata:
  name: spot-workload
spec:
  terminationGracePeriodSeconds: 25
  containers:
  - name: app
    image: gcr.io/my-project/my-app:latest

```

### Implementing preStop Hooks

To perform cleanup operations within the grace window, add a `preStop` lifecycle hook as shown in [`skills/cloud/gke-reliability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-reliability/SKILL.md) at line 207. This hook executes before the container receives SIGTERM, allowing checkpointing, connection draining, or state persistence.

```yaml
    lifecycle:
      preStop:
        exec:
          command: ["/bin/sh", "-c", "echo 'Saving checkpoint…'; /app/checkpoint.sh"]

```

## Extending Node-Level Shutdown Grace Periods (GKE 1.35+)

For clusters running GKE 1.35 or later, you can extend the node's shutdown behavior using ComputeClasses. The [`skills/cloud/gke-cluster-autoscaler/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cluster-autoscaler/SKILL.md) file at line 29 describes setting `kubeletConfig.shutdownGracePeriodSeconds` to values higher than the default, allowing the node to wait longer for Pods to terminate after receiving the preemption notice.

```yaml
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
  name: spot-with-extended-grace
spec:
  spot: true
  kubeletConfig:
    shutdownGracePeriodSeconds: 120

```

This node-level setting works in conjunction with pod-level configurations, providing up to 120 seconds for complex shutdown sequences while still respecting the 30-second notice for the initial termination signal.

## Protecting Workloads with PodDisruptionBudgets

While PodDisruptionBudgets (PDBs) cannot prevent involuntary Spot preemptions, they protect against voluntary disruptions during node drains. As noted in [`skills/cloud/gke-compute-classes/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md) at line 196, configure PDBs to ensure minimum availability during maintenance operations, though they will not block the `compute.instances.preempted` event.

```yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: spot-pdb
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: spot-workload

```

## Monitoring Spot VM Preemption Events

Verify your grace period configuration by monitoring preemption events through Cloud Logging. The [`skills/cloud/cloud-logging-query-generation/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/cloud-logging-query-generation/SKILL.md) file at line 188 provides queries for the `compute.instances.preempted` method, while GKE-specific metrics like `node_preempted` help track termination patterns.

```bash
gcloud logging read 'protoPayload.methodName="compute.instances.preempted"' \
  --project=my-project \
  --format="value(timestamp,textPayload)"

```

## Summary

- **Set `terminationGracePeriodSeconds` below 30 seconds** (e.g., 25s) on all Pods running on Spot VMs to fit within the mandatory preemption notice window.
- **Implement `preStop` hooks** to execute cleanup logic before SIGTERM is sent to containers.
- **Use ComputeClasses with `kubeletConfig.shutdownGracePeriodSeconds`** (GKE 1.35+) to extend node-level waiting periods for complex shutdown sequences.
- **Configure PodDisruptionBudgets** to protect against voluntary disruptions, though they cannot block involuntary Spot preemptions.
- **Monitor `compute.instances.preempted` events** via Cloud Logging to validate graceful shutdown behavior.

## Frequently Asked Questions

### What happens if terminationGracePeriodSeconds exceeds 30 seconds on Spot VMs?

If you set `terminationGracePeriodSeconds` to 30 seconds or higher, the kubelet will attempt to wait for the full duration, but the Compute Engine scheduler will forcefully terminate the VM after exactly 30 seconds. This results in ungraceful container shutdown and potential data loss, as the instance receives the `compute.instances.preempted` signal with no flexibility in the timeline.

### Do PodDisruptionBudgets prevent Spot VM preemptions?

No, PodDisruptionBudgets do not prevent involuntary Spot VM preemptions. According to the source code in [`skills/cloud/gke-compute-classes/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-compute-classes/SKILL.md), PDBs only enforce minimum availability during voluntary disruptions such as node drains or cluster upgrades. The Compute Engine scheduler bypasses PDB enforcement when reclaiming Spot capacity.

### How can I verify my grace period configuration is working?

Query Cloud Logging for `compute.instances.preempted` events and correlate them with your application logs to ensure cleanup completes before termination. Check that your `preStop` hook execution time plus the container's shutdown time remains under your configured `terminationGracePeriodSeconds` threshold.

### What is the difference between pod-level and node-level grace periods?

The pod-level `terminationGracePeriodSeconds` controls how long the kubelet waits for an individual Pod to terminate after receiving the preemption notice, while the node-level `kubeletConfig.shutdownGracePeriodSeconds` (available in GKE 1.35+) controls how long the entire node waits for all Pods to exit during a system shutdown event. The node-level setting must accommodate the sum of all pod-level grace periods on that node.