# Configuring GKE Reliability and SLOs: Production Best Practices

> Configure GKE reliability and SLOs using regional clusters, PDBs, health probes, and topology constraints. Define error budget alerts with PromQL burn-rate windows for production best practices.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: best-practices
- Published: 2026-08-13

---

**To configure GKE reliability and SLOs, combine regional cluster settings with workload-level Pod Disruption Budgets, health probes, and topology spread constraints, then define error-budget alert policies using PromQL-based burn-rate windows.**

The `google/skills` repository provides a comprehensive framework for production-grade Google Kubernetes Engine (GKE) deployments. According to the `gke-reliability` skill documented in [`skills/cloud/gke-reliability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-reliability/SKILL.md), achieving high availability requires both cluster-level infrastructure choices and workload-level Kubernetes manifests. This guide walks through the golden path for configuring GKE reliability and implementing measurable Service-Level Objectives (SLOs).

## Cluster-Level Reliability Foundations

Start with the infrastructure layer to establish baseline resilience. The `gke-reliability` skill recommends deploying **regional clusters** to distribute control plane components across multiple zones. Enable **auto-repair** and **auto-upgrade** features to maintain node health without manual intervention, ensuring the cluster foundation remains stable before applying workload-specific safeguards.

## Workload-Level Reliability Safeguards

With the cluster foundation configured, apply five critical patterns to your deployments as specified in [`skills/cloud/gke-reliability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-reliability/SKILL.md).

### Pod Disruption Budgets (PDBs)

Define a `PodDisruptionBudget` to guarantee minimum availability during voluntary disruptions like node upgrades or cluster autoscaler scale-downs. Specify either `minAvailable` or `maxUnavailable` to prevent excessive pod termination during maintenance windows.

```yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api-pdb
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: api-gateway

```

### Health Probes

Configure `livenessProbe`, `readinessProbe`, and `startupProbe` with explicit timeouts, initial delays, and failure thresholds. These probes enable the kubelet to restart unhealthy containers and remove pods from service endpoints before they receive traffic, preventing cascading failures.

```yaml
livenessProbe:
  httpGet:
    path: /healthz
    port: 8080
  initialDelaySeconds: 30
  timeoutSeconds: 5
  failureThreshold: 3
readinessProbe:
  httpGet:
    path: /ready
    port: 8080
  periodSeconds: 10

```

### Graceful Shutdown Configuration

Set `terminationGracePeriodSeconds` and a `preStop` hook to allow load balancers time to deregister pods before the container receives SIGTERM. This pattern prevents in-flight request failures during rolling updates or node drains.

```yaml
lifecycle:
  preStop:
    exec:
      command: ["sleep", "15"]
terminationGracePeriodSeconds: 60

```

### Topology Spread Constraints

Distribute pods across failure domains to survive zone- or host-level outages. Use `topology.kubernetes.io/zone` and `kubernetes.io/hostname` keys as documented in the reliability skill to ensure high availability.

```yaml
topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: DoNotSchedule
    labelSelector:
      matchLabels:
        app: critical-service

```

### Replica Count Guidelines

Maintain **2 or more replicas** for stateless services and **3 or more** for critical or stateful workloads. This ensures quorum maintenance and zero-downtime operations during zone failures or rolling updates.

## Defining and Monitoring SLOs for GKE

Reliability requires quantifiable measurement. The `google-cloud-slo-alert-configuration` skill in [`skills/cloud/google-cloud-slo-alert-configuration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/SKILL.md) provides a Terraform-based workflow for implementing PromQL-based SLO alert policies.

### Identifying Service-Level Indicators

First, define your **SLIs** (Service-Level Indicators). Common GKE metrics include request error rates calculated from `http_request_total` versus `http_request_error_total`, or latency percentiles. Choose an **SLO target** such as 99.9% success rate, establishing a 0.1% error budget for the measurement window.

### Configuring Burn-Rate Alert Policies

Deploy `google_monitoring_alert_policy` resources that evaluate error-burn rates. The skill generates Terraform configurations that trigger when consumption exceeds thresholds over specific temporal windows.

```hcl
resource "google_monitoring_alert_policy" "fast_burn" {
  display_name = "SLO Fast Burn Alert"
  combiner     = "OR"
  conditions {
    display_name = "Error rate exceeds budget over 1h"
    condition_prometheus_query_language {
      alert_rule = "sum(rate(http_request_error_total[1h])) / sum(rate(http_request_total[1h])) > 0.002"
    }
  }
}

```

### PromQL Templates for Error Budgets

Reference [`skills/cloud/google-cloud-slo-alert-configuration/references/promql_templates.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/references/promql_templates.md) for pre-built query snippets. Use **fast-burn** windows (1 hour) for urgent detection of sudden spikes and **slow-burn** windows (3 days) for identifying gradual degradation that threatens long-term reliability targets.

## Summary

- Deploy regional GKE clusters with auto-repair and auto-upgrade enabled per the `gke-reliability` skill.
- Implement Pod Disruption Budgets, health probes, and topology spread constraints as documented in [`skills/cloud/gke-reliability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-reliability/SKILL.md).
- Configure graceful shutdown with `preStop` hooks and adequate `terminationGracePeriodSeconds` to prevent request drops during termination.
- Define SLIs based on request error rates or latency, targeting 99.9% availability or higher to establish error budgets.
- Deploy burn-rate alert policies using the Terraform templates from [`skills/cloud/google-cloud-slo-alert-configuration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/SKILL.md).

## Frequently Asked Questions

### What is the difference between an SLI and an SLO in GKE?

An SLI (Service-Level Indicator) is a measurable metric such as request latency or error rate derived from Prometheus or Cloud Monitoring. An SLO (Service-Level Objective) defines the target reliability percentage (e.g., 99.9%) for that SLI over a specific time window, effectively quantifying your error budget.

### How do Pod Disruption Budgets improve GKE reliability?

Pod Disruption Budgets prevent voluntary disruptions—such as node upgrades or cluster autoscaler scale-downs—from reducing application availability below defined thresholds. By setting `minAvailable` or `maxUnavailable` in the PDB spec, you ensure Kubernetes maintains the minimum required pod count during maintenance operations.

### What is a burn-rate alert and why use different windows?

A burn-rate alert tracks how quickly your service consumes its error budget relative to the SLO target. Fast-burn alerts (1 hour) detect sudden error spikes requiring immediate incident response, while slow-burn alerts (3 days) identify gradual degradation that, if uncorrected, would exhaust the quarterly error budget.

### Where can I find the complete GKE reliability configuration templates?

The `google/skills` repository hosts authoritative templates in [`skills/cloud/gke-reliability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-reliability/SKILL.md) for workload configuration and [`skills/cloud/google-cloud-slo-alert-configuration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/SKILL.md) for SLO monitoring implementation, including ready-to-use PromQL examples in [`skills/cloud/google-cloud-slo-alert-configuration/references/promql_templates.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/references/promql_templates.md).