# How to Troubleshoot GKE 429 Resource Exhausted Errors with Gemini API

> Troubleshoot GKE 429 Resource Exhausted errors when calling the Gemini API. Learn how to use exponential back-off retries or Provisioned Throughput to resolve quota depletion.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-09

---

**GKE workloads calling the Gemini API receive HTTP 429 Resource Exhausted errors when the Dynamic Shared Quota (DSQ) pool is temporarily depleted, requiring exponential back-off retries or Provisioned Throughput to resolve.**

When deploying AI workloads on **Google Kubernetes Engine (GKE)** that consume the **Gemini API** through the Agent Platform, encountering HTTP 429 *Resource Exhausted* responses indicates a temporary capacity issue rather than a project-level quota violation. According to the source code in the `google/skills` repository, these errors stem from the **Dynamic Shared Quota (DSQ)** system used by Gemini and OpenMaaS models. Understanding the distinction between shared pool exhaustion and dedicated quota limits is essential for implementing the correct mitigation strategy.

## Understanding Dynamic Shared Quota (DSQ) Exhaustion

The 429 errors originate from the **Dynamic Shared Quota (DSQ)** mechanism documented in [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md). Unlike standard project quotas that limit your specific billing account, DSQ represents a pooled compute resource allocation shared across all users of Gemini and OpenMaaS models. When this shared pool saturates due to high demand, the service returns 429 *Resource Exhausted* to indicate temporary unavailability, not that your specific project has exceeded its limits.

### High-Throughput Workload Impact

Rapid-fire API calls from multiple GKE pods—or from a single pod with aggressive retry logic—can exhaust the DSQ faster than the pool replenishes. This scenario is common in autoscaling environments where many replicas initialize simultaneously and hammer the Gemini API without client-side throttling.

### Preview Model Limitations

Certain **preview models**, including Gemini 3.1, are **only** available in the `global` region according to the *Gemini Models* section of [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md). These preview models operate under stricter DSQ limits than stable regional models, making them more susceptible to 429 errors during peak usage periods.

## Resolving 429 Errors on GKE

### Implement Exponential Back-Off and Retry

The primary mitigation for DSQ exhaustion is implementing **exponential back-off** in your application code. Configure your retry loop to wait progressively longer between attempts—starting at 1 second, then 2 seconds, 4 seconds, and 8 seconds—up to a maximum of approximately 30 seconds. This approach gives the DSQ pool time to replenish while preventing your workload from contributing to the congestion.

### Enable Provisioned Throughput (PT)

For production workloads requiring guaranteed capacity and consistent latency, request **Provisioned Throughput (PT)** via the Google Cloud Console. PT allocates dedicated quota that bypasses the shared DSQ entirely, eliminating 429 errors caused by pool exhaustion. This option is ideal for business-critical inference pipelines running on GKE.

### Verify Model and Region Availability

If you are using preview models restricted to the `global` region, consider migrating to stable models available in specific regions like `us-central1`. Regional models typically have larger DSQ allocations and lower contention than the global endpoint. Check [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md) to confirm which models are available in your target regions.

### Validate GKE Cluster Configuration

Ensure your GKE cluster meets the prerequisites for AI workloads as detailed in [`skills/cloud/google-cloud-solution-guided-gke-ai-migration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-guided-gke-ai-migration/SKILL.md). Specifically, **Phase 0** and **Phase 1** of the guided migration require:

- **AI-optimized node pools** utilizing GPU or TPU accelerators
- **Workload Identity** correctly configured to allow pods to obtain Google Cloud access tokens
- Proper service account bindings for Gemini API authentication

Without Workload Identity, your pods may fail authentication entirely or retry rapidly enough to trigger rate limiting.

### Monitor DSQ Metrics

Configure Cloud Monitoring dashboards to track **Vertex AI Shared Quota** (labeled as *Dynamic Shared Quota*) metrics. Set alerting thresholds to notify your team when DSQ utilization approaches 80%, allowing you to scale down non-critical workloads or enable PT before 429 errors impact production traffic.

## Python Implementation with GenAI SDK

The following example demonstrates exponential back-off using the **GenAI SDK** (the preferred SDK for Gemini models) within a GKE pod utilizing Workload Identity:

```python
import time
import google.auth
from google.auth.transport.requests import Request
from google.genai import GenerativeModel

# Acquire a fresh ADC access token (works with Workload Identity)

creds, _ = google.auth.default()
creds.refresh(Request())

# Initialise the client (replace with your desired location)

model = GenerativeModel(
    model_name="gemini-1.5-flash",        # change to your model

    location="us-central1",               # choose a region where the model is available

    credentials=creds)

def call_gemini(prompt: str, max_retries: int = 5):
    backoff = 1   # seconds

    for attempt in range(max_retries):
        try:
            response = model.generate_content(prompt)
            return response.text
        except Exception as e:
            if "429" in str(e):
                # DSQ exhausted – back-off and retry

                time.sleep(backoff)
                backoff *= 2          # exponential back-off

                continue
            raise  # re-raise non-429 errors

    raise RuntimeError("Exceeded max retries for Gemini API")

print(call_gemini("Explain the difference between a pod and a node in Kubernetes."))

```

For **OpenAI-compatible** endpoints serving OpenMaaS models, apply the same retry pattern using the `openai` Python client with identical back-off logic.

## Summary

- **429 Resource Exhausted** indicates temporary **Dynamic Shared Quota (DSQ)** pool depletion, not project quota limits.
- **Exponential back-off** is the primary client-side mitigation for DSQ contention in development and non-critical workloads.
- **Provisioned Throughput (PT)** provides dedicated capacity that bypasses DSQ for production GKE workloads requiring guaranteed availability.
- **Preview models** (e.g., Gemini 3.1) restricted to the `global` region face stricter DSQ limits than stable regional models.
- Proper **GKE cluster configuration** with Workload Identity and AI-optimized node pools is required for secure, efficient Gemini API access.

## Frequently Asked Questions

### What is the difference between DSQ exhaustion and project quota limits?

**Dynamic Shared Quota (DSQ)** represents a pooled compute resource shared among all users of Gemini and OpenMaaS models, while project quota limits are specific to your Google Cloud billing account and project. A 429 error from DSQ exhaustion means the shared infrastructure is temporarily overloaded, whereas a project quota error (typically 429 with different messaging or 403) indicates you have exceeded your allocated requests-per-minute or tokens-per-minute limits.

### Should I implement exponential back-off or Provisioned Throughput for my GKE workload?

Implement **exponential back-off** for development environments, batch processing, or workloads tolerant of variable latency. Enable **Provisioned Throughput (PT)** for production services requiring consistent response times and guaranteed capacity. PT eliminates DSQ-related 429 errors entirely by reserving dedicated compute resources for your project.

### Why do I only receive 429 errors when using specific Gemini models?

You likely are using **preview models** that are only available in the `global` region. According to [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md), preview models like Gemini 3.1 operate under stricter DSQ limits than stable models. Switching to a stable model available in specific regions (such as `gemini-2.5-pro` in `us-central1`) typically resolves region-specific 429 issues.

### How can I verify my GKE cluster is properly configured for Gemini API calls?

Consult **Phase 0** and **Phase 1** of [`skills/cloud/google-cloud-solution-guided-gke-ai-migration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-guided-gke-ai-migration/SKILL.md) to verify your cluster utilizes **AI-optimized node pools** and has **Workload Identity** enabled. Ensure your pod service accounts are bound to Google Cloud service accounts with appropriate Vertex AI IAM roles, allowing seamless token acquisition for Gemini API authentication without manual credential management.