# Troubleshooting GKE Workload Deployment Issues: A 6-Step Diagnostic Workflow

> Solve GKE workload deployment issues with our 6-step diagnostic workflow. Debug pods, analyze events, and fix common errors like CrashLoopBackOff and ImagePullBackOff efficiently.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Troubleshooting GKE workload deployment issues requires systematic analysis of pod status, namespace events, container logs, and network policies to isolate root causes like CrashLoopBackOff, OOMKilled, and ImagePullBackOff before proposing GitOps corrections.**

The **gke-workload-troubleshooting** skill in the [google/skills](https://github.com/google/skills) repository provides a deterministic, non-interactive workflow for debugging Kubernetes workloads on Google Kubernetes Engine (GKE). This open-source diagnostic framework automates the investigation of common failure states while maintaining strict read-only safety guarantees and producing version-controlled remediation patches.

## The 6-Step Diagnostic Workflow

Defined in [`skills/cloud/gke-workload-troubleshooting/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-workload-troubleshooting/SKILL.md), the workflow processes GKE workload failures through six logical stages, from context discovery to GitOps correction. Each step exposes specific `kubectl` and `gcloud` commands that isolate failure domains without modifying cluster state.

### Step 0: Context Discovery and Time-Window Definition

The workflow begins by extracting GKE cluster identifiers and workload metadata automatically. The skill parses `project_id`, `cluster_name`, `cluster_location`, `workload_name`, and `workload_namespace` from the user prompt, active [`SETTINGS.md`](https://github.com/google/skills/blob/main/SETTINGS.md), or local environment variables, applying defaults such as `workload_namespace = default` when values are missing.

If the cluster is reachable, the skill executes `gcloud container clusters get-credentials` to establish authentication. When network isolation or permissions prevent live access, the workflow falls back to **dry-run mode**, printing the complete diagnostic command sequence for manual execution by an operator.

### Step 1: Analyze Pod Status and Conditions

This step determines why a pod is not running or repeatedly restarting. The skill retrieves the deployment selector using `kubectl get deployment {workload_name} -n {workload_namespace} -o jsonpath='{.spec.selector.matchLabels}'`, then lists matching pods to inspect their `phase` and container `state`.

The workflow branches based on detected states:

- **Pending**: Indicates scheduling or resource constraints
- **CrashLoopBackOff**: Signals container runtime failures or missing dependencies  
- **ContainerCreating**: Suggests image pull issues or volume mount problems

### Step 2: Query Namespace Events

To surface GKE-level scheduling and infrastructure events, the skill runs:

```bash
kubectl get events -n {workload_namespace} --sort-by='.metadata.creationTimestamp'

```

For historical analysis, the workflow queries Cloud Logging within a ±30-minute window around the incident timestamp:

```bash
gcloud logging read \
  "resource.type=\"k8s_cluster\" AND logName=\"projects/{project_id}/logs/events\" \
   AND jsonPayload.involvedObject.namespace=\"{workload_namespace}\"" \
  --start-time="{start_time}" --end-time="{end_time}" \
  --project="{project_id}"

```

This time-window handling narrows results to the most relevant entries, reducing noise from unrelated cluster activity.

### Step 3: Inspect Application Logs

The skill pulls recent container output to locate runtime exceptions, OOM signatures, or network timeouts. For active containers:

```bash
kubectl logs {pod_name} -n {workload_namespace} --all-containers --tail=100

```

For terminated containers in `CrashLoopBackOff` states, the workflow retrieves previous instance logs:

```bash
kubectl logs {pod_name} -n {workload_namespace} --all-containers -p --tail=100

```

### Step 4: Verify Service Connectivity and Network Policies

Network isolation verification involves checking endpoint availability and policy configurations:

```bash
kubectl get endpoints {target_service_name} -n {target_namespace}
kubectl get networkpolicies -n {workload_namespace} -o yaml

```

In dry-run mode, the skill inspects manifests for service hostnames and ports, suggesting NetworkPolicy patches when egress traffic appears blocked.

### Step 5: Propose GitOps Correction

The final step synthesizes findings into a human-readable root-cause analysis and generates a YAML manifest patch. Rather than applying changes directly to the cluster, the skill prepares PR-ready corrections—such as increasing memory limits from `256Mi` to `512Mi` for OOMKilled containers or adding missing secret references.

## Essential Diagnostic Commands

The following command templates represent the core CLI actions orchestrated by the skill. Replace placeholders (`{}`) with values discovered during Step 0.

Retrieve deployment configuration and selectors:

```bash
kubectl get deployment {workload_name} -n {workload_namespace} -o yaml
kubectl get deploy/{workload_name} -n {workload_namespace} -o jsonpath='{.spec.selector.matchLabels}'

```

List pods matching deployment labels:

```bash
kubectl get pods -l {selector_labels} -n {workload_namespace}

```

Generate a memory limit patch for GitOps remediation:

```bash
cat <<EOF > patch.yaml
spec:
  template:
    spec:
      containers:
      - name: {container_name}
        resources:
          limits:
            memory: 512Mi
EOF

```

## Safety and Architecture Principles

The **gke-workload-troubleshooting** skill implements four critical design patterns that ensure production safety:

- **Read-only diagnostics**: The skill limits itself to gathering logs, events, and manifests, preventing accidental modifications to running workloads
- **Fallback dry-run mode**: When live cluster access is unavailable, the workflow delivers complete command lists for human operators to execute manually
- **Time-window precision**: Cloud Logging queries center on a 1-hour window around the incident timestamp, filtering irrelevant historical data
- **GitOps-first remediation**: All fixes are expressed as version-controlled manifest patches, aligning with modern continuous-delivery practices and preserving audit trails

## Summary

- The **gke-workload-troubleshooting** skill in `google/skills` provides a deterministic 6-step workflow for troubleshooting GKE workload deployment issues
- **Step 0** extracts cluster context and defaults to dry-run mode when live access is unavailable
- **Steps 1-4** isolate failures through pod status analysis, event querying, log inspection, and network verification using standard `kubectl` and `gcloud` commands
- **Step 5** generates GitOps-ready manifest patches rather than modifying clusters directly
- The architecture prioritizes read-only safety, time-bound log analysis, and version-controlled remediation

## Frequently Asked Questions

### What is the difference between dry-run and live mode in the GKE troubleshooting skill?

In **live mode**, the skill executes `kubectl` and `gcloud` commands directly against the target cluster after authenticating via `gcloud container clusters get-credentials`. In **dry-run mode**, which activates automatically when cluster access fails or credentials are unavailable, the skill prints the complete diagnostic command sequence without execution, allowing operators to review and run commands manually in sandboxed or restricted environments.

### How does the workflow handle CrashLoopBackOff containers?

The skill detects `CrashLoopBackOff` states during Step 1 by inspecting container status fields. It then retrieves logs from the previously terminated container instance using `kubectl logs {pod_name} -n {workload_namespace} --all-containers -p --tail=100` to capture the error message that caused the exit, while simultaneously querying namespace events for image pull or volume mount failures that might prevent successful startup.

### What GitOps corrections does the skill generate for OOMKilled pods?

When logs contain OOM (Out of Memory) signatures, the skill calculates the discrepancy between the memory limit and actual usage—such as detecting that a `256Mi` limit is insufficient for `270Mi` actual consumption. It then generates a YAML patch file increasing the memory limit to a safe threshold (e.g., `512Mi`) and prepares this as a Pull Request against the workload repository, ensuring changes undergo code review before deployment.

### Can this workflow diagnose issues without direct cluster access?

Yes. The skill's **dry-run mode** functions without cluster connectivity by parsing local manifest files and environment variables to construct the diagnostic command tree. While it cannot retrieve live logs or events in this state, it produces a complete, copy-pasteable command sequence that operators can execute in environments with appropriate credentials, making it valuable for air-gapped or highly restricted production environments.