Troubleshooting GKE Workload Deployment Issues: A 6-Step Diagnostic Workflow
Troubleshooting GKE workload deployment issues requires systematic analysis of pod status, namespace events, container logs, and network policies to isolate root causes like CrashLoopBackOff, OOMKilled, and ImagePullBackOff before proposing GitOps corrections.
The gke-workload-troubleshooting skill in the google/skills repository provides a deterministic, non-interactive workflow for debugging Kubernetes workloads on Google Kubernetes Engine (GKE). This open-source diagnostic framework automates the investigation of common failure states while maintaining strict read-only safety guarantees and producing version-controlled remediation patches.
The 6-Step Diagnostic Workflow
Defined in skills/cloud/gke-workload-troubleshooting/SKILL.md, the workflow processes GKE workload failures through six logical stages, from context discovery to GitOps correction. Each step exposes specific kubectl and gcloud commands that isolate failure domains without modifying cluster state.
Step 0: Context Discovery and Time-Window Definition
The workflow begins by extracting GKE cluster identifiers and workload metadata automatically. The skill parses project_id, cluster_name, cluster_location, workload_name, and workload_namespace from the user prompt, active SETTINGS.md, or local environment variables, applying defaults such as workload_namespace = default when values are missing.
If the cluster is reachable, the skill executes gcloud container clusters get-credentials to establish authentication. When network isolation or permissions prevent live access, the workflow falls back to dry-run mode, printing the complete diagnostic command sequence for manual execution by an operator.
Step 1: Analyze Pod Status and Conditions
This step determines why a pod is not running or repeatedly restarting. The skill retrieves the deployment selector using kubectl get deployment {workload_name} -n {workload_namespace} -o jsonpath='{.spec.selector.matchLabels}', then lists matching pods to inspect their phase and container state.
The workflow branches based on detected states:
- Pending: Indicates scheduling or resource constraints
- CrashLoopBackOff: Signals container runtime failures or missing dependencies
- ContainerCreating: Suggests image pull issues or volume mount problems
Step 2: Query Namespace Events
To surface GKE-level scheduling and infrastructure events, the skill runs:
kubectl get events -n {workload_namespace} --sort-by='.metadata.creationTimestamp'
For historical analysis, the workflow queries Cloud Logging within a ±30-minute window around the incident timestamp:
gcloud logging read \
"resource.type=\"k8s_cluster\" AND logName=\"projects/{project_id}/logs/events\" \
AND jsonPayload.involvedObject.namespace=\"{workload_namespace}\"" \
--start-time="{start_time}" --end-time="{end_time}" \
--project="{project_id}"
This time-window handling narrows results to the most relevant entries, reducing noise from unrelated cluster activity.
Step 3: Inspect Application Logs
The skill pulls recent container output to locate runtime exceptions, OOM signatures, or network timeouts. For active containers:
kubectl logs {pod_name} -n {workload_namespace} --all-containers --tail=100
For terminated containers in CrashLoopBackOff states, the workflow retrieves previous instance logs:
kubectl logs {pod_name} -n {workload_namespace} --all-containers -p --tail=100
Step 4: Verify Service Connectivity and Network Policies
Network isolation verification involves checking endpoint availability and policy configurations:
kubectl get endpoints {target_service_name} -n {target_namespace}
kubectl get networkpolicies -n {workload_namespace} -o yaml
In dry-run mode, the skill inspects manifests for service hostnames and ports, suggesting NetworkPolicy patches when egress traffic appears blocked.
Step 5: Propose GitOps Correction
The final step synthesizes findings into a human-readable root-cause analysis and generates a YAML manifest patch. Rather than applying changes directly to the cluster, the skill prepares PR-ready corrections—such as increasing memory limits from 256Mi to 512Mi for OOMKilled containers or adding missing secret references.
Essential Diagnostic Commands
The following command templates represent the core CLI actions orchestrated by the skill. Replace placeholders ({}) with values discovered during Step 0.
Retrieve deployment configuration and selectors:
kubectl get deployment {workload_name} -n {workload_namespace} -o yaml
kubectl get deploy/{workload_name} -n {workload_namespace} -o jsonpath='{.spec.selector.matchLabels}'
List pods matching deployment labels:
kubectl get pods -l {selector_labels} -n {workload_namespace}
Generate a memory limit patch for GitOps remediation:
cat <<EOF > patch.yaml
spec:
template:
spec:
containers:
- name: {container_name}
resources:
limits:
memory: 512Mi
EOF
Safety and Architecture Principles
The gke-workload-troubleshooting skill implements four critical design patterns that ensure production safety:
- Read-only diagnostics: The skill limits itself to gathering logs, events, and manifests, preventing accidental modifications to running workloads
- Fallback dry-run mode: When live cluster access is unavailable, the workflow delivers complete command lists for human operators to execute manually
- Time-window precision: Cloud Logging queries center on a 1-hour window around the incident timestamp, filtering irrelevant historical data
- GitOps-first remediation: All fixes are expressed as version-controlled manifest patches, aligning with modern continuous-delivery practices and preserving audit trails
Summary
- The gke-workload-troubleshooting skill in
google/skillsprovides a deterministic 6-step workflow for troubleshooting GKE workload deployment issues - Step 0 extracts cluster context and defaults to dry-run mode when live access is unavailable
- Steps 1-4 isolate failures through pod status analysis, event querying, log inspection, and network verification using standard
kubectlandgcloudcommands - Step 5 generates GitOps-ready manifest patches rather than modifying clusters directly
- The architecture prioritizes read-only safety, time-bound log analysis, and version-controlled remediation
Frequently Asked Questions
What is the difference between dry-run and live mode in the GKE troubleshooting skill?
In live mode, the skill executes kubectl and gcloud commands directly against the target cluster after authenticating via gcloud container clusters get-credentials. In dry-run mode, which activates automatically when cluster access fails or credentials are unavailable, the skill prints the complete diagnostic command sequence without execution, allowing operators to review and run commands manually in sandboxed or restricted environments.
How does the workflow handle CrashLoopBackOff containers?
The skill detects CrashLoopBackOff states during Step 1 by inspecting container status fields. It then retrieves logs from the previously terminated container instance using kubectl logs {pod_name} -n {workload_namespace} --all-containers -p --tail=100 to capture the error message that caused the exit, while simultaneously querying namespace events for image pull or volume mount failures that might prevent successful startup.
What GitOps corrections does the skill generate for OOMKilled pods?
When logs contain OOM (Out of Memory) signatures, the skill calculates the discrepancy between the memory limit and actual usage—such as detecting that a 256Mi limit is insufficient for 270Mi actual consumption. It then generates a YAML patch file increasing the memory limit to a safe threshold (e.g., 512Mi) and prepares this as a Pull Request against the workload repository, ensuring changes undergo code review before deployment.
Can this workflow diagnose issues without direct cluster access?
Yes. The skill's dry-run mode functions without cluster connectivity by parsing local manifest files and environment variables to construct the diagnostic command tree. While it cannot retrieve live logs or events in this state, it produces a complete, copy-pasteable command sequence that operators can execute in environments with appropriate credentials, making it valuable for air-gapped or highly restricted production environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →