How to Troubleshoot Kubernetes Pods That Appear Healthy but Cannot Communicate
When Kubernetes pods show a Running status but cannot reach each other, the issue typically lies in the networking layer beneath the pod status, requiring systematic checks of CNI plugins, Service configurations, NetworkPolicies, and iptables rules.
This guide draws from real-world interview scenarios documented in the litu54/DevOps-Interview-Guide repository, specifically the troubleshooting patterns captured in JPMorgan/DevOps_SRE.md and Others/DevOps_Engineer_4.md. If your pods report healthy status yet fail to communicate, use this systematic approach to isolate whether the root cause is at the CNI, Service, or policy layer.
Verify Pod-Level Connectivity First
Start by confirming basic Layer 3 connectivity between the pods. As noted in JPMorgan/DevOps_SRE.md【L9】, healthy-looking pods can still suffer from underlying network partition issues.
Check that each pod has a unique IP address in the cluster network:
kubectl get pod -o wide
Test direct pod-to-pod communication by pinging from one container to another:
kubectl exec pod-a -- ping -c 3 $(kubectl get pod pod-b -o jsonpath='{.status.podIP}')
If the ping fails, the problem exists at the CNI (Container Network Interface) level. If ping succeeds but application-level calls fail, investigate the Service abstraction or higher-level networking components.
Inspect Service Configuration and Selectors
A common cause of silent communication failures is a Service misconfiguration where the selector does not match the pod labels, or the port to targetPort mapping is incorrect. According to Others/DevOps_Engineer_4.md【L32-L34】, verifying these mappings is a critical troubleshooting step.
Check your Service endpoints to ensure pods are actually registered:
kubectl get endpoints my-svc
Review the Service YAML for label alignment:
apiVersion: v1
kind: Service
metadata:
name: backend-svc
spec:
selector:
app: backend # must match pod label 'app: backend'
ports:
- port: 80
targetPort: 8080 # must match container port
Test the Service cluster IP directly from a pod:
kubectl exec pod-a -- curl -s <service-clusterIP>:<port>
Review NetworkPolicies for Silent Blocks
NetworkPolicies can silently drop traffic between pods in the same namespace without generating pod-level error events. List active policies to check for overly restrictive rules:
kubectl get networkpolicy -n <namespace>
A restrictive policy might block all ingress by default. Apply a minimal allow-all policy for intra-namespace traffic to test:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-all
spec:
podSelector: {}
ingress:
- {}
If communication resumes after applying this policy, incrementally restrict rules to identify the specific traffic pattern being blocked.
Examine CNI Plugin Health and Logs
Most clusters use Calico, Cilium, Weave, or flannel for pod networking. Verify the CNI pods are running in kube-system:
kubectl get pods -n kube-system -l k8s-app=calico-node
Check logs for IP exhaustion or VXLAN errors:
kubectl logs -n kube-system <cni-pod-name>
Look for errors like "IP-AM exhausted" or "VXLAN failure" that would prevent proper pod IP assignment despite the pod showing as Running.
Probe iptables and kube-proxy Rules
The kube-proxy component programs iptables (or IPVS) rules for Service traffic. Missing or corrupted NAT rules can cause packets to drop silently.
Inspect current iptables rules on the node:
sudo iptables-save | grep -i KUBE
Look for KUBE-SVC chains corresponding to your Service. If rules appear missing or corrupted, restart kube-proxy to force rule recreation:
kubectl rollout restart daemonset/kube-proxy -n kube-system
Confirm DNS Resolution with CoreDNS
If pods communicate via DNS names (e.g., my-svc.default.svc.cluster.local), verify CoreDNS is healthy:
kubectl get pods -n kube-system -l k8s-app=kube-dns
Test resolution from within a pod:
kubectl exec -n kube-system <coredns-pod> -- nslookup my-svc
A misconfigured resolv.conf or broken CoreDNS deployment manifests as "no error logs but communication fails" because applications hang waiting for DNS resolution.
Check Node-Level Firewall Rules
Node-level iptables or host firewall configurations can override Kubernetes networking rules. SSH into the worker node and verify no DROP rules are blocking pod CIDR ranges:
sudo iptables -L -v
Check that the node's firewall allows traffic between pod CIDR blocks and the cluster's service network.
Capture Packets with Diagnostic Tools
Use tcpdump within the pod's network namespace to confirm whether packets leave the container and arrive at the destination:
kubectl exec -it pod-a -- tcpdump -i any -nn port <target-port>
Run this on both the source and destination pods simultaneously. If packets leave the source but never arrive at the destination, the issue lies in the intermediate network path (CNI, node firewall, or NetworkPolicy).
Recreate Resources to Clear Stale State
When all else fails, stale iptables entries or cached endpoints may persist despite configuration corrections. Force a fresh network setup by recreating the Service and pods:
kubectl delete svc my-svc
kubectl delete pod pod-a pod-b
kubectl apply -f my-deployment.yaml
kubectl apply -f my-svc.yaml
This clears any lingering state in the CNI or kube-proxy.
Summary
- Verify basic connectivity first with
pingandcurlto isolate Layer 3 vs. Layer 7 issues. - Check Service selectors and port mappings, as mismatches cause empty endpoint lists without obvious errors.
- Review NetworkPolicies that may silently block intra-namespace traffic.
- Inspect CNI health and logs for IP allocation or overlay network failures.
- Validate iptables rules programmed by kube-proxy for Service NAT.
- Test DNS resolution through CoreDNS when using service names.
- Use tcpdump to confirm packet flow before concluding the issue is application-level.
Frequently Asked Questions
Why do my pods show Running but cannot communicate?
Kubernetes marks pods as Running once the container starts and probes succeed, but this status does not guarantee network connectivity. The issue usually resides in the CNI plugin, Service configuration, or NetworkPolicy rules that operate below the pod abstraction layer.
How do I check if a Kubernetes Service is blocking traffic?
Verify the Service has active endpoints using kubectl get endpoints <service-name>. If the endpoint list is empty, the selector labels do not match your pod labels, or the pods are not passing readiness probes. Also check if a NetworkPolicy is restricting ingress to the service endpoints.
What is the fastest way to test pod-to-pod networking in Kubernetes?
Execute a shell in one pod and ping the IP of another pod directly: kubectl exec pod-a -- ping <pod-b-ip>. If this fails, the problem is at the CNI or node level. If it succeeds but service calls fail, investigate the Service and kube-proxy configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →