1. Introduction
When one Kubernetes pod can't reach another Service inside the same cluster, the failure is invisible from outside. The cluster appears healthy — all pods are Running, all deployments are stable — but your application returns errors, timeouts, or 500s because an internal dependency is unreachable. This is one of the hardest problems to diagnose because it can look like an application bug when it's actually a networking or configuration issue.
Internal service communication in Kubernetes relies on three things working correctly: DNS resolution (to turn a Service name into a ClusterIP), kube-proxy (to translate the ClusterIP to a pod IP), and network connectivity (for the packet to actually reach the pod). If any of these breaks, internal communication fails.
2. What Internal Service Communication Means
When pod A makes a request to http://my-service:8080, the following happens: DNS resolves my-service to the Service's ClusterIP (e.g. 10.96.42.100), kube-proxy (via iptables or IPVS) intercepts the connection to that ClusterIP and rewrites it to a ready pod IP, and the packet travels through the CNI network to reach the target pod. A failure anywhere in this chain causes the connection to fail.
3. Common Causes
- Service selector doesn't match any pod labels — Service has no endpoints
- Service port/targetPort mismatch — traffic goes to the wrong port on the pod
- NetworkPolicy blocks ingress to the target pods or egress from the source pods
- kube-proxy is not running or is in a failed state on the node
- DNS resolution of the Service name is failing (see DNS troubleshooting guide)
- Pod is using a different namespace than the Service but not using the FQDN
- Service is headless (
clusterIP: None) and the application expects a stable IP - Target pods are running but failing readiness probes — not receiving traffic
- CNI plugin is misconfigured or pod-to-pod routing is broken on the node
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm the failure and isolate the layer
# From the source pod, test step by step:
kubectl exec -it <source-pod> -n <namespace> -- sh
# 1. DNS: resolve the service name
nslookup my-service # short name
nslookup my-service.my-namespace.svc.cluster.local # FQDN
# 2. ClusterIP: if DNS works, note the IP and test directly
# (get the ClusterIP from: kubectl get service my-service)
curl http://10.96.42.100:8080/health
# 3. Pod IP: test directly to a pod IP (bypasses kube-proxy)
# (get a pod IP from: kubectl get pod <pod> -o jsonpath='{.status.podIP}')
curl http://10.244.1.5:8080/health
Step 2: Check the Service selector and endpoints
# Verify the Service has endpoints
kubectl get endpoints my-service -n <namespace>
# ENDPOINTS should show pod IPs like: 10.244.1.5:8080, 10.244.2.3:8080
# If <none>: the selector doesn't match any pods
kubectl describe service my-service -n <namespace>
# Note the Selector field
# Check pods that should match:
kubectl get pods -n <namespace> -l app=my-app # use your selector
kubectl get pods -n <namespace> --show-labels
# Common fix: label mismatch
# Service selector: app: my-app
# Pod label: app: myapp (missing hyphen)
kubectl label pod <pod-name> -n <namespace> app=my-app --overwrite
Step 3: Check port configuration
# View the Service port configuration
kubectl get service my-service -n <namespace> -o yaml | grep -A 10 "ports:"
# Example of a misconfiguration:
# ports:
# - port: 80 <-- Service port (what you call)
# targetPort: 8080 <-- Pod port (what the container listens on)
# If the pod listens on 3000 but targetPort is 8080, connection fails
# Verify what port the pod is listening on:
kubectl exec -it <target-pod> -n <namespace> -- ss -tlnp
# or: netstat -tlnp
# Fix: correct the targetPort in the Service
kubectl edit service my-service -n <namespace>
# Update targetPort to match what the pod is listening on
Step 4: Check NetworkPolicy
# List all NetworkPolicies in relevant namespaces
kubectl get networkpolicy -n <source-namespace>
kubectl get networkpolicy -n <target-namespace>
# A NetworkPolicy that denies all ingress to the target service:
kubectl describe networkpolicy -n <target-namespace>
# If you have a default-deny policy, add an allow rule:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-from-source
namespace: <target-namespace>
spec:
podSelector:
matchLabels:
app: my-service-pods
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: <source-namespace>
- podSelector:
matchLabels:
app: source-app
policyTypes:
- Ingress
Step 5: Check kube-proxy health
# kube-proxy runs as a DaemonSet on every node
kubectl get daemonset kube-proxy -n kube-system
kubectl get pods -n kube-system -l k8s-app=kube-proxy
# Check kube-proxy logs on the affected node:
NODE=$(kubectl get pod <source-pod> -n <namespace> -o jsonpath='{.spec.nodeName}')
kubectl logs -n kube-system $(kubectl get pod -n kube-system -l k8s-app=kube-proxy --field-selector spec.nodeName=$NODE -o jsonpath='{.items[0].metadata.name}') --tail=50
# If kube-proxy is failing, restart the DaemonSet pod on that node:
kubectl delete pod -n kube-system $(kubectl get pod -n kube-system -l k8s-app=kube-proxy --field-selector spec.nodeName=$NODE -o jsonpath='{.items[0].metadata.name}')
Step 6: Test cross-namespace communication
# Services in other namespaces require the FQDN:
# <service-name>.<namespace>.svc.cluster.local
# Wrong (if source is in a different namespace):
curl http://my-service:8080/ # only works in same namespace
# Correct:
curl http://my-service.backend.svc.cluster.local:8080/
# Test from the source pod:
kubectl exec -it <source-pod> -n <source-namespace> -- curl http://my-service.target-namespace.svc.cluster.local:8080/health
Step 7: Debug pod-to-pod connectivity directly
# If ClusterIP fails but direct pod IP works = kube-proxy issue
# If direct pod IP also fails = CNI/node routing issue
# Test direct pod-to-pod connectivity:
TARGET_POD_IP=$(kubectl get pod <target-pod> -n <namespace> -o jsonpath='{.status.podIP}')
kubectl exec -it <source-pod> -n <namespace> -- curl http://$TARGET_POD_IP:8080/health
# If this fails: check CNI plugin pods
kubectl get pods -n kube-system | grep -E "calico|flannel|weave|cilium|aws-node"
kubectl logs -n kube-system <cni-pod-name> --tail=50
5. Verification Steps
# After fixing, confirm end-to-end:
kubectl exec -it <source-pod> -n <namespace> -- curl -v http://my-service:8080/health
# Expected: HTTP 200 response
# Check endpoints are populated:
kubectl get endpoints my-service -n <namespace>
# Should show pod IPs
# Run a sustained test to catch intermittent failures:
kubectl exec -it <source-pod> -n <namespace> -- sh -c 'for i in $(seq 1 10); do curl -s -o /dev/null -w "%{http_code}
" http://my-service:8080/health; done'
# All 10 should return 200
6. Common Mistakes
- Assuming "pod is Running = service is reachable" — a pod can be Running but failing readiness probes, meaning it's excluded from Service endpoints
- Not using the FQDN for cross-namespace communication — short names only work within the same namespace
- Creating a NetworkPolicy in one namespace but forgetting that the target namespace also needs an allow rule
- Checking Service port (what clients call) instead of targetPort (what the pod listens on) when debugging
- Overlooking headless Services — a Service with
clusterIP: Nonedoesn't use kube-proxy; clients must handle the pod IPs themselves
7. Prevention Tips
- Always test Service connectivity after deploying a new Service — run a debug pod in the source namespace and curl the Service
- Use consistent labelling conventions so Service selectors reliably match pods
- When adding NetworkPolicies, start with logging-only rules to understand traffic patterns before applying deny policies
- Add a basic DNS connectivity test to your pod readiness probe scripts
- For complex multi-service applications, include integration tests in CI that verify inter-service communication using
kindor a staging cluster - If you're using a service mesh (Istio, Linkerd), verify that sidecar injection is consistent — mixed-injection namespaces can cause silent communication failures
8. FAQ
The Service has endpoints but connections still time out. What else could it be?
If endpoints exist, kube-proxy knows where to route traffic, but the connection still times out: (1) The target pod is accepting connections but not responding — check pod logs, (2) A NetworkPolicy is blocking the specific port or source, (3) kube-proxy on the source node is unhealthy — the iptables rules aren't being applied, (4) The CNI plugin is misconfigured for that pod's network namespace. Test direct pod IP connectivity to isolate between kube-proxy and CNI issues.
Calls to a Service work from some namespaces but not others. Why?
Almost certainly a NetworkPolicy issue. Check if the target namespace has any NetworkPolicy with policyTypes: [Ingress] that doesn't include an allow rule from the source namespace. You can also verify by checking if direct pod IP access works — if pod IP works but ClusterIP doesn't, the issue is in kube-proxy; if pod IP also fails from that specific namespace, it's NetworkPolicy.
Everything was working, then after a deployment it stopped. What changed?
Check: (1) Did the pod labels change? A Deployment update might have changed the pod template labels, breaking the Service selector match. (2) Did the containerPort change? The Service's targetPort may no longer match. (3) Did a NetworkPolicy get applied as part of the deployment? Use kubectl rollout undo to revert and confirm whether the deployment is the cause.
9. Summary
| Test result | Diagnoses | Fix |
|---|---|---|
| DNS fails for Service name | CoreDNS or dnsPolicy issue | See DNS Resolution Issues guide |
| DNS works, ClusterIP times out | No endpoints or kube-proxy issue | Check Service selector; check kube-proxy pods |
| ClusterIP fails, pod IP works | kube-proxy iptables broken | Restart kube-proxy pod on affected node |
| Pod IP also fails | CNI / node routing issue | Check CNI pods; check node network interfaces |
| Works in namespace A, fails in B | NetworkPolicy blocking ingress | Add allow rule for source namespace/pods |
Explore More in This Category
Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.