1. Introduction

DNS resolution failures in Kubernetes can cause a wide range of symptoms that look completely unrelated to DNS: services that can't find each other, applications that fail to connect to databases, HTTP clients that hang indefinitely, and pods that appear healthy but can't reach anything by hostname. When DNS breaks in a cluster, the knock-on effects are significant because nearly all inter-service communication relies on it.

Kubernetes uses CoreDNS (since 1.13) to handle DNS for pods. If CoreDNS is degraded, if pods have incorrect resolv.conf, or if the cluster's DNS policies are misconfigured, service discovery fails. This guide covers the full diagnostic path from confirming DNS is broken through to fixing CoreDNS, ndots configuration, and search domain issues.

2. How Kubernetes DNS Works

Every pod in a Kubernetes cluster gets a /etc/resolv.conf injected by the kubelet. By default it looks like:

nameserver 10.96.0.10       # CoreDNS ClusterIP (kube-dns Service)
search default.svc.cluster.local svc.cluster.local cluster.local
options ndots:5

This means pod DNS queries go to CoreDNS first. CoreDNS resolves cluster-internal names (like my-service.my-namespace.svc.cluster.local) and forwards external names to the upstream DNS configured in the cluster (typically the VPC DNS or a public resolver).

3. Common Causes

4. Step-by-Step Diagnosis and Fix

Step 1: Confirm DNS is actually broken

# Run a DNS test pod in the affected namespace
kubectl run dns-test --image=busybox:1.35 --rm -it --restart=Never -- sh

# Inside the pod:
nslookup kubernetes.default        # should return the API server IP
nslookup kubernetes.default.svc.cluster.local  # FQDN form
nslookup google.com                # external DNS

# Test with a specific nameserver:
nslookup kubernetes.default 10.96.0.10  # query CoreDNS directly

# Check /etc/resolv.conf inside the pod:
cat /etc/resolv.conf

Step 2: Check CoreDNS pod health

# Check CoreDNS pod status
kubectl get pods -n kube-system -l k8s-app=kube-dns

# If pods are CrashLoopBackOff or not Running:
kubectl describe pod -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=50

# Check if CoreDNS is being OOMKilled (common cause of DNS flaps)
kubectl get pod -n kube-system -l k8s-app=kube-dns   -o jsonpath='{.items[*].status.containerStatuses[*].lastState}'

# If OOMKilled, increase CoreDNS memory limit:
kubectl edit deployment coredns -n kube-system
# Increase resources.limits.memory from 170Mi to 300Mi+

Step 3: Verify the kube-dns Service and endpoints

# Check the kube-dns Service
kubectl get service kube-dns -n kube-system
# NAME       TYPE        CLUSTER-IP   EXTERNAL-IP   PORT(S)
# kube-dns   ClusterIP   10.96.0.10   <none>        53/UDP,53/TCP

# Check it has endpoints pointing to CoreDNS pods
kubectl get endpoints kube-dns -n kube-system
# ENDPOINTS should NOT be <none>

# If endpoints are empty:
kubectl describe service kube-dns -n kube-system
# Look at Selector — should match CoreDNS pod labels

Step 4: Check CoreDNS ConfigMap for upstream issues

# View the CoreDNS configuration
kubectl get configmap coredns -n kube-system -o yaml

# A typical healthy Corefile looks like:
# .:53 {
#     errors
#     health { lameduck 5s }
#     ready
#     kubernetes cluster.local in-addr.arpa ip6.arpa {
#        pods insecure
#        fallthrough in-addr.arpa ip6.arpa
#     }
#     prometheus :9153
#     forward . /etc/resolv.conf {    <-- forwards external queries to VPC DNS
#        max_concurrent 1000
#     }
#     cache 30
#     loop
#     reload
#     loadbalance
# }

# If the forward directive points to unreachable DNS servers:
# Change to use a reliable upstream:
forward . 8.8.8.8 8.8.4.4 {
    max_concurrent 1000
}

Step 5: Check for network policies blocking DNS

"# List network policies in the affected namespace
kubectl get networkpolicy -n <namespace>

# A restrictive NetworkPolicy can block UDP 53 traffic from pods to CoreDNS
# The following policy allows DNS (add to your namespace's policy):
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-dns
  namespace: <your-namespace>
spec:
  podSelector: {}
  egress:
  - ports:
    - protocol: UDP
      port: 53
    - protocol: TCP
      port: 53
  policyTypes:
  - Egress

Step 6: Fix pod DNS policy misconfiguration

# Check what dnsPolicy your pods are using
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.dnsPolicy}'

# ClusterFirst (default) = use CoreDNS
# Default = use node's /etc/resolv.conf (not cluster DNS)
# None = use only dnsConfig

# If set to Default or None inadvertently, fix in the Deployment spec:
spec:
  template:
    spec:
      dnsPolicy: ClusterFirst  # restore to default

Step 7: Fix ndots and search domain issues for external DNS

# If external domain lookups are slow or failing due to ndots:5,
# override per-deployment with dnsConfig:
spec:
  template:
    spec:
      dnsPolicy: ClusterFirst
      dnsConfig:
        options:
        - name: ndots
          value: "1"    # reduces unnecessary search domain iterations
        - name: timeout
          value: "2"
        - name: attempts
          value: "3"

5. Verification Steps

# After fixing, run the DNS test again:
kubectl run dns-verify --image=busybox:1.35 --rm -it --restart=Never -- sh

# Inside:
nslookup kubernetes.default                    # should resolve
nslookup my-service.my-namespace.svc.cluster.local  # your service
nslookup google.com                            # external

# Check CoreDNS metrics (if Prometheus is available):
kubectl port-forward -n kube-system service/kube-dns 9153:9153
curl http://localhost:9153/metrics | grep coredns_dns_request_duration

6. Common Mistakes

7. Prevention Tips

8. FAQ

My pods can reach Services by IP but not by name. Is this a DNS issue?

Yes — classic symptom. The network connectivity is fine, but DNS resolution is broken. Start with nslookup kubernetes.default from inside the pod. If that fails, CoreDNS isn't reachable. If it succeeds for kubernetes.default but fails for your service name, there might be a namespace issue — try the fully qualified name: nslookup my-service.my-namespace.svc.cluster.local.

DNS works for some pods but not others in the same cluster. Why?

Check: (1) Is there a NetworkPolicy in the affected namespace that blocks egress to port 53? (2) Do the affected pods have a different dnsPolicy? (3) Are the affected pods on a specific node where the node's iptables rules are corrupted? Run the DNS test pod in the exact same namespace and node as the failing pods.

External DNS queries are very slow (2-3 seconds) even though CoreDNS is healthy. What's happening?

This is almost always the ndots:5 problem. With 5 search domains and ndots:5, a query for google.com generates 6 DNS queries before getting a result (one for each search domain, then the bare query). Reduce ndots to 1 or 2 with dnsConfig.options in your pod spec. The search domains are still used for short names like my-service.

9. Summary

SymptomCauseFix
All DNS fails in podCoreDNS down or unreachableCheck CoreDNS pods; check kube-dns endpoints
Cluster DNS fails, external worksCoreDNS misconfiguredReview CoreDNS ConfigMap kubernetes plugin
Intermittent DNS failuresCoreDNS OOMKilledIncrease CoreDNS memory limits
DNS fails in one namespace onlyNetworkPolicy blocking port 53Add egress allow rule for UDP/TCP 53
External DNS slow (2-3s)ndots:5 causing extra lookupsSet ndots:1 in pod dnsConfig

Explore More in This Category

Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.