1. Introduction
When an EKS pod can't reach the internet — whether to download packages, call an external API, or pull data from a public endpoint — the failure is usually silent. The pod starts successfully, runs, but any outbound call to an internet address hangs or times out. The pod itself is healthy; the problem is in the network path between the pod and the outside world.
On EKS, pods in private subnets don't have direct internet access. They need a NAT Gateway in a public subnet, correct route tables, and properly configured security groups. If any piece of this chain is missing or misconfigured, all outbound internet traffic silently drops. This guide walks through each layer.
2. What This Error Means
EKS worker nodes in private subnets (the recommended setup) don't have public IP addresses. For outbound internet access, traffic must flow: pod → node → NAT Gateway (in public subnet) → Internet Gateway → internet. If your pod can reach other pods and internal services but not external IPs, the failure is almost always in the NAT Gateway layer, the route table, or the security group's outbound rules.
3. Common Causes
- No NAT Gateway exists in the VPC or it's in a different availability zone
- The private subnet's route table doesn't have a
0.0.0.0/0route to the NAT Gateway - Security group outbound rules on the nodes restrict all egress (common in hardened environments)
- The public subnet's route table doesn't route
0.0.0.0/0to the Internet Gateway - DNS resolution failing —
/etc/resolv.confin the pod points to a VPC DNS server that can't resolve public names - EKS VPC CNI is using custom networking and pods are in a secondary subnet without NAT
- iptables rules on the node (applied by kube-proxy or the CNI) are blocking egress
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm the pod can't reach the internet
# Run a quick connectivity test from inside the pod
kubectl exec -it <pod-name> -n <namespace> -- sh -c "curl -s --connect-timeout 5 https://httpbin.org/ip || echo 'FAILED'"
# Test DNS resolution separately
kubectl exec -it <pod-name> -n <namespace> -- nslookup google.com
# If DNS fails, check CoreDNS
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=30
Step 2: Check the route table on the private subnet
# Find the node's subnet
NODE_NAME=<your-node-name>
aws ec2 describe-instances --filters "Name=private-dns-name,Values=$(kubectl get node $NODE_NAME -o jsonpath='{.spec.providerID}' | cut -d'/' -f5)" --query 'Reservations[].Instances[].SubnetId' --output text
# Get the route table for that subnet
SUBNET_ID=subnet-0abc12345
aws ec2 describe-route-tables --filters "Name=association.subnet-id,Values=$SUBNET_ID" --query 'RouteTables[].Routes'
# You need: DestinationCidrBlock=0.0.0.0/0, NatGatewayId=nat-xxxxxxxx
# If this entry is missing, add it:
aws ec2 create-route --route-table-id rtb-0abc12345 --destination-cidr-block 0.0.0.0/0 --nat-gateway-id nat-0abc12345
Step 3: Verify the NAT Gateway is healthy and in a public subnet
# List all NAT Gateways
aws ec2 describe-nat-gateways --query 'NatGateways[].{ID:NatGatewayId,State:State,Subnet:SubnetId}'
# Check that the NAT Gateway is in a PUBLIC subnet
# (a subnet with an IGW route)
aws ec2 describe-route-tables --filters "Name=association.subnet-id,Values=<nat-gw-subnet-id>" --query 'RouteTables[].Routes[?GatewayId!=null && starts_with(GatewayId, `igw-`)]'
# If no NAT Gateway exists, create one in a public subnet
aws ec2 create-nat-gateway --subnet-id subnet-public-0abc12345 --allocation-id eipalloc-0abc12345 # allocate an EIP first if needed
Step 4: Check security group outbound rules
By default, security groups allow all outbound traffic. But in hardened environments, outbound rules may be restricted. If your nodes have custom egress rules, verify they allow the traffic you need.
# Check node security group outbound rules
aws ec2 describe-security-groups --group-ids sg-node-0abc12345 --query 'SecurityGroups[].IpPermissionsEgress'
# If there's no 0.0.0.0/0 allow rule for port 443, add it:
aws ec2 authorize-security-group-egress --group-id sg-node-0abc12345 --protocol tcp --port 443 --cidr 0.0.0.0/0
Step 5: Check VPC DNS settings
# Verify VPC DNS support is enabled
VPC_ID=vpc-0abc12345
aws ec2 describe-vpc-attribute --vpc-id $VPC_ID --attribute enableDnsSupport
aws ec2 describe-vpc-attribute --vpc-id $VPC_ID --attribute enableDnsHostnames
# Both should return: "Value": true
# If not:
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-support
# Check CoreDNS ConfigMap for upstream DNS settings
kubectl get configmap coredns -n kube-system -o yaml
Step 6: Use a debug pod to isolate the issue
# Run a dedicated network debug pod
kubectl run netshoot --image=nicolaka/netshoot --rm -it --restart=Never -- bash
# Inside the pod, test systematically:
# 1. DNS
nslookup google.com 169.254.20.10 # CoreDNS ClusterIP
# 2. IP connectivity to a known public IP (bypasses DNS)
ping -c 3 8.8.8.8
# 3. HTTPS to a public endpoint
curl -v --max-time 10 https://httpbin.org/ip
# 4. Check routing table inside the pod
ip route show
5. Verification Steps
# After fixing, re-test from an affected pod
kubectl exec -it <pod-name> -n <namespace> -- curl -s --connect-timeout 10 https://httpbin.org/ip
# Expected output: {"origin": "x.x.x.x"}
# The IP shown will be the NAT Gateway's Elastic IP
# Confirm the NAT Gateway is being used
aws ec2 describe-nat-gateways --query 'NatGateways[?State==`available`].NatGatewayAddresses[].PublicIp'
6. Common Mistakes
- Creating a NAT Gateway but forgetting to update the route table in the private subnet
- Placing the NAT Gateway in a private subnet instead of a public subnet
- Having one NAT Gateway but nodes spread across multiple AZs — use one NAT GW per AZ to avoid cross-AZ traffic costs and single points of failure
- Confusing the Internet Gateway route (for public subnets) with the NAT Gateway route (for private subnets)
- Assuming that because a node can ping an internal IP, it can also reach the internet — these use different network paths
7. Prevention Tips
- Deploy one NAT Gateway per availability zone for production workloads — this eliminates cross-AZ traffic costs and AZ-level single points of failure
- Use Terraform or CloudFormation to provision NAT Gateways and route tables so configurations are version-controlled
- If pods only need to reach specific AWS services, use VPC endpoints instead of a NAT Gateway — cheaper and more secure
- Regularly run the Reachability Analyzer against critical connectivity paths as part of infrastructure testing
- Monitor NAT Gateway bandwidth and connection metrics in CloudWatch — a saturated NAT Gateway can also cause connectivity issues
8. FAQ
The pod can reach internal services but not external IPs. What does this tell me?
Internal traffic (pod-to-pod, pod-to-RDS) typically goes through the VPC's internal routing and doesn't need a NAT Gateway. If internal works but external doesn't, the problem is specifically in the path from the node to the NAT Gateway, or the NAT Gateway itself is missing or misconfigured. Start with the route table check.
DNS resolves but HTTPS connections still time out. Why?
DNS uses UDP on port 53 and may have a different path than TCP port 443. If DNS works (resolving to a public IP) but HTTPS times out, the issue is likely in the security group outbound rules (port 443 may be blocked) or the route table (traffic reaches the node but has no path to the NAT Gateway).
Do I need a NAT Gateway if my nodes are in a public subnet?
No. Nodes in public subnets with public IPs can reach the internet directly through the Internet Gateway. But EKS best practice is to run worker nodes in private subnets for security reasons, in which case a NAT Gateway is required for outbound internet access.
9. Summary
| Symptom | Cause | Fix |
|---|---|---|
| All outbound internet fails | No NAT Gateway or missing route | Create NAT GW in public subnet; add 0.0.0.0/0 route |
| DNS fails, IP direct works | CoreDNS misconfiguration | Check CoreDNS ConfigMap and VPC DNS attributes |
| HTTPS fails, HTTP works | Security group blocking port 443 | Add outbound TCP 443 rule to node SG |
| Works in some pods, not others | Pods in different subnets | Check route tables on each subnet independently |
| Intermittent failures | NAT GW bandwidth saturation | Add NAT GW per AZ; scale compute; use VPC endpoints |
Explore More in This Category
Explore more in this category: AWS & Cloud guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.