1. Introduction
When an EKS worker node fails to join the cluster, it simply doesn't appear in kubectl get nodes — or it appears as NotReady indefinitely. No pods get scheduled to it, and the cluster behaves as if the node doesn't exist. This is a common pain point when setting up new node groups, migrating to managed node groups, or troubleshooting autoscaling events.
The root causes are almost always in one of three layers: IAM permissions (the node can't call the EKS API to bootstrap), networking (the node can't reach the EKS control plane endpoint), or bootstrap configuration (the node's user data is wrong or the kubelet failed to start). This guide walks through all three with specific commands for both self-managed and managed node groups.
2. What "Node Not Joining" Actually Means
When an EKS worker node starts, it runs a bootstrap script (usually /etc/eks/bootstrap.sh on Amazon Linux 2) that configures the kubelet with the cluster's API endpoint and CA certificate. The kubelet then attempts to register itself with the control plane using a TLS bootstrap certificate. If the registration succeeds, the node appears in kubectl get nodes.
If the node doesn't appear, the bootstrap sequence failed somewhere. The three most common failure points are: the node's IAM role isn't in the aws-auth ConfigMap (or aws-auth is misconfigured), the node can't reach the EKS control plane endpoint, or the bootstrap script itself failed due to wrong parameters.
3. Common Causes
- The node's IAM role is not in the
aws-authConfigMap (self-managed node groups) - The node's IAM role is missing required EKS node policies (
AmazonEKSWorkerNodePolicy,AmazonEC2ContainerRegistryReadOnly,AmazonEKS_CNI_Policy) - The node can't reach the EKS control plane endpoint — VPC, security group, or private endpoint misconfiguration
- The bootstrap script received wrong parameters (wrong cluster name, wrong region)
- The user data script is failing silently — misconfigured or missing
- The node's kubelet failed to start (check systemd logs)
- The EC2 instance doesn't have internet access (or VPC endpoints) needed to pull the bootstrap configuration from EKS
- AMI mismatch — using an AMI not compatible with the cluster's Kubernetes version
- The cluster's
aws-authConfigMap has a YAML syntax error that broke all node authentication
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm the node is not joining
# List all nodes
kubectl get nodes -o wide
# Check if the node appears but is NotReady
kubectl get nodes | grep NotReady
# For a node that doesn't appear at all, check the EC2 console or AWS CLI
aws ec2 describe-instances --filters "Name=tag:kubernetes.io/cluster/<cluster-name>,Values=owned" --query 'Reservations[*].Instances[*].{ID:InstanceId,State:State.Name,Launch:LaunchTime}' --output table
Step 2: Check the aws-auth ConfigMap (most common fix for self-managed nodes)
# View the current aws-auth ConfigMap
kubectl get configmap aws-auth -n kube-system -o yaml
# Expected entry for a node group role:
# mapRoles:
# - rolearn: arn:aws:iam::123456789012:role/my-node-group-role
# username: system:node:{{EC2PrivateDNSName}}
# groups:
# - system:bootstrappers
# - system:nodes
# If your node's IAM role is not listed, add it:
kubectl edit configmap aws-auth -n kube-system
# Or use eksctl (safer — prevents YAML corruption):
eksctl create iamidentitymapping --cluster <cluster-name> --region <region> --arn arn:aws:iam::123456789012:role/my-node-group-role --username system:node:{{EC2PrivateDNSName}} --group system:bootstrappers --group system:nodes
Step 3: Verify the node's IAM role has required policies
# Get the node's IAM role ARN from EC2
aws ec2 describe-instances --instance-ids <instance-id> --query 'Reservations[*].Instances[*].IamInstanceProfile.Arn' --output text
# Get the role name from the instance profile
aws iam get-instance-profile --instance-profile-name <profile-name> --query 'InstanceProfile.Roles[0].RoleName' --output text
# List attached policies on the role
aws iam list-attached-role-policies --role-name <node-role-name>
# Required policies for EKS worker nodes:
# - arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy
# - arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryReadOnly
# - arn:aws:iam::aws:policy/AmazonEKS_CNI_Policy (or custom CNI policy)
# Attach any missing policies:
aws iam attach-role-policy --role-name <node-role-name> --policy-arn arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy
Step 4: Check bootstrap script and kubelet logs on the node
# SSH into the node (or use SSM Session Manager for private nodes)
aws ssm start-session --target <instance-id>
# On the node — check bootstrap log
sudo cat /var/log/cloud-init-output.log | tail -100
sudo journalctl -u kubelet -n 150 --no-pager
# Common kubelet errors to look for:
# "Unable to register node ... Unauthorized" → aws-auth issue
# "Failed to connect to API server" → network/endpoint issue
# "x509: certificate signed by unknown authority" → wrong CA cert / cluster name
# Check bootstrap.sh was called with correct parameters
sudo grep bootstrap /var/lib/cloud/instances/*/user-data.txt 2>/dev/null || sudo cat /var/lib/cloud/instance/user-data.txt
Step 5: Check network connectivity to the EKS control plane
# From the node, test connectivity to the API endpoint
# Get the cluster endpoint:
aws eks describe-cluster --name <cluster-name> --region <region> --query 'cluster.endpoint' --output text
# Example: https://ABCDE1234.gr7.us-east-1.eks.amazonaws.com
# On the node, test HTTPS connectivity:
curl -sk <cluster-endpoint>/healthz
# Expected: ok
# If this fails, check:
# 1. Security group on the node allows outbound HTTPS (443)
# 2. Security group on the cluster control plane allows inbound from node SG
# 3. For private-endpoint-only clusters: VPC endpoints exist for EKS, ECR, S3
# Check EKS cluster security group allows inbound from nodes:
aws ec2 describe-security-groups --group-ids <cluster-sg-id> --query 'SecurityGroups[*].IpPermissions' --output json
Step 6: Check for AMI and Kubernetes version mismatch
# Get cluster Kubernetes version
aws eks describe-cluster --name <cluster-name> --region <region> --query 'cluster.version' --output text
# Get the AMI in use by the node
aws ec2 describe-instances --instance-ids <instance-id> --query 'Reservations[*].Instances[*].ImageId' --output text
# Look up the correct EKS-optimised AMI for your version:
aws ssm get-parameter --name /aws/service/eks/optimized-ami/<k8s-version>/amazon-linux-2/recommended/image_id --region <region> --query Parameter.Value --output text
5. Verification Steps
# After applying fixes, watch for the node to appear
kubectl get nodes -w
# Expected: node transitions from (absent) to Ready
# NAME STATUS ROLES AGE VERSION
# ip-10-0-1-45.ec2.internal Ready <none> 42s v1.29.x
# Confirm the node can accept pods
kubectl describe node <node-name> | grep -A10 "Conditions:"
# Deploy a test pod explicitly to the new node
kubectl run test-pod --image=busybox --overrides='{"spec": {"nodeName": "<node-name>"}}' -- sleep 60
kubectl get pod test-pod -o wide
kubectl delete pod test-pod
6. Common Mistakes
- Editing aws-auth with
kubectl editand introducing a YAML formatting error — always useeksctlor validate YAML before applying - Using the instance profile ARN instead of the IAM role ARN in aws-auth — the
rolearnfield must be the role ARN, not the instance profile ARN - Not providing internet access or VPC endpoints for private clusters — nodes need to reach the EKS API, ECR, and S3 endpoints
- Using a community AMI or custom AMI without the EKS bootstrap script — the AMI must include
/etc/eks/bootstrap.shfor the standard bootstrap process to work - Assuming managed node groups handle aws-auth automatically for all use cases — if you've customised auth, you may need to verify the mapping is correct
- Not checking kubelet logs on the node — the exact error is almost always there and is more specific than anything
kubectlshows
7. Prevention Tips
- Use EKS Managed Node Groups wherever possible — they handle IAM role mapping, AMI compatibility, and bootstrap configuration automatically
- For self-managed nodes, use infrastructure-as-code (Terraform, CloudFormation) to manage aws-auth entries rather than editing the ConfigMap manually
- Pin node group AMIs to specific EKS-optimised AMI versions that match your cluster's Kubernetes version
- For private clusters, set up VPC endpoints for
com.amazonaws.region.eks,com.amazonaws.region.ecr.api,com.amazonaws.region.ecr.dkr, andcom.amazonaws.region.s3before launching nodes - Monitor the
aws-authConfigMap for accidental modifications — any change that corrupts its YAML will prevent all new nodes from joining - Consider migrating to EKS Pod Identity or IRSA for pod-level permissions to reduce reliance on node instance profile permissions
- If using VPC Peering or Transit Gateway, ensure routing tables allow traffic from the node's subnet to the EKS control plane endpoint
8. FAQ
The node appears in kubectl get nodes as NotReady. Is this the same problem?
Not exactly. A NotReady node has registered with the cluster (passed the IAM/auth check) but the kubelet reports it's not healthy — usually because the CNI plugin hasn't finished configuring networking, or the node is under memory/disk pressure. Run kubectl describe node <name> and check the Conditions section for the specific reason.
How do I know which IAM role my node is using?
Run aws ec2 describe-instances --instance-ids <id> --query 'Reservations[*].Instances[*].IamInstanceProfile'. This gives you the instance profile ARN. Then use aws iam get-instance-profile to get the underlying role ARN — that role is what needs to be in aws-auth.
Can I add a node to an EKS cluster without modifying aws-auth?
For Managed Node Groups, yes — AWS handles it. For Fargate profiles, there's no node to add. For self-managed nodes, you must add the IAM role to aws-auth (or use the newer EKS access entries API, which replaces aws-auth in newer cluster versions).
9. Summary
EKS nodes that don't join are almost always blocked by IAM mapping (aws-auth), missing node policies, or network connectivity. The kubelet logs on the node contain the exact error. For self-managed nodes, always use eksctl create iamidentitymapping rather than editing aws-auth directly.
| Symptom | Root cause | Fix |
|---|---|---|
| Node never appears in kubectl get nodes | IAM role not in aws-auth | Add role via eksctl create iamidentitymapping |
| Kubelet logs: Unauthorized | Missing node IAM policies | Attach EKSWorkerNodePolicy, ECRReadOnly, CNI policy |
| Kubelet logs: failed to connect to API | Network / security group issue | Allow outbound 443 from node; check VPC endpoints |
| Bootstrap script fails | Wrong cluster name or region in user data | Fix bootstrap.sh parameters in launch template |
| All nodes stopped joining suddenly | aws-auth ConfigMap corrupted | Restore from backup or fix YAML syntax |
Explore More in This Category
Explore more in this category: AWS & Cloud guides. Browse all DevOps Compass articles or jump to a related area: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.