1. Introduction

When an EKS worker node fails to join the cluster, it simply doesn't appear in kubectl get nodes — or it appears as NotReady indefinitely. No pods get scheduled to it, and the cluster behaves as if the node doesn't exist. This is a common pain point when setting up new node groups, migrating to managed node groups, or troubleshooting autoscaling events.

The root causes are almost always in one of three layers: IAM permissions (the node can't call the EKS API to bootstrap), networking (the node can't reach the EKS control plane endpoint), or bootstrap configuration (the node's user data is wrong or the kubelet failed to start). This guide walks through all three with specific commands for both self-managed and managed node groups.

2. What "Node Not Joining" Actually Means

When an EKS worker node starts, it runs a bootstrap script (usually /etc/eks/bootstrap.sh on Amazon Linux 2) that configures the kubelet with the cluster's API endpoint and CA certificate. The kubelet then attempts to register itself with the control plane using a TLS bootstrap certificate. If the registration succeeds, the node appears in kubectl get nodes.

If the node doesn't appear, the bootstrap sequence failed somewhere. The three most common failure points are: the node's IAM role isn't in the aws-auth ConfigMap (or aws-auth is misconfigured), the node can't reach the EKS control plane endpoint, or the bootstrap script itself failed due to wrong parameters.

3. Common Causes

4. Step-by-Step Diagnosis and Fix

Step 1: Confirm the node is not joining

# List all nodes
kubectl get nodes -o wide

# Check if the node appears but is NotReady
kubectl get nodes | grep NotReady

# For a node that doesn't appear at all, check the EC2 console or AWS CLI
aws ec2 describe-instances   --filters "Name=tag:kubernetes.io/cluster/<cluster-name>,Values=owned"   --query 'Reservations[*].Instances[*].{ID:InstanceId,State:State.Name,Launch:LaunchTime}'   --output table

Step 2: Check the aws-auth ConfigMap (most common fix for self-managed nodes)

# View the current aws-auth ConfigMap
kubectl get configmap aws-auth -n kube-system -o yaml

# Expected entry for a node group role:
# mapRoles:
# - rolearn: arn:aws:iam::123456789012:role/my-node-group-role
#   username: system:node:{{EC2PrivateDNSName}}
#   groups:
#   - system:bootstrappers
#   - system:nodes

# If your node's IAM role is not listed, add it:
kubectl edit configmap aws-auth -n kube-system

# Or use eksctl (safer — prevents YAML corruption):
eksctl create iamidentitymapping   --cluster <cluster-name>   --region <region>   --arn arn:aws:iam::123456789012:role/my-node-group-role   --username system:node:{{EC2PrivateDNSName}}   --group system:bootstrappers   --group system:nodes

Step 3: Verify the node's IAM role has required policies

# Get the node's IAM role ARN from EC2
aws ec2 describe-instances   --instance-ids <instance-id>   --query 'Reservations[*].Instances[*].IamInstanceProfile.Arn'   --output text

# Get the role name from the instance profile
aws iam get-instance-profile   --instance-profile-name <profile-name>   --query 'InstanceProfile.Roles[0].RoleName'   --output text

# List attached policies on the role
aws iam list-attached-role-policies   --role-name <node-role-name>

# Required policies for EKS worker nodes:
# - arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy
# - arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryReadOnly
# - arn:aws:iam::aws:policy/AmazonEKS_CNI_Policy (or custom CNI policy)

# Attach any missing policies:
aws iam attach-role-policy   --role-name <node-role-name>   --policy-arn arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy

Step 4: Check bootstrap script and kubelet logs on the node

# SSH into the node (or use SSM Session Manager for private nodes)
aws ssm start-session --target <instance-id>

# On the node — check bootstrap log
sudo cat /var/log/cloud-init-output.log | tail -100
sudo journalctl -u kubelet -n 150 --no-pager

# Common kubelet errors to look for:
# "Unable to register node ... Unauthorized"  → aws-auth issue
# "Failed to connect to API server"           → network/endpoint issue
# "x509: certificate signed by unknown authority" → wrong CA cert / cluster name

# Check bootstrap.sh was called with correct parameters
sudo grep bootstrap /var/lib/cloud/instances/*/user-data.txt 2>/dev/null ||   sudo cat /var/lib/cloud/instance/user-data.txt

Step 5: Check network connectivity to the EKS control plane

# From the node, test connectivity to the API endpoint
# Get the cluster endpoint:
aws eks describe-cluster   --name <cluster-name>   --region <region>   --query 'cluster.endpoint'   --output text
# Example: https://ABCDE1234.gr7.us-east-1.eks.amazonaws.com

# On the node, test HTTPS connectivity:
curl -sk <cluster-endpoint>/healthz
# Expected: ok

# If this fails, check:
# 1. Security group on the node allows outbound HTTPS (443)
# 2. Security group on the cluster control plane allows inbound from node SG
# 3. For private-endpoint-only clusters: VPC endpoints exist for EKS, ECR, S3

# Check EKS cluster security group allows inbound from nodes:
aws ec2 describe-security-groups   --group-ids <cluster-sg-id>   --query 'SecurityGroups[*].IpPermissions'   --output json

Step 6: Check for AMI and Kubernetes version mismatch

# Get cluster Kubernetes version
aws eks describe-cluster   --name <cluster-name>   --region <region>   --query 'cluster.version'   --output text

# Get the AMI in use by the node
aws ec2 describe-instances   --instance-ids <instance-id>   --query 'Reservations[*].Instances[*].ImageId'   --output text

# Look up the correct EKS-optimised AMI for your version:
aws ssm get-parameter   --name /aws/service/eks/optimized-ami/<k8s-version>/amazon-linux-2/recommended/image_id   --region <region>   --query Parameter.Value   --output text

5. Verification Steps

# After applying fixes, watch for the node to appear
kubectl get nodes -w

# Expected: node transitions from (absent) to Ready
# NAME                          STATUS   ROLES    AGE   VERSION
# ip-10-0-1-45.ec2.internal     Ready    <none>   42s   v1.29.x

# Confirm the node can accept pods
kubectl describe node <node-name> | grep -A10 "Conditions:"

# Deploy a test pod explicitly to the new node
kubectl run test-pod --image=busybox   --overrides='{"spec": {"nodeName": "<node-name>"}}'   -- sleep 60
kubectl get pod test-pod -o wide
kubectl delete pod test-pod

6. Common Mistakes

7. Prevention Tips

8. FAQ

The node appears in kubectl get nodes as NotReady. Is this the same problem?

Not exactly. A NotReady node has registered with the cluster (passed the IAM/auth check) but the kubelet reports it's not healthy — usually because the CNI plugin hasn't finished configuring networking, or the node is under memory/disk pressure. Run kubectl describe node <name> and check the Conditions section for the specific reason.

How do I know which IAM role my node is using?

Run aws ec2 describe-instances --instance-ids <id> --query 'Reservations[*].Instances[*].IamInstanceProfile'. This gives you the instance profile ARN. Then use aws iam get-instance-profile to get the underlying role ARN — that role is what needs to be in aws-auth.

Can I add a node to an EKS cluster without modifying aws-auth?

For Managed Node Groups, yes — AWS handles it. For Fargate profiles, there's no node to add. For self-managed nodes, you must add the IAM role to aws-auth (or use the newer EKS access entries API, which replaces aws-auth in newer cluster versions).

9. Summary

EKS nodes that don't join are almost always blocked by IAM mapping (aws-auth), missing node policies, or network connectivity. The kubelet logs on the node contain the exact error. For self-managed nodes, always use eksctl create iamidentitymapping rather than editing aws-auth directly.

SymptomRoot causeFix
Node never appears in kubectl get nodesIAM role not in aws-authAdd role via eksctl create iamidentitymapping
Kubelet logs: UnauthorizedMissing node IAM policiesAttach EKSWorkerNodePolicy, ECRReadOnly, CNI policy
Kubelet logs: failed to connect to APINetwork / security group issueAllow outbound 443 from node; check VPC endpoints
Bootstrap script failsWrong cluster name or region in user dataFix bootstrap.sh parameters in launch template
All nodes stopped joining suddenlyaws-auth ConfigMap corruptedRestore from backup or fix YAML syntax

Explore More in This Category

Explore more in this category: AWS & Cloud guides. Browse all DevOps Compass articles or jump to a related area: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.