1. Introduction
A misconfigured AWS security group is one of the most common causes of silent connectivity failures. Unlike an application error, a security group deny produces no error message on the client side — traffic simply never arrives. You'll see timeouts, not refusals, which makes the source of the problem easy to misdiagnose as an application bug.
This guide covers how to diagnose, identify, and fix security group misconfigurations across EC2, EKS, RDS, and load balancer scenarios. If you're debugging a node that won't join your EKS cluster or an EKS pod that can't reach the internet, security groups are usually the first place to check.
2. What a Security Group Misconfiguration Means
AWS security groups are stateful virtual firewalls attached to network interfaces. They evaluate inbound and outbound rules to allow or deny traffic. A misconfiguration means one or more rules are missing, use the wrong protocol or port, or reference the wrong source/destination.
The most common symptoms: connection timeouts from EC2 to RDS, pods unable to reach the internet or internal services, EKS nodes failing to reach the control plane, and health checks on load balancers perpetually failing.
3. Common Causes
- Missing inbound rule on the target resource (e.g. RDS has no rule allowing the app tier)
- Rule references the wrong security group ID or CIDR block
- Port or protocol mismatch (TCP vs UDP, wrong port number)
- Overly restrictive outbound rules blocking return traffic from new connections
- NACLs on the subnet silently blocking traffic that security groups allow
- EKS node security group missing rules for the control plane endpoint on port 443
- ALB security group not allowing health check traffic to reach target instances
- VPC Peering or Transit Gateway routing is correct but security groups block cross-VPC traffic
4. Step-by-Step Diagnosis and Fix
Step 1: Enable VPC Flow Logs to see rejected traffic
The fastest way to confirm a security group block is to look at VPC Flow Logs and filter for REJECT entries. Without Flow Logs, you're guessing.
# Enable flow logs on a VPC (sends to CloudWatch Logs)
aws ec2 create-flow-logs \
--resource-type VPC \
--resource-ids vpc-0abc12345 \
--traffic-type REJECT \
--log-destination-type cloud-watch-logs \
--log-group-name /aws/vpc/flowlogs \
--deliver-logs-permission-arn arn:aws:iam::123456789012:role/FlowLogsRole
# Query CloudWatch Logs Insights for recent REJECTs
# Useful filter: filter @message like /REJECT/ | stats count() by dstPort
Step 2: Inspect security groups on both ends
# List security groups attached to an instance
aws ec2 describe-instances --instance-ids i-0abc12345 --query 'Reservations[].Instances[].SecurityGroups'
# View all inbound rules for a security group
aws ec2 describe-security-groups --group-ids sg-0abc12345 --query 'SecurityGroups[].IpPermissions'
# View outbound rules
aws ec2 describe-security-groups --group-ids sg-0abc12345 --query 'SecurityGroups[].IpPermissionsEgress'
Step 3: Use the AWS Reachability Analyzer
The Reachability Analyzer traces a network path between two resources and tells you exactly which component — security group, NACL, route table, or gateway — is blocking the connection.
# Create a reachability analysis path
aws ec2 create-network-insights-path --source i-0source12345 --destination i-0dest12345 --protocol TCP --destination-port 5432
# Start the analysis
aws ec2 start-network-insights-analysis --network-insights-path-id nip-0abc12345
# Check results (replace with your analysis ID)
aws ec2 describe-network-insights-analyses --network-insights-analysis-ids nia-0abc12345 --query 'NetworkInsightsAnalyses[].{Reachable:NetworkPathFound,ExplainCode:Explanations[0].ExplanationCode}'
Step 4: Add or fix missing security group rules
# Add inbound rule: allow PostgreSQL from an app security group
aws ec2 authorize-security-group-ingress --group-id sg-rds-0abc12345 --protocol tcp --port 5432 --source-group sg-app-0abc12345
# Add inbound rule using a CIDR (less preferred — use SG references where possible)
aws ec2 authorize-security-group-ingress --group-id sg-0abc12345 --protocol tcp --port 443 --cidr 10.0.0.0/8
# Add outbound rule (e.g. if you locked down egress)
aws ec2 authorize-security-group-egress --group-id sg-0abc12345 --protocol tcp --port 443 --cidr 0.0.0.0/0
Step 5: Check EKS-specific security group requirements
EKS clusters have specific security group requirements for node-to-control-plane communication. If an EKS node is failing to join the cluster, start here:
# Get the cluster security group ID
aws eks describe-cluster --name my-cluster --query 'cluster.resourcesVpcConfig.clusterSecurityGroupId'
# Node SG must allow outbound to control plane on 443
# Control plane SG must allow inbound from node SG on 443 and 10250
# Node SG must allow inbound from control plane on 1025-65535 (for webhook traffic)
# Check if the recommended cluster SG rule exists
aws ec2 describe-security-group-rules --filters Name=group-id,Values=sg-cluster12345 --query 'SecurityGroupRules[?FromPort==`443`]'
Step 6: Check and fix NACLs
# Get the NACL for a subnet
aws ec2 describe-network-acls --filters Name=association.subnet-id,Values=subnet-0abc12345 --query 'NetworkAcls[].{ID:NetworkAclId,Entries:Entries}'
# NACLs are stateless — you need both inbound and outbound rules.
# Add an inbound ALLOW rule (rule number 100, TCP 5432)
aws ec2 create-network-acl-entry --network-acl-id acl-0abc12345 --rule-number 100 --protocol tcp --rule-action allow --ingress --cidr-block 10.0.1.0/24 --port-range From=5432,To=5432
5. Verification Steps
# Test connectivity from within the VPC using a bastion or SSM session
# Replace 10.0.2.50:5432 with your target IP and port
nc -zv 10.0.2.50 5432
# Or from a pod in EKS:
kubectl run net-test --image=nicolaka/netshoot --rm -it --restart=Never -- nc -zv my-rds.cluster-abc.us-east-1.rds.amazonaws.com 5432
# Re-run Reachability Analyzer after your fix
aws ec2 start-network-insights-analysis --network-insights-path-id nip-0abc12345
# Should now return: "NetworkPathFound": true
6. Common Mistakes
- Checking only one side — both the source and destination security groups must have compatible rules
- Referencing a security group from a different region or account (cross-account SG references require VPC peering or sharing)
- Forgetting that NACLs block traffic even when SG rules are correct
- Using
0.0.0.0/0for internal-only traffic when a security group reference is more precise and auditable - Updating a security group but not accounting for the few seconds of propagation delay before rules take effect
- Expecting security groups to handle east-west traffic between pods in the same EKS node — that goes through the node's network namespace, not through AWS SGs
7. Prevention Tips
- Always enable VPC Flow Logs in REJECT mode on production VPCs — the storage cost is minimal and the diagnostic value is enormous
- Define security groups in Terraform or CloudFormation so changes go through code review
- Use security group references (not CIDR ranges) for internal traffic — they update automatically when instances are added or removed
- Use AWS Config rules to detect overly permissive security groups (e.g.
0.0.0.0/0on sensitive ports) - Pair with IAM least-privilege reviews — security groups handle network-layer access, IAM handles API-layer access; both need to be correct
- For EKS, use the cluster security group recommendations from the EKS best practices guide
8. FAQ
Why does my connection time out instead of being refused?
A timeout means the traffic was silently dropped — exactly what security groups (and NACLs) do when they deny a connection. A "connection refused" error means the traffic reached the destination but nothing was listening on that port. If you're seeing timeouts, check security groups and NACLs first.
Can I use a security group from one VPC in another VPC?
Not directly. Security group references only work within the same VPC unless you're using VPC Peering with cross-VPC security group referencing enabled, or AWS Resource Access Manager (RAM) to share security groups. In most cross-VPC scenarios, you'll use CIDR ranges or configure separate security groups in each VPC.
I added the inbound rule but it still doesn't work. What else could it be?
Check in order: (1) Is the rule on the correct security group? (2) Is there a NACL on the subnet blocking traffic? (3) Is there a route table entry for the traffic? (4) Is the target service actually running and listening on that port? Use the Reachability Analyzer — it will identify the exact blocking component.
How do security groups interact with Kubernetes network policies?
Security groups operate at the EC2/ENI level (L3/L4), while Kubernetes network policies operate at the pod level using the CNI plugin. On EKS with the VPC CNI, you can optionally use security groups per pod (via Security Groups for Pods feature). Otherwise, security groups apply to the node, and network policies control pod-to-pod traffic within the cluster.
9. Summary
| Symptom | Likely Cause | Fix |
|---|---|---|
| Connection timeout (no error) | Security group or NACL blocking traffic | Enable Flow Logs; run Reachability Analyzer |
| EC2 can't reach RDS | Missing inbound rule on RDS SG | Add inbound rule from app SG on port 5432 |
| EKS node won't join | Node SG missing outbound 443 to control plane | Add outbound rule to cluster SG on 443 |
| ALB health checks failing | Instance SG not allowing ALB SG | Add inbound from ALB SG on app port |
| SG rules look right, still blocked | NACL on subnet is stateless-blocking | Add NACL allow rules for both directions |
Explore More in This Category
Explore more in this category: AWS & Cloud guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.