1. Introduction
AWS has hundreds of services. As a DevOps engineer, you'll actively use maybe a dozen of them on a regular basis. This guide focuses specifically on those: the services that show up in infrastructure decisions every week, the ones whose misconfiguration causes production incidents, and the ones you need to understand well enough to diagnose problems from first principles.
The goal isn't comprehensive AWS knowledge — it's the mental model that lets you understand what's happening in a production environment and fix it when it breaks.
2. IAM and Permissions
IAM (Identity and Access Management) controls who can do what in your AWS account. Every action in AWS — an EC2 instance calling an S3 bucket, a CI/CD pipeline pushing to ECR, a Kubernetes pod reading from Secrets Manager — goes through IAM. Getting IAM wrong is the most common cause of "it works locally but fails in production."
The core concepts: Users (human identities), Roles (machine identities that can be assumed by services), Policies (JSON documents that define allowed actions), and Groups (collections of users with shared policies). In practice, you'll spend most of your time working with roles: EC2 instance profiles, EKS service accounts, and Lambda execution roles.
The principle of least privilege means each role should have exactly the permissions it needs to do its job — no more. In practice, start with the permissions you know you need, deploy, observe what fails, and add permissions incrementally. It's faster than guessing and more secure than broad policies.
For EKS specifically, IRSA (IAM Roles for Service Accounts) lets individual pods assume specific IAM roles without sharing the node's instance profile. This is the right pattern for pod-level AWS access. See How to Use IRSA in EKS for the implementation guide.
# Check what identity a pod is using
kubectl exec -it <pod-name> -- aws sts get-caller-identity
# Simulate whether a role can perform an action
aws iam simulate-principal-policy --policy-source-arn arn:aws:iam::123456789012:role/my-role --action-names s3:GetObject --resource-arns arn:aws:s3:::my-bucket/key
3. VPC Basics
A VPC (Virtual Private Cloud) is your private network inside AWS. Every resource you create — EC2 instances, RDS databases, EKS nodes — lives in a VPC. You define IP address ranges (CIDR blocks), divide them into subnets, control routing with route tables, and filter traffic with security groups and NACLs.
The most important VPC concepts for DevOps work: public vs private subnets (public subnets have a route to the internet gateway; private subnets need a NAT Gateway for outbound internet access), route tables (which control where traffic goes), and VPC endpoints (which let private resources reach AWS services without going through the internet).
For EKS, the standard architecture is: worker nodes in private subnets, a NAT Gateway per availability zone for outbound internet access, and load balancers in public subnets. If your pods can't reach the internet or can't pull images from ECR, the NAT Gateway or route tables are almost always involved.
4. Security Groups and Routing
Security groups are stateful firewalls attached to network interfaces. They control inbound and outbound traffic by protocol, port, and source/destination. A security group misconfiguration produces a silent connection timeout — not a refused connection, not an error message — which makes it one of the hardest problems to diagnose without the right approach.
The diagnostic path: enable VPC Flow Logs and filter for REJECT entries to confirm traffic is being dropped. Use AWS Reachability Analyzer to trace the exact path and identify the blocking rule. Then add the missing inbound or outbound rule. See Fix Security Group Misconfiguration in AWS for the full guide.
Remember that NACLs (Network Access Control Lists) are stateless and operate at the subnet level — they can block traffic even when security groups allow it, and they need explicit rules for both inbound and return traffic.
5. Load Balancers
AWS offers two main load balancer types for DevOps work: ALB (Application Load Balancer) for HTTP/HTTPS traffic with host and path-based routing, and NLB (Network Load Balancer) for TCP/UDP traffic that needs lower latency or the ability to preserve client IP addresses.
In EKS, the AWS Load Balancer Controller creates ALBs from Kubernetes Ingress resources and NLBs from Services of type LoadBalancer. Common issues: the controller doesn't have IAM permissions to create load balancers, the subnets are missing the required tags (kubernetes.io/role/elb), or health checks are failing because the security group doesn't allow ALB traffic to reach the pods. See Fix AWS Load Balancer Not Routing Traffic for the diagnostic steps.
6. EKS and Kubernetes on AWS
EKS (Elastic Kubernetes Service) is AWS's managed Kubernetes service. AWS manages the control plane (API server, etcd, scheduler) — you manage the worker nodes, node groups, add-ons, and the applications running on the cluster. On EKS you typically use managed node groups (EC2 instances that AWS patches and replaces) or Fargate (serverless pods with no node management).
The EKS-specific things to understand: how nodes join the cluster (via the aws-auth ConfigMap or newer EKS access entries), how IAM integrates with Kubernetes (IRSA for pod-level access), how the EBS CSI driver provisions persistent volumes, and how the Load Balancer Controller creates AWS load balancers from Kubernetes resources. See Fix EKS Node Not Joining Cluster if nodes aren't registering correctly.
# Update your kubeconfig to connect to an EKS cluster
aws eks update-kubeconfig --region us-east-1 --name my-cluster
# Check cluster and node status
kubectl get nodes -o wide
kubectl cluster-info
7. CloudWatch Logs and Metrics
CloudWatch is AWS's built-in monitoring and logging service. Container logs from EKS pods are collected by the AWS Fluent Bit DaemonSet and sent to CloudWatch Logs. EC2 system metrics (CPU, memory, disk) appear in CloudWatch Metrics automatically. You can also define custom metrics, create dashboards, and set alarms that notify you when something exceeds a threshold.
The most useful CloudWatch features for day-to-day DevOps work: Log Insights (query your logs with a SQL-like language), metric alarms (alert when CPU exceeds 80% or error rate spikes), and Container Insights (pod-level metrics for EKS). CloudWatch is not as powerful as Prometheus/Grafana for cluster-level observability, but it's the first place to look when you need a quick answer and don't want to set up additional tooling.
# Query CloudWatch Logs Insights from the CLI:
aws logs start-query --log-group-name /aws/eks/my-cluster/application --start-time $(date -d '1 hour ago' +%s) --end-time $(date +%s) --query-string 'fields @timestamp, @message | filter @message like /ERROR/ | limit 50'
8. Common AWS Troubleshooting Paths
The AWS problems that surface most often in production environments and what to check first:
- AccessDenied errors — identify the exact IAM principal making the call (
aws sts get-caller-identityfrom inside the pod or instance), then check if the required action is in the role's policy. See Fix AWS AccessDenied Error. - Connection timeouts to AWS services — security groups and NACLs first, then VPC endpoint configuration if the resource is in a private subnet.
- EKS pods can't reach the internet — check the NAT Gateway exists, the private subnet route table has a 0.0.0.0/0 → NAT Gateway entry, and outbound security group rules allow the traffic.
- Load balancer returning 502/503 — check target group health, verify security groups allow ALB traffic to reach the pods, and check the AWS Load Balancer Controller logs.
- ECR image pull failures — the node's IAM role needs
ecr:GetAuthorizationTokenand repository-level pull permissions. For private clusters, check the VPC endpoint for ECR.
9. Where to Go Next
The AWS & Cloud category covers each of these failure modes in depth with step-by-step diagnostic guides. Once you're comfortable with the services above, the areas that give you the most leverage to improve: infrastructure as code (Terraform or CDK), cost optimisation (right-sizing, reserved instances, Spot), and multi-account strategy for separating production, staging, and development workloads.
Explore More
Explore more in this category: AWS & Cloud guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.