Networking problems in AWS are often invisible until something stops working — a pod can't pull from ECR, a service can't reach the database, or a deployment stalls because no nodes are available in the right availability zone. This category covers the networking and security decisions that affect production reliability: choosing between VPC Peering and Transit Gateway, configuring IRSA for pod-level AWS access without static credentials, and diagnosing why pods can't be scheduled due to topology constraints.
Troubleshooting
Network-related failures that prevent pods from running or reaching their dependencies.
Fix EKS worker nodes that won't join — aws-auth ConfigMap, IAM policies, bootstrap errors, VPC endpoints, and security group configuration.
Diagnose and fix security group rules causing connection timeouts — VPC flow logs, NACLs, and Reachability Analyzer.
Fix EKS pods with no outbound connectivity — NAT Gateway, route tables, DNS, and CNI plugin.
Resolve ALB/NLB routing failures — target health checks, security group rules, and listener configuration.
Fix Nginx 404s on SPA route refresh — try_files for React, Vue, Angular in standalone, Docker, and Kubernetes.
Ingress not routing traffic — Ingress controllers, IngressClass, TLS secrets, backend endpoints, and path types.
Fix XFF header not passing real client IP — Nginx, AWS ALB/NLB Proxy Protocol, and Ingress Controller ConfigMap.
CoreDNS failures, OOMKilled DNS pods, port 53 NetworkPolicy blocks, dnsPolicy issues, and ndots latency.
Pods can't reach internal Services — selector, port, NetworkPolicy, kube-proxy, and cross-namespace FQDN.
Diagnose pods stuck in Pending — covering insufficient resources, taints, node selectors, affinity rules, and PVC binding failures.
Fix expired tokens, IAM permission errors, cross-account access issues, and misconfigured credential helpers for AWS ECR.
Guides
Architecture decisions and implementation guides for secure AWS networking.
All Categories
Every article on DevOps Compass is organized into a focused category.
Pod lifecycle, workloads, probes, scheduling, and production cluster operations.
EKS, IAM, ECR, VPC architecture, and cost-aware infrastructure decisions.
Jenkins, GitLab CI, GitHub Actions, deployment strategies, and automation.
Docker, image security, registries, and container runtime debugging.
Prometheus, Grafana, Loki, alerting, and observability for production.
VPCs, IAM, TLS, network policies, and access control patterns.