1. Introduction

There are a lot of DevOps learning resources. Most of them give you a list of tools with logos and suggest you work through them in order. That approach tends to produce engineers who know which tools exist but can't diagnose production problems — because they never learned why the tools exist or how they fit together.

This roadmap takes a different approach. Each section focuses on a skill that you will use on the job, not just in a certification exam. The goal is to build a mental model of how real infrastructure works so that when something breaks — and it will — you know where to look.

2. What DevOps Actually Means on Real Teams

The term "DevOps" gets used to describe everything from a job title to a cultural philosophy to a set of automation tools. On real teams, it tends to mean: someone who can build and maintain the infrastructure that software runs on, the pipelines that deliver it, and the tooling that observes it. That usually requires knowing a bit of software, a bit of systems, and a lot about the interface between the two.

You don't need to master everything at once. Most DevOps engineers get good by going deep on one problem — a failing deployment, a misbehaving Kubernetes cluster, a CI pipeline that breaks every Friday — and following the chain of cause and effect until they understand it fully. The roadmap below gives you the sequence that makes that learning most efficient.

3. Core Foundation Skills

Linux

Almost all production infrastructure runs on Linux. You don't need to memorise every command, but you need to be comfortable navigating a server, reading logs, inspecting processes, managing files and permissions, and editing configuration files without a GUI. The specific skills that matter most on the job: systemctl, journalctl, ps, df, du, ss, netstat, grep, awk, sed, tail -f, file permissions and ownership, and understanding how processes start and stop. Learn these before anything else.

Networking

Most production incidents trace back to a network problem: a firewall rule, a DNS failure, a routing issue, or a certificate expiry. You need to understand: how DNS works (A records, CNAMEs, TTL, resolution chains), TCP/IP at a working level (IP addresses, ports, protocols, packets), HTTP and HTTPS (request/response, headers, TLS handshake), and how traffic flows from a browser through a load balancer to a container. You don't need to configure routers, but you do need to diagnose connectivity problems. The Networking & Security category has practical guides for when things go wrong.

Git

Git is how code and configuration move between people and environments. You need to be comfortable with branching strategies, rebasing, merge conflicts, tagging, and using Git in automated pipelines — not just the basics of commit and push. In DevOps specifically, you'll be writing Dockerfiles, Kubernetes manifests, Terraform configs, and CI pipeline definitions, all of which live in Git and need to be reviewed and versioned like code.

Scripting

Bash scripting is the most immediately useful skill after Linux basics. You'll use it to write deployment scripts, automation glue, health checks, and CI pipeline steps. Python is worth learning after Bash — it's better for anything that involves parsing structured data, calling APIs, or anything beyond a few dozen lines. The rule of thumb: use Bash for simple system automation, Python for anything that needs to be maintained or extended.

4. CI/CD and Automation

Continuous integration and continuous deployment pipelines are how code gets from a developer's laptop to production reliably. A typical pipeline runs automated tests, builds a container image, pushes it to a registry, and deploys it to a cluster — all triggered by a Git push. Understanding how pipelines work is essential because you'll spend a significant amount of your time keeping them healthy.

Start with one tool and learn it properly: GitLab CI, GitHub Actions, or Jenkins. Understand stages, jobs, artifacts, caching, environment variables, secrets management, and failure modes. The CI/CD category covers the failure patterns you'll encounter most often in production.

5. Containers and Kubernetes

Containers (Docker) package software and its dependencies together so it runs consistently anywhere. Kubernetes orchestrates those containers at scale — handling scheduling, restarts, scaling, service discovery, and deployment strategies. These two technologies dominate production infrastructure for most teams today.

Learn containers before Kubernetes: understand how images are built, how layers work, how networking and volumes work inside containers. Then move to Kubernetes once you understand what it's orchestrating. The Kubernetes category has the operational depth you'll need.

6. Cloud Infrastructure

AWS is the dominant cloud platform for DevOps work, though the concepts transfer to GCP and Azure. The services you'll actually use in production are narrower than the AWS catalog suggests: IAM (permissions), VPC (networking), EC2 or EKS (compute), ALB/NLB (load balancers), RDS (databases), S3 (object storage), ECR (container registry), and CloudWatch (logs and metrics). Understanding how these fit together is more important than knowing every feature of each service.

The AWS & Cloud category covers the IAM, EKS, and networking issues that surface most often. The AWS Essentials guide walks through these services from a practical DevOps perspective.

7. Monitoring and Troubleshooting

Monitoring isn't just dashboards — it's the practice of making systems observable so you can understand what's happening and diagnose problems when they occur. The three pillars are: logs (what happened), metrics (how much/how fast/how many), and traces (which path did a request take). In practice, most teams start with logs and metrics.

Learn Prometheus and Grafana for metrics, and at least one log aggregation tool (Loki, ELK, or CloudWatch). More importantly, learn how to think through an incident: start with symptoms, form a hypothesis, test it, narrow the cause, fix it, verify the fix. That diagnostic skill is what separates engineers who can work on-call from those who can't. The Monitoring & Logging category covers practical observability for Kubernetes environments.

8. Security and Reliability Basics

Security in DevOps (sometimes called DevSecOps) isn't a separate phase — it's woven into everything from how you manage secrets to how you configure IAM roles to how you scan container images. The baseline skills: never hardcode credentials (use environment variables, secrets managers, or IRSA), understand principle of least privilege in IAM, scan images for known CVEs, use TLS everywhere, and rotate credentials regularly.

Reliability engineering focuses on keeping systems available: set resource limits so one pod can't starve another, use readiness and liveness probes correctly, configure PodDisruptionBudgets, and make sure your monitoring alerts before users notice a problem. These aren't advanced topics — they're baseline practices that every team should be doing from the start.

9. Suggested Learning Path

If you're starting from scratch or consolidating foundations, this is the order that builds skills on each other most efficiently:

  1. Linux command line: A week of deliberate practice navigating, reading logs, and managing processes. Learn by doing, not by reading.
  2. Networking and DNS: Understand how a request travels end to end. Set up a small project and trace what happens at each hop.
  3. Git and scripting: Write Bash scripts to automate tasks you already do manually. Push everything to Git.
  4. Docker: Build images, run containers, understand layers, volumes, and networking. Don't skip this before Kubernetes.
  5. CI/CD: Set up a pipeline for a real project. Make it build, test, and deploy. Break it and fix it.
  6. Kubernetes: Deploy the same project to a local cluster (kind or minikube). Learn what Kubernetes is solving by operating it.
  7. AWS: Run the project on a real cloud environment. Add a load balancer, IAM roles, and CloudWatch logs.
  8. Monitoring: Add Prometheus and Grafana. Create an alert that fires before a real problem gets bad.

Explore More

Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.