1. Introduction

A runbook is a documented set of procedures for handling a specific operational task or incident. In theory, every team has them. In practice, most teams have a folder full of documents that nobody reads during an incident because they're too long, too vague, too out of date, or too hard to navigate under pressure.

The goal of a runbook is not completeness. It's usability at 2am by someone who may not be the most experienced person on the team. That changes how you write it — shorter, more direct, more command-heavy, less background explanation.

This guide covers what makes a runbook actually useful, what sections to include, what to leave out, a full template with annotated examples, and a process for keeping runbooks alive over time rather than letting them go stale.

2. When and Why to Write a Runbook

Write a runbook for any operational task that:

Common runbook categories in DevOps teams:

CategoryExamples
Incident responseService down, high error rate, database failover, DDoS response
Deployment proceduresProduction release process, hotfix deployment, rollback procedure
Infrastructure tasksScaling a node group, rotating credentials, certificate renewal
Routine maintenanceDatabase vacuuming, log rotation, index cleanup, cache flush
On-call proceduresAlert triage, escalation paths, stakeholder communication templates
Recovery proceduresRestoring from backup, DR failover, data recovery steps

3. Prerequisites

Before writing a runbook, you need:

4. The Runbook Structure

A practical runbook has eight sections. Each one serves a specific purpose for the person executing it under pressure. Here is the complete structure, followed by a detailed explanation of each section:

# Runbook: [Name of Issue or Task]

## Overview
One paragraph: what this runbook covers and when to use it.
## Severity
& Impact Who is affected, how many users, what SLO/SLA is at risk.
## Prerequisites
Access, tools, and permissions required before starting.
## Symptoms
What the engineer will see: alert name, error messages, metrics.
## Diagnosis Steps
Ordered commands and checks to identify the root cause.
## Resolution Steps
Ordered commands and actions to resolve the issue.
## Verification
How to confirm the issue is resolved. Specific checks.
## Escalation
Who to contact if this runbook doesn't resolve the issue.

Section 1: Overview

Two to four sentences. What is this runbook for? What system or service does it cover? What is the condition that triggers its use? Keep it to the minimum context a new on-call engineer needs to confirm they are looking at the right runbook.

## Overview This runbook covers the response procedure for the payment-api service reporting a high error rate (>5% 5xx responses over 5 minutes). It applies to the production environment only. For staging issues, see: runbook-payment-api-staging-errors.md Related alert: payment-api-high-error-rate (PagerDuty policy: payments-team)

Section 2: Severity and Impact

State the business and user impact clearly. Who is affected? What functionality is degraded or unavailable? What SLO or SLA is at risk and over what time horizon? This section helps the on-call engineer decide how urgently to escalate and whether to wake up additional people immediately.

## Severity & Impact Severity: SEV-2 (escalate to SEV-1 if error rate exceeds 20% for > 10 minutes) User impact:  Payment processing is degraded or unavailable.               Checkout flow fails for affected users. SLO at risk:  payment-api availability SLO (99.9% / 30-day window).               Each minute of >5% error rate consumes ~0.07% of monthly error budget. Stakeholders: Notify #incidents Slack channel immediately.               Page payments-team-lead if not resolved within 15 minutes.

Section 3: Prerequisites

List everything the engineer needs before they can start. Be specific about access levels, tools, and credentials. An engineer who discovers mid-incident that they don't have the right kubectl context or Vault access has wasted critical minutes. Better to surface that before the steps begin.

## Prerequisites - kubectl access to the production cluster (context: eks-prod-us-east-1) - AWS CLI configured with the ops-readonly role (for CloudWatch/RDS access) - Access to #incidents Slack channel for communication - PagerDuty responder access (to acknowledge and reassign alerts)
      # Verify your cluster context before starting: kubectl config current-context
      # Expected: eks-prod-us-east-1

Section 4: Symptoms

Describe exactly what the engineer will observe. If your team runs Loki or ELK, include the exact log query the engineer should run — not just "check the logs". Include the alert name, the specific metric thresholds that triggered it, example error messages they will see in logs, and any dashboard panels that are relevant. This helps the engineer confirm they are responding to the right incident and not something unrelated.

## Symptoms Alert:   payment-api-high-error-rate fires when:          rate(http_requests_total{status=~'5..', job='payment-api'}[5m]) > 0.05 Grafana: Dashboard 'Payment API Overview' — panel 'Error Rate %' showing spike          Dashboard 'Payment API Overview' — panel 'P99 Latency' may also be elevated Logs:    Loki query to confirm:          {namespace="production", container="payment-api"} | = "error" | rate [5m]          Common error signatures:          "connection refused" — downstream DB or dependency issue          "context deadline exceeded" — timeout, check latency metrics          "OOMKilled" — memory issue, check pod restarts

Section 5: Diagnosis Steps

This is the most important section. Write ordered steps with exact commands. Do not write prose — write commands. Every step should produce observable output that confirms whether the problem is in that layer or not. The engineer should be able to follow these steps without knowing the system in depth.

## Diagnosis Steps 1. Check pod status and recent restarts    kubectl get pods -n production -l app=payment-api    kubectl describe pod <pod-name> -n production | tail -30 2. Check pod logs for the error    kubectl logs -n production -l app=payment-api --tail=100 --previous    # Look for: error message, stack trace, timestamp of first failure 3. Check if the issue is a single pod or all pods    kubectl get pods -n production -l app=payment-api    # If one pod is bad: likely a code or config issue on that pod    # If all pods are bad: likely a downstream dependency or config change 4. Check downstream dependencies    # Test DB connectivity from inside a pod:    kubectl exec -it <pod-name> -n production -- \      nc -zv payment-db.internal 5432    # Test Redis:    kubectl exec -it <pod-name> -n production -- \      redis-cli -h redis.internal ping 5. Check for recent deployments    kubectl rollout history deployment/payment-api -n production    # If a recent deployment is visible, proceed to Resolution Step 1 (rollback)

Section 6: Resolution Steps

Write the resolution steps in the same way as diagnosis — ordered, command-first, with expected output. Where there are multiple possible resolutions depending on the diagnosis, number them and reference which diagnosis step leads to each. A branching structure is fine as long as the branch conditions are explicit.

## Resolution Steps --- If caused by a bad deployment (from Diagnosis Step 5) --- 1. Rollback to the previous deployment    kubectl rollout undo deployment/payment-api -n production    kubectl rollout status deployment/payment-api -n production    # Wait for rollout to complete — watch error rate in Grafana --- If caused by DB connectivity failure (from Diagnosis Step 4) --- 2. Check RDS instance status    aws rds describe-db-instances \      --db-instance-identifier payment-db-prod \      --query 'DBInstances[0].DBInstanceStatus'    # If 'available': check security group rules and VPC routing    # If 'rebooting' or 'failing-over': wait and monitor — RDS will recover 3. If connection pool exhaustion is suspected    kubectl rollout restart deployment/payment-api -n production    # Restarting pods forces new DB connections — short-term mitigation    # File a follow-up ticket to investigate pool sizing --- If caused by memory (OOMKilled from Diagnosis Step 2) --- 4. Temporarily increase memory limit and redeploy    kubectl set resources deployment/payment-api -n production \      --limits=memory=1Gi    # IMPORTANT: file a ticket to update the manifest in git    # kubectl set resources does not persist — it will be overwritten on next deploy

Section 7: Verification

Tell the engineer exactly how to confirm the issue is resolved. Name the specific metric, dashboard panel, or command output that represents a healthy state. 'Monitor it for a few minutes' is not a verification step — it leaves the engineer guessing about when to stand down.

## Verification 1. Error rate has returned below threshold    Grafana: Dashboard 'Payment API Overview' > panel 'Error Rate %'    Should show < 1% for 5 consecutive minutes 2. Pods are Running and not restarting    kubectl get pods -n production -l app=payment-api    Expected: all pods READY 1/1, RESTARTS count not increasing 3. Confirm with a synthetic transaction (if available)    curl -X POST https://api.example.com/v1/payments/health-check    Expected: HTTP 200, {"status": "ok"} 4. Confirm error budget impact    Grafana: Dashboard 'SLO Overview' > 'payment-api availability'    Note the error budget remaining and record in the incident ticket

Section 8: Escalation

Define clearly who to contact when this runbook does not resolve the issue. Include the contact method (Slack, PagerDuty, phone), not just a name. People change roles and leave teams — use a role or on-call rotation name rather than an individual wherever possible.

## Escalation If this runbook does not resolve the issue within 30 minutes: 1. Payments team lead    PagerDuty: escalate incident to 'payments-team-lead' policy    Slack: @payments-team-lead in #incidents 2. Database team (if DB issue is suspected)    PagerDuty: page 'database-oncall' policy    Slack: @database-team in #incidents 3. AWS Support (if RDS infrastructure issue confirmed)    Support plan: Business (4hr SLA for production issues)    Case URL: https://console.aws.amazon.com/support Incident commander: first responder owns the incident until handoff. Communication: post updates to #incidents every 15 minutes.

5. Complete Worked Example

Here is a condensed but complete runbook for a real-world scenario, showing how all eight sections work together. This is what a finished runbook looks like — not a template, but an actual document ready for use.

# Runbook: Kubernetes CrashLoopBackOff — payment-api Last updated: 2025-04-15 | Owner: payments-team | Severity: SEV-2
      ## Overview Use this runbook when the payment-api container is in CrashLoopBackOff in the production namespace. The container starts and exits repeatedly. This runbook does not apply to ImagePullBackOff — see: runbook-image-pull.md
      ## Severity & Impact SEV-2. Payment processing unavailable for affected pods. Escalate to SEV-1 if all pods are crashing simultaneously.
      ## Prerequisites - kubectl access: context eks-prod-us-east-1 - Read access to AWS Secrets Manager (for secret validation steps)
      ## Symptoms Alert: payment-api-pod-crash-loop (threshold: restarts > 3 in 10 min) Visible in: kubectl get pods -n production | grep payment-api
      ## Diagnosis Steps 1. kubectl logs <pod> -n production --previous    → Read last 50 lines. Note the error and exit code. 2. kubectl describe pod <pod> -n production | grep -A5 'Last State'    → Note exit code. Exit 137 = OOMKilled. Exit 1 = app error. 3. kubectl get secret payment-api-config -n production    → Confirm secret exists. Missing secret = common crash cause.
      ## Resolution Steps Exit code 1 / app error in logs:   → Check if caused by recent deploy: kubectl rollout history ...   → If yes: kubectl rollout undo deployment/payment-api -n production Exit code 137 (OOMKilled):   → kubectl set resources deployment/payment-api -n production \        --limits=memory=1Gi   → File ticket: update memory limit in git and merge to main Missing secret:   → Confirm in AWS Secrets Manager. Re-sync with External Secrets Operator.   → kubectl rollout restart deployment/payment-api -n production
      ## Verification kubectl get pods -n production -l app=payment-api → All pods: READY 1/1, RESTARTS not increasing for 5 minutes Grafana: 'Payment API Overview' > 'Error Rate %' < 1%
      ## Escalation Not resolved in 20 min → page payments-team-lead via PagerDuty

6. What to Leave Out

Runbooks fail not because they are missing information but because they contain too much of the wrong kind. These things actively make runbooks worse:

7. Keeping Runbooks Current

A runbook that is six months out of date is worse than no runbook. An engineer who follows stale steps and takes a wrong action because the commands or service topology changed has been actively harmed by the documentation.

Embed runbook review into your team's existing processes

Metadata that runbooks must always have

FieldWhy it matters
Last updated dateTells the reader whether to trust the content or treat it with caution
Owner / teamWho to ask if the runbook is wrong or incomplete
Reviewed byConfirms another engineer has validated the steps
Related runbooksLinks to connected procedures — avoids dead ends mid-incident
Related alertsMaps the runbook to the alert that triggers it for fast lookup
Change logBrief history of what changed and why — critical for post-incident context

8. Summary

A runbook that gets used during incidents shares the same characteristics: it is short, specific, command-first, and written for someone who is stressed and working fast. The structure is consistent across every runbook so engineers know where to find what they need without reading linearly.

SectionPurposeKeep it to
OverviewConfirm this is the right runbook2–4 sentences
Severity & ImpactCalibrate urgency and escalation timing5–8 lines
PrerequisitesPrevent mid-incident access failuresBullet list
SymptomsConfirm the incident matches the runbookAlert name + key signals
Diagnosis StepsIdentify the root cause systematicallyNumbered, commands only
Resolution StepsFix the issue with exact actionsNumbered, branched by cause
VerificationConfirm resolution with specific checks3–5 concrete checks
EscalationDefine next steps if the runbook doesn't workRoles + contact method

Start with your most-paged alert. Write its runbook the day after the next incident that triggers it. Repeat for the next most-paged alert. Within a quarter, you will have covered the scenarios that matter most — and your on-call rotation will be noticeably less stressful.