1. Introduction
A failing CI/CD pipeline stops your team from shipping. The error message is often vague — "exit code 1", "job failed", "build failed" — with the real cause buried several log lines above. The challenge isn't fixing the specific error once you find it; it's knowing where to look and what questions to ask.
This guide provides a systematic debugging framework that applies to any CI/CD platform: GitLab CI, GitHub Actions, Jenkins, CircleCI, or others. For platform-specific issues like GitLab runner timeouts or stuck Jenkins pipelines, see the dedicated guides.
2. What "Pipeline Failed" Actually Means
A pipeline fails when any job exits with a non-zero status code. This can happen in the build, test, lint, security scan, deploy, or post-deploy stages. The pipeline stops at the first failing job (unless configured otherwise), which means the visible failure may be a symptom — triggered by a missing artifact, environment variable, or secret from an earlier step that appeared to succeed.
3. Common Causes
- Missing or incorrect environment variable or secret — the most common cause
- Docker image pull failure — registry authentication expired or image tag doesn't exist
- Build tool cache corruption — stale cache causes dependency resolution errors
- Test failure — actual code bug or flaky test that needs isolation
- Insufficient runner resources — out of disk space, memory, or CPU
- Network timeout — downloading dependencies from an external registry
- Permissions issue — runner doesn't have access to a required secret, file, or API
- Script with poor error handling — a command fails silently and a later command produces a confusing error
- Race condition in parallel jobs — shared state, artifact not ready when needed
4. Step-by-Step Debugging Process
Step 1: Read the full log output of the failing job
# In most CI platforms: click on the failing job to expand it
# Scroll up from the red error banner to find the actual failing command
# Look for patterns like:
# - "Error: cannot find module '...'"
# - "exit status 1"
# - "permission denied"
# - "connection refused" or "timeout"
# - "No such file or directory"
# - "authentication required"
Step 2: Reproduce the failure locally
# The fastest way to debug a CI failure is to run the same command locally
# Use the exact same Docker image if possible:
# For GitLab CI
docker run --rm -it -e MY_SECRET="$MY_SECRET" -v $(pwd):/workspace -w /workspace registry.gitlab.com/my-org/my-image:tag bash -c "your-failing-command"
# For GitHub Actions (using act tool)
act push --job build --secret-file .env
# For Jenkins (using Docker agent)
docker run --rm -it jenkins/agent:latest bash
Step 3: Check environment variables and secrets
Missing or incorrectly scoped environment variables cause a large fraction of CI/CD failures. See the dedicated Fix Environment Variables Not Working in CI/CD guide for the full diagnostic. Quick check:
# Add a debug step before the failing step to print environment
# (Mask sensitive values — only print key names or non-sensitive values)
# GitLab CI:
debug-env:
stage: debug
script:
- env | grep -v 'SECRET\|TOKEN\|PASSWORD\|KEY' | sort
- echo "CI_REGISTRY=$CI_REGISTRY"
- echo "DEPLOY_ENV=$DEPLOY_ENV"
when: manual
# GitHub Actions:
- name: Debug environment
run: env | grep -v 'SECRET\|TOKEN\|PASSWORD' | sort
Step 4: Check for Docker build and push failures
If the pipeline fails at a Docker step, see Fix Docker Build Failed in CI/CD for specific diagnostics. Common quick checks:
# Check Docker login is working
docker login registry.example.com -u $CI_REGISTRY_USER -p $CI_REGISTRY_PASSWORD
# Verify the target image tag exists before pulling
docker manifest inspect registry.example.com/my-image:1.2.3
# Check disk space (Docker builds fill disk fast)
df -h
docker system df
Step 5: Check runner health and resources
# For GitLab: check runner status in the CI/CD settings
# For self-hosted runners, SSH in and check:
df -h # disk space
free -m # memory
docker ps -a # any zombie containers taking up resources
# Clean up Docker resources if disk is full
docker system prune -af
docker volume prune -f
Step 6: Isolate flaky tests
# Re-run the failing job to check if it's consistently failing
# (in most platforms you can retry a specific job)
# For test failures: run the specific test file locally
pytest tests/test_specific.py -v
jest src/components/MyComponent.test.js --verbose
# Enable more verbose output in your test framework
pytest --tb=long -v
jest --verbose --runInBand # --runInBand disables parallel test execution
Step 7: Check for race conditions in parallel jobs
# If jobs run in parallel and share artifacts, ensure dependencies are explicit
# GitLab CI: use 'needs' to declare dependencies
deploy:
needs:
- job: build
artifacts: true
script:
- deploy.sh
# GitHub Actions: use 'needs' at the job level
jobs:
deploy:
needs: [build, test]
steps: [...]
5. Verification Steps
# After fixing, push a new commit and watch the pipeline
# Focus on the previously failing job
# Verify the fix didn't break adjacent jobs:
# - Check that artifact handoff between jobs works
# - Confirm environment variables are available in downstream jobs
# - Run a full pipeline, not just the individual failing job
6. Common Mistakes
- Not reading the full log — the UI shows a summary; the cause is always in the detailed output
- Retrying the pipeline without understanding the cause — flaky tests should be fixed, not retried indefinitely
- Using
|| trueto suppress errors — this hides real failures from downstream jobs - Hard-coding credentials in pipeline configuration instead of using secrets — a security risk and ops burden
- Not adding
set -eto shell scripts — without it, failures in multi-line scripts are swallowed
7. Prevention Tips
- Always use
set -euo pipefailat the top of shell scripts in CI — makes every failure immediately visible - Pin all dependency versions in package managers — unpinned versions cause random failures when a new version breaks your build
- Validate required environment variables exist at the start of every job before running any actual work
- Monitor pipeline success rate and mean time to fix — trends in these metrics indicate systemic problems
- Keep CI images updated but use explicit version tags, not
latest - Document recurring failures in a runbook — saves time when the same issue appears at 2am
8. FAQ
The pipeline fails in CI but works locally. Why?
Environment differences are almost always the cause: different OS, different tool versions, different environment variables, missing files that exist locally but not in CI, or a service dependency (database, cache) that's available locally but not in the CI runner. Replicate the CI environment exactly using Docker to debug this.
A job fails intermittently. How do I fix a flaky test?
First, confirm it's flaky by running it multiple times. Common causes: test-order dependency (one test pollutes state for another), timing-based assertions (use retries or wait conditions instead of fixed sleeps), and external service calls that sometimes fail. Use your test runner's retry plugin as a short-term fix while you fix the root cause.
My pipeline is very slow. How do I speed it up?
Profile each job's duration first. Common improvements: cache dependencies between runs (node_modules, .gradle, pip cache), use Docker layer caching, run tests in parallel, split large jobs into smaller parallel ones, and use a faster runner (more CPU/RAM). See Fix GitLab Runner Timeout for caching specifics.
9. Summary
| Symptom | Likely Cause | Fix |
|---|---|---|
| Generic exit code 1 | Command failed — read full log | Scroll up past error banner to see actual command output |
| Works locally, fails in CI | Environment difference | Reproduce with same Docker image; check env vars |
| Intermittent failures | Flaky test or race condition | Run in isolation; use retry plugins; fix root cause |
| Disk full errors | Runner out of space | docker system prune; increase runner disk; cache smarter |
| Missing env var errors | Secret not passed to job | See Fix Env Vars guide; validate at job start |
Explore More in This Category
Explore more in this category: CI/CD guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.