1. Introduction
The Kubernetes Metrics Server is a cluster add-on that collects resource utilisation metrics (CPU and memory) from each node's kubelet and exposes them via the Metrics API. Without it, kubectl top pods and kubectl top nodes don't work, and more critically, the Horizontal Pod Autoscaler can't make scaling decisions. If your HPA is not scaling, a broken Metrics Server is the most likely cause.
This guide covers installation verification, the most common certificate and connectivity errors, and how to configure the Metrics Server correctly for production clusters.
2. What "Metrics Server Not Working" Means
The Metrics Server works by: (1) running as a Deployment in kube-system, (2) connecting to each node's kubelet on port 10250 to scrape resource metrics, (3) aggregating those metrics and registering them with the Kubernetes aggregation layer, and (4) serving them via the metrics.k8s.io API. A failure at any of these steps produces different symptoms.
3. Common Causes
- Metrics Server is not installed in the cluster
- Certificate verification failure — Metrics Server can't verify kubelet's TLS certificate
- NetworkPolicy blocking Metrics Server from reaching kubelets on port 10250
- Metrics Server pod is crashing (OOMKilled or configuration error)
- The metrics API is not registered with the aggregation layer
- Kubelet serving certificate is self-signed and Metrics Server doesn't trust it
- Node firewall blocking port 10250 from the control plane network
- Resource requests too low — Metrics Server OOMKills in large clusters
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm the error and check if Metrics Server is installed
# Test kubectl top
kubectl top nodes
# If: "error: Metrics API not available" = not installed or API not registered
# If: "Error from server (ServiceUnavailable)" = installed but not healthy
# Check if Metrics Server is deployed
kubectl get deployment metrics-server -n kube-system
kubectl get pods -n kube-system -l k8s-app=metrics-server
# Check if the metrics API is registered
kubectl api-resources | grep metrics
# Should show: nodes and pods from metrics.k8s.io
Step 2: Install Metrics Server (if missing)
# Official installation
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
# For clusters with self-signed kubelet certificates:
# Download and patch the deployment to skip certificate verification
curl -L https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml -o metrics-server.yaml
# Add --kubelet-insecure-tls to the args in the Deployment
# (Only use this if you understand the security implications)
kubectl patch deployment metrics-server -n kube-system --type=json -p='[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'
Step 3: Fix certificate issues (production approach)
# For production, don't use --kubelet-insecure-tls
# Instead, configure kubelet to serve proper certificates
# On each node, check kubelet certificate config:
# /var/lib/kubelet/config.yaml
# serverTLSBootstrap: true <-- enables automatic cert rotation
# For EKS: kubelet serving certificates are managed automatically
# No --kubelet-insecure-tls needed
# Verify Metrics Server can reach kubelets:
kubectl logs -n kube-system deployment/metrics-server --tail=50
# Look for: "Failed to scrape node"
Step 4: Fix NetworkPolicy issues
# Check if NetworkPolicy blocks Metrics Server → kubelet (port 10250)
kubectl get networkpolicy -n kube-system
# Allow Metrics Server egress to kubelets:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-metrics-server
namespace: kube-system
spec:
podSelector:
matchLabels:
k8s-app: metrics-server
egress:
- ports:
- protocol: TCP
port: 10250
policyTypes:
- Egress
Step 5: Fix resource issues for large clusters
# Check if Metrics Server is OOMKilled
kubectl describe pod -n kube-system -l k8s-app=metrics-server | grep -A 5 "Last State"
# Increase Metrics Server resources for large clusters:
kubectl patch deployment metrics-server -n kube-system --type=json -p='[{"op":"replace","path":"/spec/template/spec/containers/0/resources",
"value":{"requests":{"cpu":"200m","memory":"200Mi"},
"limits":{"cpu":"1","memory":"512Mi"}}}]'
5. Verification Steps
# These should all return data after fixing
kubectl top nodes
kubectl top pods -A --sort-by=memory | head -10
# Confirm the metrics API is working
kubectl get --raw /apis/metrics.k8s.io/v1beta1/nodes | jq '.items[].metadata.name'
# HPA should now have access to metrics
kubectl describe hpa <hpa-name> -n <namespace>
# Look for: "AbleToScale: True" and recent metric values
6. Common Mistakes
- Using
--kubelet-insecure-tlsin production clusters — this disables TLS verification for kubelet connections - Not checking if Metrics Server is installed before debugging — "Metrics API not available" just means it's missing
- Forgetting that Metrics Server metrics have a 15-30 second lag —
kubectl topshowing low usage doesn't mean the pod isn't spiking - Setting Metrics Server memory too low for large clusters — it needs ~1MB per node; 100 nodes = 100MB minimum
7. Prevention Tips
- Include Metrics Server in your cluster bootstrap — don't discover it's missing when HPA stops scaling in production
- Monitor the Metrics Server pod health and alert on CrashLoopBackOff or OOMKilled
- For large clusters (100+ nodes), allocate 512Mi+ memory to Metrics Server and enable HA (2 replicas)
- If you need HPA to work reliably, see Fix Kubernetes HPA Not Scaling — the complete HPA diagnostic guide for the full HPA diagnostic guide
- Consider Prometheus + Prometheus Adapter as a more powerful alternative to Metrics Server for custom metrics-based autoscaling
8. FAQ
kubectl top shows some nodes but not others. Why?
The Metrics Server couldn't scrape certain kubelets. Check the Metrics Server logs for "failed to scrape node" errors with the specific node names. Causes: the node's kubelet isn't reachable from the Metrics Server pod, the node's certificate is failing verification, or there's a NetworkPolicy blocking the scrape.
Metrics Server is running but HPA says "unable to get metrics". Why?
The Metrics API is available but returning no data for the HPA's target resource. Usually this means: (1) the HPA references CPU/memory metrics but the pod hasn't run long enough to have metrics, (2) the targeted pods don't have resource requests set (required for percentage-based CPU metrics), or (3) the Metrics Server has a scrape delay and HPA is reading stale data. See Fix HPA Not Scaling for the full diagnosis.
9. Summary
| Error | Cause | Fix |
|---|---|---|
| Metrics API not available | Not installed | Install via official manifests |
| ServiceUnavailable | Pod crashing or not ready | Check pod logs; fix cert or network issues |
| Failed to scrape node | Can't reach kubelet on 10250 | Check NetworkPolicy; fix cert verification |
| OOMKilled | Not enough memory for cluster size | Increase Metrics Server memory limits |
| kubectl top no data | Metrics API registered but empty | Wait for first scrape; check kubelet health |
Explore More in This Category
Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.