1. Introduction

Choosing a logging stack is one of those infrastructure decisions that feels low-stakes until you're six months in, your Elasticsearch cluster is consuming 40% of your compute budget, and your on-call engineer can't find the log line they need because the query times out.

ELK — Elasticsearch, Logstash (or Beats), and Kibana — has been the dominant open-source logging stack for over a decade. It's powerful, flexible, and battle-tested at scale. Grafana Loki is the newer challenger, built with a different philosophy: instead of indexing everything, index only the labels, store the log lines compressed and cheaply, and query them with a language that will feel familiar to anyone who uses Prometheus.

This guide compares the two stacks honestly — architecture, resource cost, query capability, Kubernetes fit, and operational overhead — and gives you a clear framework for choosing between them based on what your team actually needs.

2. Quick Summary

A side-by-side view of the key attributes before going deeper on each:

ELK StackGrafana Loki
Core componentsElasticsearch, Logstash/Beats, KibanaLoki, Promtail/Alloy, Grafana
Storage modelIndexes full log content — text and metadataIndexes only labels; stores log lines compressed
Query languageLucene / KQL — full-text search across all fieldsLogQL — label filters + line filters + metric queries
Full-text searchYes — index-based, very fast on any fieldLimited — line filter regex only, no field indexing
Resource usageHigh — Elasticsearch is memory and storage intensiveLow — label-only index dramatically reduces overhead
Kubernetes fitGood — but requires careful resource tuningExcellent — Promtail reads pod logs natively via labels
Grafana integrationVia Elasticsearch datasource pluginNative — Loki is a first-class Grafana datasource
Metrics correlationSeparate stack needed (e.g. Prometheus + Grafana)Native — correlate logs and metrics in same Grafana UI
Alert supportElasticsearch alerting (X-Pack / free tier)Grafana alerting on LogQL queries
Scaling complexityHigh — Elasticsearch cluster sizing is non-trivialMedium — stateless querier, object storage for scale
Managed cloud optionElastic Cloud, AWS OpenSearchGrafana Cloud (hosted Loki)
LicenceElastic License 2.0 (not fully open source)AGPL-3.0 (open source)
Operational costHigh — JVM tuning, shard management, snapshotsLow to medium — fewer moving parts

3. How Each Stack Works

ELK Stack architecture

Logs flow from your applications or nodes into Beats agents (Filebeat, Metricbeat) or Logstash for heavier processing, then into Elasticsearch where they are fully parsed and indexed. Kibana provides the query UI and dashboards on top.

Elasticsearch indexes every field in every log line. A log entry like {"level": "error", "message": "connection refused", "service": "payment-api", "latency_ms": 342} gets indexed across all four fields. This means you can search by any field, with full-text search, instantly — but it also means Elasticsearch must maintain a large inverted index and store both the original log and the index data.

# Typical ELK data flow Application logs   → Filebeat (lightweight shipper on each node)   → Logstash (optional: parse, filter, enrich)   → Elasticsearch (index + store)   → Kibana (search + dashboards)
      # Index storage overhead is significant:
      # Raw logs:       100 GB
      # Elasticsearch:  ~200–300 GB (index + replica + source)
      # Example KQL query in Kibana:
      # level: error AND service: payment-api AND latency_ms > 300

Loki architecture

Loki deliberately avoids full-text indexing. When Promtail (or Grafana Alloy) ships logs to Loki, only the labels attached to the log stream are indexed — things like namespace, pod name, container name, and environment. The log lines themselves are compressed and stored in chunks in object storage (S3, GCS, or local filesystem).

Queries in Loki use LogQL. You first filter by labels (which is fast, because those are indexed), then optionally apply a regex or string filter against the log lines in the matched chunks. The trade-off is explicit: fast label-based filtering, slower arbitrary full-text search.

# Typical Loki data flow Application logs (stdout/stderr on Kubernetes pods)   → Promtail / Grafana Alloy (reads /var/log/pods/*)   → Loki (index labels, store chunks in object storage)   → Grafana (query + dashboards, same UI as Prometheus)
      # Storage overhead is much lower:
      # Raw logs:  100 GB
      # Loki:      ~40–60 GB (compressed chunks + tiny label index)
      # Example LogQL queries:
      # Filter by label (fast — uses index):
      # {namespace="production", container="payment-api"} | = "error"
      # Count errors per minute (metric query):
      # rate({namespace="production"} | = "error" [5m])
      # Parse and filter a JSON log field:
      # {namespace="production"} | json | latency_ms > 300

4. Key Differences

Query power and full-text search

This is the most significant functional difference and the one that matters most for how your team will actually use the system day to day.

Elasticsearch indexes every field, so any KQL query is fast regardless of what you search on. If your logs are structured JSON and you need to search by a specific request ID, a customer ID, a transaction hash, or any field across months of data — Elasticsearch handles this well. This is its core strength.

Loki does not index log content. A search for a specific request ID in Loki requires Loki to scan and decompress the log chunks that match your label filter, then apply the string or regex filter against the raw text. For recent logs in a small time window this is fast enough. For historical searches across weeks of data without tight label filters, it can be slow or require expensive compute.

Resource consumption and cost

Elasticsearch runs on the JVM. It needs substantial heap memory — a minimum of 4–8 GB per node for anything non-trivial, and production clusters commonly run 16–64 GB nodes. Storage costs are amplified by the index overhead and replicas. For a team ingesting 50 GB/day of logs, a properly sized Elasticsearch cluster typically needs 3–5 nodes with 16+ GB RAM each.

Loki's querier and distributor components are stateless and lightweight. The index is tiny. The log chunks go to object storage — S3, GCS, or MinIO — which is orders of magnitude cheaper than block storage. The same infrastructure decisions apply when choosing how your logging VPCs connect — see VPC Peering vs Transit Gateway for network topology trade-offs. The same 50 GB/day log volume on Loki typically runs on 2–3 small nodes and pays object storage rates for persistence rather than provisioned disk on large instances.

# Rough resource comparison for 50 GB/day log ingest
      # ELK Stack (self-hosted, 7-day retention, 1 replica):
      # - 3x Elasticsearch nodes: 16 GB RAM, 500 GB SSD each
      # - 1x Logstash: 4 GB RAM
      # - 1x Kibana: 2 GB RAM
      # - Total storage: ~2.5 TB (index + replica overhead)
      # - Approx monthly cost (AWS): $800–1200
      # Loki Stack (self-hosted, 7-day retention):
      # - 1x Loki (or 2 for HA): 4 GB RAM
      # - 1x Grafana: 1 GB RAM (likely already running for metrics)
      # - Storage: ~150–200 GB in S3
      # - Approx monthly cost (AWS): $80–150

Kubernetes and container-native logging

Loki was designed with Kubernetes in mind. If your Loki or Promtail pods get stuck during deployment, the Pod Pending troubleshooting guide covers the most common scheduling failures. Promtail automatically discovers pods via the Kubernetes API, reads their stdout/stderr log files from /var/log/pods/, and attaches Kubernetes metadata as labels — namespace, pod name, container name, node name — without any log parsing configuration. For a Kubernetes cluster where most logs go to stdout, the default Promtail configuration works out of the box.

Filebeat also supports Kubernetes autodiscovery and works well, but the Elasticsearch pipeline requires more upfront configuration to parse and route Kubernetes log metadata correctly. For teams already running Grafana and Prometheus, adding Loki is often a 30-minute task. Adding ELK to a Kubernetes cluster that doesn't already have it is a multi-day project.

# Deploy Loki + Promtail on Kubernetes with Helm helm repo add grafana https://grafana.github.io/helm-charts helm repo update
      # Install Loki (simple scalable mode — S3 backend) helm install loki grafana/loki \   --namespace monitoring \   --set loki.storage.type=s3 \   --set loki.storage.s3.region=us-east-1 \   --set loki.storage.s3.bucketnames=my-loki-logs
      # Install Promtail (ships pod logs to Loki automatically) helm install promtail grafana/promtail \   --namespace monitoring \   --set config.clients[0].url=http://loki:3100/loki/api/v1/push
      # Promtail auto-discovers all pods and attaches these labels:
      # namespace, pod, container, node_name, app (from pod labels)

Grafana integration and observability correlation

This is a significant operational advantage for Loki in teams already running Prometheus and Grafana. Loki is a native Grafana datasource. Engineers can switch between their Prometheus metrics dashboard and the correlated log view in the same Grafana panel, using the same time range selector. Grafana's Explore view lets you run LogQL and PromQL side by side.

Elasticsearch integrates with Grafana via a datasource plugin, but Kibana is the native UI. Teams using ELK typically end up with two separate observability UIs — Grafana for metrics, Kibana for logs — with no built-in correlation between them. This isn't fatal, but it adds friction during incidents.

Operational overhead

Running Elasticsearch in production requires ongoing attention. JVM heap tuning, shard allocation, index lifecycle management (ILM) policies, snapshot management for backups, and cluster health monitoring are all operational responsibilities that have no equivalent in Loki. An Elasticsearch cluster that is undersized will degrade under load. Shard imbalance is a common production issue that requires manual intervention.

Loki's operational surface is smaller. The stateless components are easy to scale horizontally. The main operational concern is label cardinality — a poorly instrumented application pushing high-cardinality labels will cause index growth and query degradation. With proper label discipline, Loki is notably lower maintenance than Elasticsearch.

5. Pros and Cons

ELK Stack

ProsCons
Full-text search across every field — fast, regardless of cardinalityHigh resource consumption — Elasticsearch needs substantial RAM
Mature, battle-tested at enormous scaleComplex operational model — JVM tuning, shards, ILM
Rich query language (KQL/Lucene) with aggregationsStorage overhead from full-text indexing (2–3× raw log size)
Kibana provides powerful out-of-the-box dashboards and alertingExpensive to self-host at scale
Strong ecosystem — APM, security analytics, ML featuresElastic licence change (2021) limits some use cases
Well-supported by managed services (Elastic Cloud, AWS OpenSearch)Kibana is separate from Grafana — fragmented observability UI
Easier to search unstructured or inconsistently structured logsSteep learning curve for cluster management

Grafana Loki

ProsCons
Dramatically lower resource cost — label-only indexNo full-text field indexing — arbitrary search can be slow
Native Grafana integration — unified metrics and logs UIHigh-cardinality labels will degrade performance significantly
Excellent Kubernetes fit — Promtail autodiscoveryLogQL has a learning curve for complex queries
Object storage backend (S3/GCS) — cheap long-term retentionLess mature than Elasticsearch for deep analytics use cases
LogQL metric queries — generate metrics from log dataDebugging without tight label filters requires patience
Simple horizontal scaling with stateless componentsNo native APM or security analytics features
Fully open source (AGPL-3.0)Log parsing happens at query time — adds query latency

6. When to Use ELK

Choose ELK when...

  • Your team regularly searches logs by arbitrary fields — customer IDs, transaction IDs, request hashes, or any high-cardinality value
  • You need rich log analytics: aggregations, histograms, trend analysis across long time ranges
  • You have compliance or security requirements that benefit from Elasticsearch's ML and SIEM features
  • Your logs are inconsistently structured or unstructured text where full-text search is the only practical query method
  • You are already running Elastic Cloud or AWS OpenSearch and want a managed service with SLA guarantees
  • Your team has existing Elasticsearch expertise and the operational investment is already made
  • You need APM (Application Performance Monitoring) integrated with your log data in a single platform

A concrete ELK scenario: a SaaS product team supporting enterprise customers, where the support workflow involves searching logs by customer ID or session token across multiple microservices going back weeks. This is precisely what Elasticsearch is built for — indexed, sub-second lookups by any field, regardless of time range.

7. When to Use Loki

Choose Loki when...

  • Your primary logging use case is Kubernetes pod and container logs searchable by namespace, service, or environment
  • You are already running Prometheus and Grafana for metrics — Loki gives you logs in the same UI at minimal added cost
  • Resource cost and storage efficiency are meaningful constraints — self-hosting ELK at scale is expensive
  • Your team debugs by filtering to a specific service and time window, then reading log lines — not arbitrary field searches
  • You want a logging system that scales with Kubernetes naturally without dedicated cluster management expertise
  • You need long retention periods where object storage (S3) costs make ELK storage costs prohibitive
  • You are building a new stack from scratch and want the lowest operational overhead path to functional logging

A concrete Loki scenario: a platform engineering team running 20 microservices on Kubernetes, already using Prometheus and Grafana. Their debugging workflow is 'show me logs from service X in namespace Y for the last 30 minutes'. Loki handles this perfectly, integrates with their existing Grafana instance, and costs a fraction of what ELK would on the same cluster.

8. Final Recommendation

The honest answer is that Loki is the better default choice for most DevOps teams running Kubernetes today. Whichever stack you choose, the log queries you rely on during incidents should be captured in a runbook so they're immediately available when things go wrong. — not because ELK is bad, but because the typical Kubernetes debugging workflow maps well to Loki's label-based model, and the resource and operational cost difference is substantial.

ELK remains the right choice when full-text search across arbitrary fields is a genuine daily requirement, not a theoretical one. If your team regularly needs to find a specific customer record, transaction, or identifier across months of multi-service log data, that use case is hard to serve well with Loki.

ScenarioRecommendationKey reason
Kubernetes-native, Prometheus already runningLokiNative integration, unified Grafana UI, low cost
Arbitrary field search is a core workflowELKLoki's label model cannot replace indexed field search
Cost and resource efficiency is a priorityLokiOrder-of-magnitude cheaper to run at typical DevOps scale
Compliance, SIEM, or security analyticsELKElasticsearch has mature features Loki does not
Small team, new stack, fast time to valueLokiSimpler setup, lower operational surface
Existing ELK investment and expertiseELKSwitching cost likely outweighs Loki's resource savings
Long retention (90+ days) with cost constraintsLokiObject storage retention is dramatically cheaper

If you are starting fresh today with no existing investment in either stack, choose Loki. Set it up with Promtail, point it at your Kubernetes cluster, add it as a Grafana datasource, and you have a functional logging system in an afternoon. You can always add Elasticsearch later for specific use cases that outgrow what Loki can do.