Skip to main content
Upcoming Webinar:

From Alert Noise to Actionable Signals: Lessons from Production

September 30, 2026
11:00 AM ET
Register

Replacing the Grafana LGTM Stack with a Unified Platform

Simran Kumari
Simran Kumari
September 29, 2026
19 min read
Don't forget to share!
TwitterLinkedInFacebook

Ready to get started?

Try OpenObserve Cloud today for more efficient and performant observability.

Table of Contents
Replacing the Grafana LGTM stack with a unified observability platform

OpenObserve is an open-source (AGPL-3.0) observability platform that replaces the Grafana LGTM stack (Loki, Grafana, Tempo, Mimir) with one system. Logs, metrics, and traces are ingested natively over OpenTelemetry, stored in one engine, and queried with SQL (plus PromQL for metrics), with dashboards, alerts, RUM, SLOs, incidents, and AI built in instead of bolted on.

This guide is written for teams evaluating a move off LGTM. It starts with the problems that usually trigger the search, then compares the LGTM stack and OpenObserve feature by feature and covers a migration plan.

TL;DR

  • Best unified replacement for LGTM: OpenObserve: it consolidates Loki, Tempo, and Mimir's ingestion and storage into one system and replaces Grafana's dashboards with a built-in visualization layer, so there is one deployment to run instead of four.
  • Best for teams hitting Loki cardinality limits: OpenObserve: columnar storage indexes logs without the label-cardinality constraints that slow Loki down.
  • Best for reducing query language sprawl: OpenObserve: SQL for logs and traces, PromQL for metrics, so teams drop LogQL and TraceQL.
  • Best for a low-risk migration: OpenObserve: native OTLP and Prometheus remote write let existing collectors fan out to OpenObserve alongside LGTM, with no instrumentation rewrite.
  • Best for consolidating beyond LGTM's scope: OpenObserve: pipelines, RUM, synthetics, alerts, incidents, SLOs, AI SRE, AI Assistant, MCP server, LLM observability, database monitoring, and BYOB live in the same platform.
  • Best for deployment choice: OpenObserve: use it as managed SaaS with OpenObserve Cloud, or self-host it on your own infrastructure.

What is the Grafana LGTM stack?

LGTM is shorthand for four separate Grafana Labs projects that are commonly deployed together to cover the three core observability signals:

  • Loki: log aggregation, indexes metadata labels rather than full log content, queried with LogQL.
  • Grafana: the visualization and alerting layer, connects to Loki, Tempo, Mimir, and other data sources as a front end.
  • Tempo: distributed tracing backend, queried with TraceQL, usually correlated with logs and metrics via trace IDs.
  • Mimir: long-term Prometheus-compatible metrics storage, queried with PromQL.

Each of these is its own project with its own release cadence, its own storage backend (object storage plus a compactor/ingester tier), and its own scaling model. Grafana itself does not store any telemetry; it is purely the query and visualization front end for the other three.

The LGTM stack gives you Grafana's dashboards, but Grafana does not ingest or store logs, metrics, or traces itself. Loki, Tempo, and Mimir have to be deployed, scaled, and upgraded independently for the stack to function.

Why are teams replacing the Grafana LGTM stack?

Four systems, four operational surfaces

Running LGTM in production means running Loki's ingester/querier/compactor components, Tempo's distributor/ingester/compactor components, and Mimir's equivalent set, each with its own object storage buckets, retention configuration, and upgrade path. A change to one component (say, a Loki version bump) can require re-validating compatibility with the Grafana data source in use. Teams report multi-week ramp-up time just to get all four components production-ready with high availability.

Three query languages to maintain

LogQL for logs, PromQL for metrics, TraceQL for traces: three distinct syntaxes with different mental models for filtering, aggregating, and joining data. Correlating a trace with its logs typically means copying a trace ID from one view into a query against another store, since there is no native join across the two.

Loki's high-cardinality ceiling

Loki indexes log labels, not log content, intentionally, to keep it lightweight. In practice, adding a high-cardinality label (user ID, request ID, pod name with dynamic suffixes) can blow up the index and degrade query performance, pushing teams toward restrictive labeling schemes that limit what they can search on later.

Mimir's scaling complexity

Mimir is a capable long-term Prometheus store, but its microservices architecture (distributor, ingester, querier, store-gateway, compactor) means a production deployment is itself a small distributed system to operate, monitor, and capacity-plan for, on top of the three other components in the stack.

The stack stops at L, G, T, and M

The self-hosted LGTM stack covers logs, metrics, traces, and dashboards. Almost everything a modern on-call team needs beyond that is either a Grafana Cloud product or another tool to run: SLO management, incident response and on-call, synthetic checks, frontend monitoring backends, AI-assisted investigation, and LLM observability. Every one of those adds a contract, a UI, or a deployment to the "four systems" problem above.

The alternative: OpenObserve at a glance

OpenObserve was built to collapse that sprawl into one platform:

  • One binary handles ingestion, storage, query, dashboards, and alerting. Run it as a single node or scale it out with the Kubernetes Helm chart.
  • One storage engine keeps logs, metrics, and traces in columnar Parquet on object storage (S3, GCS, Azure Blob, MinIO).
  • One query model: SQL across every signal, plus native PromQL for metrics.
  • OpenTelemetry-native ingestion over OTLP, plus Prometheus remote write, Fluent Bit, Vector, and other common shippers.
  • Everything else built in: pipelines, RUM, synthetics, alerts, incidents, SLOs, and AI features share the same data, the same RBAC, and the same UI.

LGTM stack vs. OpenObserve: side-by-side feature comparison

The table compares Grafana (the LGTM stack) with OpenObserve. Features that are only available in Grafana Cloud, or only in OpenObserve Enterprise or OpenObserve Cloud, are marked.

Capability Grafana OpenObserve
Logs Loki: label-indexed, LogQL, high-cardinality labels degrade performance Native, full-text and SQL search, columnar storage without label-cardinality limits
Metrics Mimir: Prometheus-compatible, PromQL, multi-component deployment Native, PromQL-compatible, Prometheus remote write, also queryable with SQL
Traces Tempo: TraceQL, correlated to logs via trace IDs across stores Native OTLP traces, SQL-queryable, service graph, one-click pivot to logs and metrics
Dashboards Grafana: separate front end wired to each backend Built in: 19 chart types plus custom ECharts, SQL/PromQL, drilldowns, variables
Prebuilt dashboards Large community dashboard library Prebuilt dashboards for common sources, plus a Dashboard Migrator that converts Grafana JSON
Pipelines Processing happens in the collector (Grafana Alloy / OTel Collector) before data reaches each backend Built-in real-time and scheduled pipelines: VRL functions, enrichment tables, redaction, routing
RUM (real user monitoring) Faro Web SDK collects browser telemetry; RUM views and session replay via Frontend Observability (Grafana Cloud only) Native RUM: session replay, Core Web Vitals, error tracking, correlated to backend traces
Synthetic monitoring Blackbox exporter for HTTP, TCP, ICMP, and DNS probes; managed k6-based Synthetic Monitoring (Grafana Cloud only) Built-in API, TLS, and Playwright browser checks from multiple regions, plus status pages
Alerts Grafana Alerting plus Loki and Mimir rulers SQL and PromQL alerts, real-time and scheduled, deduplication
Incidents Grafana IRM for incident response and on-call (Grafana Cloud only) Automatic alert correlation into incidents, alert graph, correlated telemetry, activity timeline
SLOs Grafana SLO (Grafana Cloud only); self-hosted teams generate recording rules with tools like Sloth or Pyrra Native SLOs with error budgets, multi-window burn-rate alerts, per-dimension SLIs
AI SRE AI investigation features (Grafana Cloud only) AI SRE agent investigates the moment an alert fires, writes an evidence-backed RCA, bring your own LLM
AI Assistant Grafana Assistant (Grafana Cloud only) AI Assistant turns plain English into queries, dashboards, and alerts on your data
MCP server Open-source Grafana MCP server queries through Grafana data sources Built-in MCP server with RBAC-scoped, audited access to logs, metrics, and traces
AI / LLM observability AI Observability (Grafana Cloud only) LLM and agent tracing with token cost per span; online and offline evaluations and datasets
Database monitoring Prometheus exporters plus community dashboards; Database Observability (Grafana Cloud only) Integrations for PostgreSQL, MySQL, MongoDB, Redis, Oracle, and more, with prebuilt dashboards and alerts
Bring Your Own Bucket Self-hosted: your own buckets, one set per backend (Loki, Mimir, Tempo); Grafana Cloud stores data for you Self-hosted: one bucket for all signals; OpenObserve Cloud Enterprise can keep data in your S3 or Azure account
Query languages 3 (LogQL, PromQL, TraceQL) SQL for everything, plus PromQL for metrics
Deployment units 4 separate systems 1 binary/container
License AGPLv3 (Loki, Grafana, Tempo, Mimir) AGPL-3.0

Grafana Cloud capabilities change frequently; check Grafana's documentation for current availability. The point for a buyer is structural: on LGTM, each capability past the core four signals is another product or another deployment, while in OpenObserve it is a menu item in the same UI, running on the same data.

Feature-by-feature: what you get when you replace LGTM

Logs: search everything, not just labels

OpenObserve stores logs in columnar format and lets you search them with full-text search or SQL, so you are not forced to decide up front which fields become labels. High-cardinality fields such as request IDs and user IDs are just columns. Learn more on the OpenObserve logs page or in the log search docs.

Powerful log search in OpenObserve with SQL and full-text queries

Metrics: keep PromQL, add SQL

PromQL support in OpenObserve is scoped to metrics. OpenObserve accepts Prometheus remote write and runs PromQL natively, so existing Prometheus and Mimir dashboards, recording rules, and alert expressions carry over largely unchanged. The same metrics are also queryable with SQL when an investigation needs joins or correlation with logs and traces. This mirrors how LGTM already splits things (PromQL was always Mimir's language, never Loki's or Tempo's), but OpenObserve replaces LogQL and TraceQL with a single SQL engine.

SQL and PromQL query support in OpenObserve

See the OpenObserve metrics platform or the benchmarked OpenObserve vs Prometheus and Mimir comparison.

Traces: correlated with logs and metrics by default

Logs, metrics, and traces share one resource model in OpenObserve, so correlation works out of the box. From a slow trace, jump directly to the log lines and metrics from that same request, without copying a trace ID between separate UIs. Explore distributed tracing in OpenObserve or the traces docs.

Cross-signal correlation from a trace to its logs and metrics in OpenObserve

Dashboards and prebuilt dashboards

Dashboards are part of the same binary rather than a separately wired front end. Build panels by dragging fields against logs, metrics, and traces using SQL or PromQL, add variables to filter every panel at once, and use drilldowns to jump from a data point into the underlying logs or traces.

Drill-down analysis in OpenObserve dashboards

You do not have to start from a blank canvas. OpenObserve ships prebuilt dashboards for common sources such as Kubernetes, AWS services, databases, and Docker, which you import as JSON. For your existing Grafana dashboards, the OpenObserve Dashboard Migrator converts Grafana dashboard JSON automatically. See visualization and dashboards for the full chart library and the dashboards docs to build your own.

Metrics dashboard in OpenObserve

Pipelines: shape data before it lands

In LGTM, parsing, redaction, and routing live in the collector, configured separately for each backend. OpenObserve has pipelines built into the platform: real-time pipelines that parse, enrich, redact, and route data as it arrives, and scheduled pipelines that aggregate data on a timer (for example, turning logs into metrics). Pipelines use VRL functions, CSV enrichment tables, and conditions, built in a visual editor. See the pipelines docs.

OpenObserve pipeline editor with source stream, VRL function, and destination nodes

Real user monitoring (RUM) and session replay

OpenObserve captures frontend telemetry natively: session replay with privacy masking, Core Web Vitals (LCP, INP, CLS), and JavaScript error detection, all queryable with SQL and correlated against backend traces and logs in the same store. In the Grafana ecosystem, the equivalent backend is a separate Grafana Cloud product. See the RUM docs or the hands-on Datadog vs OpenObserve RUM comparison.

Session replay for real user monitoring in OpenObserve

Synthetic monitoring

Scheduled API checks, TLS checks, and real Playwright browser checks run from multiple regions, with every result landing next to your logs, metrics, and traces instead of in a separate synthetic monitoring silo. Synthetics is currently in Beta on OpenObserve Cloud; self-hosted support is not yet available. See synthetic monitoring, the synthetics docs, or how it compares in the top synthetic monitoring tools in 2026.

Unified synthetic check results across regions in OpenObserve

Alerts

OpenObserve alerts use SQL for logs and traces and PromQL for metrics, run in real time or on a schedule, and deliver to Slack, email, PagerDuty-style webhooks, and more. Aggregation windows and silence periods cut noise, and ML-based anomaly detection catches unusual patterns without hand-tuned thresholds. Instead of spreading rules across Grafana Alerting and separate Loki and Mimir rulers, every alert lives in one place. See the alerts docs.

OpenObserve alert types: scheduled, real-time, and anomaly detection

Incidents

When an alert is set to create incidents, OpenObserve groups related firings into a single incident using shared dimensions such as service, cluster, and namespace. Each incident has an overview dashboard, an alert graph that highlights likely root-cause alerts, correlated logs, metrics, and traces filtered to the incident window, and an activity timeline for collaboration. Notifications fire when an incident is created or escalates, not on every repeat alert. See the incidents docs.

OpenObserve incident grouping related alerts into a single incident

SLOs and error budgets

OpenObserve SLOs let you define "good" as a query, pick a target, and get an error budget for the difference, on the same engine that already holds your logs, metrics, and traces. Burn-rate alerts use a long and a short window, so they page on sustained damage and resolve when the incident does. You can group an SLO by region, endpoint, or tenant and get a separate budget per group. Read the SLOs launch notes or the SLO docs for details.

OpenObserve SLO list showing budget remaining, burn rate, and status against target

AI SRE agent

The AI SRE agent starts investigating the moment an incident is created: it queries the correlated logs, metrics, and traces, follows service dependencies, compares against past incidents, and writes a structured root cause analysis with the evidence it used. You bring your own LLM provider (or a self-hosted model), so telemetry can stay inside your infrastructure. Our guide to AI incident management and automated root cause analysis covers the approach in depth, and the AI SRE agent docs cover setup.

OpenObserve AI root cause analysis report with incident summary, timeline, and evidence

AI Assistant

The AI Assistant turns plain-English requests into real SQL, PromQL, or VRL queries, dashboards, and alerts grounded in your actual data. It is context-aware: on the traces page, for example, it picks up the current stream, time range, and selected trace.

O2 Assistant turning a natural-language request into a query in OpenObserve

MCP server

The OpenObserve MCP server gives AI agents such as Claude Code, Cursor, and other MCP clients direct access to your telemetry. Agents can search logs, run metric and trace queries, and create alerts or dashboards, with every call authenticated, scoped by OpenObserve RBAC, and recorded in the audit trail. See the MCP setup docs.

Claude Code using the OpenObserve MCP server to report service health and create alerts

AI and LLM observability

If you run LLM applications or agents, OpenObserve captures every prompt, tool call, and agent handoff as an OpenTelemetry span, with input and output token cost computed per span, on Cloud and self-hosted deployments. AI observability adds quality scoring on live traffic (LLM-as-a-judge or your own scorer), offline experiments against datasets, and human review queues; evaluations are an enterprise feature. Because LLM spans sit in the same store as infrastructure spans, "model problem or infra problem?" is one click. See LLM observability for tracing details, or the LLM tracing and LLM evaluations docs.

OpenObserve LLM trace showing prompts, tool calls, and token cost per span

Database monitoring

OpenObserve has database monitoring integrations for PostgreSQL, MySQL, MongoDB, Redis, Oracle, Cassandra, DynamoDB, Snowflake, and more, collected with the OpenTelemetry Collector. Prebuilt dashboards and alert rules cover connections, slow queries, resource use, and errors, and database telemetry correlates with the application traces that called it. See the database integration docs.

OpenObserve database dashboard for Amazon RDS showing connections, memory, and transaction log metrics

Bring Your Own Bucket (BYOB)

Self-hosted OpenObserve always writes to your own object storage, but with one bucket for all signals instead of one per backend. For OpenObserve Cloud customers, Bring Your Own Bucket connects your own S3 bucket or Azure Blob container, so telemetry stays in your account, in your region, under your own access controls, while OpenObserve runs ingestion, compaction, and queries. The bucket must be in the same region and cloud provider as your OpenObserve Cloud deployment; see the BYOB documentation.

Bring Your Own Bucket in OpenObserve Cloud keeping telemetry in your own storage account

How to get started with OpenObserve

OpenObserve gives you two ways to get started:

  • OpenObserve Cloud: the fastest way to try it. Sign up for OpenObserve Cloud and start sending logs, metrics, and traces with nothing to install or run.
  • OpenObserve self-hosted: run it on your own infrastructure. Download OpenObserve and deploy it as a single binary or on Kubernetes with the Helm chart.

The quickstart guide walks through both options. See OpenObserve pricing for current plan details.

How to evaluate OpenObserve as your LGTM replacement

Use these questions to structure a proof of concept. Each one maps to a common reason teams leave LGTM:

  1. Operations: How many components do we run today for logs, metrics, traces, and dashboards? OpenObserve replaces them with one binary or one Helm release.
  2. Query experience: Can engineers answer a cross-signal question (trace to logs to metrics) in one place? Test it in OpenObserve with the correlation and drilldown flows above.
  3. Cardinality: Which fields did we avoid indexing in Loki? Ingest them into OpenObserve and search on them directly.
  4. Scope beyond LGTM: Which Grafana Cloud products or extra tools do we pay for or run for SLOs, incidents, synthetics, RUM, or AI? Check each against the comparison table.
  5. Data control: Where must telemetry live? Pick self-hosted, Cloud with BYOB, or Enterprise in your own cloud.
  6. Cost: What do Loki and Mimir retention cost today? Compare against OpenObserve's storage footprint on the same data.

How do you migrate from the LGTM stack to OpenObserve without rewriting instrumentation?

The migration path does not require touching application code. If you already instrument with OpenTelemetry, the OTel Collector can send telemetry to LGTM and OpenObserve at the same time:

  1. Add OpenObserve as a second OTLP exporter destination in your existing OTel Collector, alongside your current Loki, Tempo, and Mimir endpoints.
  2. For metrics scraped by Prometheus and remote-written to Mimir, add OpenObserve's Prometheus remote-write endpoint as a second remote_write target.
  3. Import prebuilt dashboards for your common sources, and convert your most-used Grafana dashboards with the OpenObserve Dashboard Migrator.
  4. Recreate critical alerts and define SLOs for your key services, then turn on incident creation for the alerts that matter.
  5. Run both stacks in parallel long enough to validate query results and alert parity.
  6. Cut traffic over and decommission Loki, Tempo, and Mimir once OpenObserve is handling production queries.

Because OpenObserve speaks OTLP and Prometheus remote write natively, none of this changes how services emit logs, metrics, or traces, only where that data is sent.

Architecture: four backends vs. one

OpenObserve unified architecture compared to Grafana LGTM stack

In the LGTM stack, Grafana is a front end wired to Loki, Tempo, and Mimir separately, each with its own storage and scaling tier. In OpenObserve, ingestion, storage, query, visualization, and alerting live in the same binary, deployable as a single node for smaller workloads or via Helm chart in Kubernetes for high availability. See the architecture docs.

How much can you save by replacing the LGTM stack with OpenObserve?

Savings come from two places: fewer infrastructure components to run (no separate ingester, compactor, and store-gateway tiers for three different systems) and OpenObserve's columnar storage compression, which gives approximately 140x lower storage costs in typical log workloads compared to Elasticsearch-based stacks (see the benchmark); actual results vary based on data entropy and cardinality. How much that translates to against an existing LGTM deployment depends on your current retention windows and node counts for Loki and Mimir, since those are usually the largest cost centers. Consolidating the extra tools you run beyond LGTM (SLOs, incidents, synthetics, RUM, AI) adds to those savings. See the cost of self-hosting observability on ClickHouse for how storage engine choice affects self-hosted economics more broadly.

Conclusion

The LGTM stack works, and plenty of teams run it successfully, but "works" comes with the overhead of operating four distributed systems, training engineers on three query languages, and adding more products for everything past logs, metrics, traces, and dashboards. OpenObserve collapses that into one open-source platform: one binary, native OpenTelemetry ingestion, SQL for querying, and pipelines, RUM, synthetics, alerts, incidents, SLOs, AI SRE, the AI Assistant, the MCP server, LLM observability, database monitoring, and BYOB built in. You can run it as managed SaaS, free and self-hosted, or with enterprise features on your own infrastructure.

If you are evaluating the broader field, see Top 10 Grafana Alternatives in 2026, the direct OpenObserve vs Grafana comparison, or the full OpenObserve as a Grafana alternative breakdown.

Getting started takes minutes, not weeks. Try OpenObserve in the cloud with a free trial, or self-host it as a single binary or via the Kubernetes Helm chart.

Frequently Asked Questions

About the Author

Simran Kumari

Simran Kumari

LinkedIn

Passionate about observability, AI systems, and cloud-native tools. All in on DevOps and improving the developer experience.

Follow OpenObserve on Google

Add OpenObserve as a preferred source to see more of our articles in Google Search and Top Stories.

Latest From Our Blogs

View all posts