Upcoming Webinar:

Getting Started with OpenObserve

August 27, 2026
11:00 AM ET

SLOs and Error Budgets

Alert on how fast you are spending, not on every error. Running on the same engine that already holds your logs, metrics, and traces.

Read the docsRead the docs
Image
Measure what you promised icon

Measure what you promised

Define good as a query, pick a target, get an error budget for the difference.

Alert on speed, not noise icon

Alert on speed, not noise

Burn-rate alerts fire on sustained damage and ignore the two minute spike.

One engine, not three tools icon

One engine, not three tools

SLOs run on the same backend already holding your logs, metrics, and traces.

SLOs in OpenObserve

Define what good means

Good is a query, not a checkbox

A scope filter sets the denominator. A good-when predicate picks the numerator out of it. A live preview splits good from bad on real data as you type, so you find out the definition is wrong before you save it, not a week later.

A real number on day one

Count SLIs cover request success rates. Time slice SLIs cover latency and freshness. Windows are rolling 7, 30, or 90 days, and backfill runs the moment you save, newest first, so you get a real number now instead of waiting a full window to see one. Alerts stay frozen until coverage clears the floor, so nobody gets paged by a half-filled window.

Define what good means

Alert on burn rate

Burn rate puts failures in proportion

Burn rate is the multiple of budget-neutral spending. At 1 you finish the window having used exactly your allowance. At 14.4 a 30 day budget is gone in about two days. A two minute spike barely moves the number, a sustained degradation eventually consumes it.

Two windows, so alerts resolve when incidents do

The long window establishes the problem is real and sustained. The short window confirms it is still happening. Long runs at twelve times short, which is what stops a burn-rate alert from hanging around for hours after the incident is over.

Alert on burn rate

Group by dimension

One series per group

One number tells you something is wrong, not what. Group by region, endpoint, or tenant and each one gets its own SLI, its own budget, and its own burn rate.

The overall number is measured, not summed

Group totals can be capped or incomplete. The headline SLI is computed on its own, so it never silently becomes the sum of whichever groups fit.

Group by dimension

One measurement, many alerts

Measurement and paging stay apart

An SLO only measures. Alerting on one is an ordinary alert, carrying the same destinations, silence, and severity as everything else you run.

Three urgencies, one objective

Fast burn pages you, because ×14.4 empties a 30 day budget in two days. Mid burn notifies a channel, slow burn files a ticket.

One measurement, many alerts

FAQs

Ready to get started?

Stop letting customers find your outages first.

Schedule DemoSchedule Demo