# Your AI SRE. Root Cause Before You Wake Up.

> AI SRE Agent: an always-available Site Reliability Engineer in OpenObserve Enterprise that investigates incidents the moment an alert fires to cut MTTR.

Source: https://openobserve.ai/ai-sre/

---

Meet your 24/7 AI SRE. It starts investigating the second an alert fires, so you arrive to root causes instead of raw telemetry.

- [Start Free Cloud Trial](https://cloud.openobserve.ai/web/login/)
- [Talk to a Human](/demo/)

### Investigating Before You Log In

The agent starts the moment an alert fires, not when someone opens a ticket.

### Verifiable Evidence

Get a complete audit trail of every log, trace, and metric during the agent’s investigation.

### Bring Your Own AI Provider

Connect your own LLM provider with your API key, so your team stays in control of security, governance, and spend.

## How AI SRE Agent Works

### Systemic Intelligence

- **Signal Analysis Across All Telemetry** - Analyze logs, metrics, and traces across your entire environment automatically, investigating every signal and dependency exactly like a senior SRE.
- **Structured Findings with Context** - The agent automatically documents actionable remediation plans, delivering a full incident breakdown including diagnosis, root cause, and fix.

[Learn More](https://openobserve.ai/docs/user-guide/analytics/incidents/#correlated-telemetry)

### AI Analysis

- **Complete Evidence Chain for Every Finding** - Verify findings with a complete evidence chain. Review correlated logs, metrics, and traces, inspect service topology graphs, and trace the incident’s timeline.
- **Automated Correlation & Impact Mapping** - Map dependencies across distributed services instantly. The agent identifies upstream causes and downstream effects.

[Learn More](https://openobserve.ai/docs/user-guide/analytics/incidents/#ai-powered-root-cause-analysis)

### Agentic Control

- **Autonomous Tool Execution Without Human Triggers** - The agent uses OpenObserve’s own tooling via MCP, the same way a person would navigate the UI, except it never misses a step.
- **Evidence and Reasoning at Every Step** - Unlike black-box systems, OpenObserve’s AI SRE shows exactly what data it analyzed and how it reached conclusions, helping engineers validate recommendations and learn from AI decision-making.

[Learn More](https://openobserve.ai/docs/user-guide/analytics/incidents/#getting-started)

### Incident Automation

- **Immediate Event-Driven Response** - Triggered instantly, no delay. The agent initiates the investigation cycle the moment the alert fires.
- **Never Forgets a Past Incident** - Link current anomalies to historical incident data, and every incident becomes part of the knowledge base.

[Learn More](https://openobserve.ai/docs/user-guide/analytics/incidents/#incident-overview-dashboard)

## Measured against industry leaders

Same telemetry, same workloads, one platform. Every number is OpenObserve against a named vendor - not an industry average.

8x cost reduction means you can unify your observability into a single platform.

140x storage means longer retention doesn't necessarily mean expensive bills.

5x to 15x faster queries mean dashboards load in milliseconds, not minutes.

See how much you would save switching today.

- [See all comparisons](/comparison/)

## Teams trust OpenObserve to investigate alongside them

## AI SRE FAQs

### What does the OpenObserve AI SRE agent do?

The AI SRE is a background service that powers intelligent workflows in OpenObserve Enterprise. When an alert fires, the agent immediately begins investigating - no human trigger required - examining logs, metrics, and traces across the environment, mapping upstream causes and downstream effects through service dependencies, and linking the anomaly to historical incidents it has seen before. It then delivers a structured root cause analysis with diagnosis, evidence, and an actionable remediation plan before an engineer even opens the ticket. Every finding carries a complete evidence chain - the correlated telemetry, service topology graphs, and incident timeline it examined - so engineers can verify the reasoning and learn from it rather than trusting a black box.

### Will my logs and metrics be sent to a third-party LLM?

Data routing depends entirely on your configured provider. The agent assembles a context window of relevant logs, metrics, and traces to send to your chosen model. If you use a self-hosted or private cloud LLM, all data remains within your infrastructure. OpenObserve supports any OpenAI-compatible endpoint to ensure complete data sovereignty.

### How does the AI SRE learn the context of my infrastructure?

The agent uses a three-phase pipeline for every incident. It starts with Context Assembly to gather logs, metrics, and dependency maps. It then performs Historical Pattern Matching before using LLM Analysis to generate a structured RCA. Accuracy improves with every incident handled without requiring manual configuration.

### Which LLM providers are supported?

The AI SRE Agent supports OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, DeepSeek, and OpenRouter. You can also connect any OpenAI-compatible endpoint, including self-hosted models, to meet specific cost or compliance requirements.

### How is the AI SRE different from tools like Datadog or Grafana?

Traditional tools often force teams to sample or drop telemetry to manage costs, leaving AI models with incomplete data to reason over. OpenObserve stores full-fidelity data at up to 140x lower storage cost, so the agent analyzes the complete dataset - every log line, span, and metric around the incident - rather than a sampled fraction. The agent also works through the Model Context Protocol (MCP), navigating OpenObserve's own tooling exactly like a human engineer would, which makes every step reproducible and auditable instead of hidden inside a proprietary pipeline. And unlike bundled AI assistants, you bring your own LLM provider and API key, keeping model choice, governance, and spend under your control. The result is a verifiable investigation, not an opaque suggestion.

### How is the AI SRE Agent priced?

O2 Cloud Pricing is based on AI Credits, which power the agent’s autonomous investigations and reasoning cycles. $0.50 per AI Credit. New accounts receive 20 free credits (one-time) to start investigating immediately.

### How does the OpenObserve AI SRE agent compare to Dash0 Agent0 or Datadog Bits AI?

All three aim at automated root-cause analysis. The difference is where they run and what they cost: Dash0's Agent0 and Datadog's Bits AI are only available inside closed-source, cloud-only platforms billed per signal or per host. OpenObserve's AI SRE agent runs on the same open-source platform you can self-host, reasons over logs, metrics, and traces stored in your own object storage, and is priced on ingestion rather than per seat or per signal.

## Explore guides, videos, and articles

to help you get the most out of AI SRE.

### AI SRE setup guide

[Learn more](https://openobserve.ai/docs/administration/deployment/sre-agent-setup-guide/)

### Incident Correlation

[Learn more](https://openobserve.ai/blog/incident-correlation/)

### AI Powered Root Cause Analysis

[Learn more](https://openobserve.ai/webinars-videos/sre-agent/)

- [Explore All Blogs](/blog/)

## Ready to get started?

Try OpenObserve today for more efficient and performant observability.

- Get Started For Free
- [Schedule Demo](/demo/)
