# AI & LLM monitoring: See what your agents are really doing

> Trace every agent, tool call, and model request. Score quality on live traffic and attribute cost per token, OpenTelemetry-native and priced per GB.

Source: https://openobserve.ai/ai-llm-monitoring/

---

Trace every agent, tool call, and model request. Score quality on live traffic and attribute cost to the token, in the same platform that runs the rest of your stack. OpenTelemetry-native, deploy anywhere, priced per GB instead of per span.

- [Start free](https://cloud.openobserve.ai/web/login/)
- [Read the docs](https://openobserve.ai/docs/integration/ai/)

## Purpose-built for agents. Priced for reality.

Most LLM tools are a silo you bolt onto your real observability stack, metered per span, locked to one cloud. OpenObserve is one platform for your AI and everything under it.

### Unified, not bolt-on

LLM traces land next to your logs, metrics, traces, and RUM. When an agent is slow, see the pod, database, or vector store behind it. No second tool, no swivel-chair.

### Deploy anywhere

Cloud in four regions, self-hosted, bring-your-own-cloud, or bring-your-own-bucket. Federated search spans regions and clouds while keeping egress under control.

### Radically efficient

A Rust engine on columnar Parquet storage, billed per GB. Agentic apps fan out into hundreds of spans per request. You shouldn't pay for each one.

## Instrument what you already run

Frameworks, model providers, gateways, and no-code builders, all through standard OpenTelemetry. If it emits gen_ai spans, it works today.

## From a runaway agent to the trace behind it

### Trace & map

**Know exactly what your agents touch.** Every agent request is a distributed trace. OpenObserve maps each agent to the models, tools, services, and datastores it calls, with request counts and error health on every edge, so a runaway loop or a failing tool is obvious at a glance.

- **Agent Graph** - Renders the full call tree across sub-agents, tools, and models.
- **Health at a glance** - Border colors flag healthy, degraded, and critical paths by error rate.
- **Compare releases** - Filter by environment, agent, and version to compare releases.

### Debug sessions

**Debug a bad answer in one view.** Open any session and replay the whole conversation: every turn, every tool call, the model behind it, and where the cost and latency actually went. Stop grepping logs to reconstruct what an agent did.

- **Session ribbons** - Break down cost, duration, and tokens per turn.
- **Tool, cost, and latency hotspots** - Surface the expensive, slow steps instantly.
- **One hop to the trace** - Jump straight from a turn to its full distributed trace.

### Evaluate in production

**Measure quality on real traffic, continuously.** Online evaluations score live spans, traces, and whole sessions the moment they arrive. Use LLM-as-judge with your own provider, or call your remote scorer. Set a healthy threshold per metric and watch quality on a live dashboard instead of a one-off notebook.

- **Built-in scorers** - Relevance, hallucination, toxicity, bias, and more.
- **Score at any scope** - Span, trace, or session scope, on a sample or everything.
- **Versioned score configs** - A Quality view that flags what needs attention.

### Close the loop

**From production trace to eval set.** Route real traces into review queues, let humans score them alongside the automatic evaluators, and distill the good and bad ones into datasets you can test future versions against. Discovery surfaces the failures worth reviewing in the first place.

- **Annotation queues** - Reviewer scores layered over system scores.
- **Distill to a dataset** - One click to turn a reviewed trace into a dataset.
- **Agent Behavior** - Catches loops and groups failures by kind.

## AI is a first-class layer in one unified stack

Your AI and LLM traffic is just another source, flowing through the same correlation engine as your frontend, APIs, databases, and infrastructure. Traces, metrics, logs, LLM observability, evals, and AI SRE live in one platform, queried together with SQL and PromQL.

## Your AI data, where you want it

Prompts and completions carry your most sensitive data. Run OpenObserve however your security and cost model demands, and query across all of it.

### Managed Cloud

Fully hosted, zero ops. Four regions today, with data residency where you need it.

### Self-hosted

A single Rust binary from laptop to petabyte scale. Open source, no license wall.

### Bring your own Bucket

Point storage at your own S3 or object storage bucket, or let us operate a managed deployment inside your cloud account. Your perimeter, your data; nothing locked in.

## Go deeper on AI observability

Guides, walkthroughs, and conversations on tracing agents, evaluating LLMs, and running it all on OpenTelemetry.

### Instrument the OpenAI Agents SDK with OpenTelemetry

[Learn more](https://openobserve.ai/blog/instrument-openai-agents-sdk-opentelemetry/)

### What Is LLM Observability? The 2026 Guide

[Learn more](https://openobserve.ai/blog/what-is-llm-observability/)

### The State of AI Observability: A Panel Discussion

[Learn more](https://youtu.be/E1_UQWd0aPY)

- [Explore all resources](/resources/)

## AI & LLM Observability FAQs

### Do I have to re-instrument my app?

No. OpenObserve ingests standard OpenTelemetry gen_ai spans and OpenInference conventions, so instrumentation you already have keeps working. It covers LangChain, CrewAI, LlamaIndex, OpenAI, Anthropic, LiteLLM, and 80+ more frameworks, providers, and gateways. Instrument once and route to OpenObserve, another backend, or both.

### How do evaluations work in production?

Online eval jobs score live spans, traces, or full sessions as they arrive, using LLM-as-judge with your own provider or a remote HTTP scorer. You define a score config with a healthy threshold, choose a sampling rate, and watch results roll up in the Quality dashboard. Human reviewers can score the same traces in annotation queues, and reviewed traces distill into datasets.

### What are my deployment options?

Managed cloud in four regions (US East, US West, Europe, India), self-hosted as a single binary, bring-your-own-cloud (we run a managed deployment in your account), or bring-your-own-bucket (data stays in your object storage). Federated search then queries across regions and clouds while keeping egress controlled.

### Can I manage it as code?

Yes. A Terraform / OpenTofu provider lets you provision and manage OpenObserve as code, reviewed in pull requests and reproducible across environments. Enterprise controls include RBAC and SSO.

### How is pricing different from Datadog or Langfuse?

OpenObserve bills per GB ingested and queried, not per LLM span, per unit, or per seat. Agentic applications that fan out into hundreds of spans per request stay predictable instead of spiking your bill, and users are unlimited. You can also sample eval jobs and redact sensitive fields with VRL pipelines before storage.

## Ready to get started with AI & LLM monitoring?

Point your OpenTelemetry exporter at OpenObserve and correlate traces, costs, and eval scores in one place. Pick the plan that fits your team.

### Cloud Free trial

Fully managed OpenObserve in the cloud. Start with a free trial - no infrastructure to run.

[Sign up free](https://cloud.openobserve.ai/web/login/)

### Enterprise Professional

Self-hosted Enterprise, free up to 50GB. Advanced features for growing teams.

[Download for free](/downloads/)

### Enterprise Premium

Enterprise-grade support, SSO, and SLAs for large-scale, multi-region deployments.

[Contact sales](/contact-sales/)

- [See Pricing Details](/pricing/)
