> For the complete documentation index, see [llms.txt](https://docs.e6data.com/query-engine/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.e6data.com/query-engine/guides/operations/monitoring/metrics-and-alerts.md).

# Metrics and alerts

The metrics e6data exposes, the dashboards worth building, and the alerts to set on your observability stack.

e6data exposes operational telemetry that you route to your own observability stack via OpenTelemetry - see [Observability export](/query-engine/guides/operations/monitoring/observability-export.md) for setup. This page covers what to dashboard and which alerts to set.

## Key metrics

| Area              | Metrics                                                |
| ----------------- | ------------------------------------------------------ |
| Cluster health    | Executor count, CPU and memory, queue depth            |
| Query performance | P50/P95/P99 query latency, slow-query count            |
| Workload          | Queries per minute, concurrent queries, cache hit rate |
| Errors            | Failed-query rate, error-type breakdown, timeout rate  |

## Dashboards to build

* **Cluster health** - executor count, CPU/memory, and queue depth per cluster.
* **Query performance** - latency percentiles and slow-query count over time.
* **Workload patterns** - throughput, concurrency, and cache hit rate.
* **Errors** - failed-query rate and the breakdown by error type.

If your observability backend has prebuilt e6data dashboards, start from those.

## Alerts to set

| Alert                           | Suggested condition                                          |
| ------------------------------- | ------------------------------------------------------------ |
| Cluster fails to resume         | More than 5 minutes in `Pending` after a resume is triggered |
| Failed-query rate spike         | Above your established baseline for the cluster              |
| Query-timeout spike             | A sudden increase in timed-out queries                       |
| P99 latency breach              | Above your latency threshold, per cluster                    |
| Cache hit rate drop             | Below the expected baseline                                  |
| Workspace disabled unexpectedly | Any unplanned workspace disable                              |

Set these in your observability backend on the exported metrics; e6data doesn't send alerts itself.

## See also

* [Observability export](/query-engine/guides/operations/monitoring/observability-export.md)
* [Monitoring](/query-engine/guides/operations/monitoring.md)
* [Health checks](/query-engine/guides/operations/monitoring/health-checks.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.e6data.com/query-engine/guides/operations/monitoring/metrics-and-alerts.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
