> For the complete documentation index, see [llms.txt](https://docs.e6data.com/query-engine/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.e6data.com/query-engine/guides/operations/upgrades-and-releases/rollbacks-failure-modes-recovery.md).

# Rollbacks, failure modes, and recovery

How cluster and platform rollouts behave, how automatic rollback works, expected downtime, and what to do when an upgrade fails.

This page explains what happens when you trigger an upgrade - how rollouts proceed, how rollback works, what downtime to expect, and how to respond when something fails. For *where* to trigger upgrades, see [Product and cluster upgrades](/query-engine/guides/operations/upgrades-and-releases/product-and-cluster-upgrades.md).

## Cluster rollouts

When you change a cluster's version, the upgrade happens **without downtime** and **recovers on its own if it fails**.

| Stage                          | Cluster status                                 | Activity History                                                                        |
| ------------------------------ | ---------------------------------------------- | --------------------------------------------------------------------------------------- |
| You save the new version       | **Updating**                                   | An **Update** entry appears, marked **In Progress**                                     |
| Upgrade completes successfully | **Running** (on the new version)               | The **Update** entry becomes **Success**                                                |
| Upgrade fails                  | **Running** (reverted to the previous version) | The **Update** entry becomes **Failed**, with a message explaining the automatic revert |

What the platform guarantees during the upgrade:

* **No downtime.** Your cluster keeps serving queries on its current version the whole time the new version is being brought up. Traffic only moves to the new version once it's healthy.
* **In-flight queries are protected.** Queries already running when the switch happens are allowed to finish.
* **A new (empty) cluster comes up directly** on its first version, with nothing to switch from.

## Automatic rollback (clusters)

If a cluster's new version fails to come up, the platform automatically reverts to the version that was working before. You don't have to do anything to recover:

* The cluster ends up **Running on its previous version** - it does not get stuck broken or offline.
* The **Activity History** shows the **Update** as **Failed**.
* The cluster row shows a message noting the most recent update failed and was automatically reverted, and that the cluster is running on the previous version.

{% hint style="info" %}
Because recovery is automatic, there is no separate "roll back" button for clusters. To intentionally move back to an earlier version, edit the cluster and select that version - it's just a normal version change.
{% endhint %}

## Platform rollout

Platform releases deploy the shared workspace components together. The deployment status progresses through:

| Status       | Meaning                                                                                                       |
| ------------ | ------------------------------------------------------------------------------------------------------------- |
| Not deployed | No platform release has been deployed through this section yet. This does **not** mean your platform is down. |
| Pending      | Selected, not yet started                                                                                     |
| Deploying    | Rollout in progress                                                                                           |
| Deployed     | Successfully live                                                                                             |
| Failed       | One or more components failed                                                                                 |
| Rolled Back  | The deploy failed and the platform was automatically reverted to the previous version                         |

The Platform Components section polls and updates this status in near real time while a deployment is in progress, and surfaces per-component results when a release fails.

{% hint style="info" %}
"Not deployed" on a healthy workspace is normal. Your platform components are set up and running from the moment your workspace is created - independently of this section. The workspace keeps running on what it was provisioned with until you select and deploy a release. See [What a new workspace starts with](/query-engine/guides/operations/upgrades-and-releases.md#what-a-new-workspace-starts-with).
{% endhint %}

## Platform rollback

Like clusters, a platform release rolls back automatically if it fails: if the new version doesn't come up healthy, the platform reverts to the previous version on its own and the status shows **Rolled Back**. You don't have to do anything to recover from a failed platform deploy.

You can also move back to an earlier release yourself - for example, to step back from a release that deployed successfully but you no longer want:

1. Go to **Settings → Version Upgrades** and find the **Platform Components** section.
2. In the release dropdown, select the earlier release you want.
3. Select **Upgrade** to deploy it.

{% hint style="info" %}
The currently-active release is excluded from the dropdown (you can't re-deploy what's already running), and automatically-managed releases are hidden.
{% endhint %}

## Downtime expectations

| Upgrade type           | Expected downtime                                                                                                    |
| ---------------------- | -------------------------------------------------------------------------------------------------------------------- |
| Cluster version change | None by design - traffic switches only after the new version is healthy, and in-flight queries are allowed to finish |
| Product upgrade        | Brief, service-dependent transition while the new version comes up                                                   |
| Platform upgrade       | A short transition is possible as shared components roll; plan a maintenance window for major platform changes       |

For clusters, because the cluster keeps serving its current version until the new one is healthy, a failed upgrade never takes it offline - it reverts to the version that was already serving.

## Failure scenarios and what to do

| Scenario                                | What happens                                                                                                            | What to do                                                                                                    |
| --------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| Cluster new version won't start         | Automatic rollback to previous version; a status message notes it                                                       | Check **Activity History** for the error; correct the cause (for example, resource sizing) before retrying    |
| Cluster edit is blocked                 | The cluster is mid-deployment or in a failed state                                                                      | Wait for it to reach a steady state, then edit again                                                          |
| Product upgrade fails                   | Row shows **Failed** with a notification                                                                                | Re-open the dropdown and retry, or select a different version; check the status detail                        |
| Platform release fails                  | Automatically rolled back to the previous version (status **Rolled Back**); a banner and per-component detail are shown | No action needed to recover; review the component-level failure detail before retrying                        |
| A newly published version isn't listed  | The list hasn't synced yet                                                                                              | Use the **refresh** control on the cluster Version field to pull the latest list (it changes nothing running) |
| Expected version is missing from a menu | The version may have been deprecated (retired) by e6data                                                                | Deprecated versions are hidden from the upgrade menus; pick a current version instead                         |

## Operational checklist before upgrading

* **Confirm the surface.** Cluster version → edit the cluster; product → Version Upgrades (product rows); platform → Version Upgrades → Platform Components section.
* **Check current status is steady.** Don't start an upgrade on a cluster that's mid-deployment or failed.
* **Know your rollback path.** Clusters and platform releases both recover automatically if an upgrade fails; for platform you can additionally re-deploy an earlier release yourself.
* **For platform changes,** consider a maintenance window for anything significant.
* **Watch the status** through completion - Activity History for clusters, the tab status and banners for products and platform.

## See also

* [Status signals](/query-engine/guides/operations/upgrades-and-releases/status-signals.md) - where each surface reports progress.
* [Zero-downtime upgrades](/query-engine/guides/clusters/zero-downtime-upgrades.md) - the cluster upgrade mechanism in depth.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.e6data.com/query-engine/guides/operations/upgrades-and-releases/rollbacks-failure-modes-recovery.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
