> For the complete documentation index, see [llms.txt](https://docs.e6data.com/query-engine/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.e6data.com/query-engine/get-started/architecture/detailed-architecture-walkthrough/the-components.md).

# The components

Each component is an independent service with a narrow contract, communicating over gRPC and Thrift. This is what lets the engine scale — and be operated — one piece at a time.

***

#### 01 — Core Engine

**Query Planner** · *Adaptive · learning optimizer*\
An adaptive, learning optimizer built for workload drift. The Calcite-based optimizer uses a custom cost model robust to sparse statistics. A LEO feedback loop captures runtime cardinalities and plan outcomes, associating them with query templates through query-graph signatures — improving estimates, avoiding repeat mistakes, and adapting to workload drift. Common CTEs are detected and reused automatically, with early-stage automatic materialized-view detection; hypergraph join enumeration with A\* search handles very large join graphs, and SQL UDFs are inlined with recursive CTE decorrelation. Optimized for BI workloads and tools such as Tableau and Power BI, the planner plans everything and executes nothing — so planning capacity scales independently of compute.\
\&#xNAN;*Tags: Apache Calcite · custom cost model · LEO feedback loop · hypergraph join enumeration*

**Executor** · *Rust · Arrow-centric*\
Stateless Rust workers, Arrow-centric end to end. Plan fragments run as pipelines of vectorized columnar operators — scan, filter, hash and partitioned joins, aggregation, window, sort — on an async runtime, with data staying in Arrow's columnar form from scan to result. Execution is distributed first: exchange operators repartition data across the fleet so large joins and aggregations scale out rather than concentrating on a single worker, with tracked memory and spill-to-disk for workloads that exceed memory.\
\&#xNAN;*Tags: Rust · Apache Arrow · vectorized operators · distributed exchange · spill-to-disk · stateless pods*

**Distribution Services** · *Discovery · coordination · workload mgmt*\
Service discovery keeps planners and executors addressable as pods are added and removed; scheduling algorithms place plan fragments across executors to maximize throughput and manage the data exchange between them; and query workload management — queueing, admission control, executor allocation — keeps concurrent workloads sharing the cluster predictably, with configurable guardrails that prevent one bad query from disproportionately impacting others.\
\&#xNAN;*Tags: service discovery · coordination · workload management · queueing*

**Metadata Services** · *Catalog abstraction*\
One interface over many catalogs. Metadata services abstract Hive Metastore, AWS Glue, Unity Catalog, Polaris, and other Iceberg REST (IRC) catalogs behind a single schema and statistics layer, and keep it current by discovering and syncing tables automatically — the planner sees one consistent view of your lake, whatever runs underneath.\
\&#xNAN;*Tags: catalog abstraction · Iceberg REST · schema registry · statistics · auto sync*

**Transpiler** · *Any dialect → E6 SQL*\
Makes the engine multi-dialect. The planner hands incoming SQL to the transpiler, which converts other engines' dialects into the E6 dialect before planning. Existing queries, dashboards, and scheduled jobs run as written, without a rewrite.\
\&#xNAN;*Tags: dialect translation · planner-integrated · zero rewrite*

**Governance** · *Ranger · Unity · more*\
Plugs into the governance you already run. Access control integrates with Apache Ranger, Unity Catalog governance, and similar systems, alongside built-in authentication, RBAC authorization, audit trails, and query history — so lake-side permissions are enforced end to end.\
\&#xNAN;*Tags: Apache Ranger · Unity governance · RBAC · audit*

***

#### 02 — Supporting Services

**Protocol Layer** · *Entry point*\
The entry point for every query. Terminates client connections over the PostgreSQL wire protocol, the Trino protocol, the Spark protocol, and custom gRPC — plus the E6 UI — authenticates sessions, and routes each statement into planning.\
\&#xNAN;*Tags: pg-wire · Trino protocol · Spark protocol · gRPC*

**Platform Operators** · *Kubernetes control*\
Kubernetes operators manage the fleet: provisioning workspaces, scaling executor pools, rolling out upgrades, and wiring monitoring. The engine is cloud-portable because the platform layer speaks Kubernetes, not any one cloud's control plane.\
\&#xNAN;*Tags: K8s operators · autoscaling · observability*

***

#### 03 — Data Plane — In the Customer's Environment

**Object Storage** · *S3 · GCS · ADLS*\
The engine reads directly from your buckets across AWS, GCP, and Azure through a unified cloud-storage layer. Storage scales with your lake, not with the engine.

**Open Table Formats** · *Iceberg · Delta · Hudi · Parquet*\
First-class readers for Apache Iceberg, Delta Lake, and Hudi — including transaction log handling, snapshot isolation, and partition/file pruning — plus raw Parquet. Format metadata drives planning, so the optimizer skips data before compute ever touches it.

**Existing Catalogs** · *No migration*\
e6data attaches to the catalogs you already run — Hive Metastore, AWS Glue, Unity Catalog, and more — as a reader. Your governance, lineage, and access tooling keep working unchanged.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.e6data.com/query-engine/get-started/architecture/detailed-architecture-walkthrough/the-components.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
