> For the complete documentation index, see [llms.txt](https://docs.e6data.com/query-engine/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.e6data.com/query-engine/guides/deployment/azure-in-vpc/configure-registry-kubernetes-networking.md).

# Configure registry, Kubernetes, and networking

Build the Azure infrastructure (VNet, AKS, identities, ADLS storage) and install the Kubernetes platform components for an In-VPC e6data deployment on AKS.

This page builds the Azure infrastructure for an In-VPC deployment and installs the cluster-wide Kubernetes components the e6data workspace runs on. Run these steps with the Azure CLI (`az`), `kubectl`, and Helm.

Part 1 provisions Azure infrastructure (run once per cluster). Part 2 installs the Kubernetes platform components (also once per cluster). Per-workspace deployment is covered on the next page, [Deploy workspace and e6data](/query-engine/guides/deployment/azure-in-vpc/deploy-workspace-and-e6data.md).

## Infrastructure requirements

Review these before you begin. They apply whether you create a new VNet (Part 1, Step 1) or reuse an existing one.

### VNet options

| Option            | When to use                                                                                        |
| ----------------- | -------------------------------------------------------------------------------------------------- |
| **New VNet**      | Recommended for an isolated e6data deployment. Follow Part 1, Step 1.                              |
| **Existing VNet** | When e6data must share the network with other workloads. Skip Step 1 and meet the checklist below. |

### VNet and subnet sizing

With **Azure CNI overlay** (used here), pods get IPs from the **pod CIDR** (`10.244.0.0/16`), not the subnet - so the AKS subnet only needs addresses for **nodes**, not pods.

| Resource         | Minimum | Recommended         | Notes                                                  |
| ---------------- | ------- | ------------------- | ------------------------------------------------------ |
| **VNet CIDR**    | /20     | /16 (`10.0.0.0/16`) | Must not overlap pod/service CIDRs or peered networks. |
| **AKS subnet**   | /24     | /17 (`10.0.0.0/17`) | Nodes only (pods are on the overlay).                  |
| **Pod CIDR**     | -       | `10.244.0.0/16`     | Virtual (overlay); must not overlap the VNet or peers. |
| **Service CIDR** | -       | `172.20.0.0/16`     | Cluster-internal; must not overlap the VNet or peers.  |

### Availability zones

| Requirement | Minimum | Recommended       |
| ----------- | ------- | ----------------- |
| **Zones**   | 2       | 3 (`1`, `2`, `3`) |

e6data spreads workloads across zones for high availability - the engine NodePool lists zones `1`, `2`, and `3`.

### Azure quotas (vCPU)

Quota is per VM family, per region. Check and request increases before deploying - increases can take 24–48 hours.

| Family                   | Used by                                                 |
| ------------------------ | ------------------------------------------------------- |
| **Dpsv6 / Dpsv5 (ARM)**  | System pool (`Standard_D2ps_v6`) and the operator pool. |
| **D / E / F / L (ARM)**  | Engine nodes (Karpenter).                               |
| **Total Regional vCPUs** | All of the above.                                       |

Sponsorship and free subscriptions often start with **0 quota for ARM families** - request an increase first if `az aks create` fails with a quota or SKU-not-available error.

### VM families

| Component        | VM families        | Architecture | Notes                                                     |
| ---------------- | ------------------ | ------------ | --------------------------------------------------------- |
| **System pool**  | `Standard_D2ps_v6` | arm64        | AKS system pods; use `Standard_D4as_v5` for x86.          |
| **e6-operator**  | `D`, `E`           | arm64        | Operator and cert-manager (`e6operator` NodePool).        |
| **Engine nodes** | `D`, `E`, `F`, `L` | arm64        | Query engine (Karpenter); the `L` family adds local NVMe. |

The e6data engine is optimized for **ARM64** for price-performance. Confirm your region has the chosen families available.

### Existing VNet checklist

If you reuse an existing VNet, confirm:

* [ ] No CIDR overlap between the VNet, pod CIDR (`10.244.0.0/16`), service CIDR (`172.20.0.0/16`), or any peered network.
* [ ] The AKS subnet has room for your node count (with CNI overlay, pods don't consume subnet IPs).
* [ ] Nodes have outbound internet (a NAT Gateway is recommended).
* [ ] `Microsoft.Storage` and `Microsoft.ContainerRegistry` service endpoints are enabled on the subnet.

### Existing AKS cluster checklist

If you already have an AKS cluster, **skip Part 1 Steps 1–3** and verify the cluster meets the requirements below first:

```bash
az aks show -g <rg> -n <cluster> \
  --query "{k8s:kubernetesVersion, oidc:oidcIssuerProfile.enabled, wi:securityProfile.workloadIdentity.enabled, network:networkProfile.networkPlugin, outbound:networkProfile.outboundType}" -o yaml
```

| Requirement                         | Expected                                                                  | If not met                                                                           |
| ----------------------------------- | ------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ |
| **OIDC issuer + Workload Identity** | both `true`                                                               | `az aks update -g <rg> -n <cluster> --enable-oidc-issuer --enable-workload-identity` |
| **Kubernetes version**              | 1.29+ (1.33 ideal)                                                        | Upgrade the cluster.                                                                 |
| **Outbound egress**                 | reaches `e6labs.azurecr.io`, `mcr.microsoft.com`, `quay.io`, `app.e6.run` | Allow-list these (Azure Firewall / UDR).                                             |
| **Deployer permissions**            | User Access Administrator                                                 | Grant it - required for the role assignments.                                        |

{% hint style="warning" %}
OIDC issuer + Workload Identity is the setting that silently breaks everything if it's off. Check it first.
{% endhint %}

Then skip to **Step 4** (cluster identity) and continue through Parts 1 and 2 as written.

## Part 1 - Azure infrastructure

By the end of this part you will have a VNet with a NAT Gateway, an AKS cluster with OIDC issuer and Workload Identity enabled, a system node pool and an engine node pool strategy, an ADLS Gen2 storage account, and a workspace Managed Identity federated to the e6data service accounts.

### Step 0: Set variables

```bash
# Naming and placement
export PREFIX="e6data"                  # resource name prefix (lowercase, no spaces)
export LOCATION="eastus"
export RESOURCE_GROUP="${PREFIX}-rg"
export CLUSTER_NAME="${PREFIX}-aks"

# Network
export VNET_CIDR="10.0.0.0/16"
export AKS_SUBNET_CIDR="10.0.0.0/17"     # first half of the VNet

# Cluster
export K8S_VERSION="1.33"

# Workspace (also used as the Kubernetes namespace)
export WORKSPACE_NAME="my-workspace"

# Subscription context
export SUBSCRIPTION_ID=$(az account show --query id -o tsv)
```

### Step 1: Create the resource group, VNet, and NAT Gateway

```bash
az group create --name $RESOURCE_GROUP --location $LOCATION

az network vnet create \
  --resource-group $RESOURCE_GROUP \
  --name ${PREFIX}-vnet \
  --address-prefix $VNET_CIDR \
  --location $LOCATION

# AKS subnet (nodes + overlay pods). Service endpoints keep Storage/ACR traffic on the Azure backbone.
az network vnet subnet create \
  --resource-group $RESOURCE_GROUP \
  --vnet-name ${PREFIX}-vnet \
  --name ${PREFIX}-aks \
  --address-prefixes $AKS_SUBNET_CIDR \
  --service-endpoints Microsoft.Storage Microsoft.ContainerRegistry

# NAT Gateway for stable egress (static public IP, 30-minute idle timeout)
az network public-ip create \
  --resource-group $RESOURCE_GROUP \
  --name ${PREFIX}-nat-pip \
  --sku Standard \
  --allocation-method Static \
  --location $LOCATION

az network nat gateway create \
  --resource-group $RESOURCE_GROUP \
  --name ${PREFIX}-nat \
  --public-ip-addresses ${PREFIX}-nat-pip \
  --idle-timeout 30 \
  --location $LOCATION

az network vnet subnet update \
  --resource-group $RESOURCE_GROUP \
  --vnet-name ${PREFIX}-vnet \
  --name ${PREFIX}-aks \
  --nat-gateway ${PREFIX}-nat

# Capture IDs for cluster creation
export AKS_SUBNET_ID=$(az network vnet subnet show \
  --resource-group $RESOURCE_GROUP --vnet-name ${PREFIX}-vnet --name ${PREFIX}-aks \
  --query id -o tsv)

export VNET_ID=$(az network vnet show \
  --resource-group $RESOURCE_GROUP --name ${PREFIX}-vnet --query id -o tsv)

echo "AKS_SUBNET_ID=$AKS_SUBNET_ID"
```

### Step 2: Create the AKS cluster

Create the cluster with OIDC issuer and Workload Identity enabled, Azure CNI overlay with Cilium, and an ARM64 system pool tainted `CriticalAddonsOnly`:

```bash
az aks create \
  --resource-group $RESOURCE_GROUP \
  --name $CLUSTER_NAME \
  --location $LOCATION \
  --kubernetes-version $K8S_VERSION \
  --tier standard \
  --node-resource-group ${CLUSTER_NAME}-nodes \
  --enable-managed-identity \
  --enable-oidc-issuer \
  --enable-workload-identity \
  --network-plugin azure \
  --network-plugin-mode overlay \
  --network-policy cilium \
  --network-dataplane cilium \
  --pod-cidr 10.244.0.0/16 \
  --service-cidr 172.20.0.0/16 \
  --dns-service-ip 172.20.0.10 \
  --outbound-type loadBalancer \
  --load-balancer-sku standard \
  --vnet-subnet-id $AKS_SUBNET_ID \
  --nodepool-name system \
  --node-vm-size Standard_D2ps_v6 \
  --node-count 2 \
  --enable-cluster-autoscaler \
  --min-count 2 \
  --max-count 4 \
  --node-osdisk-size 50 \
  --max-pods 110 \
  --nodepool-taints CriticalAddonsOnly=true:NoSchedule \
  --generate-ssh-keys
```

Notes:

* `Standard_D2ps_v6` is ARM64. All e6data system images are multi-arch; for x86 use `Standard_D4as_v5`.
* The NAT Gateway on the subnet takes precedence for egress even with `--outbound-type loadBalancer`.
* To use Entra ID (Azure AD) RBAC for cluster auth, append `--enable-aad --enable-azure-rbac --disable-local-accounts --aad-admin-group-object-ids <group-object-id>`.

### Step 3: Connect and capture cluster outputs

```bash
az aks get-credentials --resource-group $RESOURCE_GROUP --name $CLUSTER_NAME

# OIDC issuer URL - needed for every federated identity credential below
export AKS_OIDC_ISSUER=$(az aks show \
  --resource-group $RESOURCE_GROUP --name $CLUSTER_NAME \
  --query "oidcIssuerProfile.issuerUrl" -o tsv)

# Node resource group (where AKS puts VMs/VMSS/LBs)
export NODE_RESOURCE_GROUP=$(az aks show \
  --resource-group $RESOURCE_GROUP --name $CLUSTER_NAME \
  --query nodeResourceGroup -o tsv)

echo "AKS_OIDC_ISSUER=$AKS_OIDC_ISSUER"
kubectl get nodes
```

AKS hosts the OIDC issuer for you - there is nothing to create. You only reference its URL when creating federated credentials.

### Step 4: Grant the cluster identity network access

The cluster's control-plane identity needs **Network Contributor** on the VNet so the Azure cloud provider can manage load balancer backend pools for nodes in your subnet:

```bash
CLUSTER_IDENTITY_PRINCIPAL_ID=$(az aks show \
  --resource-group $RESOURCE_GROUP --name $CLUSTER_NAME \
  --query identity.principalId -o tsv)

az role assignment create \
  --assignee-object-id $CLUSTER_IDENTITY_PRINCIPAL_ID \
  --assignee-principal-type ServicePrincipal \
  --role "Network Contributor" \
  --scope $VNET_ID
```

### Step 5: Container registry access

e6data images are pulled from `e6labs.azurecr.io`. Choose one option and confirm it with your onboarding engineer (they configure the matching side).

**Option A - AcrPull role** (if e6data grants your kubelet identity access, or you mirror images to your own ACR):

```bash
KUBELET_OBJECT_ID=$(az aks show \
  --resource-group $RESOURCE_GROUP --name $CLUSTER_NAME \
  --query identityProfile.kubeletidentity.objectId -o tsv)

# Scope = your mirrored ACR resource ID, or the e6data-shared ACR ID from onboarding
az role assignment create \
  --assignee-object-id $KUBELET_OBJECT_ID \
  --assignee-principal-type ServicePrincipal \
  --role AcrPull \
  --scope "<acr-resource-id>"
```

**Option B - Pull secret** (e6data provides a scoped ACR token). The secret must exist in **both** the `e6operator` namespace and the workspace namespace before components install, so create the namespaces now:

```bash
for NS in e6operator ${WORKSPACE_NAME}; do
  kubectl create namespace ${NS} --dry-run=client -o yaml | kubectl apply -f -
  kubectl create secret docker-registry e6data-acr-pull \
    --namespace ${NS} \
    --docker-server=e6labs.azurecr.io \
    --docker-username="<token-name-from-e6data>" \
    --docker-password="<token-password-from-e6data>"
done
```

With Option B, also reference `e6data-acr-pull` wherever the guide notes `imagePullSecrets` (operator values in Step 2.4, NamespaceConfig in the next page).

### Step 6: Engine node pool

Engine pods need a user node pool to scale on. Pick one provisioner. The `workspace-name` label and taint are mandatory - engine pods schedule onto nodes by exactly this label/taint pair.

**Option A - Static user node pool (simplest):**

```bash
az aks nodepool add \
  --resource-group $RESOURCE_GROUP \
  --cluster-name $CLUSTER_NAME \
  --name engine \
  --mode User \
  --node-vm-size Standard_D8as_v5 \
  --enable-cluster-autoscaler \
  --min-count 0 --max-count 10 \
  --node-osdisk-size 100 \
  --zones 1 2 3 \
  --labels workspace-name=${WORKSPACE_NAME} \
  --node-taints workspace-name=${WORKSPACE_NAME}:NoSchedule
```

**Option B - Node Auto Provisioning (NAP, managed Karpenter):**

```bash
az aks update --resource-group $RESOURCE_GROUP --name $CLUSTER_NAME \
  --node-provisioning-mode Auto
```

**Option C - Self-hosted Karpenter** (recommended for scale-to-zero engine nodes). This uses the AKS Karpenter provider published in Microsoft's container registry. It has five sub-steps: identity, cluster values, bootstrap token, Helm install, and the NodeClass/NodePool.

#### 6C.1 - Karpenter Managed Identity and roles

```bash
az identity create \
  --resource-group $RESOURCE_GROUP \
  --name ${CLUSTER_NAME}-karpenter-identity \
  --location $LOCATION

KARPENTER_PRINCIPAL_ID=$(az identity show -g $RESOURCE_GROUP \
  --name ${CLUSTER_NAME}-karpenter-identity --query principalId -o tsv)
export KARPENTER_IDENTITY_CLIENT_ID=$(az identity show -g $RESOURCE_GROUP \
  --name ${CLUSTER_NAME}-karpenter-identity --query clientId -o tsv)

# Federated credential -> kube-system/karpenter service account
az identity federated-credential create \
  --name ${CLUSTER_NAME}-karpenter-federated \
  --identity-name ${CLUSTER_NAME}-karpenter-identity \
  --resource-group $RESOURCE_GROUP \
  --issuer $AKS_OIDC_ISSUER \
  --subject "system:serviceaccount:kube-system:karpenter" \
  --audiences "api://AzureADTokenExchange"

# Role assignments
NODE_RG_ID="/subscriptions/${SUBSCRIPTION_ID}/resourceGroups/${NODE_RESOURCE_GROUP}"
RG_ID="/subscriptions/${SUBSCRIPTION_ID}/resourceGroups/${RESOURCE_GROUP}"

az role assignment create --assignee-object-id $KARPENTER_PRINCIPAL_ID \
  --assignee-principal-type ServicePrincipal --role "Virtual Machine Contributor" --scope $NODE_RG_ID
az role assignment create --assignee-object-id $KARPENTER_PRINCIPAL_ID \
  --assignee-principal-type ServicePrincipal --role "Network Contributor" --scope $VNET_ID
az role assignment create --assignee-object-id $KARPENTER_PRINCIPAL_ID \
  --assignee-principal-type ServicePrincipal --role "Managed Identity Operator" --scope $NODE_RG_ID
az role assignment create --assignee-object-id $KARPENTER_PRINCIPAL_ID \
  --assignee-principal-type ServicePrincipal --role "Reader" --scope $RG_ID
```

#### 6C.2 - Gather cluster values

```bash
export KARPENTER_VERSION="1.7.1"
export CLUSTER_ENDPOINT="https://$(az aks show -g $RESOURCE_GROUP -n $CLUSTER_NAME --query fqdn -o tsv):443"
export NODE_IDENTITIES=$(az aks show -g $RESOURCE_GROUP -n $CLUSTER_NAME \
  --query identityProfile.kubeletidentity.resourceId -o tsv)
export VNET_GUID=$(az network vnet show -g $RESOURCE_GROUP -n ${PREFIX}-vnet --query resourceGuid -o tsv)
export SSH_PUBLIC_KEY="$(cat ~/.ssh/id_rsa.pub)"
```

#### 6C.3 - Bootstrap token

New nodes use this token to join the cluster. It must be in `<6-char-id>.<16-char-secret>` format:

```bash
TOKEN_ID=$(LC_ALL=C tr -dc 'a-z0-9' < /dev/urandom | head -c 6)
TOKEN_SECRET=$(LC_ALL=C tr -dc 'a-z0-9' < /dev/urandom | head -c 16)
export BOOTSTRAP_TOKEN="${TOKEN_ID}.${TOKEN_SECRET}"

kubectl create secret generic "bootstrap-token-${TOKEN_ID}" \
  --namespace kube-system \
  --type=bootstrap.kubernetes.io/token \
  --from-literal=token-id="$TOKEN_ID" \
  --from-literal=token-secret="$TOKEN_SECRET" \
  --from-literal=usage-bootstrap-authentication=true \
  --from-literal=usage-bootstrap-signing=true
```

#### 6C.4 - Install Karpenter

Create `karpenter-values.yaml`, substituting the variables set above:

```yaml
replicas: 2
controller:
  env:
    - { name: LEADER_ELECT, value: "true" }
    - { name: CLUSTER_NAME, value: "${CLUSTER_NAME}" }
    - { name: CLUSTER_ENDPOINT, value: "${CLUSTER_ENDPOINT}" }
    - { name: KUBELET_BOOTSTRAP_TOKEN, value: "${BOOTSTRAP_TOKEN}" }
    - { name: NETWORK_PLUGIN, value: "azure" }
    - { name: NETWORK_POLICY, value: "cilium" }
    - { name: NODE_IDENTITIES, value: "${NODE_IDENTITIES}" }
    - { name: ARM_SUBSCRIPTION_ID, value: "${SUBSCRIPTION_ID}" }
    - { name: LOCATION, value: "${LOCATION}" }
    - { name: VNET_SUBNET_ID, value: "${AKS_SUBNET_ID}" }
    - { name: SSH_PUBLIC_KEY, value: "${SSH_PUBLIC_KEY}" }
    - { name: VNET_GUID, value: "${VNET_GUID}" }
    - { name: ARM_USE_CREDENTIAL_FROM_ENVIRONMENT, value: "true" }
    - { name: ARM_USE_MANAGED_IDENTITY_EXTENSION, value: "false" }
    - { name: ARM_USER_ASSIGNED_IDENTITY_ID, value: "" }
    - { name: AZURE_NODE_RESOURCE_GROUP, value: "${NODE_RESOURCE_GROUP}" }
  resources:
    requests: { cpu: 200m, memory: 256Mi }
    limits: { cpu: "1", memory: 1Gi }
serviceAccount:
  name: karpenter
  annotations:
    azure.workload.identity/client-id: "${KARPENTER_IDENTITY_CLIENT_ID}"
podLabels:
  azure.workload.identity/use: "true"
tolerations:
  - { key: CriticalAddonsOnly, operator: Exists, effect: NoSchedule }
logLevel: info
```

```bash
envsubst < karpenter-values.yaml | helm upgrade --install karpenter \
  oci://mcr.microsoft.com/aks/karpenter/karpenter \
  --version "${KARPENTER_VERSION}" \
  --namespace kube-system --wait --timeout 5m -f -

kubectl rollout status deployment/karpenter -n kube-system --timeout=120s
```

#### 6C.5 - AKSNodeClass and NodePool for engine nodes

The taint `workspace-name=<workspace>` is what the e6data engine pods tolerate, so nodes are dedicated to the workspace:

```bash
cat << EOF | kubectl apply -f -
apiVersion: karpenter.azure.com/v1alpha2
kind: AKSNodeClass
metadata:
  name: ${WORKSPACE_NAME}-nodeclass
  labels:
    app: e6data
    e6data-workspace-name: ${WORKSPACE_NAME}
spec:
  imageFamily: Ubuntu2204
  osDiskSizeGB: 100
  tags:
    Name: ${WORKSPACE_NAME}
    app: e6data
    namespace: ${WORKSPACE_NAME}
    type: internal-compute
    ManagedBy: karpenter
EOF

cat << EOF | kubectl apply -f -
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: ${WORKSPACE_NAME}-nodepool
  labels:
    workspace-name: ${WORKSPACE_NAME}
    type: ${WORKSPACE_NAME}
spec:
  disruption:
    budgets:
      - nodes: 100%
        reasons: [Empty]
      - nodes: "0"
        reasons: [Drifted]
    consolidateAfter: 30s
    consolidationPolicy: WhenEmpty
  limits:
    cpu: "5000"
  template:
    metadata:
      labels:
        workspace-name: ${WORKSPACE_NAME}
        type: ${WORKSPACE_NAME}
    spec:
      expireAfter: 720h
      nodeClassRef:
        group: karpenter.azure.com
        kind: AKSNodeClass
        name: ${WORKSPACE_NAME}-nodeclass
      requirements:
        - key: karpenter.azure.com/sku-family
          operator: In
          values: [D, E, F, L]
        - key: karpenter.sh/capacity-type
          operator: In
          values: [on-demand, spot]
        - key: kubernetes.io/arch
          operator: In
          values: [arm64]
      taints:
        - { effect: NoSchedule, key: workspace-name, value: ${WORKSPACE_NAME} }
EOF

kubectl get aksnodeclasses.karpenter.azure.com
kubectl get nodepools.karpenter.sh
```

Notes:

* `kubernetes.io/arch: arm64` with sku-families `D/E/F/L` makes Karpenter pick ARM variants. For x86, change arch to `amd64`.
* `capacity-type` includes `spot`; remove it for on-demand only.
* A `Warning: use v1beta1.AKSNodeClass instead` message is a Karpenter 1.7.x deprecation notice - `v1alpha2` still works and matches the e6data production templates. Safe to ignore.

### Step 7: Create the metadata storage (ADLS Gen2)

e6data stores workspace metadata in an ADLS Gen2 account (hierarchical namespace enabled) with one private container per workspace:

```bash
# Storage account names: lowercase alphanumeric only, max 24 chars
export STORAGE_ACCOUNT=$(echo "${PREFIX}workspaces" | tr -d '-' | tr '[:upper:]' '[:lower:]' | cut -c1-24)

az storage account create \
  --name $STORAGE_ACCOUNT \
  --resource-group $RESOURCE_GROUP \
  --location $LOCATION \
  --kind StorageV2 --sku Standard_LRS \
  --min-tls-version TLS1_2 --https-only true \
  --allow-blob-public-access false \
  --enable-hierarchical-namespace true

# Metadata container. --auth-mode login requires your user to hold "Storage Blob Data
# Contributor" on the account (Owner alone is not enough for data-plane operations).
az storage container create \
  --account-name $STORAGE_ACCOUNT \
  --name ${WORKSPACE_NAME}-metadata \
  --auth-mode login

export STORAGE_ACCOUNT_ID=$(az storage account show \
  --name $STORAGE_ACCOUNT --resource-group $RESOURCE_GROUP --query id -o tsv)
```

{% hint style="warning" %}
Do **not** enable blob versioning or soft delete on this account - both break ABFS/Hadoop compatibility, which the engine relies on.
{% endhint %}

The storage backend URI used later in the workspace config is `abfss://${WORKSPACE_NAME}-metadata@${STORAGE_ACCOUNT}.dfs.core.windows.net/`.

### Step 8: Create the workspace Managed Identity and federated credentials

All e6data workspace components share **one** user-assigned Managed Identity, with one federated credential per e6data service account that needs Azure access. The core components are **engine** and **console**:

```bash
az identity create \
  --resource-group $RESOURCE_GROUP \
  --name ${PREFIX}-${WORKSPACE_NAME}-engine-identity \
  --location $LOCATION

export ENGINE_IDENTITY_CLIENT_ID=$(az identity show -g $RESOURCE_GROUP \
  --name ${PREFIX}-${WORKSPACE_NAME}-engine-identity --query clientId -o tsv)
export ENGINE_IDENTITY_PRINCIPAL_ID=$(az identity show -g $RESOURCE_GROUP \
  --name ${PREFIX}-${WORKSPACE_NAME}-engine-identity --query principalId -o tsv)

# One federated credential per core service account
for SA in engine console; do
  az identity federated-credential create \
    --name ${PREFIX}-${WORKSPACE_NAME}-${SA}-federated \
    --identity-name ${PREFIX}-${WORKSPACE_NAME}-engine-identity \
    --resource-group $RESOURCE_GROUP \
    --issuer $AKS_OIDC_ISSUER \
    --subject "system:serviceaccount:${WORKSPACE_NAME}:${WORKSPACE_NAME}-${SA}" \
    --audiences "api://AzureADTokenExchange"
done

echo "ENGINE_IDENTITY_CLIENT_ID=$ENGINE_IDENTITY_CLIENT_ID"
```

{% hint style="info" %}
When you later enable add-ons (Laminar, Copilot, Metriq), add a federated credential for each of their service accounts against this same identity. The monitoring service account intentionally has no identity binding - it pushes to GreptimeDB with username and password.
{% endhint %}

### Step 9: Role assignments for storage access

```bash
# Read/write on the metadata storage account
az role assignment create \
  --assignee-object-id $ENGINE_IDENTITY_PRINCIPAL_ID \
  --assignee-principal-type ServicePrincipal \
  --role "Storage Blob Data Contributor" \
  --scope $STORAGE_ACCOUNT_ID

# Read-only on each storage account that holds your data lake (repeat per account)
az role assignment create \
  --assignee-object-id $ENGINE_IDENTITY_PRINCIPAL_ID \
  --assignee-principal-type ServicePrincipal \
  --role "Storage Blob Data Reader" \
  --scope "/subscriptions/${SUBSCRIPTION_ID}/resourceGroups/<data-rg>/providers/Microsoft.Storage/storageAccounts/<data-storage-account>"
```

### Part 1 verification

```bash
az aks show -g $RESOURCE_GROUP -n $CLUSTER_NAME \
  --query "{k8s:kubernetesVersion, oidc:oidcIssuerProfile.enabled, wi:securityProfile.workloadIdentity.enabled, network:networkProfile.networkPlugin}" -o table

kubectl get nodes -o wide

az identity federated-credential list \
  --identity-name ${PREFIX}-${WORKSPACE_NAME}-engine-identity \
  --resource-group $RESOURCE_GROUP \
  --query "[].{name:name, subject:subject}" -o table
```

Record `AKS_OIDC_ISSUER`, `NODE_RESOURCE_GROUP`, `ENGINE_IDENTITY_CLIENT_ID`, `STORAGE_ACCOUNT`, and the storage backend URI - the workspace deployment uses them.

## Part 2 - Kubernetes platform components

This part installs the cluster-level components: an `e6operator` NodePool, cert-manager, the e6-operator CRDs, the e6-operator controller, and the service-token verifier secret. Karpenter itself was installed in Part 1 Step 6C. e6data provides the operator chart and CRDs during onboarding.

{% hint style="info" %}
The recommended layout pins cert-manager and the operator onto dedicated `e6operator` nodes (label `app: e6operator`, taint `workload=e6operator`). Step 2.1 creates that NodePool if you run Karpenter. If you use static or NAP node pools instead, omit the `nodeSelector`/`tolerations` blocks in Steps 2.2 and 2.4 and let the components run on the system pool.
{% endhint %}

### Step 2.1: e6-operator NodePool (Karpenter only)

```bash
cat << 'EOF' | kubectl apply -f -
apiVersion: karpenter.azure.com/v1alpha2
kind: AKSNodeClass
metadata:
  name: e6operator
spec:
  imageFamily: Ubuntu2204
  osDiskSizeGB: 50
  tags:
    ManagedBy: karpenter
    Name: e6operator
    app: e6data
    workload: e6operator
---
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: e6operator
spec:
  disruption:
    consolidationPolicy: WhenEmpty
    consolidateAfter: 30s
  limits:
    cpu: "100"
    memory: 100Gi
  template:
    metadata:
      labels:
        workload: e6operator
        app: e6operator
    spec:
      expireAfter: Never
      nodeClassRef:
        group: karpenter.azure.com
        kind: AKSNodeClass
        name: e6operator
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: [on-demand]
        - key: kubernetes.io/arch
          operator: In
          values: [arm64]
        - key: karpenter.azure.com/sku-family
          operator: In
          values: [D, E]
      taints:
        - { effect: NoSchedule, key: workload, value: e6operator }
EOF
```

### Step 2.2: Install cert-manager

The operator's admission webhooks get their serving certificates from cert-manager, so it is required:

```bash
export CERT_MANAGER_VERSION="v1.17.2"

cat > cert-manager-values.yaml << 'EOF'
crds:
  enabled: true
prometheus:
  enabled: false
webhook:
  timeoutSeconds: 4
  tolerations:
    - { key: workload, operator: Equal, value: e6operator, effect: NoSchedule }
  nodeSelector:
    app: e6operator
tolerations:
  - { key: workload, operator: Equal, value: e6operator, effect: NoSchedule }
nodeSelector:
  app: e6operator
cainjector:
  tolerations:
    - { key: workload, operator: Equal, value: e6operator, effect: NoSchedule }
  nodeSelector:
    app: e6operator
startupapicheck:
  tolerations:
    - { key: workload, operator: Equal, value: e6operator, effect: NoSchedule }
  nodeSelector:
    app: e6operator
EOF

helm upgrade --install cert-manager oci://quay.io/jetstack/charts/cert-manager \
  --version "${CERT_MANAGER_VERSION}" \
  --namespace cert-manager --create-namespace --wait \
  -f cert-manager-values.yaml

kubectl get pods -n cert-manager
```

### Step 2.3: Install the e6-operator CRDs

The CRDs are applied with `kubectl`, **not** Helm, because the combined payload exceeds Helm's 1 MB release-secret limit. `--server-side` is required because client-side apply hits the 256 KB annotation limit:

```bash
# Extract the chart provided by e6data
tar -xzf e6-operator-crds.tgz

kubectl apply --server-side --force-conflicts -f ./e6-operator-crds/crds/

# Expect ~25 CRDs in the e6data.io group
kubectl get crds | grep e6data.io
```

Key CRDs used in workspace deployment: `namespaceconfigs`, `queryrouters`, `monitoringservices`, `metadataservices`, `queryservices`, `accesstokens`, `e6catalogs`, `governances`.

### Step 2.4: Install the e6-operator controller

```bash
# Extract the chart; the helm path must point to the folder containing Chart.yaml
tar -xzf e6-operator.tgz

export E6_OPERATOR_IMAGE_TAG="<VERSION_PROVIDED_BY_E6DATA>"

cat > operator-values.yaml << EOF
replicaCount: 2
image:
  repository: e6labs.azurecr.io/e6data/e6-operator
  tag: "${E6_OPERATOR_IMAGE_TAG}"
  pullPolicy: Always
# Option B (pull secret): uncomment the next two lines.
# imagePullSecrets:
#   - name: e6data-acr-pull
controllerManager:
  logLevel: info
  leaderElection:
    enabled: true
env:
  - name: LOG_MODE
    value: "production"
# The chart key is `webhook` (singular)
webhook:
  enabled: true
  certManager:
    enabled: true
resources:
  requests: { cpu: 200m, memory: 256Mi }
  limits: { cpu: "1", memory: 1Gi }
priorityClassName: system-cluster-critical
tolerations:
  - { key: workload, operator: Equal, value: e6operator, effect: NoSchedule }
nodeSelector:
  app: e6operator
karpenter:
  enabled: true        # set false if using static pools or NAP
# clusterMonitoring activates only when greptime.endpoint is set. In-VPC tenants run
# their own cluster monitoring, so leave it unset (stays off).
EOF

helm upgrade --install e6operator ./e6-operator \
  --namespace e6operator --create-namespace --wait --timeout 8m \
  -f operator-values.yaml

kubectl rollout status deployment/e6operator -n e6operator --timeout=300s
kubectl get pods -n e6operator
```

### Step 2.5: Create the service-token verifier secret

The console authenticates `e6_svc_` service tokens against the e6data control plane using a shared verifier secret. The e6-operator reflects this secret into every workspace namespace as `<workspace>-cp-verifier`.

{% hint style="warning" %}
If this secret is missing, service-token authentication silently fails with `401`. Create it now with the value from your onboarding engineer.
{% endhint %}

```bash
export CP_VERIFIER_SECRET_VALUE="<value-provided-by-e6data>"

kubectl create secret generic cp-verifier-shared \
  --namespace e6operator \
  --from-file=SERVICE_CREDENTIAL_VERIFIER_SECRET=<(printf '%s' "$CP_VERIFIER_SECRET_VALUE") \
  --dry-run=client -o yaml | kubectl apply -f -
```

## Next

Continue to [Deploy workspace and e6data](/query-engine/guides/deployment/azure-in-vpc/deploy-workspace-and-e6data.md).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.e6data.com/query-engine/guides/deployment/azure-in-vpc/configure-registry-kubernetes-networking.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
