> ## Documentation Index
> Fetch the complete documentation index at: https://support.lilt.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Install System (AWS EKS)

This article describes how to install LILT on Amazon Elastic Kubernetes Service from the EKS install package. It covers configuration, the phased install, secrets, verification, and recovery.

## Overview

An EKS deployment has two layers:

1. **The AWS infrastructure**: a VPC and subnets, the EKS cluster and its node groups, an RDS MySQL instance, an S3 bucket, an SQS queue, a container registry, and the IAM roles the workloads assume.

2. **The LILT application stack**, deployed by `install-lilt-eks.sh`: Istio and the Gateway API routes, the data stores, and the LILT Helm charts.

Choose one of two paths for the first layer:

| Path | Use it when | Guide |
| - | - | - |
| Terraform module | You want LILT to provision the infrastructure. The `lilt-aws-env` module ships in the install package. | [Provision AWS infrastructure with Terraform](/self-managed/v6.1/install-aws-terraform-module) |
| Bring your own | You already have a cluster, or you provision AWS resources with your own tooling. | [Install on an existing EKS cluster](/self-managed/v6.1/install-aws-existing-cluster) |

Both paths converge here for the second layer.

<Note>
  For a bare-metal or virtual-machine deployment, use the on-premises install package instead. See [Install System (Amazon Linux 2023 or Rocky 8/9)](/self-managed/v6.1/install-system-amazon-linux-2023-or-rocky-8-9).
</Note>

## Prerequisites

1. An EKS cluster, version 1.32 or later, with the EBS CSI driver installed. Version 1.35 is the tested default, and the LILT Terraform module creates 1.35.

2. Helm 4, plus `kubectl`, `aws`, and `jq` on your PATH.

   <Warning>
     Check your Helm version before you start.

     ```bash theme={null}
     helm version --short   # must report v4.x
     ```

     Several component scripts use flags that exist only in Helm 4, among them `--rollback-on-failure` for ArgoCD and Argo Rollouts and `--server-side` for the Elasticsearch cluster and the LILT charts. On releases whose `install-lilt-eks.sh` does not check the version at startup, Helm 3 gets through the prerequisites and the data stores and then fails an hour in, reporting an unrecognised flag rather than the Helm version.
   </Warning>

   Add [crane](https://github.com/google/go-containerregistry) if you populate your registry with `seed-registry.sh`. The installer does not install it, and the script stops with `ERROR: crane not on PATH.` Skip it only if you mirror the release images by some other means.

3. An RDS MySQL instance, an S3 bucket, and an SQS queue that the cluster can reach.

4. A container registry the cluster can pull from, seeded with the release images. See [Seed your container registry](/self-managed/v6.1/install-package#seed-your-container-registry).

5. IAM roles for the workloads that need AWS access, bound with EKS Pod Identity or IRSA.

6. The extracted install package. See [Install package](/self-managed/v6.1/install-package).

<Note>
  Run the installer from a machine that can reach the Kubernetes API. Clusters created by the LILT Terraform module have a private API endpoint, so the install runs from inside the VPC. See [Install from a bastion](/self-managed/v6.1/install-from-a-bastion).
</Note>

## Node Kernel Settings

The EKS-optimized Amazon Linux 2023 AMI already carries the kernel modules, `sysctl` values and swap and SELinux settings that Kubernetes needs, so a cluster on the stock AMI needs nothing here.

Two settings are not part of that baseline and matter to LILT specifically. Check them on your nodes whichever AMI you use:

```bash theme={null}
# Istio's ztunnel restarts under load without a raised file-descriptor limit.
ulimit -n                        # want 131072

# Elasticsearch fails to start below this.
sysctl vm.max_map_count          # want 262144
```

To set them, on each node:

```bash theme={null}
cat <<EOF >> /etc/security/limits.conf
soft nofile 131072
hard nofile 131072
EOF

echo "DefaultLimitNOFILE=131072" >> /etc/systemd/system.conf

echo "vm.max_map_count=262144" | sudo tee -a /etc/sysctl.conf
sysctl --system
```

<Accordion title="Custom or non-EKS-optimized AMI: the full baseline">
  If you build your own AMI, or run a node OS that is not the EKS-optimized image, the Kubernetes baseline is yours to provide as well. Bake this into the image rather than applying it to running nodes, so replacements and scale-ups inherit it.

  ```bash theme={null}
  # Kernel modules for kube-proxy and the Istio redirect rules.
  cat <<EOF | sudo tee /etc/modules-load.d/k8s.conf
  overlay
  br_netfilter
  nf_nat
  xt_REDIRECT
  xt_owner
  iptable_nat
  iptable_mangle
  iptable_filter
  EOF

  for m in overlay br_netfilter nf_nat xt_REDIRECT xt_owner \
           iptable_nat iptable_mangle iptable_filter; do
    modprobe "$m"
  done

  cat <<EOF | sudo tee /etc/sysctl.d/k8s.conf
  net.bridge.bridge-nf-call-iptables  = 1
  net.bridge.bridge-nf-call-ip6tables = 1
  net.ipv4.ip_forward                 = 1
  EOF

  sysctl --system

  # The kubelet does not start with swap enabled.
  swapoff -a
  sed -e '/swap/s/^/#/g' -i /etc/fstab

  # SELinux, on distributions that enforce it.
  sudo setenforce 0
  sudo sed -i 's/^SELINUX=enforcing$/SELINUX=permissive/' /etc/selinux/config
  ```

  An unset `xt_REDIRECT` or `xt_owner` is worth calling out: the mesh's traffic redirection depends on them, and without them pods start and then fail to reach each other, which reads as a networking or DNS fault rather than a missing kernel module.
</Accordion>

## Configure the Install

Copy the template and edit it. Every variable the installer reads is documented there with its default.

```bash theme={null}
cp install.env.example install.env
chmod 0600 install.env
```

The file is organised by intent: identity, pre-existing resources, DNS and TLS, image registry, feature toggles, and the secret backend. Leave a line commented to keep the default shown next to it. Every variable is listed in the [install.env reference](/self-managed/v6.1/install-env-reference).

The settings almost every install needs:

| Variable | Purpose |
| - | - |
| `SUBDOMAIN`, `DNS_DOMAIN` | Every ingress hostname is `<name>-<SUBDOMAIN>.<DNS_DOMAIN>`. Both are required. |
| `AWS_REGION`, `AWS_ACCOUNT_ID` | Resource ARNs and registry hosts. The account ID defaults to your current caller identity. |
| `REGISTRY_BASE` | The registry to pull images from. |
| `S3_BUCKET`, `SQS_QUEUE_URL`, `DB_NAME` | The pre-existing resources to use. |
| `CONNECTORS_ADMIN_PASSWORD` | The `admin` login for the `/connectors-admin` console. The install stops with an error if this is unset, so choose one before you start. |
| `SECRET_BACKEND` | `seed`, `vault`, or `aws`. Defaults to `seed`. |
| `IAM_AUTH_METHOD` | `irsa` or `pod_identity`. Defaults to `irsa`. It must match how your IAM roles are actually bound, or every AWS call from a pod fails. |
| `PREFIX` | Optional. Set it to the Terraform module's `prefix` and the installer derives the resource names for you. |

## How the Installer Works

`install-lilt-eks.sh` is a flat sequence of component scripts. Each one:

* Prints `--- Installing <Component> ---` before it runs, so the last such line names the component that failed.

* Checks its own `ENABLE_<NAME>` variable and exits without doing anything when that variable is not `true`.

* Uses `helm upgrade --install`, which is a no-op when the chart and its values are unchanged.

Together these three properties make the whole script safe to re-run from the top. That is the intended recovery path.

### Component Flags

Set any of these to `false` in `install.env` to skip that component.

**Optional components, enabled by default:**

`ENABLE_CLUSTER_AUTOSCALER`, `ENABLE_MONGODB`, `ENABLE_CLICKHOUSE`, `ENABLE_DRAGONFLY_OPERATOR`, `ENABLE_CACHE_DRAGONFLY`, `ENABLE_RABBITMQ`, `ENABLE_ECK_OPERATOR`, `ENABLE_ECK_CLUSTER`, `ENABLE_ARGOCD`, `ENABLE_ARGO_ROLLOUTS`, `ENABLE_ARGO_WORKFLOWS`, `ENABLE_PROMETHEUS`, `ENABLE_METRICS_SERVER`, `ENABLE_ISTIO_INGRESSGATEWAY`, `ENABLE_QDRANT`, `ENABLE_PROXY_SQL`, `ENABLE_WSO2`, `ENABLE_MYSQL_CLIENT`, `ENABLE_RAY_OPERATOR_CRDS`, `ENABLE_RAY_OPERATOR`, `ENABLE_RAYMAN`, `ENABLE_LILT_CHARTS`, `ENABLE_EXTERNAL_SECRETS`, `ENABLE_VAULT`, `ENABLE_NVIDIA_GPU_OPERATOR`.

**Optional components, disabled by default:**

| Flag | Component |
| - | - |
| `ENABLE_DNS` | Route53 record creation |
| `ENABLE_CERT_MANAGER` | cert-manager with Route53 DNS-01 |
| `ENABLE_SEED_REGISTRY` | Copying the release images into your registry |

<Note>
  `install.env.example` also lists `ENABLE_DATADOG` and `ENABLE_OTELCOL`. Neither entrypoint calls an installer for them, and neither component ships in the tarball, so setting these has no effect on a customer install.
</Note>

**Prerequisites, enabled by default:**

`ENABLE_NAMESPACE`, `ENABLE_CLEANUP_STALE_PODS`, `ENABLE_CLUSTER_PREREQS`, `ENABLE_RDS_SECRETS`, `ENABLE_S3_IRSA`, `ENABLE_RDS_LILT_USER`, `ENABLE_DATABASE_BOOTSTRAP`, `ENABLE_PREINSTALL_ISTIO_CRDS`, `ENABLE_GATEWAY_API_CRDS`, `ENABLE_ISTIOD`, `ENABLE_LILT_NETWORKING`.

<Warning>
  Prerequisites are foundational, and everything after them depends on them. Disable one only on a cluster where that step has already completed.
</Warning>

Dependency order matters among the optional components too. The ECK operator has to exist before its cluster, the Ray operator before Rayman, and ProxySQL fronts the database.

### Phased Bring-Up

Bringing the infrastructure and data stores up first, and the application afterwards, makes a failure much easier to attribute.

**Phase 1: infrastructure and data stores.** Add these lines to `install.env`:

```bash theme={null}
ENABLE_ARGOCD=false
ENABLE_ARGO_ROLLOUTS=false
ENABLE_ARGO_WORKFLOWS=false
ENABLE_PROMETHEUS=false
ENABLE_QDRANT=false
ENABLE_PROXY_SQL=false
ENABLE_WSO2=false
ENABLE_MYSQL_CLIENT=false
ENABLE_RAY_OPERATOR_CRDS=false
ENABLE_RAY_OPERATOR=false
ENABLE_RAYMAN=false
ENABLE_LILT_CHARTS=false
```

Run the installer, then confirm that the data stores are healthy:

```bash theme={null}
sh install-lilt-eks.sh
kubectl get pods -n lilt
```

**Phase 2: the full stack.** Remove those lines so each component reverts to its default, and run the installer again.

## Secrets

`SECRET_BACKEND` selects where the installer reads application secrets from: `seed` (the default, from the file in the tarball), `vault`, or `aws`. LILT also ships an in-cluster HashiCorp Vault as the default External Secrets Operator backend.

<Warning>
  The seed file holds placeholder values, most of them the literal string `dummypass`. Do not run a production install with them. On EKS the file is overwritten from the seed on every run while `SECRET_BACKEND=seed`, so edits there do not persist.
</Warning>

See [Secrets and Vault](/self-managed/v6.1/secrets-and-vault) for the backend comparison, the full Vault path and key table, the per-secret path overrides, `seed-vault-secrets.sh`, and how to point LILT at a Vault or AWS Secrets Manager you already run.

## Override Values

Every `.yaml` file under `lilt/environments/lilt/values.d/` is applied after the environment overlay. Drop a whole override file in there rather than editing the values files the installer ships. Your overrides then survive an upgrade.

GPU sizing and single sign-on are configured the same way, by copying a shipped profile into that directory. See [GPU profiles and values overlays](/self-managed/v6.1/gpu-profiles-and-values-overlays) for the profile tables, the memory and compute-capability floors, and the rule for combining files.

## Single Sign-On

LILT ships with external SSO disabled, so password sign-in works out of the box and no install points at an identity provider you do not control.

To connect a provider, copy a file from `lilt/sso-profiles/` into `lilt/environments/lilt/values.d/`, then add the client secret as the `singleOidcClientSecret` key of the `lilt-secrets` secret. See [GPU profiles and values overlays](/self-managed/v6.1/gpu-profiles-and-values-overlays#single-sign-on-profiles) for the full value list.

<Warning>
  Set `SINGLE_OIDC_PROVIDER_ENABLED=true` only when every other `SINGLE_OIDC_PROVIDER_*` value is filled in. With one of them empty, `front` fails its startup validation and the pod crashes.
</Warning>

## Run the Install

```bash theme={null}
sh install-lilt-eks.sh 2>&1 | tee /var/log/lilt-install.log
```

Budget roughly two hours. Keep the log: when something fails 90 minutes in, it is the only account of what happened.

## Verify

```bash theme={null}
# Anything not Running or Completed needs a look.
kubectl get pods -n lilt | grep -v Running | grep -v Completed

# The ingress load balancer should have an address.
kubectl get svc -n istio-ingressgateway

# The application should answer through the ingress.
curl -sS -o /dev/null -w '%{http_code} verify=%{ssl_verify_result}\n' https://<SUBDOMAIN>.<DNS_DOMAIN>/
```

Read them outwards from the cluster. The first command filters to pods that need attention, so no output is the result you want. The second should show a hostname in `EXTERNAL-IP`, not `<pending>`. The third tests DNS, the load balancer, the gateway and the certificate in one request: any status code means it reached the gateway, and `verify=0` means the certificate chain validated. A chain missing its intermediate still satisfies a browser that has cached it, and fails later in a stricter client.

For the same checks with the failure modes spelled out, see [Install from a bastion](/self-managed/v6.1/install-from-a-bastion#step-6-verify).

## Resuming a Failed Install

The installer runs with `set -e`, so it stops at the first failing component and the last `--- Installing <Component> ---` line names it.

1. Inspect the component that failed:

   ```bash theme={null}
   kubectl get pods -n lilt
   kubectl describe pod <pod> -n lilt
   kubectl logs <pod> -n lilt
   ```

2. Fix the underlying cause: a missing value, an unreachable dependency, or a resource limit.

3. Run the installer again from the top. Components that already succeeded are no-ops.

To skip a component you have already dealt with, set its flag for that run only:

```bash theme={null}
ENABLE_MONGODB=false sh install-lilt-eks.sh
```

### Common Failures

**A pod stays in `ContainerCreating` for more than 15 minutes.** The image is usually still pulling. The neural images are tens of gigabytes. Check `kubectl describe pod` for a pull error before you intervene.

**`UPGRADE FAILED: timed out waiting for the condition`.** A component took longer to become ready than its timeout allowed. Confirm the pods are healthy and run the installer again.

**A pod cannot reach S3.** Check that the service account carries no `eks.amazonaws.com/role-arn` annotation if the cluster uses EKS Pod Identity, and that no static AWS credentials are set in the secrets file. Either one overrides workload identity.

## Uninstall

To remove LILT and leave the data in place:

```bash theme={null}
sh uninstall-lilt.sh
```

The script tears the stack down in dependency order: application, then data services, then platform. It also removes the objects Helm never owned, which a bare `helm uninstall` leaves behind.

Data volumes are kept by default. To destroy them as well, opt in on the command line:

```bash theme={null}
WIPE_DATA=true sh uninstall-lilt.sh
```

To remove the installer's custom resource definitions too:

```bash theme={null}
PURGE_CRDS=true sh uninstall-lilt.sh
```

<Warning>
  Pass `WIPE_DATA` and `PURGE_CRDS` on the command line only. Never put them in `install.env`, which persists across installs.
</Warning>

<Warning>
  To decommission an environment entirely, run the uninstall before you destroy the AWS resources. Destroying them first leaves the load balancer and the volumes behind persistent volume claims still attached to network interfaces, and the VPC deletion hangs.
</Warning>

Neither form removes the AWS resources themselves. Delete those separately, or with `terraform destroy` if you created them with the module.

## Related Articles

* [Install package](/self-managed/v6.1/install-package)
* [Provision AWS infrastructure with Terraform](/self-managed/v6.1/install-aws-terraform-module)
* [Install on an existing EKS cluster](/self-managed/v6.1/install-aws-existing-cluster)
* [Install from a bastion](/self-managed/v6.1/install-from-a-bastion)
* [Self-managed hardware requirements](/self-managed/v6.1/self-managed-hardware-requirements)
* [Infrastructure release notes](/self-managed/v6.1/infrastructure-release-notes)
