Skip to content

AWS

AWS is one of two implemented cloud lanes — see GCP for the other, and Cloud support for what each one runs. Every stage below is an OpenTofu stack orchestrated by Terramate; each stack declares the ones it depends on (after in its stack.tm.hcl), so Terramate always applies them in the right order even when you run one command that spans several stages.

Configure before deploying

Every stack already has a variables.tfvars committed, holding values that deploy against the reference environment. You are editing a working configuration, not authoring one — which also means the fastest way to see what a stack expects is to read the file already sitting next to it.

  1. Edit opentofu/config.tm.hcl — region, EKS cluster name, the Helm chart versions used by the bootstrap (cilium_version, flux_operator_version, flux_instance_version), and openbao_url.

  2. Edit the variables.tfvars in each stack directory (opentofu/aws/network/, opentofu/aws/openbao/cluster/, opentofu/aws/openbao/management/, opentofu/aws/eks/init/, opentofu/aws/eks/configure/), replacing the reference values with yours — principally cluster_name, env, private_domain_name and public_domain_name — and above all flux_sync_url in opentofu/aws/eks/configure/variables.tfvars: the URL of your fork, which is what Flux actually syncs from.

  3. Edit opentofu/shared/tailscale/variables.tfvars — tailnet and admin_users are the reference tailnet’s identity, and this stack is the first thing the root deploy applies; point them at your own tailnet.

  4. Export the one secret Terraform needs from the environment rather than a file:

    export TF_VAR_tailscale_api_key=<YOUR_TAILSCALE_API_KEY>

Expiring old generation archives is not a manual step

Each CNPG cluster generation writes to its own xplane-* prefix and nothing reclaims them by hand — the bucket’s own claim carries a BucketLifecycleConfiguration that expires them automatically, scoped so it cannot touch the dated seeds (named zitadel-* / harbor-*, never xplane-). See infrastructure/aws-0/cloudnative-pg/s3-bucket.yaml.

Deploy

cd opentofu
terramate script run deploy

That is the whole deploy. One command, from opentofu/, for all three stages — that is what Terramate is for. It resolves the dependency graph and applies every stack in order; there is no stage you have to drive by hand. The one exception is the very first deploy on a new platform:

The first deploy on a new platform ends red at stage5-verify-openbao-oidc, and that is expected. OpenBao’s management stack runs before ZITADEL has issued OpenBao’s OIDC client, so it cannot create the oidc/ mount yet. Create it once, then resume — from the repository root:

terramate -C opentofu/aws/openbao/management script run deploy
cd opentofu && terramate script run deploy

On a TM_CLOUD=aws,gcp run, TM_CLOUD=gcp terramate -C opentofu/gcp/gke/init script run deploy resumes just the GCP stacks the halt skipped. See OIDC client rotation.

No cloud flag either: TM_CLOUD defaults to aws, so this builds the AWS lane and skips GCP. Set TM_CLOUD=aws,gcp (or all) to build both clouds in the same run — see Commands.

What that one command does, in order:

shared/tailscale first — the tailnet-wide singletons network declares in its after list. The same run also applies shared/aws-gcp-federation, which belongs to the shared lane rather than either cloud: the GCP lane needs it and it costs nothing idle here.

Stage 1 — the network. A VPC across three availability zones, public and private subnets, a Route53 private hosted zone, VPC endpoints, and the Tailscale subnet router EC2 instance that gives you private access to everything built after this point.

Stage 2 — OpenBao. The lineage (a persistent seal key and the snapshot bucket), then the cluster behind a Network Load Balancer, then its configuration. As committed, opentofu/aws/openbao/cluster/variables.tfvars sets mode = "dev": a single t3.micro on single-node Raft, enough to follow these guides and not highly available. Set mode = "ha" for the five-node Raft cluster on spot instances — the same configuration steps apply either way.

Either way the cluster is auto-unsealed via AWS KMS with the lineage’s key, and ./scripts/provision/openbao-config.sh rehydrate brings its contents back from the newest snapshot — so there is no manual bao operator init / unseal step, and on the first deploy of a lineage the root token and recovery keys are written to two separate AWS Secrets Manager entries. The management stack then layers a three-tier PKI (root → intermediate → leaf) and the policies each cluster’s JWT auth roles bind. Taking the AWS root offline is decided but not yet performed — see the callout on PKI & Secrets and ADR-0033.

Stage 3 — Kubernetes. aws/eks/init runs a two-stage bootstrap internally: the EKS cluster comes up with the temporary VPC-CNI bootstrap addon, then that is replaced with Cilium (which also replaces kube-proxy) and the Flux Operator and Instance are installed — the point at which the cluster starts reconciling the rest of this repository from Git. A third internal step recycles any node-group node whose ENIs predate Cilium so it can pick up prefix delegation (see opentofu/aws/eks/init/workflows.tm.hcl).

aws/eks/configure is applied twice — once by eks/init’s stage 2, which shells into it, and once as its own stack when Terramate reaches it. The second apply is a no-op, so this is waste rather than breakage, but it is why the run reports one more stack than you might expect. GKE does the same thing.

Deploying one stack on its own

Rarely needed, and never required by the flow above — but each stack’s script also runs from its own directory, which is useful when re-running a single stage after a failure:

cd opentofu/aws/eks/init
terramate script run deploy          # just the Kubernetes stage

Verify

The API endpoint is private, so the Tailscale subnet router from stage 1 must be up first (tailscale status):

aws eks update-kubeconfig --region eu-west-3 --name aws-0
kubectl get nodes
flux get kustomizations

Once the deploy finishes, Flux takes over and reconciles the rest without any further command from you: Security (External Secrets, cert-manager, Kyverno), Infrastructure (Cilium policies, Gateway API, Karpenter), Observability (VictoriaMetrics, VictoriaLogs, Grafana) and Tooling (Harbor, Headlamp, Homepage).

That is where a bootstrap actually succeeds or quietly stalls, so there is a page for checking it layer by layer: Verify the cluster.

ZITADEL comes up with nothing configured in it, so Google login does not work until you run the setup steps in Set up single sign-on — the same steps on both clouds, with different values.

See also Access for reaching the VPN, OpenBao, the cluster and the dashboards, and Teardown when you are done.