AWS
AWS is one of two implemented cloud lanes — see
GCP for the other, and
Cloud support for
what each one runs. Every stage below is an OpenTofu stack orchestrated by
Terramate; each stack declares the ones it depends on
(after in its stack.tm.hcl), so Terramate always applies them in the right
order even when you run one command that spans several stages.
Configure before deploying
Every stack already has a variables.tfvars committed, holding values that
deploy against the reference environment. You are editing a working
configuration, not authoring one — which also means the fastest way to see what
a stack expects is to read the file already sitting next to it.
Edit
opentofu/config.tm.hcl— region, EKS cluster name, the Helm chart versions used by the bootstrap (cilium_version,flux_operator_version,flux_instance_version), andopenbao_url.Edit the
variables.tfvarsin each stack directory (opentofu/aws/network/,opentofu/aws/openbao/cluster/,opentofu/aws/openbao/management/,opentofu/aws/eks/init/,opentofu/aws/eks/configure/), replacing the reference values with yours — principallycluster_name,env,private_domain_nameandpublic_domain_name— and above allflux_sync_urlinopentofu/aws/eks/configure/variables.tfvars: the URL of your fork, which is what Flux actually syncs from.Edit
opentofu/shared/tailscale/variables.tfvars—tailnetandadmin_usersare the reference tailnet’s identity, and this stack is the first thing the root deploy applies; point them at your own tailnet.Export the one secret Terraform needs from the environment rather than a file:
export TF_VAR_tailscale_api_key=<YOUR_TAILSCALE_API_KEY>
Expiring old generation archives is not a manual step
Each CNPG cluster generation writes to its own xplane-* prefix and nothing
reclaims them by hand — the bucket’s own claim carries a
BucketLifecycleConfiguration that expires them automatically, scoped so it
cannot touch the dated seeds (named zitadel-* / harbor-*, never
xplane-). See infrastructure/aws-0/cloudnative-pg/s3-bucket.yaml.
Deploy
cd opentofu
terramate script run deployThat is the whole deploy. One command, from opentofu/, for all three
stages — that is what Terramate is for. It resolves the dependency graph and
applies every stack in order; there is no stage you have to drive by hand. The
one exception is the very first deploy on a new platform:
The first deploy on a new platform ends red at stage5-verify-openbao-oidc,
and that is expected. OpenBao’s management stack runs before ZITADEL has
issued OpenBao’s OIDC client, so it cannot create the oidc/ mount yet. Create
it once, then resume — from the repository root:
terramate -C opentofu/aws/openbao/management script run deploy
cd opentofu && terramate script run deployOn a TM_CLOUD=aws,gcp run, TM_CLOUD=gcp terramate -C opentofu/gcp/gke/init script run deploy
resumes just the GCP stacks the halt skipped. See
OIDC client rotation.
No cloud flag either: TM_CLOUD defaults to aws, so this builds the AWS lane
and skips GCP. Set TM_CLOUD=aws,gcp (or all) to build both clouds in the same
run — see Commands.
What that one command does, in order:
shared/tailscale first — the tailnet-wide singletons network declares in
its after list. The same run also applies shared/aws-gcp-federation, which
belongs to the shared lane rather than either cloud: the GCP lane needs it and
it costs nothing idle here.
Stage 1 — the network. A VPC across three availability zones, public and private subnets, a Route53 private hosted zone, VPC endpoints, and the Tailscale subnet router EC2 instance that gives you private access to everything built after this point.
Stage 2 — OpenBao. The lineage (a persistent seal key and the snapshot
bucket), then the cluster behind a Network Load Balancer, then its
configuration. As committed, opentofu/aws/openbao/cluster/variables.tfvars sets
mode = "dev": a single t3.micro on single-node Raft, enough to follow these
guides and not highly available. Set mode = "ha" for the five-node Raft cluster
on spot instances — the same configuration steps apply either way.
Either way the cluster is auto-unsealed via AWS KMS with the lineage’s key, and
./scripts/provision/openbao-config.sh rehydrate brings its contents back from the newest
snapshot — so there is no manual bao operator init / unseal step, and on the
first deploy of a lineage the root token and recovery keys are written to two
separate AWS Secrets Manager entries. The management stack then layers a
three-tier PKI (root → intermediate → leaf) and the policies each cluster’s JWT
auth roles bind. Taking the AWS root offline is decided but not yet performed —
see the callout on
PKI & Secrets and
ADR-0033.
Stage 3 — Kubernetes. aws/eks/init runs a two-stage bootstrap internally:
the EKS cluster comes up with the temporary VPC-CNI bootstrap addon, then that is
replaced with Cilium (which also replaces kube-proxy) and the Flux Operator and
Instance are installed — the point at which the cluster starts reconciling the
rest of this repository from Git. A third internal step recycles any node-group
node whose ENIs predate Cilium so it can pick up prefix delegation (see
opentofu/aws/eks/init/workflows.tm.hcl).
aws/eks/configure is applied twice — once by eks/init’s stage 2, which shells
into it, and once as its own stack when Terramate reaches it. The second apply is
a no-op, so this is waste rather than breakage, but it is why the run reports one
more stack than you might expect. GKE does the same thing.Deploying one stack on its own
Rarely needed, and never required by the flow above — but each stack’s script also runs from its own directory, which is useful when re-running a single stage after a failure:
cd opentofu/aws/eks/init
terramate script run deploy # just the Kubernetes stageVerify
The API endpoint is private, so the Tailscale subnet router from stage 1 must be
up first (tailscale status):
aws eks update-kubeconfig --region eu-west-3 --name aws-0
kubectl get nodes
flux get kustomizationsOnce the deploy finishes, Flux takes over and reconciles the rest without any further command from you: Security (External Secrets, cert-manager, Kyverno), Infrastructure (Cilium policies, Gateway API, Karpenter), Observability (VictoriaMetrics, VictoriaLogs, Grafana) and Tooling (Harbor, Headlamp, Homepage).
That is where a bootstrap actually succeeds or quietly stalls, so there is a page for checking it layer by layer: Verify the cluster.
ZITADEL comes up with nothing configured in it, so Google login does not work until you run the setup steps in Set up single sign-on — the same steps on both clouds, with different values.
See also Access for reaching the VPN, OpenBao, the cluster and the dashboards, and Teardown when you are done.