Skip to content

Prerequisites

The tools, the GitHub App and the Tailscale account below apply to both cloud lanes. The account and state-bucket sections are per cloud.

Public DNS needs an AWS account with Route 53, whichever cloud you deploy. The public domain is a cross-cluster concern — one zone for the whole platform, not one per cloud — and this reference implementation hosts that zone on Route 53; gcp-0 reaches it by assuming an AWS IAM role (ADR-0019). The domain itself can be registered anywhere — Route 53, Cloudflare, Gandi… — as long as it delegates the public zone to Route 53.

Accounts and access

  • AWS account with admin-level permissions (VPC, EKS, IAM, S3, Route53, Secrets Manager, KMS) and credentials configured locally (~/.aws/credentials or environment variables). Required for the AWS lane, and required in a reduced form for the GCP lane’s public DNS.

  • GCP project and organisation — only for the GCP lane, plus a second project that holds nothing but OpenTofu state. gcloud authenticated.

  • A registered domain — any registrar works, as long as its public zone is delegated to Route 53 (see the callout above). OpenTofu also creates a private hosted zone under it for internal service DNS.

  • GitHub account — Flux needs a way to pull this repository: a personal access token or a GitHub App.

  • Tailscale account and API key — provisions the subnet router that gives you private access to the cluster.

  • A GitHub App, and its credentials in AWS Secrets Manager — Flux authenticates to pull this repository as a GitHub App, and opentofu/aws/eks/configure reads its credentials from Secrets Manager at apply time (var.github_app_secret_name, default github/flux-app); if the secret does not exist, Stage 3 of the AWS deploy fails. Create the App per the Flux GitHub App docs, then publish its credentials:

    jq -n --arg key "$(cat your-githubapp.private-key.pem)" \
      '{githubAppID: "<app_id>", githubAppInstallationID: "<installation_id>", githubAppPrivateKey: $key}' \
      > flux-ghapp.json
    
    aws secretsmanager create-secret \
      --name github/flux-app \
      --description "FluxCD Github App" \
      --region eu-west-3 \
      --secret-string file://flux-ghapp.json

    GCP lane — opentofu/gcp/gke/configure reads the same credentials from GCP Secret Manager instead (var.flux_github_app_secret_name, default flux-github-app), so the GCP bootstrap never depends on AWS credentials. Publish the same JSON payload there too:

    gcloud secrets create flux-github-app --replication-policy=automatic \
      --project=<your-project> --data-file=flux-ghapp.json

State backend — create this bucket first

Nothing in this repository can plan until a state bucket exists. It is the one prerequisite OpenTofu cannot create for you: every stack stores its state in it, so a stack that created it would have nowhere to record that it had. That chicken-and-egg is why this step is manual, and why it is easy to forget — the failure on a fresh clone is a backend error from the first tofu init, not a message telling you to read this page.

Which bucket you need differs by lane:

  • AWS lane — one S3 bucket, created below. That is all.
  • GCP lane — a GCS bucket in a project that holds nothing else (created on Get Started → GCP, along with GCP’s two other hand-created prerequisites), plus the S3 bucket below anyway: the shared stacks that belong to neither cloud — the tailnet singletons in opentofu/shared/tailscale and the AWS↔GCP DNS federation stack — keep their state in S3, and the GCP lane cannot deploy without them.

State for cloud-owned stacks is per cloud: AWS stacks use the S3 bucket, GCP stacks the GCS bucket. The principle is that state lives outside the blast radius of what it manages, without coupling the clouds — the history behind that split and its trade-offs are in ADR-0018.

The AWS bucket

BUCKET=demo-smana-remote-backend   # must match the backend blocks; see below
REGION=eu-west-3

aws s3api create-bucket \
  --bucket "$BUCKET" --region "$REGION" \
  --create-bucket-configuration LocationConstraint="$REGION"

# Versioning is the ONLY recovery path if state is corrupted or wrongly
# overwritten. Turn it on before the first apply, not after.
aws s3api put-bucket-versioning \
  --bucket "$BUCKET" \
  --versioning-configuration Status=Enabled

aws s3api put-bucket-encryption \
  --bucket "$BUCKET" \
  --server-side-encryption-configuration \
  '{"Rules":[{"ApplyServerSideEncryptionByDefault":{"SSEAlgorithm":"aws:kms","KMSMasterKeyID":"alias/aws/s3"}}]}'

aws s3api put-public-access-block \
  --bucket "$BUCKET" \
  --public-access-block-configuration \
  "BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true"

State objects contain credentials and private keys in plaintext, so the encryption and public-access settings above are not optional hardening — treat the bucket as a secret store.

No DynamoDB table is needed. The backends set use_lockfile = true, which uses native S3 locking via a .tflock object; the older DynamoDB locking table that Terraform guides describe is obsolete here.

If you use a different bucket name, you must edit it in every stack, because an OpenTofu backend block cannot take a variable. Find them all with:

grep -rn 'bucket ' --include=backend.tf opentofu/

One cross-stack reader hardcodes it a second time — the terraform_remote_state data source in opentofu/aws/llm-platform/data.tf. Changing the backends and missing it leaves the reader pointing at a bucket that no longer receives writes: stale reads, no error. (The GCS bucket has its own pair of hardcoded readers, noted where that bucket is created.)

Tools

This repository pins every CLI version it depends on in mise.toml — install mise, then run:

mise install

That single command installs OpenTofu, Terramate, the Flux CLI, Helm, Kustomize, the Google Cloud SDK (gcloud), and Trivy (the config scanner every preview/deploy/drift detect script runs) at the exact versions this repository is built against. mise.toml is the source of truth for those versions — check it directly rather than trusting a number written in prose, here or anywhere else.

A few tools mise does not manage — install these separately:

  • the AWS CLI, authenticated
  • kubectl
  • gke-gcloud-auth-plugin (GCP lane) — gcloud components install gke-gcloud-auth-plugin; kubectl cannot talk to a GKE cluster without it
  • the OpenBao CLI (bao) — see openbao.org
  • jq
  • the Tailscale client, to check tailscale status once Stage 1 is up

With accounts in place and tools installed, continue to your cloud lane — AWS or GCP. Both are implemented; they differ in how much of the platform reconciles on top, which Cloud support lays out side by side.