Skip to content
Repository Layout

Repository Layout

Every deployable domain in this repository follows the same shape: <domain>/base/<component>/ holds the cloud-neutral Kubernetes manifests, <domain>/<cluster>/ overlays them for one cluster, and a Flux Kustomization under clusters/<cluster>/ (or a sibling clusters/<cluster>-*/ for an opt-in umbrella) wires the overlay into the reconciliation graph.

There are two clusters: aws-0 and gcp-0. Both read the same base/ directories, so a component is written once and overlaid twice. Infrastructure that predates Kubernetes — the VPC, the cluster itself, OpenBao — lives under opentofu/<cloud>/ instead, orchestrated by Terramate.

The rule that keeps this honest: anything in base/ must work on both clouds. A manifest only one cloud can use belongs in that cloud’s overlay, even when both clouds happen to want a file of the same name — see infrastructure/aws-0/gapi/platform-public-gateway.yaml, and clustersecretstore.yaml, which exists under both security/aws-0/openbao/ and security/gcp-0/openbao/ because one names the AWS Secrets Manager provider and the other GCP Secret Manager.

PathWhat lives thereReconciled by
opentofu/Everything that has to exist before a Kubernetes API doesOpenTofu / Terramate
opentofu/aws/network/VPC, subnets, Route53, the AWS-side Tailscale subnet router
opentofu/aws/openbao/lineage/What outlives every rebuild — the snapshot bucket, its KMS key, the GitHub OIDC role
opentofu/aws/openbao/cluster/The active OpenBao instance, rebuilt from its newest snapshot on every deploy
opentofu/aws/openbao/management/What is configured inside it — the private PKI, policies, the operator login
opentofu/aws/eks/init/Stage 1 — EKS cluster, node groups, bootstrap addons, Gateway API CRDs
opentofu/aws/eks/configure/Stage 2 — Cilium replaces the CNI, the Flux Operator is installed
opentofu/aws/llm-platform/opt-in — S3 Files and IAM for the self-hosted LLM platform
opentofu/gcp/network/VPC, subnets, Cloud DNS, the GCP-side Tailscale subnet router
opentofu/gcp/gke/init/Stage 1 — GKE Standard cluster with a private-only endpoint
opentofu/gcp/gke/configure/Stage 2 — Cilium displaces GKE's CNI, the Flux Operator is installed
opentofu/gcp/openbao/lineage/The GCS snapshot bucket, and the transfer job mirroring the AWS one into it
opentofu/gcp/openbao/cluster/The GCP standby — restores the mirrored snapshot under the same AWS seal key
opentofu/gcp/openbao/management/What is configured inside it — the private PKI, its issuing role, the two policies
opentofu/shared/tailscale/The tailnet itself — ACLs, nameservers, search paths — owned by neither cloud
opentofu/shared/aws-gcp-federation/Workload Identity Federation letting gcp-0 write Route53 records
clusters/The Flux Kustomizations themselves — the entry point the FluxInstance points atflux-system
clusters/aws-0/One Kustomization per domain, carrying the dependsOn graph
clusters/aws-0-llm-platform/The suspended LLM umbrella's children, kept a sibling so the recursive sync cannot reach them
clusters/gcp-0/Same shape as aws-0 — one Kustomization per domain, over the GCP overlays
clusters/gcp-0-llm-platform/gcp-0's suspended LLM umbrella children, same sibling pattern
flux/Flux managing itself — sources, notifications, previews, and its own dashboardsflux-* (six)
flux/operator/Flux Operator and the FluxInstance
flux/sources/Every GitRepository, HelmRepository and OCIRepository — deliberately unsharded
flux/artifact-generators/The ArtifactGenerator that slices the repository into one ExternalArtifact per domain
flux/notifications/Alert and Provider resources — reconciliation failures to Slack
flux/observability/Flux's own Grafana dashboards and alert rules
flux/previews/ResourceSet-driven preview environments
namespaces/Namespace manifests — first in the dependency chain, because everything else lands in onenamespaces
crds/Custom Resource Definitions applied ahead of the controllers that consume themcrds
infrastructure/Cilium policies, Crossplane, Karpenter, Gateway API, External DNS, CSI drivers, KEDAinfrastructure, crossplane-*, karpenter
security/cert-manager, External Secrets, Kyverno, ZITADEL, the EKS Pod Identities, RBACsecurity, eks-pod-identities, zitadel
observability/VictoriaMetrics, VictoriaLogs, VictoriaTraces, Grafana, and the RunLore SRE agentobservability-*
tooling/Harbor, Headlamp, Homepage, self-hosted runners (off by default)tooling
apps/App composition claims — the tenant-facing workload API, and the demo applicationsapps
container-images/Dockerfiles for the images this repository builds and publishes to ghcr.io— built in CI, not reconciled
scripts/validate-manifests.sh, validate-links.sh, verify-doc-paths.sh, openbao-config.sh, and the rest—
docs/Diagram sources, the platform constitution, the superpowers designs and plans, and the read-only spec archive—
website/This documentation site — Hugo, Hextra, and the content you are reading—

The base / overlay pattern

Within infrastructure/, security/, observability/ and tooling/, each component gets its own directory under base/ — typically a HelmRelease or a handful of raw manifests plus a kustomization.yaml. The cluster’s own kustomization.yaml lists which of those base components are actually included; a component present under base/ but absent (or commented out) from the overlay is not deployed. tooling/base/gha-runners is the clearest example — present under base/, commented out in tooling/aws-0/kustomization.yaml, so the self-hosted CI runners stay off by default.

This is also why the two clusters run different amounts of platform: gcp-0’s overlays simply include fewer base components today. Which ones, and why, is on Cloud support.

A handful of components bypass the shared overlay and get their own top-level Flux Kustomization directly under clusters/<cluster>/<domain>/ instead — on aws-0: crossplane-controller, karpenter, grafana-operator, victoria-metrics, victoria-traces and zitadel; gcp-0 adds gapi/gapi-public and four security-* Kustomizations (OpenBao, its snapshots, public certs, Tailscale) — usually because they need their own dependsOn ordering rather than sharing the domain’s overlay lifecycle.

Two deployment models in one tree

  • OpenTofu / Terramate (opentofu/) provisions everything that has to exist before a Kubernetes API server does: the VPC, the cluster itself, and the OpenBao cluster. Split opentofu/aws/, opentofu/gcp/ and opentofu/shared/ — the last holding the tailnet and the AWS↔GCP DNS federation, which belong to neither cloud. See Commands.
  • Flux / Kustomize (everything else) reconciles the cluster once it exists. clusters/aws-0/ and clusters/gcp-0/ are the entry points each cluster’s FluxInstance points at; from there the dependency chain runs Namespaces → CRDs → Crossplane → workload identities → Security → Infrastructure → Observability → Applications.

Opt-in surfaces

Several parts of the tree are deliberately inert by default, each gated independently of the base/overlay mechanism above:

  • opentofu/aws/llm-platform/ — a Terramate stack tagged opt-in; its scripts no-op unless TM_LLM_PLATFORM_ENABLED=true.
  • clusters/aws-0-llm-platform/ — an umbrella Flux Kustomization (clusters/aws-0/llm-platform.yaml) with spec.suspend: true, kept a sibling of clusters/aws-0/ specifically so flux-system’s recursive sync does not pick up its children and bypass the suspend.
  • clusters/gcp-0-llm-platform/ — the same suspended-umbrella pattern for gcp-0, gated by clusters/gcp-0/llm-platform.yaml.
  • opentofu/<lane>/** — the directory is the cloud selector. Stacks under aws/ and gcp/ run only when TM_CLOUD names their lane (it defaults to aws); anything under shared/ is owned by neither cloud and always runs. So an AWS deploy from opentofu/ never builds GCP as a side effect, and a GCP one never rebuilds aws-0.

On aws-0 both LLM gates have to be released for an end-to-end deploy; on gcp-0 the umbrella is the only LLM gate — see the clusters/AGENTS.md.