Cloud support
The platform runs on two clouds: aws-0 on EKS and gcp-0 on GKE Standard.
They are not one abstraction with two backends. They are two implementations of
the same three-stage model, sharing every Kubernetes-layer component and
diverging exactly where the clouds themselves diverge.
This page is the map of that divergence: what each cloud uses, where the two deliberately meet, and what is still missing on GCP.
Status at a glance
| Layer | aws-0 | gcp-0 |
|---|---|---|
| Network stack | ✅ opentofu/aws/network | ✅ opentofu/gcp/network |
| Secrets / PKI stack | ✅ opentofu/aws/openbao/{cluster,management} | ✅ opentofu/gcp/openbao/{cluster,management} |
| Kubernetes stack | ✅ opentofu/aws/eks/{init,configure} | ✅ opentofu/gcp/gke/{init,configure} |
| Namespaces · CRDs · Flux | ✅ | ✅ |
| Crossplane | ✅ provider-aws | ✅ provider-gcp |
| Security (cert-manager, ESO, Kyverno, Tailscale) | ✅ | ✅ |
| Infrastructure (Cilium, Gateway API, external-dns) | ✅ | ✅ |
| Observability (VictoriaMetrics, Grafana, RunLore) | ✅ | ✅ same stack |
| Tooling (Harbor) | ✅ | ✅ Harbor on GCS with Workload Identity |
| Applications | ✅ | ✅ podinfo · basic · App Wizard · image-gallery (GCS through Workload Identity) |
| LLM platform | ⏸️ opt-in, suspended | ⏸️ opt-in, suspended |
| Flux extras (alerts, dashboards) | ✅ | ✅ minus flux-previews |
One nuance in the Infrastructure row: Cilium itself is OpenTofu-owned on both
clouds (Stage 2 of the Kubernetes stack), not Flux-managed — and the Cilium
extras aws-0 wires through Flux (infrastructure/base/cilium: Hubble UI
route, dashboards, scrape configs) have no gcp-0 entry.
Matching managed services
The Kubernetes layer is identical on both clouds — same Cilium, same Flux, same OpenBao, same VictoriaMetrics, same Gateway API. Everything below is the layer where a cloud’s own service is unavoidable, and what stands in for what.
Compute and networking
| Concern | AWS | GCP | Notes |
|---|---|---|---|
| Managed Kubernetes | EKS | GKE Standard | Standard, not Autopilot — Autopilot forbids the DaemonSet privileges Cilium needs (ADR-0005) |
| CNI | Cilium (replaces VPC-CNI) | Cilium (displaces GKE’s) | Same chart, same version, both self-managed (ADR-0009) |
| Node autoscaling | Karpenter | Node Auto-Provisioning + ComputeClass | (ADR-0006) |
| Node OS | Bottlerocket | Container-Optimized OS (cos_containerd) | |
| Load balancer | ELB / NLB | Google Cloud Load Balancing | Both fronted by Gateway API, not consumed directly |
| Private access | Tailscale subnet router | Tailscale subnet router | One tailnet spans both clouds — opentofu/shared/tailscale |
| Encryption in transit | Cilium WireGuard | not required | The WireGuard workaround is an AWS prefix-delegation issue; GKE does not hit it |
Storage
| Concern | AWS | GCP | Notes |
|---|---|---|---|
| Block storage class | gp3 (EBS) | standard-rwo (pd-balanced) | Supplied to manifests as ${storage_class} — no PVC hardcodes either |
| Object storage | S3 | Cloud Storage | |
| Model weights (LLM) | Amazon S3 Files | Cloud Storage FUSE CSI | (ADR-0004, ADR-0021) |
| Registry storage | Harbor → S3 driver | Harbor → GCS driver | Harbor itself is self-hosted on both (ADR-0020) |
| OpenTofu state | S3 bucket | GCS bucket, dedicated project | Deliberately not shared (ADR-0018) |
Identity and secrets
| Concern | AWS | GCP | Notes |
|---|---|---|---|
| Workload identity | EKS Pod Identity | GKE Workload Identity Federation | Never IRSA, never a static key (ADR-0002) |
| Crossplane claim | EPI | GCPWorkloadIdentity | Cloud-shaped on purpose (ADR-0007) |
| Bootstrap secret store | AWS Secrets Manager | Google Secret Manager | Read at apply time by the cluster stack |
| Runtime secret store (ESO) | AWS Secrets Manager | GCP Secret Manager | The one ClusterSecretStore every ExternalSecret reads (ADR-0025) |
| Private PKI | OpenBao | OpenBao | Same PKI model, one instance per cloud |
| OpenBao auto-unseal | AWS KMS | Cloud KMS |
DNS and certificates
This is the one place the two clouds are deliberately not symmetric, and the asymmetry is the point.
| Concern | AWS | GCP |
|---|---|---|
| Private zone | Route 53 private hosted zone — priv.aws.ogenki.io | Cloud DNS private zone — priv.gcp.ogenki.io |
| Public zone | Route 53 — cloud.ogenki.io | Route 53, via AWS IAM OIDC federation |
| external-dns (private) | provider: aws | provider: google |
| external-dns (public) | provider: aws | provider: aws — a second instance |
| Public certificate issuance | Let’s Encrypt DNS-01 → Route 53 | Let’s Encrypt DNS-01 → Route 53, federated |
Private DNS is native on each cloud. Only the public zone is centralised —
and only because cloud.ogenki.io is a Route 53 zone this repository does not
manage, while Let’s Encrypt must resolve _acme-challenge publicly to issue a
certificate. Delegating a subdomain to Cloud DNS or moving the zone were both
considered and rejected; what made federation acceptable is that it needs no
static AWS credential on GCP — cert-manager and external-dns assume an AWS
role using a projected Kubernetes ServiceAccount token, the same
identity-by-token model both clouds already use internally. The full argument,
including what the dependency costs, is in
ADR-0019.
Data services
Neither cloud’s managed database is used. PostgreSQL is CloudNativePG and key/value is Valkey — both self-hosted, both driven by a Crossplane claim, and both now work on either cloud:
| Claim | AWS | GCP |
|---|---|---|
SQLInstance (PostgreSQL) | ✅ CloudNativePG, barman backups to S3 | ✅ CloudNativePG, barman backups to Cloud Storage |
KVStore (Valkey) | ✅ | ✅ cloud-neutral Composition — no cloud resources, so it works unchanged |
SQLInstance was the last claim to reach GCP, and until
crossplane-configuration v0.4.0 its GCP Composition was a deliberate dead-end —
it failed evaluation rather than composing nothing, so a claim said why instead
of hanging Ready=Unknown. That was consistent with
ADR-0007: a claim that
cannot be honoured should say so at reconcile time.
It is a real implementation now. Both clouds render from the same KCL module,
differing in where barman writes (gs:// with googleCredentials.gkeEnvironment
rather than s3:// with s3Credentials.inheritFromIAMRole) and in the identity
that writes — one GCPWorkloadIdentity, bucket-scoped, in place of four AWS IAM
resources. The CloudNativePG operator runs on both clouds; gcp-0 pulls the same
whole base directory aws-0 does (infrastructure/gcp-0/cloudnative-pg/
references ../../base/cloudnative-pg plus its own gcs-bucket.yaml) — including
the base’s three Grafana dashboards, since both of their preconditions hold on
gcp-0 too: crds/base installs the GrafanaDashboard CRD, and
clusters/gcp-0/observability/observability-grafana-operator.yaml runs the
operator that reconciles it.
Decisions that shaped the split
Seven records carry the multicloud reasoning. Read in this order they explain why the platform looks the way it does on a second cloud:
The shared layer
Two things belong to neither cloud and are provisioned once:
opentofu/shared/tailscale— one tailnet, both clusters. Its state stays in S3 precisely because the tailnet is not a GCP resource or an AWS one.opentofu/shared/aws-gcp-federation— the AWS IAM OIDC provider that trusts the GKE issuer, which is what makes the DNS row above work.
Known gaps
gcp-0 now runs the same layers as aws-0. Four components are still
excluded, and the distinction that matters is why — three of them are not
gaps at all.
Excluded by design, not missing
flux-previews— PR preview environments. Running them on both clusters would double-provision every preview and both would write the same public DNS records. Previews belong to one cluster by nature. (It also hardcodescluster_name: aws-0while living in the sharedflux/tree, which is worth fixing regardless.)karpenter/karpenter-nodepools— GCP uses Node Auto-Provisioning with ComputeClasses instead (ADR-0006).eks-pod-identities—GCPWorkloadIdentityis the counterpart, a different Kind rather than a second Composition (ADR-0002).
Genuinely not portable yet
- Harbor’s database has no backups on
gcp-0. The claim side is ready — the Composition renders barman’sObjectStoreand a bucket-scoped identity as soon asbackupis set. The cluster side is not: the barman plugin ships aCiliumNetworkPolicywhose egress is atoFQDNsallowlist of S3 endpoints, so deploying it would let the plugin start and then silently drop every connection tostorage.googleapis.com.
Needs a human, on both clouds
Two ExternalSecrets read from the cloud’s managed secret store and nothing
seeds them — ./scripts/provision/secret-store.sh check --cloud gcp (or aws) lists
what is missing, and its seed command creates the generatable ones:
- Harbor’s admin and Valkey passwords, at
harbor-admin-password - Flux’s Slack token — the key named in
flux/notifications/externalsecret-flux-slack-app.yaml
Until they exist, Harbor waits on its secret and Flux alerts are dropped. Reconciliation itself is unaffected.
None of the above are cloud-abstraction failures. The shared layer reconciles identically on both clouds; what is left is one application with a cloud baked into its code, one policy to port, and secrets to seed.
One identity provider, or two? — settled
ZITADEL is exposed at auth.${public_domain_name}, and that variable is
per-cluster — cloud.ogenki.io on aws-0, gcp.cloud.ogenki.io on gcp-0 —
so there is no DNS collision either way.
ADR-0022 first
made ZITADEL a singleton on aws-0, reachable across the cloud boundary.
ADR-0024 superseded it:
the IdP is now a per-cloud deployable component, defaulting to AWS, so a
GCP-only platform can authenticate without an AWS cluster running. The accepted
cost is one user directory per cloud, with no federation between them. Only
public DNS stays AWS-owned (ADR-0019).
gcp-0 does not take that opt-out: AWS is the primary cloud
(ADR-0027), so aws-0 hosts
the one instance and gcp-0 consumes it. The opt-out exists for a GCP-only
platform, where the singleton relocates rather than being duplicated — two
clouds each running a directory is ruled out, since a grant means nothing
without knowing which directory issued it.
Placement has two halves, and only one of them is typed:
| Gate | Where | Value today |
|---|---|---|
| 1 · which URL consumers read | deploy_identity_provider, derived from primary_cloud in opentofu/config.tm.hcl | false on gcp-0 |
| 2 · whether an instance runs | spec.suspend in clusters/gcp-0/security/zitadel.yaml | true |
Gate 1 cannot disagree with the declaration, because it is the declaration.
Gate 2 is committed Flux state — Flux never reads Terramate globals — so it is
verified instead, by ./scripts/ci/validate-idp-topology.sh in CI. Changing which
cloud hosts is a migration,
not a toggle: the database seed, admin credential and OIDC clients travel with
it.
Adding a third cloud
The mechanics of extending this — which values are cloud-neutral, which are per-cluster, and what a new lane must supply — are in Guides → Add a cloud provider.