GCP
GCP instantiates the three-stage model
as six OpenTofu stacks, mirroring AWS
stack for stage. Each stack’s stack.tm.hcl declares what it runs after, so
Terramate applies them in this order:
| Stack | Model stage | Owns |
|---|---|---|
opentofu/gcp/network/ | Network | VPC, node/pod/service ranges, the control-plane CIDR, a private Cloud DNS zone, the Tailscale subnet router |
opentofu/gcp/openbao/lineage/ | Security | Snapshot bucket (also the mirror of AWS snapshots), node and drill identities, Storage Transfer job, GitHub WIF pool. Persistent |
opentofu/gcp/openbao/cluster/ | Security | OpenBao on Compute Engine, single-node Raft, Cloud KMS or the AWS seal (standby) |
opentofu/gcp/openbao/management/ | Security | The three-tier PKI, the lineage/ marker mount, policies, and the rehydrate step. Persistent |
opentofu/gcp/gke/init/ | Kubernetes (Stage 1) | The GKE cluster, the static node pool, Gateway API CRDs, IAM, the flux-system namespace and secrets |
opentofu/gcp/gke/configure/ | Kubernetes (Stage 2) | Cilium, Flux Operator + Instance |
Bootstrap is two stacks here for exactly the reason it is on AWS: a provider
cannot be configured from an attribute of a resource the same configuration is
about to create, so the Cilium and Flux helm_releases must live in a second
root module that reads the finished cluster back with a data source. That
argument is written out in full on the
AWS page
and is not repeated here.
Prerequisites are not the same as AWS. GCP has three hand-created
bootstrap items — a state bucket in its own project, a Cloud KMS key ring, and
a Tailscale OAuth client. The sequence is in docs/gcp-bootstrap.md in the
repository. See also Get Started → GCP.
The committed shape
| Value | |
|---|---|
| Location | europe-west4-a — zonal, not regional |
| Cluster | GKE Standard (ADR-0005) |
| Node image | COS_CONTAINERD |
| Static pool | 2 × e2-standard-4, spot |
| Control plane | Private endpoint, private nodes — reachable only over the tailnet |
| Autoscaling | Node Auto-Provisioning + ComputeClass (ADR-0006) |
Zonal and spot are deliberate: this is a reference platform that is torn down after every run, so a regional control plane and on-demand nodes would buy availability nobody is using. It is not a production posture, and neither is the AWS side.
Two create-time settings that decide everything
The whole reason GKE can run this platform at all comes down to two fields in
opentofu/gcp/gke/init/main.tf, and both are create-time only — changing
either later replaces the cluster.
datapath_provider = "DATAPATH_PROVIDER_UNSPECIFIED"
network_policy = falseDATAPATH_PROVIDER_UNSPECIFIED selects GKE’s legacy datapath rather than
Dataplane V2. That reads like choosing the older option, and it is the
important one: Dataplane V2 is Cilium — Google’s build of it — and it ships
without the CiliumGatewayClassConfig CRD and the io.cilium/gateway-controller
this platform depends on. Adopting it would break both Tailscale gateways and
the Envoy access-log pipeline into VictoriaLogs. So the platform takes the
legacy datapath and installs its own Cilium on top.
network_policy = false turns off GKE’s own policy controller, because Cilium
is the policy engine. Leaving it on would put two enforcers in one datapath.
This pairing is why ADR-0005 rules out Autopilot entirely: Autopilot does not permit the privileged DaemonSet a self-managed CNI requires, so the choice is Standard or nothing.
Where GKE is simpler than EKS
Two of the most awkward parts of the AWS lane have no counterpart here.
No CNI to disable first. On EKS, Stage 2 patches the aws-node and
kube-proxy DaemonSets onto no nodes before Cilium can take over. On GKE,
Cilium’s cni.exclusive displaces GKE’s CNI config directly — /etc/cni/net.d
ends up holding only 05-cilium.conflist, with GKE’s renamed .cilium_bak —
and kubeProxyReplacement handles kube-proxy without touching a managed
DaemonSet the addon manager would revert anyway.
No WireGuard. On AWS, encryption.type: wireguard is load-bearing: it works
around an open Cilium bug that breaks the Gateway API L7 proxy under ENI prefix
delegation. GCP does not use prefix delegation and does not hit the bug, so the
tunnel overhead is simply not paid here. This was verified on a live cluster
rather than assumed.
Both differences were checked during the ADR-0005 validation, not inferred from documentation.
Identity, DNS and storage
These are the three places the GCP lane genuinely diverges from AWS rather than just renaming things. The full side-by-side is on Cloud support; the short version:
- Identity — GKE Workload Identity Federation, claimed through
GCPWorkloadIdentityrather thanEPI. No static service-account keys. - DNS — the private zone is Cloud DNS; the public zone is Route 53,
reached by federating an AWS IAM role onto a projected Kubernetes
ServiceAccount token (ADR-0019).
gcp-0therefore runs two external-dns instances, one per provider. - Storage — the default block class is
standard-rwo(pd-balanced, despite the name —standardis the HDD tier). Object storage is Cloud Storage, and model weights mount through the Cloud Storage FUSE CSI driver (ADR-0021).
What is excluded here, and why
gcp-0 reconciles the same layers as aws-0 — observability, tooling and
applications included. What it leaves out is deliberate: flux-previews
(previews belong to one cluster by nature). The
Cloud support page
has the full map of what runs where, and what closing each exclusion would take.