Skip to content

Teardown

Destroying by hand, stack by stack, is easy to get wrong: the EKS cluster holds resources (Gateways, PVC-backed volumes, IAM access keys) that need cleaning up in a specific order before OpenTofu can even start deleting.

EKS only

This is the safe, documented path — always tear the cluster down this way:

cd opentofu/aws/eks/init
terramate script run destroy

Five steps, defined in opentofu/aws/eks/init/workflows.tm.hcl:

  1. prepare-destroy — runs scripts/ops/aws/eks-prepare-destroy.sh (see below).
  2. stage2-destroy-addons — attempts to destroy the eks/configure stack (Cilium, Flux) via scripts/ops/teardown/destroy-stage2.sh in its attempt mode. Never fatal: everything that stack manages lives inside the cluster step 3 deletes anyway, so a failure here must not strand the one billable resource (why).
  3. stage1-destroy-cluster — destroys the eks/init stack (the cluster itself).
  4. stage3-sweep-orphaned-volumes — deletes the EBS volumes that were still detaching when step 1 swept (why).
  5. stage4-reconcile-state — drops any stage-2 state entries left behind, now that the cluster holding those objects is provably gone.

Full teardown

Tears down every stack — EKS, OpenBao, Network — in one command:

cd opentofu
terramate script run --reverse destroy          # aws (the default)
TM_CLOUD=all terramate script run --reverse destroy   # both clouds

Reverse dependency order, with a single confirmation prompt (scripts/ops/teardown/terramate-destroy-confirm.sh) cached for 10 minutes so the whole sweep only asks once. TM_DESTROY_CONFIRMED=true skips it for CI.

Why eks/configure shows [skip]. It is a registered stack (after = ["/opentofu/aws/eks/init"]), so a reverse walk reaches it before eks/init — the opposite of the order a cluster needs. It used to destroy itself there, which brought Cilium and Flux down raw: Flux never suspended, admission webhooks left admitting, and PVCs, NodePools and IAM access keys never cleaned up. eks/init’s prepare-destroy then ran against a cluster whose networking was already gone, and this page carried a warning telling people not to use the command that ought to work.

Its destroy script is now a no-op that says so. The stack is still destroyed — by eks/init’s stage2-destroy-addons job, at the point in the sequence where it is safe. Ownership of the ordering lives in one place because the stack graph cannot express it.

What eks-prepare-destroy.sh does first

Before OpenTofu deletes anything, the script:

  • Suspends every Flux Kustomization.
  • Disables Kyverno’s and the Cilium operator’s blocking admission webhooks — once their pods are evicted with the nodes, every subsequent delete would otherwise fail against a webhook with no live endpoint.
  • Reclaims CSI-provisioned EBS volumes, by calling scripts/ops/k8s/reclaim-csi-volumes.sh — the same script the GKE teardown calls, since every step of it is plain Kubernetes. It patches every PV’s persistentVolumeReclaimPolicy to Delete — including PVs deliberately set to Retain — deletes CloudNativePG Cluster resources so the operator releases their PVCs cleanly, scales down every Deployment/StatefulSet that mounts a PVC, deletes remaining PVC-mounting pods, then runs kubectl delete pvc --all --all-namespaces and waits up to 300s for reclaim. This step is unconditional — nothing in the script gates it, and it deletes PVC data regardless of the reclaim policy a PV was created with. If you need to keep data, back it up out of band before running eks-prepare-destroy.sh; there is no flag that skips this step.
  • Separately, sweeps EBS volumes orphaned by earlier teardown runs — volumes in available state, tagged for this cluster, whose PV no longer exists. EKS_DESTROY_KEEP_VOLUMES=true skips only this sweep of already-orphaned volumes; it has no effect on the PV/PVC reclaim above. This sweep runs before the destroy, so it can only see what has finished detaching by then — see the second sweep for the rest.
  • Reclaims the IAM access keys Harbor’s S3 registry storage uses — AWS caps AccessKeysPerUser at 2, so without this a second or third rebuild’s Harbor can fail to start on a full quota.
  • Deletes Karpenter NodePools, Gateway API resources (HTTPRoutes → Gateways → GatewayClasses, with finalizers stripped once controllers are gone), Envoy Gateway / AI Gateway extension resources, InferencePool resources, and EKS Pod Identity associations.
  • Strips finalizers from any Crossplane composite resource stuck terminating, so the namespace delete that follows doesn’t hang forever.

Stage 2 never gates the cluster

The eks/configure stack manages Cilium, the Flux Operator and the Flux Instance — all of them objects inside the cluster that stage 1 deletes moments later. Its teardown is therefore tidiness, never a prerequisite, and scripts/ops/teardown/destroy-stage2.sh enforces that: attempt reports a failure and exits 0.

Both clouds proved why the hard version is wrong:

  • 2026-08-23, GKE — the helm and kubectl providers held an access token acquired at plan time. It expired mid-destroy, stage 2 failed Unauthorized, and the cluster and both nodes were still RUNNING afterwards. Finished by hand, which diverged state and needed tofu state rm.
  • 2026-08-29, EKS — a teardown reached stage 2 with the cluster already deleted. Every in-cluster delete returned the server has asked for the client to provide credentials, against an endpoint that no longer resolved in DNS. Errors about objects that had ceased to exist, failing the run.

The usual reason to be running a destroy at all is that the cluster is broken, or its private endpoint is unreachable because the tailnet is down — exactly when stage 2 cannot succeed and exactly when you most need the cluster gone.

stage4-reconcile-state cleans up afterwards. It runs last, once stage 1 has provably deleted the cluster: anything still in stage 2’s state describes an object that lived there, so it cannot exist any more. Clearing state earlier would be unsafe in the other direction — if the cluster destroy then failed, state would have been emptied for resources that still exist.

The sweep before the destroy

terramate script run destroy opens with stage0-sweep-teardown-blockers, which runs scripts/ops/aws/sweep-teardown-blockers.sh. It clears the two things that make tofu destroy fail, neither of which Terraform owns:

  • ExternalDNS records. Route53 refuses DeleteHostedZone while any record other than the zone’s own NS/SOA remains. ExternalDNS writes an A record and a TXT ownership record per exposed service, and once the cluster is gone nothing reclaims them.
  • The EKS-managed cluster security group (eks-cluster-sg-<name>-*). EKS creates it, Terraform never owned it, and it outlives the cluster — so DeleteVpc fails with DependencyViolation naming a group that appears nowhere in the state file.

It runs before the destroy, the opposite of the volume sweep below, for the opposite reason: these two block the destroy, so a sweep running afterwards would never be reached.

The cost of skipping this is larger than it looks. terramate script run --reverse destroy stops at the first failing stack. On 2026-09-02 the AWS stacks failed on exactly these blockers, so the sweep never reached the GCP stacks at all — a GKE cluster kept running because a leftover DNS record two stacks away blocked a hosted zone. The teardown reported failure, and anyone reading only the exit code would still have been billed for a whole second cloud.

The script refuses to touch anything shared: NS and SOA are never deleted, the default security group is never deleted, the group is matched by the EKS-generated name for the cluster you name, and a group still holding network interfaces is reported and skipped — because that means the cluster is not actually gone.

The sweep after the destroy

terramate script run destroy ends with stage3-sweep-orphaned-volumes, which runs scripts/ops/aws/sweep-orphaned-volumes.sh once the cluster is gone.

It exists because the pre-destroy sweep above runs at the wrong moment to be complete. It fires moments after the PVCs are deleted, so a volume still detaching is not yet available and is skipped — and the script says so: “may still be detaching — the next run retries”. The next run is the next teardown, which is a whole rebuild away, so a volume that detached a second too late bills for the entire gap, and forever if the cluster is never rebuilt. That is the shape of the accumulation: 62 volumes (~518 GiB) by 2026-07, each teardown leaving a few for the following one to find.

After tofu destroy returns, every node is terminated, so every volume of this cluster is unambiguously detached and there is no in-flight state left to race. The same three filters apply — available, tagged kubernetes.io/cluster/<name>=owned, and tagged kubernetes.io/created-for/pvc/name — so it can neither touch a second live cluster’s storage nor a hand-made volume.

Run it by hand if you tore the cluster down some other way. It is a dry run unless you pass --apply:

./scripts/ops/aws/sweep-orphaned-volumes.sh --cluster-name aws-0 --region eu-west-3

GCP has the same step as stage2-sweep-orphaned-disks — see the GCP teardown.

Verify against AWS, not against the exit code

A Terramate destroy can exit 0 having destroyed nothing. The safeguards refuse silently, a job that fails before tofu runs still ends the run, and a backgrounded invocation reports the status of whatever came last in the pipeline. The GCP side says the same for the same reason. Ask the provider:

aws eks list-clusters --region eu-west-3
aws ec2 describe-instances --region eu-west-3 \
  --filters Name=instance-state-name,Values=running \
  --query 'Reservations[].Instances[].[InstanceId,Tags[?Key==`Name`].Value|[0]]' --output text
aws ec2 describe-volumes --region eu-west-3 \
  --filters Name=status,Values=available --query 'Volumes[].[VolumeId,Size]' --output text
aws elbv2 describe-load-balancers --region eu-west-3 --query 'LoadBalancers[].LoadBalancerName' --output text
aws ec2 describe-nat-gateways --region eu-west-3 \
  --filter Name=state,Values=available --query 'NatGateways[].NatGatewayId' --output text
aws ec2 describe-addresses --region eu-west-3 --query 'Addresses[].PublicIp' --output text

Empty output from all six is the only thing that means the platform is gone. NAT gateways and unattached Elastic IPs bill by the hour whether or not anything uses them, and available volumes bill by the GiB-month — those three are what an “successful” teardown most often leaves behind.

What is not deleted

The platform constitution withholds delete permissions from Crossplane for stateful services — so xplane-* IAM roles, policies, and S3 buckets outlive the cluster. They cost nothing to leave behind and are re-adopted by name on the next deploy. To remove them by hand once you are sure you are done:

aws iam list-roles --query 'Roles[?starts_with(RoleName, `xplane-`)].RoleName' --output text

Non-interactive

Both destroy scripts accept TM_DESTROY_CONFIRMED=true to skip the interactive y/n prompt. It exists for CI; skip it on a first manual teardown so you get the confirmation.