Serving
An OpenAI-compatible inference platform on EKS: vLLM on L4 spot GPUs, fronted
by Envoy AI Gateway, scaled by KEDA on vLLM saturation signals, and declared
as a single Crossplane
InferenceService claim
per model.
terramate script run deploy and a
plain Flux reconciliation both leave the cluster LLM-free.At a glance
| Engine | vLLM, one Deployment per model, port 8000 |
| Gateway | Envoy Gateway + Envoy AI Gateway 1.1.0 |
| Routing | AIGatewayRoute, keyed on the x-ai-eg-model header |
| Prompt routing | vLLM Semantic Router, as a gRPC ext_proc filter — only acts on model: MoM |
| Autoscaling | KEDA, three vLLM saturation triggers OR-combined, min=1 (always warm) |
| Weights | Amazon S3 Files (POSIX over S3), RWX PVC shared by a preload Job and the serving pod |
| GPUs | Karpenter gpu-l4 NodePool — single-GPU g6 spot-first instances, Bottlerocket NVIDIA AMI, capped at 4 GPUs |
| Composition | crossplane-inference-service KCL module 0.9.0, pinned inside crossplane-configuration-aws:v0.4.6 |
Turning it on
The two gates are deliberately independent, so neither one accidentally brings the other along:
# Gate 1 — AWS side (S3 Files filesystem + IAM). Terramate stack tagged
# `opt-in`; skipped unless TM_LLM_PLATFORM_ENABLED=true (verified in
# opentofu/aws/llm-platform/workflows.tm.hcl — unset or != "true" echoes [skip]
# and exits 0).
TM_LLM_PLATFORM_ENABLED=true terramate -C opentofu/aws/llm-platform script run deploy
# Gate 2 — Kubernetes side. The umbrella Flux Kustomization ships suspended
# (spec.suspend: true, clusters/aws-0/llm-platform.yaml).
flux resume kustomization llm-platform -n flux-systemThe umbrella aggregates 8 child Flux Kustomizations under
clusters/aws-0-llm-platform/:
| Child | Renders | Path |
|---|---|---|
vllm-semantic-router | Prompt-classification router (MoM virtual model) | infrastructure/base/vllm-semantic-router |
runtimeclass-nvidia | RuntimeClass nvidia | infrastructure/base/runtimeclass-nvidia |
llm-platform-gpu-nodepools | Karpenter gpu-l4 NodePool + EC2NodeClass | infrastructure/base/karpenter-nodepools-gpu |
envoy-gateway | Envoy Gateway controller | infrastructure/base/envoy-gateway |
envoy-ai-gateway | Envoy AI Gateway + the Semantic Router EnvoyPatchPolicy | infrastructure/base/envoy-ai-gateway |
llm-platform-apps | The InferenceService claims + OpenWebUI | apps/llm |
llm-platform-security-epi | The preload Job’s EKS Pod Identity | security/base/epis-llm |
llm-platform-promptfoo | Nightly agent-eval CronJob | tooling/base/promptfoo |
That directory is a sibling of clusters/aws-0/, not a child, on
purpose: flux-system syncs clusters/aws-0/ recursively, so a nested
path would be auto-discovered and applied — bypassing the suspend gate
entirely.
On gcp-0
gcp-0 has one gate, not two: the weights bucket is a Crossplane claim
rather than an OpenTofu stack, so there is no TM_LLM_PLATFORM_ENABLED —
the only gate is the umbrella Kustomization clusters/gcp-0/llm-platform.yaml
(spec.suspend: true). Weights are served from a GCS bucket over the Cloud
Storage FUSE CSI driver instead of an S3 Files POSIX mount — see
ADR-0021
for why, including what it gives up. Why the umbrella is still suspended, and
what the first resume proved, is on the
status page.
Security posture
- Zero trust by default. Every workload carries a default-deny
CiliumNetworkPolicy. The serving pod’s egress is kube-dns only; only the bounded preload Job is grantedworld:443, because it is short-lived and the serving pod cannot reach HuggingFace even in principle. - No credentials in Git. API keys and the HuggingFace token come from AWS Secrets Manager through External Secrets — see PKI & Secrets.
- Read-only IAM on the serving pod. Each claim’s serving ServiceAccount carries a per-claim EKS Pod Identity scoped to read its own weights prefix, rendered by the composition; only the shared preload Job’s identity can write to the bucket.
- Private ingress only. Reachable exclusively from the tailnet — see Private Access.
Known gaps live on the status page.