Skip to content

Serving

An OpenAI-compatible inference platform on EKS: vLLM on L4 spot GPUs, fronted by Envoy AI Gateway, scaled by KEDA on vLLM saturation signals, and declared as a single Crossplane InferenceService claim per model.

This platform is off by default. Two independent gates must both be released before anything LLM-related exists on the cluster — see Turning it on. A plain terramate script run deploy and a plain Flux reconciliation both leave the cluster LLM-free.

At a glance

EnginevLLM, one Deployment per model, port 8000
GatewayEnvoy Gateway + Envoy AI Gateway 1.1.0
RoutingAIGatewayRoute, keyed on the x-ai-eg-model header
Prompt routingvLLM Semantic Router, as a gRPC ext_proc filter — only acts on model: MoM
AutoscalingKEDA, three vLLM saturation triggers OR-combined, min=1 (always warm)
WeightsAmazon S3 Files (POSIX over S3), RWX PVC shared by a preload Job and the serving pod
GPUsKarpenter gpu-l4 NodePool — single-GPU g6 spot-first instances, Bottlerocket NVIDIA AMI, capped at 4 GPUs
Compositioncrossplane-inference-service KCL module 0.9.0, pinned inside crossplane-configuration-aws:v0.4.6

Turning it on

The two gates are deliberately independent, so neither one accidentally brings the other along:

# Gate 1 — AWS side (S3 Files filesystem + IAM). Terramate stack tagged
# `opt-in`; skipped unless TM_LLM_PLATFORM_ENABLED=true (verified in
# opentofu/aws/llm-platform/workflows.tm.hcl — unset or != "true" echoes [skip]
# and exits 0).
TM_LLM_PLATFORM_ENABLED=true terramate -C opentofu/aws/llm-platform script run deploy

# Gate 2 — Kubernetes side. The umbrella Flux Kustomization ships suspended
# (spec.suspend: true, clusters/aws-0/llm-platform.yaml).
flux resume kustomization llm-platform -n flux-system

The umbrella aggregates 8 child Flux Kustomizations under clusters/aws-0-llm-platform/:

ChildRendersPath
vllm-semantic-routerPrompt-classification router (MoM virtual model)infrastructure/base/vllm-semantic-router
runtimeclass-nvidiaRuntimeClass nvidiainfrastructure/base/runtimeclass-nvidia
llm-platform-gpu-nodepoolsKarpenter gpu-l4 NodePool + EC2NodeClassinfrastructure/base/karpenter-nodepools-gpu
envoy-gatewayEnvoy Gateway controllerinfrastructure/base/envoy-gateway
envoy-ai-gatewayEnvoy AI Gateway + the Semantic Router EnvoyPatchPolicyinfrastructure/base/envoy-ai-gateway
llm-platform-appsThe InferenceService claims + OpenWebUIapps/llm
llm-platform-security-epiThe preload Job’s EKS Pod Identitysecurity/base/epis-llm
llm-platform-promptfooNightly agent-eval CronJobtooling/base/promptfoo

That directory is a sibling of clusters/aws-0/, not a child, on purpose: flux-system syncs clusters/aws-0/ recursively, so a nested path would be auto-discovered and applied — bypassing the suspend gate entirely.

On gcp-0

gcp-0 has one gate, not two: the weights bucket is a Crossplane claim rather than an OpenTofu stack, so there is no TM_LLM_PLATFORM_ENABLED — the only gate is the umbrella Kustomization clusters/gcp-0/llm-platform.yaml (spec.suspend: true). Weights are served from a GCS bucket over the Cloud Storage FUSE CSI driver instead of an S3 Files POSIX mount — see ADR-0021 for why, including what it gives up. Why the umbrella is still suspended, and what the first resume proved, is on the status page.

Security posture

  • Zero trust by default. Every workload carries a default-deny CiliumNetworkPolicy. The serving pod’s egress is kube-dns only; only the bounded preload Job is granted world:443, because it is short-lived and the serving pod cannot reach HuggingFace even in principle.
  • No credentials in Git. API keys and the HuggingFace token come from AWS Secrets Manager through External Secrets — see PKI & Secrets.
  • Read-only IAM on the serving pod. Each claim’s serving ServiceAccount carries a per-claim EKS Pod Identity scoped to read its own weights prefix, rendered by the composition; only the shared preload Job’s identity can write to the bucket.
  • Private ingress only. Reachable exclusively from the tailnet — see Private Access.

Known gaps live on the status page.