Skip to content

Gateways

Every model call on this platform crosses one of two gateways, chosen by who is calling. Both speak the OpenAI API; they differ in how they know the caller and in what they let it do. What runs today, and the planned move to agentgateway, is on the status page.

ai-gatewayAgent gateway
CallersHumans and coding clients: OpenCode, Continue, OpenWebUI, the nightly Promptfoo evalAgent runs, through each run’s identity-proxy sidecar
IdentityAn API key per client, from AWS Secrets ManagerA projected token per run (JWT), verified on every call
Model choiceNamed by the client, or picked from the prompt by the Semantic Router (model: MoM)An alias, mapped statically to one backend; nothing selects per request
Also carries—MCP tool calls, the room tools, and the token exchange (sts) with octo-sts
SoftwareEnvoy Gateway + Envoy AI Gateway (Agent Router) 1.1.0Agent Router 1.1.0 on Envoy Gateway; agentgateway selected to replace it

The two gateways side by side. Top: humans and coding clients reach ai-gateway over the tailnet with an API key; the Semantic Router may pick the model, and the request lands on a vLLM model in the serving fleet. Bottom: an agent run’s identity-proxy sidecar attaches the run’s own token; the agent gateway verifies it, meters the run’s tokens and routes the model alias to a frontier provider (Z.ai GLM, Claude), the MCP tools, or octo-sts for a GitHub token. Both gateways write access logs and gen_ai metrics to the Victoria stack

Source: docs/architecture/ai-platform.drawio, page 2.

ai-gateway: humans and coding clients

The platform speaks the OpenAI API. A client points at one endpoint, names a model — or asks the platform to choose one — and never learns which pod answered.

Sending a request

# What's available
curl -sS https://llm.priv.aws.ogenki.io/v1/models \
  -H "Authorization: Bearer $LLM_API_KEY" | jq '.data[].id'

# Name a model explicitly — skips classification entirely
curl -sS https://llm.priv.aws.ogenki.io/v1/chat/completions \
  -H "Authorization: Bearer $LLM_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
        "model": "xplane-qwen-coder",
        "messages": [{"role": "user", "content": "Write a Go worker pool."}]
      }'

# Let the Semantic Router choose from the prompt
curl -sS https://llm.priv.aws.ogenki.io/v1/chat/completions \
  -H "Authorization: Bearer $LLM_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
        "model": "MoM",
        "messages": [{"role": "user", "content": "Write a Go worker pool."}]
      }'

The two calls reach the same pod. The difference is roughly 250–300 ms of classification on the MoM path — worth it when the client cannot know which model suits the prompt, wasted when it can. See Coding Clients for which client pins which model, and why.

The request path

The hop-by-hop request path shows each stage below in one picture.

Ingress. External clients arrive over Tailscale at the Cilium Gateway platform-tailscale-general and are forwarded to the Envoy AI Gateway data plane Service. In-cluster clients — OpenWebUI, the nightly Promptfoo eval — address that Service directly. Nothing is reachable from the public internet; see Private access.

Authentication. A SecurityPolicy (apiKeyAuth) targets the Gateway, so every route inherits it rather than each declaring its own. Keys come from AWS Secrets Manager through External Secrets, and the gateway strips the Authorization header before forwarding — vLLM never sees a credential.

Prompt classification (ext_proc, filter index 0). The Semantic Router is wired in as an Envoy ext_proc gRPC filter, inserted ahead of the AI Gateway’s own extproc by an EnvoyPatchPolicy — a raw xDS JSONPatch. The ordering is not optional and an EnvoyExtensionPolicy cannot express it, because that API can only append filters. The router rewrites body.model only when the client sent model: MoM (or the literal auto); an explicit xplane-* name passes through untouched.

Routing. The AI Gateway extproc derives the x-ai-eg-model header from the (possibly rewritten) body and emits gen_ai_* telemetry. An AIGatewayRoute matches that header and forwards through an AIServiceBackend → Backend → the model’s Service on port 8000.

Not every claim routes this way yet: three of the four still use a hand-written route, listed under known gaps.

Semantic routing — model: MoM

Sending model: MoM lets the Semantic Router pick from the prompt. Its decision list, highest priority first (infrastructure/base/vllm-semantic-router/helmrelease.yaml):

PriorityDecisionTarget
110code + reasoningxplane-qwen-coder (use_reasoning: true)
100codexplane-qwen-coder
90reasoning (math / physics)xplane-qwen3-8b (use_reasoning: true)
80multilingualxplane-qwen3-8b
50general (the default)xplane-qwen3-8b
xplane-llamaguard3-1b and the two LoRA adapters on xplane-qwen-coder appear in no rule above. They hold serving capacity and are reachable only by naming them directly. The router’s own in-pod prompt_guard classifier blocks jailbreak attempts, but that is a filter, not a routing decision — there is no automatic guardrail dispatch (see known gaps).

LoRA canaries

xplane-qwen-coder carries two LoRA adapters, each addressable as a model name of its own, and sends 10% of its traffic to one of them:

  gateway:
    enabled: true
    canaries:
      - adapter: xplane-qwen-coder-sql-dpo
        weightPercent: 10

The adapter name is matched verbatim against loraAdapters[].name — it is not derived from the base model’s name, so a typo produces a route to nothing rather than a validation error.

Canaries are mutually exclusive with the Gateway API Inference Extension’s endpoint picker. See the roadmap for where the endpoint picker stands and what turning it on would take.

The agent gateway: agent runs

Every call an agent makes, to a model, a tool or the token exchange, goes through this gateway under the run’s own identity. The harness never sees that token: the run’s identity-proxy sidecar attaches it (see Agent runtime).

ComponentSoftwareWhat it doesWhy this software
GatewayAgent Router 1.1.0 (Envoy AI Gateway) on Envoy GatewayVerifies each run’s token (JWT), attributes and meters every request to its run, routes the model alias to a provider. Per-run and fleet token budgets, and routing by tierOne gateway for models, tools and token exchange, with per-run identity in every access-log line
ModelsZ.ai GLM-5.3 for public runs; Anthropic Claude for internal runs, through Amazon Bedrock on aws-0 and Vertex AI on gcp-0The providers the router sends model calls to. Agents ask for an alias, never for a providerSwapping or adding a provider changes the router, not the agents
Tool serversMCP servers for Flux Operator, VictoriaMetrics and VictoriaLogs, read-only, and the room-broker’s room_* toolspublic runs get documentation tools only; cluster, metric and log reads are for internal runs. The room tools are routed per roleAgents investigate with the data humans use, under the same identity checks

The token-exchange listener names each repository’s audiences, at most eight per listener; adding a repository is described on the agents overview.

Decided 2026-10-01: agentgateway was selected after its proof of concept on gcp-0 to replace Agent Router as the agents’ gateway (models, MCP and the sts listener). An ADR superseding programme ADR-0042 and ADR-0050’s Option 1 (on the programme branches, not yet on main) for the agent router follows. The ai-gateway stays on Envoy Gateway and Agent Router.

What the two share, and what they keep apart

From the model routing and budgets design (SP4):

ConcernDecisionWhy
Semantic RouterOn ai-gateway only; agents use their own GatewayThe router can be neither a single point of failure nor a prompt reader for agents
Agent model choiceEvery agent alias maps to one backend: no weights, canaries or fallbackNothing re-routes a trajectory mid-run; escalation is a new run at the next tier
BackendsOne set per gateway, each with its own provider keyRevoking the agents’ key never breaks chat or RunLore, and provider spend splits by key
BudgetsToken budgets at both gateways, in token units, through one global rate limit backed by ValkeyOne mechanism for every principal; the factory still revokes a run at its exact cap