Gateways
Every model call on this platform crosses one of two gateways, chosen by who is calling. Both speak the OpenAI API; they differ in how they know the caller and in what they let it do. What runs today, and the planned move to agentgateway, is on the status page.
ai-gateway | Agent gateway | |
|---|---|---|
| Callers | Humans and coding clients: OpenCode, Continue, OpenWebUI, the nightly Promptfoo eval | Agent runs, through each run’s identity-proxy sidecar |
| Identity | An API key per client, from AWS Secrets Manager | A projected token per run (JWT), verified on every call |
| Model choice | Named by the client, or picked from the prompt by the Semantic Router (model: MoM) | An alias, mapped statically to one backend; nothing selects per request |
| Also carries | — | MCP tool calls, the room tools, and the token exchange (sts) with octo-sts |
| Software | Envoy Gateway + Envoy AI Gateway (Agent Router) 1.1.0 | Agent Router 1.1.0 on Envoy Gateway; agentgateway selected to replace it |
Source: docs/architecture/ai-platform.drawio, page 2.
ai-gateway: humans and coding clients
The platform speaks the OpenAI API. A client points at one endpoint, names a model — or asks the platform to choose one — and never learns which pod answered.
Sending a request
# What's available
curl -sS https://llm.priv.aws.ogenki.io/v1/models \
-H "Authorization: Bearer $LLM_API_KEY" | jq '.data[].id'
# Name a model explicitly — skips classification entirely
curl -sS https://llm.priv.aws.ogenki.io/v1/chat/completions \
-H "Authorization: Bearer $LLM_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "xplane-qwen-coder",
"messages": [{"role": "user", "content": "Write a Go worker pool."}]
}'
# Let the Semantic Router choose from the prompt
curl -sS https://llm.priv.aws.ogenki.io/v1/chat/completions \
-H "Authorization: Bearer $LLM_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "MoM",
"messages": [{"role": "user", "content": "Write a Go worker pool."}]
}'The two calls reach the same pod. The difference is roughly 250–300 ms of
classification on the MoM path — worth it when the client cannot know which
model suits the prompt, wasted when it can. See
Coding Clients
for which client pins which model, and why.
The request path
The hop-by-hop request path shows each stage below in one picture.
Ingress. External clients arrive over Tailscale at the Cilium Gateway
platform-tailscale-general and are forwarded to the Envoy AI Gateway data
plane Service. In-cluster clients — OpenWebUI, the nightly Promptfoo eval —
address that Service directly. Nothing is reachable from the public internet;
see Private access.
Authentication. A SecurityPolicy (apiKeyAuth) targets the Gateway, so
every route inherits it rather than each declaring its own. Keys come from AWS
Secrets Manager through External Secrets, and the gateway strips the
Authorization header before forwarding — vLLM never sees a credential.
Prompt classification (ext_proc, filter index 0). The Semantic Router is
wired in as an Envoy ext_proc gRPC filter, inserted ahead of the AI
Gateway’s own extproc by an EnvoyPatchPolicy — a raw xDS JSONPatch. The
ordering is not optional and an EnvoyExtensionPolicy cannot express it,
because that API can only append filters. The router rewrites body.model
only when the client sent model: MoM (or the literal auto); an explicit
xplane-* name passes through untouched.
Routing. The AI Gateway extproc derives the x-ai-eg-model header from the
(possibly rewritten) body and emits gen_ai_* telemetry. An AIGatewayRoute
matches that header and forwards through an AIServiceBackend → Backend →
the model’s Service on port 8000.
Not every claim routes this way yet: three of the four still use a hand-written route, listed under known gaps.
Semantic routing — model: MoM
Sending model: MoM lets the Semantic Router pick from the prompt. Its
decision list, highest priority first
(infrastructure/base/vllm-semantic-router/helmrelease.yaml):
| Priority | Decision | Target |
|---|---|---|
| 110 | code + reasoning | xplane-qwen-coder (use_reasoning: true) |
| 100 | code | xplane-qwen-coder |
| 90 | reasoning (math / physics) | xplane-qwen3-8b (use_reasoning: true) |
| 80 | multilingual | xplane-qwen3-8b |
| 50 | general (the default) | xplane-qwen3-8b |
xplane-llamaguard3-1b and the two LoRA adapters on xplane-qwen-coder
appear in no rule above. They hold serving capacity and are reachable only
by naming them directly. The router’s own in-pod prompt_guard classifier
blocks jailbreak attempts, but that is a filter, not a routing decision —
there is no automatic guardrail dispatch (see
known gaps).LoRA canaries
xplane-qwen-coder carries two LoRA adapters, each addressable as a model name
of its own, and sends 10% of its traffic to one of them:
gateway:
enabled: true
canaries:
- adapter: xplane-qwen-coder-sql-dpo
weightPercent: 10The adapter name is matched verbatim against loraAdapters[].name — it is
not derived from the base model’s name, so a typo produces a route to nothing
rather than a validation error.
Canaries are mutually exclusive with the Gateway API Inference Extension’s endpoint picker. See the roadmap for where the endpoint picker stands and what turning it on would take.
The agent gateway: agent runs
Every call an agent makes, to a model, a tool or the token exchange, goes through this gateway under the run’s own identity. The harness never sees that token: the run’s identity-proxy sidecar attaches it (see Agent runtime).
| Component | Software | What it does | Why this software |
|---|---|---|---|
| Gateway | Agent Router 1.1.0 (Envoy AI Gateway) on Envoy Gateway | Verifies each run’s token (JWT), attributes and meters every request to its run, routes the model alias to a provider. Per-run and fleet token budgets, and routing by tier | One gateway for models, tools and token exchange, with per-run identity in every access-log line |
| Models | Z.ai GLM-5.3 for public runs; Anthropic Claude for internal runs, through Amazon Bedrock on aws-0 and Vertex AI on gcp-0 | The providers the router sends model calls to. Agents ask for an alias, never for a provider | Swapping or adding a provider changes the router, not the agents |
| Tool servers | MCP servers for Flux Operator, VictoriaMetrics and VictoriaLogs, read-only, and the room-broker’s room_* tools | public runs get documentation tools only; cluster, metric and log reads are for internal runs. The room tools are routed per role | Agents investigate with the data humans use, under the same identity checks |
The token-exchange listener names each repository’s audiences, at most eight per listener; adding a repository is described on the agents overview.
Decided 2026-10-01: agentgateway was selected after its proof of
concept on gcp-0 to replace Agent Router as the agents’ gateway (models, MCP and the sts
listener). An ADR superseding programme ADR-0042 and ADR-0050’s Option 1 (on the programme
branches, not yet on main) for the agent router follows. The ai-gateway stays on Envoy Gateway
and Agent Router.
What the two share, and what they keep apart
From the model routing and budgets design (SP4):
| Concern | Decision | Why |
|---|---|---|
| Semantic Router | On ai-gateway only; agents use their own Gateway | The router can be neither a single point of failure nor a prompt reader for agents |
| Agent model choice | Every agent alias maps to one backend: no weights, canaries or fallback | Nothing re-routes a trajectory mid-run; escalation is a new run at the next tier |
| Backends | One set per gateway, each with its own provider key | Revoking the agents’ key never breaks chat or RunLore, and provider spend splits by key |
| Budgets | Token budgets at both gateways, in token units, through one global rate limit backed by Valkey | One mechanism for every principal; the factory still revokes a run at its exact cap |