Skip to content
0046 · Frontier providers

Frontier models through Z.ai and keyless Anthropic per cloud (Bedrock on aws-0, Vertex on gcp-0)

Status: Accepted Date: 2026-09-25 Deciders: Smana (Platform Owner) Related Spec: SP4 — LLM complexity routing


Context

Agents need a capable model with zero GPUs, and hard chat prompts need one above the local 7–8B fleet. Until now the only frontier caller was RunLore, holding a Z.ai key in its own pod. The agent factory adds a data-class rule: public agent work may go to a SaaS model, while internal data (cluster reads, RunLore findings) may reach only EU-resident Anthropic or self-hosted models.

Decision Drivers

  • No provider key in any workload pod; keys live only where the gateways read them.
  • Internal data stays in EU regions.
  • Both client formats: OpenAI (OpenWebUI, OpenCode) and Anthropic (Claude Code).
  • Cost: GLM-5.2 is $1.40 / $4.40 per 1M tokens (GLM-5.3, which replaced it on 2026-09-27, lists at the same price); Claude Opus 5.5 is $4 / $20.

Considered Options

Option 1: Z.ai GLM + Anthropic on Bedrock (Pod Identity) / Vertex (Workload Identity)

Pros:

  • Bedrock needs no key: the data plane’s ServiceAccount assumes a role scoped to the eu.anthropic.* inference profiles.
  • Agent Router translates both OpenAI and Anthropic input to Bedrock’s AWSAnthropic schema.

Cons:

  • EU geo profiles route across EU regions (Frankfurt, Paris, Stockholm, Milan, Spain, Ireland), not Paris alone.
  • Bedrock is billed through AWS Marketplace, and model access needs a one-time subscription.

Option 2: A native Anthropic API key

Pros:

  • The simplest setup, and it gets new models first.

Cons:

  • Agent Router v1.1.0 has no OpenAI → native-Anthropic translator (PR #2127 is open), so OpenAI-format clients cannot reach it.
  • A long-lived key to hold and rotate.

Option 3: OpenRouter or another aggregator

Pros:

  • One key and many models.

Cons:

  • A third party sees every prompt, and data residency is the aggregator’s choice.
  • It adds a hop and a markup.

Option 4: Self-hosted only

Pros:

  • No data leaves the cluster.

Cons:

  • The fleet’s 7–8B models are not agent-capable. It also ties agents to GPU capacity, which the programme exists to avoid.

Option 5: One Anthropic provider for both clouds (Vertex-only)

Pros:

  • Unified provider: aws-0 and gcp-0 both call Vertex.

Cons:

  • aws-0’s internal data leaves AWS via cross-cloud identity federation (EKS token exchange to GCP). This breaks ADR-0007’s rule that each cloud uses its own native service.

Decision Outcome

Chosen option: “Option 1”

Rationale: It is the only option that serves both client formats with no key in a workload pod, and keeps internal data in the EU. Z.ai serves public work at roughly a third of Claude Opus 5.5’s list input price.

Consequences

Positive

  • Separate keys per Gateway split both spend and blast radius: the platform key under platform/llm/zai, and the agents’ own key, read only through their agents-secrets store, at zai on the dedicated agents kv-v2 mount. Until SP4 PR 6 moves the platform key to platform/llm/zai, it is read from platform/runlore/credentials, where it already lives, so no bootstrap has to copy it.
  • Bedrock credentials rotate themselves and cannot be exfiltrated as a string.

Negative

  • Z.ai’s processing and retention terms are unverified (research, open question 10). That is why only public data may reach it.
  • A Bedrock Marketplace subscription is an owner action per account.

Neutral

  • gcp-0 reaches the same Claude models through Vertex with Workload Identity (GCPAnthropic), in a follow-up.

Implementation Notes

StepScopeState
SP4 PR 1The platform Z.ai backend and tier-frontier on ai-gatewayBuilt
SP4 PR 2The Bedrock EPIs, claude-* on ai-gateway, and the agent tiers on agent-routerNot built
Follow-upVertex (GCPAnthropic) on gcp-0Not started

So the keyless Anthropic path does not exist yet on either cloud. Agents reach one model, agent-default → Z.ai GLM-5.3, on the public listener; the internal listener has no model backend, so internal work has no model to call until PR 2 or the Vertex follow-up lands.


References