Skip to content

AI Platform

Two halves, one platform. Serving runs open-weights models on the cluster’s own GPUs behind an OpenAI-compatible gateway. Agents put models to work: a labelled issue becomes a sandboxed agent run that opens a pull request, under its own identity and in a room humans can watch and steer.

Serving is off by default, and the agents are a work in progress. What runs where, what is proven live and what is planned lives on one page: Status and roadmap. Every other page here describes the design.

How the parts fit

The AI Platform in one view. On the serving side, developers and coding clients send OpenAI-compatible requests over the tailnet to ai-gateway (Envoy AI Gateway with an API key and the Semantic Router), which routes them to vLLM models declared as InferenceService claims, on L4 GPUs and scaled by KEDA. On the agents side, a maintainer labels a GitHub issue; the Agent Factory turns it into a task, starts an agent run in a gVisor sandbox and opens a room the maintainer watches and steers. Every call the agent makes goes through the agent gateway with a per-run identity, to frontier models, MCP tools and octo-sts for a GitHub token; the agent pushes a branch and opens a PR. Both halves send traces, logs and metrics to the Victoria stack and Grafana

Source: docs/architecture/ai-platform.drawio, page 1.

PartWhat it doesPage
ServingvLLM, one Crossplane InferenceService claim per model, on Karpenter L4 GPUs, scaled by KEDAServing
Gatewaysai-gateway for humans and coding clients; the agent gateway for agent runs, with a per-run identityGateways
Coding clientsOpenCode, Continue and OpenWebUI pointed at ai-gatewayCoding clients
AgentsThe factory, rooms and the sandboxed runtimeAgents
ObservabilityvLLM, KEDA and gateway metrics; per-run traces, step logs and dashboardsObservability

Why self-host at all

A hosted API is cheaper, faster to adopt, and better at the frontier. This platform exists for the things a hosted API cannot give you:

  • Prompts never leave the tailnet. Every request path here is private — there is no public endpoint, and no third party sees the code being completed. That is the whole reason the coding fleet exists.
  • The model is pinned to a commit. model.revision is a HuggingFace commit SHA, so the model behind an endpoint cannot change under you. A hosted endpoint’s weights move when the provider decides.
  • Latency is a scheduling problem, not a queue you do not control. Tab-completion needs sub-200 ms; that is achievable when the replica is yours and always warm, and not negotiable with a shared API.
  • It is a real workload for the platform to carry. GPUs, spot reclaim, saturation-based autoscaling, a shared RWX filesystem and an ext_proc-filtered gateway exercise parts of this platform that a stateless web app never touches.

And the honest side of the ledger: there is no scale-to-zero — see Autoscaling & GPUs for the four-GPU cost floor that implies, and why it is a deadlock rather than a missing feature.