AI Platform
Two halves, one platform. Serving runs open-weights models on the cluster’s own GPUs behind an OpenAI-compatible gateway. Agents put models to work: a labelled issue becomes a sandboxed agent run that opens a pull request, under its own identity and in a room humans can watch and steer.
Serving is off by default, and the agents are a work in progress. What runs where, what is proven live and what
is planned lives on one page: Status and roadmap.
Every other page here describes the design.
How the parts fit
Source: docs/architecture/ai-platform.drawio, page 1.
| Part | What it does | Page |
|---|---|---|
| Serving | vLLM, one Crossplane InferenceService claim per model, on Karpenter L4 GPUs, scaled by KEDA | Serving |
| Gateways | ai-gateway for humans and coding clients; the agent gateway for agent runs, with a per-run identity | Gateways |
| Coding clients | OpenCode, Continue and OpenWebUI pointed at ai-gateway | Coding clients |
| Agents | The factory, rooms and the sandboxed runtime | Agents |
| Observability | vLLM, KEDA and gateway metrics; per-run traces, step logs and dashboards | Observability |
Why self-host at all
A hosted API is cheaper, faster to adopt, and better at the frontier. This platform exists for the things a hosted API cannot give you:
- Prompts never leave the tailnet. Every request path here is private — there is no public endpoint, and no third party sees the code being completed. That is the whole reason the coding fleet exists.
- The model is pinned to a commit.
model.revisionis a HuggingFace commit SHA, so the model behind an endpoint cannot change under you. A hosted endpoint’s weights move when the provider decides. - Latency is a scheduling problem, not a queue you do not control. Tab-completion needs sub-200 ms; that is achievable when the replica is yours and always warm, and not negotiable with a shared API.
- It is a real workload for the platform to carry. GPUs, spot reclaim, saturation-based autoscaling, a shared RWX filesystem and an ext_proc-filtered gateway exercise parts of this platform that a stateless web app never touches.
And the honest side of the ledger: there is no scale-to-zero — see Autoscaling & GPUs for the four-GPU cost floor that implies, and why it is a deadlock rather than a missing feature.
Turning the opt-in serving platform on, what it deploys, and its security posture.
One model, one YAML file — a complete claim with its reasoning intact, what it renders, and every field it accepts.
Three KEDA triggers on leading vLLM signals, the scale-to-zero deadlock, the gpu-l4 NodePool and S3 Files weights.
ai-gateway for humans and coding clients, the agent gateway for agent runs: identity, routing and what they share.
Connecting OpenCode, Continue and OpenWebUI to the gateway — authentication, model IDs, and troubleshooting.
Autonomous coding agents that run sandboxed under their own identity, collaborate with humans in rooms, and ship small changes.
The gVisor sandbox, per-run identity, octo-sts and the rulesets that confine every run.
The append-only log of a task: live view, steering, room tools and approvals.
From a labelled issue to a merged PR: intake, triage, teams, revise, the merge gate and the kill switch.
How a developer gives work to agents, follows it, steers it and stops it.
Serving metrics and alerts, and per-run traces, step logs, gen_ai metrics and dashboards.
The one place state lives: what serves, what is built, deployed and proven live, and the roadmap.