Skip to content

Agent Factory

Work in progress. This section describes the target design. The runtime and identity layer, and the agent router’s per-run identity and model route, are built and proven live on aws-0; only the repository’s trust policies are on main. Rooms, the factory and the per-run observability are planned. Only these pages and the design documents are on main: the code merges once everything is built and its user experience is signed off after a live walkthrough.

What it is

The LLM platform serves models; the agent factory puts them to work. A maintainer labels an issue, and agents (an implementer, then a reviewer) work it on their own branch in a gVisor sandbox. They open a pull request, and the factory narrates each step on the issue. Humans steer through GitHub reviews, or by joining the agents’ room. Only policy-defined low-risk changes merge themselves; everything else waits for a human.

New here? The ideas in two minutes

TermIn plain words
AgentA language model in a loop: it reads a task, runs tools (shell, git, file edits, queries), looks at the result and decides the next step, until the task is done or its budget runs out
HarnessThe program that runs that loop inside the sandbox. Here it is OpenHands, wrapped by a small agent-run entrypoint that clones the repository, starts the conversation and prints a step log
RunOne agent, one role, one task, one branch, with a deadline. Declared as an AgentRun object in Kubernetes
RoleWhat a run is allowed to do. An implementer pushes to its own agent/<id> branch and opens a PR; a reviewer, tester or triager only reads and reports to the room, which posts a reviewer’s verdict on the PR
SandboxThe pod a run lives in, isolated by gVisor (a user-space kernel), with no long-lived credential and a network policy that denies everything not named
Run identityA token issued for that run only. Every call to a model or a tool carries it, and the run exchanges it for a short-lived GitHub token, so every action is attributed to the run that made it
agent-routerThe gateway every agent call goes through: it checks the run’s token, meters its tokens and routes to the model. Budgets are planned
Room(planned) A shared, append-only log of a task: what each agent did, what humans said, which approvals were given. Humans watch it live and can post into it
Factory(planned) The controller that turns a labelled issue into runs, narrates progress on the issue, enforces budgets and owns the kill switch
Merge gate(planned) The rule that decides which agent PRs may merge themselves: only low-risk classes, only with green CI

The agent in the loop is not a trusted component. Every control sits outside the sandbox:

  • network policy;
  • per-run identity;
  • scoped, short-lived GitHub tokens;
  • branch rulesets;
  • token budgets;
  • the merge gate.

Architecture

The Agent Factory on one page. Two triggers: an issue label or a PR review in the target GitHub repository, and a human with task agent:run (roomctl is planned). The planned factory’s Task controller snapshots, triages, queues and meters runs, narrates on the issue, opens a room on the planned room-broker (an append-only log on CNPG Postgres), and hands low-risk PRs to the merge gate, policy-bot with a merger GitHub App. The built runtime turns an AgentRun claim, through Crossplane, into a default-deny CiliumNetworkPolicy and a gVisor Sandbox pod holding the OpenHands harness and an Envoy identity-proxy. Model, tool and token-exchange calls leave through the proxy with a per-run JWT to the agent router, an Envoy AI Gateway that meters tokens per run and routes to the models (Z.ai GLM-5.3 today, Anthropic Claude planned), the MCP servers for Flux, VictoriaMetrics and VictoriaLogs, and octo-sts, which mints a token for the agents’ GitHub App that can push only to branches under agent/ in that repository; token budgets and model tiers are planned. Git pushes and the pull request go straight to GitHub with the installation token. The harness streams events to the room, the room posts the verdict comment back to GitHub, and the pod and the router send logs, metrics and spans to VictoriaLogs, VictoriaMetrics and VictoriaTraces, with a planned Grafana page per run

Source: docs/architecture/agent-factory.drawio.

Components and software

Everything that runs on the cluster is open source. The external services are GitHub and the model providers. Each group is listed with its status.

Runtime and identity: built, proven live

One AgentRun object becomes a fully isolated, fully attributed run. This part is proven end to end: an agent took issue #2112 to PR #2114, which was merged.

ComponentSoftwareWhat it doesWhy this software
Run APICrossplane v2 composition, written in KCLTurns one AgentRun claim into everything a run needs: ServiceAccount, task ConfigMap, network policy, Sandbox. Projects the run’s phase, PR and token usage back into its statusThe platform’s standard for self-service APIs; one claim, one lifecycle, deleted as a whole
Sandbox lifecycleagent-sandboxA Sandbox resource: one pod with a stable identity and a clean start, and no restarts that hide failuresKubernetes-native and built for agent workloads; the same building block as AWS’s agents-on-EKS blueprint
IsolationgVisor (runsc) on a dedicated Karpenter node poolRuns the agent’s commands against gVisor’s user-space kernel, so an exploit has to break gVisor before it reaches the node’s kernelStrong isolation without VMs, and it runs on ordinary EKS nodes (Kata would need bare metal or nested virtualisation)
HarnessOpenHands agent-server and SDK, wrapped by a small agent-run entrypointThe agent loop: shell, editor, git, MCP tools. agent-run clones the repository, starts the conversation, prints the step log and revokes the GitHub token at the endOpen source, headless (an HTTP API rather than an IDE), model-agnostic, with MCP support
Identity proxyEnvoy sidecarAttaches the run’s own short-lived token to every model, tool and token-exchange call. The harness never sees that tokenThe agent cannot leak a gateway token it never sees. The one credential it holds is its GitHub token: in memory, one repository, one role, ≤ 1 h, revoked when the run ends
Network policyCilium CiliumNetworkPolicyDefault deny, per run: egress only to named hosts (GitHub, the router, optional package registries)FQDN-aware policy, plus Hubble to see every dropped flow
GitHub accessocto-sts and a GitHub App, plus a repository rulesetExchanges the run’s identity for a GitHub token scoped to one repository and its role’s permissions, valid ≤ 1 h and revoked when the run ends. The ruleset lets the App push only agent/** branchesNo long-lived GitHub token anywhere; the rules live in each repository’s trust policies
SecretsOpenBao and External SecretsHolds the few platform secrets (App keys, provider keys); none reaches a sandboxThe platform’s secret store, nothing agent-specific

Agent router: identity and routing built; budgets and tiers planned

ComponentSoftwareWhat it doesWhy this software
GatewayAgent Router (formerly Envoy AI Gateway) on Envoy GatewayVerifies each run’s token (JWT), attributes and meters every request to its run, routes the model alias to a provider. (Planned) per-run and fleet token budgets, and routing by tierOne gateway for models, tools and token exchange, with per-run identity in every access-log line
ModelsZ.ai GLM-5.3 for public runs today; (planned) Anthropic Claude through Amazon Bedrock for internal runsThe providers the router sends model calls to. Agents ask for an alias, never for a providerSwapping or adding a provider changes the router, not the agents
Tool serversMCP servers for Flux Operator, VictoriaMetrics and VictoriaLogs, read-onlypublic runs get documentation tools only; cluster, metric and log reads are for internal runs (planned: they await a model route)Agents investigate with the data humans use, under the same identity checks

Rooms: planned

ComponentSoftwareWhat it doesWhy this software
Room brokerA small Go serviceKeeps each task’s append-only log (agent steps, human messages, handoffs, approvals), serves it live, and posts a reviewer’s verdict on the PRA purpose-built log: the room is the audit trail, so it must be append-only and attributed
Log storagePostgreSQL through CloudNativePG, with Valkey for fan-outDurable, append-only storage; Valkey tells every broker replica that there is something newThe platform’s standard database and key-value store
Web viewA small TypeScript UI behind oauth2-proxy and ZITADEL SSOWatch a room live, post a message for the next run, approve an actionSingle sign-on with the platform’s identity provider; no framework, strict content security policy
Room toolsMCP tools served by the brokerLet agents post, hand over to another role, or record a verdict in their roomAgents collaborate through the log, never by prompting each other
roomctlA CLIThe same room from a terminalFor people who live in the shell

Agent factory: planned

ComponentSoftwareWhat it doesWhy this software
Task controllerA Go controller (controller-runtime)Turns a labelled issue into a task: snapshot, triage, a room, a team of runs on one branch; narrates on the issue; turns “Request changes” into a new runThe only component that creates runs, so every run has a task and a budget
AdmissionKueueQueues sandboxes so a burst of tasks waits instead of overloading the node poolThe Kubernetes-native job queue, with quotas
Run meter and kill switchPart of the controllerRevokes a run that spends its token budget; one label on a pinned issue stops everythingControls that act from outside the sandbox
Merge gatepolicy-bot and a merger GitHub AppDecides which agent PRs may merge themselves (only low-risk classes, green CI), then arms GitHub’s auto-merge. A separate App holds that right, and only itThe policy lives in the repository and is reviewable; the right to merge is isolated from everything else
Admission policyKyvernoDenies AgentRun creation to anyone but the factoryOne path in, so no run escapes its budget

Observability: designed, next to build

ComponentSoftwareWhat it does
LogsVictoriaLogsEvery run’s step log and every gateway call, attributed to the run
MetricsVictoriaMetricsTokens, cost, latency and errors per run
TracesVictoriaTraces, fed by OpenTelemetryOne trace per run: steps, model calls and tool calls. Metadata only: no prompts or outputs
DashboardsGrafanaOne page per run and a fleet overview

One repository at first

The design targets one repository at first. Nothing in it is tied to that repository: the agents' App, the trust policies and the branch ruleset live in the repository they protect, and every token audience names its repository. A second repository is a planned extension, not a step you can take today. It would need:

  1. The agents’ GitHub App installed on it, and the factory’s App for narration.
  2. The octo-sts trust policies for the roles it allows, in that repository.
  3. The agent/** branch ruleset applied to it.
  4. Two platform changes: its four audiences added to the gateway’s token-exchange listener (at most eight per listener), and a factory intake that polls more than one repository.

Runs are per repository: a task never spans two.

Security boundaries

BoundaryMechanism
Code executiongVisor sandbox, restricted pod security, no service-account token in the harness
NetworkDefault-deny CNP per run; egress only to named FQDNs and the gateway
IdentityTwo projected tokens per run, one for the gateway and one for token exchange; each audience names the run’s role and its data class or repository; both live until the run’s deadline
GitHubShort-lived installation tokens from octo-sts, scoped to one repository and the role’s permissions; a ruleset lets the agents’ App push only agent/**
SpendPer-run deadline; token budgets at the gateway; the factory’s run meter enforces maxTokens
MergeOnly low-risk classes (docs-links, revert) auto-merge, through policy-bot and a separate merger App; everything else waits for a human
StopOne label on a pinned control issue stops intake, refuses new runs and revokes every running one

Design documents