Agents
gcp-0 today, what is
proven live and what is planned is on the
status page.What it is
The LLM platform serves models; the agent factory puts them to work. A maintainer labels an issue, and agents (an implementer, then a reviewer) work it on their own branch in a gVisor sandbox. They open a pull request, and the factory narrates each step on the issue. Humans steer through GitHub reviews, or by joining the agents’ room. Only policy-defined low-risk changes merge themselves; everything else waits for a human.
New here? The ideas in two minutes
| Term | In plain words |
|---|---|
| Agent | A language model in a loop: it reads a task, runs tools (shell, git, file edits, queries), looks at the result and decides the next step, until the task is done or its budget runs out |
| Harness | The program that runs that loop inside the sandbox. Here it is OpenHands, wrapped by a small agent-run entrypoint that clones the repository, starts the conversation and prints a step log |
| Run | One agent, one role, one task, one branch, with a deadline. Declared as an AgentRun object in Kubernetes |
| Role | What a run is allowed to do. An implementer pushes to its own agent/<id> branch and opens a PR; a reviewer, tester or triager only reads and reports to the room, which posts a reviewer’s verdict on the PR |
| Sandbox | The pod a run lives in, isolated by gVisor (a user-space kernel), with no long-lived credential and a network policy that denies everything not named |
| Run identity | A token issued for that run only. Every call to a model or a tool carries it, and the run exchanges it for a short-lived GitHub token, so every action is attributed to the run that made it |
| agent-router | The gateway every agent call goes through: it checks the run’s token, meters its tokens and routes to the model |
| Room | A shared, append-only log of a task: what each agent did, what humans said, the handoffs between roles. Humans watch it live, post into it, steer the running agent and approve its actions |
| Factory | The controller that turns a labelled issue into runs, narrates progress on the issue, meters each run’s tokens and owns the kill switch |
| Merge gate | The rule that decides which agent PRs may merge themselves: only low-risk classes, only with green CI |
The agent in the loop is not a trusted component. Every control sits outside the sandbox:
- network policy;
- per-run identity;
- scoped, short-lived GitHub tokens;
- branch and tag rulesets;
- token budgets;
- the merge gate.
Architecture
At a glance
One task, end to end. A maintainer labels an issue; the factory starts a sandboxed agent run and
opens a room; every call the agent makes goes through the gateway under the run’s own identity;
the agent pushes an agent/** branch and opens a PR; a human review decides the merge.
Source: docs/architecture/agent-factory-overview.drawio.
In detail
The diagram below shows the target architecture: the whole programme once built. Its legend marks
each box as deployed on gcp-0 (noting where its live gate is pending), built but not yet
deployed, or planned.
Source: docs/architecture/agent-factory.drawio.
The parts
| Part | What it covers |
|---|---|
| Runtime | The gVisor sandbox on each cloud, per-run identity, octo-sts, the branch and tag rulesets, network policy |
| Rooms | The broker’s append-only log, the bridge, the web view, steering, room tools, approvals |
| Factory | Intake, triage, teams, revise, the merge gate and the kill switch |
| Gateways | The agent gateway: per-run identity, models, MCP tools, token exchange |
| Observability | Per-run traces, step logs, gen_ai metrics, AgentRun state, dashboards |
| User guide | What a developer does with it |
Everything that runs on the cluster is open source. The external services are GitHub and the model providers.
One repository at first
The design targets one repository at first. Nothing in it is tied to that repository: the agents' App, the trust policies and the rulesets live in the repository they protect, and every token audience names its repository. A second repository is a planned extension, not a step you can take today. It would need:
- The agents’ GitHub App installed on it, and the factory’s App for narration.
- The octo-sts trust policies for the roles it allows, in that repository.
- The
agent/**branch ruleset and the tag ruleset applied to it. - Two platform changes: its four audiences added to the gateway’s token-exchange listener (at most eight per listener), and a factory intake that polls more than one repository.
Runs are per repository: a task never spans two.
Design documents
- Programme design: the contracts between the sub-projects, and the owner decisions.
- Runtime and identity · Rooms · Factory · Model routing and budgets · Observability
- User guide: what a developer does with it.
- Status: what is built, reviewed and proven live, and what waits on the owner.