Factory
The factory is the controller that turns a labelled issue into runs, narrates progress on the issue, meters each run’s tokens and owns the kill switch. It is the only component that creates runs, so every run has a task and a budget. This page describes the design; what runs today is on the status page.
Source: docs/architecture/ai-platform.drawio, page 5.
From issue to merge
| Step | What happens |
|---|---|
| Intake | A maintainer adds factory/ready to an issue. The factory takes a snapshot of it, so a later edit does not change a task already under way, and comments with the run id, branch, budget and a watch link. A task already running refuses a second label |
| Triage | Once per task: which class it is (for example docs-links), which team of roles works it, and how big a budget it gets. A refusal is explained in a comment. What triage predicts only chooses the team and the budget; risk is enforced on the diff at merge time |
| Teams | Roles run as sequential AgentRuns in one room, on one agent/<task> branch: an implementer, then a reviewer, with a tester in larger teams. Only the implementer writes |
| Revise | A GitHub review with “Request changes” becomes the input of a new run on the same branch. /factory retry tries again after a failure |
| Merge gate | Only low-risk classes, and only with green CI and the policy agreeing: docs-links and revert (the factory’s revert of an auto-merged docs-links change). Everything else waits for a human. Until everything is built and merged, the gate runs in shadow mode: it says what it would merge, and merges nothing |
| Kill switch | One task: the label factory/stop. Every task: the stop ConfigMap, or a label on a pinned control issue |
The developer’s view of the same journey is in the user guide.
Components and software
| Component | Software | What it does | Why this software |
|---|---|---|---|
| Task controller | A Go controller (controller-runtime) | Turns a labelled issue into a task: snapshot, a room, the team’s runs on its branch; narrates on the issue. Starts a reviewer after the implementer, and turns “Request changes” into a new run | The only component that creates runs, so every run has a task and a budget |
| Admission | Kueue | Queues sandboxes so a burst of tasks waits instead of overloading the node pool. Until it lands, the factory’s own caps bound concurrency | The Kubernetes-native job queue, with quotas |
| Run meter and kill switch | Part of the controller | The run meter revokes any run, hand-launched ones included, at its token cap. The agent-factory-stop ConfigMap pauses intake and stops every task; so does a label on a pinned control issue | Controls that act from outside the sandbox |
| Merge gate | policy-bot and a merger GitHub App | Decides which agent PRs may merge themselves (only low-risk classes, green CI), then arms GitHub’s auto-merge. A separate App holds that right, and only it | The policy lives in the repository and is reviewable; the right to merge is isolated from everything else |
| Admission policy | Kyverno | Denies AgentRun creation to anyone but the factory | One path in, so no run escapes its budget |
Controls
| Boundary | Mechanism |
|---|---|
| Spend | Per-run deadline; the factory’s run meter revokes a run at its token cap; token budgets at the gateway; at most 20 tasks a day |
| Merge | Only low-risk classes (docs-links, revert) auto-merge, through policy-bot and a separate merger App; everything else waits for a human |
| Stop | kubectl -n agent-system create configmap agent-factory-stop pauses intake and stops every task; one label on a pinned control issue also refuses new runs and revokes every running one |
The sandbox, identity and GitHub boundaries are on Runtime → Security boundaries. The factory design has the full merge policy, gate paths and the five-layer kill switch.