Skip to content
0048 · Factory orchestrator

A custom Task controller with Kueue admission orchestrates the agent factory

Status: Accepted Date: 2026-09-27 Deciders: Smana Related Spec: SP3 dark factory


Context

The dark factory turns triggers (a maintainer’s issue label, RunLore findings, schedules) into sequential agent runs that open a pull request, then watches CI, policy and humans for hours until the PR merges, closes or is reverted. It must enforce three levels of token budget, honour a kill switch, and be the only creator of AgentRuns so that a principal’s daily budget has one enforcement point.


Decision Drivers

  • The lifecycle is a reconciliation against GitHub state that lasts hours
  • Budgets, caps and the kill switch are domain logic every option would still need
  • No new stateful service: the cluster is rebuilt routinely
  • Configuration in Git, state at runtime, audit through the room log that already records every run

Considered Options

Option 1: A custom controller and a Task CRD, Kueue for admission

Pros:

  • Waiting on GitHub is a requeue; state lives in the CR; no database
  • Budgets and the kill switch are first-class code, tested with envtest
  • Kueue caps and drains sandbox pods independently of the factory

Cons:

  • We own the code (one CRD, one config file, libraries)

Option 2: Argo Workflows

Pros:

  • DAGs, retries, semaphores, a UI and an archive

Cons:

  • Long GitHub waits need suspend/resume or polling pods; a UI to secure; no budget concept

Option 3: Tekton, Temporal, a Crossplane Task XR, or gh-aw

Pros:

  • Tekton has GitHub interceptors; Temporal has durable timers; an XR matches AgentRun; gh-aw has “safe outputs”

Cons:

  • Tekton is CI-shaped with no global queue; Temporal adds a server and a database; an XR has no timers, events or backoff; gh-aw runs on GitHub runners, outside the sandbox and identity design

Decision Outcome

Chosen option: “Option 1”

Rationale: Every alternative still needs the budget, cap and kill-switch logic written by us, and each adds a component or loses the reconciliation model. A controller-runtime controller reconciling Tasks, shipped as a signed chart from Smana/agent-platform beside the room broker, is the smallest thing that fits.


Consequences

Positive

  • One Deployment with leader election; Tasks are inspectable with kubectl
  • The factory is the single creator of runs, so the caps it checks before creating one hold. Its run meter revokes a run over its tier’s runTokens; a gateway-side per-run ceiling arrives with SP4 PR 2 (ADR-0050), not yet built

Negative

  • Our code is on the critical path. Mitigated by envtest suites and by kill-switch layers that do not depend on the factory: the GitHub App today, Kueue from SP3 phase 4

Neutral

  • Tasks are runtime objects, never in Git, so the validation catalog is unaffected

Implementation Notes

Code in Smana/agent-platform (the factory packages, the agent-factory binary and chart); manifests in tooling/base/agent-factory/, an agent-platform umbrella child on both clusters. Helm never upgrades a chart’s crds/, so the HelmRelease replaces the Task CRD on upgrade.

What phase 1 builds, against the decision above:

DecidedBuilt today
Kueue admits and drains sandbox podsNot deployed; it arrives with SP3 phase 4 (triage and teams). Until then the factory’s own caps bound concurrency: 3 active tasks, 4 concurrent runs, 20 tasks a day
Three levels of token budgetThe run budget only: the run meter revokes a run at its cap (budget-run). Task tokens are recorded, not enforced; the principal and fleet budgets follow
A kill switchThe agent-factory-stop ConfigMap stops everything; factory/stop stops one task

References