Alternatives considered
What alternatives did we consider for each layer of the agent factory, and why did the current choices win? A new candidate gets a row in the table. If it could replace a whole layer, it also gets a section under projects evaluated in detail.
In short:
- No candidate evaluated so far covers more than one or two layers. The factory, the rooms and the controls around them would stay ours whatever we adopted.
- A choice can be deferred rather than refused. Each detailed evaluation lists the concrete triggers that would make us reconsider it.
How candidates are evaluated
- Against code, not announcements. Claims are checked in the project’s repository at a pinned commit, and linked.
- On the work still ahead. What is already built does not count in its favour; what adopting a candidate would still leave us to build does.
- Against the platform’s requirements:
- it runs on both EKS and GKE, on spot nodes;
- every workload gets its own identity and network policy;
- no long-lived credential reaches a sandbox;
- open source first.
Where each tool fits
An agent platform has several layers. Most “agent platforms” cover one or two of them; ours needs all five.
flowchart TB
subgraph L5["Factory: issue → runs → PR, budgets, merge gate, stop switch"]
F5["Ours: agent-factory controller + Kueue + policy-bot"]
T5["Temporal: durable workflows (watched)"]
end
subgraph L4["Collaboration: shared sessions with humans and agents"]
F4["Ours: rooms (room-broker + Postgres log)"]
K4["kagent sessions: single owner"]
end
subgraph L3["Gateway: models, tools, per-run identity and budgets"]
F3["agentgateway (chosen) · Agent Router (today)"]
end
subgraph L2["Harness: the agent loop"]
F2["OpenHands (chosen)"]
A2["ax: Antigravity, Gemini only"]
K2["kagent: Claude Code, Codex, own ADKs"]
end
subgraph L1["Runtime: isolation, lifecycle, identity"]
F1["agent-sandbox + gVisor (chosen)"]
S1["Agent Substrate (google/ax and kagent v1 run on it)"]
end
L5 --> L4 --> L3 --> L2 --> L1
Decisions at a glance
| Layer | Chosen | Alternatives considered | Why |
|---|---|---|---|
| Runtime | agent-sandbox + gVisor, one pod per run | Agent Substrate (with google/ax or kagent), Kata/Firecracker, OpenHands Enterprise, Coder, E2B/Daytona | Open source; runs on EKS and GKE, on spot nodes; each run gets its own ServiceAccount and Cilium policy, so the constitution’s rules apply per run. Kata/Firecracker needs bare-metal hosts or nested virtualisation: the smallest AWS metal host has 128 vCPU, about $1.5/h spot or $5.4/h on demand (Vantage), against a few cents an hour for a run’s spot node. An agent step mostly waits on the model (about 11 s a reply), so gVisor’s system-call overhead does not show |
| Harness | OpenHands agent-server | Headless Claude Code, kagent’s harnesses, ax’s Antigravity agent | Open source, headless (an HTTP API), works with any OpenAI-compatible model, supports MCP. Claude Code is proprietary; Antigravity is Gemini-only |
| Gateway | agentgateway (decided 2026-10-01; Agent Router runs today) | Agent Router 1.1 on Envoy Gateway | See gateways |
| Rooms | A log we own: room-broker on PostgreSQL | OpenHands shared conversations, ACP, A2A through a gateway, Valkey Streams, NATS JetStream, ax or Substrate as the session layer, kagent sessions | Two different questions. Protocols (OpenHands shared conversations, ACP, A2A, ax or Substrate, kagent sessions): none records who said what among several people and agents, or who may approve. Brokers (Valkey Streams, NATS JetStream): an event is written to Postgres first, which fixes its place in the log, then announced with LISTEN/NOTIFY. At a few events per second a second stateful system adds nothing, and delivering live before the write would let watchers see events the audit log never records |
| Factory | A small custom controller + Kueue | Argo Workflows, Tekton, Temporal, a Crossplane Task composition, gh-aw | The factory’s state lives in Kubernetes objects and GitHub. Waiting hours on CI or a review is a periodic re-check of that state, so a restart loses nothing. Budgets and the kill switch are domain logic any engine would still need. Argo has no budget concept and a Crossplane composition has no timers. Temporal suits workflows whose state lives in code, and is the one to watch (below) |
| Merge gate | policy-bot behind a repository ruleset | Required reviews, rulesets alone, a custom check, Prow/tide, Mergify, Kodiak | The policy lives in the repository and is reviewable, and the right to merge sits with one dedicated App |
The full records are in the programme’s design documents, listed under Sources.
What would make us reconsider
| Current choice | Alternative that could win | Reconsider when |
|---|---|---|
| Custom factory controller | Temporal | The task lifecycle needs multi-step compensation, timed human waits multiply, retry and resume rules grow, or the factory’s bugs cluster around timers and retries |
Postgres LISTEN/NOTIFY for room fan-out | NATS JetStream, for fan-out only: the log stays in Postgres and is written first | Notification latency or database load shows on the room dashboards, or rooms need many broker replicas |
| gVisor | Kata or Firecracker microVMs | Concurrency reaches about 20 sandboxes, enough to fill a metal host; a gVisor incompatibility blocks a class of task; or nested virtualisation becomes available on our node types |
| agent-sandbox, one pod per run | Agent Substrate | See Reconsidering Substrate |
Projects evaluated in detail
Candidates that could replace a whole layer. Each section states when it was checked.
google/ax
Checked on 2026-10-04 at ac23328.
What it is. Google’s agent orchestrator: it runs each task as an agent on Agent Substrate and bootstraps Google’s Antigravity agent inside it.
Why not.
- Its control plane has no authentication or authorization. The issue is open; its reporter says Google’s security programme rated it critical (ax#376).
- An unvalidated branch name reaches
git fetch, a code-execution bug (ax#363). - Gemini only, with the API key copied into every task (client.go, reconciler.go).
- Still being rewritten: v0.3.0 dropped its event log and replaced its harness, and the queue and controller were replaced a week later (dc4f36c, ac23328).
Adopting ax would replace the harness as well as the runtime, and lock the platform to one model provider.
Reconsider if ax closes #376 and #363 and supports other model providers without putting keys in the sandbox, on top of the Substrate triggers.
Agent Substrate
Checked on 2026-10-04 at 16b863a.
What it is. A runtime that packs many gVisor-isolated agents into shared worker pods and can suspend and resume an agent, memory included, from snapshots.
What it would bring. Suspend and resume is the one capability we lack and cannot build cheaply (see below).
Why not yet.
| Constraint | Evidence |
|---|---|
| Several agents share one pod. Each gets its own egress policy, but Kubernetes identity and Cilium policy apply to the shared pod | glossary, EgressPolicy |
| Authorization is experimental and off by default | --experimental-enable-authz |
| A reclaimed worker leaves its awake agents permanently crashed; the docs warn against spot nodes | setup guide |
| Documented on kind and GKE only; needs Kubernetes certificate APIs that are stable only from 1.37 | README, Kubernetes 1.37 |
Pre-1.0; EKS-related fixes landed on main in October, unreleased | v0.3.0, #1898, #432 |
Reconsider: see the triggers.
kagent
Checked on 2026-10-04 at bf8afa56 (v1) and v0.10.3.
What it is. A CNCF Sandbox project, mostly maintained by Solo.io, and really two products:
- v1 (alpha, where the project invests) runs only on its own fork of Agent Substrate, with pluggable harnesses (Claude Code, Codex, its own, or yours) and sessions.
- v0.10 (stable, called “legacy” by the project) runs one long-running Deployment per agent.
Why not.
| Gap | Evidence |
|---|---|
| A session has one owner and one agent. Anyone joining through a share link is recorded as the owner, and no message records its author | share.go, schema |
| The open-source controller authorizes every action and takes the user’s identity from a header, so anything reaching it directly can act as anyone | auth.go |
| v1 cannot run without Substrate, and gives an agent no ServiceAccount or pod spec of its own | app.go, AGENTS.md |
| An agent identifies itself to the controller with an unsigned header until Substrate ships per-agent tokens | substrate#1660 |
| No GitHub intake, PR handling, merge gate, budgets, daily cap or stop switch, in either edition | Research |
| It dropped kubernetes-sigs/agent-sandbox in June: “substrate is the thing we’ll go with” | #2049 |
Worth learning from: forking a session together with a snapshot of the running agent, which beats our fork. It depends on Substrate.
Reconsider: see the triggers.
Temporal
Assessed on 2026-10-04 against the factory’s needs; no proof of concept.
What it is. A durable-execution engine (temporalio/temporal, MIT). Workflows are written as code, and the engine keeps their state through crashes, with durable timers, signals, retry policies and a full history of every execution.
What it would bring.
- Timed human waits, such as remind after a day, escalate after two, close after a week, without re-check logic.
- Retry and resume policies declared per step.
- An execution history and UI per task.
It can keep its state in PostgreSQL, so on this platform the database would be a SQLInstance
claim, not new infrastructure.
Why not now. The factory’s waits are re-checks of state that already lives in Kubernetes objects and GitHub, and a controller restart loses nothing. Temporal would add its server services and workflows that are versioned as code, which is heavy at today’s 20 tasks a day.
Reconsider under the conditions in the table above. The first candidate is the automatic resume of runs lost to spot reclamation, now being designed: if resume rules multiply, a workflow engine becomes worth its weight.
Would rebuilding on kagent + Substrate be simpler?
It is the strongest challenge to this design, so it is judged on the work still ahead, not on what is already built.
What it would simplify:
| Today | On Substrate |
|---|---|
Four of the 18 issues found live on gcp-0 come from one pod per run with a bridge sidecar: a lost pod restarting the conversation, transcripts lost at exit | Durable storage, suspend and resume, and a durable task log remove those classes of bug |
| A run waiting on a human approval holds a pod or ends | A parked agent costs nothing and resumes with its memory |
| A fork copies the log and starts a fresh run | A fork carries a snapshot of the running agent |
| We maintain an identity sidecar to keep provider keys out of the sandbox | Substrate’s egress gateway injects them (static keys only) |
| Every run starts cold and pulls its image | Agents start from a pre-warmed snapshot |
What we would still build:
- The factory, whole.
- The core of rooms: several participants, roles, handoff and who may approve.
- Authorization for kagent itself, which means compiling our own controller around kagent’s Go library: in effect, maintaining a fork of an alpha.
- Short-lived GitHub tokens per run and role, through a custom Substrate credential provider.
- Budgets, at the gateway.
What blocks it today: no EKS until Kubernetes 1.37 lands there, no spot nodes, a possible snapshot-restore failure for a Python harness such as OpenHands (kagent reports one), shared pods that would need an amendment to the constitution, and two alphas changing fast.
Net: today it adds moving parts we do not control rather than removing them. If we move,
Substrate as a backend behind AgentRun beats kagent + Substrate: kagent adds little we lack
and brings its authorization gap.
Reconsidering Substrate: not now, deliberately open
We intend to reconsider Substrate, and kagent with it, in the near future. Until then, a run on a reclaimed spot node fails and is resumed from its branch; making that resume automatic is being designed. The design keeps the door open:
AgentRunis the abstraction, so a Substrate backend would be a new composition behind it, not a rewrite.- Run identity accepts tokens from any issuer.
- The room bridge avoids WebSocket, which Substrate’s egress blocks.
Next step: a two-day spike at the next gcp-0 rebuild. Run OpenHands as a Substrate agent on a
gVisor worker pool, then suspend it, resume it and restore it on another node, and compare its start
time with our sandbox pod.
- If OpenHands restores and continues its conversation, we design the Substrate backend.
- If it crashes, the path stays closed for OpenHands until it is fixed upstream.
We reconsider on the spike’s result, on 2026-12-15, or as soon as any of these holds:
- Substrate ships per-agent tokens (#1660) and enforces authorization by default;
- EKS serves the Kubernetes certificate APIs as stable (EKS 1.37);
- Substrate gains a story for spot nodes, such as suspending agents on preemption;
- for kagent: open-source authorization, sessions with several attributed participants, per-agent egress (#3019), and a v1 release with a migration path.