Skip to content

Alternatives considered

Work in progress. The agent ecosystem moves fast. Each detailed evaluation states when it was checked and what would reopen it, and the page grows as new candidates appear.

What alternatives did we consider for each layer of the agent factory, and why did the current choices win? A new candidate gets a row in the table. If it could replace a whole layer, it also gets a section under projects evaluated in detail.

In short:

  • No candidate evaluated so far covers more than one or two layers. The factory, the rooms and the controls around them would stay ours whatever we adopted.
  • A choice can be deferred rather than refused. Each detailed evaluation lists the concrete triggers that would make us reconsider it.

How candidates are evaluated

  • Against code, not announcements. Claims are checked in the project’s repository at a pinned commit, and linked.
  • On the work still ahead. What is already built does not count in its favour; what adopting a candidate would still leave us to build does.
  • Against the platform’s requirements:
    • it runs on both EKS and GKE, on spot nodes;
    • every workload gets its own identity and network policy;
    • no long-lived credential reaches a sandbox;
    • open source first.

Where each tool fits

An agent platform has several layers. Most “agent platforms” cover one or two of them; ours needs all five.

    flowchart TB
  subgraph L5["Factory: issue → runs → PR, budgets, merge gate, stop switch"]
    F5["Ours: agent-factory controller + Kueue + policy-bot"]
    T5["Temporal: durable workflows (watched)"]
  end
  subgraph L4["Collaboration: shared sessions with humans and agents"]
    F4["Ours: rooms (room-broker + Postgres log)"]
    K4["kagent sessions: single owner"]
  end
  subgraph L3["Gateway: models, tools, per-run identity and budgets"]
    F3["agentgateway (chosen) · Agent Router (today)"]
  end
  subgraph L2["Harness: the agent loop"]
    F2["OpenHands (chosen)"]
    A2["ax: Antigravity, Gemini only"]
    K2["kagent: Claude Code, Codex, own ADKs"]
  end
  subgraph L1["Runtime: isolation, lifecycle, identity"]
    F1["agent-sandbox + gVisor (chosen)"]
    S1["Agent Substrate (google/ax and kagent v1 run on it)"]
  end
  L5 --> L4 --> L3 --> L2 --> L1
  

Decisions at a glance

LayerChosenAlternatives consideredWhy
Runtimeagent-sandbox + gVisor, one pod per runAgent Substrate (with google/ax or kagent), Kata/Firecracker, OpenHands Enterprise, Coder, E2B/DaytonaOpen source; runs on EKS and GKE, on spot nodes; each run gets its own ServiceAccount and Cilium policy, so the constitution’s rules apply per run. Kata/Firecracker needs bare-metal hosts or nested virtualisation: the smallest AWS metal host has 128 vCPU, about $1.5/h spot or $5.4/h on demand (Vantage), against a few cents an hour for a run’s spot node. An agent step mostly waits on the model (about 11 s a reply), so gVisor’s system-call overhead does not show
HarnessOpenHands agent-serverHeadless Claude Code, kagent’s harnesses, ax’s Antigravity agentOpen source, headless (an HTTP API), works with any OpenAI-compatible model, supports MCP. Claude Code is proprietary; Antigravity is Gemini-only
Gatewayagentgateway (decided 2026-10-01; Agent Router runs today)Agent Router 1.1 on Envoy GatewaySee gateways
RoomsA log we own: room-broker on PostgreSQLOpenHands shared conversations, ACP, A2A through a gateway, Valkey Streams, NATS JetStream, ax or Substrate as the session layer, kagent sessionsTwo different questions. Protocols (OpenHands shared conversations, ACP, A2A, ax or Substrate, kagent sessions): none records who said what among several people and agents, or who may approve. Brokers (Valkey Streams, NATS JetStream): an event is written to Postgres first, which fixes its place in the log, then announced with LISTEN/NOTIFY. At a few events per second a second stateful system adds nothing, and delivering live before the write would let watchers see events the audit log never records
FactoryA small custom controller + KueueArgo Workflows, Tekton, Temporal, a Crossplane Task composition, gh-awThe factory’s state lives in Kubernetes objects and GitHub. Waiting hours on CI or a review is a periodic re-check of that state, so a restart loses nothing. Budgets and the kill switch are domain logic any engine would still need. Argo has no budget concept and a Crossplane composition has no timers. Temporal suits workflows whose state lives in code, and is the one to watch (below)
Merge gatepolicy-bot behind a repository rulesetRequired reviews, rulesets alone, a custom check, Prow/tide, Mergify, KodiakThe policy lives in the repository and is reviewable, and the right to merge sits with one dedicated App

The full records are in the programme’s design documents, listed under Sources.

What would make us reconsider

Current choiceAlternative that could winReconsider when
Custom factory controllerTemporalThe task lifecycle needs multi-step compensation, timed human waits multiply, retry and resume rules grow, or the factory’s bugs cluster around timers and retries
Postgres LISTEN/NOTIFY for room fan-outNATS JetStream, for fan-out only: the log stays in Postgres and is written firstNotification latency or database load shows on the room dashboards, or rooms need many broker replicas
gVisorKata or Firecracker microVMsConcurrency reaches about 20 sandboxes, enough to fill a metal host; a gVisor incompatibility blocks a class of task; or nested virtualisation becomes available on our node types
agent-sandbox, one pod per runAgent SubstrateSee Reconsidering Substrate

Projects evaluated in detail

Candidates that could replace a whole layer. Each section states when it was checked.

google/ax

Checked on 2026-10-04 at ac23328.

What it is. Google’s agent orchestrator: it runs each task as an agent on Agent Substrate and bootstraps Google’s Antigravity agent inside it.

Why not.

  • Its control plane has no authentication or authorization. The issue is open; its reporter says Google’s security programme rated it critical (ax#376).
  • An unvalidated branch name reaches git fetch, a code-execution bug (ax#363).
  • Gemini only, with the API key copied into every task (client.go, reconciler.go).
  • Still being rewritten: v0.3.0 dropped its event log and replaced its harness, and the queue and controller were replaced a week later (dc4f36c, ac23328).

Adopting ax would replace the harness as well as the runtime, and lock the platform to one model provider.

Reconsider if ax closes #376 and #363 and supports other model providers without putting keys in the sandbox, on top of the Substrate triggers.

Agent Substrate

Checked on 2026-10-04 at 16b863a.

What it is. A runtime that packs many gVisor-isolated agents into shared worker pods and can suspend and resume an agent, memory included, from snapshots.

What it would bring. Suspend and resume is the one capability we lack and cannot build cheaply (see below).

Why not yet.

ConstraintEvidence
Several agents share one pod. Each gets its own egress policy, but Kubernetes identity and Cilium policy apply to the shared podglossary, EgressPolicy
Authorization is experimental and off by default--experimental-enable-authz
A reclaimed worker leaves its awake agents permanently crashed; the docs warn against spot nodessetup guide
Documented on kind and GKE only; needs Kubernetes certificate APIs that are stable only from 1.37README, Kubernetes 1.37
Pre-1.0; EKS-related fixes landed on main in October, unreleasedv0.3.0, #1898, #432

Reconsider: see the triggers.

kagent

Checked on 2026-10-04 at bf8afa56 (v1) and v0.10.3.

What it is. A CNCF Sandbox project, mostly maintained by Solo.io, and really two products:

  • v1 (alpha, where the project invests) runs only on its own fork of Agent Substrate, with pluggable harnesses (Claude Code, Codex, its own, or yours) and sessions.
  • v0.10 (stable, called “legacy” by the project) runs one long-running Deployment per agent.

Why not.

GapEvidence
A session has one owner and one agent. Anyone joining through a share link is recorded as the owner, and no message records its authorshare.go, schema
The open-source controller authorizes every action and takes the user’s identity from a header, so anything reaching it directly can act as anyoneauth.go
v1 cannot run without Substrate, and gives an agent no ServiceAccount or pod spec of its ownapp.go, AGENTS.md
An agent identifies itself to the controller with an unsigned header until Substrate ships per-agent tokenssubstrate#1660
No GitHub intake, PR handling, merge gate, budgets, daily cap or stop switch, in either editionResearch
It dropped kubernetes-sigs/agent-sandbox in June: “substrate is the thing we’ll go with”#2049

Worth learning from: forking a session together with a snapshot of the running agent, which beats our fork. It depends on Substrate.

Reconsider: see the triggers.

Temporal

Assessed on 2026-10-04 against the factory’s needs; no proof of concept.

What it is. A durable-execution engine (temporalio/temporal, MIT). Workflows are written as code, and the engine keeps their state through crashes, with durable timers, signals, retry policies and a full history of every execution.

What it would bring.

  • Timed human waits, such as remind after a day, escalate after two, close after a week, without re-check logic.
  • Retry and resume policies declared per step.
  • An execution history and UI per task.

It can keep its state in PostgreSQL, so on this platform the database would be a SQLInstance claim, not new infrastructure.

Why not now. The factory’s waits are re-checks of state that already lives in Kubernetes objects and GitHub, and a controller restart loses nothing. Temporal would add its server services and workflows that are versioned as code, which is heavy at today’s 20 tasks a day.

Reconsider under the conditions in the table above. The first candidate is the automatic resume of runs lost to spot reclamation, now being designed: if resume rules multiply, a workflow engine becomes worth its weight.

Would rebuilding on kagent + Substrate be simpler?

It is the strongest challenge to this design, so it is judged on the work still ahead, not on what is already built.

What it would simplify:

TodayOn Substrate
Four of the 18 issues found live on gcp-0 come from one pod per run with a bridge sidecar: a lost pod restarting the conversation, transcripts lost at exitDurable storage, suspend and resume, and a durable task log remove those classes of bug
A run waiting on a human approval holds a pod or endsA parked agent costs nothing and resumes with its memory
A fork copies the log and starts a fresh runA fork carries a snapshot of the running agent
We maintain an identity sidecar to keep provider keys out of the sandboxSubstrate’s egress gateway injects them (static keys only)
Every run starts cold and pulls its imageAgents start from a pre-warmed snapshot

What we would still build:

  • The factory, whole.
  • The core of rooms: several participants, roles, handoff and who may approve.
  • Authorization for kagent itself, which means compiling our own controller around kagent’s Go library: in effect, maintaining a fork of an alpha.
  • Short-lived GitHub tokens per run and role, through a custom Substrate credential provider.
  • Budgets, at the gateway.

What blocks it today: no EKS until Kubernetes 1.37 lands there, no spot nodes, a possible snapshot-restore failure for a Python harness such as OpenHands (kagent reports one), shared pods that would need an amendment to the constitution, and two alphas changing fast.

Net: today it adds moving parts we do not control rather than removing them. If we move, Substrate as a backend behind AgentRun beats kagent + Substrate: kagent adds little we lack and brings its authorization gap.

Reconsidering Substrate: not now, deliberately open

We intend to reconsider Substrate, and kagent with it, in the near future. Until then, a run on a reclaimed spot node fails and is resumed from its branch; making that resume automatic is being designed. The design keeps the door open:

  • AgentRun is the abstraction, so a Substrate backend would be a new composition behind it, not a rewrite.
  • Run identity accepts tokens from any issuer.
  • The room bridge avoids WebSocket, which Substrate’s egress blocks.

Next step: a two-day spike at the next gcp-0 rebuild. Run OpenHands as a Substrate agent on a gVisor worker pool, then suspend it, resume it and restore it on another node, and compare its start time with our sandbox pod.

  • If OpenHands restores and continues its conversation, we design the Substrate backend.
  • If it crashes, the path stays closed for OpenHands until it is fixed upstream.

We reconsider on the spike’s result, on 2026-12-15, or as soon as any of these holds:

  • Substrate ships per-agent tokens (#1660) and enforces authorization by default;
  • EKS serves the Kubernetes certificate APIs as stable (EKS 1.37);
  • Substrate gains a story for spot nodes, such as suspending agents on preemption;
  • for kagent: open-source authorization, sessions with several attributed participants, per-agent egress (#3019), and a v1 release with a migration path.

Sources