Rooms are an AHP-shaped log we own, stored in CNPG, fanned out with Postgres LISTEN/NOTIFY
Status: Accepted Date: 2026-09-27 Deciders: Smana (Platform Owner) Related Spec: SP2 — Collaboration rooms
Context
Humans and agent runs must share one ordered session: watch it live, steer the running agent, hand work to the next role, approve actions, fork it (programme D6, D7). The session is also SP3’s audit trail of record (C4), so it must outlive every pod and be impossible to rewrite. Each participant authenticates as itself: runs with their own ServiceAccount token (C2), humans through ZITADEL.
Decision Drivers
- One order for every participant, replayable after any disconnect
- Identity per event, stamped by the server, never claimed by a client
- The transcript survives the sandbox, redacted, and nothing can rewrite it
- Nothing dials into a sandbox (C4)
- Survives a spot interruption of the broker without losing an open approval
Considered Options
Option 1: An AHP-shaped log we own, sequential runs, CNPG + LISTEN/NOTIFY
The broker assigns a gapless seq per room. Events use the frozen C4 envelope. Runs push through a
bridge sidecar. Agents never talk to each other: a run records room_handoff or room_verdict, and
the orchestrator starts the next run with a brief built from the log. Each append also issues a
pg_notify in its transaction, and every broker replica holds one LISTEN connection that wakes its
viewers; a replica that loses it catches each room up from its last seq, so a notification is only
ever a wake-up, never data.
Pros: authorisation and identity are ours; replay is a range read; append-only is enforced by grants and triggers, testable; fan-out rides the database that already holds the log, so no infrastructure beyond the platform’s own claims. Cons: we own the protocol, and an AHP facade is later work.
Option 2: AHP v0.9 on the wire
Pros: an emerging standard for multi-client agent sessions. Cons: pre-1.0; leaves authentication and agent-to-agent out of scope, which are the hard part here.
Option 3: OpenHands conversations as the room API
Pros: already in the sandbox. Cons: no identity; the store dies with the pod; one conversation per server.
Option 4: ACP, or A2A through agentgateway
Cons: ACP is 1:1 editor-to-agent. Nothing here speaks A2A, and A2A has no humans.
Note (2026-10-01): agentgateway is no longer the only A2A path. Agent Router ships an A2A
capability in Preview, on Envoy’s native A2A filter, and an A2ARoute CRD is proposed upstream
(agent-router#2070). If an agent
outside the cluster must join a room, try it on the gateway the platform already runs first
(ecosystem re-check).
Option 5: Valkey Streams, NATS JetStream, or one in-memory broker as the log
Cons: KVStore is cache semantics by its XRD; JetStream is new infrastructure; one replica makes a spot interruption an outage for every open approval.
Option 6: Valkey pub/sub hints beside the CNPG log (design S8)
A KVStore carries “room X reached seq N” hints between broker replicas; the log stays in CNPG.
Pros: the design’s original choice; decouples fan-out load from the database.
Cons: one more dependency and one more CNP, and the KVStore’s CNP admits
the whole namespace; the log is already in Postgres, which can carry the same hint with
LISTEN/NOTIFY. Reversible until phase 2’s fan-out code lands.
Option 7: google/ax or Agent Substrate as the session substrate
Re-checked on 2026-09-27, one day after ADR-0041 rejected them as the runtime. ax v0.3.0 deleted its durable event log; Substrate emits actor lifecycle events only. Neither has sequencing, multi-client replay, per-message identity or approvals. Substrate’s parking and fork are ahead, but waking a parked actor is ingress-shaped (against C4), and a memory fork freezes the parent’s credentials into the child (against C2).
Decision Outcome
Chosen option: “Option 1”.
Rationale: the decision drivers are identity, order and durability. Only a log we own gives all three without delegating authorisation to a pre-1.0 protocol or a runtime we rejected. Fan-out uses the database that already holds the log rather than adding Valkey: one dependency fewer.
Consequences
Positive
- The transcript and the end reason of every run survive the pod.
- Reviewers, testers and triagers have a destination for their output.
UPDATE eventsas the broker’s role fails: history is append-only by grant and trigger (the schema also refusesUPDATEandTRUNCATEfrom its owner), not by convention.
Negative
- The protocol is ours to maintain. An AHP facade can sit on the log at AHP 1.0.
- A rebuild recovers the log up to the last promoted seed only.
- Fan-out shares the log’s database: each broker replica holds one
LISTENconnection, and a lost notification is recovered by reading the log from the lastseq, never from the notification.
Neutral
- Next re-check of ax and Substrate: 2026-12-15, or sooner when the cluster the platform runs on
serves v1
ClusterTrustBundleandPodCertificateRequest(gcp-0’s GKE version first; EKS 1.37 if aws-0 hosts again) and Substrate closes #1898, lifts its no-spot rule (#1528) and fixes #1657.
Implementation Notes
Code: Smana/agent-platform (OD-4). Manifests: infrastructure/base/room-broker/. Plan:
docs/superpowers/plans/2026-09-27-agent-collaboration-rooms-plan.md.
- Seams only. The core packages (envelope, redaction, store, wire) hold no platform constants:
platform facts enter through config or a consumer-side interface
(agent-platform
AGENTS.md). It keeps the core reusable outside this platform. - Spin-out deferred. Whether agent-platform becomes a standalone project is decided at the plan’s Phase 7 UX sign-off, not before: the choice needs the finished UX to judge.