Design decisions

The decisions behind the first release and the evidence for each.

The design document left seven choices to be settled during implementation. Each is recorded here with its evidence and consequences.

#DecisionStatus
1Claude Code as the first real harnessAccepted
2Direct Kubernetes execution backendAccepted
3Slack (Socket Mode) and Linear as first adaptersAccepted
4One schema and one validatorAccepted
5Workspace fencing and checkpointsAccepted
6Capacity, retention and backupInterim defaults
7Aggregate resource size limitAccepted

1. Claude Code as the first harness

The first real harness is Claude Code 2.1.289, pinned exactly in build/seat-claudecode.Dockerfile.

Feasibility evidence. All of the following were tested against a stub model API, with no user credentials:

The harness conformance suite passes against the real binary inside the seat image: read-only root, uid 1000, no capabilities.

Not yet verified. A real model through the real proxy (the live test), very long sessions with auto-compaction, and transcript compatibility across Claude Code versions (a version change is treated as a format change).

2. Direct Kubernetes backend

The design asked for a comparison of three options against the backend contract.

Agentcontainers (Kubedoll Heavy Industries)

Evaluated at commit abf4c01 (10 July 2026):

Wrapping it in a Pod would put a container runtime socket or privileged host access next to a seat, which the design forbids. Not adopted. We keep its ideas (explicit capability lists, secrets from providers, digest pinning), and the sandbox profile leaves room for a future node-level enforcer feature.

Kubernetes Agent Sandbox

It offers a Sandbox CRD with templates, claims and warm pools, plus stable identity, persistent storage, runtimeClass and "hibernation". Its documentation does not specify whether volumes are retained when a Sandbox is deleted, or what the hibernation guarantee is. Our retirement semantics and honest suspension claims depend on both. Deferred: it can become a second implementation of the backend interface, and its controller would then be reused rather than duplicated.

Chosen: direct implementation

Each seat gets a ServiceAccount without an API token, a projected token with the gateway audience, a retained volume claim with no owner, a manifest ConfigMap, a default-deny NetworkPolicy and a StatefulSet with 0 or 1 replicas.

3. Slack and Linear

Slack over Socket Mode. No public ingress is needed, which suits private clusters and kind. Events are persisted before they are acknowledged and deduplicated by event ID. One app installation serves every representative. chat.postMessage has no idempotency key, so an ambiguous send is marked unknown and never resent.

Linear over GraphQL. Created records embed steadmesh-op:<id>, so a lost response can be resolved by reading back.

Both default to ownership = "external". Destroying an organisation never deletes a workspace, app, project or issue. Setup is described in Connect real services.

4. One schema and one validator

pkg/spec is embedded in the CRD and mirrored field for field by the provider's typed schema; a test checks that the two stay in step. pkg/compile runs at plan time, where errors carry attribute paths, and again in the controller (Configured=False/InvalidSpec). Recompiling a resolved specification gives the same result, which is also tested.

5. Fencing and checkpoints

Three layers keep a single writer:

  1. Kubernetes runs at most one Pod per seat, with a RWO volume, and the controller never force-deletes.
  2. An execution lease with an incrementing generation; the controller's fence call is made only after the old Pod is confirmed gone.
  3. Every fenced write compares the caller's generation in the same transaction.

Checkpoints are application-level: a harness session reference plus a portable handoff (objective, open questions, and the relevant records, messages and operations). Persistent disk survives a restart; running processes and open network sessions do not, and the API says so. Process and microVM snapshots are not supported in this release.

6. Capacity, retention and backup (interim)

ItemInterim default
Seat resourcesCPU 250m request, 2 limit; memory 512Mi request, 2Gi limit; 5 GiB workspace
Inbox5 attempts (1s, 8s, 64s, then 5m), then dead-letter; at most 1000 pending internal messages per seat
Payloads16 KiB tool results, 12 KiB message bodies, 8 KiB event payloads
Harness turns20 minute limit
RetentionRetain by default, for organisations, seats, memory and workspaces

A backup must cover the database and the workspace volumes together, with the infrastructure revision; metadata alone is not enough. The backup and restore procedure, recovery objectives and volume reclaim policy are Milestone 5 work. Measured activation costs are Milestone 6.

7. Aggregate resource size

One steadmesh_organization owns one object. Its specification is limited to 256 KiB, well under etcd's object limit, and instructions are always referenced, never inlined. At around 200 seats, or when teams need independent ownership, the plan is to split into per-team or per-seat resources through a state migration. At every step each field keeps a single owner, never competing writers.