Architecture
Components, ownership, the seat Pod contract, authentication and fencing, messaging, the gateway and the readiness flow.
This page is the binding description of how the components fit together: who owns what, the contract every seat Pod follows, and how messages, tools and readiness work. The design rationale is in Design decisions.
Components
| Component | Code | Responsibility |
|---|---|---|
| Specification and compiler | pkg/spec, pkg/compile | The schema, plus the single validator used by both the provider and the controller. Resolves templates, applies versioned defaults, expands grants and routes, and computes per-seat revisions. |
| Kubernetes API | api/v1alpha1 | AgentOrganization (declared) and AgentSeat (owned by the controller). |
| Terraform provider | provider/ | Writes the resolved organisation with server-side apply (field manager terraform-provider-steadmesh, never forced). Waits for readiness, imports, detects drift. |
| Controller | controller/, runtime/ | Reconciles organisations into seats, runs seat lifecycle (wake, idle stop, safe restarts, fenced recovery, retirement), runs the sandbox backend and reports readiness. |
| Platform service | services/ | The trusted runtime: identity, durable inbox and outbox, memory, the tools gateway, the operation ledger, the scheduler, the model proxy, Slack ingress, verification and probes. |
| Connectors | connectors/ | Slack (Socket Mode), Linear (GraphQL), model endpoints (Anthropic, OpenAI and any compatible endpoint, behind the model proxy), and secret resolvers for Kubernetes and Vault. |
| Seat runtime | cmd/seat-runner, harnesses/, cmd/steadmesh-tools | The in-Pod supervisor (lease, inbox, checkpoints, probes), the harness adapters and the tool client. |
| Packaging | charts/platform, build/ | The Helm chart (CRDs, controller, platform service) and container images. |
The controller and the platform service never execute agent-written code inside their own processes. Agents run only in seat sandboxes.
Ownership
Each piece of configuration has exactly one writer:
- The provider owns the declared fields of an
AgentOrganization. - The controller owns generated
AgentSeatobjects and their Kubernetes children, and writes only the status of organisations. - Helm installs CRDs and software, but does not also manage the organisation object.
- Agents cannot change declarations through their runtime tools.
| Information | Authoritative location | Writer |
|---|---|---|
| Organisation declaration | Infrastructure source, applied to the Kubernetes spec | Deployment workflow |
| Readiness and effective revision | Kubernetes status | Controller |
| Seat identity and retirement | Platform database | Controller, through the platform |
| Messages and pending wakes | Platform database | Ingress and runtime services |
| Memory | Platform database | Agents, through tools |
| Workspace files | Per-seat persistent volume | The seat's process |
| Projects and tasks | The work tracker | Agents, through connectors |
| External operation attempts | Operation ledger | Gateway |
| Credentials | Kubernetes Secrets or Vault | Credential owners |
Terraform state holds identifiers, declared values and computed outputs. It never holds conversations, memory or credentials, and runtime activity never causes drift.
Namespaces and services
- Control-plane namespace (
steadmesh-system): the controller (ServiceAccountsteadmesh-controller), the platform service (ServiceAccountsteadmesh-platform, Servicesteadmesh-platform:8080) and, in the example foundation, Postgres. - Organisation namespace: created by the infrastructure repository. It holds one
AgentOrganization, the seats and the instruction ConfigMaps. - Platform URL for seats and the controller:
http://steadmesh-platform.steadmesh-system.svc:8080.
Secrets
k8s:<name>refers to a Secret in the control-plane namespace, readable only by the platform ServiceAccount.vault:<path>refers to Vault KV v2. The platform logs in with Kubernetes auth, using its own ServiceAccount token.- Seats never receive external credentials. They use tools through the gateway and the model through the proxy, which a loopback forwarder in the seat runner reaches with the seat's current token. The proxy allows only the seat's configured model. Credential values never appear in prompts, logs, manifests or tool configuration.
- The resolvers are in
connectors/secretref.
Deterministic names
For seat key k in organisation o, the base name is seat- plus k lowercased with _ replaced by -, truncated to 40 characters. The suffix - plus the first 8 hex characters of sha256(o + "/" + k) is then appended. The result stays under the 52-character limit for StatefulSet names. The code is in pkg/names.
| Object | Name |
|---|---|
| AgentSeat, StatefulSet, ServiceAccount, NetworkPolicy | name |
| Pod | name-0 |
| Workspace volume claim (no owner reference, so it is retained) | ws-name |
| Manifest ConfigMap | name-manifest |
Labels: steadmesh.io/organization, steadmesh.io/seat and steadmesh.io/component=seat. ServiceAccounts are annotated with steadmesh.io/seat-id and steadmesh.io/organization-id.
Seat Pod contract
- Workload. A StatefulSet with 0 or 1 replicas and one container running
/usr/local/bin/seat-runner. It runs as uid 1000 and non-root, with a read-only root filesystem, all capabilities dropped and theRuntimeDefaultseccomp profile.automountServiceAccountTokenis false. - Identity. A projected token at
/var/run/steadmesh/tokenwith audiencesteadmesh-gateway, valid for 3600s and rotated by the kubelet. This is the seat's only credential, and the Kubernetes API rejects it. - Volumes.
/seatis the persistent workspace:workspace/is the harness's working directory,home/isHOMEand holds harness sessions, andrunner/holds runner state./etc/steadmesh/manifestis read-only and containsmanifest.json(the resolved seat manifest) andinstructions.md(the rendered instructions, in order)./tmpis an emptyDir. The runner creates the seat directories itself and fails clearly if they are not writable. - Environment.
STEADMESH_PLATFORM_URL,STEADMESH_TOKEN_FILE,STEADMESH_ORG_ID,STEADMESH_SEAT_ID,STEADMESH_SEAT_KEY,STEADMESH_CONFIG_REVISION,STEADMESH_HARNESS,STEADMESH_MANIFEST_DIR, andPOD_UIDfrom the downward API. - Probes.
:8081/healthz;:8081/readyzreports ready once the runner holds the lease and has loaded its bootstrap. - Termination. Grace period 300s. On SIGTERM the runner enters Quiescing: it lets the turn finish (interrupting it after 240s), checkpoints, reports Stopped and releases the lease.
- Network. Ingress is denied. Egress is allowed only to the platform Pods on TCP 8080 and to DNS. The controller proves this with a probe Pod before reporting
SandboxEnforced, and never starts a seat without it. - Restarts. The StatefulSet uses the
OnDeleteupdate strategy, so only the controller decides when a Pod restarts.
Authentication and fencing
There are three kinds of caller, each with its own credential: infrastructure management (the Kubernetes API), the controller (/internal/v1/*, using its ServiceAccount token, which must match an allow-listed username), and seats (/v1/*).
- A seat's token is checked with a TokenReview for audience
steadmesh-gateway. The resultingsystem:serviceaccount:<ns>:<sa>identity maps to exactly one active seat. The Pod UID comes from the token's bound claims. A client cannot pick a different principal. POST /v1/lease/acquiresucceeds when the lease is free, expired, or already held by the same Pod. It increments the generation. The TTL is 30s, and the runner renews every 10s.- Every mutating call carries
X-Steadmesh-Generation. The platform compares it with the current generation in the same transaction as the write, and stale or missing generations get409 fenced. This covers inbox leases and acks, memory writes, events, checkpoints, handoffs and connector operations. - A runner whose renewal is fenced kills its harness immediately and exits.
- The controller calls
/internal/v1/seats/{id}/fenceonly after it has confirmed that the previous Pod no longer exists. It never force-deletes a Pod on a node it cannot reach, so availability waits until fencing is reliable.
Messages, inbox and execution
- Envelope. Each message carries its ID, organisation, conversation, origin (human, seat, system, schedule or probe), sender and recipient, parent and correlation IDs, external event ID, body and reply route. Trusted ingress assigns the origin fields; agent text cannot set them.
- Ingress. A Slack event is authorised from the verified team and user IDs matching a channel binding. It is stored before the event is acknowledged, and deduplicated by connection and event ID. The request handler never waits for a model.
- Inbox.
GET /v1/inbox/nextlong-polls, then atomically leases the oldest eligible delivery. Order is preserved within a conversation, and the call creates an execution record with the seat's configuration revision. A delivery lease lasts 15 minutes. If it expires, the message is requeued with backoff (1s, 8s, 64s, then 5m), and after 5 attempts it is dead-lettered so it never blocks the seat. Each seat holds at most 1000 pending internal messages. - Outbox. Replies to humans are written to an outbox in the same transaction as the reply. Each send is recorded as a connector operation, and an ambiguous send is marked unknown rather than resent.
- Probes. The runner checks a workspace write and read. A harness with a model then proves itself with a real turn: it must reach the model through the proxy (
model), call the platformselftool through its tool bridge (tool) and complete the turn (turn). A harness without a model callsselfthroughsteadmesh-toolsinstead. See Harnesses and models.
Gateway and operation ledger
connections.invoke works as follows:
- Check the grant: the seat's capability for
connection:<k>must include the operation and match the target restrictions. - Insert a
pendingledger row. The idempotency key defaults to a hash of the seat, operation and canonical parameters; read-only operations get a fresh key each time. - If the key already exists, return the recorded operation instead of running it again.
- Call the adapter and record the outcome:
- success becomes
succeeded, with the receipt; - a permanent error becomes
failed; - a retryable error is retried with bounded backoff;
- an ambiguous or unclassified error triggers a read-back. The effect is either found (
succeeded) or not confirmed (unknown), and an unknown operation is never re-run automatically.
- success becomes
The effective authority is the intersection of three things: the platform grant, what the adapter implements, and what the external account allows.
Readiness flow
On each reconcile of the current generation, the controller works through these steps in order:
- Compile the specification and verify instruction digests (
Configured). - Sync identities with the platform. This starts retiring removed seats (they wind down, then retire; see Retiring a seat) and activates the new policy revision before returning (
IdentitiesReady). - Ensure each seat's objects: ServiceAccount, volume, manifest, NetworkPolicy and StatefulSet (
StorageReady,HarnessCompatible). - Run the NetworkPolicy enforcement probe (
SandboxEnforced). - Have the platform verify connections, Slack ingress and bound users (
ConnectionsAuthenticated,IngressReady,BindingsValid). - Probe each seat once per configuration revision, waking it if needed (
RoutesExecutable). - Aggregate the results into
OperationalReady, which is true only when every condition is true for the current generation.
orgctl verify annotates the organisation with steadmesh.io/verify-request=<nonce>. The controller runs a completely fresh pass and echoes the nonce in steadmesh.io/verify-observed only once every check has reached a terminal result. Transient states such as a conflicting update or a probe Pod that is still running keep the pass open.
Lifecycle decisions
One poller per organisation reads /internal/v1/organizations/{id}/runtime every 2s and feeds a pure decision function:
- Wake: pending deliveries and 0 replicas lead to 1 replica.
- Idle: a Warm seat with no pending work, past
last_activity + idle_timeout, scales to 0 replicas, and the runner quiesces. - Revision change: when the adopted revision differs from the declared one, the Pod is deleted at a safe boundary.
- Admin suspension:
spec.adminSuspendedkeeps the seat at 0 replicas and Blocked. - Failure: the seat goes to Recovering. The controller confirms the old Pod is gone, then fences and restarts.
Runtime API
Seat paths (/v1) | Controller paths (/internal/v1) |
|---|---|
POST lease/acquire, lease/renew, lease/releasePOST state; GET self, bootstrapGET tools; POST tools/{name}GET inbox/next?wait=; POST inbox/{id}/ackPOST executions/{id}/events; PUT checkpoint/v1/model/{connection}/… (model proxy)
|
POST organizations:syncGET organizations/{id}/runtimePOST organizations/{id}/verifyDELETE organizations/{id}?retention=POST seats/{id}/fencePOST seats/{id}/probe; GET seats/{id}/probe/{probe}
|
The wire types are in pkg/runtimeapi. Errors are JSON {code, message}, with codes unauthenticated, forbidden, fenced, conflict, not_found, invalid, blocked and unavailable. /healthz (liveness) and /readyz (database reachable) are separate, and /metrics serves Prometheus.
Durable data model
All durable data is in Postgres (migrations in services/store/migrations), scoped by organisation in every query and index:
- organizations, seats (immutable IDs; one active key per organisation; retired rows kept)
- execution leases, executions and events
- conversations, messages (unique external event ID), deliveries, outbox
- memory stores, records (full-text search) and append-only revisions
- sessions (harness checkpoints) and handoffs
- connector operations (unique idempotency key)
- wake schedules and connection checks
Uniqueness and optimistic concurrency are enforced by database constraints, not only by application code.