Operating an organisation
Check readiness, understand seat lifecycle, change configuration safely, retire seats and troubleshoot.
The deployment workflow
An infrastructure repository applies three stages in order. The example roots live in examples/:
| Stage | Creates | Typical owner |
|---|---|---|
foundation | Cluster prerequisites: namespaces, Postgres, the secret manager integration, credential Secrets | Platform/infra team |
platform | The Helm chart: CRDs, controller, platform service | Platform/infra team |
organisation | The steadmesh_organization resource and its instruction bundles | Whoever owns the organisation's structure |
| Command | What it does |
|---|---|
make -C examples apply | Plans and applies each stage in order, then runs a fresh orgctl verify. The first failing stage stops the run. |
make -C examples plan | Plans every stage. Later stages need the earlier ones applied. |
make -C examples status | Organisations, seats and readiness. |
make -C examples verify | A fresh readiness check with no Terraform involved. |
make -C examples teardown | Destroys organisation, platform and foundation. The cluster stays. |
make -C examples apply-<stage> | One stage only. |
Every saved plan, rendered plan, log and output goes to examples/out/ for review.
Checking readiness
bin/orgctl status --context kind-steadmesh --namespace steadmesh-example
bin/orgctl verify --context kind-steadmesh --namespace steadmesh-example --timeout 10m
status prints the current conditions, the seats and the representative endpoints. verify asks the controller for a completely fresh pass: it re-verifies connections and re-probes every seat. It exits non-zero unless the organisation is OperationalReady. A no-change terraform apply is not a health check, so run verify after every deploy.
Conditions
The conditions are evaluated in this order. The first one that is not True becomes the reason for OperationalReady.
| Condition | True when | Common reasons for False |
|---|---|---|
Configured | The specification compiles, and all instruction digests match their content. | InvalidSpec, digest mismatch |
IdentitiesReady | Every seat has a durable identity in the platform. | PlatformUnavailable |
ConnectionsAuthenticated | Every required connection authenticates. | ConnectionUnauthenticated with the adapter error |
IngressReady | The Slack Socket Mode connection is established. | IngressUnavailable: connecting |
BindingsValid | Every bound human exists in the workspace. | Unknown or deleted user |
StorageReady | Every seat's workspace volume is bound. | Pending PVC, no storage class |
HarnessCompatible | The harness profiles are supported. | Unsupported adapter or capability |
SandboxEnforced | A probe Pod has proved that the NetworkPolicy actually blocks traffic, and any requested RuntimeClass exists. | CNI without NetworkPolicy support, missing RuntimeClass |
RoutesExecutable | Every seat passed a synthetic probe for its current revision: a tool call, a workspace write/read and, if it has a model connection, a model ping. | ProbePending, ProbeFailed |
A seat that is stopped but has passed its probe still counts as ready. The probe never creates business work and never sends messages to people.
IntegrationsDegraded is informational and is not part of OperationalReady. It is True while an optional connection, such as a work tracker, is failing. Agents keep working without it, and work publication catches up when it recovers. Which connections are optional is described under connections.
Seat lifecycle
Seats have technical execution states, separate from any project's business status:
| State | Meaning |
|---|---|
Provisioning | The workspace, identity or Pod is being prepared. |
Warm | The runtime is up and waiting for work. |
Executing | The harness is processing a message under an execution lease. |
Quiescing | The runtime is finishing the current turn and checkpointing before it stops. |
Stopped | The Pod is gone, but every durable piece of state remains. |
Recovering | The runtime failed. The old writer is being fenced off before a new one starts. |
Blocked | A dependency or required capability is unavailable. Messages are kept, and the reason is shown. |
Retiring | The seat was removed from the declaration and is winding down. It takes no new messages and is retired when its retirement turn is done or its grace period ends (see Retiring a seat). |
Retired | The seat was removed from the declaration and has retired. Its data is retained by default. |
The controller acts on these rules:
- Wake. A pending message for a stopped seat scales it to one replica. This usually takes a few seconds.
- Idle stop. With
idle_policy = "warm_then_stop", a warm seat with no pending work is stopped onceidle_timeouthas passed since its last activity. - Configuration change. When a seat's configuration revision changes, its Pod is replaced at a safe boundary: the runner finishes or interrupts the turn, checkpoints, then exits.
- Crash. The controller confirms the old Pod is gone, never force-deleting it, before the new runner's lease takes over. Stale writers are rejected by lease generation.
Changing an organisation
- Instructions or templates. The plan's
resolved_seatsshows which seats get a new configuration revision. Each execution records the revision it ran with. - Grants. A revocation takes effect for the next tool call, before the apply reports ready. Information a model has already read cannot be un-read. Revocation stops future access only.
- Seats. Adding a seat creates a new identity. Removing one retires it gracefully (below). Renaming a seat key is a retirement plus a new seat. To carry data over, set
adopt_from = "<old seat id>"on the new seat once the old one has retired. - Harness. Changing
harness_profilekeeps identity, memory, workspace and messages. A native session that cannot be converted is replaced by a portable handoff, with a recovery note.
Retiring a seat
Removing a seat from the declaration never cuts it off mid-task. The apply returns at once, and the seat retires in three steps:
- Retiring. The seat takes no new work. Sending it a message fails with
unavailableand nothing is queued. Messages already waiting for it go back to their senders as not delivered, with the original quoted. Wake schedules and readiness probes stop. Its current turn, if any, carries on, and it keeps its tools, memory and sandbox access. - Wind-down. The seat gets one last message: a retirement notice. It finishes or pauses its work, saves a handoff with
handoff.update, releases or hands over the work items it owns, and tells the seats it works with. This is bounded by the grace period, 10 minutes by default. When the period ends, the running turn is interrupted. - Retired. Once the notice turn is done, or the grace period ends, the platform retires the seat. Whatever it still holds is handed back:
- Unfinished work items it owns are released to
ready, with a log entry pointing at its handoff. - Messages it never handled are returned to their senders.
- The organisation's representatives get a summary: whether it finished, its handoff, the released work items and the returned messages.
- Unfinished work items it owns are released to
Adding the seat back before it retires cancels the retirement. Messages already returned stay returned. The grace period is the platform's RETIREMENT_GRACE: the chart value retirement.grace, or seat_retirement_grace on the modules/platform Terraform module. 0s retires removed seats at once, and what they held is still handed back. Deleting the whole organisation retires every seat immediately.
# Watch a seat wind down
kubectl --context kind-steadmesh -n steadmesh-example get agentseats
kubectl --context kind-steadmesh -n steadmesh-example get agentseat <seat-name> -o jsonpath='{.status.retiringUntil}'
Agents can change the organisation only the way people do: through the infrastructure repository and its deployment workflow, and only if you grant them access to that repository. Seat credentials have no access to the Kubernetes API.
Administrative controls
These are infrastructure controls. Normal business steering goes through representatives.
# Stop a seat's execution without editing its prompt or workspace
kubectl --context kind-steadmesh -n steadmesh-example patch agentseat <seat-name> \
--type merge -p '{"spec":{"adminSuspended":true}}'
# Resume it
kubectl --context kind-steadmesh -n steadmesh-example patch agentseat <seat-name> \
--type merge -p '{"spec":{"adminSuspended":false}}'
# Inspect a seat
kubectl --context kind-steadmesh -n steadmesh-example get agentseat <seat-name> -o yaml
kubectl --context kind-steadmesh -n steadmesh-example logs <seat-name>-0
Seat object names look like seat-representative-sean-bd156bfb. Find them with kubectl get agentseats -l steadmesh.io/seat=<seat key>.
Retention and deletion
- Retiring a seat stops new deliveries, revokes its capabilities and removes its Pod. Its workspace volume and database records are kept, because the volume has no garbage-collecting owner.
- Destroying the organisation with
data_retention = "retain"removes the runtime and revokes everything, but keeps volumes and records. With"delete", the volumes go too. - External resources are never deleted. That covers Slack workspaces, Linear projects and accounts, because connections are
ownership = "external". - Deleting a namespace or database from outside the platform bypasses retention. Protect them in the infrastructure repository, and use a
Retainreclaim policy for production storage classes.
Observability
Every component logs structured JSON. Each line carries organisation, seat, execution and message IDs, so one inbound message can be traced from acceptance, through wake and harness execution, to tool calls and the reply. The platform exports Prometheus metrics on :8080/metrics:
| Metric | Meaning |
|---|---|
steadmesh_ingress_events_total | Inbound events by outcome: accepted, duplicate, unknown user, over cap |
steadmesh_messages_accepted_total, steadmesh_messages_rejected_total | Message acceptance |
steadmesh_inbox_oldest_pending_seconds | Queue age per seat |
steadmesh_seats | Seats by execution state |
steadmesh_deliveries_dead_lettered_total | Poison messages set aside after 5 attempts |
steadmesh_memory_conflicts_total | Stale-revision memory writes |
steadmesh_connector_operations_total, steadmesh_connector_retries_total, steadmesh_connector_unknown_outcomes_total | External operations, retries and outcomes the platform could not confirm |
steadmesh_tool_calls_total, steadmesh_automation_runs_total | Tool usage and automation runs queued |
steadmesh_http_requests_total, steadmesh_http_request_duration_seconds | API traffic |
The console
The Steadmesh Console is an optional, read-only browser view of a running organisation. It is off by default. When it is off, nothing is deployed and the platform does not serve the console API. To turn it on, set enable_console in the platform stage and apply:
terraform -chdir=examples/platform apply -var enable_console=true
orgctl console --context kind-steadmesh # forwards http://127.0.0.1:8090
Without orgctl, run kubectl --context kind-steadmesh -n steadmesh-system port-forward deploy/steadmesh-console 8090. The console listens only on its Pod's loopback interface and has no login of its own, so it can be reached only by someone allowed to port-forward in the control-plane namespace.
| View | Shows |
|---|---|
| Organisation | Teams and seats with their role and state (working, queued, waiting, starting, blocked, offline). It shows whether a runner holds each seat's lease separately from when the seat last made progress. Each seat page lists its permanent seat ID, every run across Pod replacements with the lease generation it ran under, its handoff note, its checkpoint and its memory metadata. Memory record contents are never shown. |
| Work | What seats produced: connector operations such as Linear issues, with their receipts, plus artifacts and runs. Failed operations and operations with unknown outcomes appear first, as blockers. |
| Activity | A live timeline of messages, runs, tool calls, errors and connector operations, next to a map of seats and their declared routes. A route lights up only when a message is actually sent on it. Selecting a message opens its conversation: delivery state per recipient, the run it triggered, that run's tool calls and the reply. |
| Readiness | Organisation conditions; for each seat, the requested configuration revision against the one actually running, the probe result, the workspace claim and failing conditions. Connections show only where their secret reference points and whether the last check succeeded, never the secret value. |
The console's ServiceAccount can only read AgentOrganizations, AgentSeats, Pods and PersistentVolumeClaims. The platform accepts that one identity on /console/v1, and every route there is a GET. Tool inputs and connector results shown in run details have credential-like fields (token, secret, password, API key, authorization) replaced with [redacted].
Rotating credentials
Connections refer to credentials by secret_ref: a Kubernetes Secret in the control-plane namespace (k8s:<name>) or a Vault KV v2 entry (vault:<path>). To rotate one, replace the value in place and keep the reference. You don't need to run Terraform or restart anything.
kubectl -n steadmesh-system patch secret slack-credentials \
-p '{"stringData":{"bot_token":"xoxb-new","app_token":"xapp-new"}}'
vault kv put secret/steadmesh/linear api_key=lin_api_new
| Step | What happens |
|---|---|
| Detection | Kubernetes Secrets are watched, and a change is picked up within seconds. Vault is polled every credentials.refreshInterval (Helm value, default 60s), and a poll of every secret also backs up the watch. orgctl verify refreshes immediately. A change that leaves the values the same, such as a new Vault version with identical data or a label change, does nothing. |
| Validation | The new credential is checked against the service before it is used: auth.test for Slack, a read for Linear, the model endpoint for model connections. |
| Switch-over | New requests use the new credential. Requests and model streams already in progress finish on the old one. Slack's Socket Mode connection is replaced make-before-break: the new socket connects before the old one closes. Events are acknowledged only after they are stored, and duplicates are discarded by event ID, so none are lost or processed twice. |
| Revoked before it was replaced | When a service rejects a request as unauthorised, the platform re-reads the secret straight away, at most once every 30 seconds per connection. If the credential changed, it retries the request once with the new credential. A retry happens only when the rejection proves the request had no effect. An operation whose outcome is unknown is never retried, as before. A reply to a person that is rejected this way stays queued and is retried with backoff instead of being dropped. |
When a replacement is wrong. A new credential that the service rejects is never used. If the previous credential still works, it stays in use for credentials.grace (default 10 minutes). After that, or immediately if the previous credential is also rejected, the connection becomes unavailable and the reason is shown. Putting a valid value in the secret recovers it automatically.
When the secret is deleted, the previous credential stays in use for the grace period, and then the connection becomes unavailable. When the secret store can't be read, for example because Vault is sealed or unreachable, the current credential stays in use for credentials.maxStale (default 24 hours) while refreshes keep retrying.
Status. Each connection's credential state appears in the console's Readiness view and in the platform's verify response; orgctl status and orgctl verify show it in the ConnectionsAuthenticated condition message. The states are current, replacement_rejected, secret_missing, refresh_failing and unavailable, with the secret's version and the reason. While a required connection still works but is degraded, ConnectionsAuthenticated stays true with reason CredentialRefreshDegraded. No secret value is ever shown or logged.
Vault login. The platform's own Vault login is renewed before it expires, and it logs in again if renewal fails. That is separate from reloading the credentials stored in Vault. It is reported by the steadmesh_vault_login_ok and steadmesh_vault_login_expiry_seconds metrics.
Keeping values out of Terraform. Secrets that Terraform creates (the quickstart's slack_bot_token and similar) keep their values in that stage's state. To avoid that, create the Secrets or Vault entries yourself and pass only references in existing_secret_refs.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| Apply times out on create, and the next plan wants to replace the organisation | Terraform marks a resource tainted when its create fails. Fix the reported condition, then run terraform -chdir=examples/organisation untaint steadmesh_organization.this. Timeouts on update don't taint. |
SandboxEnforced=False naming the NetworkPolicy | The cluster's CNI doesn't enforce NetworkPolicy, or the probe Pod is still running. Seats are deliberately not started without enforcement. Use a CNI that enforces it (kind's default does). |
IngressReady=False: connecting right after a Slack outage | Socket Mode is reconnecting with backoff. Re-run orgctl verify after a short wait. |
ConnectionUnauthenticated: invalid_auth | A wrong or revoked token in the Secret named by secret_ref. Put a valid value in the Secret; it is picked up without a restart (see Rotating credentials). Run orgctl verify to check immediately. |
| A seat crash-loops, logging "seat directory … is not writable" | The workspace volume's directories belong to another user. Use a storage class that honours fsGroup, or fix ownership on the volume. |
terraform init: "local package … doesn't match any of the checksums" | A rebuilt dev provider. The example Makefile regenerates the lock file on every init; if you run Terraform by hand, delete .terraform.lock.hcl. |
A reply never arrives, and the seat shows Blocked | Read the seat's condition message. The message is kept in its inbox and is delivered once the dependency recovers. |
An external operation shows unknown | The external API may have applied it, but the response was lost and could not be confirmed. The platform never replays it blindly. The agent sees it in its recovery context and reconciles it. |