Operating an organisation

Check readiness, understand seat lifecycle, change configuration safely, retire seats and troubleshoot.

The deployment workflow

An infrastructure repository applies three stages in order. The example roots live in examples/:

StageCreatesTypical owner
foundationCluster prerequisites: namespaces, Postgres, the secret manager integration, credential SecretsPlatform/infra team
platformThe Helm chart: CRDs, controller, platform servicePlatform/infra team
organisationThe steadmesh_organization resource and its instruction bundlesWhoever owns the organisation's structure
CommandWhat it does
make -C examples applyPlans and applies each stage in order, then runs a fresh orgctl verify. The first failing stage stops the run.
make -C examples planPlans every stage. Later stages need the earlier ones applied.
make -C examples statusOrganisations, seats and readiness.
make -C examples verifyA fresh readiness check with no Terraform involved.
make -C examples teardownDestroys organisation, platform and foundation. The cluster stays.
make -C examples apply-<stage>One stage only.

Every saved plan, rendered plan, log and output goes to examples/out/ for review.

Checking readiness

bin/orgctl status --context kind-steadmesh --namespace steadmesh-example
bin/orgctl verify --context kind-steadmesh --namespace steadmesh-example --timeout 10m

status prints the current conditions, the seats and the representative endpoints. verify asks the controller for a completely fresh pass: it re-verifies connections and re-probes every seat. It exits non-zero unless the organisation is OperationalReady. A no-change terraform apply is not a health check, so run verify after every deploy.

Conditions

The conditions are evaluated in this order. The first one that is not True becomes the reason for OperationalReady.

ConditionTrue whenCommon reasons for False
ConfiguredThe specification compiles, and all instruction digests match their content.InvalidSpec, digest mismatch
IdentitiesReadyEvery seat has a durable identity in the platform.PlatformUnavailable
ConnectionsAuthenticatedEvery required connection authenticates.ConnectionUnauthenticated with the adapter error
IngressReadyThe Slack Socket Mode connection is established.IngressUnavailable: connecting
BindingsValidEvery bound human exists in the workspace.Unknown or deleted user
StorageReadyEvery seat's workspace volume is bound.Pending PVC, no storage class
HarnessCompatibleThe harness profiles are supported.Unsupported adapter or capability
SandboxEnforcedA probe Pod has proved that the NetworkPolicy actually blocks traffic, and any requested RuntimeClass exists.CNI without NetworkPolicy support, missing RuntimeClass
RoutesExecutableEvery seat passed a synthetic probe for its current revision: a tool call, a workspace write/read and, if it has a model connection, a model ping.ProbePending, ProbeFailed

A seat that is stopped but has passed its probe still counts as ready. The probe never creates business work and never sends messages to people.

IntegrationsDegraded is informational and is not part of OperationalReady. It is True while an optional connection, such as a work tracker, is failing. Agents keep working without it, and work publication catches up when it recovers. Which connections are optional is described under connections.

Seat lifecycle

Seats have technical execution states, separate from any project's business status:

StateMeaning
ProvisioningThe workspace, identity or Pod is being prepared.
WarmThe runtime is up and waiting for work.
ExecutingThe harness is processing a message under an execution lease.
QuiescingThe runtime is finishing the current turn and checkpointing before it stops.
StoppedThe Pod is gone, but every durable piece of state remains.
RecoveringThe runtime failed. The old writer is being fenced off before a new one starts.
BlockedA dependency or required capability is unavailable. Messages are kept, and the reason is shown.
RetiringThe seat was removed from the declaration and is winding down. It takes no new messages and is retired when its retirement turn is done or its grace period ends (see Retiring a seat).
RetiredThe seat was removed from the declaration and has retired. Its data is retained by default.

The controller acts on these rules:

Changing an organisation

Retiring a seat

Removing a seat from the declaration never cuts it off mid-task. The apply returns at once, and the seat retires in three steps:

  1. Retiring. The seat takes no new work. Sending it a message fails with unavailable and nothing is queued. Messages already waiting for it go back to their senders as not delivered, with the original quoted. Wake schedules and readiness probes stop. Its current turn, if any, carries on, and it keeps its tools, memory and sandbox access.
  2. Wind-down. The seat gets one last message: a retirement notice. It finishes or pauses its work, saves a handoff with handoff.update, releases or hands over the work items it owns, and tells the seats it works with. This is bounded by the grace period, 10 minutes by default. When the period ends, the running turn is interrupted.
  3. Retired. Once the notice turn is done, or the grace period ends, the platform retires the seat. Whatever it still holds is handed back:
    • Unfinished work items it owns are released to ready, with a log entry pointing at its handoff.
    • Messages it never handled are returned to their senders.
    • The organisation's representatives get a summary: whether it finished, its handoff, the released work items and the returned messages.
    Its lease is fenced, credentials delivered to its sandbox are revoked, and its personal memory store is retired. The controller then stops the Pod and removes the runtime, keeping the workspace volume.

Adding the seat back before it retires cancels the retirement. Messages already returned stay returned. The grace period is the platform's RETIREMENT_GRACE: the chart value retirement.grace, or seat_retirement_grace on the modules/platform Terraform module. 0s retires removed seats at once, and what they held is still handed back. Deleting the whole organisation retires every seat immediately.

# Watch a seat wind down
kubectl --context kind-steadmesh -n steadmesh-example get agentseats
kubectl --context kind-steadmesh -n steadmesh-example get agentseat <seat-name> -o jsonpath='{.status.retiringUntil}'

Agents can change the organisation only the way people do: through the infrastructure repository and its deployment workflow, and only if you grant them access to that repository. Seat credentials have no access to the Kubernetes API.

Administrative controls

These are infrastructure controls. Normal business steering goes through representatives.

# Stop a seat's execution without editing its prompt or workspace
kubectl --context kind-steadmesh -n steadmesh-example patch agentseat <seat-name> \
  --type merge -p '{"spec":{"adminSuspended":true}}'

# Resume it
kubectl --context kind-steadmesh -n steadmesh-example patch agentseat <seat-name> \
  --type merge -p '{"spec":{"adminSuspended":false}}'

# Inspect a seat
kubectl --context kind-steadmesh -n steadmesh-example get agentseat <seat-name> -o yaml
kubectl --context kind-steadmesh -n steadmesh-example logs <seat-name>-0

Seat object names look like seat-representative-sean-bd156bfb. Find them with kubectl get agentseats -l steadmesh.io/seat=<seat key>.

Retention and deletion

Observability

Every component logs structured JSON. Each line carries organisation, seat, execution and message IDs, so one inbound message can be traced from acceptance, through wake and harness execution, to tool calls and the reply. The platform exports Prometheus metrics on :8080/metrics:

MetricMeaning
steadmesh_ingress_events_totalInbound events by outcome: accepted, duplicate, unknown user, over cap
steadmesh_messages_accepted_total, steadmesh_messages_rejected_totalMessage acceptance
steadmesh_inbox_oldest_pending_secondsQueue age per seat
steadmesh_seatsSeats by execution state
steadmesh_deliveries_dead_lettered_totalPoison messages set aside after 5 attempts
steadmesh_memory_conflicts_totalStale-revision memory writes
steadmesh_connector_operations_total, steadmesh_connector_retries_total, steadmesh_connector_unknown_outcomes_totalExternal operations, retries and outcomes the platform could not confirm
steadmesh_tool_calls_total, steadmesh_automation_runs_totalTool usage and automation runs queued
steadmesh_http_requests_total, steadmesh_http_request_duration_secondsAPI traffic

The console

The Steadmesh Console is an optional, read-only browser view of a running organisation. It is off by default. When it is off, nothing is deployed and the platform does not serve the console API. To turn it on, set enable_console in the platform stage and apply:

terraform -chdir=examples/platform apply -var enable_console=true
orgctl console --context kind-steadmesh     # forwards http://127.0.0.1:8090

Without orgctl, run kubectl --context kind-steadmesh -n steadmesh-system port-forward deploy/steadmesh-console 8090. The console listens only on its Pod's loopback interface and has no login of its own, so it can be reached only by someone allowed to port-forward in the control-plane namespace.

ViewShows
OrganisationTeams and seats with their role and state (working, queued, waiting, starting, blocked, offline). It shows whether a runner holds each seat's lease separately from when the seat last made progress. Each seat page lists its permanent seat ID, every run across Pod replacements with the lease generation it ran under, its handoff note, its checkpoint and its memory metadata. Memory record contents are never shown.
WorkWhat seats produced: connector operations such as Linear issues, with their receipts, plus artifacts and runs. Failed operations and operations with unknown outcomes appear first, as blockers.
ActivityA live timeline of messages, runs, tool calls, errors and connector operations, next to a map of seats and their declared routes. A route lights up only when a message is actually sent on it. Selecting a message opens its conversation: delivery state per recipient, the run it triggered, that run's tool calls and the reply.
ReadinessOrganisation conditions; for each seat, the requested configuration revision against the one actually running, the probe result, the workspace claim and failing conditions. Connections show only where their secret reference points and whether the last check succeeded, never the secret value.

The console's ServiceAccount can only read AgentOrganizations, AgentSeats, Pods and PersistentVolumeClaims. The platform accepts that one identity on /console/v1, and every route there is a GET. Tool inputs and connector results shown in run details have credential-like fields (token, secret, password, API key, authorization) replaced with [redacted].

Rotating credentials

Connections refer to credentials by secret_ref: a Kubernetes Secret in the control-plane namespace (k8s:<name>) or a Vault KV v2 entry (vault:<path>). To rotate one, replace the value in place and keep the reference. You don't need to run Terraform or restart anything.

kubectl -n steadmesh-system patch secret slack-credentials \
  -p '{"stringData":{"bot_token":"xoxb-new","app_token":"xapp-new"}}'
vault kv put secret/steadmesh/linear api_key=lin_api_new
StepWhat happens
DetectionKubernetes Secrets are watched, and a change is picked up within seconds. Vault is polled every credentials.refreshInterval (Helm value, default 60s), and a poll of every secret also backs up the watch. orgctl verify refreshes immediately. A change that leaves the values the same, such as a new Vault version with identical data or a label change, does nothing.
ValidationThe new credential is checked against the service before it is used: auth.test for Slack, a read for Linear, the model endpoint for model connections.
Switch-overNew requests use the new credential. Requests and model streams already in progress finish on the old one. Slack's Socket Mode connection is replaced make-before-break: the new socket connects before the old one closes. Events are acknowledged only after they are stored, and duplicates are discarded by event ID, so none are lost or processed twice.
Revoked before it was replacedWhen a service rejects a request as unauthorised, the platform re-reads the secret straight away, at most once every 30 seconds per connection. If the credential changed, it retries the request once with the new credential. A retry happens only when the rejection proves the request had no effect. An operation whose outcome is unknown is never retried, as before. A reply to a person that is rejected this way stays queued and is retried with backoff instead of being dropped.

When a replacement is wrong. A new credential that the service rejects is never used. If the previous credential still works, it stays in use for credentials.grace (default 10 minutes). After that, or immediately if the previous credential is also rejected, the connection becomes unavailable and the reason is shown. Putting a valid value in the secret recovers it automatically.

When the secret is deleted, the previous credential stays in use for the grace period, and then the connection becomes unavailable. When the secret store can't be read, for example because Vault is sealed or unreachable, the current credential stays in use for credentials.maxStale (default 24 hours) while refreshes keep retrying.

Status. Each connection's credential state appears in the console's Readiness view and in the platform's verify response; orgctl status and orgctl verify show it in the ConnectionsAuthenticated condition message. The states are current, replacement_rejected, secret_missing, refresh_failing and unavailable, with the secret's version and the reason. While a required connection still works but is degraded, ConnectionsAuthenticated stays true with reason CredentialRefreshDegraded. No secret value is ever shown or logged.

Vault login. The platform's own Vault login is renewed before it expires, and it logs in again if renewal fails. That is separate from reloading the credentials stored in Vault. It is reported by the steadmesh_vault_login_ok and steadmesh_vault_login_expiry_seconds metrics.

Keeping values out of Terraform. Secrets that Terraform creates (the quickstart's slack_bot_token and similar) keep their values in that stage's state. To avoid that, create the Secrets or Vault entries yourself and pass only references in existing_secret_refs.

Troubleshooting

SymptomCause and fix
Apply times out on create, and the next plan wants to replace the organisationTerraform marks a resource tainted when its create fails. Fix the reported condition, then run terraform -chdir=examples/organisation untaint steadmesh_organization.this. Timeouts on update don't taint.
SandboxEnforced=False naming the NetworkPolicyThe cluster's CNI doesn't enforce NetworkPolicy, or the probe Pod is still running. Seats are deliberately not started without enforcement. Use a CNI that enforces it (kind's default does).
IngressReady=False: connecting right after a Slack outageSocket Mode is reconnecting with backoff. Re-run orgctl verify after a short wait.
ConnectionUnauthenticated: invalid_authA wrong or revoked token in the Secret named by secret_ref. Put a valid value in the Secret; it is picked up without a restart (see Rotating credentials). Run orgctl verify to check immediately.
A seat crash-loops, logging "seat directory … is not writable"The workspace volume's directories belong to another user. Use a storage class that honours fsGroup, or fix ownership on the volume.
terraform init: "local package … doesn't match any of the checksums"A rebuilt dev provider. The example Makefile regenerates the lock file on every init; if you run Terraform by hand, delete .terraform.lock.hcl.
A reply never arrives, and the seat shows BlockedRead the seat's condition message. The message is kept in its inbox and is delivered once the dependency recovers.
An external operation shows unknownThe external API may have applied it, but the response was lost and could not be confirmed. The platform never replays it blindly. The agent sees it in its recovery context and reconciles it.