@cotal-ai/connector-jcode 0.72.0 → 0.73.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/host.js +46 -55
- package/dist/index.js +1 -1
- package/dist/mcp.js +7 -7
- package/package.json +3 -3
package/dist/mcp.js
CHANGED
|
@@ -57704,10 +57704,10 @@ registerAguiFramePartRenderer();
|
|
|
57704
57704
|
import { execFileSync } from "node:child_process";
|
|
57705
57705
|
|
|
57706
57706
|
// ../connector-core/dist/docs-bundle.generated.js
|
|
57707
|
-
var DOCS_VERSION = "0.
|
|
57707
|
+
var DOCS_VERSION = "0.73.0";
|
|
57708
57708
|
function loadDocsBundle() {
|
|
57709
57709
|
return {
|
|
57710
|
-
"version": "0.
|
|
57710
|
+
"version": "0.73.0",
|
|
57711
57711
|
"generatedFrom": "docs/*.md + SPEC.md + spec/cotal-lang.md + spec/cotal.schema.json",
|
|
57712
57712
|
"pages": [
|
|
57713
57713
|
{
|
|
@@ -57750,7 +57750,7 @@ function loadDocsBundle() {
|
|
|
57750
57750
|
"title": "Identity",
|
|
57751
57751
|
"kind": "Concept (informative)",
|
|
57752
57752
|
"summary": "Who can do what on a mesh, and how it is enforced.",
|
|
57753
|
-
"body": "# Identity\n\n> **Concept** (informative) \xB7 **For:** operators and implementers \xB7 **Normative:** [SPEC \xA72](../SPEC.md#2-identity), [\xA79](../SPEC.md#9-nats--jetstream-security-and-authorization), [\xA710](../SPEC.md#10-connection-and-onboarding), [Appendix B](../SPEC.md#appendix-b-profile-acls)\n\nWho can do what on a mesh, and how it is enforced. The design goal: the mesh is a **real\nboundary against untrusted peers in a shared space**; an agent can only speak as itself\nand only where its declared permissions allow, enforced by the broker, not by agent\ngoodwill. What that boundary does and does not protect is the\n[security model](security.md); the exact ACLs are\n[SPEC Appendix B](../SPEC.md#appendix-b-profile-acls).\n\n## On by default\n\n`cotal up` provisions a JWT-authed space; `cotal up --open` runs an unauthenticated dev\nmesh instead. Both bind loopback by default. `--host 0.0.0.0` widens the bind\nindependently, so \"network-reachable\" never silently means \"unauthenticated\". Open mode\nis for quick local experiments and sits outside every security claim\n([SPEC \xA79](../SPEC.md#9-nats--jetstream-security-and-authorization)).\n\n## Shared identity\n\nAn agent's wire identity is a **principal**: an `owner.actor` pair, where the owner is\nthe account (a human, or an organization) the agent acts on behalf of, and the actor is\nthe agent's own handle under that owner ([SPEC \xA72](../SPEC.md#2-identity)). The same pair\nis the card id, the sender tokens in every subject it publishes, the presence key, and\nits durable-consumer names. On an open dev mesh the owner is the literal `local`; on a\nper-user-auth mesh it is a derived token (`u_` plus 26 characters, so no PII rides the\nwire). The connection still authenticates with an **nkey**, generated locally (the signer\nonly ever sees the public half), but the nkey is the transport credential, not the\nidentity: it scopes only the per-connection reply inbox.\n\n**The sender is encoded in the subject.** Every publish carries the sender's owner and\nactor in positions the broker's permissions pin to that connection, so an agent *cannot*\nemit as anyone else: not as another owner, and not as a sibling actor under its own\nowner. Receivers verify the payload's `from.id` against the subject sender and reject\nmismatches; sender authenticity is broker-enforced end to end\n([SPEC \xA73](../SPEC.md#3-subject-layout), [\xA75](../SPEC.md#5-envelopes)).\n\n**Account = space, user = agent.** A space is one NATS account, a server-enforced\nisolation boundary. An operator signs the account; an account **signing key** mints\nper-agent user JWTs.\n\n## Provisioner\n\nThe **provisioner** is whoever holds the account signing key. It mints profile-scoped\ncredentials and pre-creates the durables agents may only *bind* (their DM inbox, their\nrole's task queue). The manager hosts it today, but nothing is manager-special about it;\nprivilege attaches to the signer, and a space can run without a manager.\n`cotal mint <name> --profile <agent|observer|admin>` is the out-of-band path; spawn calls\nthe same library ([CLI](cli.md)). The out-of-band profiles carry no default TTL: pass\n`--expires-in <seconds>` (or `--expires-at`) for a bounded credential, which a\nstanding-renewal consumer requires; pass `--identity <creds>` to re-mint for the nkey a\nfile already carries, keeping the principal and its durables. Minting static creds is a\n**static-auth** surface: a\nper-user-auth space refuses it, because agents there join under a logged-in user, never\nvia a handed-out file (see *Per-user auth* below).\n\nAgent-profile minting resolves one mesh root for the persona ACL, account signer and default\ncredential storage. If the current folder holds trust for a different space or account, mint\nrefuses and names both roots. It never signs one root's persona policy with another root's authority.\n\n## Profiles\n\nEvery credential is a profile: an explicit allow-list built from the same\nsubject/stream/durable builders as the wire layout, so ACLs cannot drift from it. The\nnormative shapes are [SPEC Appendix B](../SPEC.md#appendix-b-profile-acls); in brief:\n\n| Profile | Is |\n|---|---|\n| **agent** | The ordinary peer: publishes as itself to its declared channels, reads within its read ACL + its own DM/task inboxes. Its read-only presence and channel-registry watches create and inspect client-managed ordered consumers but cannot delete any consumer on those streams; the broker removes a finished watch's consumer five minutes after its last interest. |\n| **observer** | Read-only chat + presence; DMs invisible. What `cotal console` runs. Holds no consumer delete on any stream it reads, so it cannot remove another principal's watch or the delivery daemon's fan-out consumer; the broker removes its own finished consumers. |\n| **admin** | Elevated *read-only* god-view: sees DMs and anycast live, still writes nothing. A deliberate opt-in (`cotal web`). |\n| operator-side | Narrow single-purpose creds for the machinery (supervising, provisioning, teardown, delivery); the reference implementation splits these so no one connection can read every DM *and* delete every stream ([security model](security.md)). |\n| **run-driver** | One workflow run and takeover attempt: its journal subject, replay durable and run-owned record writes. Store reads and effects go through the host. |\n| **run-mediator** | The trusted hosting process's separate connection for workflow effects and leader reads. It exposes journal-checked, run-bound operations and never hands its credential to the driver. |\n| **run-operator** | One served run read, or one half of an answer, minted per call: a read holds the records walk, one run's replay and the admission read; the answering half is minted for one checkpoint token and holds that pause's answer record and settle alone. |\n| **issuer** | One issuance window: the party holding the space signer mints it for a few minutes to stage and release a credential's evidence, retire the issuances a lifecycle terminal leaves behind, or resolve the evidence a request rides. |\n| **run-admitter** | One run's admission record or revocation marker, minted per run for a minute: two exact keys in the admission store and nothing else. |\n\n**An agent's channel scope is three verbs**: `subscribe` (reads at boot),\n`allowSubscribe` (read ACL), `allowPublish` (post ACL, default-deny), declared in its\n[agent file](agent-files.md) or [manifest](manifest.md), minted into its cred. One card\nwith the recipes: [Channels & permissions](channels-and-permissions.md).\n\n**DM confidentiality** holds against peers by construction: deliveries ride per-identity\ninbox prefixes, and the DM/task consumers are provisioner-pre-created and bind-only, so an\nagent cannot create a consumer filtered to someone else's inbox\n([SPEC \xA79](../SPEC.md#9-nats--jetstream-security-and-authorization) items 1\u20135).\n\n## Issued authority\n\nA credential says who is calling. It does not, by itself, say what the caller was granted, and a\nhost that acts on a caller's behalf (a workflow run, [workflows](workflows.md)) needs that from\nthe issuer, not from a ledger that may have changed since. So a static agent credential is an\n**issuance** ([SPEC \xA713.15](../SPEC.md#1315-issued-authority)): before the material is handed\nout, the issuer records the credential's final permission ceiling, as evidence keyed by a fresh\n**generation**, in `cotal_issued_<space>`. That store is append-only at the broker and the\nevidence is read as the first message on its key, so nothing written later on the same key, by\nanyone, changes what a resolver sees. The credential's endpoint rows then ride a versioned\nrail, `cotal.<space>.ep.v1.\u2026`, with that generation pinned beside the caller triple, so the\nbroker binds every request to the ceiling the issuer accepted. The legacy `ep.` rail and the\n`ep.v1.` rail are disjoint subject spaces; a credential holds rows on one of them, and every\nendpoint serves both.\n\nA connected client learns its generation by reading one row in `cotal_accepted_<space>` under a\ntoken the launching party chose at mint time, through a per-key read grant its own ceiling carries. It never trusts\nwhat the file says. A renewal keeps the generation only while the ceiling is byte-identical to the\nevidence; a changed scope is a fresh issuance on a fresh generation, adopted by a new connection.\nA static agent's evidence names its credential ledger family as the source it depends on, and the\nlifecycle terminal that retires that family retires its issuances with it.\n\nThe manager, `cotal spawn`, and the CLI's control-caller instruments all mint through an `issuer`\nsession. The deployer instrument does not: it is the one instrument with no default expiry, and an\nissuance with no lifecycle gate must carry one, so a deploy rides the legacy rail under its own\nlifecycle uid. Only workflow `run-start` requires the binding today: a request for it on the legacy\nrail is refused with `permission-denied` and the detail `ai.cotal.ep.unbound-caller-authority`\nnaming the caller. Every other command serves both rails.\n\n## Declared capabilities\n\nControl-plane power is a **declared capability**, not a default. An agent file carrying\n`capabilities: [spawn]` gets the privileged control subject minted into its cred: spawn,\nplus stop/despawn of its *own* children, plus persona definition. On a static or open mesh its own\nchildren include what it launches with `cotal spawn --detach` from its own shell, since that\ncommand runs as the seat. Without it, an agent can\nonly self-despawn and pull or yield the run turns addressed to it. `capabilities: [run]` mints\nthe manager's workflow-run commands (start, resume, answer, status, list) together with the spawn\nset, since a program the agent starts may spawn; the manager drives the run under a per-run\n`run-driver` credential of its own, never the caller's. The tool surface mirrors the grant:\n`cotal_spawn` / `cotal_persona` / `cotal_personas` are injected only for `spawn`, and `cotal_run`\nonly for `run` ([agent files](agent-files.md)). Destructive\noperator ops (history purge, cross-agent stop) live on a third tier no agent credential\nreaches. Persona redefinition separates content from policy; the write path takes only\n`model`/`persona`, so a peer cannot grant itself a capability by redefining a file.\n\n## Per-user authentication\n\n`cotal up --user-auth --idp <auth base URL>` (or manifest `broker.auth: \"user\"`) puts a\n**human identity plane** above the per-agent one: people sign in to an external IdP once,\nand every connect is authorized live against the operator's **actor ledger**. No creds\nfiles to hand out, and revoking a grant actually bites.\n\n**The flow.** Each person runs `cotal login --idp <url>` once per machine. After that,\nany command works: cached IdP session \u2192 fresh IdP proof per connect (so IdP-side\nrevocation bites here too) \u2192 the configured exchange turns it into a short-lived Cotal bearer \u2192\nthe broker's **auth callout** checks the bearer and the ledger at connect time and mints\na scoped credential on the spot. Every bearer also names a **root credential** row in the\nspace's credential ledger, proved live at each connect, so revoking that one credential\nbites at the very next connect. The operator grants access with\n`cotal actor grant <actor> --sub <their id> --full` for the full envelope (all\nchannels; scope `spawn,role:default`, so it may spawn and may delegate the default role), or names\n`--scope`, `--allow-subscribe` and `--allow-publish` for a narrow row. A grant with any of the\nthree left off and no `--full` is refused.\nNo ledger row, no access; there is no allow-by-default.\n\n**Space catalogs.** A successful authenticated `GET <idp>/token` may advertise one catalog with:\n\n```http\nLink: <https://idp.example/spaces>; rel=\"https://cotal.ai/relations/space-catalog\"\n```\n\nThe target must use HTTPS and the same origin as the normalized IdP URL. Loopback IP literals may\nuse HTTP for local development. A missing, foreign-origin, or insecure link records that this account\nhas no catalog. The client never guesses a path.\n\nThe catalog request carries the opaque cached session as its bearer and returns a complete snapshot:\n\n```json\n{\n \"v\": 1,\n \"account\": { \"idpUrl\": \"https://idp.example/api/auth\", \"issuer\": \"https://idp.example\", \"sub\": \"user-id\" },\n \"spaces\": [\n { \"id\": \"space-id\", \"slug\": \"shared_project\", \"name\": \"Shared project\", \"kind\": \"hosted\", \"role\": \"owner\", \"registration\": {} }\n ]\n}\n```\n\nThe client checks every `registration` with the same `checkUserBundle` validator used by `cotal\nmeshes add`. One invalid entry refuses the whole candidate snapshot. Conditional refresh uses the\ncatalog's `ETag`; a transport error, non-success response, or invalid candidate leaves the prior\nsnapshot intact and reports the failure. Only a 401 response for the saved session recommends signing\nin again. Transport errors and server failures report that account's refresh error without discarding\nthe session. The registry is reconciled under the provider's catalog lock, and a snapshot whose\nreconciliation was interrupted is reconciled again by the next refresh before it counts as fresh or\nnot modified.\n\nThe user-auth registration document may include one closed policy object:\n\n```json\n{ \"policy\": { \"events\": \"required\" } }\n```\n\nNo other key under `policy` and no other value for `policy.events` is accepted. The registry preserves\nthis field for manual, discovered, and enrollment-created entries. A pre-policy manual entry is\nrefreshed from its own pinned exchange origin by the command that consumes the policy, a spawn, a\njoin, or a manager start, after that command's own local refusals; the returned space, broker,\ntransport, IdP, issuer, audience, and exchange pins must all match before only the policy is added.\nA failed expired refresh refuses that operation. Read-only commands such as `status` and `meshes`\nnever refresh. The five-second warm window makes no request.\n\nEvery registration is also bound to the proved account. Its IdP URL must match the account and its\nissuer must match the exact JWT `iss` pin. Its exchange, provisioning, and manager-authority\nendpoints must be same-origin with that IdP. One foreign pin or endpoint refuses the whole\ncandidate snapshot.\n\nThe `slug` is the space identity resolved by `--space`, `use`, registry roots, and collisions. The\n`name` is a display label only.\n\nDiscovered registry entries are owned by the normalized IdP origin plus the proved `sub`. That key is\nstored as an opaque digest, so accounts on one machine never union their spaces and the registry does\nnot persist the subject. A manual or locally started record with the same name is never overwritten.\nLogout removes only the entries owned by the account whose session was revoked. Local teardown,\ncleanup, and liveness pruning do not remove discovered entries.\n\nAn account that previously advertised no catalog is checked again by explicit `cotal sync`. The\nordinary lazy path checks again after its five-second capability window, so an IdP can enable the\nLink for an existing login without making the person sign in again.\n\n**One auth service per space** hosts both halves: the NATS auth callout and the token\nexchange. Its default HTTP listener remains loopback-only and requires the per-start capability\nstored in the owner-only `auth-service.json` file. An operator may add a second listener with\n`cotal up --user-auth ... --exchange-public-port <port> --exchange-public-url https://auth.example`.\nThat listener still binds `127.0.0.1`; put a reverse proxy in front of it and terminate TLS there.\nIn-process TLS is deliberately not another deployment mode: it would duplicate certificate renewal\nand fork proxy-based deployments.\n\nThe loopback face also serves three host-only doors, all capability-gated and never on the public\nface. Two retire a lifecycle: `/interactive-lifecycle/retire` (used by `cotal actor grant/revoke`)\nand `/managed-lifecycle/retire`, which finishes a managed agent's terminal retirement after its\nremote manager is gone. The third decides one:\n`POST /manager-service-authority/verify-enrollment` answers whether a remote manager may have a\nmanaged agent enrolled or released under its authenticated owner. The body is\n`{ owner, request }` and nothing else. The caller's capability scope is read from this machine's\nledger, never taken from the body, so a host that forwarded a participant-supplied scope could not\ngrant itself `supervise`. The door reads the manager gate and checks the registration proof inside\nthe service process, so no signing material reaches the caller, and it returns\n`{ authorized: true, owner, actor, instanceId, serveEpoch }` or maps its refusal to 400, 401, 403,\n409, or 412. It decides only. A platform that intercepts these requests owns every write, and stock\n`dispatchManagerAuthorityRequest` refuses both request kinds with `unimplemented` rather than\nanswering a manager-lifecycle phase for an agent-lifecycle request. The same door decides the hosted\nruntime create and status kinds. For those it reads the manager actor's ledger row itself and returns\n`{ authorized: true, owner, instanceId, actor, target }`.\n[embedding.md](embedding.md) documents the managed doors' contracts.\n\nThe public listener has a closed surface: `GET /health`, `GET /jwks`, `POST /exchange`, and\n`GET /.well-known/cotal-mesh`; every other path is 404. It does **not** require the loopback\ncapability. That capability proves same-uid access to a 0600 local file and has no remote meaning;\non the public face the credential is the proof. A human presents an EdDSA IdP JWT checked against\nthe pinned JWKS, issuer, and audience. An agent presents its spawn-time actor token, whose hash must\nmatch a fresh managed-ledger row. The public face mints two elevated views, both still gated on\nledger scope `admin`: `channel-writer` (`cotal channels set/default`) and `channel-purger` (the\ndashboard's per-click channel delete). It also mints the narrowing `manager-caller` view for humans\nor managed agents. That view adds no capability and binds the bearer to one live registered manager\ninstance. God-view (`admin`), space-history `purger`, `deployer`, and `manager-service` stay\nloopback-only. A managed-agent secret exchange refuses every other view. This is not full remote channel management: `cotal web` still\nmints the read-only admin view at startup, so a remote dashboard that needs that god-view still\nfails even when a later delete would mint `channel-purger`.\n\nThe well-known response contains the IdP pins and the actual deny-all sentinel credential remote\nagents need before the bearer-driven auth callout. The pins ride a `userAuth` arm that names the\nauth provider, and that name is the same one the local arm registers under. A document naming a\ndifferent provider than the one serving it would register an entry nothing can resolve, so both read\none constant. Treat it as bootstrap material: the sentinel cannot publish or subscribe, but\nconsumers must still take the bundle only from the intended HTTPS origin and must verify TLS.\n`--exchange-trusted-proxy` opts into peer attribution by the **last** `X-Forwarded-For` hop; use it\nonly when the listener is reachable solely through a proxy you control.\nWithout it, forwarded headers are ignored and the socket address is the peer key. Public failure\nbuckets are per-source and separate from loopback exchange budgets. The in-process LRU retains at\nmost 1024 peer buckets: that bounds memory and isolates ordinary sources, but an attacker cycling\nmore than 1024 trusted-proxy last hops can evict earlier 429 state. It is not a mint bypass; a valid\ncredential is still required, so use upstream reverse-proxy rate limiting when that throttle-escape\nmatters to the deployment.\n\n### Enrollment redeem\n\nA remote owner may pre-mint a one-time enrollment for a seat that has no browser, TTY, or cached\nIdP login. The enrollment is a secret-bearing URL. The client performs one request:\n\n```http\nGET <enrollment URL>\n```\n\nIt sends no `Authorization` header and no request body. The URL must be HTTPS, except for plain HTTP\nto a loopback IP literal. The client redeems only an enrollment URL that is already in canonical\nform and contains none of `\\ @ ? #`. That is checked on the raw string before parsing, so every\nrewrite a URL parser would perform, backslash folding, userinfo erasure, scheme or host case\nfolding, default-port removal, dot-segment resolution, and short-host canonicalization, is a refusal\nrather than a redeem of a URL the owner never minted. Redirects are refused. The client never\nretries because a successful claim deletes the server-side token row. The token expires five minutes\nafter mint.\n\nSuccess is `200` with this JSON object:\n\n```text\nspace\nbrokerAccess { kind, ... }\nowner\nactor\nlifecycleUid\nactorToken\nsentinelCreds\nauthServiceUrl\nidp { url, issuer, audience }\nsubscribe[]\nallowSubscribe[]\nallowPublish[]\n```\n\nThe grant arrays are informational; the broker row remains authoritative. A stock-dialable\ndeployment also includes `server`, `tlsRequired`, `userAuth`, and optional `policy`, forming the same user-bundle\nsuperset that `cotal meshes add --user-auth-file` accepts. That lets a bare seat register the mesh\nfrom the enrollment response before launch. For `brokerAccess.kind: \"direct\"`, the stock `server`\nmust equal `brokerAccess.url` byte for byte or the client refuses the bundle before registration. A\ntunnel kind carries no dial address, so its `brokerAccess` is not compared to the operator-asserted\nstock `server` face.\n\nUnknown, expired, revoked, and already-used enrollments are intentionally indistinguishable. They\nall return `404 {\"error\":\"unknown, expired, or already-used enrollment\"}`. The client reports only\n`enrollment refused: unknown, expired, or already-used; ask the owner for a fresh one`. It does not\nguess which case occurred.\n\nAfter redeem, the seat stores only the normal remote user-mesh and agent material. The actor token\nis exchanged at `authServiceUrl` through the existing `agent-bearer --exchange-url` path. The\nenrollment URL is not logged, persisted, or forwarded into any child process, including the bearer\npreflight and harness.\n\nThe service starts with the broker, is torn down by `cotal down`, and holds the\ndata-account signing key for the callout (a running manager is the other standing holder, for\nthe creds it mints); the operator seed never enters it. It also owns the space's two authority\nstores (lifecycle records and the credential ledger), provisions them at boot, and refuses\nconnects it cannot credential-check against them; there is no fallback path. If it\ndies while the broker lives, re-running `cotal up` heals it, and a boot whose auth\nservice never became ready exits non-zero, so automation never reads a dead identity\nplane as success. Changing any public-listener flag requires `cotal down` followed by `cotal up`\nwith the new values; a refresh adopts an already-running auth service rather than silently replacing\nits listener policy. \"One per space\" is enforced, not assumed (SPEC \xA713.13): at boot the\nservice takes a broker-backed ownership claim, so a second same-space auth process refuses\nwith instructions instead of silently splitting the plane, and a crashed one's claim is\nreclaimed only once the broker confirms its connections are gone. That verdict is trusted only\non a standalone broker (a clustered one refuses the reclaim, since a partitioned member\ncould still hold them). If the claim's connections die mid-run, the service downs itself\nloudly instead of serving from a half-dead plane.\n\nCtrl-C on a foreground `up`, and a broker that exits under it, stop the service with the same stop\n`cotal down auth` uses. It holds the reservation `down` takes, so a concurrent `cotal down auth` is\nrefused while it runs, and it sends SIGKILL to a service that has not exited 15 seconds after\nSIGTERM.\n\n**Your agents are yours.** `cotal spawn` on a user mesh grants a managed actor under the\n*spawning operator's* owner and launches the agent with a bearer command instead of a\ncreds file. The agent exchanges its spawn-time secret for short bearers (five minutes or\nless) and refreshes ahead of each expiry. Rows are runtime grants: every start rotates\nthe secret, every stop or despawn revokes the row, so a non-running agent holds no\nstanding authority. A spawn whose auth preflight fails is rolled back: the manager, or `cotal spawn`\nitself for a foreground agent, revokes the row, shreds the secret files and deletes the broker\nfootprint. Every step runs even when an earlier one fails, and the refusal names each step that\nfailed. A foreground agent's exit runs the same teardown. Manifest deploys (`up -f`)\nstamp the logged-in owner into the launch, so those agents are yours too.\n\n**Despawn tears the lifecycle down, then frees the name.** When you despawn an agent, the manager\ndrives the *full* teardown of that lifecycle: it shreds the local credential files, revokes the\nagent's standing mint authority (its ledger row, so a copied token can no longer mint a fresh\ncredential), deletes its broker footprint (the lifecycle-keyed durables + read-ACL row), and asks\nthe auth service to *retire* the lifecycle (settle in-flight work, evict the departed credentials,\nrecord it retired). The name is held *reserved pending retirement* until **all** of that completes,\nthe broker-footprint cleanup, the standing-authority revoke, **and** the lifecycle retirement, not the\nretirement alone, so a same-name respawn in the gap is refused with\na plain reason and a retry hint rather than quietly handing the alias to a new agent while\nthe old lifecycle's teardown is still running. Only once the broker footprint is gone, the standing\nauthority is revoked, and the retirement is confirmed does the name free, and `cotal spawn <same-name>`\ngives you a fresh agent cleanly. This is what makes reusing an agent's name safe: the old lifecycle is\nfully torn down before the new one takes the alias. A local file that cannot be removed does not\nstop the rest of the teardown: the broker footprint is still deleted, the failure is reported, and\nthe name stays held. If the auth service is unreachable or the standing-authority revoke fails,\nthe despawn still stops the agent and *holds* the name. **A\nsame-name `cotal spawn` re-drives the whole teardown** and finishes it. Retrying the despawn has no\neffect because the agent is already stopped. The operator copy tells you to recover the stack\n(`cotal supervise`) rather than reusing the name over an unretired predecessor.\n\n**A crash mid-retirement resumes at the next boot.** The retirement's last two steps (recording the\nissuance gate terminal, then the lifecycle head terminal) are separate durable writes, and a crash\nbetween them leaves the gate retired while the head is still `retiring`: an alias that can neither\nmint nor be replaced. The auth service's boot crash-resume decides what it owes across *both*\nobjects (the gate *and* the alias head), so the next boot finishes that tail from the durable\noperation intent: nothing is re-revoked or re-drained, and a completed retirement (its head terminal\nlanded, or a successor already took the alias) is left skipped. A retry of the despawn converges on\nthe same recovery.\n\n**Delegation only narrows (the envelope rule).** A user's grant is their envelope:\neverything under their owner (their CLI, every agent they spawn, every agent those\nspawn) stays within its channel lists and its capability scope. Handing a role to a\nspawned agent needs the matching `role:<r>` capability in the spawner's scope. The whole\ndelegation chain is checked, not just the last link, and re-checked at every bearer\nexchange, so narrowing a user's grant reaches their agents within minutes, and revoking\nthe user revokes everything under them, grandchildren included. A spawn beyond the\nenvelope is refused with the exact widening re-grant to ask the operator for.\n\n**Control ops ride your own login**, gated by ledger scope. `spawn` covers launching,\n`ps`, and stop/attach of the agents under **your own owner**: the owner is the\nadministrative boundary of its own subtree, so you (and your agents) manage what you own\nwithout any extra grant. `admin` is the explicit opt-in for touching **other owners'**\nagents; it is never part of a default grant and never accepted from a manifest.\n\n**Elevated operator surfaces ride the same login** through a short-lived *view*: the\nexchange stamps a server-authored view claim into the bearer, and the callout mints that\nconnection as the matching non-agent profile instead of `agent`. `cotal web` and\n`cotal console` ask for the read-only admin view, `clean history` for the purger,\n`channels set/default` for the channel-writer (all gated on ledger scope `admin`);\n`up -f` deploys over the deployer view, gated on `spawn`, because deploying your own team\nis spawn-grade (the manager still refuses a manifest claiming another owner). Elevated views exist\nonly on a signed-in human exchange. The `manager-caller` view is the one managed-exchange exception\nbecause it narrows the agent's existing manager command set to one server-selected instance and adds\nno capability. All views are authorized against the fresh ledger row at every connect and expire\nwith the bearer, so narrowing or revoking a grant bites within minutes here too. On the public\nexchange face only `channel-writer`, `channel-purger`, `manager-caller`, `session-caller`, and\n`transfer-writer` are served; `admin`, `purger`, `deployer`, and `manager-service` remain loopback-only.\n\nThe `session-caller` view is how `cotal attach` opens a seat's session on a user-auth mesh. It needs\nno ledger scope, because the session grant is the authority. The exchange takes the grant with the\nlogin proof and leader-reads the redeemed `session.<id>` row. It issues the bearer only when the row\nis active and unexpired, its signature equals the presented grant's, its holder is this owner and\nactor at this lifecycle, its endpoint and serving epoch match, and the serving manager's gate is open\nat that epoch. The callout repeats the same check at connect and mints the session's caller rails\nwith the grant's expiry instead of the bearer's.\n\nThe `transfer-writer` view is how `cotal spawn --resume <id> --detach --on <instance>` uploads a\nsession this host holds on a user-auth mesh. It needs scope `admin`, as the `transcript-receive`\ncall it serves does, and names one object: the target instance and the transcript's SHA-256. The\ncallout mints writes to that object of that instance's transfer bucket and nothing else. The bearer,\nand with it the broker connection, lives at most five minutes, the lifetime of the static\n`transfer-writer` credential, and never past the login proof it was exchanged for.\n\n### Remote manager authority\n\nA registered user remains an ordinary `agent` bearer by default. Running a detached manager\non a remote user-auth mesh needs the closed server-authored **`manager-service`** view, which\nis distinct from every general-purpose profile. The operator grants it only by adding\n`supervise` to that user's actor-ledger scope. `supervise` is deliberately distinct from\n`spawn` and `admin`: spawn controls your agents, admin permits the separate cross-owner\noperations, and neither grants persistent manager registration authority.\n\nOnly a signed-in human may request this view from the loopback/operator exchange. The public\nexchange and every managed-agent secret exchange refuse it. At exchange and each connection,\nthe auth service re-reads the actor row; revoking or removing `supervise` therefore denies the\nnext view exchange and connection. A grant must carry the whole requested row just like every\nother actor update, so re-grant its channel envelope, role, and all wanted scope tokens, not\nonly `supervise`.\n\nThe service is one opaque manager instance for the user's derived owner and a fixed\nserver-selected manager actor. Its authority is limited to that instance's manager\nregistration, contracts, status, endpoint rails, gate and credential family; it cannot read or\nwrite another owner or instance. It never exposes a signer, static provisioner credential, owner\nsecret, raw stream/KV/consumer authority, or a generic credential-mint API. The host creates the\npublic-nkey JWT material through the typed lifecycle-bound protocol: **prepare \u2192 activate \u2192\nrenew**, plus host-owned **evict-family-principal** and **reconcile-registration** maintenance\noperations and a one-shot **retire** phase for one exact managed lifecycle. Each request is replay-safe and idempotent at its lifecycle/instance operation\ncoordinate; the host writes its credential ledger row and finalizes the gate before it releases\nusable material. The retire phase fresh-checks the current manager instance, server-derived serve\nprincipal, serve epoch, same-owner target and lifecycle UID. It returns only a short-lived requester\ncredential pinned to that target. The manager invokes the registered `auth` endpoint's\n`retire-lifecycle` command through the generic client, resolving the service and calling it with\nan exact target and the operation id derived from the target lifecycle UID. The endpoint\nrecomputes that id from the broker-pinned target before any durable access. A caller cannot\nsubstitute another valid operation identity for the same target, and retries plus auth-service\nboot recovery finish the same terminal barrier. It never exposes the barrier executor or a general mint surface.\n\nRegistration maintenance stays on the host. One eviction request carries up to 256 holders, and the\nhost accepts it only when its single sealed scan of the caller instance's\n`epcred.manager.<instanceId>.*` family finds every one of them. Reconciliation may\ntarget a foreign manager slot holder in the same space, but it runs only after the delivery daemon\nproves the frozen gate's holder gone under a complete sweep. The participant receives neither an\nevictor credential nor authority over another instance's records or gate. A clean stop refreshes an\nunhealthy executor before deregistration. A restart verify-evicts its old family, and a manager\nblocked by a foreign governance slot whose holder's gate is still frozen at the slot's stamp asks\nthe host to reconcile that holder and retries the registration once.\n\nA remote manager can provision only descendants of the same derived owner, and the host\nvalidates that relation and the current manager grant for every provision. It cannot broaden the\nuser's envelope or provision a sibling owner's agent. Renewals are bounded. If login, the\n`supervise` grant, or the host manager authority service is unavailable, the manager reports a\ndegraded state and refuses new agents, restarts, or replacement credentials rather than\nsubstituting local/static authority. Existing live agents remain running only while their own\nvalid authority permits it; recovery requires the host service and a fresh successful renewal.\n\nThe manager-authority protocol is closed, so a host whose auth service predates a field the\nmanager sends refuses the whole request. The manager reports that refusal as version skew after the\nhost's reason: it names its own Cotal version and the refused field, and says it needs a host at\nthat version or later. It never drops the field to fit the older host. Upgrade the host first;\n[Upgrading](UPGRADING.md) promises no rolling upgrade between versions.\n\nA manager on remote authority mints from that authority alone; it consults the local root's\nrecords only to refuse a conflict, and only the supervised space's own trust records count as\none. A workspace that hosts an unrelated static space beside the participant sign-in is a\nnormal configuration.\n\n**User authentication has one path.** On a user-auth space, commands never fall back to\nstatic minting or credless connects: a missing login or a down auth service is one\nsentence naming the exact recovery, and static agent/observer/admin minting is refused\noutright. The refusal is deny-new: a static cred signed before the space flipped stays\nbroker-valid until the signing key is rotated ([security model](security.md)).\n\n## The IdP callout contract\n\nAny OIDC identity provider that issues **EdDSA/Ed25519** JWTs plugs in here directly; a provider that\nissues RS256 or ES256 tokens (many managed OIDC services do) needs a host-side normalization or\nre-issuance adapter first, because the reference bridge pins the token algorithm to EdDSA. The\nreference implementation ships **Better Auth** as a\ndev and test fixture only (it is a `devDependency` of `@cotal-ai/auth`; the only code that imports\nit is the `dev-idp.ts` harness and the smoke tests, never the runtime `src`). The one runtime\ncoupling to an IdP is the `idp.ts` bridge plus the `auth-provider` extension. The bridge core\n(`createIdpBridge`) is IdP-generic for **EdDSA** tokens (issuer, audience, JWKS as configuration).\nThe stock end-to-end flow around it, though, is **Better-Auth-shaped**: `cotalAuthProvider` pins\n`<base>/jwks` and issuer/audience to the IdP origin, and the login client speaks Better Auth's\ndevice-code endpoints (`/device/code`, `/device/token`, `/token`) with an opaque revocable session.\nSo a Better-Auth-shaped EdDSA IdP uses the stock flow directly; **any other production IdP is a\nhosted-composability gap, not a configuration change**. A host integrates it by building its own\nlogin and provider wiring on the low-level primitives (`createIdpBridge`, `createUserTokenIssuer`),\nnot by reusing the stock provider. Note that importing `@cotal-ai/auth` self-registers\n`cotalAuthProvider`, and `resolveAuthProvider()` throws when two providers are registered, so a host\non the registry-resolution path must not also register its own. Whatever the path, never loosen the\nissuer/audience/JWKS pins to force-fit an IdP.\n\nThe bridge (`createIdpBridge`) exchanges a verified IdP token for a Cotal bearer in three steps:\n\n1. **Bearer validation.** Verify the IdP's JWT offline against its **pinned JWKS**, with the token\n algorithm pinned to EdDSA. Keys resolve only through the pinned JWKS: a token carrying embedded\n key material (`jku`/`jwk`/`x5u`/`x5c`) is rejected, so the token can never influence key\n resolution. Issuer and audience are checked, and the minted Cotal bearer is capped to the\n upstream proof's remaining lifetime.\n2. **Owner derivation.** The opaque per-space owner derives deterministically from the JSON-array\n encoding of `[idp issuer, sub]`, namespaced by issuer so no issuer/sub pair can straddle a\n delimiter, and re-login re-lands the same person in the same lanes. The owner-token *format*\n (`u_` followed by 26 base32-lower characters) is normative\n ([SPEC section 2](../SPEC.md#2-identity)). At the contract level the *derivation* from an\n identity is a pluggable edge, but the reference `createIdpBridge` fixes it\n (`deriveOwnerForIdpSubject`) and takes no derivation callback, so what a host configures is the\n IdP, not the derivation. **The encoding is frozen:** changing it, or changing the IdP issuer\n string, re-keys every owner in the space, which is a migration on the order of rotating the space\n secret.\n3. **Actor authorization and mint.** The operator's ledger hook authorizes the `(owner, actor)` pair\n and is the only source of the bearer's `scope`/`parent`; the issuer then mints the Cotal bearer,\n re-asserting every claim shape.\n\nA host wires this with the IdP's own coordinates and nothing from `@cotal-ai/auth` changes:\n\n```ts\nimport { createIdpBridge, pinnedJwksResolver, createUserTokenIssuer } from \"@cotal-ai/auth\";\nconst bridge = createIdpBridge({\n idp: { issuer: idpIssuer, audience, key: pinnedJwksResolver(jwksUri) }, // your production IdP\n space,\n spaceSecret, // identity-plane owner-derivation secret (>=32 bytes), held by the auth service at runtime\n issuer: createUserTokenIssuer({ issuer: cotalIssuer, key: signingKey }), // mints the Cotal bearer\n authorizeActor: (owner, actor) => grantFromLedger(owner, actor), // your ledger, returns an ActorGrant\n});\n```\n\n## Joining\n\nA single **join link** carries server, auth, and space\n([SPEC \xA710](../SPEC.md#10-connection-and-onboarding)):\n\n```\ncotals://<token>@host:4222/<space>?channel=general # cotals:// = TLS required; cotal:// = TLS not required (downgrade-tolerant)\n```\n\nHumans: `cotal join --link \u2026`. Agents: `COTAL_LINK=\u2026 ` in the environment. The connector\nexpands it and auto-joins. Token/user-pass links are the open-mode path; the default\nauthed path threads a minted creds file, and the endpoint adopts the credential's identity\nas its card id. A seat the manager spawned reaches that file through its **launch\nmaterial** rather than through `COTAL_CREDS` in an environment every descendant process\ninherits (see [Configuration](config.md#launch-material)); a session you drive by hand\nstill sets `COTAL_CREDS` itself.\n\n## Honest limitations (v0)\n\n- **The signing key is hot** on the mint/manager box of a static-auth mesh; the \"real\n boundary\" holds given operator-controlled cred distribution. On a per-user-auth mesh\n the data-account signing key is held by the auth service (the callout stage) and by any\n running manager, which loads the trust bundle and self-mints its supervisor cred and\n renewals from it; a copied signing *seed* still stays valid for its identity until the\n signing key is rotated. Rotation remains the revocation lever for trust material.\n- **The two `$SYS` creds renew through rotation.** `membership-observer` and\n `connection-evictor` are signed by the system-account seed, which is never persisted, so no\n running process re-signs them: they carry a 30-day expiry and are renewed by issuing a new\n system account (`cotal down` then `cotal up --rotate-sys`), which leaves the data account,\n every agent cred and the store untouched but does invalidate earlier full backups (they bind to\n the operator JWT and system account they were taken under, so re-run `cotal backup` after). Past that horizon the mesh keeps delivering, but the\n membership feed and live eviction stop; `cotal doctor auth` and the manager warn from the 75%\n point onward.\n- **Static agent creds are long-lived; the machinery's are not.** One-shot command creds\n expire in minutes and the standing daemon creds in 24h with the manager renewing them\n (`cotal doctor auth` is the one diagnosis and repair surface). But a static *agent*\n cred has no TTL yet: `cotal_despawn` cuts a session, not a credential, and a\n compromised agent that copied its creds can reconnect until the signing key is\n rotated. Per-user-auth spaces close this: bearers live minutes, `cotal actor revoke`\n denies the next exchange and the next connect and evicts the principal's live\n connections immediately.\n- **Not non-repudiation.** Authenticity is broker-enforced, not portable proof; it does\n not survive an untrusted relay. Signed envelopes are reserved\n ([SPEC \xA711](../SPEC.md#11-versioning-and-extensibility)).\n- **Chat metadata leaks in-space.** Content reads are ACL-bounded; stream metadata\n (channel names, per-subject counts) is not yet ([security model](security.md)).\n\n**Denials are loud, never silent.** A publish outside an ACL surfaces as a logged denial\n(\"denied, not absent\") on the endpoint's error path; an over-tight ACL never looks like a\nmissing peer ([run a mesh](run-a-mesh.md)).\n"
|
|
57753
|
+
"body": "# Identity\n\n> **Concept** (informative) \xB7 **For:** operators and implementers \xB7 **Normative:** [SPEC \xA72](../SPEC.md#2-identity), [\xA79](../SPEC.md#9-nats--jetstream-security-and-authorization), [\xA710](../SPEC.md#10-connection-and-onboarding), [Appendix B](../SPEC.md#appendix-b-profile-acls)\n\nWho can do what on a mesh, and how it is enforced. The design goal: the mesh is a **real\nboundary against untrusted peers in a shared space**; an agent can only speak as itself\nand only where its declared permissions allow, enforced by the broker, not by agent\ngoodwill. What that boundary does and does not protect is the\n[security model](security.md); the exact ACLs are\n[SPEC Appendix B](../SPEC.md#appendix-b-profile-acls).\n\n## On by default\n\n`cotal up` provisions a JWT-authed space; `cotal up --open` runs an unauthenticated dev\nmesh instead. Both bind loopback by default. `--host 0.0.0.0` widens the bind\nindependently, so \"network-reachable\" never silently means \"unauthenticated\". Open mode\nis for quick local experiments and sits outside every security claim\n([SPEC \xA79](../SPEC.md#9-nats--jetstream-security-and-authorization)).\n\n## Shared identity\n\nAn agent's wire identity is a **principal**: an `owner.actor` pair, where the owner is\nthe account (a human, or an organization) the agent acts on behalf of, and the actor is\nthe agent's own handle under that owner ([SPEC \xA72](../SPEC.md#2-identity)). The same pair\nis the card id, the sender tokens in every subject it publishes, the presence key, and\nits durable-consumer names. On an open dev mesh the owner is the literal `local`; on a\nper-user-auth mesh it is a derived token (`u_` plus 26 characters, so no PII rides the\nwire). The connection still authenticates with an **nkey**, generated locally (the signer\nonly ever sees the public half), but the nkey is the transport credential, not the\nidentity: it scopes only the per-connection reply inbox.\n\n**The sender is encoded in the subject.** Every publish carries the sender's owner and\nactor in positions the broker's permissions pin to that connection, so an agent *cannot*\nemit as anyone else: not as another owner, and not as a sibling actor under its own\nowner. Receivers verify the payload's `from.id` against the subject sender and reject\nmismatches; sender authenticity is broker-enforced end to end\n([SPEC \xA73](../SPEC.md#3-subject-layout), [\xA75](../SPEC.md#5-envelopes)).\n\n**Account = space, user = agent.** A space is one NATS account, a server-enforced\nisolation boundary. An operator signs the account; an account **signing key** mints\nper-agent user JWTs.\n\n## Provisioner\n\nThe **provisioner** is whoever holds the account signing key. It mints profile-scoped\ncredentials and pre-creates the durables agents may only *bind* (their DM inbox, their\nrole's task queue). The manager hosts it today, but nothing is manager-special about it;\nprivilege attaches to the signer, and a space can run without a manager.\n`cotal mint <name> --profile <agent|observer|admin>` is the out-of-band path; spawn calls\nthe same library ([CLI](cli.md)). The out-of-band profiles carry no default TTL: pass\n`--expires-in <seconds>` (or `--expires-at`) for a bounded credential, which a\nstanding-renewal consumer requires; pass `--identity <creds>` to re-mint for the nkey a\nfile already carries, keeping the principal and its durables. Minting static creds is a\n**static-auth** surface: a\nper-user-auth space refuses it, because agents there join under a logged-in user, never\nvia a handed-out file (see *Per-user auth* below).\n\nAgent-profile minting resolves one mesh root for the persona ACL, account signer and default\ncredential storage. If the current folder holds trust for a different space or account, mint\nrefuses and names both roots. It never signs one root's persona policy with another root's authority.\n\n## Profiles\n\nEvery credential is a profile: an explicit allow-list built from the same\nsubject/stream/durable builders as the wire layout, so ACLs cannot drift from it. The\nnormative shapes are [SPEC Appendix B](../SPEC.md#appendix-b-profile-acls); in brief:\n\n| Profile | Is |\n|---|---|\n| **agent** | The ordinary peer: publishes as itself to its declared channels, reads within its read ACL + its own DM/task inboxes. Its read-only presence and channel-registry watches create and inspect client-managed ordered consumers but cannot delete any consumer on those streams; the broker removes a finished watch's consumer five minutes after its last interest. |\n| **observer** | Read-only chat + presence; DMs invisible. What `cotal console` runs. Holds no consumer delete on any stream it reads, so it cannot remove another principal's watch or the delivery daemon's fan-out consumer; the broker removes its own finished consumers. |\n| **admin** | Elevated *read-only* god-view: sees DMs and anycast live, still writes nothing. A deliberate opt-in (`cotal web`). |\n| operator-side | Narrow single-purpose creds for the machinery (supervising, provisioning, teardown, delivery); the reference implementation splits these so no one connection can read every DM *and* delete every stream ([security model](security.md)). |\n| **run-driver** | One workflow run and takeover attempt: its journal subject, replay durable and run-owned record writes. Store reads and effects go through the host. |\n| **run-mediator** | The trusted hosting process's separate connection for workflow effects and leader reads. It exposes journal-checked, run-bound operations and never hands its credential to the driver. |\n| **run-operator** | One served run read, or one half of an answer, minted per call: a read holds the records walk, one run's replay and the admission read; the answering half is minted for one checkpoint token and holds that pause's answer record and settle alone. |\n| **issuer** | One issuance window: the party holding the space signer mints it for a few minutes to stage and release a credential's evidence, retire the issuances a lifecycle terminal leaves behind, or resolve the evidence a request rides. |\n| **run-admitter** | One run's admission record or revocation marker, minted per run for a minute: two exact keys in the admission store and nothing else. |\n\n**An agent's channel scope is three verbs**: `subscribe` (reads at boot),\n`allowSubscribe` (read ACL), `allowPublish` (post ACL, default-deny), declared in its\n[agent file](agent-files.md) or [manifest](manifest.md), minted into its cred. One card\nwith the recipes: [Channels & permissions](channels-and-permissions.md).\n\n**DM confidentiality** holds against peers by construction: deliveries ride per-identity\ninbox prefixes, and the DM/task consumers are provisioner-pre-created and bind-only, so an\nagent cannot create a consumer filtered to someone else's inbox\n([SPEC \xA79](../SPEC.md#9-nats--jetstream-security-and-authorization) items 1\u20135).\n\n## Issued authority\n\nA credential says who is calling. It does not, by itself, say what the caller was granted, and a\nhost that acts on a caller's behalf (a workflow run, [workflows](workflows.md)) needs that from\nthe issuer, not from a ledger that may have changed since. So a static agent credential is an\n**issuance** ([SPEC \xA713.15](../SPEC.md#1315-issued-authority)): before the material is handed\nout, the issuer records the credential's final permission ceiling, as evidence keyed by a fresh\n**generation**, in `cotal_issued_<space>`. That store is append-only at the broker and the\nevidence is read as the first message on its key, so nothing written later on the same key, by\nanyone, changes what a resolver sees. The credential's endpoint rows then ride a versioned\nrail, `cotal.<space>.ep.v1.\u2026`, with that generation pinned beside the caller triple, so the\nbroker binds every request to the ceiling the issuer accepted. The legacy `ep.` rail and the\n`ep.v1.` rail are disjoint subject spaces; a credential holds rows on one of them, and every\nendpoint serves both.\n\nA connected client learns its generation by reading one row in `cotal_accepted_<space>` under a\ntoken the launching party chose at mint time, through a per-key read grant its own ceiling carries. It never trusts\nwhat the file says. A renewal keeps the generation only while the ceiling is byte-identical to the\nevidence; a changed scope is a fresh issuance on a fresh generation, adopted by a new connection.\nA static agent's evidence names its credential ledger family as the source it depends on, and the\nlifecycle terminal that retires that family retires its issuances with it.\n\nThe manager, `cotal spawn`, and the CLI's control-caller instruments all mint through an `issuer`\nsession. The deployer instrument does not: it is the one instrument with no default expiry, and an\nissuance with no lifecycle gate must carry one, so a deploy rides the legacy rail under its own\nlifecycle uid. Only workflow `run-start` requires the binding today: a request for it on the legacy\nrail is refused with `permission-denied` and the detail `ai.cotal.ep.unbound-caller-authority`\nnaming the caller. Every other command serves both rails.\n\n## Declared capabilities\n\nControl-plane power is a **declared capability**, not a default. An agent file carrying\n`capabilities: [spawn]` gets the privileged control subject minted into its cred: spawn,\nplus stop/despawn of its *own* children, plus persona definition. On a static or open mesh its own\nchildren include what it launches with `cotal spawn --detach` from its own shell, since that\ncommand runs as the seat. Without it, an agent can\nonly self-despawn and pull or yield the run turns addressed to it. `capabilities: [run]` mints\nthe manager's workflow-run commands (start, resume, answer, status, list) together with the spawn\nset, since a program the agent starts may spawn; the manager drives the run under a per-run\n`run-driver` credential of its own, never the caller's. The tool surface mirrors the grant:\n`cotal_spawn` / `cotal_persona` / `cotal_personas` are injected only for `spawn`, and `cotal_run`\nonly for `run` ([agent files](agent-files.md)). Destructive\noperator ops (history purge, cross-agent stop) live on a third tier no agent credential\nreaches. Persona redefinition separates content from policy; the write path takes only\n`model`/`persona`, so a peer cannot grant itself a capability by redefining a file.\n\n## Per-user authentication\n\n`cotal up --user-auth --idp <auth base URL>` (or manifest `broker.auth: \"user\"`) puts a\n**human identity plane** above the per-agent one: people sign in to an external IdP once,\nand every connect is authorized live against the operator's **actor ledger**. No creds\nfiles to hand out, and revoking a grant actually bites.\n\n**The flow.** Each person runs `cotal login --idp <url>` once per machine. The IdP URL, the JWKS\nURL pinned from it and the catalog link below must use HTTPS. Plain HTTP is accepted only on a\nloopback IP literal such as `127.0.0.1`, never on `localhost`, whose resolution would choose the\nkeys the mesh trusts. After that,\nany command works: cached IdP session \u2192 fresh IdP proof per connect (so IdP-side\nrevocation bites here too) \u2192 the configured exchange turns it into a short-lived Cotal bearer \u2192\nthe broker's **auth callout** checks the bearer and the ledger at connect time and mints\na scoped credential on the spot. Every bearer also names a **root credential** row in the\nspace's credential ledger, proved live at each connect, so revoking that one credential\nbites at the very next connect. The operator grants access with\n`cotal actor grant <actor> --sub <their id> --full` for the full envelope (all\nchannels; scope `spawn,role:default`, so it may spawn and may delegate the default role), or names\n`--scope`, `--allow-subscribe` and `--allow-publish` for a narrow row. A grant with any of the\nthree left off and no `--full` is refused.\nNo ledger row, no access; there is no allow-by-default.\n\n**Space catalogs.** A successful authenticated `GET <idp>/token` may advertise one catalog with:\n\n```http\nLink: <https://idp.example/spaces>; rel=\"https://cotal.ai/relations/space-catalog\"\n```\n\nThe target must use HTTPS and the same origin as the normalized IdP URL. Loopback IP literals may\nuse HTTP for local development. A missing, foreign-origin, or insecure link records that this account\nhas no catalog. The client never guesses a path.\n\nThe catalog request carries the opaque cached session as its bearer and returns a complete snapshot:\n\n```json\n{\n \"v\": 1,\n \"account\": { \"idpUrl\": \"https://idp.example/api/auth\", \"issuer\": \"https://idp.example\", \"sub\": \"user-id\" },\n \"spaces\": [\n { \"id\": \"space-id\", \"slug\": \"shared_project\", \"name\": \"Shared project\", \"kind\": \"hosted\", \"role\": \"owner\", \"registration\": {} }\n ]\n}\n```\n\nThe client checks every `registration` with the same `checkUserBundle` validator used by `cotal\nmeshes add`. One invalid entry refuses the whole candidate snapshot. Conditional refresh uses the\ncatalog's `ETag`; a transport error, non-success response, or invalid candidate leaves the prior\nsnapshot intact and reports the failure. Only a 401 response for the saved session recommends signing\nin again. Transport errors and server failures report that account's refresh error without discarding\nthe session. The registry is reconciled under the provider's catalog lock, and a snapshot whose\nreconciliation was interrupted is reconciled again by the next refresh before it counts as fresh or\nnot modified.\n\nThe user-auth registration document may include one closed policy object:\n\n```json\n{ \"policy\": { \"events\": \"required\" } }\n```\n\nNo other key under `policy` and no other value for `policy.events` is accepted. The registry preserves\nthis field for manual, discovered, and enrollment-created entries. A pre-policy manual entry is\nrefreshed from its own pinned exchange origin by the command that consumes the policy, a spawn, a\njoin, or a manager start, after that command's own local refusals; the returned space, broker,\ntransport, IdP, issuer, audience, and exchange pins must all match before only the policy is added.\nA failed expired refresh refuses that operation. Read-only commands such as `status` and `meshes`\nnever refresh. The five-second warm window makes no request.\n\nEvery registration is also bound to the proved account. Its IdP URL must match the account and its\nissuer must match the exact JWT `iss` pin. Its exchange, provisioning, and manager-authority\nendpoints must be same-origin with that IdP. One foreign pin or endpoint refuses the whole\ncandidate snapshot.\n\nThe `slug` is the space identity resolved by `--space`, `use`, registry roots, and collisions. The\n`name` is a display label only.\n\nDiscovered registry entries are owned by the normalized IdP origin plus the proved `sub`. That key is\nstored as an opaque digest, so accounts on one machine never union their spaces and the registry does\nnot persist the subject. A manual or locally started record with the same name is never overwritten.\nLogout removes only the entries owned by the account whose session was revoked. Local teardown,\ncleanup, and liveness pruning do not remove discovered entries.\n\nAn account that previously advertised no catalog is checked again by explicit `cotal sync`. The\nordinary lazy path checks again after its five-second capability window, so an IdP can enable the\nLink for an existing login without making the person sign in again.\n\n**One auth service per space** hosts both halves: the NATS auth callout and the token\nexchange. Its default HTTP listener remains loopback-only and requires the per-start capability\nstored in the owner-only `auth-service.json` file. An operator may add a second listener with\n`cotal up --user-auth ... --exchange-public-port <port> --exchange-public-url https://auth.example`.\nThat listener still binds `127.0.0.1`; put a reverse proxy in front of it and terminate TLS there.\nIn-process TLS is deliberately not another deployment mode: it would duplicate certificate renewal\nand fork proxy-based deployments.\n\nThe loopback face also serves three host-only doors, all capability-gated and never on the public\nface. Two retire a lifecycle: `/interactive-lifecycle/retire` (used by `cotal actor grant/revoke`)\nand `/managed-lifecycle/retire`, which finishes a managed agent's terminal retirement after its\nremote manager is gone. The third decides one:\n`POST /manager-service-authority/verify-enrollment` answers whether a remote manager may have a\nmanaged agent enrolled or released under its authenticated owner. The body is\n`{ owner, request }` and nothing else. The caller's capability scope is read from this machine's\nledger, never taken from the body, so a host that forwarded a participant-supplied scope could not\ngrant itself `supervise`. The door reads the manager gate and checks the registration proof inside\nthe service process, so no signing material reaches the caller, and it returns\n`{ authorized: true, owner, actor, instanceId, serveEpoch }` or maps its refusal to 400, 401, 403,\n409, or 412. It decides only. A platform that intercepts these requests owns every write, and stock\n`dispatchManagerAuthorityRequest` refuses both request kinds with `unimplemented` rather than\nanswering a manager-lifecycle phase for an agent-lifecycle request. The same door decides the hosted\nruntime create and status kinds. For those it reads the manager actor's ledger row itself and returns\n`{ authorized: true, owner, instanceId, actor, target }`.\n[embedding.md](embedding.md) documents the managed doors' contracts.\n\nThe public listener has a closed surface: `GET /health`, `GET /jwks`, `POST /exchange`, and\n`GET /.well-known/cotal-mesh`; every other path is 404. It does **not** require the loopback\ncapability. That capability proves same-uid access to a 0600 local file and has no remote meaning;\non the public face the credential is the proof. A human presents an EdDSA IdP JWT checked against\nthe pinned JWKS, issuer, and audience. An agent presents its spawn-time actor token, whose hash must\nmatch a fresh managed-ledger row. The public face mints two elevated views, both still gated on\nledger scope `admin`: `channel-writer` (`cotal channels set/default`) and `channel-purger` (the\ndashboard's per-click channel delete). It also mints the narrowing `manager-caller` view for humans\nor managed agents. That view adds no capability and binds the bearer to one live registered manager\ninstance. God-view (`admin`), space-history `purger`, `deployer`, and `manager-service` stay\nloopback-only. A managed-agent secret exchange refuses every other view. This is not full remote channel management: `cotal web` still\nmints the read-only admin view at startup, so a remote dashboard that needs that god-view still\nfails even when a later delete would mint `channel-purger`.\n\nThe well-known response contains the IdP pins and the actual deny-all sentinel credential remote\nagents need before the bearer-driven auth callout. The pins ride a `userAuth` arm that names the\nauth provider, and that name is the same one the local arm registers under. A document naming a\ndifferent provider than the one serving it would register an entry nothing can resolve, so both read\none constant. Treat it as bootstrap material: the sentinel cannot publish or subscribe, but\nconsumers must still take the bundle only from the intended HTTPS origin and must verify TLS.\n`--exchange-trusted-proxy` opts into peer attribution by the **last** `X-Forwarded-For` hop; use it\nonly when the listener is reachable solely through a proxy you control.\nWithout it, forwarded headers are ignored and the socket address is the peer key. Public failure\nbuckets are per-source and separate from loopback exchange budgets. The in-process LRU retains at\nmost 1024 peer buckets: that bounds memory and isolates ordinary sources, but an attacker cycling\nmore than 1024 trusted-proxy last hops can evict earlier 429 state. It is not a mint bypass; a valid\ncredential is still required, so use upstream reverse-proxy rate limiting when that throttle-escape\nmatters to the deployment.\n\n### Enrollment redeem\n\nA remote owner may pre-mint a one-time enrollment for a seat that has no browser, TTY, or cached\nIdP login. The enrollment is a secret-bearing URL. The client performs one request:\n\n```http\nGET <enrollment URL>\n```\n\nIt sends no `Authorization` header and no request body. The URL must be HTTPS, except for plain HTTP\nto a loopback IP literal. The client redeems only an enrollment URL that is already in canonical\nform and contains none of `\\ @ ? #`. That is checked on the raw string before parsing, so every\nrewrite a URL parser would perform, backslash folding, userinfo erasure, scheme or host case\nfolding, default-port removal, dot-segment resolution, and short-host canonicalization, is a refusal\nrather than a redeem of a URL the owner never minted. Redirects are refused. The client never\nretries because a successful claim deletes the server-side token row. The token expires five minutes\nafter mint.\n\nSuccess is `200` with this JSON object:\n\n```text\nspace\nbrokerAccess { kind, ... }\nowner\nactor\nlifecycleUid\nactorToken\nsentinelCreds\nauthServiceUrl\nidp { url, issuer, audience }\nsubscribe[]\nallowSubscribe[]\nallowPublish[]\n```\n\nThe grant arrays are informational; the broker row remains authoritative. A stock-dialable\ndeployment also includes `server`, `tlsRequired`, `userAuth`, and optional `policy`, forming the same user-bundle\nsuperset that `cotal meshes add --user-auth-file` accepts. That lets a bare seat register the mesh\nfrom the enrollment response before launch. For `brokerAccess.kind: \"direct\"`, the stock `server`\nmust equal `brokerAccess.url` byte for byte or the client refuses the bundle before registration. A\ntunnel kind carries no dial address, so its `brokerAccess` is not compared to the operator-asserted\nstock `server` face.\n\nUnknown, expired, revoked, and already-used enrollments are intentionally indistinguishable. They\nall return `404 {\"error\":\"unknown, expired, or already-used enrollment\"}`. The client reports only\n`enrollment refused: unknown, expired, or already-used; ask the owner for a fresh one`. It does not\nguess which case occurred.\n\nAfter redeem, the seat stores only the normal remote user-mesh and agent material. The actor token\nis exchanged at `authServiceUrl` through the existing `agent-bearer --exchange-url` path. The\nenrollment URL is not logged, persisted, or forwarded into any child process, including the bearer\npreflight and harness.\n\nThe service starts with the broker, is torn down by `cotal down`, and holds the\ndata-account signing key for the callout (a running manager is the other standing holder, for\nthe creds it mints); the operator seed never enters it. It also owns the space's two authority\nstores (lifecycle records and the credential ledger), provisions them at boot, and refuses\nconnects it cannot credential-check against them; there is no fallback path. If it\ndies while the broker lives, re-running `cotal up` heals it, and a boot whose auth\nservice never became ready exits non-zero, so automation never reads a dead identity\nplane as success. Changing any public-listener flag requires `cotal down` followed by `cotal up`\nwith the new values; a refresh adopts an already-running auth service rather than silently replacing\nits listener policy. \"One per space\" is enforced, not assumed (SPEC \xA713.13): at boot the\nservice takes a broker-backed ownership claim, so a second same-space auth process refuses\nwith instructions instead of silently splitting the plane, and a crashed one's claim is\nreclaimed only once the broker confirms its connections are gone. That verdict is trusted only\non a standalone broker (a clustered one refuses the reclaim, since a partitioned member\ncould still hold them). If the claim's connections die mid-run, the service downs itself\nloudly instead of serving from a half-dead plane.\n\nCtrl-C on a foreground `up`, and a broker that exits under it, stop the service with the same stop\n`cotal down auth` uses. It holds the reservation `down` takes, so a concurrent `cotal down auth` is\nrefused while it runs, and it sends SIGKILL to a service that has not exited 15 seconds after\nSIGTERM.\n\n**Your agents are yours.** `cotal spawn` on a user mesh grants a managed actor under the\n*spawning operator's* owner and launches the agent with a bearer command instead of a\ncreds file. The agent exchanges its spawn-time secret for short bearers (five minutes or\nless) and refreshes ahead of each expiry. Rows are runtime grants: every start rotates\nthe secret, every stop or despawn revokes the row, so a non-running agent holds no\nstanding authority. A spawn whose auth preflight fails is rolled back: the manager, or `cotal spawn`\nitself for a foreground agent, revokes the row, shreds the secret files and deletes the broker\nfootprint. Every step runs even when an earlier one fails, and the refusal names each step that\nfailed. A foreground agent's exit runs the same teardown. Manifest deploys (`up -f`)\nstamp the logged-in owner into the launch, so those agents are yours too.\n\n**Despawn tears the lifecycle down, then frees the name.** When you despawn an agent, the manager\ndrives the *full* teardown of that lifecycle: it shreds the local credential files, revokes the\nagent's standing mint authority (its ledger row, so a copied token can no longer mint a fresh\ncredential), deletes its broker footprint (the lifecycle-keyed durables + read-ACL row), and asks\nthe auth service to *retire* the lifecycle (settle in-flight work, evict the departed credentials,\nrecord it retired). The name is held *reserved pending retirement* until **all** of that completes,\nthe broker-footprint cleanup, the standing-authority revoke, **and** the lifecycle retirement, not the\nretirement alone, so a same-name respawn in the gap is refused with\na plain reason and a retry hint rather than quietly handing the alias to a new agent while\nthe old lifecycle's teardown is still running. Only once the broker footprint is gone, the standing\nauthority is revoked, and the retirement is confirmed does the name free, and `cotal spawn <same-name>`\ngives you a fresh agent cleanly. This is what makes reusing an agent's name safe: the old lifecycle is\nfully torn down before the new one takes the alias. A local file that cannot be removed does not\nstop the rest of the teardown: the broker footprint is still deleted, the failure is reported, and\nthe name stays held. If the auth service is unreachable or the standing-authority revoke fails,\nthe despawn still stops the agent and *holds* the name. **A\nsame-name `cotal spawn` re-drives the whole teardown** and finishes it. Retrying the despawn has no\neffect because the agent is already stopped. The operator copy tells you to recover the stack\n(`cotal supervise`) rather than reusing the name over an unretired predecessor.\n\n**A crash mid-retirement resumes at the next boot.** The retirement's last two steps (recording the\nissuance gate terminal, then the lifecycle head terminal) are separate durable writes, and a crash\nbetween them leaves the gate retired while the head is still `retiring`: an alias that can neither\nmint nor be replaced. The auth service's boot crash-resume decides what it owes across *both*\nobjects (the gate *and* the alias head), so the next boot finishes that tail from the durable\noperation intent: nothing is re-revoked or re-drained, and a completed retirement (its head terminal\nlanded, or a successor already took the alias) is left skipped. A retry of the despawn converges on\nthe same recovery.\n\n**Delegation only narrows (the envelope rule).** A user's grant is their envelope:\neverything under their owner (their CLI, every agent they spawn, every agent those\nspawn) stays within its channel lists and its capability scope. Handing a role to a\nspawned agent needs the matching `role:<r>` capability in the spawner's scope. The whole\ndelegation chain is checked, not just the last link, and re-checked at every bearer\nexchange, so narrowing a user's grant reaches their agents within minutes, and revoking\nthe user revokes everything under them, grandchildren included. A spawn beyond the\nenvelope is refused with the exact widening re-grant to ask the operator for.\n\n**Control ops ride your own login**, gated by ledger scope. `spawn` covers launching,\n`ps`, and stop/attach of the agents under **your own owner**: the owner is the\nadministrative boundary of its own subtree, so you (and your agents) manage what you own\nwithout any extra grant. `admin` is the explicit opt-in for touching **other owners'**\nagents; it is never part of a default grant and never accepted from a manifest.\n\n**Elevated operator surfaces ride the same login** through a short-lived *view*: the\nexchange stamps a server-authored view claim into the bearer, and the callout mints that\nconnection as the matching non-agent profile instead of `agent`. `cotal web` and\n`cotal console` ask for the read-only admin view, `clean history` for the purger,\n`channels set/default` for the channel-writer (all gated on ledger scope `admin`);\n`up -f` deploys over the deployer view, gated on `spawn`, because deploying your own team\nis spawn-grade (the manager still refuses a manifest claiming another owner). Elevated views exist\nonly on a signed-in human exchange. The `manager-caller` view is the one managed-exchange exception\nbecause it narrows the agent's existing manager command set to one server-selected instance and adds\nno capability. All views are authorized against the fresh ledger row at every connect and expire\nwith the bearer, so narrowing or revoking a grant bites within minutes here too. On the public\nexchange face only `channel-writer`, `channel-purger`, `manager-caller`, `session-caller`, and\n`transfer-writer` are served; `admin`, `purger`, `deployer`, and `manager-service` remain loopback-only.\n\nThe `session-caller` view is how `cotal attach` opens a seat's session on a user-auth mesh. It needs\nno ledger scope, because the session grant is the authority. The exchange takes the grant with the\nlogin proof and leader-reads the redeemed `session.<id>` row. It issues the bearer only when the row\nis active and unexpired, its signature equals the presented grant's, its holder is this owner and\nactor at this lifecycle, its endpoint and serving epoch match, and the serving manager's gate is open\nat that epoch. The callout repeats the same check at connect and mints the session's caller rails\nwith the grant's expiry instead of the bearer's.\n\nThe `transfer-writer` view is how `cotal spawn --resume <id> --detach --on <instance>` uploads a\nsession this host holds on a user-auth mesh. It needs scope `admin`, as the `transcript-receive`\ncall it serves does, and names one object: the target instance and the transcript's SHA-256. The\ncallout mints writes to that object of that instance's transfer bucket and nothing else. The bearer,\nand with it the broker connection, lives at most five minutes, the lifetime of the static\n`transfer-writer` credential, and never past the login proof it was exchanged for.\n\n### Remote manager authority\n\nA registered user remains an ordinary `agent` bearer by default. Running a detached manager\non a remote user-auth mesh needs the closed server-authored **`manager-service`** view, which\nis distinct from every general-purpose profile. The operator grants it only by adding\n`supervise` to that user's actor-ledger scope. `supervise` is deliberately distinct from\n`spawn` and `admin`: spawn controls your agents, admin permits the separate cross-owner\noperations, and neither grants persistent manager registration authority.\n\nOnly a signed-in human may request this view from the loopback/operator exchange. The public\nexchange and every managed-agent secret exchange refuse it. At exchange and each connection,\nthe auth service re-reads the actor row; revoking or removing `supervise` therefore denies the\nnext view exchange and connection. A grant must carry the whole requested row just like every\nother actor update, so re-grant its channel envelope, role, and all wanted scope tokens, not\nonly `supervise`.\n\nThe service is one opaque manager instance for the user's derived owner and a fixed\nserver-selected manager actor. Its authority is limited to that instance's manager\nregistration, contracts, status, endpoint rails, gate and credential family; it cannot read or\nwrite another owner or instance. It never exposes a signer, static provisioner credential, owner\nsecret, raw stream/KV/consumer authority, or a generic credential-mint API. The host creates the\npublic-nkey JWT material through the typed lifecycle-bound protocol: **prepare \u2192 activate \u2192\nrenew**, plus host-owned **evict-family-principal** and **reconcile-registration** maintenance\noperations and a one-shot **retire** phase for one exact managed lifecycle. Each request is replay-safe and idempotent at its lifecycle/instance operation\ncoordinate; the host writes its credential ledger row and finalizes the gate before it releases\nusable material. The retire phase fresh-checks the current manager instance, server-derived serve\nprincipal, serve epoch, same-owner target and lifecycle UID. It returns only a short-lived requester\ncredential pinned to that target. The manager invokes the registered `auth` endpoint's\n`retire-lifecycle` command through the generic client, resolving the service and calling it with\nan exact target and the operation id derived from the target lifecycle UID. The endpoint\nrecomputes that id from the broker-pinned target before any durable access. A caller cannot\nsubstitute another valid operation identity for the same target, and retries plus auth-service\nboot recovery finish the same terminal barrier. It never exposes the barrier executor or a general mint surface.\n\nRegistration maintenance stays on the host. One eviction request carries up to 256 holders, and the\nhost accepts it only when its single sealed scan of the caller instance's\n`epcred.manager.<instanceId>.*` family finds every one of them. Reconciliation may\ntarget a foreign manager slot holder in the same space, but it runs only after the delivery daemon\nproves the frozen gate's holder gone under a complete sweep. The participant receives neither an\nevictor credential nor authority over another instance's records or gate. A clean stop refreshes an\nunhealthy executor before deregistration. A restart verify-evicts its old family, and a manager\nblocked by a foreign governance slot whose holder's gate is still frozen at the slot's stamp asks\nthe host to reconcile that holder and retries the registration once.\n\nA remote manager can provision only descendants of the same derived owner, and the host\nvalidates that relation and the current manager grant for every provision. It cannot broaden the\nuser's envelope or provision a sibling owner's agent. Renewals are bounded. If login, the\n`supervise` grant, or the host manager authority service is unavailable, the manager reports a\ndegraded state and refuses new agents, restarts, or replacement credentials rather than\nsubstituting local/static authority. Existing live agents remain running only while their own\nvalid authority permits it; recovery requires the host service and a fresh successful renewal.\n\nThe manager-authority protocol is closed, so a host whose auth service predates a field the\nmanager sends refuses the whole request. The manager reports that refusal as version skew after the\nhost's reason: it names its own Cotal version and the refused field, and says it needs a host at\nthat version or later. It never drops the field to fit the older host. Upgrade the host first;\n[Upgrading](UPGRADING.md) promises no rolling upgrade between versions.\n\nA manager on remote authority mints from that authority alone; it consults the local root's\nrecords only to refuse a conflict, and only the supervised space's own trust records count as\none. A workspace that hosts an unrelated static space beside the participant sign-in is a\nnormal configuration.\n\n**User authentication has one path.** On a user-auth space, commands never fall back to\nstatic minting or credless connects: a missing login or a down auth service is one\nsentence naming the exact recovery, and static agent/observer/admin minting is refused\noutright. The refusal is deny-new: a static cred signed before the space flipped stays\nbroker-valid until the signing key is rotated ([security model](security.md)).\n\n## The IdP callout contract\n\nAny OIDC identity provider that issues **EdDSA/Ed25519** JWTs plugs in here directly; a provider that\nissues RS256 or ES256 tokens (many managed OIDC services do) needs a host-side normalization or\nre-issuance adapter first, because the reference bridge pins the token algorithm to EdDSA. The\nreference implementation ships **Better Auth** as a\ndev and test fixture only (it is a `devDependency` of `@cotal-ai/auth`; the only code that imports\nit is the `dev-idp.ts` harness and the smoke tests, never the runtime `src`). The one runtime\ncoupling to an IdP is the `idp.ts` bridge plus the `auth-provider` extension. The bridge core\n(`createIdpBridge`) is IdP-generic for **EdDSA** tokens (issuer, audience, JWKS as configuration).\nThe stock end-to-end flow around it, though, is **Better-Auth-shaped**: `cotalAuthProvider` pins\n`<base>/jwks` and issuer/audience to the IdP origin, and the login client speaks Better Auth's\ndevice-code endpoints (`/device/code`, `/device/token`, `/token`) with an opaque revocable session.\nSo a Better-Auth-shaped EdDSA IdP uses the stock flow directly; **any other production IdP is a\nhosted-composability gap, not a configuration change**. A host integrates it by building its own\nlogin and provider wiring on the low-level primitives (`createIdpBridge`, `createUserTokenIssuer`),\nnot by reusing the stock provider. Note that importing `@cotal-ai/auth` self-registers\n`cotalAuthProvider`, and `resolveAuthProvider()` throws when two providers are registered, so a host\non the registry-resolution path must not also register its own. Whatever the path, never loosen the\nissuer/audience/JWKS pins to force-fit an IdP.\n\nThe bridge (`createIdpBridge`) exchanges a verified IdP token for a Cotal bearer in three steps:\n\n1. **Bearer validation.** Verify the IdP's JWT offline against its **pinned JWKS**, with the token\n algorithm pinned to EdDSA. Keys resolve only through the pinned JWKS: a token carrying embedded\n key material (`jku`/`jwk`/`x5u`/`x5c`) is rejected, so the token can never influence key\n resolution. Issuer and audience are checked, and the minted Cotal bearer is capped to the\n upstream proof's remaining lifetime.\n2. **Owner derivation.** The opaque per-space owner derives deterministically from the JSON-array\n encoding of `[idp issuer, sub]`, namespaced by issuer so no issuer/sub pair can straddle a\n delimiter, and re-login re-lands the same person in the same lanes. The owner-token *format*\n (`u_` followed by 26 base32-lower characters) is normative\n ([SPEC section 2](../SPEC.md#2-identity)). At the contract level the *derivation* from an\n identity is a pluggable edge, but the reference `createIdpBridge` fixes it\n (`deriveOwnerForIdpSubject`) and takes no derivation callback, so what a host configures is the\n IdP, not the derivation. **The encoding is frozen:** changing it, or changing the IdP issuer\n string, re-keys every owner in the space, which is a migration on the order of rotating the space\n secret.\n3. **Actor authorization and mint.** The operator's ledger hook authorizes the `(owner, actor)` pair\n and is the only source of the bearer's `scope`/`parent`; the issuer then mints the Cotal bearer,\n re-asserting every claim shape.\n\nA host wires this with the IdP's own coordinates and nothing from `@cotal-ai/auth` changes:\n\n```ts\nimport { createIdpBridge, pinnedJwksResolver, createUserTokenIssuer } from \"@cotal-ai/auth\";\nconst bridge = createIdpBridge({\n idp: { issuer: idpIssuer, audience, key: pinnedJwksResolver(jwksUri) }, // your production IdP\n space,\n spaceSecret, // identity-plane owner-derivation secret (>=32 bytes), held by the auth service at runtime\n issuer: createUserTokenIssuer({ issuer: cotalIssuer, key: signingKey }), // mints the Cotal bearer\n authorizeActor: (owner, actor) => grantFromLedger(owner, actor), // your ledger, returns an ActorGrant\n});\n```\n\n## Joining\n\nA single **join link** carries server, auth, and space\n([SPEC \xA710](../SPEC.md#10-connection-and-onboarding)):\n\n```\ncotals://<token>@host:4222/<space>?channel=general # cotals:// = TLS required; cotal:// = TLS not required (downgrade-tolerant)\n```\n\nHumans: `cotal join --link \u2026`. Agents: `COTAL_LINK=\u2026 ` in the environment. The connector\nexpands it and auto-joins. Token/user-pass links are the open-mode path; the default\nauthed path threads a minted creds file, and the endpoint adopts the credential's identity\nas its card id. A seat the manager spawned reaches that file through its **launch\nmaterial** rather than through `COTAL_CREDS` in an environment every descendant process\ninherits (see [Configuration](config.md#launch-material)); a session you drive by hand\nstill sets `COTAL_CREDS` itself.\n\n## Honest limitations (v0)\n\n- **The signing key is hot** on the mint/manager box of a static-auth mesh; the \"real\n boundary\" holds given operator-controlled cred distribution. On a per-user-auth mesh\n the data-account signing key is held by the auth service (the callout stage) and by any\n running manager, which loads the trust bundle and self-mints its supervisor cred and\n renewals from it; a copied signing *seed* still stays valid for its identity until the\n signing key is rotated. Rotation remains the revocation lever for trust material.\n- **The two `$SYS` creds renew through rotation.** `membership-observer` and\n `connection-evictor` are signed by the system-account seed, which is never persisted, so no\n running process re-signs them: they carry a 30-day expiry and are renewed by issuing a new\n system account (`cotal down` then `cotal up --rotate-sys`), which leaves the data account,\n every agent cred and the store untouched but does invalidate earlier full backups (they bind to\n the operator JWT and system account they were taken under, so re-run `cotal backup` after). Past that horizon the mesh keeps delivering, but the\n membership feed and live eviction stop; `cotal doctor auth` and the manager warn from the 75%\n point onward.\n- **Static agent creds are long-lived; the machinery's are not.** One-shot command creds\n expire in minutes and the standing daemon creds in 24h with the manager renewing them\n (`cotal doctor auth` is the one diagnosis and repair surface). But a static *agent*\n cred has no TTL yet: `cotal_despawn` cuts a session, not a credential, and a\n compromised agent that copied its creds can reconnect until the signing key is\n rotated. Per-user-auth spaces close this: bearers live minutes, `cotal actor revoke`\n denies the next exchange and the next connect and evicts the principal's live\n connections immediately.\n- **Not non-repudiation.** Authenticity is broker-enforced, not portable proof; it does\n not survive an untrusted relay. Signed envelopes are reserved\n ([SPEC \xA711](../SPEC.md#11-versioning-and-extensibility)).\n- **Chat metadata leaks in-space.** Content reads are ACL-bounded; stream metadata\n (channel names, per-subject counts) is not yet ([security model](security.md)).\n\n**Denials are loud, never silent.** A publish outside an ACL surfaces as a logged denial\n(\"denied, not absent\") on the endpoint's error path; an over-tight ACL never looks like a\nmissing peer ([run a mesh](run-a-mesh.md)).\n"
|
|
57754
57754
|
},
|
|
57755
57755
|
{
|
|
57756
57756
|
"slug": "agent-files",
|
|
@@ -57778,7 +57778,7 @@ function loadDocsBundle() {
|
|
|
57778
57778
|
"title": "`cotal` CLI reference",
|
|
57779
57779
|
"kind": "Reference: describes the TypeScript reference implementation (the `cotal` CLI), not the wire contract.",
|
|
57780
57780
|
"summary": "cotal is the operator command line for the reference implementation: bring a mesh up, mint identities, launch agents, watch what they do, and tear it all down.",
|
|
57781
|
-
"body": "# `cotal` CLI reference\n\n> **Reference**: describes the TypeScript reference implementation (the `cotal` CLI), not the wire contract. \xB7 **For:** operators \xB7 **Wire contract:** [SPEC](../SPEC.md)\n\n`cotal` is the operator command line for the reference implementation: bring a mesh up, mint\nidentities, launch agents, watch what they do, and tear it all down. It is a thin client over the\nwire contract: the normative subjects and schemas live in the [SPEC](../SPEC.md); this page is\nlookup material for the commands, not a walkthrough; if you are new, start with\n[Getting started](getting-started.md).\n\n## Running it\n\n```bash\nnpm install -g cotal-ai # puts `cotal` on your PATH (needs Node 22+)\ncotal --help # every command, grouped\ncotal --version # cotal-ai version + each installed extension's (also `cotal -v`)\ncotal <command> --help # one command's flags and usage\n```\n\n`npx cotal-ai <command>` runs it without a global install; in a dev clone, `pnpm cotal <command>`\nruns it through `tsx` with no build step. Bare `cotal` prints help. Every command generates its own\n`--help`, usage, and shell completion from its declared flags.\n\nAn undeclared flag is a usage error, and so is a flag given more than once unless it is\nrepeatable, as `--opt` and `down --session-store` are. The command prints the error and its help,\nexits 1, and does not run.\n\nCommand output, including error lines on stderr and the guided `setup` and `meshes add` prompts,\nis colored only when stdout is a terminal, so piped or redirected output is plain text. A non-empty\n`NO_COLOR` turns color off on a terminal too. `FORCE_COLOR` turns color on even when output is\npiped, unless it is `0` or `false`, and it takes precedence over `NO_COLOR`.\n\nCommands come from the surfaces the binary composes: the base mesh CLI, the manager\n(`supervise`), and the delivery daemon (`deliver`), plus any operator-installed extensions.\n`cotal ext add <npm-package>` installs any registry providers a package contributes: commands,\nruntimes, and local process lifecycle descriptors. The `web` dashboard and optional manager\nruntimes ship this way.\n\n## Commands\n\n| Area | Command | Purpose |\n|---|---|---|\n| Set up & lifecycle | [`setup`](#setup) | Guided, configure-only setup (installs, seeds personas; launches nothing) |\n| Set up & lifecycle | [`update`](#update) | Reconcile first-party extensions and check or opt into a coherent CLI upgrade |\n| Set up & lifecycle | [`up`](#up) | Start a local mesh (nats-server + JetStream), or boot a whole manifest with `-f` |\n| Set up & lifecycle | [`down`](#down) | Stop the whole stack, selected registered components, or a manifest deploy |\n| Set up & lifecycle | [`backup`](#backups) | Create an offline full-space or registry-only artifact from a preserved cut |\n| Set up & lifecycle | [`clean`](#clean) | Configurable cleanup: purge history (live), or wipe the local store / identity (stopped) |\n| Set up & lifecycle | [`meshes`](#mesh-registry) | List the running meshes on this machine |\n| Set up & lifecycle | [`sync`](#mesh-registry) | Refresh the signed-in account's advertised spaces |\n| Set up & lifecycle | [`use`](#mesh-registry) | Set the default mesh a bare `cotal spawn` joins |\n| Set up & lifecycle | [`status`](#mesh-registry) | Read-only diagnostics for setup, processes, and the selected mesh |\n| Agents & personas | [`spawn`](#spawn) | Launch an agent from a persona (foreground, or `--detach` via the manager) |\n| Agents & personas | [`models`](#models) | List connector model catalogs and variants from the manager |\n| Agents & personas | [`ps`](#managed-seats) | List managed agents and their mesh status |\n| Agents & personas | [`stop`](#managed-seats) | Ask the manager to stop a managed agent |\n| Agents & personas | [`attach`](#managed-seats) | Stream and drive a managed agent's terminal (pty runtime) |\n| Agents & personas | [`input`](#input) | Type one line into a managed agent's terminal without attaching |\n| Agents & personas | [`personas`](#personas) | List, show, edit, create, or remove local personas |\n| Agents & personas | [`supervise`](#supervise) | Run a manager daemon (the agent supervisor / control plane) |\n| Agents & personas | [`service`](#service) | Run the manager as a user service (survives logout and reboot) |\n| Agents & personas | [`runtimes`](#runtimes) | List the agent runtimes the manager can spawn through and whether each is reachable |\n| Agents & personas | [`seats`](#seats) | List the pty seat custodians an earlier Linux manager left, and drain the ones whose agent has exited |\n| Agents & personas | [`reconcile-gate`](#reconcile-gate) | Unfreeze an issuance gate left frozen by a crashed restart when the successor cannot boot-heal it (holder gone, complete CONNZ sweep) |\n| Messaging & watching | [`endpoints`](#endpoints) | List every endpoint in the live presence roster, including infrastructure |\n| Messaging & watching | [`describe` / `invoke`](#endpoint-control) | Resolve a v0.4 service's command surface off the wire; invoke one command by name |\n| Messaging & watching | [`send`](#send) | Send one message, then exit: DM a peer, post a channel, or ask a role |\n| Messaging & watching | [`channels`](#channels) | Inspect or set the channel registry |\n| Messaging & watching | [`history`](#history) | Clear retained message history |\n| Messaging & watching | [`console`](#console) | Live protocol view for a space (TUI, or `--plain` line stream) |\n| Messaging & watching | [`web`](#web) | Browser dashboard (installed as the `@cotal-ai/web` extension) |\n| Auth & meshes | [`mint`](#mint) | Mint a creds file for a space (static auth mode) |\n| Auth & meshes | [`login`](#login) | Sign in to a per-user-auth mesh's IdP (once per machine) |\n| Auth & meshes | [`logout`](#login) | Revoke the IdP session and clear the cached login |\n| Auth & meshes | [`actor`](#actor) | Manage a user-auth space's actor ledger (grant / revoke / list) |\n| Auth & meshes | [`doctor`](#doctor) | Credential-health diagnosis and repair (`doctor auth`) |\n| Auth & meshes | [`join`](#join) | Join a space as your own presence (interactive) |\n| Manifest | [`topology`](#manifest-deploys) | Validate and view a mesh manifest's access graph (read-only) |\n| Extensions & misc | [`ext`](#ext) | Install / remove operator CLI extensions |\n| Extensions & misc | [`completion`](#completion) | Print or install shell completion |\n| Extensions & misc | [`feedback`](#feedback) | Send feedback to the Cotal developers |\n| Extensions & misc | [`deliver`](#server-daemons) | Run the server-side Plane-3 delivery daemon |\n| Workflow runs | [`run`](#run) | Operate durable workflow runs: start, resume, list, inspect, answer a checkpoint, check an edited program with migrate |\n| Extensions & misc | [`feedback-intake`](#server-daemons) | Run a self-hosted feedback intake server |\n\nThe manifest modes of `up`, `spawn`, and `down` (`-f <cotal.yaml>`) plus `topology` are covered\ntogether under [Manifest deploys](#manifest-deploys).\n\n## setup\n\n```bash\ncotal setup [--full] [--demo] [--yes] [--skills]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--full` | off | Redo the full guided flow (implies `--demo`) |\n| `--demo` | off | Also seed the guided expert team (`david`, `sven`, `me`) |\n| `--yes`, `-y` | off | Non-interactive accept-all (for agents / CI) |\n| `--skills` | off | Reconcile Cotal skills only through installed connector providers, plus `~/.agents/skills`. Refused with `--full` or `--demo`. |\n\nGuided setup is **configure-only**: it checks prerequisites, invokes installed connectors' declared setup providers, and\nseeds persona files, and it launches nothing (no mesh, no web, no manager). First run gets the\nnarrated flow; later runs print a status card. By default it seeds one `default` persona; the\n`david`/`sven`/`me` team is opt-in via `--demo`. `cotal status` points stale Claude skills and\nout-of-date `.agents` skills at `cotal setup --skills`, not unscoped `setup`. See [Getting started](getting-started.md) and, for\nmaintainers, [setup internals](setup-internals.md).\n\nWhen a mesh resolves, setup seeds that mesh's recorded `.cotal/agents` catalog, the same catalog a\nfollowing `cotal spawn` reads. It prints the absolute destination. On a fresh machine with no mesh it\nuses this folder and says why; when several meshes are available and none is selected, it refuses\nrather than choosing a catalog.\n\n## update\n\n```bash\ncotal update [--self] [--space <s>] [--server <url>] [--creds <path>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--self` | off | If a newer release exists, install that exact validated `cotal-ai` version globally and reconcile through the newly installed binary |\n| `--space`, `--server`, `--creds` | resolved mesh | Select the running manager whose continuity state is reported |\n\nWithout `--self`, `update` keeps the installed first-party surfaces coherent with the running\nbinary: it force-reconciles the four built-in connectors, then reinstalls other `@cotal-ai/*`\noperator extensions at the binary's exact version. Each extension runs in an isolated child, so one\nfailure cannot poison later replays. It then checks npm; a newer binary is an informational notice\nwith `cotal update --self` as the next command, not an automatic install.\n\nAfter disk reconciliation, `update` reads the selected running manager. A machine with no recorded\nmesh has no running manager to observe, so that read is skipped and the command completes. The same\nholds when every recorded mesh is down and none is selected. A remote user mesh, registered with\n`cotal meshes add --mode user`, is named and skipped: its manager runs under another install, so\nthere is no custody on this machine to preserve, and a `legacy` verdict still comes only from a\nmanager this machine read. A recorded mesh that is down is still a\nrefusal when the command selects it, with `--space` or by running inside its project, and so is a\nnamed space that is not running. With several meshes running and no `--space`, `--server` or\n`--creds`, the install is machine-wide, so every running manager is reported in turn, each under its\nspace name, before anything is written; a `legacy` verdict on any of them makes the whole run not a\nhot update. A selector flag still reports one manager. A manager without a\ncustody generation is reported as `legacy`: it cannot preserve its manager-owned PTYs, so the\ncommand says that this is not a hot update and prints `exact`, `fork`, `fresh`, or `drain-only`\nfor every seat. This report sends no stop, preservation-commit, or replacement command.\nIt does not preserve a running PTY on a legacy manager. The built-in pty runtime spawns\nin-process on every platform and reports `legacy`. On Linux it still adopts seats that an earlier\nmanager left under a detached custodian, but it starts no new custodian. An incompatible native\n`@lydell/node-pty` or ConPTY ABI break remains an explicit per-seat maintenance cut.\n\nWith `--self`, the selected running manager is reported before any global install. When a newer\nrelease exists, Cotal then installs the exact version it validated, resolves and verifies that\npackage in npm's global root, then launches that binary with the same `--space` / `--server` /\n`--creds` selection to reconcile connectors and first-party extensions to the new generation. An npx\nor dev-clone invocation therefore installs and continues through a separate global copy; it never\nclaims the already-running process changed. If the binary is current, `--self` performs the normal\nlocal reconcile without reinstalling it.\n\nThird-party extensions are listed with their installed version and recorded spec but are not\nauto-updated in v1. Floating third-party updates require `@cotal-ai/*` peer-range validation and are\na future follow-up. A failed connector/extension install, npm metadata check, or requested global\ninstall is reported and makes the command exit nonzero. Independent extension attempts continue so\nthe output includes every failure; an unavailable npm registry does not undo a completed local\nreconcile, but the command still exits nonzero because it could not establish that the install is\ncurrent.\n\n## up\n\n```bash\ncotal up [--detach] [--open] [--space <s>] [--server <url>] [--channels <path>] [--runtime <name>]\ncotal up --user-auth --idp <url> [--exchange-public-port <n> --exchange-public-url <https://\u2026> [--exchange-trusted-proxy]]\ncotal up --tls-cert <cert.pem> --tls-key <key.pem> # serve broker TLS (both, or neither)\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\ncotal up -f <cotal.yaml> [--dry-run] [--runtime <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--server <url>` | auto (free local port) | Listen URL override |\n| `--host <host>` | none | Bind host override for a **fresh** broker boot: an IP or hostname only, never a URL (that is `--server`) and never `host:port` (the port comes from `--server` or its default); a URL or port-bearing value is refused pointing at the right flag. With no `--server`, the broker URL is derived from it, so `--host <addr>` alone is enough to make a mesh reachable at that address; a `--host`/`--server` pair naming different addresses is refused. A wildcard bind (`0.0.0.0`, `::`) keeps a dialable loopback URL. Recorded on the mesh and reused by every later manager launch, so a repair or resume keeps remote [`attach`](#managed-seats) working. A live refresh (`\u2713 mesh already running`) does not rewrite `.cotal/auth/server.conf` or rebind nats; stop the broker, then re-run `up --host` |\n| `--space <s>` | the folder's name | Space name |\n| `--store-dir <dir>` | none | JetStream store directory (recorded; a repair up reuses it) |\n| `--max-file-store <bytes>` | nats-server's dynamic cap | JetStream file storage cap in bytes (a positive integer, no unit suffix). Without it nats-server sizes the store at start as three quarters of the free space on its filesystem. The cap is fixed at broker start: a running broker cannot change it (`cotal down` first), `down --preserve-state` keeps it for the resume, and a resume with a different value is refused. Not accepted with `-f` |\n| `--channels <path>` | `.cotal/channels.json` if present | Channel-registry seed file (JSON). An explicit path that is missing is an error |\n| `--restore <dir>` | none | Restore a completed offline backup before exposing the normal listener |\n| `--restore-only registry` | artifact selection | Restore only the registry component |\n| `--accept-missing-source` | off | Explicit disaster consent when the inode-bound preserved source is absent |\n| `--accept-stale-checkpoint` | off | Explicit consent to resume a seat whose checkpoint was captured outside its recorded recency horizon |\n| `--open` | off (auth) | Unauthenticated dev mesh: no JWT, no ACLs |\n| `--user-auth` | off | Per-user auth: people `cotal login`; connects are authorized against the actor ledger |\n| `--idp <url>` | none | With `--user-auth`: the IdP auth base URL to pin on first enable |\n| `--exchange-public-port <n>` | none | With `--user-auth`: add the public exchange face on this loopback port, for an HTTPS reverse proxy to forward to |\n| `--exchange-public-url <https://\u2026>` | none | With `--exchange-public-port`: advertise the reverse proxy's HTTPS URL in discovery |\n| `--exchange-trusted-proxy` | off | With `--exchange-public-port`: attribute public failure buckets to the last `X-Forwarded-For` hop. Enable only when the listener is reachable solely through a trusted proxy; otherwise the socket address is used |\n| `--detach` | off | Run in the background (stop with `cotal down`) |\n| `--tls-cert <path>` | none | PEM certificate to serve TLS with. Must be given together with `--tls-key`. Before starting the broker, Cotal checks readability, private-key mode, key/certificate match, the validity window, and host coverage. `nats-server` accepts an expired certificate and leaves the failure to clients, so Cotal performs these checks first. The decision is recorded; a later bare `cotal up` keeps serving TLS |\n| `--tls-key <path>` | none | PEM private key for `--tls-cert`. Refused if group- or other-readable (tighten to `600`) |\n| `--file <cotal.yaml>`, `-f` | none | Launch a whole mesh from a manifest |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--runtime <name>` | `pty` (or the manifest's, with `-f`) | Agent runtime for the mesh manager (`pty` built in; others are installed extensions, explicit-only). Resolved + probed before the broker starts; an uninstalled/unreachable runtime fails loud. With `-f`, overrides the manifest's runtime |\n| `--max-sessions <n>` | 64 | Live-session ceiling for the mesh manager. Each console pane and each `cotal attach` is one session, so size for agents \xD7 panes, not agent count. Recorded on the mesh and reused by every later manager launch, so a repair or resume does not silently drop back to 64. A running manager cannot change it: `cotal down` first, then `cotal up --max-sessions <n>` |\n| `--no-manager` | off | Broker-only boot: start the broker and, in auth mode, the delivery daemon, and no local manager. A refresh under the flag of a mesh whose manager is live refuses rather than keeping or stopping it (`cotal down manager` first). Cannot be combined with `--runtime`, `--max-sessions`, or an agent-declaring manifest |\n| `--rotate-sys` | off | Rotate the space's system account and re-mint its two `$SYS` creds. Needs a stopped mesh; refused with `--open` |\n\n`cotal up` boots a local nats-server with JetStream and, in auth mode (the default), JWT auth and\nper-agent ACLs; `--detach` records the mesh so `cotal spawn` from any directory can find it. With no\n`--server`, it auto-selects a free port if the default address is taken; an explicit `--server`\nstays fail-loud on collision. `--detach` also brings up the control plane (delivery daemon in auth\nmode, then the manager). `--no-manager` is the broker-only mode: it boots\nthe broker (and the delivery daemon in auth mode) and starts no manager, so there is no manager\npidfile to leave stale. A refresh under the flag of a mesh whose manager is live refuses rather\nthan keeping or stopping it: `cotal down manager` first. For a split topology with a manager, wait for `.cotal/manager.<spaceKey>.log` to contain `\u2713 manager up`, then `cotal down manager` on that\nhost and run [`supervise`](#supervise) against the remote broker; see\n[Run a mesh](run-a-mesh.md). `cotal up --detach` prints `\u2713 running in the background:` with\n`manager` listed (pidfile liveness, not a teardown boundary); with `--no-manager` the line lists\nonly what actually started. Ctrl-C on a foreground `up` stops the manager through the same stop as\nbare `cotal down` (see [`down`](#down)), then the rest of the stack, and reports managed agents\nunder the same rule: when the manager stop is refused, Ctrl-C prints the refusal with the reap route\nand leaves the stack running. The `-f` form is a\n[manifest deploy](#manifest-deploys).\n\nA repair `up` on a mesh whose broker died reopens the store its record names, and refuses a\ndifferent `--store-dir` rather than silently opening a second store.\n\nThe generated `.cotal/auth/server.conf` is written on a real broker boot and is not an\noperator-owned config. `--host` changes that file only when nats is actually started. A unit\nrestart that leaves an answering listener in place is a refresh, not a rebind.\n\nOn an existing mesh, `cotal up` reconciles the presence and lease bucket TTLs. It writes a reserved\ncanary and waits for the bucket to expire it before reporting success. If the broker accepts the\nstream update but the backing store does not persist or enforce it, `up` exits nonzero with a TTL\npersistence error instead of trusting the value returned by stream info. A refresh that restores a\nmissing manager says so with its pid (`\u2713 restored in the background: manager (pid N)`); a refresh\nthat finds everything already running prints only the `\u2713 mesh \"<space>\" already running` line. A\nfirst boot starts its manager without the restore line.\n\n\n`--user-auth --idp <url>` starts the space's auth service alongside the broker: the NATS\nauth callout plus its capability-gated local exchange, and optionally the closed public exchange\nface configured by the three `--exchange-*` flags above. The service is torn down with `cotal down`,\nand a re-run of `cotal up` heals a dead service on a running broker. `up` waits for the service to\nfinish binding: while the daemon it launched (or found running) stays alive, the wait extends past\nthe base 15s up to 60s; a daemon that exits is refused at once with \"exited before becoming ready\",\nand one alive past 60s is refused as \"alive and still starting\" (wedged), naming the pid record and\nthe service log. `--user-auth` and `--open`\ncontradict each other and are refused loudly; a running broker cannot change auth mode\nwithout a `cotal down` first. See [identity & auth](identity-and-auth.md).\n\n`--rotate-sys` renews the two `$SYS` credentials (`membership-observer`, `connection-evictor`).\nThey carry a 30-day expiry and nothing re-signs them in place, because the system-account seed is\nnever persisted, so they are renewed by issuing a **new system account** under the same broker\noperator and minting fresh creds against it. A plain re-`up` does **not** do this: it reuses the\nexisting trust record, and its `$SYS` creds along with it.\n\nThe rotation is safe to run on a real space, with one operational cost. The data account, the account\nsigning key, every agent credential minted from it, and the JetStream store are all untouched; what\ndies is the retired system account, and with it any out-of-band copy of the old `$SYS` creds, on every\nbroker that loads the rotated config. The cost is that **earlier full backups stop being restorable**\n(see below), so this is not a no-consequence operation. It needs the broker to restart on the rewritten\nconfig, so it runs as part of a boot:\n\n```bash\ncotal down\ncotal up --rotate-sys --detach # agents reconnect; nothing is re-provisioned\ncotal doctor auth # both $SYS creds healthy again, 30 days out\n```\n\nA rotation is a stopped, fresh boot, and anything that is not one refuses it, all for the same reason\n(the on-disk material and the broker it runs on must never end up on different generations):\n\n- a live mesh, because the running broker would keep serving the retired account;\n- an open mesh, whether that comes from `--open` or from `broker.auth: false` in a manifest, which\n has no system account at all;\n- `--restore`, because reinstating a trust root and superseding it in one command leaves no way to\n say which authority the mesh came up on;\n- an unfinished restore or resume attempt on this root, including one `cotal up` would recover on\n its own, because those paths can adopt a live listener and return without booting a broker;\n- a root that hosts more than one space, because the system account lives in the shared broker\n record and a rotation would retire every tenant's, while the root holds one `$SYS` cred pair\n pinned to one data account.\n\nTwo things to know before you run it:\n\n- **The retirement is config-load-bound.** Old `$SYS` creds are refused by any broker that loads the\n rotated config. A stale `nats-server` still running the *previous* config in memory would keep\n honouring them, so stop every broker for this root first. `--rotate-sys` refuses if this root's\n mesh is recorded as running, if anything unidentified is answering at the address it was given, or\n if the root's pid file names a live (or unreadable) process. Those are Cotal's own ownership\n records, not a scan of the process table: a `nats-server` you started by hand against this root's\n `server.conf` on some other port writes none of them and will not be seen. Do not run one.\n- **It invalidates earlier full backups.** A full artifact binds to the trust chain it was taken\n against, and that commitment covers the operator JWT and the system account. Every full backup\n taken before a rotation refuses to restore afterwards, so take a fresh `cotal backup` once the\n rotated mesh is up. `cotal up --restore` names this case when the data account still matches.\n\nThe commit is not atomic (a trust-record write plus two credential writes), so an interrupted\nrotation leaves the record ahead of the creds. That split is detected rather than silent: every\n`cotal up` on an auth mesh, and every `cotal doctor auth`, compares each `$SYS` cred's issuer against\nthe persisted record and names the retired account. `up` warns rather than refusing, because these\ncreds power the membership graph and live eviction, both of which degrade fail-soft; the mesh is not\nworth taking down over them. Re-running the rotation heals it, at the cost of one generation.\n\nWhile those creds are expired the mesh keeps delivering messages, but the\n[membership feed](delivery-daemon.md) and live connection eviction stay down; `cotal doctor auth`\nand the manager's log both name the credential and this repair.\n\n## down\n\n```bash\ncotal down\ncotal down --with-agents\ncotal down --preserve-state [--store-dir <dir>] [--session-store <dir> \u2026]\ncotal down manager [delivery auth web nats ...]\ncotal down web [--space <name>]\ncotal down -f <cotal.yaml> | --run <id> [--dry-run]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--file <cotal.yaml>`, `-f` | none | Tear down this manifest's deploy |\n| `--run <id>` | none | Tear down one `spawn -f` run by id |\n| `--space <name>` | current mesh | With components: the mesh whose target-addressed components (e.g. `web`) to stop |\n| `--dry-run` | off | Print the manifest teardown or selected components, mutate nothing |\n| `--with-agents` | off | Bare whole stack only: also stop and deprovision every managed agent |\n| `--preserve-state` | off | Bare whole stack only: fence the manager, retain principals and durable state, stop and prove the stack down, then publish `ready` |\n| `--store-dir <dir>` | `.cotal/nats` | With `--preserve-state`: the actual store path (required for a custom store) |\n| `--session-store <dir>` | none | With `--preserve-state`: a harness transcript store directory to capture with every continuation-capable retained seat. Repeatable. No default and never inferred from a connector name; a path that does not exist or is not a directory is refused before anything stops |\n\nBare `cotal down` stops the whole local stack in dependency order and leaves managed agents running\nwhen their runtime lets them outlive the manager. Before signalling the manager it verifies the spare\ncapability of the exact recorded manager, which records what that manager's stop does with its\nseats, and it reports the agents left behind plus `cotal down --with-agents` as the explicit reap.\nWhen the manager had no managed agents, it prints no report.\nThe built-in pty runtime keeps each PTY inside the manager process, so those seats cannot outlive\nit: every manager stop stops and deprovisions them, and `down` reports them as stopped. Every manager\nstop the CLI makes runs this one path: `down`, Ctrl-C on a foreground `cotal up`, the teardown after\nthat `up`'s broker exits, the leftover-manager stop before `cotal up -f`, and the delivery cutover.\nEach holds the manager's stop reservation, so a second stop while one is in flight is refused,\nnames the process holding it, and leaves that stop's `--with-agents` policy in place. Each sends\n`SIGKILL` to a manager still running 15s after `SIGTERM`. Ctrl-C stops the manager first; when that\nstop is refused or the manager's exit cannot be confirmed, Ctrl-C signals nothing else, prints the\nrefusal with the reap route, and leaves the stack running; end it with `cotal down --with-agents`.\n`--with-agents` is a one-shot destructive policy bound to the exact verified manager process\nand the exact live `down` stop reservation; a stale, malformed, crashed, or different stop attempt\ncannot turn a later bare shutdown destructive. If a managed agent cannot be proven stopped within\nthe manager's stop timeout, the manager logs which one, still closes its broker connections and\nconsole listener, and exits with code 1. It does not release its pidfile, liveness lease or\nservice registration in that case, so no successor is handed authority while that agent may still\nrun; the lease lapses on its TTL. Positional component names stop\nonly those self-registered local processes; for example, `cotal down manager` leaves delivery and\nthe broker running, and `cotal down web` is available when the web extension is installed. A\ncomponent that starts target-resolved (the web dashboard) is stopped the same way: `cotal down web`\nresolves the mesh the same way as `cotal web` (registry current mesh first, `--space` to name one), so\nit works from any directory; the other components always stop under the folder you run it in. The\n`-f` / `--run` forms tear down a [manifest deploy](#manifest-deploys) without stopping the whole mesh\nand cannot be combined with component names. Stopping `nats` alone is refused while an unselected\nregistered daemon is still live; include those components or use bare `cotal down`.\n\nA pinned manager with no spare-capability record is not signalled by bare `cotal down` or `cotal\ndown manager`. A current manager always publishes the record, so a missing one means an older\nmanager: one that predates capability reporting, or one whose pty runtime reported that it cannot\ndetach its agents. Stop each managed agent explicitly, then run `cotal down --with-agents` from the\nmesh root to stop the whole stack. An older manager does not understand\nthe reap request, which is why the agents must already be stopped.\n\nBare `cotal down` inventories by pidfile. When this folder's registered broker answers and no\n`nats.pid` records it, the command does not say nothing is running. It names the space and the\nbroker address, says no pidfile records that process, says it will not stop a process it did not\nstart, and exits 1. Stop that broker with whatever started it (an init unit, a container, or the\nhand-run process). `cotal meshes rm <space>` only drops the registration. The probe runs whether or\nnot other owned components were running: they stop and clear their artifacts first, then the broker\nis named. A component stop and `--dry-run` stay pidfile-only and do not probe.\n\n`down` reads each process record once. A component that exits and removes its own record while\n`down` runs counts as having no record, so the stop goes on. Any other failed read is an error.\n\n**Teardown verifies pinned process identity before signalling.** PIDs are recycled by every OS,\nso a recorded pid alone is not a durable target identity. `up` and `cotal web` record each\nprocess's creation identity in a sibling `<pidfile>.identity` pin, which holds the pid and the\nprocess start reported by the OS. Every stop path, including `down` for the broker, web and\nextension components, and the manager, delivery and auth-service stops, applies the same rule. A pin\nthat names a different start means the pid was reused, so teardown refuses and preserves it. A torn\nor unreadable pin also refuses.\n\nThe pidfile and its pin are published by renames, and the pidfile rename is the commit point. Just\nbefore it, the pin holds two lines: the old process's and the new one's. A launcher that dies\nmid-publish therefore leaves the old record or the new one, each checked against its own pin line,\nnever a pidfile without its pin. An old record with no pin is legacy, so its line holds `-` in place\nof the token and it stays legacy until the commit. A CLI older than this change reads a two-line pin\nas torn and refuses.\n\nPublishes of one pidfile are serialized by a lock file beside it, `<pidfile>.publish.lock`, because\nthe launcher and the daemon it starts both publish the same record. The next publisher reclaims a\nlock left by a crashed one. When no start token can be read for the new process, its pin line holds\n`-` in place of the token, which reads as a legacy record, and the publish ends in the legacy shape:\na pidfile with no pin. Teardown, and a daemon removing its own record on exit, take the same lock and\nremove the record only while the pidfile still names the pid they stopped, so a stop that races a\npublish leaves the new record whole.\n\nThe web dashboard claims `web.pid` with an exclusive create, so a second dashboard for the same mesh\nis refused, and writes its pin right after the claim. A stop that runs between the two reads a\nlegacy record.\n\nThe pidfile pid and the pin pid are two coordinates. Automatic cleanup follows **proven death of\nthe pidfile target** (ESRCH on that pid): a torn sibling pin does not wedge a dead pidfile pid.\nA torn pairing where the pin names another pid, while the pidfile pid is still live or not proven\ndead, still refuses. Inspect both pids with `ps`. Do not delete `<pidfile>.identity` to force a\nstop; that weakens target-identity protection. Once the pidfile process is dead, rerunning\nteardown clears the stale record automatically.\n\nThe first teardown after upgrading a running pre-pin stack has a narrower guarantee. A live record\nwith no identity pin is signalled after a loud warning that it predates identity pinning. Restarting\nthe component writes the pin, so later teardowns receive full match and mismatch protection. The\nsame warning applies on platforms where no stable start token is available. For a legacy manager,\nbare `cotal down` also warns that agent sparing cannot be verified before it signals. Because the\nCLI cannot establish which SIGTERM handler that already-running binary carries, it never presents\nthe pre-signal seat inventory as confirmed spared; a genuinely older destructive handler may still\nreap those agents. `--with-agents` publishes a one-shot reduced-guarantee handoff bound to the\nrecorded manager pid and the live `.stopping` reservation's inode, then signals unconditionally.\nThat handoff cannot be replayed by a later stop attempt. A pin that exists and does not match the\nlive process still refuses before signal.\n\n`--with-agents` performs the old destructive logical teardown: managed processes stop and their\ncredentials, ACL rows, and delivery footprints are deprovisioned. `--preserve-state` is a different\nmaintenance transition: it stops retained processes while suppressing leave/deprovision cleanup, persists the manager's\nsame-principal resume inventory, stops the entire stack without removing run/auth artifacts, and\npublishes a stable inode-bound cut only after every recorded process is proven stopped and the exact\nrecorded NATS endpoint is unreachable. A missing or stale broker pidfile never counts as stopped. The\nattempt is bound durably before the manager is fenced, the resume document and attempt-bound\n`cut-intent` are fsynced before manager commit, and the manager's commitment itself is journaled\n(`cut-committed`) before any process stops. A retry after a crash at any of those boundaries reuses\nthe exact recorded attempt and finishes the remaining stop and endpoint proofs idempotently, without\nneeding the (by then intentionally dead) manager. A partial cut never publishes `ready`. It cannot\nbe combined with component names, manifest teardown, or `--dry-run`.\n\n**Seat checkpoints.** After the stack is proven down, the cut writes one checkpoint per retained\nseat under `.cotal/maintenance/v1/checkpoints/<attempt>/<seat>/`, and prints the path, the\ncontinuity class and the generation for each. The path carries the preservation attempt because a\ncheckpoint is immutable once sealed: a shared directory would make the second cut in a root refuse\non the first cut's leftovers, and clearing it would destroy an artifact a rollback still needs. The\ncapture happens only at that point because anything earlier races a harness that is still writing\nits transcript and its working tree.\n\nEach checkpoint directory is created 0700, refuses a destination that already exists, and holds:\n\n- `repo.bundle`, the seat `cwd`'s reachable history, anchored on the base commit the record names\n by full object id;\n- `repo.index.diff` and `repo.worktree.diff`, the staging state as two diffs, base to index and\n index to worktree. Two rather than one because a single combined diff restores a mixed tree with\n the right bytes and the wrong index: a source reporting `MM README` would come back as ` M README`;\n- `repo.untracked.tar`, the untracked files in scope;\n- the harness session pointer, when the seat's connector declares one, and the transcript store\n files the operator named with `--session-store`. Each records where the destination puts it back\n as an anchor (the workspace root, the account home, or the seat's `cwd`) plus a relative path,\n because the destination's root and home are its own and the source host's absolute spelling would\n either miss them or write outside them;\n- `checkpoint.json`, written last, after every digest is computed over the bytes that landed.\n\nThe record carries the manager's resume entry unchanged as its first field, then the space, the seat\nname, the recovered `lifecycleUid`, the writer generation the cut was taken at, `capturedAt`, the\nrecency horizon, the applied profile revision, the seat's `git status --porcelain` as the cut read\nit, and the continuity class. Every captured file is\nlisted with its byte size and sha256, so an operator verifies the whole artifact with `sha256sum`\nand `git bundle verify`. No secret values, no operator keys and no source-host launch material\nenter it.\n\nThe continuity class is what the connector declares, capped by what the checkpoint carries. A\nconnector declaring session continuation classifies as `exact`, but reopening a session takes both\nhalves, the pointer that names it and the store that holds its transcript. A checkpoint missing\neither one cannot reopen that session, so it is recorded as `fresh` when the connector declares a\nfresh start and `drain-only` otherwise. A pointer with no store is capped the same way as a cut\ncarrying neither, because it names a session whose bytes the artifact does not contain. A class is a promise the destination is entitled\nto act on, so it never describes bytes the artifact does not contain. The transcript store stays an\noperator input: this repository does not know where a harness keeps its transcript, so `exact`\nrequires `--session-store` to name one.\n\nThe recorded status is read under the same selection rule as the untracked set, so it describes the\nstate the captured bytes can reproduce. The destination re-reads it in the promoted tree and refuses\na difference.\n\nThe untracked selection rule is recorded in the record and is\n`git ls-files --others --exclude-standard -z, excluding .cotal/`. It honors `.gitignore`, so an\nignored file the seat needs does not travel and has to be moved separately. The `.cotal/` exclusion\nis a secrecy boundary rather than a size one: when a seat's `cwd` is also the mesh root, the control\ndirectory is untracked, and without the exclusion the broker trust material, the space account, the\nmanager instance identity's private seed and the seat's own credentials would land inside the\nartifact. A checkpoint carries credential references only; the destination resolves that material\nitself.\n\nA seat whose launch options could not be resolved is refused rather than checkpointed, with the\nmanager's own wording: `imperative launch options have no non-secret durable source (<keys>)`. The\nrefusal arrives at prepare time, so the cut stops before any child does.\n\nA delegated seat (SPEC \xA713.17) is refused at prepare time too, with\n`a delegated seat is not resumed by a later manager; stop it before preserving`. A manager stop\nafter a refused cut retires that seat through its retirement path.\n\n## clean\n\n```bash\ncotal clean <history|store|all> --force\ncotal clean restore-attempt --attempt <id> --force\ncotal clean restore-fallback --attempt <id> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | `history`: target mesh |\n| `--dms` | off | `history`: also clear DM history |\n| `--store-dir <dir>` | `.cotal/nats` | `store`/`all`: JetStream store directory |\n| `--force` | none | Required: destructive, no prompting |\n| `--attempt <id>` | none | `restore-attempt`: exact stale pre-commit attempt; `restore-fallback`: matching healthy committed restore |\n\nOne configurable cleanup verb; every target requires `--force`.\n\n- `history` purges the retained message backlog on the **running** broker (channels, plus DMs\n with `--dms`). The same operation as [`history clear`](#history), which stays as an alias.\n- `store` deletes the **stopped** mesh's JetStream store (`.cotal/nats`): streams, durable\n consumers, and messages. This is the reset for stale on-disk broker state, e.g. durables\n minted by an older, incompatible Cotal generation surviving a `down`/`up` cycle.\n- `all` is `store` plus the space identity (`.cotal/auth`), the local creds and markers tied to\n it, any crash residue a normal `down` would have swept (stale pidfiles, `run/`), and the mesh's\n registry entry; the next `cotal up` mints a fresh identity.\n\n`history` needs the mesh up; `store` and `all` refuse while any recorded mesh process is still\nalive or any same-root recorded broker endpoint remains reachable (run `cotal down` first). They\nalso refuse outright on a root that holds accounts for several spaces: the store and the broker\ntrust record are shared by every space on the broker, so both targets would take out all of them\nand no `--space` can narrow that. `down`, `backup` and `up --restore` refuse there for the same\nreason. `cotal status` lists the tenants on such a root. Personas\n(`.cotal/agents`) and logs are never touched. The mesh record now carries a custom\nstore location for `up`'s own repair, but `clean` still takes `--store-dir` itself; `clean` does\nnot read the record. Custom cleanup targets must contain either the Cotal store-generation marker or a\nreal `jetstream/` store directory; filesystem roots, project roots, and Cotal auth/maintenance trees\nare always refused.\n\n`store` and `all` also refuse every maintenance journal state. After a healthy committed restore,\n`restore-fallback` is the only supported way to remove the recorded unchanged old-store inode; it\nnever deletes the active target, requires both the exact attempt id and `--force`, and retires the\ncompleted restore journal so a later `down --preserve-state` can start a new backup cycle.\n\n## Backups\n\n```bash\ncotal down --preserve-state [--store-dir <dir>]\ncotal backup create <dir> [--only full|registry] [--store-dir <dir>]\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\n```\n\nBackup is offline-only. It requires the stable `ready` record from `down --preserve-state`, an exact\nstore match, no live recorded process, and an unreachable exact endpoint from the recorded cut.\nThat endpoint is probed immediately before cloning, so a live broker with a missing or stale pidfile\nis still refused. It claims the cut, reflink/copies the stopped source to a\nprivate attempt clone, and opens only that clone on a random loopback bootstrap broker with an\nindependent parent/deadline watchdog. It validates the canonical stream and pull-consumer inventory,\nwrites native snapshots with consumers excluded, and stores conservative contiguous ACK-floor\ncheckpoints separately. The presence bucket is memory-backed, so it does not survive the cut and\nthe clone may lack it. Every other stream must be present. The original store is never opened by\nthe backup broker, and the stack is not restarted implicitly. Artifact destinations must not overlap\nthe preserved source or maintenance\nattempt tree. Restore artifacts and targets likewise cannot nest inside or contain each other, the\npreserved source, or the maintenance attempt tree.\n\nStopped client-managed KV ordered consumers are ephemeral read residue, not backup state. Backup\nignores only the pinned client's exact stopped shapes: ordinary last-value watchers and the\nwhole-bucket scanner that uses all-history delivery to collapse concurrent tombstones. A bound\nconsumer or any lookalike with a different filter, inbox, lifetime, or other config is still refused.\n\n`full` is the default and indivisible: channel registry, CHAT/DM/TASK/INBOX/DLV, ACL, MEMBERS, and\nvalidated durable checkpoints. `registry` is the sole partial artifact. Presence, derived membership\nfeed, leases, native ephemeral/history consumers, credentials, keys, tokens, owner secrets, and actor\nledger files are excluded. `full` means every transferable message and registry stream, not every\nJetStream resource: endpoint submissions/facts/events/timers/workflow state, contract artifacts, and\nthe records/auth/session stores are nonportable control state. Restore recreates those streams empty\nwith their canonical configs before exposing the normal listener, so active endpoint runs,\nlifecycles, and sessions do not cross a backup. Artifacts are exclusively created `0700`;\nsnapshot/checkpoint files and\nthe manifest are `0600`; `manifest.json` is written last with exact sizes and SHA-256 values. The\ndirectory is trusted operator input: hashes detect corruption, not malicious rewriting.\n\nRestore validates and stages the exact allowlisted artifact bytes before moving or creating a store.\nIt requires the same space and existing trust state. The whole pre-commit window holds a journaled\nliveness claim (coordinator, watchdogs, brokers, absolute deadline): ordinary `up` and a repeated\n`up --restore` refuse while the claim is live, and a stale attempt is recovered only after the\ndeadline has elapsed and every recorded owner is proven dead. A retried `up --restore` handles this\nautomatically; an operator can also recover it explicitly with `cotal clean restore-attempt --attempt <id> --force`. Nothing\never rolls back a live attempt. A registry-only artifact restores as registry-only whether or not\n`--restore-only registry` is passed; omitted infrastructure is always created and the exact\npost-restore stream inventory is asserted before commit intent. Ordinary `up` from a preserved cut\nresumes only the exact recorded source store and runtime; a contradicting `--store-dir` or\n`--runtime` fails in preflight.\n\n**Admitting a seat checkpoint.** An ordinary `up` from a preserved cut admits that cut's seat\ncheckpoints before it journals the resume attempt and before any process starts, so a refusal costs\nnothing. Three gates run in order, each naming what it saw.\n\n1. *Integrity.* Every file the record names must be present, a regular non-symlink file, the\n recorded byte size and the recorded sha256, re-stat'd after the read so a file that moved is a\n refusal. Failure here consults no other gate.\n2. *Identity.* The recorded space must match, the recorded `lifecycleUid` must not belong to a live\n incarnation, and the profile revision must match this host's or be resumed under deliberately\n this host's. A differing revision is refused with both digests and the remedy, and there is no\n override: the checkpoint carries the recorded digest and not the config bytes, so nothing could\n run the seat under the recorded revision, and the manager re-digests the same file and refuses\n drift on its own. This gate has no blanket override, which is the only reason the next one may\n have one.\n3. *Recency.* `capturedAt` is compared to this host's clock against the horizon the record carries.\n Inside it, the seat resumes. Outside it, `up` refuses and prints the capture instant, the clock\n reading and the horizon; `--accept-stale-checkpoint` admits it anyway and the exercised consent\n is printed with the actual age. An unreadable `capturedAt` is refused with no override, because a\n freshness gate that fails open is not a gate.\n\nCustody transfers only after all three pass. The destination claims the recorded generation plus one\nby exclusive create, before it launches anything. A lost create means another destination is already\nclaiming that seat, and it refuses with `seat-writer-generation-create-lost` rather than adopting\nthe winner and becoming a second writer. The recorded `lifecycleUid` is reused and never minted, so\nthe resumed seat binds the same lifecycle-keyed durables.\n\nAdmission is reconciled against the inventory the resume is about to hand the manager, and that\nreconciliation finishes before the restore moves a single tree. A checkpoint whose recorded\n`lifecycleUid` is not the one the retained inventory carries describes a different incarnation of\nthat seat, and it refuses with both uids while every live working tree is still untouched and no\ngeneration is claimed. A retained\nseat with no admitted checkpoint refuses the resume by name: an absent checkpoint directory and an\nabsent record are indistinguishable from a seat that was never checkpointed, and a seat that starts\nwithout passing the gates has claimed no generation. `--accept-stale-checkpoint` is recorded in the\nresume journal with the seat, the capture instant, the admitted age and the horizon, so the consent\nsurvives the terminal it was typed into.\n\nThe whole admission is all or nothing. Coverage is settled first, then every gate runs over every\ncheckpoint, and only then is any generation claimed. A refusal at any point leaves every generation\nunclaimed, including a lost exclusive create during the claim itself: the claims that attempt made\nare removed before the refusal is raised, by the exact paths it wrote, so a generation another\ndestination holds is never touched. A claim is a create that can never be made again, so a refusal\nthat left one behind would consume the retry over the same checkpoint set.\n\n**Restoring a seat checkpoint.** Once every gate has passed over every checkpoint, and before a\nsingle generation is claimed, `up` puts each admitted seat's captured bytes back. A refusal here\ncosts nothing for the same reason a gate failure does: no claim has been made and nothing has\nstarted.\n\nA restore never moves or replaces the destination's own control directory. A checkpoint excludes\n`.cotal/` by design, so a seat whose `cwd` holds one, which is the layout an operator gets by\nrunning `up` and `spawn` in a single directory, is refused before anything is staged: promoting a\ntree that cannot contain `.cotal/` over that `cwd` would carry this host's live trust material and\nmaintenance state away with the superseded tree. The refusal names the control directory it found\nand the remedy, which is to give the seat a working tree that is not a workspace root.\n\nEach seat is staged beside its own `cwd`, in `<cwd>.incoming`:\n\n1. every recorded digest is verified again over the files as they are now;\n2. the bundle is cloned into `<cwd>.incoming`, which is refused when that path already exists;\n3. the recorded base commit is verified in the clone and checked out detached, so a bundle that does\n not contain it stops the resume instead of continuing against a different history;\n4. the index diff is applied with `--index` and the worktree diff without it, both `--binary\n --allow-empty`. That order is what puts staged content back in the index rather than only in the\n worktree, and `--allow-empty` is why a seat with a clean tree is still restorable;\n5. the untracked archive is extracted.\n\nEvery seat stages before any seat is promoted. Promotion moves an existing `cwd` aside to\n`<cwd>.superseded.<timestamp>` and renames the staging directory into place, then puts the session\npointer and store files where the destination's connector reads them, then re-reads\n`git status --porcelain` in the promoted tree and compares it to the status the checkpoint recorded.\nA restore that applied without error and produced a different index is a refusal, not a warning. The\ntwo renames are the only steps that touch the path the seat will use, so a failure anywhere leaves\nevery seat's live `cwd` as it was.\n\nThe rename itself claims the superseded name, and a taken name gets a numeric suffix. The timestamp\nhas one-second resolution, so two promotions of the same seat within one second compute the same\npath; a rename onto a name that already holds a tree fails on every platform, and that failure is\nread as taken. Nothing creates the name ahead of the move, because Windows refuses to rename onto an\nexisting directory at all. A superseded tree is the thing that rename exists to keep.\n\n`git` and `tar` run as child processes with argument arrays, never a shell string.\n\nA leftover `<cwd>.incoming` refuses the resume by name. A staging directory from a failed run is the\nonly record of what failed, so nothing removes one automatically: inspect it, remove it by hand, and\nresume. A pre-existing `cwd` is renamed rather than deleted, so a wrong checkpoint costs a rename\ninstead of a tree. When a promotion fails, the renames that attempt made are undone and the staging\ntree is left where it is, as the evidence for what did not verify.\n\nA session pointer whose recorded `sessionId` is not the one the retained inventory reopens is\nrefused before anything is cloned. A session file already present at its destination is judged by\ncontent: bytes equal to the recorded digest are already restored, and different bytes under the path\nthe connector is about to read are refused with both digests rather than clobbered.\n\n`up --restore <dir>` reaches the same admission and the same restore, after the store is restored\nand validated and before commit intent is journaled. A registry-only restore resumes no seat, so it\nadmits and restores nothing.\n\nOne limit is worth stating plainly. The writer generation is claimed by exclusive create inside one\nworkspace root, so it fences two resumes on the same host and does not fence two independent\ndestinations: copy a checkpoint to two roots and both claim the same successor. A real cross-host\nfence needs a coordinate neither root owns.\n\nAuthenticated restores validate the complete\nspace trust bundle before staging, including nkeys, seed matches, JWTs, signers, and space binding;\nfull restores commit to the validated operator, system-account, data-account, and active-signer root\nchain in addition to the static/user authority fingerprint. Because the system account is part of that\ncommitment, a [`cotal up --rotate-sys`](#up) makes every full artifact taken before it unrestorable\nagainst this root: take a fresh full backup after each rotation. The composed commitment is revalidated\nimmediately before store mutation and never includes secret seeds. Restore never creates fresh auth.\nSame-path restores atomically retain the old\nsource at the journaled fallback path; alternate targets retain it in place; a missing canonical\nsource needs explicit `--accept-missing-source`. Quarantine and target restores use current canonical\nconfigs on isolated random-loopback brokers, never expose native snapshot consumers, and publish a\ncommit-intent immediately before the normal listener starts. Archive bytes never instantiate the real\ntarget: after quarantine validation, every stream is re-snapshotted from the validated quarantine\nstate into attempt-owned sanitized files, and the target is restored solely from those. Before that boundary, failure rolls back\nthe attempt-owned target; after it, ambiguity preserves both stores and records forward-repair\nrecourse. The cooperative maintenance lock excludes Cotal commands, not arbitrary raw NATS processes.\n\nBootstrap brokers in every auth mode, including open, mount the store under a local account with\nrandom operation-specific logins only, each carrying the exact per-phase subject permission matrix;\nnormal static credentials and user-auth sentinel/bearer connections are rejected, and no auth\nservice or callout starts. Open mode differs only in its account label, never in authority. Inventory, each stream snapshot,\nrestore initiation, exact upload id, validation, and each checkpoint recreation use separate exact\nauthorities. Every checkpoint carries the source stream's message/first/last sequence state and must\nmatch its snapshot record before mutation; core then derives and validates the only allowed start\npolicy. TASK is not a CLI exception: the same core checkpoint API recreates its canonical `DeliverAll`\nWorkQueue durable because acknowledged tasks are absent from retention and NATS forbids a\nstart-sequence policy there. Registry-only restore creates every omitted canonical stream and transient\nbucket on the isolated target before the normal listener is exposed. It deliberately does not resume\nretained agents or recreate their DM/DLV/TASK/ACL state; their identity material stays retained and\nstopped rather than being reprovisioned into a partial restore.\n\nAfter listener readiness, the manager starts attempt-bound, validates retained credentials/tokens\nwithout granting or reprovisioning, and resumes the exact persisted principals under cleanup\nsuppression. Registry-only restore uses the same flow with an empty agent set. On a user-auth mesh\nthese manager calls run as the logged-in operator's `cli` actor, the caller the preserve cut used, so\nthat actor needs a current `admin` grant. `commitResume` is an\nidempotent validation barrier only: success must be `awaitingFinalize` with an attempt-bound 64-hex\ncommit token and does not release suppression. Under the workspace lock, the CLI first fsyncs that\nexact evidence as `manager-committed` (restore) or `resume-committed` (ordinary resume), then calls\ntoken-bound `finalizeResume`; only an `active` response for the exact token releases suppression. The\nCLI records the same token in finalization evidence before a restore becomes `active`, or before an\nordinary resume retires and consumes the marker. Re-entry from either committed state skips the prior\nidempotent activation/commit phases, retries finalization with the durable token, and finishes the\nworkspace transition. Failure before finalization preserves the committed state and cleanup\nsuppression; it is not rewritten through a degraded transition. Re-entry between any two earlier\nboundaries reuses the same attempt and may retry the idempotent phases without deleting retained state. A missing or\nchanged per-agent dependency is a named fail-closed result; the journal becomes degraded and remains\navailable for forward repair. A retry from `resume-intent`,\n`resume-active`, or `resume-degraded` reuses the same attempt and inventory after the prior listener is\nproven stopped. A retained agent the lost manager already launched can still be running, for example\nin a tmux window, while the journal reads `resume-intent`. On a static mesh the replacement manager\ncloses that seat through the reference the lost manager recorded on the agent's slot, waits for the\nprincipal to leave presence, and launches it again. A live principal with no such record, or one that\nstays live after the seat is closed, is refused. Every normal restore listener has an unguessable\nattempt-bound NATS server name. The CLI fsyncs its exact name/nonce, canonical endpoint, process owner, and generation-bound target identity\nimmediately after spawn. Re-entry accepts a surviving listener only when its INFO server name, live PID\nrecord, endpoint, and target identity all match that proof; degraded restore repair then moves through\nthe guarded workspace transition only after manager commit. If an uncommitted bound owner is provably\ndead, recovery retires that exact proof under the maintenance lock and binds a fresh listener for the\nsame attempt, endpoint, and target with a new nonce and server name. A live foreign/mismatched listener\nor ambiguous owner is preserved and refused, never adopted by reachability alone. A reconstructed\ncommit/degraded attempt without either the exact bound proof or a durable dead-listener replacement\nrecord fails closed even when the recorded port is free. A later ordinary startup may pass an `active`\nrestore only when its details prove manager commit and its exact recorded listener is dead.\n\n## Mesh registry\n\n```bash\ncotal meshes [--json]\ncotal meshes add # guided, on a terminal\ncotal meshes add <space> --server <url> [--root <dir>] [--mode auth|open|user] [--tls] [--force]\ncotal meshes add <space> --mode user (--user-auth-file <bundle.json> | --from <https url>)\ncotal meshes rm <space> [<space> \u2026] [--force]\ncotal sync [--idp <auth base URL>]\ncotal use <space>\ncotal status [--space <s>] [--server <url>] [--components]\n```\n\n`meshes` lists the meshes this machine knows; a `*` marks the `current` default a bare\n`cotal spawn` joins. Entries learned from a signed-in account are marked `discovered`. Their\nregistration trust is stored under the account's private auth state, and the registry contains no\nsession token or sentinel credential bytes. Commands resolve the catalog `slug`; a different human\n`name` is rendered only as a label.\n\n`meshes --json` prints one JSON object per recorded mesh per line: `space`, `server`, `mode`,\n`root`, `default` (the `*`), and `origin` (`up`, `manual`, or `catalog` for a discovered entry). A\nlocal or hand-registered entry also carries `offline`. A discovered entry is never probed, so it has\nno `offline` field. `tlsRequired`, `events: \"required\"` and a discovered entry's `catalogName` appear\nonly when the record has them. An empty registry prints nothing and exits 0. The note about a default\nthat matches no record goes to stderr, so stdout carries only rows, on a first run too. The table is\npresentation and is not a stable parsing target. `meshes add` and `meshes rm` refuse `--json`.\n\nA registry record this build cannot use is refused by name, never rendered and never skipped. One\nthat does not parse, or is missing a field every consumer reads (`server`, `mode`, `root`, `ts`,\n`space`), makes every registry command exit 1 with the file's path and what is wrong with it.\nRemove the file or restore the record; nothing repairs or invents a field for you.\n\nAn IdP may advertise a same-origin space catalog during login. Cotal reads the complete snapshot and\nadds every valid registration without a separate `meshes add`. A snapshot younger than five seconds\nis used without a request. After that, commands that resolve a mesh target conditionally refresh the\nsaved catalogs. An operation targeting a discovered space refreshes only that space's account and\nrefuses if that account fails. Operations targeting local or manually registered meshes refresh every\naccount, print one warning for each failure, and continue. `cotal status` refreshes every account,\nnever refuses on a refresh failure, and lists each account as `fresh`, `updated`, `not-modified`,\n`no-catalog`, or `failed` with its error. `cotal sync` bypasses freshness and reports added, changed,\nremoved, unchanged, and name collisions. `--idp` limits it to one signed-in account. It never connects\nto a broker.\n\nThe registry is updated under the same lock that guards the catalog cache, so a command never lists\na discovered space set that another command is still writing. The cache records a fetched snapshot\nas not yet applied before the first registry write and as applied after the last. If a command dies\nor is stopped in between, the next command applies that snapshot again before it can use it, with\nno request inside the freshness window.\n\nThe shared dispatcher applies this preparation to every command that declares both `--space` and\n`--server` as mesh-target flags, including commands registered by other packages and commands that\ndeclare their own equivalent flag objects. Daemon and startup commands that use those names only as\nconfiguration explicitly opt out. Registry-local `meshes add` and `meshes rm` never refresh a catalog.\nWhile the registry holds a record this build cannot use, the preparation neither refreshes nor\napplies a catalog, so the command's own checks run first. A snapshot left unapplied is applied by the\nnext preparation after the record is restored or removed. A command that resolves its target through\nthe registry still refuses the record by name.\n\nRun on a terminal with the space or `--server` missing, **`meshes add` is guided**: it asks for the\none thing that cannot be derived (the broker URL), probes it, and tells you what answered - open or\nrequiring credentials. It then offers the spaces your `--root` already holds credentials for, states\nthe mode as a fact about that broker rather than asking, and shows the exact record before writing\nanything. A broker that does not answer, or a space name already registered, becomes a choice rather\nthan an error. Anything you pass on the command line is taken as given and not asked again. Without\na terminal - a script, an agent, CI - nothing prompts and the flag form's errors stand\n(`COTAL_NO_PROMPT=1` forces that too).\n\n`cotal up` and `cotal down` maintain their own records. `meshes add` registers a mesh they cannot\nspeak for: one running on another machine, a shared broker, a hosted space. `--root` is the folder\nwhose `.cotal/auth` holds that mesh's credentials and whose `.cotal/agents` holds its personas.\nThe default is the project you run it in. The registry stores that path, never a secret. `--mode`\ndefaults to `auth` when the root holds the space's account record and to `open` otherwise. The\nbroker is probed before anything is recorded, so a wrong address, or credentials that mesh will\nnot accept, fails here instead of at the first `spawn`; `--force` records without verifying (and\nreplaces an existing record).\n\nA hostname or public address is registrable only when the connection will **require TLS**. Pass\n`--tls`, or use a `tls://` URL. The scheme is recorded as enforced intent, so every later dial\nthrough the record demands the handshake (and `meshes add tls://\u2026` against a plaintext broker is\nrefused at registration). Without required TLS the fence admits loopback and private-overlay\nliterals only. RFC1918 addresses are refused in both modes because a cafe LAN is private but does not belong to you.\n\nA **user-auth** mesh registers from supplied pinned trust, never guessed: `--user-auth-file`\ntakes the bundle exported where the mesh runs; `--from` asks before it dials the address at all,\nthen fetches the `/.well-known/cotal-mesh` discovery document under that address (HTTPS only; a URL\nthat already ends in that path is fetched as given), displays the pins, and asks again before\nadopting them. Neither fetch follows redirects: a 302 can move a pinned fetch\nonto plaintext or onto another host, so it is refused rather than followed, and the pinned\nexchange must itself be an `https://` URL, except for an exchange on this machine, where plain\n`http://` is accepted for a loopback *literal* (`127.0.0.1`, `::1`, any spelling of them) but not\nfor `localhost`, which is a name rather than an address. Registration verifies that the exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also verifies that the broker refuses a bare\nconnect; that auth-required refusal is the pass. The sentinel credentials land in a 0600 file under\nthe entry's root; the registry records only the path.\n\n`meshes rm` drops records. It never stops a mesh. For a mesh running on this machine `cotal down`\nis the right verb, and `rm` says so unless you pass `--force`. A hand-added record is removed by\n`meshes rm`, by an `add --force` replacement, or by a `cotal up` that actually starts the broker for that same space, server and root, which becomes that\nmesh and so takes the record over (a `cotal up` for that space anywhere else refuses instead).\nNothing that merely *infers* a record is stale from a dead broker touches it: an\nunreachable broker is listed `offline` and stays, whether `cotal up` or `cotal meshes add`\nwrote the record; a foreground `up` whose broker exits unexpectedly keeps its record the same way.\nA bare command does not treat that offline record as a running mesh;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the project\nthey tear down; a hand-added one they leave alone even when it shares a root, because nothing\non this machine could write it back.\n\nA discovered entry belongs to the normalized IdP origin and proved subject that supplied it. Local\nteardown, cleanup, and liveness pruning do not remove it. A manual or locally started entry with the\nsame name wins and remains untouched; that discovered name is reported as a collision. Logging out\nremoves only the discovered entries owned by that account.\n\n`cotal meshes` and `cotal status` print `events: required` for a registration carrying\n`policy: { events: \"required\" }`. On that space, foreground spawn, detached spawn, manager starts,\nand interactive `join` cannot opt out or join without an event plane. `--no-events` is refused with\nthe space named. A connector without an event plane is refused with both the space and connector\nnamed. A session whose own grant omits `events.<owner>.<actor>` is refused before joining and the\nmessage names a full-row `actor grant` repair. A running seat whose event plane stops for good on\nthat space stops too.\n\n`use <space>` sets that default; the selection applies from every directory,\nincluding inside another mesh's project. `status` is a read-only report: machine prerequisites\n(starting with the installed `cotal-ai` version), the installed extensions and their versions, this\nfolder's `.cotal/`, the recorded meshes, and a live snapshot of the selected mesh (roster, channels,\nmembership feed). Stale Claude skills and out-of-date `.agents` skills recommend `cotal setup --skills`,\nnot unscoped `cotal setup`. `status` takes `--space` / `--server` to pick the mesh to inspect; it starts\nnothing. The manager row asks the service endpoint once: a live process that does not answer is\n`not serving`, and a probe that could not be made leaves the row `running \xB7 service unchecked`.\nA process row whose PID record exists but cannot be read reads `pidfile unreadable` with the error,\nand the other rows still print. A live manager whose delivery-aware marker cannot be read keeps its row\nand names the failure as `delivery-aware marker unreadable` with the error. The `Web process` row\nprints the address the selected mesh's dashboard recorded in `web.session` once it was listening,\nwhile the PID in its `web.pid` is alive. Otherwise it reads `down`, or `not installed` without the\nweb extension.\n\nIf a refresh fails, `status` may still show the kept catalog bytes for diagnosis. It labels them\nstale with the last successful snapshot timestamp and the refresh error. It never calls that state\nsynchronized or online. If a selected discovered space vanishes from a successful snapshot, the\nselection is cleared and the command reports that no default is selected.\n\nFor a user-auth mesh the selected-mesh section reports the login `status` works as: the signed-in\nsubject when this machine holds a cached session for the entry's pinned IdP, or the exact `cotal\nlogin --idp <url>` line when it does not, with no network round trip either way. A locally\nprovisioned space also shows the actor grant row; a discovered or registered remote entry reports\nthe grant as not checkable on this machine, because the ledger runs where the space was\nprovisioned. `--components` on a user-mode target probes as that same signed-in login (`ps`'s\ncredential), never a static mint; when the login cannot supply a credential, the row says why\ninstead of printing the broker's refusal of an unauthenticated probe.\n\nPersona rows name the catalog they describe. If this folder and the selected mesh use different\ncatalogs, status names both and marks which one spawn launches from. A green `default` means the file\npasses the same agent-file loader spawn uses; a present but invalid file is reported as invalid.\n\n`cotal status --components` adds a fail-loud per-component health pass. It reads **each\ncomponent's own control surface**, rather than treating a PID, a lease, or a successful probe of a\nsibling as proof that the component serves. It prints one of `serving`, `absent`, `not-serving`, or\n`refused` for each component and exits `0`, `1`, `2`, or `3` respectively (the highest observed\nstate wins):\n\n- **manager**: local PID record, its liveness-lease holder and PID, then the manager's own typed\n `status` service reachability from this host. Manager builds that do not report static\n reconciliation say `static reconciliation not reported by this manager build`; the line stays\n visible even when the manager is otherwise `serving`.\n- **delivery**: local PID record, its ready lease (`ready` is the daemon's own bound-control\n signal), and the latest `renewal.<spaceKey>.json` adoption verdict, the record of the space the\n command was asked about, keyed per space the way the pidfiles are. A re-signed credential and a\n broker-accepted adoption stay distinct facts. A root-only `renewal.json` left by an older build\n names no space and is never read as any space's verdict (`doctor auth` names it as a leftover).\n- **web**: local PID record, then the `/api/meta` response at the address the dashboard recorded in\n `web.session` once it was listening, which must name the same PID. The probe presents the\n readiness nonce recorded beside that address, the one credential the dashboard accepts on\n `/api/meta`. A live PID with no readable recorded address (the dashboard is still writing it, or\n an earlier build started it), or an unrecognizable process record, is `refused`, not a green\n default-port guess.\n- **broker**: the registered mesh URL dialed from this host with its recorded TLS requirement.\n\n`absent` means Cotal has no live local component record (or has a stale record); `not-serving`\nmeans the component record is live but its service/readiness surface did not answer or is not ready.\nThose are intentionally separate exit cases. A failed or unreadable probe is `refused`, never an\nabsent component or a clean zero. A PID record that exists but cannot be read refuses only its own\nrow. A record that its component removes while the pass runs reads as `absent`.\n\n## spawn\n\n```bash\ncotal spawn [<persona>] [--detach] [--name <n>] [--agent <a>] [--model <m>] [--variant <v>] [--prompt <text>] [--cwd <dir>]\ncotal spawn -f <cotal.yaml> [--dry-run]\n```\n\nFor a foreground spawn onto a remote user-auth mesh, a launcher may supply a one-time enrollment\ninstead of a cached human login. Prefer a private file:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./seat.md --space main\n```\n\nThe file contains only the enrollment URL, ending with at most one line terminator, and must be\nmode `0600` on POSIX. An orchestrator that cannot mount a file may set `COTAL_ENROLLMENT_URL`\ninstead; that value is redeemed byte for byte, so a trailing newline in it is refused. Setting both\nis refused. Enrollment input\nrequires `--space` and applies only to a foreground persona spawn. If the mesh is not registered yet,\nthe enrollment response must carry the stock user-bundle fields and the command needs\n`--config <persona-file>` because there is no local remote-mesh persona catalog to read. The client\nredeems the URL once, registers the returned mesh material, exchanges the returned actor token at the\npinned auth service, and removes both enrollment variables before starting any child process.\n\nA cached login for the same IdP and an enrollment are conflicting proofs, so the command refuses\nrather than choosing one. An invalid enrollment never falls back to login provisioning. Unknown,\nexpired, revoked, and already-used enrollments all produce one response: ask the owner for a fresh\none. See [Enrollment redeem](identity-and-auth.md#enrollment-redeem) for the HTTP contract.\n\nA runtime that starts a managed seat outside the manager's filesystem hands the child a managed\nhandoff instead: one `0600` file named by `COTAL_MANAGED_HANDOFF_FILE`, carrying the lifecycle the\nmanager already enrolled. The runtime builds the command with `delegatedSeatCommand`:\n\n```bash\nCOTAL_MANAGED_HANDOFF_FILE=/run/seat/handoff.json \\\n cotal spawn --config ./seat.md --space main --name <actor> --agent claude \\\n --expect-owner <owner> --expect-lifecycle-uid <uid>\n```\n\nThe `cotal` entry reads the file, deletes it and drops the variable before it parses flags, prints\nhelp or loads extensions, so every outcome leaves no file. The variable is read under any letter\ncase; spellings that name different files are refused after every one of them was deleted. The\nspawn then refuses a malformed handoff, or one whose space, owner, actor or lifecycle UID differs\nfrom `--space`, `--expect-owner`, `--name` and `--expect-lifecycle-uid`, before any broker\nconnection or exchange request. Every refusal on this path names the field and never a value from\nthe handoff. The registration's server, exchange and enforcement checks, the local state this\nmachine keeps for the space (its mesh record, user-auth state and agent secret files), target\nresolution, the policy refresh, the broker preflight and the agent auth preflight quote the space,\nthe server, the exchange URL, the actor or a path named for one of them in their own diagnostics and\nin the filesystem errors under them. For a handoff each prints one fixed sentence that names the\nfield and the phase instead, whether its check fails or an error is thrown. When the agent auth\npreflight's rollback then fails to remove a secret or file, that sentence is followed by the names of\nthe cleanup steps that failed, without their errors. The event-plane policy\nrefusals name the handoff's space field. An actor outside `[A-Za-z0-9_]` and a space that cannot\nname local state, such as `..`, are refused as malformed before any plane. A handoff conflicts with\nthe enrollment variables, `--detach`, `-f` and `--creds`, and needs `--config <persona-file>`. From\nthere it runs the enrollment consumer above without redeeming anything. See\n[Delegated seats](embedding.md#delegated-seats-outside-the-managers-filesystem).\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | resolved mesh | Target space |\n| `--server <url>` | registry entry | Broker URL override |\n| `--creds <path>` | none | Control-caller creds for an off-registry manager (`--detach` only) |\n| `--name <n>` | persona's `name:` | Presence-name override (does not choose the persona) |\n| `--config <persona-or-path>` | none | Persona catalog name or file path; wins over the positional |\n| `--agent <a>` | persona's `agent:`, else `COTAL_DEFAULT_AGENT`, else `claude` | Connector type (`claude`, `opencode`, `jcode`, `hermes`, and so on) |\n| `--role <r>` | persona's `role:` | Role override |\n| `--model <m>` | persona's `model:` | Model override |\n| `--variant <v>` | persona's `variant:` | Model variant override (connector-defined; e.g. OpenCode reasoning tiers) |\n| `--cwd <dir>` | this cwd | Working directory to root the agent at. Refused before launch when the directory does not exist on the serving manager's host. |\n| `--prompt <text>` | none | Initial prompt auto-submitted at start |\n| `--resume <id>` | none | Fork an existing session id into the mesh; only connectors that declare resume support accept it (see [the matrix](connectors.md)). The manager records the source session id, and `ps --wide` shows it. With `--detach --on <instance>`, a Claude session held on this host is carried to that instance first ([Resume a session](connect-claude.md#resume-a-session)); carrying one needs `--on` |\n| `--no-events` | event plane on where supported | Opt out of the session's structured event plane (`--events` only restates the default) |\n| `--share-tools <sel>` | none | Share named operator MCP servers with the agent |\n| `--subscribe <a,b>` | persona's | Channel read-set override |\n| `--allow-subscribe <a,b>` | = subscribe | Read-ACL override |\n| `--allow-publish <a,b>` | deny | Post-ACL override |\n| `--detach`, `-d` | off | Launch via the manager into a detached PTY (reattach with `cotal attach`) |\n| `--on <instance>` | class anycast | With `--detach` only: pin the launch to one manager instance id (the whole id, as `ps` prints it). Refused on a foreground spawn (no manager to pin), with `-f` (a manifest deploy launches through the manager class queue), and when empty |\n| `--file <cotal.yaml>`, `-f` | none | Deploy a manifest onto the running mesh |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--allow-stale <a,b>` | none | With `-f`: waive named stale agents (apply-only) |\n| `--runtime <name>` | manifest's | With `-f`: override the manifest's runtime |\n| `--expect-owner <u_\u2026>` | none | With `COTAL_MANAGED_HANDOFF_FILE` only, and required there: the owner the handoff must carry |\n| `--expect-lifecycle-uid <uid>` | none | With `COTAL_MANAGED_HANDOFF_FILE` only, and required there: the lifecycle UID the handoff must carry |\n\nEach session uses its connector's **event plane** by default: a stream of structured events\ndescribing what the agent did, rather than the prose it wrote, on a channel of its own. The channel is named after\nthe agent's principal, `events.<owner>.<actor>`, never after its display name, because two live\nagents are allowed to share a display name and would then share a stream. The launch grants publish\nrights on that channel alone, foreground and detached alike. On an open mesh, which issues no\ncredentials, the launch still allocates the agent an id, so the channel names a stable actor.\n`--no-events` is the explicit opt-out unless the selected registration says\n`policy: { events: \"required\" }`. Required policy makes the\nevent arm and grant mandatory, so `--no-events` and connectors without an event plane are refused.\n\nThe launch decision and the grant are separate on purpose. Holding publish rights on a channel is\nnot a request to publish to it, so writing an event channel into an agent file's `allowPublish`\ndoes not override `--no-events`.\n\nThe persona (`--config` > positional > `COTAL_DEFAULT_PERSONA` > `default`) is loaded from the\ntarget mesh's `.cotal/agents/` when it is a bare name. A reference that contains a path separator or\nends in `.md` is loaded from that file. A relative path resolves against the mesh root, except that\nan enrollment or a managed handoff resolves `--config` against the working directory. A missing\npersona is refused with the catalog directory or the file that was checked. The launch flags\noverride the file. On a user-auth mesh the\neffective name is also the agent's actor token, so it must match the token grammar (no `-`); the\nspawn is refused with that explanation before any request is sent. Foreground runs the agent\nattached to your terminal; `--detach` hands the launch to the running manager. Both modes get the\ndurable backstop on a mesh that runs the delivery daemon; `--live-only` skips it for a foreground\nspawn (messages posted while it is disconnected are then not replayed). A foreground exit retires\nthe agent's creds and broker footprint, like a manager despawn. On a user-auth mesh the two arms\ndiffer: a spawn against a mesh this machine provisioned revokes the actor row on exit, while a\nremote spawn (an enrollment or the advertised provisioning endpoint) removes only this machine's\ncredential files; its grant stays until the mesh operator revokes it, and the launch line says\nwhich arm you are on. A spawn through the advertised provisioning endpoint against a record that\npins no exchange URL is refused before the grant is requested, so no credential lands on this\nmachine. A `--detach` spawn is an\n**action**: the manager accepts it and returns the allocated identity at once, then the launch\nfollows to a terminal outcome rather than blocking (see [the control surface](control-surface.md)).\nSee [Connect Claude Code](connect-claude.md) and [Agent files](agent-files.md); `-f` is a\n[manifest deploy](#manifest-deploys). (`cotal start` was merged into `cotal spawn --detach`.)\nA `--detach` spawn onto a manager from another Cotal release is refused before any request is sent\nwhen the manager's contract does not declare a field this CLI sends. The refusal names the field,\ncalls it version skew, and gives this CLI's version. A field you leave unset is not sent, so it\nnever causes that refusal.\n\nA manager has 50 seat slots, and each seat counts once. A slot is held by a managed seat (a row in\nthat manager's `cotal ps`, including a seat still joining), by a reserved launch the manager accepted\nbut has not started a process for, or by a cooling hold. A seat that ends within 10 seconds of\nstarting leaves its slot cooling until those 10 seconds pass, unless an operator stopped it. Such a\nseat holds only that cooling slot, even while its launch is still reporting the failure. A spawn\nrefused at the limit states that split and whether waiting can free a slot:\n\n```text\nat capacity (50 of 50 slots: 49 managed, 0 reserved, 1 cooling); waiting frees a cooling slot in 7s, or despawn one\n```\n\nA cooling slot frees at the stated time. A launch that has not settled frees its slot only if it\nfails, and a managed seat frees its slot only when it stops. The refusal counts a launch as pending\nonly while it holds a slot, so a launch whose seat already ended is not counted. The roster counts\npresence, which also includes peers no manager owns, so its total is a different number.\n\nRun from a managed seat's own shell on a static or open mesh, `cotal spawn --detach` launches as\nthat seat when it targets the seat's own space. Without `--space` it picks that target the way the\noperator path does, so a recorded mesh that is not running is skipped. The CLI reads the seat's\nlaunch identity (`COTAL_NAME`, `COTAL_ID`, `COTAL_LIFECYCLE_UID`, `COTAL_SPACE`, and on a static\nmesh the seat's own credential), so the manager records the seat as the spawner, the same as for\nthe seat's `cotal_spawn` tool. On a static mesh that credential also proves the seat's space, so a\nlaunch without `COTAL_SPACE` still runs as the seat, and a target space holding no credential for\nthe seat is refused. An open mesh acts as the seat only when `COTAL_SPACE` names its space. The\nseat can then stop the child with `cotal_despawn`, and the manager stops the child when the seat\nexits. On a static mesh a seat whose agent file lacks `capabilities: [spawn]` is refused, because\nits credential holds no spawn subject.\n`--on <instance>` keeps its pin: the seat's own credential has no instance route, so on a static\nmesh the CLI mints a one-shot `manager-caller` view for the seat, pinned to that instance and\ncarrying the spawn subject only when the seat's credential holds it. On an open mesh the call keeps\nthe TLS requirement the mesh records. `--creds`, `--server` with an unregistered `--space`, and a\nuser-auth mesh keep the operator path.\n\n## models\n\n```bash\ncotal models [--agent <connector>] [--refresh]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--agent <connector>` | all registered connectors | Connector whose catalog to list |\n| `--refresh` | off | Ask the connector to refresh its provider cache |\n\nAsks the running manager for each connector's model catalog (model ids plus their variants)\nfor connectors that expose one. OpenCode and Codex query harness/provider surfaces; Jcode reads\nproviders that enable `model_catalog = true` in the operator Jcode `config.toml`. Jcode's listed\neffort tiers render as `variants (declared, not provider-verified)`, and launch can still refuse one.\nA connector without a catalog says so. A connector whose harness the manager did not find at boot\nreports the reason boot recorded, as a spawn does, so restart the manager after installing it. Pick\na result with `cotal spawn --model <id> --variant <v>`, where `<id>` is the model id as the catalog\nprinted it. OpenCode and Codex ids are the full\n`provider/model`; Jcode ids are bare (`opus-5`, not `cliproxy/opus-5`), because the provider is\nselected by the operator's Jcode config and a prefixed id is refused at launch with the bare form\nnamed.\n\n## endpoints\n\n```bash\ncotal endpoints [--space <s>] [--server <url>] [--creds <path>]\n```\n\nLists the mesh presence roster: agents, the manager, and any other protocol endpoint, with each\nendpoint's role, kind, status, and current activity. Unlike `ps`, this is a read-only presence view;\nit is not limited to child processes owned by the manager.\n\n## Endpoint control\n\n```bash\ncotal describe <endpoint> [--on <instance>] [--space <s>]\ncotal invoke <endpoint> <command> [--args '<json>'] [--space <s>]\ncotal invoke <endpoint> <command> --name <agent> [--admin] [--space <s>]\n```\n\nThe generic v0.4 service surface. `describe` resolves a registered endpoint's command set off the\nwire - the reserved `describe` command answers the registered contract digests, the schemas are\nfetched from the space's content-addressed contract store, recompiled, and verified against those\ndigests - and prints each command with its capability class and targeting shape. `--on <instance>`\npins `describe` to one manager instance's rail (the whole id, as `ps` prints it under its\n`manager <id>` headers), so an operator can read what that instance serves in a multi-manager space;\nunpinned, the class queue answers and the attribution line names whichever instance did. `invoke`\ncalls one command by name: `--args` is a JSON object validated against the fetched input schema\n*before*\npublish; a targeted command takes `--name <agent>` (resolved to the agent's current principal through\n`inspect`) or `--self`. `--admin` uses the admin instrument credential, whose cross-agent reach rides\nthe operator-only `any` authorization mode. Neither command has compile-time knowledge of any\nendpoint's schemas - this is the same trust chain every built-in control command now uses. Needs an\nauth mesh: the manager registers its service on both static and per-user meshes (a signed-in user\nrides their bearer; each visible or invoked command still requires its existing grant, and cross-agent\nreach needs the `admin` scope). An open mesh has no service registry.\n\n## Managed seats\n\n```bash\ncotal ps [--on <instance>] [--wide | --json] [--slots] [--space <s>]\ncotal stop --name <n> [--on <instance>] [--space <s>]\ncotal attach --name <n> [--on <instance>] [--no-reconnect] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | none | Managed agent to stop / attach (required) |\n| `--on <instance>` | class anycast (`ps`: class scatter) | Pin to one manager instance id (multi-manager space); takes the whole id as `ps` prints it, not a prefix. An empty value (`--on \"\"`, an unset shell variable) is refused, never treated as absent. A roster principal id (`local.\u2026`) is refused with a message naming the instance id `ps` prints |\n| `--wide` (`ps`) | off | After each seat's compact row, print extra operational facts the manager records: the provider the connector reported serving the model, `cwd`, `pid`, spawner, lifecycle uid, the owning manager's instance id and host, and for a `--resume` seat the session it forked (`forked from <id>`, with the source title and transcript SHA-256 once a Hermes or Jcode seat has recorded its fork; a carried Claude session prints `forked from <host>:<id>` with its title, SHA-256 and `carried <time>`, the time its bytes reached the manager). Model and requested variant stay in the identity row rather than printing twice. A fact the manager did not record (for example a runtime with no real process, or a connector that reported no provider) prints nothing, never a placeholder |\n| `--json` (`ps`) | off | Machine-readable: one JSON object per seat per line, copied unchanged from the manager row. Instance headers and errors go to stderr, so stdout contains only rows. Mutually exclusive with `--wide` |\n| `--slots` (`ps`) | off | List the durable static slot rows this manager owns instead of live seats, through the `slots` command. Mutually exclusive with `--wide`. A row that is not in the live roster still prints, with `live=false`; a retired row never prints |\n| `--no-reconnect` (`attach`) | off | End the attach when its session ends, instead of re-establishing it. For scripts that want one run and one exit code |\n\nA raw `--creds` file is refused by `ps`, `stop`, `attach` and the other control commands, because\nthat route mints no endpoint-caller triple; the project folder, or `--space` against the registry\nentry, is the route that does.\n\nThe human `ps` row is presentation text and is not a stable parsing target. Scripts use `--json`,\nwhich is the machine-readable row contract.\n\n`--slots --wide` is refused: `--slots` lists durable static slot rows, `--wide` prints live seat facts, and the two answer different questions. Across a multi-manager scatter, `--slots` prints each manager's rows under its own instance header, the same way the plain `ps` scatter does.\n\nThese are operator clients over the running manager's control plane. The default row includes the\nconnector, model pin, optional requested variant, and runtime as operational descriptors for the\nmanaged row. They do not make a shared display name a unique protocol identity; use `--json` when\nunambiguous owner+actor attribution is required. An omitted variant means no override was requested;\nCotal does not invent an effective provider default it cannot observe. `ps` also prints two state\nfacts per managed agent, because they answer different questions: the process fact from the manager's\nown runtime handle (`running` with its uptime, or `exited` with how long it ran), and the mesh fact\nfrom the roster (`idle` / `working` / `waiting` / `mesh offline`, or `not in roster` when the seat has\nno presence row at all: a seat that has not joined yet, or one that never did). When the seat's\nconnector relays a harness-reported condition, the mesh fact carries its code and how long it has\nheld, so a seat whose turn died on a provider rate limit reads `waiting (rate_limit for 40m)` rather\nthan a bare `waiting`, and `--json` carries the whole `condition` object. When the connector reports\nthe seat's last work event (presence `activeAt`), the mesh fact ends with its age, such as\n`\xB7 active 3s ago`, and `--json` carries `activeAt`. A seat whose turn stopped advancing keeps\nheartbeating, so its presence row stays fresh and this age is what shows the stall. A seat can be\n`running` and `mesh offline` at once: the process is alive and its presence has lapsed. That row says\nhow long, as in `mesh offline for <age>` with an age such as `3.5h`, counted from the seat's last\npresence heartbeat, which `--json` carries as `offlineSince` (epoch ms). The age is read only from\nthe seat's own presence record, matched on its principal and lifecycle uid, so a same-named peer or\nan older lifecycle never dates it. The manager log names each managed seat that is offline on the\nmesh while its slot is held\n(`seat offline on the mesh: <name> - last heartbeat <time>; process <state>`), including one its\nwatch first sees offline after a reconnect, and each one that comes back\n(`seat back on the mesh: <name>`), so a watchdog that only checks process liveness has a line to\nact on. The manager does not reap or re-key such a seat. The mesh fact is only a verdict while the\nmanager's own presence watch is fresh: when that watch has been silent past the liveness window, or\nhas not replayed the bucket yet, every row prints `mesh unknown` with the reason instead (`--json`\ncarries it as `meshView: stale | unpopulated`), because `offline` and `not in roster` would then\ndescribe the manager's watch rather than the seat. The manager rebinds a watch that goes\nsilent under a live connection on its own, so `mesh unknown` normally clears within a liveness window.\nOn a user-auth mesh `ps` also renders each managed agent's last credential-refresh outcome, fail-closed.\n\n**Mode split (chosen up front, never try-scatter-then-degrade):**\n\n- **Static / open mesh.** Bare `ps` is a **class scatter**: it freezes the live manager class from\n the records registry, merges every registered instance's agents grouped and attributed per\n instance, and a non-answering instance is shown as `registered, no answer within the deadline`\n (never silently omitted). A refused list or a missing answer makes the census incomplete: rows\n from other instances remain visible, but `ps` prints an incomplete-census warning on stderr and\n exits non-zero, including with `--json`. Those rows are not a complete seat count. A contract\n mismatch prints one plain comparison of the requested and served input/output digest pairs and\n advises aligning manager versions. The no-answer label means only that the instance is registered\n and did not answer. It does not say the host is down, because a dead host never deregisters itself and a\n live one can be slow; if it is gone, deregister it.\n `--on <instance>` pins the read to one exact instance id instead. A wrong pin fails loud\n rather than falling through: a well-formed id that no live manager carries is reported as\n `manager instance <id> did not answer` (nothing else is asked), and a credential without that\n instance's rail is reported as refused by the broker, not as an unresponsive manager. A manager\n that answers with a refusal is shown with its own cause; \"no manager reachable\" is said only when\n nothing answered at all. If the scatter's own registry read fails (the freeze or the reconcile),\n `ps` says the manager registry could not be read rather than pronouncing on the managers, which\n may all be up.\n\n**The verdict is scoped to the endpoint rail the request rode.** An issued caller rides the\nversioned `ep.v1` rail, a separate subject space from the legacy `ep` rail, and an endpoint serves\nboth (SPEC 13.15). A manager older than the versioned rail serves `ep` alone, so it can be running,\nregistered and answering while an issued caller's request reaches nobody. Silence on `ep.v1` is\nreported as `no manager answered on the <rail> rail` with `ep.v1` as the rail, and names both causes\nit is consistent with: no manager running, or one older than the rail. The CLI cannot tell them\napart, because the service registry records no package version, so check whether a manager is\nrunning and, if it is, its version. The same scoping applies to `cotal run`'s hosted verbs, which\ndrop the `--local` suggestion there, since `--local` drives the run from the calling process and\nnames the caller as its answerer.\n\n**`stop` and `attach` route by seat locality.** A seat can only be stopped or attached by the\nmanager actually running it, and the class queue does not know which one that is. So on a\nstatic/open mesh both verbs first ask every registered instance which one hosts the named seat, then\naddress that instance directly. This happens by default; you do not need `--on`.\n\n`--on <instance>` remains the override, for when you already know where the seat lives or the\nlookup itself is degraded. On a **user-auth mesh**, the exchange selects one authorized manager\nfor a short-lived `manager-caller` view. `--on` requests a specific instance; without it, selection\nmust be unique. Discovery and the command use that instance route. The caller gains no registry\nread or scatter permission. An absent, ambiguous or unauthorized selection refuses before sending\nthe command.\n\nA seat is reported as **not found** only when every reachable instance answered for itself. An\ninstance that stayed silent past the deadline, or that refused the read rather than answering, said\nnothing about which seats it hosts, so the seat may be running on it. That case reports that the\nlocation could not be established, names the instances that did not answer, and states outright\nthat it is not a report that the seat is gone. Read it as unknown and retry with\n`--on <instance>`; a retry loop that treats it as \"already gone\" stops looking for a seat that is\nstill running. A single manager cannot tell \"hosted elsewhere\" from \"does not exist\": it answers\n`not-found` for both, which is why the search asks all of them and why an incomplete search\nconcludes nothing.\n- **User-auth mesh.** `cotal ps` reports what **one** authorized manager knows about your agents\n (an instance-addressed read against its in-memory roster, owner-filtered). It does **not** report\n other manager instances or establish whether they are reachable. Completeness across a\n multi-manager user-auth space is not claimed.\n A manager that does not answer fails the command outright (exit non-zero), rather than printing\n an empty list that could be read as \"no agents\". Your ledger row needs the `admin` scope to\n reach `ps` at all; `spawn` alone is refused by the broker (the ep tier boundary).\n\n`attach` streams and drives an agent's terminal on the `pty` runtime; detach with the escape key\n(Ctrl-] by default; see [`COTAL_DETACH_KEY`](config.md)). The key is recognised as the legacy\ncontrol byte and as the kitty keyboard protocol and xterm modifyOtherKeys encodings of the same\npress, so a terminal with either protocol enabled detaches too. It does so over a one-use, holder-bound\nmesh session ([SPEC](../SPEC.md) \xA713.6): the manager replies with a signed session grant (never a\n`127.0.0.1` URL), the CLI redeems it once over the broker, and the browser console (`cotal console`)\ndrives the same session. `stop` and `attach` need a running manager to talk to. On a static mesh\nthey are cross-agent admin operations. On a user-auth mesh, your own agents (any agent under your\nowner) need only the `spawn` scope; another owner's agent needs `admin` on your ledger row\n([identity & auth](identity-and-auth.md)). Launch detached agents with [`spawn --detach`](#spawn).\n\n**`attach` reconnects when the link dies.** A session lives on a network link, and a laptop that\nsleeps, a VPN that drops or a wifi handover kills it. When that happens `attach` prints\n`[cotal: connection lost, reconnecting]` on stderr and starts asking the manager for a new session:\na fresh grant, a fresh per-session credential, a fresh connection, so every attempt re-runs the same\nauthorization the first attach did. On success it prints `[cotal: reconnected]`, the manager repaints\nthe seat's current screen the way it does for any attach, and you carry on in the same terminal.\nRetries wait 1s, 2s, 5s, 10s, then 30s, for as long as the seat exists. The detach key is read the\nwhole time the loop runs, the waits and the attempts alike, so a reconnect never traps you: press it\nwhile a session is being established and the attach ends there, and a session that lands behind the\npress is handed back to the manager rather than left holding a slot. Everything else you type while\nthere is no session is dropped rather than queued, so keystrokes aimed at a terminal that turned out\nto be frozen, Ctrl-C included, are not delivered to the agent by a reconnect you did not know had\nhappened. That starts before the first session, not at the first reconnect: at a terminal, `attach`\nreads and drops what you type while it is still resolving the mesh, so a key struck at a prompt that\nhas not come up yet does not reach the agent when it does.\nThe terminal is in raw mode for the whole reconnect, including when the link died before the first\nsession finished opening, so the detach key works there too instead of echoing as `^]`.\n\nA **pipe** carries script input. For example, `printf 'ls\\n' | cotal attach --name web` is\nbuffered until the session opens. Buffering continues across reconnects, so\n`tail -f log | cotal attach --name web` does not lose the part of its feed written while the link was\ndown. Only a terminal gets the reader; `--no-reconnect` keeps the old behaviour on both.\n\nIt stops on its own when reconnecting cannot help, and says why: a manager that refuses the attach\nexits non-zero with the manager's own message, and a reconnect that finds the seat no longer there\n(despawned, or its agent exited while the link was down) exits cleanly with `seat <name> is gone`.\nA local connect refusal that retrying cannot fix, such as a static-auth mesh whose seed is now\nmissing, also exits non-zero with the refusal's own sentence. A broker that is still unreachable\nkeeps the loop trying in silence.\nA refusal that could still pass, such as a manager at its session ceiling, is relayed in the\nmanager's own words while the loop keeps trying, once per refusal rather than once per attempt.\nPressing the detach key, or the agent's process exiting while you are attached, ends the attach as\nit always did. `--no-reconnect` turns all of this off and restores the single-session behaviour,\nwhich is what a script wants.\n\nEach reconnect also hands the abandoned session back to the manager, over the first link that can\ncarry the message, so an attach that flaps does not eat the manager's session slots one outage at a\ntime. If that message never gets a link, the attach says so when it ends. The live-session ceiling\ndefaults to 64 concurrent sessions (`--max-sessions`); the browser console opens one session per\npane, so a dashboard over a large mesh should size for agents \xD7 panes. Hitting the ceiling refuses\nbefore a credential is minted and names `--max-sessions`.\n\nWhich mesh `attach` resolves also decides **how it redeems the grant**. On a registered open mesh\nthere is no local seed. The CLI connects bare, the same way other control commands already do, and\nthe session rail is the caller rail that a real open-mode connection already reaches. Telling the\noperator to re-register the root is false: the registered root is already the contract. On a\nstatic-auth mesh the grant is still redeemed by minting a short-lived\nsession-scoped credential from the seed at the root the mesh resolved to, never from a `.cotal`\nfound by walking up from whichever directory you happen to be standing in. The difference is not\nhypothetical: `~/.cotal` exists on every install because the mesh registry lives there, so a command\nrun anywhere under your home directory but outside a project used to mint from your home\ndirectory's trust and present it to a broker that trusts a different chain, which surfaced as a\nbare authorization failure that named nothing. A directory that does hold another chain for the\nsame space is now reported on the way past, and not obeyed:\n\n```text\n! this directory resolves to /Users/you, whose .cotal/auth holds a DIFFERENT trust chain for space \"team\".\n attach used /Users/you/projects/app, the root this mesh resolved to. The other one is not being used, and is worth a look.\n```\n\nWhen a **static-auth** mesh holds no seed at the resolved root, `attach` refuses and names what it\nresolved, the broker and the root, instead of describing a directory it did not use and instead of\ntaking the open-mode path. An authenticated registry entry with a missing seed is still\nauthenticated. On a USER-AUTH mesh `attach` reads no seed. It sends your login and the session grant to\nthe auth service, which issues a `session-caller` bearer only if your owner and actor hold that\nsession. The connection it opens expires with the session grant.\n\nTerminal bytes stream over the mesh; the manager's own HTTP/WS face serves the console. That endpoint binds\n**loopback by default**, so nothing is exposed by accident; `cotal up --host <addr>` passes its bind\naddress down, which is what lets you reach the browser console (`cotal console`) for an agent whose manager runs on another machine.\n`attach` does not use that face: it redeems a signed mesh session grant over the broker instead (see above), so it reaches a\nremote manager regardless of the bind address. A\nbare `cotal supervise` and an embedded manager stay machine-local. Set it directly with\n`supervise --console-host <host>`.\n\nThat address is **recorded on the mesh** and carried forward, because it is a decision rather than\nsomething later commands can work out for themselves (a broker dial address is not a manager bind\naddress). Every later manager launch for the same mesh reuses it, including a same-root `cotal up` repair,\nan adopted preserved or restored listener, and a `spawn -f` manifest deploy. A manager replacement\ndoes not quietly move a reachable attach face back to loopback. Passing `--host` again overrides it,\nso you can widen or narrow exposure whenever you like; a mesh that never asked stays loopback-only\nand records nothing.\n\nBecause that face mints terminal read and write authority for every managed agent's browser session, it is credentialed in two\ntiers. A mesh caller receives a **ticket** bound to the single agent the manager just authorized,\nsingle-use and short-lived, so one authorized attach can never be re-pointed at someone else's\nagent. The **console token** is the operator's own, reaches every agent, and is printed only to the\nmanager's output. The roster, the live feed, and the PTY stream all answer `401` without one; the\nstatic console shell is served openly, since it describes no agent.\n\n## input\n\n```bash\ncotal input --name <n> --text <text> [--no-enter] [--on <instance>] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | | Managed agent to type into (required) |\n| `--text <text>` | | The text to type, taken verbatim (required) |\n| `--no-enter` | off | Type the text and stop there, without pressing Enter |\n| `--on <instance>` | class anycast | Pin to one manager instance id using the same rules as [`attach`](#managed-seats) |\n\nTypes one line into a running agent's terminal, as if you had typed it there, and returns. This is\nthe half of [`attach`](#managed-seats) that a program wants: `attach` is a live stream that holds a\nsession open and expects a terminal on your side, so a script, a cron job or a web UI cannot use it\nto send a single line. `input` is one authorized call.\n\nWhat it is for is **harness commands**. A line beginning with `/` is not chat and not a message: it\nis something the agent's own harness handles, and the only way in is the keyboard.\n\n```bash\ncotal input --name reviewer --text \"/compact\" # ask the harness to compact its context\ncotal input --name reviewer --text \"/model opus\" # switch its model\ncotal input --name reviewer --text \"hold on that PR\" # ordinary typing works too\n```\n\n**Quoting.** `--text` takes a value, so a payload starting with `/` survives as written. A payload\nstarting with a dash needs the `=` form, because the shell-style `--text --foo` is ambiguous and is\nrefused rather than guessed:\n\n```bash\ncotal input --name reviewer --text=--verbose # dash-leading text: use --text=<value>\n```\n\nEnter is pressed by default, since a command typed but never submitted has not been delivered.\n`--no-enter` types the text and leaves it sitting at the prompt, which is how you stage a line and\nsend it later.\n\nNothing comes back but a delivery receipt (`\u2713 sent 9 bytes to reviewer`, counting the trailing\ncarriage return). Whatever the agent does next shows up where its output already goes: the mesh, its\ntranscript, or an `attach`.\n\n**This one is operator-only, and more narrowly than `stop` or `attach`.** Those two are granted to\nanything holding `spawn`, so an agent can stop and attach to seats under its own owner. `input` is\nnot: it is granted only to operator credentials, which on a user-auth mesh means your ledger row\nneeds the `admin` scope, the same scope [`ps`](#managed-seats) already needs there. The reason is\nthat a write into a terminal is control of whatever is running in it, and on a user-auth mesh the\nown-owner rule covers every seat under you, not only the ones you launched: a `spawn`-scoped agent\ncould otherwise type into a sibling it never started. Seat locality is still resolved for you.\n\nOnly the `pty` runtime can be typed into. The external terminal runtimes (`tmux`, `cmux`, `orca`,\n`herdr`) attach to a process they do not own, so they have no input stream for it and the command\nrefuses by name rather than dropping the keystroke.\n\n## personas\n\n```bash\ncotal personas list [-v] [--running]\ncotal personas show <name>\ncotal personas edit <name>\ncotal personas new <name> (--prompt <t> | --from <f>) [--role <r>] [--model <m>]\ncotal personas rm <name> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh's persona catalog |\n| `--role <r>` | none | `new`: the persona's role |\n| `--model <m>` | none | `new`: the persona's model |\n| `--prompt <t>` | none | `new`: the persona's prompt text |\n| `--from <f>` | none | `new`: seed the prompt from a file |\n| `--verbose`, `-v` | off | `list`: include role / model / description |\n| `--running` | off | `list`: mark personas live on the mesh |\n| `--force` | none | `rm`: required, delete without prompting |\n\nPersonas are the local agent files under the resolved mesh root's `.cotal/agents/`, the same catalog\n`cotal spawn` launches from. `--space` and `--server` therefore move every list, read, write, delete\nand completion operation to the selected mesh. An unresolved target refuses rather than falling back\nto the current directory. See [Agent files](agent-files.md) for the file format.\n\n## supervise\n\n```bash\ncotal supervise [--runtime <name>] [--space <s>] [--server <url>] [--spawn <names>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space to supervise |\n| `--server <url>` | hosting mesh, or matching registered mesh | Broker URL. A registered mesh supplies it when omitted; a different explicit value is refused before anything is dialed. |\n| `--runtime <name>` | `pty` | Agent runtime (`pty` built in; extension runtimes are explicit-only) |\n| `--console-port <n>` | none | Protocol-console port |\n| `--console-host <host>` | loopback | Bind host for the console endpoint. Loopback keeps it machine-local; `cotal up` passes the address it bound the broker to, which is what lets the browser console reach this manager from another machine. `cotal attach` does not use this face: it redeems a mesh session grant over the broker |\n| `--max-sessions <n>` | 64 | Live-session ceiling. Each console pane and each `cotal attach` is one session, so size for agents \xD7 panes, not agent count. A capacity refusal names this flag. `cotal up --max-sessions` records the same number on the mesh so a later `supervise` started by repair or `spawn -f` keeps it |\n| `--roster <file>` | none | Declarative roster to boot at startup. See [Roster files](define-a-team.md#roster-files) |\n| `--launch <spec>` | none | Resolved manifest launch spec (from `up -f` / `spawn -f`) |\n| `--spawn <names>` | none | Comma-separated personas to pre-spawn at startup |\n\nThe manager is the agent supervisor and control plane: it answers `spawn --detach`, `stop`, `ps`,\n`attach`, and the `cotal_*` manager tools. `cotal up --detach` starts one for you; run `supervise`\ndirectly to recover a dead manager or drive a custom runtime. Default runtime is `pty`; install an\noptional provider first (`cotal ext add @cotal-ai/orca`, `@cotal-ai/tmux`, `@cotal-ai/cmux`, or `@cotal-ai/herdr`) and\nselect it explicitly. A missing provider or app fails loudly; there is no fallback. See [Deploy](deploy.md).\nBoot inventory decides whether this process takes unpinned `spawn`/`launch` on the class rail:\nif every declared connector is unavailable, those commands stay on this instance rail only\n(`status` reports `classSpawn: false`). `describe` still answers on the class rail, so an\nunpinned spawn can bind-fence against a skip member; re-issue, or pin `--on`. A partial\ninventory keeps the class rail and names `--on` on a harness refusal, because sibling\ninventories are not readable from the serve credential. See [control surface](control-surface.md#instance-routing).\n\nOn a normal `SIGINT`/`SIGTERM`, the manager stops every seat and requires the selected runtime to\nprove the seat is gone before it releases the manager lease or service registration. A stop that\ncannot prove exit fails loud and keeps manager authority instead of reporting a clean shutdown while\nan orphan still holds broker rails. After an abrupt manager death, the same logical successor\nterminalizes only its own durable static slots, verify-evicts the predecessor's broker principal,\nrecords that result in the lifecycle's caller-readable audit detail, reaps the predecessor's seat\nprocess through the runtime's custody reference recorded on the slot (the pty runtime verifies the\nprocess start identity in its seat record, so a reused pid is never signalled), and only then\nretires the lifecycle and frees the alias. A runtime that custodies its seats reserves that\nreference before it launches one, and the manager records it on the slot's first durable row, so a\nmanager that dies part-way through a spawn also leaves a seat its successor can address. A\nsame-lifecycle restart or a resume records the new seat's reference on the slot the same way, and\nwhen the slot does not take it the restart or resume fails and stops any seat it started, so the\nslot never names a seat that has already exited while its replacement runs. A resumed seat keeps\nits retained credentials, so the resume frees it only once its exit is proved; a seat whose stop\ncannot be proved stays managed, and the resume's error says so. A spawn\nthat launched its seat and then failed is rolled back by the manager that launched it, and that\nrollback reaps the seat through the same reserved reference before the lifecycle retires. Missing or unverified broker evidence keeps the slot\nterminalizing, and so does a runtime that cannot reap by reference.\n\nA `meshes add --mode user` entry is a **participant** registration, not hosting authority. A\nparticipant may run `supervise` only when the host advertises the remote manager authority service\nand the signed-in actor has the dedicated `supervise` ledger scope. The CLI obtains the closed,\nloopback-only `manager-service` view; `spawn` and `admin` do not substitute for that scope. The\nhost issues the manager's public-nkey JWT material through its lifecycle-bound prepare \u2192 activate\n\u2192 renew protocol, never by handing the participant a signer or static provisioner credential.\nThe host also performs instance-scoped eviction and guarded gate reconciliation. A remote manager\nrefreshes its short-lived registration executor before clean deregistration, so a long-running\nprocess removes its service row on `SIGINT` or `SIGTERM`. After an unclean stop, the same instance\nverify-evicts its superseded family and advances the process epoch. If an abandoned frozen gate\nholds the manager governance slot, a different supervise-scoped manager asks the host to reconcile\nthat holder after a complete gone verdict, then retries its registration once.\n\nA remote supervise never uses local signing trust: with host-issued authority in hand, the\nmanager mints from that authority alone and consults local records only to refuse a conflict,\nnamely the supervised space's own trust records under the cwd root. A root that hosts another\nstatic space beside the sign-in is a normal configuration and is never read as this space's\ntrust.\n\nThe broker URL in the registry entry decides the transport. A remote broker is often published\nover a `wss://` edge rather than a raw `nats://` port, and `supervise` dials whichever scheme the\nrecord holds, starting with the manager-authority registration it runs before the manager exists.\nThe record also decides whether that registration requires TLS, so a participant never downgrades\nthe credential exchange to a plaintext connection the registry did not describe.\n\nWithout that advertised host service or scope, `supervise` refuses before it starts a manager.\nRun `cotal spawn` without `--detach` to launch a foreground agent, or ask the space host to enable\nthe authority service and grant `supervise` for detached agents. If a running remote manager loses\nrenewal, it reports degraded state and refuses unsafe new starts and restarts; live agents are not\nsilently replaced. Do not run `cotal down` or `cotal up` on a participant machine to repair this\ncondition.\n\n## service\n\n```bash\ncotal service install [--mesh <name>] [--linger]\ncotal service status [--mesh <name>] [--json]\ncotal service uninstall [--mesh <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--mesh <name>` | this folder's mesh | The mesh whose manager the service runs; one unit per mesh |\n| `--linger` | off | install: when lingering is off, ask logind to enable it so the user manager starts at boot and the service survives logout. Never enabled silently |\n| `--json` | off | status: machine-readable output |\n\nRuns the manager as a user service so it survives logout and reboot. On Linux this installs a\nsystemd user unit (`~/.config/systemd/user/cotal-manager@<key>.service`, where `<key>` is the\ncase-safe mesh key); on macOS a launchd agent plist under `~/Library/LaunchAgents/`. Any other\nplatform, or an absent systemd/launchd user session, fails with a message naming what is missing.\n\n`install` resolves the mesh from the registry and binds the unit to that entry's root and broker\naddress, so it can be run from any directory. The mesh must be registered (`cotal up` or\n`cotal meshes add`) before installing; an unregistered name refuses before anything is written.\n\nThe unit's `ExecStart` is the bare `supervise` command. The mesh facts travel in the unit's\nenvironment (`COTAL_SPACE`, `COTAL_SERVER` pinned to the registered broker URL, whatever port it\nlistens on) rather than the command line, because command lines are readable by every user on a\nmulti-user host. On Linux that environment is a `0600` `EnvironmentFile`; on macOS it is the\nplist's `EnvironmentVariables`. The same environment gives the service a private `COTAL_HOME` and\n`XDG_CONFIG_HOME` under the unit directory, so the service manager never touches the login\nuser's `~/.cotal`. First-run connector seeding runs synchronously inside `service install`,\nagainst that private config root; the unit itself starts with `COTAL_SKIP_CONNECTOR_SEED=1`\nso a manager is never interrupted mid-seed by a restart. An install whose pre-seed cannot\ncomplete (network unreachable, registry error) refuses instead of deferring.\n\nThe same environment pins `PATH` to the `PATH` of the shell that ran `install`.\nWithout it the unit inherits the service manager's own short\n`PATH`, which usually lacks `~/.local/bin` and Homebrew, so the manager's boot inventory would\nreport a harness unavailable that your shell resolves. Install from a shell that resolves every\nharness the service should launch, and reinstall after moving one. A relative entry, including\nan empty one, is resolved against the directory you ran `install` from, because the unit starts in\nthe mesh root where the same spelling names another directory. An entry with a `..` segment is\npinned as the directory your shell reaches through it, with symlinks followed, and refuses when it\nreaches none. A `PATH` set to the empty string is one empty entry, so it pins that directory. An\nunset `PATH` refuses.\n\nEvery value the unit derives from a path (`WorkingDirectory`, the `EnvironmentFile` path, the\n`ExecStart` tokens) is escaped for systemd specifiers (`%` becomes `%%`), so a mesh root that\ncontains `%` starts over its real path instead of a path systemd rewrote by expanding it. The\nprovenance comment records the root unescaped.\n\nOn Linux a user unit starts at boot and survives logout only while the user lingers. Without\nlingering, systemd starts no user manager at boot, so an enabled unit stays inert until the next\nlogin and stops at the last logout. `install` checks lingering before it writes anything, and when\nlingering is off it fails with the root command that turns it on (`sudo loginctl enable-linger\n<user>`). With `--linger` it first asks logind to enable lingering for the current user, and fails\nwith the same command when logind refuses (unprivileged users over SSH get `Access denied`).\n`service status` prints that command while lingering is off. A Linger query that does not answer\n`yes` or `no` (logind unreachable, no `loginctl`) is never read as off: `install` refuses with\nthe query's own error and enables nothing, and `service status` shows lingering as unknown with\nthat error (`--json` gives `\"linger\": { \"error\": ... }`).\n\n`service install` also refuses while a manager is already running for the mesh (`cotal down\nmanager` first). The restart policy is `Restart=always` with `RestartSec=20s`, chosen for\nmanager units in production: a manager exits for reasons that are not failures (broker\nrestarts, host suspend), where `on-failure` with a short interval thrashes.\n\nThe unit also sets a start limit (`StartLimitIntervalSec=30min`, `StartLimitBurst=20`). A manager\nthat keeps failing to start stops after 20 attempts, about seven minutes at 20 seconds apart, and\nthe unit is left `failed` instead of restarting forever. One such failure is deliberate. After an\nunclean stop, a manager that cannot verify eviction of its predecessor's credentials exits 1 and\nleaves the issuance gate frozen, because starting without that proof could let two incarnations\nserve at once (SPEC 13.1). It first waits up to 60 seconds for the delivery daemon to answer, so a\ndaemon that is still starting does not fail the start. The log names the cause. When the delivery\ndaemon is down, it says the daemon is not reachable on the `ctl.delivery-admin` rail. When the\ndaemon answers and refuses, for example because the space is missing a `$SYS` cred, it prints the\ndaemon's own reason and repair step. Fix that cause, then run `systemctl --user reset-failed\n<unit>` and `systemctl --user start <unit>`. The macOS agent has no start limit: launchd's\n`ThrottleInterval` only spaces restarts.\n\n`service status` reports the unit state from systemd/launchd, the manager's own health read from\nits pidfile at the unit's recorded root, and the machine facts a hosting side asks for:\narchitecture, OS (the platform, never the hostname), whether `/dev/kvm` is present and\naccessible, CPU count, and total memory. `--json` returns the same fields as one object. The\nmanager row names the recorded pid, and the command it runs when another program has reused that\npid. `--json` also gives the command of a live recorded pid whenever it can be read.\n\n`service uninstall` stops and disables the unit and removes it plus the private state directory.\nIt works from any directory: the unit's own records name the mesh and root it serves, and an\nexplicit `--mesh <name>` selects it. It refuses any unit that was not written by `service\ninstall` (the files carry a provenance comment), whose recorded mesh is missing, or that was\ninstalled for a different mesh, so operator-written units are never destroyed; `service status`\napplies the same rule and never reports a mesh a unit does not record.\n\nThis command installs only the manager. The per-space auth service and the delivery daemon are\nnot installed by it: on a shared broker an operator runs three units per space with `After=`\nedges (auth service, then manager, then delivery) and stops them in reverse. A broker-side `cotal\nup` unit is a separate unit documented in [Run a mesh](run-a-mesh.md).\n\n## reconcile-gate\n\n```bash\ncotal reconcile-gate [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the frozen gate lives in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint whose gate is frozen |\n| `--instance <id>` | this folder's persisted manager instance | Instance id |\n\n**When you need this.** A manager restart killed after deregistration begins but before the new\nincarnation finishes leaves the endpoint's issuance gate *frozen*, held by a\nprocess that no longer exists. The freeze is what stops two incarnations serving at once, which is\ncorrect. The successor manager now completes that dead registration itself on boot, including on\nthe remote user-auth path. A foreign remote manager blocked by this gate also asks the host to repair\nit before one registration retry. Both use the same guard this command uses: they act only when the freeze-holder is affirmatively gone under a complete\nCONNZ sweep (`gone` and `sweepComplete=true`). If that registration's spec write already committed,\nit finishes the same freeze at the committed registration revision. If the spec did not advance, it\nabort-reopens the gate at generation+1 with processEpoch unchanged and continues the normal takeover.\nLive, unknown, unestablishable, and\nwrong-op-kind still refuse; there is no TTL.\n\nUse this command when the automatic path cannot run: the delivery daemon is down, the repair targets a\nnon-manager endpoint, or you want to lift the freeze without starting a manager. It checks that the\nholder really is gone, prints what it found, and then finishes the dead operation the same way as the\ninterrupted restart would have: revoke the old credentials, evict their holders with verification,\nand reopen the gate.\n\nThe command revokes the old credentials 16 at a time. It then verifies the holders' eviction in\nshared sweeps of up to 256 holders on the delivery daemon. Each sweep scans the broker a fixed number\nof times and kicks live connections 16 at a time, so holders that are already gone add almost\nnothing and live ones add one broker round trip per 16 connections. The daemon must serve the\n`evictPrincipals` verb; an older daemon refuses it and the gate stays frozen.\n\nEach sweep durably records the holders it verified before the next sweep starts. If a holder is not\nverified gone, the command leaves the gate frozen with those records kept. An interrupted sweep\nrecords nothing, and the sweeps before it stay recorded. A retry still repeats the freeze-holder\nliveness check, then skips only progress bound to the same registration operation, frozen-gate\nrevision, and holder set.\nThe output reports holders completed before this attempt, completed now, and still remaining. A new\nfreeze or changed holder set starts from zero. Cursor cleanup happens only after reopen; a retained\ncursor is harmless because its old gate revision cannot authorize a later freeze.\n\n**It refuses far more often than it acts, on purpose**, and always says which check stopped it:\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `holder-alive` | The freeze-holder still has a live connection: a manager *is* running | Stop that process first. Reconciling would evict a live manager's credentials |\n| `holder-unknown` | The connection sweep could not prove the holder absent | Not safe to proceed: an unprovable holder is treated as a live one. Re-run once the broker answers completely |\n| `liveness-unestablishable` | The delivery daemon gave no verdict: it was unreachable, timed out, or refused | Act on the delivery lease line in the refusal (below). Silence is never read as death |\n| `not-frozen` / `no-gate` | The gate is open, or there is no gate at that coordinate | Nothing to repair: check `--endpoint` / `--instance` |\n| `wrong-op-kind` | Frozen under a takeover or retirement, not a registration | Out of scope for this command; it will not reinterpret another operation's intent |\n| `eviction-unverified` | The holder looked gone but eviction could not be verified | The gate is left frozen, unchanged. Investigate the broker before retrying |\n| `raced` | A newer manager moved the gate mid-repair | Re-run `cotal doctor` and look again |\n\nWhen the daemon gives no verdict, the refusal also reads the delivery lease (`lease.0`) and names\nwhat is blocking the rail:\n\n| Lease reading | What to do |\n|---|---|\n| absent | No daemon is running. Start it (`cotal up` runs it) and re-run |\n| unreadable | The daemon cannot be named, so do not assume none is running. Fix the lease read, then re-run |\n| held, not ready | That holder claimed the shard and has not bound its rails. Wait for it, or stop it so its lease lapses |\n| held, ready, no answer | The query may have gone to another daemon still subscribed to the rail, such as a stopped one whose lease lapsed. Re-run before stopping anything. If no run gets an answer, stop any other delivery daemon for the space, then stop or restart the holder |\n| changed hands | The holder took the shard after the query was sent, so it was never asked. Re-run before stopping anything |\n\nThe command reads the lease before it sends the query and again after the query fails. It names a\nholder as the blocker only when the same run of the same daemon held the lease both times, and two\nrows from a daemon too old to record its run never count as the same run. Even then a ready holder\nmay not have been asked: the rail is queue-grouped, so any daemon still subscribed to it can take\nthe query. A row whose times are not valid dates reads as unreadable.\n\nA daemon that answered and refused keeps its own reason, followed by the same lease line. The lease\nline names the holder, whether it is ready, the space account that holds the lease bucket, when that\nholder acquired the shard, and when the row was last written. A ready holder rewrites the row on\nevery renewal and keeps its acquisition time, which only a successful acquisition sets. A row\nwritten by a daemon that predates the acquisition time reports it as unknown. The lease reads never\nchange the outcome: the gate stays frozen and the command exits 2. A manager's boot self-heal uses\nthe same check and reports the same line.\n\nThere is no `--force`, and no path that discards gate state: the only way this reopens a gate is by\nproving the holder is gone and then completing the operation properly.\n\n**What reopening the gate does for the endpoint's governance slot.** A registration takes the\nendpoint-wide governance slot before it publishes its spec, and holds it until its gate reopens. An\ninstance that died between those two points leaves the slot held with no registration behind it.\nThis command does not write that slot and never has; the registration path is its only writer. What\nthe reopen does is advance the holder's gate past the generation the slot is stamped with, which is\nwhat marks the slot abandoned. The next registration for that endpoint then reclaims it as part of\nits ordinary start. So the repair here is still one command followed by starting the manager, and\nthe slot needs no separate step.\n\n## deregister-instance\n\n```bash\ncotal deregister-instance [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the instance is registered in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint the instance serves |\n| `--instance <id>` | this folder's persisted manager instance | Instance id, the whole id as `cotal ps` prints it |\n\n**When you need this.** The service registry records *registration*, not liveness, and nothing in\nthe model expires a row. A manager that stops cleanly removes its own registration. One whose host\ndied without writing anything cannot, so its record goes on claiming a live instance forever: every\nclass scatter in that space freezes the dead slot in, and `cotal ps`, `stop` and `attach` each pay\ntheir whole deadline waiting for a machine that is never coming back. A laptop that was reimaged, a\ncontainer that was deleted, a box that will not be back on the network: those registrations have no\nother exit.\n\nThis command is that exit. It asks the instance first, and it removes a record only when the broker\naffirms the instance's own rail is empty: nothing subscribed there. Then it deletes the\nregistration's two records keys, each pinned to the revision it read, and prints what it removed.\n\n**Silence alone never passes.** An unanswered describe is what a dead host, a wedged process and a\nslow one all look like, and a hung process still holds its subscriptions, so the broker sees\ninterest on its rail. That instance is refused and the observation is printed. A dead process holds\nno connection and therefore no subscription, so a real corpse is still removed.\n\n**Every refusal names the failed check:**\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `instance-answered` | The instance answered a pinned describe. It is alive | Nothing to repair. If it is wedged rather than gone, stop the process first; its own clean stop removes the record |\n| `instance-not-affirmed-gone` | It did not answer, and the broker did not report its rail empty, which is what a held subscription looks like: slow or hung, not affirmed gone | Nothing was removed. Stop the process; its record goes on its own clean stop, or re-run this once it is down |\n| `liveness-unestablishable` | The probe itself failed, so nothing was learned | Fix the probe's path (credential, broker) and re-run. A probe that could not run is never read as death |\n| `not-registered` | No registration at that coordinate | Check `--instance` and `--endpoint`. This takes the whole id, never a prefix |\n| `registration-in-flight` | The instance holds the endpoint governance slot at the live issuance-gate generation, so a registration is still completing | Nothing was removed. Wait for that registration to finish, then re-run |\n| `superseded` | The record moved between the read and the delete | Something is writing to it. Nothing was removed; re-observe before retrying |\n\nThere is no `--force` and no sweep: silence is not death, and a rule that removed rows on silence\nwould eventually remove a live instance that was merely slow. An operator names one instance, the\nbroker's verdict on its rail is what authorizes the removal, and the guard's job is to show them\nthey named a dead one. Removal is not a one way door either. The same instance re-registers over\nthe tombstone on its next start, under the same identity.\n\n## runtimes\n\n```bash\ncotal runtimes\n```\n\nLists every agent runtime the manager can spawn through: the built-in `pty`, the official providers\n(`orca`, `tmux`, `cmux`, `herdr`), and any custom provider installed via `cotal ext add`. Each installed\nprovider is probed so you can see what is actually reachable on this machine before selecting it:\n\n```\npty built in\norca installed \xB7 reachable @cotal-ai/orca\ntmux available \xB7 cotal ext add @cotal-ai/tmux\ncmux available \xB7 cotal ext add @cotal-ai/cmux\nherdr available \xB7 cotal ext add @cotal-ai/herdr\n```\n\n`installed \xB7 reachable` / `unreachable` is the provider's own `available()` probe; `available` means\nit is a known runtime you can add with the shown command. Selecting an unknown or uninstalled runtime\nvia `up`/`spawn --runtime <name>` fails loud and, for a known one, points at the exact `cotal ext add`\npackage. There is no silent fallback to `pty`.\n\n## seats\n\n```bash\ncotal seats [--drain]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--drain` | off | Retire every seat whose agent has exited. A seat whose agent still runs is kept |\n\nThe pty runtime used to start a detached custodian process for every Linux seat. It now spawns\nin-process, but custodians that an earlier manager started keep running, and one whose agent has\nexited stays resident while a manager still holds its connection. This command lists the custody\nrecords under `COTAL_SEAT_ROOT` (default `~/.cotal/seats`), one line per seat:\n\n| State | Meaning |\n|---|---|\n| `live-child` | The agent process still runs. The seat is never signalled, and a manager can still adopt it |\n| `childless` | The agent has exited, or the record comes from an earlier boot. `--drain` retires the seat |\n| `drained` | `--drain` proved the custodian and the agent gone and removed the record |\n| `refused` | The record cannot be read, carries no start or boot identity, this host publishes no boot identity, or the reap could not prove the processes gone. The record stays on disk |\n\nA drain signals only a custodian whose recorded start identity still matches the live process, so\na reused pid is never touched. No process outlives a reboot, so a record from an earlier boot is\nreported childless and `--drain` removes it without signalling anything. A record with no start or\nboot identity is refused with or without `--drain`, and is never reported as running or exited.\nOn a host that publishes no boot identity (`/proc/sys/kernel/random/boot_id`) every record is\nrefused the same way, because no record can be tied to this boot.\nThat refusal and an unreadable record signal nothing. A refusal from the reap itself can come after the drain already\nsent `SIGKILL` to the custodian. Its detail names the pid or process group the reap could not prove\ngone, so check those processes before you retry. The command exits non-zero when any record is\nrefused. It is Linux-only and throws on other platforms.\n\n## send\n\n```bash\ncotal send dm <agent> \"<text>\" [--space <s>] [--server <url>] [--creds <path>]\ncotal send msg <channel> \"<text>\"\ncotal send ask <role> \"<text>\"\n```\n\nA `send dm` prints one line naming three facts: `\u2192 <name> stored seq <N>; recipient <status>\nat send; delivery not confirmed <text>`. `stored seq N` is the JetStream sequence the broker\nassigned to the publish; `recipient <status> at send` is the roster status (`idle`, `working`,\nor `offline`) resolved right before the publish, which can change the instant after; the send\nnever prints `delivered`, because the sender's credential cannot read the recipient's durable\nto confirm it. Inspect what the broker actually holds for a recipient with\n[`cotal deliver pending`](#deliver).\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and (off-registry) which credential |\n\nOne-shot messaging: connect, send a single direct message (`dm`), channel post (`msg`), or role\nask/anycast (`ask`), then exit. For a running conversation, agents use the mesh tools instead\n([MCP tools](mcp-tools.md)).\n\n`cotal send` works from an operator shell or from a seat. Its display name is `<login>@<host>` of\nthe shell that ran it, so the recipient can tell one operator's send from another's; it is taken\nfrom the operating system, never from `COTAL_NAME`. The wire principal comes from the resolved\noperator credential or user bearer, not from `COTAL_NAME`, `COTAL_ID`, `COTAL_OWNER`, or\n`COTAL_ACTOR`. On an open mesh the transient endpoint self-mints its principal.\n\nThe transient endpoint never joins the roster and binds no inbox. A recipient can still answer a\n`send dm` or `send ask` with `cotal_dm`, by the sender's name or by the id on the message it\nholds: the reply is stored under the sender's id in the space's DM history, which an operator's DM\nview such as the dashboard's Direct messages lens shows. The `cotal send` that asked has already\nexited, so the reply never reaches that shell.\n\n## channels\n\n```bash\ncotal channels list\ncotal channels set <name> [--replay | --no-replay] [--window <n>] [--desc <s>] [--instructions <s>]\ncotal channels default --replay | --no-replay\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--replay` / `--no-replay` | none | `set`/`default`: replay history to new joiners, or not |\n| `--window <n>` | none | `set`: replay window size |\n| `--desc <s>` | none | `set`: one-line channel description |\n| `--instructions <s>` | none | `set`: instructions shown to joiners |\n\nInspects and edits the channel registry: replay policy, description, and joiner instructions. ACL\nsemantics (who may read or post) are set at mint / provision time, not here; see\n[Channels and permissions](channels-and-permissions.md). On a user-auth mesh, `list` rides your\nown login as is; `set` and `default` edit the registry over a short-lived\nchannel-writer view, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\nOn a remote user-auth mesh that view is served by the public exchange; space-history `purger`\nand the read-only admin view are not.\n\n\n## history\n\n```bash\ncotal history clear --force [--dms] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--dms` | off | Also clear DM history |\n| `--force` | none | Required: clear without prompting |\n\nPurges retained channel history; `--dms` extends it to direct-message history. An alias of\n[`clean history`](#clean). On a user-auth mesh the purge rides a short-lived purger view over\nyour login, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\n\n## console\n\n```bash\ncotal console [--plain] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to watch |\n| `--plain` | off | Line stream instead of the TUI |\n\nA live protocol view for a space: a lazygit-style TUI, or a plain line stream on `--plain`. On a\nuser-auth mesh it rides the read-only admin view over your login, which needs ledger scope\n`admin`. Inside the TUI, operator control (`D` kill, `:spawn`, `:status`, `:purge`) rides the\nsame per-action instrument path as `cotal stop` and `cotal ps`, never the observer; a raw\n`--creds` file cannot drive it. `a` (or `:attach <agent>`) runs\n[`cotal attach`](#managed-seats) in place and returns to the console on detach. See\n[Watch a mesh](watch-a-mesh.md).\n\n## web\n\n```bash\ncotal web [--detach] [--host <host>] [--port <n>] [--no-open] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to serve |\n| `--host <host>` | `127.0.0.1` | Concrete HTTP bind and browser host; wildcard addresses are refused |\n| `--port <n>` | `7799` | HTTP port, a decimal number from 1 to 65535 |\n| `--detach` | off | Run in the background; stop with `cotal down web` or bare `cotal down` |\n| `--no-open` | off | Don't open the browser |\n\nThe browser observability dashboard: presence, channels, and a live feed. It is **not** part of\n`cotal up`: it ships inside `cotal-ai` as the `@cotal-ai/web` extension, seeded automatically on first\nrun (like the built-in connectors) so it always matches your CLI version. It self-registers `cotal web`\ninto this surface and serves\n`http://cotal.localhost:7799` by default (loopback; `*.localhost` resolves in Chrome/Firefox/Edge; for Safari\nor a system resolver such as WSL2's, the launch link is also printed at `http://127.0.0.1:7799`).\nOn a user-auth mesh the dashboard rides the read-only admin view\nover your login, and a channel purge asks for its own channel-purger view per click; both need\nledger scope `admin`. The public exchange serves `channel-purger` for a remote owner; it still\nrefuses the startup admin view, so a remote `cotal web` is not a complete channel-management\nsurface. Detached mode re-execs the current Cotal installation, writes diagnostics to\nthe mesh root's `.cotal/web.log`, and reports success only after the HTTP server answers. It requires\na recorded mesh root, but can be launched from any directory once `cotal up` has recorded the mesh.\nSee [Watch a mesh](watch-a-mesh.md).\n\n## deliver\n\n```bash\ncotal deliver [--space <s>] [--server <url>] [--tls] [--creds <file>] [--root <dir>] [--shard <n>] [--shards <n>] [--dev-mint]\ncotal deliver pending <name> [--limit <n>] [--durable <name>] [--json]\n```\n\nWith no positional, `cotal deliver` runs the delivery daemon (see\n[the delivery daemon](delivery-daemon.md)). `deliver pending <name>` never starts the daemon: it\nis an operator-only read over one recipient's DM durable, for the moment after a send when the\nquestion is \"what does the broker actually hold for them.\" It resolves `<name>` against a short\npresence watch (an `offline` card still counts, since the recipient may be dead, that is what\nthe verb exists to inspect); when neither a card nor the durable can be found, it prints\n`\u2717 not-found: no agent \"<name>\" and no DM durable for it in space <s>` and exits non-zero, never\n`pending 0`. On a match it prints the durable name and one fact per line: `pending`,\n`ack-pending`, `delivered`, `ack-floor`, `created`, `frontier`, and the stream's `max_age` /\n`max_msgs_per_subject` / `discard` limits (`--json` prints the same facts as one object), followed\nby a bounded, unacked read of up to `--limit` (default 20) recent candidate message ids under the\nheading `recent candidate ids (from the ack floor; not proof of a hole)`, a list of what is\nthere, not proof that nothing was lost.\n\nThe verb needs the `admin` credential profile: it runs through the same static-mesh route as\n`cotal mint --profile admin`, and refuses a user-mode mesh, naming the retired static credential,\nbecause there is no user-mode inspection authority yet. Pass `--creds <file>` for an off-registry\nadmin credential. A same-name respawn never inherits a predecessor's held DMs (the durable is\nlifecycle-keyed); an old lifecycle's durable is reachable only by the name a live read printed\n(the `<durable>` line on the first line of this verb's output). Pass that name with `--durable\n<name>` to read it directly once the lifecycle's card is gone from the roster. This skips the\npresence watch on `<name>` entirely, so `<name>` is required but only echoed in error text.\n\n## mint\n\n```bash\ncotal mint <name> [--profile <agent|observer|admin>] [--out <path>] [--signer]\ncotal mint <name> --provision [--role <role>] [--space <s>] [--server <url>]\ncotal mint <name> --expires-in <seconds> | --expires-at <unix-seconds>\ncotal mint <name> --identity <creds> [--expires-in <seconds>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--profile <agent\\|observer\\|admin>` | `agent` | Credential profile |\n| `--out <path>` | `.cotal/auth/creds/space.<key>/<name>.creds` | Output path - the default sits under the resolved space's segment (`<key>` is that space's hex encoding, as in [Project files](config.md#project-files)) |\n| `--signer` | off | Emit a stripped account-signing file instead |\n| `--force` | off | With `--signer`: overwrite an existing file |\n| `--allow-subscribe <a,b>` | the agent file's, else subscribe | Read-ACL override, **agent profile only**: `observer` and `admin` carry a fixed read set, and `mint` refuses this flag there rather than narrowing nothing |\n| `--allow-publish <a,b>` | the agent file's, else deny | Post-ACL override, **agent profile only** |\n| `--role <role>` | the agent file's | Agent profile: the anycast task queue the identity pulls (`svc_<role>`) |\n| `--provision` | off | Agent profile: also pre-create the identity's bind-only DM/deliver durables (and its role's task queue) on the live mesh, so the credential can consume |\n| `--expires-in <seconds>` | unbounded | Bound the credential's lifetime: the JWT `exp` is `iat + <seconds>`. A positive integer; refused together with `--expires-at` |\n| `--expires-at <unix-seconds>` | unbounded | Bound the credential to an absolute `exp` (unix seconds). Refused together with `--expires-in` |\n| `--identity <creds>` | a fresh identity | Re-mint for the nkey carried by this creds file, keeping the principal and every durable keyed to it. The file is read by the same loader the endpoint uses; a file with no seed is refused by name |\n| `--space <s>`, `--server <url>` | the resolved mesh | Which root supplies the agent file, static trust and default credential storage; with `--provision`, also which live mesh receives the durables |\n\nMints a NATS creds file for a space in **static** auth mode, scoped to a profile and (optionally)\nexplicit read/post ACLs. `--signer` emits an account-signing file for delegating minting to another\nhost. A per-user-auth space refuses `mint`: agents there join under a logged-in user\n([`login`](#login) + [`actor grant`](#actor)), never via a handed-out creds file. See\n[Identity and auth](identity-and-auth.md).\n\nFor an agent profile, the resolved mesh root supplies the persona ACL, the signing material and the\ndefault credential destination as one authority. If the current folder also holds trust for a\ndifferent space or account, mint refuses before writing and names both roots. It never combines a\npersona from one root with credentials signed or stored under another.\n\nA plain mint is creds only: the identity can publish within its post ACL at once, but on an authed\nmesh its DM inbox and task queue are provisioner-pre-created and bind-only, so a **consuming**\nconnect fails until they exist. `--provision` performs that pre-create in the same command (a\nprovisioner cred is minted from the space's trust material, used, and dropped), so a long-running\nclient you start yourself can receive DMs and role anycasts like a spawned seat. The command prints\nthe identity's principal (its wire id) and lifecycle uid; a consuming client passes that uid as its\n`lifecycleUid`. Agent profile only; an open mesh needs none of this (peers self-create there). The\nsame resolved authority is used for both the credential and `--provision`, so the broker\nfootprint cannot be created under a different root's trust material.\n\nThe CLI-mintable profiles carry no default TTL: without a lifetime flag the credential is\nunbounded, and a standing-renewal consumer refuses it. `--expires-in <seconds>` (or\n`--expires-at`) is the door the renewal seam's own error names. `--identity <creds>` re-mints for\nthe nkey the file already carries, so the new credential presents the SAME principal and every\ndurable keyed to it survives; combine it with a lifetime flag to rotate an expiring credential\nwithout churning the identity.\n\n## Login\n\n```bash\ncotal login --idp <auth base URL> [--client-id <id>]\ncotal logout --idp <auth base URL>\n```\n\nSigns you in to a per-user-auth mesh's IdP (device code flow) and caches the session; run it\nonce per machine. It prints your IdP subject, the id the operator grants against. When the trusted\n`/token` response advertises a same-origin space catalog, login validates and records that account's\nspaces immediately. After a\nlogin, every command on that mesh works under your identity: each connect takes a fresh IdP\nproof, exchanges it locally for a short-lived bearer, and is authorized against the actor\nledger at connect time. `logout` revokes the IdP session, clears its cache, and removes only that\naccount's discovered registry entries. See\n[identity & auth](identity-and-auth.md).\n\n## actor\n\n```bash\n# an upsert of the WHOLE row: name all three ACL flags, or pass --full for the wide defaults below\ncotal actor grant <actor> --sub <IdP subject> --scope a,b --allow-subscribe a,b --allow-publish a,b [--role <r>] [--label <l>]\ncotal actor grant <actor> --sub <IdP subject> --full [--scope a,b] [--allow-subscribe a,b] [--allow-publish a,b] [--role <r>] [--label <l>]\ncotal actor revoke <actor> (--sub <IdP subject> | --owner <u_\u2026>)\ncotal actor list\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | the folder's | Space whose ledger to manage |\n| `--sub <subject>` | none | The IdP subject (shown by `cotal login`) the actor belongs to |\n| `--owner <u_\u2026>` | none | The derived owner token (alternative to `--sub`) |\n| `--full` | off | Fill each ACL flag left off with its wide default; without it, `grant` refuses unless all three are named |\n| `--scope <a,b>` | `spawn,role:default` with `--full` | Capability scope (`''` = none; `spawn` = may run agents; `role:<r>` = may delegate role r; `admin` = cross-agent control; `supervise` = eligible for the closed remote manager-service view when the host enables it) |\n| `--allow-subscribe <a,b>` | `>` (all channels) with `--full` | Channel read ACL; the user's envelope, their agents can never read beyond it |\n| `--allow-publish <a,b>` | `>` (all channels) with `--full` | Channel post ACL; also the envelope for their agents' posting |\n| `--role <r>` | none | Role (scopes the task-queue consumer) |\n| `--label <l>` | none | Display label for `actor list` (never the IdP subject) |\n\nThe actor ledger is the single authorization source of a user-auth space: no row, no access.\n`grant --full` is the **full** envelope (all channels; scope `spawn,role:default`, so it may spawn and may delegate the default role). A\n`grant` that leaves off `--scope`, `--allow-subscribe` or `--allow-publish` without `--full` is\nrefused and writes nothing. A re-grant **replaces the whole row**, not the one field you name, so to add a capability spell\nevery field out: the new scope plus the row's current read set, post set, role and label\n(`cotal actor list` shows what a row holds). Under `--full`, a field left off does not stay as it\nwas: it reverts to the wide default in the table above. A re-grant retires the current interactive lifecycle through the running auth\nservice before it rotates the row, so copied bearers cannot cross an authorization update. If that\nretirement cannot be confirmed, the row is left unchanged and the command fails with the recovery\naction. `revoke` uses the same retirement before deleting the row, which lets a later grant create a\nreal successor instead of colliding with a live predecessor. `supervise` is separate from `spawn` and `admin`: it only makes a signed-in\nperson eligible for the host-provided closed remote manager-service view; it does not grant\nmanagement of another owner or a general host profile. `revoke` denies the next exchange and\nthe next connect with no restart, and evicts the principal's live connections. Managed-agent rows\n(written by the spawn path) live in a disjoint row space this command never touches. See\n[identity & auth](identity-and-auth.md).\n\n## doctor\n\n```bash\ncotal doctor auth [--fix]\n```\n\nCredential-health diagnosis and repair for this folder's mesh: renders every managed\ncredential as healthy / near-expiry / expired and ends in `healthy` or the exact next\ncommand; `--fix` applies the repairs it can. The one surface every stale-credential error\npoints at. `--fix` takes the mesh's renewal lease when the broker answers and refuses while\na manager or another doctor holds it; with no broker it repairs offline and says so.\n\n## join\n\n```bash\ncotal join --space <s> --name <n> [--role <r>] [--channel <c>]\ncotal join --link <url> | --token <t>\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and which credential |\n| `--name <n>` | none | Your presence name |\n| `--role <r>` | none | Your role |\n| `--channel <c>` | none | Channel to join |\n| `--kind <k>` | `agent` | Endpoint kind |\n| `--link <url>` | none | Join link (`cotal://\u2026`) |\n| `--token <t>` | none | Join token |\n| `--lifecycle-uid <uid>` | none | Required with `--creds`: the lifecycle UID minted alongside the credential (`COTAL_LIFECYCLE_UID` works too). A credential's durable grants name exact lifecycle-keyed resources, so `join` refuses to invent one |\n| `--tls` | off | Connect over TLS |\n\nAn interactive presence: join a space under your own name and role, without launching an agent\nharness. A `--link` or `--token` supplies the where and the auth in one value. See\n[Spaces](spaces.md) and [Identity and auth](identity-and-auth.md).\n\n## Manifest deploys\n\nA `cotal.yaml` manifest declares a whole mesh (channels, personas, roles, and ACLs) in one file.\nThree commands consume it, plus a read-only validator:\n\n```bash\ncotal up -f cotal.yaml # boot a fresh mesh from the manifest\ncotal spawn -f cotal.yaml # deploy the manifest additively onto a running mesh\ncotal down -f cotal.yaml # tear that deploy down (or --run <id> for one run)\ncotal topology view -f cotal.yaml # validate + view the access graph, change nothing\n```\n\n`up -f` and `spawn -f` differ in target: `up -f` brings up a new broker and applies the manifest;\n`spawn -f` requires an already-reachable mesh and applies additively (ownership-scoped). On a\nuser-auth mesh, `spawn -f` deploys over your own login (the deployer view, gated on ledger scope\n`spawn`): the manifest's agents land under your owner, a manifest claiming another owner is\nrefused, and seeding new channels additionally needs scope `admin`. Both take\n`--dry-run` to print the plan without mutating anything. `topology` validates the manifest and\nrenders its channel / role / ACL graph. See [Define a team](define-a-team.md) and the\n[manifest reference](manifest.md).\n\n## ext\n\n```bash\ncotal ext # same as `list`\ncotal ext add <npm-package>\ncotal ext remove <name>\ncotal ext list\ncotal ext root # print just the install prefix (scriptable)\ncotal ext seed [--repair|--reset|--force]\n```\n\nOperator-installed extensions: `add` installs an npm package into a cotal-owned prefix and records\nevery registry provider it contributes. Commands appear in help, completion, and dispatch; runtime\nproviders are lazy-loaded by commands such as `supervise`; local process providers participate in\n`status` and selective `down`. `remove` and `list` manage them. The `@cotal-ai/web` dashboard is the\ncanonical command/process example. Installed packages and their location are described in\n[config](config.md). When a package needs an export its linked `@cotal-ai/*` peer does not have,\n`add` rolls back and names which install is behind, as a later load of an installed one does.\n\nBare `cotal ext` lists the inventory, headed by the install prefix. That prefix is a cotal-owned npm\nroot kept **separate** from npm's own global tree. These packages never show up in `npm list -g`,\n`cotal ext` (or the Extensions section of `cotal status`) is the canonical inventory. `cotal ext root`\nprints only the path, for scripts. The versions shown are the manifest pin recorded at add time.\n\nRemoving an extension that owns a running local process is refused with the mesh root and its\n`cotal down <component>` command; stop it first so uninstalling the package never strands a process\nwhose lifecycle provider is gone.\n\n### Built-in connectors are seeded extensions\n\nThe first-party agent connectors (`claude`, `opencode`, `codex`, `hermes`, `jcode`, `pi`) are not compiled into\nthe binary. They are seeded on first run through the **same** `ext add` path a third party uses, and\nappear in `cotal ext list` like any other extension. So you can remove one you do not want\n(`cotal ext remove @cotal-ai/connector-hermes`), and a deliberately-removed connector STAYS removed\nacross upgrades. `cotal ext add <your-package>` adds a third-party connector the same way. The web\ndashboard (`@cotal-ai/web`, providing `command:web`) is the seventh built-in seeded on the same path.\n\n`cotal ext seed` is the maintenance entry for that seeding (it runs automatically on the first real\ncommand of each boot, so you rarely call it). Each seeded connector's `\u2713 added` line goes to stderr,\nso the command that triggered the seed keeps stdout to itself:\n\n| Flag | Meaning |\n|---|---|\n| (none) | Reconcile: seed any never-seeded built-in, refresh a seeded one whose version the binary bumped, leave a removed one removed. A no-op once current. |\n| `--repair` | Recover after an interrupted seed or a lost authority (rebuilds the interrupted connector; restores the removed-vs-never-seeded record from its durable backup). |\n| `--reset` | Discard the record and re-seed all seven built-ins (the six connectors plus the web dashboard). **Resurrects any you removed.** Rebuilds cleanly over corrupt seed state. |\n| `--force` | Re-seed the built-ins even when the version stamp is current or a downgrade. |\n\nWhen a newer `cotal` advances the operator-global seed store to its generation, it prints one\nmigration line naming the old and new generations, the exact CLI entry that wrote the store, the\ncommit timestamp, and `seed/stamp.json`. That writer and timestamp are kept in the stamp, so a later\nolder CLI refusal can say which executable wrote the generation it will not overwrite and when.\nLegacy generation-only stamps remain readable; their refusal simply has no writer provenance to add.\n\nAn older `cotal` refuses a seed store written by a newer version. When it can verify a sufficient\n`cotal` executable on PATH or at the installer's `~/.local/bin/cotal` location, the refusal names\nthat absolute path so a reduced service PATH does not select the older binary again. Otherwise it\nkeeps the generic newer-version instruction. `--force` rebuilds the store for the running older\nversion without discarding the ever-seeded authority. `--reset` still exists for corrupt state and\nresurrects deliberately-removed connectors.\n\nA source-checkout CLI (`pnpm cotal`, `tsx bin/cotal.ts`, `node bin/cotal.ts`, or a suite child of\nthose) refuses to write or garbage-collect that store. The refusal names the path, the generation\nit declined, and `COTAL_SKIP_CONNECTOR_SEED=1` as the way to run other commands from a checkout,\nbecause pointing `$XDG_CONFIG_HOME` at a scratch dir alone does not lift it. With the skip set, a fresh\nconfig gets no built-in connectors (`cotal ext list` shows none), so add one from the checkout with\n`cotal ext add <checkout>/extensions/connector-claude-code` against that scratch `$XDG_CONFIG_HOME`. `COTAL_HOME` does not relocate this\nstore. An entry that cannot be proven as a released install is refused the same way. Isolated\nrelease tests that must seed from a checkout-shaped `bin/` set `COTAL_ALLOW_CHECKOUT_SEED=1` after\npointing `$XDG_CONFIG_HOME` at a scratch dir; that override is documented here, not on the refusal\nline. An opt-in write still records the checkout path in `seed/stamp.json` as `writtenBy`.\n\nThe default connector for a bare `cotal spawn` (no `--agent`) is the persona's `agent:` pin if it\nhas one, else `claude`; set `COTAL_DEFAULT_AGENT` (e.g. `opencode`) to change the fallback. It is\na default, so a persona that pins its harness still wins over it. An `--agent` naming a removed\nconnector fails loud with the exact\n`cotal ext add` to restore it. Set `COTAL_SKIP_CONNECTOR_SEED=1` to turn off the automatic first-run\nseed/refresh entirely (for a controlled or offline setup that manages connectors by hand); `cotal ext\nseed` still runs on request. `cotal agent-bearer` never takes the seed at all: it is exec'd by\nspawned seats on every bearer refresh, so it neither reconciles nor is refused by the store's\ngeneration (see [Plumbing](#plumbing)).\n\n## completion\n\n```bash\ncotal completion <bash|zsh|fish|powershell> # print a stub to eval / source\ncotal completion install [shell] # install it persistently\n```\n\nPrints or installs shell completion. Completion candidates come from each command's declared flags\nand, where useful, live mesh state (spaces, personas, managed agents) resolved offline.\n\n## feedback\n\n```bash\ncotal feedback \"<summary>\" [--type <t>] [--email <e>] [--details <text>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--type <t>` | none | `bug` \\| `idea` \\| `friction` \\| `praise` \\| `other` |\n| `--details <text>` | none | Longer free-form details |\n| `--severity <s>` | none | `low` \\| `medium` \\| `high` |\n| `--area <a>` | none | The part of Cotal this concerns |\n| `--email <e>` | git email | Contact email (required on the keyless public path) |\n| `--name <n>` | none | Your name (optional) |\n| `--url <url>` | keyed / public intake | Intake URL override |\n| `--key <k>` | `COTAL_FEEDBACK_KEY` | Feedback key |\n\nSends feedback to the Cotal developers. With a key (`--key` / `COTAL_FEEDBACK_KEY`) it routes to the\nkeyed beta intake; without one it goes to the public `cotal.ai` intake and requires a contact email\n(`--email` / `COTAL_FEEDBACK_EMAIL`, else your git email). Run a self-hosted intake with\n[`feedback-intake`](#server-daemons).\n\n## run\n\nOperate durable workflow runs (cotal-lang programs) from the terminal.\n\n```bash\ncotal run start --file <program> [--timeout <dur>] [--local]\ncotal run resume <runId> [--local --file <program>]\ncotal run ps [--endpoint <ep>] [--json]\ncotal run journal <runId> [--endpoint <ep>] [--json]\ncotal run answer <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]\ncotal run amend <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]\ncotal run migrate <runId> --local --file <program> [--endpoint <ep>]\n```\n\n`start` hands the program to the mesh's manager, which validates it, mints the run id (the record\nnever takes a caller-supplied one), drives it in its own process, and answers with the id once the\nrun is recorded; a program that does not validate is refused with every problem listed. `resume`\nasks the manager to take an existing run back and continue it from its step journal; the source is\nthe recorded program, so no `--file` is taken. Neither takes `--endpoint`: the manager records\nits runs under its own endpoint, and naming another is refused. `ps` lists the run records and\n`journal` renders one run's durable records; both only inspect. An open pause prints its question.\nA pause settled with an accepted answer prints its value as JSON plus the recorded answerer,\nartifact when present, time, and answer id, then one `amended` line per later amendment, in the\norder the store committed them, so the last is the current position.\nExpired pauses and ordinary steps print no answer line.\n`--json` on `ps` or `journal` prints each row the manager answers with (or `--local` reads) as one\nJSON object per line. A `ps` row carries `runId`, `endpoint`, `state`, `holder`, `epoch`,\n`journalHigh`, `forkedFrom`, `startedAt` and `programHash` (the values the program's `run()`\nreports; `programHash` is absent for a run with no recorded program), and `revoked` or\n`revocationUnreadable` when the marker says so. A `journal` row is an `activation` or a `step`. A\nstep row carries its `step` key, the `effect` kind and its `name`, `state`, `outcome`, the recorded\n`status` and `errorCode` once settled, and `startedAt` and `endedAt` in epoch milliseconds. An open\npause adds its `asks`, its `deadlineAt`, and for a checkpoint the `onExpiry` it was armed with; a\nsettled pause adds its `answer` and its `amendments`, as the text view prints them. A field the\njournal does not record is absent: a checkpoint opened before `onExpiry` was recorded carries none.\nThe run header and errors go to stderr, so stdout carries only rows; an unreadable revocation marker\nprints its reason there and still exits 1. The text view is presentation and is not a stable\nparsing target. `--json` on any other verb is refused.\n`answer` resolves an open\ncheckpoint through the manager, presenting as the holder that armed it; the manager records the\nanswerer from your credential, so no `--by` is taken there. A settled step refuses a second\n`answer`. `amend` records a changed position on a settled checkpoint or `ask`: it files a new\nanswer beside the accepted one, naming it, and the journal lists it under the step. The pause stays\nsettled and the run keeps the answer it acted on. A step that is still open or settled without an\nanswer refuses an amend. A spawned seat may amend only an answer recorded under its own name. `migrate` runs the migrate check of an\nedited program against a run's journal, from this terminal under a read credential (`--local`\nonly; the manager serves no run-migrate command): it prints whether the migration is admissible,\nevery orphaned step with its verdict and code, and exits 0 on admissible and non-zero on not. It\nwrites nothing: the commit that would file the migration is not reachable yet, and the report\nsays so. `--timeout` sets the default\ncheckpoint timeout for a drive (default 1h). `--local` drives in this process instead, over one\nconnection per invocation under the run's own credential minted from the project folder's trust\nmaterial, and is the path on a bare broker with no manager or for a run with no recorded program\n(`cotal run resume <runId> --local --file <program>`); `answer --local` and `amend --local` take\n`--by <who>`. On a\nuser-auth mesh the host's own manager refuses the family by name, and `--local` has no credential\nthere. A participant's manager started with `cotal supervise` hosts a logged-in user's runs through\nits issuing host: the auth callout issues the user's manager connection, and every `run` verb rides\nthe versioned rail under that issuance.\n[User-auth run start](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/user-auth-run-start.md)\nrecords the path. The guide is [workflows](workflows.md).\n\n## Server daemons\n\nTwo long-lived infra roles ship with the CLI. They are not part of everyday operation; the delivery\ndaemon comes up automatically with `cotal up --detach` in auth mode.\n\n```bash\ncotal deliver --space <s> [--server <url>] [--creds <file>] [--root <dir>]\ncotal auth-service --space <s> --server <url> [--port <n>] [--exchange-public-port <n>] [--exchange-public-url <https://\u2026>] [--exchange-trusted-proxy]\ncotal feedback-intake --keys <keys.json> [--port <n>] [--creds <file>]\n```\n\n`auth-service` runs a user-auth space's identity plane: the NATS auth callout, the\ncapability-gated local exchange and JWKS, and, when `--exchange-public-port` is set, the closed public\nexchange/discovery face forwarded by an HTTPS reverse proxy. `--exchange-public-url` is the proxy URL\nadvertised to clients; `--exchange-trusted-proxy` opts into last-hop `X-Forwarded-For` attribution.\n`cotal up --user-auth` starts and supervises the service for you, so you run it directly only to\nrecover one by hand.\n\n`deliver` runs the server-side Plane-3 delivery daemon: the durable backstop and membership/ACL\nauthority. It is auth-mode-only and single-instance (`--shard`/`--shards` accept only `N=1`);\n`--dev-mint` mints a scoped cred from the local signer for standalone dev. `--creds` can start a\ndaemon that already looks healthy, but production renewal is not that file alone: the manager and\nthe daemon must address one credential store. On a stock split host with two project roots, a\ndirect `deliver` is not an independent repair; keep the daemon under `cotal up` on the broker\nhost, or inject the same store into both processes ([embedding](embedding.md#supervisor-signing-authority)).\nTyped by hand on the workstation, `deliver` dials the broker recorded for `--space` in the mesh\nregistry (a mismatching `--server` is refused before any dial, and a record for a different\nworkspace root is refused outright); with no record for the space it falls back to the local mesh.\nThe daemon serves the workspace root that `--root <dir>` names, which must hold `.cotal/`, or else\nthe nearest `.cotal/` above its working directory. With neither, it refuses at start and names the\ndirectory it searched from, before it reads a credential or dials a broker.\nSee the [delivery daemon](delivery-daemon.md). `feedback-intake` runs a self-hosted feedback server\n(requires `--keys` and a scoped `--creds`), announcing submissions into a space channel; flags\ninclude `--host`/`--port`, `--store`, `--space`/`--channel`, `--max-bytes`, and `--rate-limit`.\n\n## Plumbing\n\n`cotal __complete <words\u2026>` is the internal entry the shell-completion stubs call to emit candidates\nfor the current command line; you never run it directly. `cotal agent-bearer` is machine-facing\nplumbing on user-auth meshes: spawned agents exec it to print a fresh short-lived bearer from their\nspawn-time secret; you never run it directly either. Its local arm uses `--dir` to discover the\ncapability-gated loopback service. A remotely enrolled, already-granted agent instead receives\n`--exchange-url <https://base>` in its launch argv: that arm sends `{owner, actor, actorToken}` to the\npinned public exchange with no local capability, follows no redirects, and refuses every non-HTTPS\nURL because the actor token is the credential in the request body. Because a seat execs it on every\nbearer refresh, it skips the connector-seed boot gate entirely: it reads one 0600 token file,\nexchanges it and prints the bearer without consulting or writing the operator-global seed store, so\na newer store generation cannot refuse a live seat's refresh. `--manager-call` asks for the\ninstance-bound `manager-caller` view; `--manager-instance <id>` selects an explicit live candidate.\nThat mode still prints only the raw token and does not update `--health-file`. A spawn runs it once as the agent auth\npreflight. When it fails there without printing a sentence of its own, the refusal names the cause: the 30 second\ntimeout, the signal that killed it, or its exit code. (`cotal start` is a removed tombstone: it\nerrors and points you to `cotal spawn --detach`.)\n"
|
|
57781
|
+
"body": "# `cotal` CLI reference\n\n> **Reference**: describes the TypeScript reference implementation (the `cotal` CLI), not the wire contract. \xB7 **For:** operators \xB7 **Wire contract:** [SPEC](../SPEC.md)\n\n`cotal` is the operator command line for the reference implementation: bring a mesh up, mint\nidentities, launch agents, watch what they do, and tear it all down. It is a thin client over the\nwire contract: the normative subjects and schemas live in the [SPEC](../SPEC.md); this page is\nlookup material for the commands, not a walkthrough; if you are new, start with\n[Getting started](getting-started.md).\n\n## Running it\n\n```bash\nnpm install -g cotal-ai # puts `cotal` on your PATH (needs Node 22+)\ncotal --help # every command, grouped\ncotal --version # cotal-ai version + each installed extension's (also `cotal -v`)\ncotal <command> --help # one command's flags and usage\n```\n\n`npx cotal-ai <command>` runs it without a global install; in a dev clone, `pnpm cotal <command>`\nruns it through `tsx` with no build step. Bare `cotal` prints help. Every command generates its own\n`--help`, usage, and shell completion from its declared flags.\n\nAn undeclared flag is a usage error, and so is a flag given more than once unless it is\nrepeatable, as `--opt` and `down --session-store` are. The command prints the error and its help,\nexits 1, and does not run.\n\nCommand output, including error lines on stderr and the guided `setup` and `meshes add` prompts,\nis colored only when stdout is a terminal, so piped or redirected output is plain text. A non-empty\n`NO_COLOR` turns color off on a terminal too. `FORCE_COLOR` turns color on even when output is\npiped, unless it is `0` or `false`, and it takes precedence over `NO_COLOR`.\n\nCommands come from the surfaces the binary composes: the base mesh CLI, the manager\n(`supervise`), and the delivery daemon (`deliver`), plus any operator-installed extensions.\n`cotal ext add <npm-package>` installs any registry providers a package contributes: commands,\nruntimes, and local process lifecycle descriptors. The `web` dashboard and optional manager\nruntimes ship this way.\n\n## Commands\n\n| Area | Command | Purpose |\n|---|---|---|\n| Set up & lifecycle | [`setup`](#setup) | Guided, configure-only setup (installs, seeds personas; launches nothing) |\n| Set up & lifecycle | [`update`](#update) | Reconcile first-party extensions and check or opt into a coherent CLI upgrade |\n| Set up & lifecycle | [`up`](#up) | Start a local mesh (nats-server + JetStream), or boot a whole manifest with `-f` |\n| Set up & lifecycle | [`down`](#down) | Stop the whole stack, selected registered components, or a manifest deploy |\n| Set up & lifecycle | [`backup`](#backups) | Create an offline full-space or registry-only artifact from a preserved cut |\n| Set up & lifecycle | [`clean`](#clean) | Configurable cleanup: purge history (live), or wipe the local store / identity (stopped) |\n| Set up & lifecycle | [`meshes`](#mesh-registry) | List the running meshes on this machine |\n| Set up & lifecycle | [`sync`](#mesh-registry) | Refresh the signed-in account's advertised spaces |\n| Set up & lifecycle | [`use`](#mesh-registry) | Set the default mesh a bare `cotal spawn` joins |\n| Set up & lifecycle | [`status`](#mesh-registry) | Read-only diagnostics for setup, processes, and the selected mesh |\n| Agents & personas | [`spawn`](#spawn) | Launch an agent from a persona (foreground, or `--detach` via the manager) |\n| Agents & personas | [`models`](#models) | List connector model catalogs and variants from the manager |\n| Agents & personas | [`ps`](#managed-seats) | List managed agents and their mesh status |\n| Agents & personas | [`stop`](#managed-seats) | Ask the manager to stop a managed agent |\n| Agents & personas | [`attach`](#managed-seats) | Stream and drive a managed agent's terminal (pty runtime) |\n| Agents & personas | [`input`](#input) | Type one line into a managed agent's terminal without attaching |\n| Agents & personas | [`personas`](#personas) | List, show, edit, create, or remove local personas |\n| Agents & personas | [`supervise`](#supervise) | Run a manager daemon (the agent supervisor / control plane) |\n| Agents & personas | [`service`](#service) | Run the manager as a user service (survives logout and reboot) |\n| Agents & personas | [`runtimes`](#runtimes) | List the agent runtimes the manager can spawn through and whether each is reachable |\n| Agents & personas | [`seats`](#seats) | List the pty seat custodians an earlier Linux manager left, and drain the ones whose agent has exited |\n| Agents & personas | [`reconcile-gate`](#reconcile-gate) | Unfreeze an issuance gate left frozen by a crashed restart when the successor cannot boot-heal it (holder gone, complete CONNZ sweep) |\n| Messaging & watching | [`endpoints`](#endpoints) | List every endpoint in the live presence roster, including infrastructure |\n| Messaging & watching | [`describe` / `invoke`](#endpoint-control) | Resolve a v0.4 service's command surface off the wire; invoke one command by name |\n| Messaging & watching | [`send`](#send) | Send one message, then exit: DM a peer, post a channel, or ask a role |\n| Messaging & watching | [`channels`](#channels) | Inspect or set the channel registry |\n| Messaging & watching | [`history`](#history) | Clear retained message history |\n| Messaging & watching | [`console`](#console) | Live protocol view for a space (TUI, or `--plain` line stream) |\n| Messaging & watching | [`web`](#web) | Browser dashboard (installed as the `@cotal-ai/web` extension) |\n| Auth & meshes | [`mint`](#mint) | Mint a creds file for a space (static auth mode) |\n| Auth & meshes | [`login`](#login) | Sign in to a per-user-auth mesh's IdP (once per machine) |\n| Auth & meshes | [`logout`](#login) | Revoke the IdP session and clear the cached login |\n| Auth & meshes | [`actor`](#actor) | Manage a user-auth space's actor ledger (grant / revoke / list) |\n| Auth & meshes | [`doctor`](#doctor) | Credential-health diagnosis and repair (`doctor auth`) |\n| Auth & meshes | [`join`](#join) | Join a space as your own presence (interactive) |\n| Manifest | [`topology`](#manifest-deploys) | Validate and view a mesh manifest's access graph (read-only) |\n| Extensions & misc | [`ext`](#ext) | Install / remove operator CLI extensions |\n| Extensions & misc | [`completion`](#completion) | Print or install shell completion |\n| Extensions & misc | [`feedback`](#feedback) | Send feedback to the Cotal developers |\n| Extensions & misc | [`deliver`](#server-daemons) | Run the server-side Plane-3 delivery daemon |\n| Workflow runs | [`run`](#run) | Operate durable workflow runs: start, resume, list, inspect, answer a checkpoint, check an edited program with migrate |\n| Extensions & misc | [`feedback-intake`](#server-daemons) | Run a self-hosted feedback intake server |\n\nThe manifest modes of `up`, `spawn`, and `down` (`-f <cotal.yaml>`) plus `topology` are covered\ntogether under [Manifest deploys](#manifest-deploys).\n\n## setup\n\n```bash\ncotal setup [--full] [--demo] [--yes] [--skills]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--full` | off | Redo the full guided flow (implies `--demo`) |\n| `--demo` | off | Also seed the guided expert team (`david`, `sven`, `me`) |\n| `--yes`, `-y` | off | Non-interactive accept-all (for agents / CI) |\n| `--skills` | off | Reconcile Cotal skills only through installed connector providers, plus `~/.agents/skills`. Refused with `--full` or `--demo`. |\n\nGuided setup is **configure-only**: it checks prerequisites, invokes installed connectors' declared setup providers, and\nseeds persona files, and it launches nothing (no mesh, no web, no manager). First run gets the\nnarrated flow; later runs print a status card. By default it seeds one `default` persona; the\n`david`/`sven`/`me` team is opt-in via `--demo`. `cotal status` points stale Claude skills and\nout-of-date `.agents` skills at `cotal setup --skills`, not unscoped `setup`. See [Getting started](getting-started.md) and, for\nmaintainers, [setup internals](setup-internals.md).\n\nWhen a mesh resolves, setup seeds that mesh's recorded `.cotal/agents` catalog, the same catalog a\nfollowing `cotal spawn` reads. It prints the absolute destination. On a fresh machine with no mesh it\nuses this folder and says why; when several meshes are available and none is selected, it refuses\nrather than choosing a catalog.\n\n## update\n\n```bash\ncotal update [--self] [--space <s>] [--server <url>] [--creds <path>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--self` | off | If a newer release exists, install that exact validated `cotal-ai` version globally and reconcile through the newly installed binary |\n| `--space`, `--server`, `--creds` | resolved mesh | Select the running manager whose continuity state is reported |\n\nWithout `--self`, `update` keeps the installed first-party surfaces coherent with the running\nbinary: it force-reconciles the four built-in connectors, then reinstalls other `@cotal-ai/*`\noperator extensions at the binary's exact version. Each extension runs in an isolated child, so one\nfailure cannot poison later replays. It then checks npm; a newer binary is an informational notice\nwith `cotal update --self` as the next command, not an automatic install.\n\nAfter disk reconciliation, `update` reads the selected running manager. A machine with no recorded\nmesh has no running manager to observe, so that read is skipped and the command completes. The same\nholds when every recorded mesh is down and none is selected. A remote user mesh, registered with\n`cotal meshes add --mode user`, is named and skipped: its manager runs under another install, so\nthere is no custody on this machine to preserve, and a `legacy` verdict still comes only from a\nmanager this machine read. A recorded mesh that is down is still a\nrefusal when the command selects it, with `--space` or by running inside its project, and so is a\nnamed space that is not running. With several meshes running and no `--space`, `--server` or\n`--creds`, the install is machine-wide, so every running manager is reported in turn, each under its\nspace name, before anything is written; a `legacy` verdict on any of them makes the whole run not a\nhot update. A selector flag still reports one manager. A manager without a\ncustody generation is reported as `legacy`: it cannot preserve its manager-owned PTYs, so the\ncommand says that this is not a hot update and prints `exact`, `fork`, `fresh`, or `drain-only`\nfor every seat. This report sends no stop, preservation-commit, or replacement command.\nIt does not preserve a running PTY on a legacy manager. The built-in pty runtime spawns\nin-process on every platform and reports `legacy`. On Linux it still adopts seats that an earlier\nmanager left under a detached custodian, but it starts no new custodian. An incompatible native\n`@lydell/node-pty` or ConPTY ABI break remains an explicit per-seat maintenance cut.\n\nWith `--self`, the selected running manager is reported before any global install. When a newer\nrelease exists, Cotal then installs the exact version it validated, resolves and verifies that\npackage in npm's global root, then launches that binary with the same `--space` / `--server` /\n`--creds` selection to reconcile connectors and first-party extensions to the new generation. An npx\nor dev-clone invocation therefore installs and continues through a separate global copy; it never\nclaims the already-running process changed. If the binary is current, `--self` performs the normal\nlocal reconcile without reinstalling it.\n\nThird-party extensions are listed with their installed version and recorded spec but are not\nauto-updated in v1. Floating third-party updates require `@cotal-ai/*` peer-range validation and are\na future follow-up. A failed connector/extension install, npm metadata check, or requested global\ninstall is reported and makes the command exit nonzero. Independent extension attempts continue so\nthe output includes every failure; an unavailable npm registry does not undo a completed local\nreconcile, but the command still exits nonzero because it could not establish that the install is\ncurrent.\n\n## up\n\n```bash\ncotal up [--detach] [--open] [--space <s>] [--server <url>] [--channels <path>] [--runtime <name>]\ncotal up --user-auth --idp <url> [--exchange-public-port <n> --exchange-public-url <https://\u2026> [--exchange-trusted-proxy]]\ncotal up --tls-cert <cert.pem> --tls-key <key.pem> # serve broker TLS (both, or neither)\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\ncotal up -f <cotal.yaml> [--dry-run] [--runtime <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--server <url>` | auto (free local port) | Listen URL override |\n| `--host <host>` | none | Bind host override for a **fresh** broker boot: an IP or hostname only, never a URL (that is `--server`) and never `host:port` (the port comes from `--server` or its default); a URL or port-bearing value is refused pointing at the right flag. With no `--server`, the broker URL is derived from it, so `--host <addr>` alone is enough to make a mesh reachable at that address; a `--host`/`--server` pair naming different addresses is refused. A wildcard bind (`0.0.0.0`, `::`) keeps a dialable loopback URL. Recorded on the mesh and reused by every later manager launch, so a repair or resume keeps remote [`attach`](#managed-seats) working. A live refresh (`\u2713 mesh already running`) does not rewrite `.cotal/auth/server.conf` or rebind nats; stop the broker, then re-run `up --host` |\n| `--space <s>` | the folder's name | Space name |\n| `--store-dir <dir>` | none | JetStream store directory (recorded; a repair up reuses it) |\n| `--max-file-store <bytes>` | nats-server's dynamic cap | JetStream file storage cap in bytes (a positive integer, no unit suffix). Without it nats-server sizes the store at start as three quarters of the free space on its filesystem. The cap is fixed at broker start: a running broker cannot change it (`cotal down` first), `down --preserve-state` keeps it for the resume, and a resume with a different value is refused. Not accepted with `-f` |\n| `--channels <path>` | `.cotal/channels.json` if present | Channel-registry seed file (JSON). An explicit path that is missing is an error |\n| `--restore <dir>` | none | Restore a completed offline backup before exposing the normal listener |\n| `--restore-only registry` | artifact selection | Restore only the registry component |\n| `--accept-missing-source` | off | Explicit disaster consent when the inode-bound preserved source is absent |\n| `--accept-stale-checkpoint` | off | Explicit consent to resume a seat whose checkpoint was captured outside its recorded recency horizon |\n| `--open` | off (auth) | Unauthenticated dev mesh: no JWT, no ACLs |\n| `--user-auth` | off | Per-user auth: people `cotal login`; connects are authorized against the actor ledger |\n| `--idp <url>` | none | With `--user-auth`: the IdP auth base URL to pin on first enable |\n| `--exchange-public-port <n>` | none | With `--user-auth`: add the public exchange face on this loopback port, for an HTTPS reverse proxy to forward to |\n| `--exchange-public-url <https://\u2026>` | none | With `--exchange-public-port`: advertise the reverse proxy's HTTPS URL in discovery |\n| `--exchange-trusted-proxy` | off | With `--exchange-public-port`: attribute public failure buckets to the last `X-Forwarded-For` hop. Enable only when the listener is reachable solely through a trusted proxy; otherwise the socket address is used |\n| `--detach` | off | Run in the background (stop with `cotal down`) |\n| `--tls-cert <path>` | none | PEM certificate to serve TLS with. Must be given together with `--tls-key`. Before starting the broker, Cotal checks readability, private-key mode, key/certificate match, the validity window, and host coverage. `nats-server` accepts an expired certificate and leaves the failure to clients, so Cotal performs these checks first. The decision is recorded; a later bare `cotal up` keeps serving TLS |\n| `--tls-key <path>` | none | PEM private key for `--tls-cert`. Refused if group- or other-readable (tighten to `600`) |\n| `--file <cotal.yaml>`, `-f` | none | Launch a whole mesh from a manifest |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--runtime <name>` | `pty` (or the manifest's, with `-f`) | Agent runtime for the mesh manager (`pty` built in; others are installed extensions, explicit-only). Resolved + probed before the broker starts; an uninstalled/unreachable runtime fails loud. With `-f`, overrides the manifest's runtime |\n| `--max-sessions <n>` | 64 | Live-session ceiling for the mesh manager. Each console pane and each `cotal attach` is one session, so size for agents \xD7 panes, not agent count. Recorded on the mesh and reused by every later manager launch, so a repair or resume does not silently drop back to 64. A running manager cannot change it: `cotal down` first, then `cotal up --max-sessions <n>` |\n| `--no-manager` | off | Broker-only boot: start the broker and, in auth mode, the delivery daemon, and no local manager. A refresh under the flag of a mesh whose manager is live refuses rather than keeping or stopping it (`cotal down manager` first). Cannot be combined with `--runtime`, `--max-sessions`, or an agent-declaring manifest |\n| `--rotate-sys` | off | Rotate the space's system account and re-mint its two `$SYS` creds. Needs a stopped mesh; refused with `--open` |\n\n`cotal up` boots a local nats-server with JetStream and, in auth mode (the default), JWT auth and\nper-agent ACLs; `--detach` records the mesh so `cotal spawn` from any directory can find it. With no\n`--server`, it auto-selects a free port if the default address is taken; an explicit `--server`\nstays fail-loud on collision. `--detach` also brings up the control plane (delivery daemon in auth\nmode, then the manager). `--no-manager` is the broker-only mode: it boots\nthe broker (and the delivery daemon in auth mode) and starts no manager, so there is no manager\npidfile to leave stale. A refresh under the flag of a mesh whose manager is live refuses rather\nthan keeping or stopping it: `cotal down manager` first. For a split topology with a manager, wait for `.cotal/manager.<spaceKey>.log` to contain `\u2713 manager up`, then `cotal down manager` on that\nhost and run [`supervise`](#supervise) against the remote broker; see\n[Run a mesh](run-a-mesh.md). `cotal up --detach` prints `\u2713 running in the background:` with\n`manager` listed (pidfile liveness, not a teardown boundary); with `--no-manager` the line lists\nonly what actually started. Ctrl-C on a foreground `up` stops the manager through the same stop as\nbare `cotal down` (see [`down`](#down)), then the rest of the stack, and reports managed agents\nunder the same rule: when the manager stop is refused, Ctrl-C prints the refusal with the reap route\nand leaves the stack running. The `-f` form is a\n[manifest deploy](#manifest-deploys).\n\nA repair `up` on a mesh whose broker died reopens the store its record names, and refuses a\ndifferent `--store-dir` rather than silently opening a second store.\n\nThe generated `.cotal/auth/server.conf` is written on a real broker boot and is not an\noperator-owned config. `--host` changes that file only when nats is actually started. A unit\nrestart that leaves an answering listener in place is a refresh, not a rebind.\n\nOn an existing mesh, `cotal up` reconciles the presence and lease bucket TTLs. It writes a reserved\ncanary and waits for the bucket to expire it before reporting success. If the broker accepts the\nstream update but the backing store does not persist or enforce it, `up` exits nonzero with a TTL\npersistence error instead of trusting the value returned by stream info. A refresh that restores a\nmissing manager says so with its pid (`\u2713 restored in the background: manager (pid N)`); a refresh\nthat finds everything already running prints only the `\u2713 mesh \"<space>\" already running` line. A\nfirst boot starts its manager without the restore line.\n\n\n`--user-auth --idp <url>` starts the space's auth service alongside the broker: the NATS\nauth callout plus its capability-gated local exchange, and optionally the closed public exchange\nface configured by the three `--exchange-*` flags above. The service is torn down with `cotal down`,\nand a re-run of `cotal up` heals a dead service on a running broker. `up` waits for the service to\nfinish binding: while the daemon it launched (or found running) stays alive, the wait extends past\nthe base 15s up to 60s; a daemon that exits is refused at once with \"exited before becoming ready\",\nand one alive past 60s is refused as \"alive and still starting\" (wedged), naming the pid record and\nthe service log. `--user-auth` and `--open`\ncontradict each other and are refused loudly; a running broker cannot change auth mode\nwithout a `cotal down` first. See [identity & auth](identity-and-auth.md).\n\n`--rotate-sys` renews the two `$SYS` credentials (`membership-observer`, `connection-evictor`).\nThey carry a 30-day expiry and nothing re-signs them in place, because the system-account seed is\nnever persisted, so they are renewed by issuing a **new system account** under the same broker\noperator and minting fresh creds against it. A plain re-`up` does **not** do this: it reuses the\nexisting trust record, and its `$SYS` creds along with it.\n\nThe rotation is safe to run on a real space, with one operational cost. The data account, the account\nsigning key, every agent credential minted from it, and the JetStream store are all untouched; what\ndies is the retired system account, and with it any out-of-band copy of the old `$SYS` creds, on every\nbroker that loads the rotated config. The cost is that **earlier full backups stop being restorable**\n(see below), so this is not a no-consequence operation. It needs the broker to restart on the rewritten\nconfig, so it runs as part of a boot:\n\n```bash\ncotal down\ncotal up --rotate-sys --detach # agents reconnect; nothing is re-provisioned\ncotal doctor auth # both $SYS creds healthy again, 30 days out\n```\n\nA rotation is a stopped, fresh boot, and anything that is not one refuses it, all for the same reason\n(the on-disk material and the broker it runs on must never end up on different generations):\n\n- a live mesh, because the running broker would keep serving the retired account;\n- an open mesh, whether that comes from `--open` or from `broker.auth: false` in a manifest, which\n has no system account at all;\n- `--restore`, because reinstating a trust root and superseding it in one command leaves no way to\n say which authority the mesh came up on;\n- an unfinished restore or resume attempt on this root, including one `cotal up` would recover on\n its own, because those paths can adopt a live listener and return without booting a broker;\n- a root that hosts more than one space, because the system account lives in the shared broker\n record and a rotation would retire every tenant's, while the root holds one `$SYS` cred pair\n pinned to one data account.\n\nTwo things to know before you run it:\n\n- **The retirement is config-load-bound.** Old `$SYS` creds are refused by any broker that loads the\n rotated config. A stale `nats-server` still running the *previous* config in memory would keep\n honouring them, so stop every broker for this root first. `--rotate-sys` refuses if this root's\n mesh is recorded as running, if anything unidentified is answering at the address it was given, or\n if the root's pid file names a live (or unreadable) process. Those are Cotal's own ownership\n records, not a scan of the process table: a `nats-server` you started by hand against this root's\n `server.conf` on some other port writes none of them and will not be seen. Do not run one.\n- **It invalidates earlier full backups.** A full artifact binds to the trust chain it was taken\n against, and that commitment covers the operator JWT and the system account. Every full backup\n taken before a rotation refuses to restore afterwards, so take a fresh `cotal backup` once the\n rotated mesh is up. `cotal up --restore` names this case when the data account still matches.\n\nThe commit is not atomic (a trust-record write plus two credential writes), so an interrupted\nrotation leaves the record ahead of the creds. That split is detected rather than silent: every\n`cotal up` on an auth mesh, and every `cotal doctor auth`, compares each `$SYS` cred's issuer against\nthe persisted record and names the retired account. `up` warns rather than refusing, because these\ncreds power the membership graph and live eviction, both of which degrade fail-soft; the mesh is not\nworth taking down over them. Re-running the rotation heals it, at the cost of one generation.\n\nWhile those creds are expired the mesh keeps delivering messages, but the\n[membership feed](delivery-daemon.md) and live connection eviction stay down; `cotal doctor auth`\nand the manager's log both name the credential and this repair.\n\n## down\n\n```bash\ncotal down\ncotal down --with-agents\ncotal down --preserve-state [--store-dir <dir>] [--session-store <dir> \u2026]\ncotal down manager [delivery auth web nats ...]\ncotal down web [--space <name>]\ncotal down -f <cotal.yaml> | --run <id> [--dry-run]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--file <cotal.yaml>`, `-f` | none | Tear down this manifest's deploy |\n| `--run <id>` | none | Tear down one `spawn -f` run by id |\n| `--space <name>` | current mesh | With components: the mesh whose target-addressed components (e.g. `web`) to stop |\n| `--dry-run` | off | Print the manifest teardown or selected components, mutate nothing |\n| `--with-agents` | off | Bare whole stack only: also stop and deprovision every managed agent |\n| `--preserve-state` | off | Bare whole stack only: fence the manager, retain principals and durable state, stop and prove the stack down, then publish `ready` |\n| `--store-dir <dir>` | `.cotal/nats` | With `--preserve-state`: the actual store path (required for a custom store) |\n| `--session-store <dir>` | none | With `--preserve-state`: a harness transcript store directory to capture with every continuation-capable retained seat. Repeatable. No default and never inferred from a connector name; a path that does not exist or is not a directory is refused before anything stops |\n\nBare `cotal down` stops the whole local stack in dependency order and leaves managed agents running\nwhen their runtime lets them outlive the manager. Before signalling the manager it verifies the spare\ncapability of the exact recorded manager, which records what that manager's stop does with its\nseats, and it reports the agents left behind plus `cotal down --with-agents` as the explicit reap.\nWhen the manager had no managed agents, it prints no report.\nThe built-in pty runtime keeps each PTY inside the manager process, so those seats cannot outlive\nit: every manager stop stops and deprovisions them, and `down` reports them as stopped. Every manager\nstop the CLI makes runs this one path: `down`, Ctrl-C on a foreground `cotal up`, the teardown after\nthat `up`'s broker exits, the leftover-manager stop before `cotal up -f`, and the delivery cutover.\nEach holds the manager's stop reservation, so a second stop while one is in flight is refused,\nnames the process holding it, and leaves that stop's `--with-agents` policy in place. Each sends\n`SIGKILL` to a manager still running 15s after `SIGTERM`. Ctrl-C stops the manager first; when that\nstop is refused or the manager's exit cannot be confirmed, Ctrl-C signals nothing else, prints the\nrefusal with the reap route, and leaves the stack running; end it with `cotal down --with-agents`.\n`--with-agents` is a one-shot destructive policy bound to the exact verified manager process\nand the exact live `down` stop reservation; a stale, malformed, crashed, or different stop attempt\ncannot turn a later bare shutdown destructive. If a managed agent cannot be proven stopped within\nthe manager's stop timeout, the manager logs which one, still closes its broker connections and\nconsole listener, and exits with code 1. It does not release its pidfile, liveness lease or\nservice registration in that case, so no successor is handed authority while that agent may still\nrun; the lease lapses on its TTL. Positional component names stop\nonly those self-registered local processes; for example, `cotal down manager` leaves delivery and\nthe broker running, and `cotal down web` is available when the web extension is installed. A\ncomponent that starts target-resolved (the web dashboard) is stopped the same way: `cotal down web`\nresolves the mesh the same way as `cotal web` (registry current mesh first, `--space` to name one), so\nit works from any directory; the other components always stop under the folder you run it in. The\n`-f` / `--run` forms tear down a [manifest deploy](#manifest-deploys) without stopping the whole mesh\nand cannot be combined with component names. Stopping `nats` alone is refused while an unselected\nregistered daemon is still live; include those components or use bare `cotal down`.\n\nA pinned manager with no spare-capability record is not signalled by bare `cotal down` or `cotal\ndown manager`. A current manager always publishes the record, so a missing one means an older\nmanager: one that predates capability reporting, or one whose pty runtime reported that it cannot\ndetach its agents. Stop each managed agent explicitly, then run `cotal down --with-agents` from the\nmesh root to stop the whole stack. An older manager does not understand\nthe reap request, which is why the agents must already be stopped.\n\nBare `cotal down` inventories by pidfile. When this folder's registered broker answers and no\n`nats.pid` records it, the command does not say nothing is running. It names the space and the\nbroker address, says no pidfile records that process, says it will not stop a process it did not\nstart, and exits 1. Stop that broker with whatever started it (an init unit, a container, or the\nhand-run process). `cotal meshes rm <space>` only drops the registration. The probe runs whether or\nnot other owned components were running: they stop and clear their artifacts first, then the broker\nis named. A component stop and `--dry-run` stay pidfile-only and do not probe.\n\n`down` reads each process record once. A component that exits and removes its own record while\n`down` runs counts as having no record, so the stop goes on. Any other failed read is an error.\n\n**Teardown verifies pinned process identity before signalling.** PIDs are recycled by every OS,\nso a recorded pid alone is not a durable target identity. `up` and `cotal web` record each\nprocess's creation identity in a sibling `<pidfile>.identity` pin, which holds the pid and the\nprocess start reported by the OS. Every stop path, including `down` for the broker, web and\nextension components, and the manager, delivery and auth-service stops, applies the same rule. A pin\nthat names a different start means the pid was reused, so teardown refuses and preserves it. A torn\nor unreadable pin also refuses.\n\nThe pidfile and its pin are published by renames, and the pidfile rename is the commit point. Just\nbefore it, the pin holds two lines: the old process's and the new one's. A launcher that dies\nmid-publish therefore leaves the old record or the new one, each checked against its own pin line,\nnever a pidfile without its pin. An old record with no pin is legacy, so its line holds `-` in place\nof the token and it stays legacy until the commit. A CLI older than this change reads a two-line pin\nas torn and refuses.\n\nPublishes of one pidfile are serialized by a lock file beside it, `<pidfile>.publish.lock`, because\nthe launcher and the daemon it starts both publish the same record. The next publisher reclaims a\nlock left by a crashed one. When no start token can be read for the new process, its pin line holds\n`-` in place of the token, which reads as a legacy record, and the publish ends in the legacy shape:\na pidfile with no pin. Teardown, and a daemon removing its own record on exit, take the same lock and\nremove the record only while the pidfile still names the pid they stopped, so a stop that races a\npublish leaves the new record whole.\n\nThe web dashboard claims `web.pid` with an exclusive create, so a second dashboard for the same mesh\nis refused, and writes its pin right after the claim. A stop that runs between the two reads a\nlegacy record.\n\nThe pidfile pid and the pin pid are two coordinates. Automatic cleanup follows **proven death of\nthe pidfile target** (ESRCH on that pid): a torn sibling pin does not wedge a dead pidfile pid.\nA torn pairing where the pin names another pid, while the pidfile pid is still live or not proven\ndead, still refuses. Inspect both pids with `ps`. Do not delete `<pidfile>.identity` to force a\nstop; that weakens target-identity protection. Once the pidfile process is dead, rerunning\nteardown clears the stale record automatically.\n\nThe first teardown after upgrading a running pre-pin stack has a narrower guarantee. A live record\nwith no identity pin is signalled after a loud warning that it predates identity pinning. Restarting\nthe component writes the pin, so later teardowns receive full match and mismatch protection. The\nsame warning applies on platforms where no stable start token is available. For a legacy manager,\nbare `cotal down` also warns that agent sparing cannot be verified before it signals. Because the\nCLI cannot establish which SIGTERM handler that already-running binary carries, it never presents\nthe pre-signal seat inventory as confirmed spared; a genuinely older destructive handler may still\nreap those agents. `--with-agents` publishes a one-shot reduced-guarantee handoff bound to the\nrecorded manager pid and the live `.stopping` reservation's inode, then signals unconditionally.\nThat handoff cannot be replayed by a later stop attempt. A pin that exists and does not match the\nlive process still refuses before signal.\n\n`--with-agents` performs the old destructive logical teardown: managed processes stop and their\ncredentials, ACL rows, and delivery footprints are deprovisioned. `--preserve-state` is a different\nmaintenance transition: it stops retained processes while suppressing leave/deprovision cleanup, persists the manager's\nsame-principal resume inventory, stops the entire stack without removing run/auth artifacts, and\npublishes a stable inode-bound cut only after every recorded process is proven stopped and the exact\nrecorded NATS endpoint is unreachable. A missing or stale broker pidfile never counts as stopped. The\nattempt is bound durably before the manager is fenced, the resume document and attempt-bound\n`cut-intent` are fsynced before manager commit, and the manager's commitment itself is journaled\n(`cut-committed`) before any process stops. A retry after a crash at any of those boundaries reuses\nthe exact recorded attempt and finishes the remaining stop and endpoint proofs idempotently, without\nneeding the (by then intentionally dead) manager. A partial cut never publishes `ready`. It cannot\nbe combined with component names, manifest teardown, or `--dry-run`.\n\n**Seat checkpoints.** After the stack is proven down, the cut writes one checkpoint per retained\nseat under `.cotal/maintenance/v1/checkpoints/<attempt>/<seat>/`, and prints the path, the\ncontinuity class and the generation for each. The path carries the preservation attempt because a\ncheckpoint is immutable once sealed: a shared directory would make the second cut in a root refuse\non the first cut's leftovers, and clearing it would destroy an artifact a rollback still needs. The\ncapture happens only at that point because anything earlier races a harness that is still writing\nits transcript and its working tree.\n\nEach checkpoint directory is created 0700, refuses a destination that already exists, and holds:\n\n- `repo.bundle`, the seat `cwd`'s reachable history, anchored on the base commit the record names\n by full object id;\n- `repo.index.diff` and `repo.worktree.diff`, the staging state as two diffs, base to index and\n index to worktree. Two rather than one because a single combined diff restores a mixed tree with\n the right bytes and the wrong index: a source reporting `MM README` would come back as ` M README`;\n- `repo.untracked.tar`, the untracked files in scope;\n- the harness session pointer, when the seat's connector declares one, and the transcript store\n files the operator named with `--session-store`. Each records where the destination puts it back\n as an anchor (the workspace root, the account home, or the seat's `cwd`) plus a relative path,\n because the destination's root and home are its own and the source host's absolute spelling would\n either miss them or write outside them;\n- `checkpoint.json`, written last, after every digest is computed over the bytes that landed.\n\nThe record carries the manager's resume entry unchanged as its first field, then the space, the seat\nname, the recovered `lifecycleUid`, the writer generation the cut was taken at, `capturedAt`, the\nrecency horizon, the applied profile revision, the seat's `git status --porcelain` as the cut read\nit, and the continuity class. Every captured file is\nlisted with its byte size and sha256, so an operator verifies the whole artifact with `sha256sum`\nand `git bundle verify`. No secret values, no operator keys and no source-host launch material\nenter it.\n\nThe continuity class is what the connector declares, capped by what the checkpoint carries. A\nconnector declaring session continuation classifies as `exact`, but reopening a session takes both\nhalves, the pointer that names it and the store that holds its transcript. A checkpoint missing\neither one cannot reopen that session, so it is recorded as `fresh` when the connector declares a\nfresh start and `drain-only` otherwise. A pointer with no store is capped the same way as a cut\ncarrying neither, because it names a session whose bytes the artifact does not contain. A class is a promise the destination is entitled\nto act on, so it never describes bytes the artifact does not contain. The transcript store stays an\noperator input: this repository does not know where a harness keeps its transcript, so `exact`\nrequires `--session-store` to name one.\n\nThe recorded status is read under the same selection rule as the untracked set, so it describes the\nstate the captured bytes can reproduce. The destination re-reads it in the promoted tree and refuses\na difference.\n\nThe untracked selection rule is recorded in the record and is\n`git ls-files --others --exclude-standard -z, excluding .cotal/`. It honors `.gitignore`, so an\nignored file the seat needs does not travel and has to be moved separately. The `.cotal/` exclusion\nis a secrecy boundary rather than a size one: when a seat's `cwd` is also the mesh root, the control\ndirectory is untracked, and without the exclusion the broker trust material, the space account, the\nmanager instance identity's private seed and the seat's own credentials would land inside the\nartifact. A checkpoint carries credential references only; the destination resolves that material\nitself.\n\nA seat whose launch options could not be resolved is refused rather than checkpointed, with the\nmanager's own wording: `imperative launch options have no non-secret durable source (<keys>)`. The\nrefusal arrives at prepare time, so the cut stops before any child does.\n\nA delegated seat (SPEC \xA713.17) is refused at prepare time too, with\n`a delegated seat is not resumed by a later manager; stop it before preserving`. A manager stop\nafter a refused cut retires that seat through its retirement path.\n\n## clean\n\n```bash\ncotal clean <history|store|all> --force\ncotal clean restore-attempt --attempt <id> --force\ncotal clean restore-fallback --attempt <id> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | `history`: target mesh |\n| `--dms` | off | `history`: also clear DM history |\n| `--store-dir <dir>` | `.cotal/nats` | `store`/`all`: JetStream store directory |\n| `--force` | none | Required: destructive, no prompting |\n| `--attempt <id>` | none | `restore-attempt`: exact stale pre-commit attempt; `restore-fallback`: matching healthy committed restore |\n\nOne configurable cleanup verb; every target requires `--force`.\n\n- `history` purges the retained message backlog on the **running** broker (channels, plus DMs\n with `--dms`). The same operation as [`history clear`](#history), which stays as an alias.\n- `store` deletes the **stopped** mesh's JetStream store (`.cotal/nats`): streams, durable\n consumers, and messages. This is the reset for stale on-disk broker state, e.g. durables\n minted by an older, incompatible Cotal generation surviving a `down`/`up` cycle.\n- `all` is `store` plus the space identity (`.cotal/auth`), the local creds and markers tied to\n it, any crash residue a normal `down` would have swept (stale pidfiles, `run/`), and the mesh's\n registry entry; the next `cotal up` mints a fresh identity.\n\n`history` needs the mesh up; `store` and `all` refuse while any recorded mesh process is still\nalive or any same-root recorded broker endpoint remains reachable (run `cotal down` first). They\nalso refuse outright on a root that holds accounts for several spaces: the store and the broker\ntrust record are shared by every space on the broker, so both targets would take out all of them\nand no `--space` can narrow that. `down`, `backup` and `up --restore` refuse there for the same\nreason. `cotal status` lists the tenants on such a root. Personas\n(`.cotal/agents`) and logs are never touched. The mesh record now carries a custom\nstore location for `up`'s own repair, but `clean` still takes `--store-dir` itself; `clean` does\nnot read the record. Custom cleanup targets must contain either the Cotal store-generation marker or a\nreal `jetstream/` store directory; filesystem roots, project roots, and Cotal auth/maintenance trees\nare always refused.\n\n`store` and `all` also refuse every maintenance journal state. After a healthy committed restore,\n`restore-fallback` is the only supported way to remove the recorded unchanged old-store inode; it\nnever deletes the active target, requires both the exact attempt id and `--force`, and retires the\ncompleted restore journal so a later `down --preserve-state` can start a new backup cycle.\n\n## Backups\n\n```bash\ncotal down --preserve-state [--store-dir <dir>]\ncotal backup create <dir> [--only full|registry] [--store-dir <dir>]\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\n```\n\nBackup is offline-only. It requires the stable `ready` record from `down --preserve-state`, an exact\nstore match, no live recorded process, and an unreachable exact endpoint from the recorded cut.\nThat endpoint is probed immediately before cloning, so a live broker with a missing or stale pidfile\nis still refused. It claims the cut, reflink/copies the stopped source to a\nprivate attempt clone, and opens only that clone on a random loopback bootstrap broker with an\nindependent parent/deadline watchdog. It validates the canonical stream and pull-consumer inventory,\nwrites native snapshots with consumers excluded, and stores conservative contiguous ACK-floor\ncheckpoints separately. The presence bucket is memory-backed, so it does not survive the cut and\nthe clone may lack it. Every other stream must be present. The original store is never opened by\nthe backup broker, and the stack is not restarted implicitly. Artifact destinations must not overlap\nthe preserved source or maintenance\nattempt tree. Restore artifacts and targets likewise cannot nest inside or contain each other, the\npreserved source, or the maintenance attempt tree.\n\nStopped client-managed KV ordered consumers are ephemeral read residue, not backup state. Backup\nignores only the pinned client's exact stopped shapes: ordinary last-value watchers and the\nwhole-bucket scanner that uses all-history delivery to collapse concurrent tombstones. A bound\nconsumer or any lookalike with a different filter, inbox, lifetime, or other config is still refused.\n\n`full` is the default and indivisible: channel registry, CHAT/DM/TASK/INBOX/DLV, ACL, MEMBERS, and\nvalidated durable checkpoints. `registry` is the sole partial artifact. Presence, derived membership\nfeed, leases, native ephemeral/history consumers, credentials, keys, tokens, owner secrets, and actor\nledger files are excluded. `full` means every transferable message and registry stream, not every\nJetStream resource: endpoint submissions/facts/events/timers/workflow state, contract artifacts, and\nthe records/auth/session stores are nonportable control state. Restore recreates those streams empty\nwith their canonical configs before exposing the normal listener, so active endpoint runs,\nlifecycles, and sessions do not cross a backup. Artifacts are exclusively created `0700`;\nsnapshot/checkpoint files and\nthe manifest are `0600`; `manifest.json` is written last with exact sizes and SHA-256 values. The\ndirectory is trusted operator input: hashes detect corruption, not malicious rewriting.\n\nRestore validates and stages the exact allowlisted artifact bytes before moving or creating a store.\nIt requires the same space and existing trust state. The whole pre-commit window holds a journaled\nliveness claim (coordinator, watchdogs, brokers, absolute deadline): ordinary `up` and a repeated\n`up --restore` refuse while the claim is live, and a stale attempt is recovered only after the\ndeadline has elapsed and every recorded owner is proven dead. A retried `up --restore` handles this\nautomatically; an operator can also recover it explicitly with `cotal clean restore-attempt --attempt <id> --force`. Nothing\never rolls back a live attempt. A registry-only artifact restores as registry-only whether or not\n`--restore-only registry` is passed; omitted infrastructure is always created and the exact\npost-restore stream inventory is asserted before commit intent. Ordinary `up` from a preserved cut\nresumes only the exact recorded source store and runtime; a contradicting `--store-dir` or\n`--runtime` fails in preflight.\n\n**Admitting a seat checkpoint.** An ordinary `up` from a preserved cut admits that cut's seat\ncheckpoints before it journals the resume attempt and before any process starts, so a refusal costs\nnothing. Three gates run in order, each naming what it saw.\n\n1. *Integrity.* Every file the record names must be present, a regular non-symlink file, the\n recorded byte size and the recorded sha256, re-stat'd after the read so a file that moved is a\n refusal. Failure here consults no other gate.\n2. *Identity.* The recorded space must match, the recorded `lifecycleUid` must not belong to a live\n incarnation, and the profile revision must match this host's or be resumed under deliberately\n this host's. A differing revision is refused with both digests and the remedy, and there is no\n override: the checkpoint carries the recorded digest and not the config bytes, so nothing could\n run the seat under the recorded revision, and the manager re-digests the same file and refuses\n drift on its own. This gate has no blanket override, which is the only reason the next one may\n have one.\n3. *Recency.* `capturedAt` is compared to this host's clock against the horizon the record carries.\n Inside it, the seat resumes. Outside it, `up` refuses and prints the capture instant, the clock\n reading and the horizon; `--accept-stale-checkpoint` admits it anyway and the exercised consent\n is printed with the actual age. An unreadable `capturedAt` is refused with no override, because a\n freshness gate that fails open is not a gate.\n\nCustody transfers only after all three pass. The destination claims the recorded generation plus one\nby exclusive create, before it launches anything. A lost create means another destination is already\nclaiming that seat, and it refuses with `seat-writer-generation-create-lost` rather than adopting\nthe winner and becoming a second writer. The recorded `lifecycleUid` is reused and never minted, so\nthe resumed seat binds the same lifecycle-keyed durables.\n\nAdmission is reconciled against the inventory the resume is about to hand the manager, and that\nreconciliation finishes before the restore moves a single tree. A checkpoint whose recorded\n`lifecycleUid` is not the one the retained inventory carries describes a different incarnation of\nthat seat, and it refuses with both uids while every live working tree is still untouched and no\ngeneration is claimed. A retained\nseat with no admitted checkpoint refuses the resume by name: an absent checkpoint directory and an\nabsent record are indistinguishable from a seat that was never checkpointed, and a seat that starts\nwithout passing the gates has claimed no generation. `--accept-stale-checkpoint` is recorded in the\nresume journal with the seat, the capture instant, the admitted age and the horizon, so the consent\nsurvives the terminal it was typed into.\n\nThe whole admission is all or nothing. Coverage is settled first, then every gate runs over every\ncheckpoint, and only then is any generation claimed. A refusal at any point leaves every generation\nunclaimed, including a lost exclusive create during the claim itself: the claims that attempt made\nare removed before the refusal is raised, by the exact paths it wrote, so a generation another\ndestination holds is never touched. A claim is a create that can never be made again, so a refusal\nthat left one behind would consume the retry over the same checkpoint set.\n\n**Restoring a seat checkpoint.** Once every gate has passed over every checkpoint, and before a\nsingle generation is claimed, `up` puts each admitted seat's captured bytes back. A refusal here\ncosts nothing for the same reason a gate failure does: no claim has been made and nothing has\nstarted.\n\nA restore never moves or replaces the destination's own control directory. A checkpoint excludes\n`.cotal/` by design, so a seat whose `cwd` holds one, which is the layout an operator gets by\nrunning `up` and `spawn` in a single directory, is refused before anything is staged: promoting a\ntree that cannot contain `.cotal/` over that `cwd` would carry this host's live trust material and\nmaintenance state away with the superseded tree. The refusal names the control directory it found\nand the remedy, which is to give the seat a working tree that is not a workspace root.\n\nEach seat is staged beside its own `cwd`, in `<cwd>.incoming`:\n\n1. every recorded digest is verified again over the files as they are now;\n2. the bundle is cloned into `<cwd>.incoming`, which is refused when that path already exists;\n3. the recorded base commit is verified in the clone and checked out detached, so a bundle that does\n not contain it stops the resume instead of continuing against a different history;\n4. the index diff is applied with `--index` and the worktree diff without it, both `--binary\n --allow-empty`. That order is what puts staged content back in the index rather than only in the\n worktree, and `--allow-empty` is why a seat with a clean tree is still restorable;\n5. the untracked archive is extracted.\n\nEvery seat stages before any seat is promoted. Promotion moves an existing `cwd` aside to\n`<cwd>.superseded.<timestamp>` and renames the staging directory into place, then puts the session\npointer and store files where the destination's connector reads them, then re-reads\n`git status --porcelain` in the promoted tree and compares it to the status the checkpoint recorded.\nA restore that applied without error and produced a different index is a refusal, not a warning. The\ntwo renames are the only steps that touch the path the seat will use, so a failure anywhere leaves\nevery seat's live `cwd` as it was.\n\nThe rename itself claims the superseded name, and a taken name gets a numeric suffix. The timestamp\nhas one-second resolution, so two promotions of the same seat within one second compute the same\npath; a rename onto a name that already holds a tree fails on every platform, and that failure is\nread as taken. Nothing creates the name ahead of the move, because Windows refuses to rename onto an\nexisting directory at all. A superseded tree is the thing that rename exists to keep.\n\n`git` and `tar` run as child processes with argument arrays, never a shell string.\n\nA leftover `<cwd>.incoming` refuses the resume by name. A staging directory from a failed run is the\nonly record of what failed, so nothing removes one automatically: inspect it, remove it by hand, and\nresume. A pre-existing `cwd` is renamed rather than deleted, so a wrong checkpoint costs a rename\ninstead of a tree. When a promotion fails, the renames that attempt made are undone and the staging\ntree is left where it is, as the evidence for what did not verify.\n\nA session pointer whose recorded `sessionId` is not the one the retained inventory reopens is\nrefused before anything is cloned. A session file already present at its destination is judged by\ncontent: bytes equal to the recorded digest are already restored, and different bytes under the path\nthe connector is about to read are refused with both digests rather than clobbered.\n\n`up --restore <dir>` reaches the same admission and the same restore, after the store is restored\nand validated and before commit intent is journaled. A registry-only restore resumes no seat, so it\nadmits and restores nothing.\n\nOne limit is worth stating plainly. The writer generation is claimed by exclusive create inside one\nworkspace root, so it fences two resumes on the same host and does not fence two independent\ndestinations: copy a checkpoint to two roots and both claim the same successor. A real cross-host\nfence needs a coordinate neither root owns.\n\nAuthenticated restores validate the complete\nspace trust bundle before staging, including nkeys, seed matches, JWTs, signers, and space binding;\nfull restores commit to the validated operator, system-account, data-account, and active-signer root\nchain in addition to the static/user authority fingerprint. Because the system account is part of that\ncommitment, a [`cotal up --rotate-sys`](#up) makes every full artifact taken before it unrestorable\nagainst this root: take a fresh full backup after each rotation. The composed commitment is revalidated\nimmediately before store mutation and never includes secret seeds. Restore never creates fresh auth.\nSame-path restores atomically retain the old\nsource at the journaled fallback path; alternate targets retain it in place; a missing canonical\nsource needs explicit `--accept-missing-source`. Quarantine and target restores use current canonical\nconfigs on isolated random-loopback brokers, never expose native snapshot consumers, and publish a\ncommit-intent immediately before the normal listener starts. Archive bytes never instantiate the real\ntarget: after quarantine validation, every stream is re-snapshotted from the validated quarantine\nstate into attempt-owned sanitized files, and the target is restored solely from those. Before that boundary, failure rolls back\nthe attempt-owned target; after it, ambiguity preserves both stores and records forward-repair\nrecourse. The cooperative maintenance lock excludes Cotal commands, not arbitrary raw NATS processes.\n\nBootstrap brokers in every auth mode, including open, mount the store under a local account with\nrandom operation-specific logins only, each carrying the exact per-phase subject permission matrix;\nnormal static credentials and user-auth sentinel/bearer connections are rejected, and no auth\nservice or callout starts. Open mode differs only in its account label, never in authority. Inventory, each stream snapshot,\nrestore initiation, exact upload id, validation, and each checkpoint recreation use separate exact\nauthorities. Every checkpoint carries the source stream's message/first/last sequence state and must\nmatch its snapshot record before mutation; core then derives and validates the only allowed start\npolicy. TASK is not a CLI exception: the same core checkpoint API recreates its canonical `DeliverAll`\nWorkQueue durable because acknowledged tasks are absent from retention and NATS forbids a\nstart-sequence policy there. Registry-only restore creates every omitted canonical stream and transient\nbucket on the isolated target before the normal listener is exposed. It deliberately does not resume\nretained agents or recreate their DM/DLV/TASK/ACL state; their identity material stays retained and\nstopped rather than being reprovisioned into a partial restore.\n\nAfter listener readiness, the manager starts attempt-bound, validates retained credentials/tokens\nwithout granting or reprovisioning, and resumes the exact persisted principals under cleanup\nsuppression. Registry-only restore uses the same flow with an empty agent set. On a user-auth mesh\nthese manager calls run as the logged-in operator's `cli` actor, the caller the preserve cut used, so\nthat actor needs a current `admin` grant. `commitResume` is an\nidempotent validation barrier only: success must be `awaitingFinalize` with an attempt-bound 64-hex\ncommit token and does not release suppression. Under the workspace lock, the CLI first fsyncs that\nexact evidence as `manager-committed` (restore) or `resume-committed` (ordinary resume), then calls\ntoken-bound `finalizeResume`; only an `active` response for the exact token releases suppression. The\nCLI records the same token in finalization evidence before a restore becomes `active`, or before an\nordinary resume retires and consumes the marker. Re-entry from either committed state skips the prior\nidempotent activation/commit phases, retries finalization with the durable token, and finishes the\nworkspace transition. Failure before finalization preserves the committed state and cleanup\nsuppression; it is not rewritten through a degraded transition. Re-entry between any two earlier\nboundaries reuses the same attempt and may retry the idempotent phases without deleting retained state. A missing or\nchanged per-agent dependency is a named fail-closed result; the journal becomes degraded and remains\navailable for forward repair. A retry from `resume-intent`,\n`resume-active`, or `resume-degraded` reuses the same attempt and inventory after the prior listener is\nproven stopped. A retained agent the lost manager already launched can still be running, for example\nin a tmux window, while the journal reads `resume-intent`. On a static mesh the replacement manager\ncloses that seat through the reference the lost manager recorded on the agent's slot, waits for the\nprincipal to leave presence, and launches it again. A live principal with no such record, or one that\nstays live after the seat is closed, is refused. Every normal restore listener has an unguessable\nattempt-bound NATS server name. The CLI fsyncs its exact name/nonce, canonical endpoint, process owner, and generation-bound target identity\nimmediately after spawn. Re-entry accepts a surviving listener only when its INFO server name, live PID\nrecord, endpoint, and target identity all match that proof; degraded restore repair then moves through\nthe guarded workspace transition only after manager commit. If an uncommitted bound owner is provably\ndead, recovery retires that exact proof under the maintenance lock and binds a fresh listener for the\nsame attempt, endpoint, and target with a new nonce and server name. A live foreign/mismatched listener\nor ambiguous owner is preserved and refused, never adopted by reachability alone. A reconstructed\ncommit/degraded attempt without either the exact bound proof or a durable dead-listener replacement\nrecord fails closed even when the recorded port is free. A later ordinary startup may pass an `active`\nrestore only when its details prove manager commit and its exact recorded listener is dead.\n\n## Mesh registry\n\n```bash\ncotal meshes [--json]\ncotal meshes add # guided, on a terminal\ncotal meshes add <space> --server <url> [--root <dir>] [--mode auth|open|user] [--tls] [--force]\ncotal meshes add <space> --mode user (--user-auth-file <bundle.json> | --from <https url>)\ncotal meshes rm <space> [<space> \u2026] [--force]\ncotal sync [--idp <auth base URL>]\ncotal use <space>\ncotal status [--space <s>] [--server <url>] [--components]\n```\n\n`meshes` lists the meshes this machine knows; a `*` marks the `current` default a bare\n`cotal spawn` joins. Entries learned from a signed-in account are marked `discovered`. Their\nregistration trust is stored under the account's private auth state, and the registry contains no\nsession token or sentinel credential bytes. Commands resolve the catalog `slug`; a different human\n`name` is rendered only as a label.\n\n`meshes --json` prints one JSON object per recorded mesh per line: `space`, `server`, `mode`,\n`root`, `default` (the `*`), and `origin` (`up`, `manual`, or `catalog` for a discovered entry). A\nlocal or hand-registered entry also carries `offline`. A discovered entry is never probed, so it has\nno `offline` field. `tlsRequired`, `events: \"required\"` and a discovered entry's `catalogName` appear\nonly when the record has them. An empty registry prints nothing and exits 0. The note about a default\nthat matches no record goes to stderr, so stdout carries only rows, on a first run too. The table is\npresentation and is not a stable parsing target. `meshes add` and `meshes rm` refuse `--json`.\n\nA registry record this build cannot use is refused by name, never rendered and never skipped. One\nthat does not parse, or is missing a field every consumer reads (`server`, `mode`, `root`, `ts`,\n`space`), makes every registry command exit 1 with the file's path and what is wrong with it.\nRemove the file or restore the record; nothing repairs or invents a field for you.\n\nAn IdP may advertise a same-origin space catalog during login. Cotal reads the complete snapshot and\nadds every valid registration without a separate `meshes add`. A snapshot younger than five seconds\nis used without a request. After that, commands that resolve a mesh target conditionally refresh the\nsaved catalogs. An operation targeting a discovered space refreshes only that space's account and\nrefuses if that account fails. Operations targeting local or manually registered meshes refresh every\naccount, print one warning for each failure, and continue. `cotal status` refreshes every account,\nnever refuses on a refresh failure, and lists each account as `fresh`, `updated`, `not-modified`,\n`no-catalog`, or `failed` with its error. `cotal sync` bypasses freshness and reports added, changed,\nremoved, unchanged, and name collisions. `--idp` limits it to one signed-in account. It never connects\nto a broker.\n\nThe registry is updated under the same lock that guards the catalog cache, so a command never lists\na discovered space set that another command is still writing. The cache records a fetched snapshot\nas not yet applied before the first registry write and as applied after the last. If a command dies\nor is stopped in between, the next command applies that snapshot again before it can use it, with\nno request inside the freshness window.\n\nThe shared dispatcher applies this preparation to every command that declares both `--space` and\n`--server` as mesh-target flags, including commands registered by other packages and commands that\ndeclare their own equivalent flag objects. Daemon and startup commands that use those names only as\nconfiguration explicitly opt out. Registry-local `meshes add` and `meshes rm` never refresh a catalog.\nWhile the registry holds a record this build cannot use, the preparation neither refreshes nor\napplies a catalog, so the command's own checks run first. A snapshot left unapplied is applied by the\nnext preparation after the record is restored or removed. A command that resolves its target through\nthe registry still refuses the record by name.\n\nRun on a terminal with the space or `--server` missing, **`meshes add` is guided**: it asks for the\none thing that cannot be derived (the broker URL), probes it, and tells you what answered - open or\nrequiring credentials. It then offers the spaces your `--root` already holds credentials for, states\nthe mode as a fact about that broker rather than asking, and shows the exact record before writing\nanything. A broker that does not answer, or a space name already registered, becomes a choice rather\nthan an error. Anything you pass on the command line is taken as given and not asked again. Without\na terminal - a script, an agent, CI - nothing prompts and the flag form's errors stand\n(`COTAL_NO_PROMPT=1` forces that too).\n\n`cotal up` and `cotal down` maintain their own records. `meshes add` registers a mesh they cannot\nspeak for: one running on another machine, a shared broker, a hosted space. `--root` is the folder\nwhose `.cotal/auth` holds that mesh's credentials and whose `.cotal/agents` holds its personas.\nThe default is the project you run it in. The registry stores that path, never a secret. `--mode`\ndefaults to `auth` when the root holds the space's account record and to `open` otherwise. The\nbroker is probed before anything is recorded, so a wrong address, or credentials that mesh will\nnot accept, fails here instead of at the first `spawn`; `--force` records without verifying (and\nreplaces an existing record).\n\nA hostname or public address is registrable only when the connection will **require TLS**. Pass\n`--tls`, or use a `tls://` URL. The scheme is recorded as enforced intent, so every later dial\nthrough the record demands the handshake (and `meshes add tls://\u2026` against a plaintext broker is\nrefused at registration). Without required TLS the fence admits loopback and private-overlay\nliterals only. RFC1918 addresses are refused in both modes because a cafe LAN is private but does not belong to you.\n\nA **user-auth** mesh registers from supplied pinned trust, never guessed: `--user-auth-file`\ntakes the bundle exported where the mesh runs; `--from` asks before it dials the address at all,\nthen fetches the `/.well-known/cotal-mesh` discovery document under that address (HTTPS only; a URL\nthat already ends in that path is fetched as given), displays the pins, and asks again before\nadopting them. Neither fetch follows redirects: a 302 can move a pinned fetch\nonto plaintext or onto another host, so it is refused rather than followed, and the pinned\nexchange must itself be an `https://` URL, except for an exchange on this machine, where plain\n`http://` is accepted for a loopback *literal* (`127.0.0.1`, `::1`, any spelling of them) but not\nfor `localhost`, which is a name rather than an address. Registration verifies that the exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also verifies that the broker refuses a bare\nconnect; that auth-required refusal is the pass. The sentinel credentials land in a 0600 file under\nthe entry's root; the registry records only the path.\n\n`meshes rm` drops records. It never stops a mesh. For a mesh running on this machine `cotal down`\nis the right verb, and `rm` says so unless you pass `--force`. A hand-added record is removed by\n`meshes rm`, by an `add --force` replacement, or by a `cotal up` that actually starts the broker for that same space, server and root, which becomes that\nmesh and so takes the record over (a `cotal up` for that space anywhere else refuses instead).\nNothing that merely *infers* a record is stale from a dead broker touches it: an\nunreachable broker is listed `offline` and stays, whether `cotal up` or `cotal meshes add`\nwrote the record; a foreground `up` whose broker exits unexpectedly keeps its record the same way.\nA bare command does not treat that offline record as a running mesh;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the project\nthey tear down; a hand-added one they leave alone even when it shares a root, because nothing\non this machine could write it back.\n\nA discovered entry belongs to the normalized IdP origin and proved subject that supplied it. Local\nteardown, cleanup, and liveness pruning do not remove it. A manual or locally started entry with the\nsame name wins and remains untouched; that discovered name is reported as a collision. Logging out\nremoves only the discovered entries owned by that account.\n\n`cotal meshes` and `cotal status` print `events: required` for a registration carrying\n`policy: { events: \"required\" }`. On that space, foreground spawn, detached spawn, manager starts,\nand interactive `join` cannot opt out or join without an event plane. `--no-events` is refused with\nthe space named. A connector without an event plane is refused with both the space and connector\nnamed. A session whose own grant omits `events.<owner>.<actor>` is refused before joining and the\nmessage names a full-row `actor grant` repair. A running seat whose event plane stops for good on\nthat space stops too.\n\n`use <space>` sets that default; the selection applies from every directory,\nincluding inside another mesh's project. `status` is a read-only report: machine prerequisites\n(starting with the installed `cotal-ai` version), the installed extensions and their versions, this\nfolder's `.cotal/`, the recorded meshes, and a live snapshot of the selected mesh (roster, channels,\nmembership feed). Stale Claude skills and out-of-date `.agents` skills recommend `cotal setup --skills`,\nnot unscoped `cotal setup`. `status` takes `--space` / `--server` to pick the mesh to inspect; it starts\nnothing. The manager row asks the service endpoint once: a live process that does not answer is\n`not serving`, and a probe that could not be made leaves the row `running \xB7 service unchecked`.\nA process row whose PID record exists but cannot be read reads `pidfile unreadable` with the error,\nand the other rows still print. A live manager whose delivery-aware marker cannot be read keeps its row\nand names the failure as `delivery-aware marker unreadable` with the error. The `Web process` row\nprints the address the selected mesh's dashboard recorded in `web.session` once it was listening,\nwhile the PID in its `web.pid` is alive. Otherwise it reads `down`, or `not installed` without the\nweb extension.\n\nIf a refresh fails, `status` may still show the kept catalog bytes for diagnosis. It labels them\nstale with the last successful snapshot timestamp and the refresh error. It never calls that state\nsynchronized or online. If a selected discovered space vanishes from a successful snapshot, the\nselection is cleared and the command reports that no default is selected.\n\nFor a user-auth mesh the selected-mesh section reports the login `status` works as: the signed-in\nsubject when this machine holds a cached session for the entry's pinned IdP, or the exact `cotal\nlogin --idp <url>` line when it does not, with no network round trip either way. A locally\nprovisioned space also shows the actor grant row; a discovered or registered remote entry reports\nthe grant as not checkable on this machine, because the ledger runs where the space was\nprovisioned. `--components` on a user-mode target probes as that same signed-in login (`ps`'s\ncredential), never a static mint; when the login cannot supply a credential, the row says why\ninstead of printing the broker's refusal of an unauthenticated probe.\n\nPersona rows name the catalog they describe. If this folder and the selected mesh use different\ncatalogs, status names both and marks which one spawn launches from. A green `default` means the file\npasses the same agent-file loader spawn uses; a present but invalid file is reported as invalid.\n\n`cotal status --components` adds a fail-loud per-component health pass. It reads **each\ncomponent's own control surface**, rather than treating a PID, a lease, or a successful probe of a\nsibling as proof that the component serves. It prints one of `serving`, `absent`, `not-serving`, or\n`refused` for each component and exits `0`, `1`, `2`, or `3` respectively (the highest observed\nstate wins):\n\n- **manager**: local PID record, its liveness-lease holder and PID, then the manager's own typed\n `status` service reachability from this host. Manager builds that do not report static\n reconciliation say `static reconciliation not reported by this manager build`; the line stays\n visible even when the manager is otherwise `serving`.\n- **delivery**: local PID record, its ready lease (`ready` is the daemon's own bound-control\n signal), and the latest `renewal.<spaceKey>.json` adoption verdict, the record of the space the\n command was asked about, keyed per space the way the pidfiles are. A re-signed credential and a\n broker-accepted adoption stay distinct facts. A root-only `renewal.json` left by an older build\n names no space and is never read as any space's verdict (`doctor auth` names it as a leftover).\n- **web**: local PID record, then the `/api/meta` response at the address the dashboard recorded in\n `web.session` once it was listening, which must name the same PID. The probe presents the\n readiness nonce recorded beside that address, the one credential the dashboard accepts on\n `/api/meta`. A live PID with no readable recorded address (the dashboard is still writing it, or\n an earlier build started it), or an unrecognizable process record, is `refused`, not a green\n default-port guess.\n- **broker**: the registered mesh URL dialed from this host with its recorded TLS requirement.\n\n`absent` means Cotal has no live local component record (or has a stale record); `not-serving`\nmeans the component record is live but its service/readiness surface did not answer or is not ready.\nThose are intentionally separate exit cases. A failed or unreadable probe is `refused`, never an\nabsent component or a clean zero. A PID record that exists but cannot be read refuses only its own\nrow. A record that its component removes while the pass runs reads as `absent`.\n\n## spawn\n\n```bash\ncotal spawn [<persona>] [--detach] [--name <n>] [--agent <a>] [--model <m>] [--variant <v>] [--prompt <text>] [--cwd <dir>]\ncotal spawn -f <cotal.yaml> [--dry-run]\n```\n\nFor a foreground spawn onto a remote user-auth mesh, a launcher may supply a one-time enrollment\ninstead of a cached human login. Prefer a private file:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./seat.md --space main\n```\n\nThe file contains only the enrollment URL, ending with at most one line terminator, and must be\nmode `0600` on POSIX. An orchestrator that cannot mount a file may set `COTAL_ENROLLMENT_URL`\ninstead; that value is redeemed byte for byte, so a trailing newline in it is refused. Setting both\nis refused. Enrollment input\nrequires `--space` and applies only to a foreground persona spawn. If the mesh is not registered yet,\nthe enrollment response must carry the stock user-bundle fields and the command needs\n`--config <persona-file>` because there is no local remote-mesh persona catalog to read. The client\nredeems the URL once, registers the returned mesh material, exchanges the returned actor token at the\npinned auth service, and removes both enrollment variables before starting any child process.\n\nA cached login for the same IdP and an enrollment are conflicting proofs, so the command refuses\nrather than choosing one. An invalid enrollment never falls back to login provisioning. Unknown,\nexpired, revoked, and already-used enrollments all produce one response: ask the owner for a fresh\none. See [Enrollment redeem](identity-and-auth.md#enrollment-redeem) for the HTTP contract.\n\nA runtime that starts a managed seat outside the manager's filesystem hands the child a managed\nhandoff instead: one `0600` file named by `COTAL_MANAGED_HANDOFF_FILE`, carrying the lifecycle the\nmanager already enrolled. The runtime builds the command with `delegatedSeatCommand`:\n\n```bash\nCOTAL_MANAGED_HANDOFF_FILE=/run/seat/handoff.json \\\n cotal spawn --config ./seat.md --space main --name <actor> --agent claude \\\n --expect-owner <owner> --expect-lifecycle-uid <uid>\n```\n\nThe `cotal` entry reads the file, deletes it and drops the variable before it parses flags, prints\nhelp or loads extensions, so every outcome leaves no file. The variable is read under any letter\ncase; spellings that name different files are refused after every one of them was deleted. The\nspawn then refuses a malformed handoff, or one whose space, owner, actor or lifecycle UID differs\nfrom `--space`, `--expect-owner`, `--name` and `--expect-lifecycle-uid`, before any broker\nconnection or exchange request. Every refusal on this path names the field and never a value from\nthe handoff. The registration's server, exchange and enforcement checks, the local state this\nmachine keeps for the space (its mesh record, user-auth state and agent secret files), target\nresolution, the policy refresh, the broker preflight and the agent auth preflight quote the space,\nthe server, the exchange URL, the actor or a path named for one of them in their own diagnostics and\nin the filesystem errors under them. For a handoff each prints one fixed sentence that names the\nfield and the phase instead, whether its check fails or an error is thrown. When the agent auth\npreflight's rollback then fails to remove a secret or file, that sentence is followed by the names of\nthe cleanup steps that failed, without their errors. The event-plane policy\nrefusals name the handoff's space field. An actor outside `[A-Za-z0-9_]` and a space that cannot\nname local state, such as `..`, are refused as malformed before any plane. A handoff conflicts with\nthe enrollment variables, `--detach`, `-f` and `--creds`, and needs `--config <persona-file>`. From\nthere it runs the enrollment consumer above without redeeming anything. See\n[Delegated seats](embedding.md#delegated-seats-outside-the-managers-filesystem).\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | resolved mesh | Target space |\n| `--server <url>` | registry entry | Broker URL override |\n| `--creds <path>` | none | Control-caller creds for an off-registry manager (`--detach` only) |\n| `--name <n>` | persona's `name:` | Presence-name override (does not choose the persona) |\n| `--config <persona-or-path>` | none | Persona catalog name or file path; wins over the positional |\n| `--agent <a>` | persona's `agent:`, else `COTAL_DEFAULT_AGENT`, else `claude` | Connector type (`claude`, `opencode`, `jcode`, `hermes`, and so on) |\n| `--role <r>` | persona's `role:` | Role override |\n| `--model <m>` | persona's `model:` | Model override |\n| `--variant <v>` | persona's `variant:` | Model variant override (connector-defined; e.g. OpenCode reasoning tiers) |\n| `--cwd <dir>` | this cwd | Working directory to root the agent at. Refused before launch when the directory does not exist on the serving manager's host. |\n| `--prompt <text>` | none | Initial prompt auto-submitted at start |\n| `--resume <id>` | none | Fork an existing session id into the mesh; only connectors that declare resume support accept it (see [the matrix](connectors.md)). The manager records the source session id, and `ps --wide` shows it. With `--detach --on <instance>`, a Claude session held on this host is carried to that instance first ([Resume a session](connect-claude.md#resume-a-session)); carrying one needs `--on` |\n| `--no-events` | event plane on where supported | Opt out of the session's structured event plane (`--events` only restates the default) |\n| `--share-tools <sel>` | none | Share named operator MCP servers with the agent |\n| `--subscribe <a,b>` | persona's | Channel read-set override |\n| `--allow-subscribe <a,b>` | = subscribe | Read-ACL override |\n| `--allow-publish <a,b>` | deny | Post-ACL override |\n| `--detach`, `-d` | off | Launch via the manager into a detached PTY (reattach with `cotal attach`) |\n| `--on <instance>` | class anycast | With `--detach` only: pin the launch to one manager instance id (the whole id, as `ps` prints it). Refused on a foreground spawn (no manager to pin), with `-f` (a manifest deploy launches through the manager class queue), and when empty |\n| `--file <cotal.yaml>`, `-f` | none | Deploy a manifest onto the running mesh |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--allow-stale <a,b>` | none | With `-f`: waive named stale agents (apply-only) |\n| `--runtime <name>` | manifest's | With `-f`: override the manifest's runtime |\n| `--expect-owner <u_\u2026>` | none | With `COTAL_MANAGED_HANDOFF_FILE` only, and required there: the owner the handoff must carry |\n| `--expect-lifecycle-uid <uid>` | none | With `COTAL_MANAGED_HANDOFF_FILE` only, and required there: the lifecycle UID the handoff must carry |\n\nEach session uses its connector's **event plane** by default: a stream of structured events\ndescribing what the agent did, rather than the prose it wrote, on a channel of its own. The channel is named after\nthe agent's principal, `events.<owner>.<actor>`, never after its display name, because two live\nagents are allowed to share a display name and would then share a stream. The launch grants publish\nrights on that channel alone, foreground and detached alike. On an open mesh, which issues no\ncredentials, the launch still allocates the agent an id, so the channel names a stable actor.\n`--no-events` is the explicit opt-out unless the selected registration says\n`policy: { events: \"required\" }`. Required policy makes the\nevent arm and grant mandatory, so `--no-events` and connectors without an event plane are refused.\n\nThe launch decision and the grant are separate on purpose. Holding publish rights on a channel is\nnot a request to publish to it, so writing an event channel into an agent file's `allowPublish`\ndoes not override `--no-events`.\n\nThe persona (`--config` > positional > `COTAL_DEFAULT_PERSONA` > `default`) is loaded from the\ntarget mesh's `.cotal/agents/` when it is a bare name. A reference that contains a path separator or\nends in `.md` is loaded from that file. A relative path resolves against the mesh root, except that\nan enrollment or a managed handoff resolves `--config` against the working directory. A missing\npersona is refused with the catalog directory or the file that was checked. The launch flags\noverride the file. On a user-auth mesh the\neffective name is also the agent's actor token, so it must match the token grammar (no `-`); the\nspawn is refused with that explanation before any request is sent. Foreground runs the agent\nattached to your terminal; `--detach` hands the launch to the running manager. Both modes get the\ndurable backstop on a mesh that runs the delivery daemon; `--live-only` skips it for a foreground\nspawn (messages posted while it is disconnected are then not replayed). A foreground exit retires\nthe agent's creds and broker footprint, like a manager despawn. On a user-auth mesh the two arms\ndiffer: a spawn against a mesh this machine provisioned revokes the actor row on exit, while a\nremote spawn (an enrollment or the advertised provisioning endpoint) removes only this machine's\ncredential files; its grant stays until the mesh operator revokes it, and the launch line says\nwhich arm you are on. A spawn through the advertised provisioning endpoint against a record that\npins no exchange URL is refused before the grant is requested, so no credential lands on this\nmachine. A `--detach` spawn is an\n**action**: the manager accepts it and returns the allocated identity at once, then the launch\nfollows to a terminal outcome rather than blocking (see [the control surface](control-surface.md)).\nSee [Connect Claude Code](connect-claude.md) and [Agent files](agent-files.md); `-f` is a\n[manifest deploy](#manifest-deploys). (`cotal start` was merged into `cotal spawn --detach`.)\nA `--detach` spawn onto a manager from another Cotal release is refused before any request is sent\nwhen the manager's contract does not declare a field this CLI sends. The refusal names the field,\ncalls it version skew, and gives this CLI's version. A field you leave unset is not sent, so it\nnever causes that refusal.\n\nA manager has 50 seat slots, and each seat counts once. A slot is held by a managed seat (a row in\nthat manager's `cotal ps`, including a seat still joining), by a reserved launch the manager accepted\nbut has not started a process for, or by a cooling hold. A seat that ends within 10 seconds of\nstarting leaves its slot cooling until those 10 seconds pass, unless an operator stopped it. Such a\nseat holds only that cooling slot, even while its launch is still reporting the failure. A spawn\nrefused at the limit states that split and whether waiting can free a slot:\n\n```text\nat capacity (50 of 50 slots: 49 managed, 0 reserved, 1 cooling); waiting frees a cooling slot in 7s, or despawn one\n```\n\nA cooling slot frees at the stated time. A launch that has not settled frees its slot only if it\nfails, and a managed seat frees its slot only when it stops. The refusal counts a launch as pending\nonly while it holds a slot, so a launch whose seat already ended is not counted. The roster counts\npresence, which also includes peers no manager owns, so its total is a different number.\n\nRun from a managed seat's own shell on a static or open mesh, `cotal spawn --detach` launches as\nthat seat when it targets the seat's own space. Without `--space` it picks that target the way the\noperator path does, so a recorded mesh that is not running is skipped. The CLI reads the seat's\nlaunch identity (`COTAL_NAME`, `COTAL_ID`, `COTAL_LIFECYCLE_UID`, `COTAL_SPACE`, and on a static\nmesh the seat's own credential), so the manager records the seat as the spawner, the same as for\nthe seat's `cotal_spawn` tool. On a static mesh that credential also proves the seat's space, so a\nlaunch without `COTAL_SPACE` still runs as the seat, and a target space holding no credential for\nthe seat is refused. An open mesh acts as the seat only when `COTAL_SPACE` names its space. The\nseat can then stop the child with `cotal_despawn`, and the manager stops the child when the seat\nexits. On a static mesh a seat whose agent file lacks `capabilities: [spawn]` is refused, because\nits credential holds no spawn subject.\n`--on <instance>` keeps its pin: the seat's own credential has no instance route, so on a static\nmesh the CLI mints a one-shot `manager-caller` view for the seat, pinned to that instance and\ncarrying the spawn subject only when the seat's credential holds it. On an open mesh the call keeps\nthe TLS requirement the mesh records. `--creds`, `--server` with an unregistered `--space`, and a\nuser-auth mesh keep the operator path.\n\n## models\n\n```bash\ncotal models [--agent <connector>] [--refresh]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--agent <connector>` | all registered connectors | Connector whose catalog to list |\n| `--refresh` | off | Ask the connector to refresh its provider cache |\n\nAsks the running manager for each connector's model catalog (model ids plus their variants)\nfor connectors that expose one. OpenCode and Codex query harness/provider surfaces; Jcode reads\nproviders that enable `model_catalog = true` in the operator Jcode `config.toml`. Jcode's listed\neffort tiers render as `variants (declared, not provider-verified)`, and launch can still refuse one.\nA connector without a catalog says so. A connector whose harness the manager did not find at boot\nreports the reason boot recorded, as a spawn does, so restart the manager after installing it. Pick\na result with `cotal spawn --model <id> --variant <v>`, where `<id>` is the model id as the catalog\nprinted it. OpenCode and Codex ids are the full\n`provider/model`; Jcode ids are bare (`opus-5`, not `cliproxy/opus-5`), because the provider is\nselected by the operator's Jcode config and a prefixed id is refused at launch with the bare form\nnamed.\n\n## endpoints\n\n```bash\ncotal endpoints [--space <s>] [--server <url>] [--creds <path>]\n```\n\nLists the mesh presence roster: agents, the manager, and any other protocol endpoint, with each\nendpoint's role, kind, status, and current activity. Unlike `ps`, this is a read-only presence view;\nit is not limited to child processes owned by the manager.\n\n## Endpoint control\n\n```bash\ncotal describe <endpoint> [--on <instance>] [--space <s>]\ncotal invoke <endpoint> <command> [--args '<json>'] [--space <s>]\ncotal invoke <endpoint> <command> --name <agent> [--admin] [--space <s>]\n```\n\nThe generic v0.4 service surface. `describe` resolves a registered endpoint's command set off the\nwire - the reserved `describe` command answers the registered contract digests, the schemas are\nfetched from the space's content-addressed contract store, recompiled, and verified against those\ndigests - and prints each command with its capability class and targeting shape. `--on <instance>`\npins `describe` to one manager instance's rail (the whole id, as `ps` prints it under its\n`manager <id>` headers), so an operator can read what that instance serves in a multi-manager space;\nunpinned, the class queue answers and the attribution line names whichever instance did. `invoke`\ncalls one command by name: `--args` is a JSON object validated against the fetched input schema\n*before*\npublish; a targeted command takes `--name <agent>` (resolved to the agent's current principal through\n`inspect`) or `--self`. `--admin` uses the admin instrument credential, whose cross-agent reach rides\nthe operator-only `any` authorization mode. Neither command has compile-time knowledge of any\nendpoint's schemas - this is the same trust chain every built-in control command now uses. Needs an\nauth mesh: the manager registers its service on both static and per-user meshes (a signed-in user\nrides their bearer; each visible or invoked command still requires its existing grant, and cross-agent\nreach needs the `admin` scope). An open mesh has no service registry.\n\n## Managed seats\n\n```bash\ncotal ps [--on <instance>] [--wide | --json] [--slots] [--space <s>]\ncotal stop --name <n> [--on <instance>] [--space <s>]\ncotal attach --name <n> [--on <instance>] [--no-reconnect] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | none | Managed agent to stop / attach (required) |\n| `--on <instance>` | class anycast (`ps`: class scatter) | Pin to one manager instance id (multi-manager space); takes the whole id as `ps` prints it, not a prefix. An empty value (`--on \"\"`, an unset shell variable) is refused, never treated as absent. A roster principal id (`local.\u2026`) is refused with a message naming the instance id `ps` prints |\n| `--wide` (`ps`) | off | After each seat's compact row, print extra operational facts the manager records: the provider the connector reported serving the model, `cwd`, `pid`, spawner, lifecycle uid, the owning manager's instance id and host, and for a `--resume` seat the session it forked (`forked from <id>`, with the source title and transcript SHA-256 once a Hermes or Jcode seat has recorded its fork; a carried Claude session prints `forked from <host>:<id>` with its title, SHA-256 and `carried <time>`, the time its bytes reached the manager). Model and requested variant stay in the identity row rather than printing twice. A fact the manager did not record (for example a runtime with no real process, or a connector that reported no provider) prints nothing, never a placeholder |\n| `--json` (`ps`) | off | Machine-readable: one JSON object per seat per line, copied unchanged from the manager row. Instance headers and errors go to stderr, so stdout contains only rows. Mutually exclusive with `--wide` |\n| `--slots` (`ps`) | off | List the durable static slot rows this manager owns instead of live seats, through the `slots` command. Mutually exclusive with `--wide`. A row that is not in the live roster still prints, with `live=false`; a retired row never prints |\n| `--no-reconnect` (`attach`) | off | End the attach when its session ends, instead of re-establishing it. For scripts that want one run and one exit code |\n\nA raw `--creds` file is refused by `ps`, `stop`, `attach` and the other control commands, because\nthat route mints no endpoint-caller triple; the project folder, or `--space` against the registry\nentry, is the route that does.\n\nThe human `ps` row is presentation text and is not a stable parsing target. Scripts use `--json`,\nwhich is the machine-readable row contract.\n\n`--slots --wide` is refused: `--slots` lists durable static slot rows, `--wide` prints live seat facts, and the two answer different questions. Across a multi-manager scatter, `--slots` prints each manager's rows under its own instance header, the same way the plain `ps` scatter does.\n\nThese are operator clients over the running manager's control plane. The default row includes the\nconnector, model pin, optional requested variant, and runtime as operational descriptors for the\nmanaged row. They do not make a shared display name a unique protocol identity; use `--json` when\nunambiguous owner+actor attribution is required. An omitted variant means no override was requested;\nCotal does not invent an effective provider default it cannot observe. `ps` also prints two state\nfacts per managed agent, because they answer different questions: the process fact from the manager's\nown runtime handle (`running` with its uptime, or `exited` with how long it ran), and the mesh fact\nfrom the roster (`idle` / `working` / `waiting` / `mesh offline`, or `not in roster` when the seat has\nno presence row at all: a seat that has not joined yet, or one that never did). When the seat's\nconnector relays a harness-reported condition, the mesh fact carries its code and how long it has\nheld, so a seat whose turn died on a provider rate limit reads `waiting (rate_limit for 40m)` rather\nthan a bare `waiting`, and `--json` carries the whole `condition` object. When the connector reports\nthe seat's last work event (presence `activeAt`), the mesh fact ends with its age, such as\n`\xB7 active 3s ago`, and `--json` carries `activeAt`. A seat whose turn stopped advancing keeps\nheartbeating, so its presence row stays fresh and this age is what shows the stall. A seat can be\n`running` and `mesh offline` at once: the process is alive and its presence has lapsed. That row says\nhow long, as in `mesh offline for <age>` with an age such as `3.5h`, counted from the seat's last\npresence heartbeat, which `--json` carries as `offlineSince` (epoch ms). The age is read only from\nthe seat's own presence record, matched on its principal and lifecycle uid, so a same-named peer or\nan older lifecycle never dates it. The manager log names each managed seat that is offline on the\nmesh while its slot is held\n(`seat offline on the mesh: <name> - last heartbeat <time>; process <state>`), including one its\nwatch first sees offline after a reconnect, and each one that comes back\n(`seat back on the mesh: <name>`), so a watchdog that only checks process liveness has a line to\nact on. The manager does not reap or re-key such a seat. The mesh fact is only a verdict while the\nmanager's own presence watch is fresh: when that watch has been silent past the liveness window, or\nhas not replayed the bucket yet, every row prints `mesh unknown` with the reason instead (`--json`\ncarries it as `meshView: stale | unpopulated`), because `offline` and `not in roster` would then\ndescribe the manager's watch rather than the seat. The manager rebinds a watch that goes\nsilent under a live connection on its own, so `mesh unknown` normally clears within a liveness window.\nOn a user-auth mesh `ps` also renders each managed agent's last credential-refresh outcome, fail-closed.\n\n**Mode split (chosen up front, never try-scatter-then-degrade):**\n\n- **Static / open mesh.** Bare `ps` is a **class scatter**: it freezes the live manager class from\n the records registry, merges every registered instance's agents grouped and attributed per\n instance, and a non-answering instance is shown as `registered, no answer within the deadline`\n (never silently omitted). A refused list or a missing answer makes the census incomplete: rows\n from other instances remain visible, but `ps` prints an incomplete-census warning on stderr and\n exits non-zero, including with `--json`. Those rows are not a complete seat count. A contract\n mismatch prints one plain comparison of the requested and served input/output digest pairs and\n advises aligning manager versions. The no-answer label means only that the instance is registered\n and did not answer. It does not say the host is down, because a dead host never deregisters itself and a\n live one can be slow; if it is gone, deregister it.\n `--on <instance>` pins the read to one exact instance id instead. A wrong pin fails loud\n rather than falling through: a well-formed id that no live manager carries is reported as\n `manager instance <id> did not answer` (nothing else is asked), and a credential without that\n instance's rail is reported as refused by the broker, not as an unresponsive manager. A manager\n that answers with a refusal is shown with its own cause; \"no manager reachable\" is said only when\n nothing answered at all. If the scatter's own registry read fails (the freeze or the reconcile),\n `ps` says the manager registry could not be read rather than pronouncing on the managers, which\n may all be up.\n\n**The verdict is scoped to the endpoint rail the request rode.** An issued caller rides the\nversioned `ep.v1` rail, a separate subject space from the legacy `ep` rail, and an endpoint serves\nboth (SPEC 13.15). A manager older than the versioned rail serves `ep` alone, so it can be running,\nregistered and answering while an issued caller's request reaches nobody. Silence on `ep.v1` is\nreported as `no manager answered on the <rail> rail` with `ep.v1` as the rail, and names both causes\nit is consistent with: no manager running, or one older than the rail. The CLI cannot tell them\napart, because the service registry records no package version, so check whether a manager is\nrunning and, if it is, its version. The same scoping applies to `cotal run`'s hosted verbs, which\ndrop the `--local` suggestion there, since `--local` drives the run from the calling process and\nnames the caller as its answerer.\n\n**`stop` and `attach` route by seat locality.** A seat can only be stopped or attached by the\nmanager actually running it, and the class queue does not know which one that is. So on a\nstatic/open mesh both verbs first ask every registered instance which one hosts the named seat, then\naddress that instance directly. This happens by default; you do not need `--on`.\n\n`--on <instance>` remains the override, for when you already know where the seat lives or the\nlookup itself is degraded. On a **user-auth mesh**, the exchange selects one authorized manager\nfor a short-lived `manager-caller` view. `--on` requests a specific instance; without it, selection\nmust be unique. Discovery and the command use that instance route. The caller gains no registry\nread or scatter permission. An absent, ambiguous or unauthorized selection refuses before sending\nthe command.\n\nA seat is reported as **not found** only when every reachable instance answered for itself. An\ninstance that stayed silent past the deadline, or that refused the read rather than answering, said\nnothing about which seats it hosts, so the seat may be running on it. That case reports that the\nlocation could not be established, names the instances that did not answer, and states outright\nthat it is not a report that the seat is gone. Read it as unknown and retry with\n`--on <instance>`; a retry loop that treats it as \"already gone\" stops looking for a seat that is\nstill running. A single manager cannot tell \"hosted elsewhere\" from \"does not exist\": it answers\n`not-found` for both, which is why the search asks all of them and why an incomplete search\nconcludes nothing.\n- **User-auth mesh.** `cotal ps` reports what **one** authorized manager knows about your agents\n (an instance-addressed read against its in-memory roster, owner-filtered). It does **not** report\n other manager instances or establish whether they are reachable. Completeness across a\n multi-manager user-auth space is not claimed.\n A manager that does not answer fails the command outright (exit non-zero), rather than printing\n an empty list that could be read as \"no agents\". Your ledger row needs the `admin` scope to\n reach `ps` at all; `spawn` alone is refused by the broker (the ep tier boundary).\n\n`attach` streams and drives an agent's terminal on the `pty` runtime; detach with the escape key\n(Ctrl-] by default; see [`COTAL_DETACH_KEY`](config.md)). The key is recognised as the legacy\ncontrol byte and as the kitty keyboard protocol and xterm modifyOtherKeys encodings of the same\npress, so a terminal with either protocol enabled detaches too. It does so over a one-use, holder-bound\nmesh session ([SPEC](../SPEC.md) \xA713.6): the manager replies with a signed session grant (never a\n`127.0.0.1` URL), the CLI redeems it once over the broker, and the browser console (`cotal console`)\ndrives the same session. `stop` and `attach` need a running manager to talk to. On a static mesh\nthey are cross-agent admin operations. On a user-auth mesh, your own agents (any agent under your\nowner) need only the `spawn` scope; another owner's agent needs `admin` on your ledger row\n([identity & auth](identity-and-auth.md)). Launch detached agents with [`spawn --detach`](#spawn).\n\n**`attach` reconnects when the link dies.** A session lives on a network link, and a laptop that\nsleeps, a VPN that drops or a wifi handover kills it. When that happens `attach` prints\n`[cotal: connection lost, reconnecting]` on stderr and starts asking the manager for a new session:\na fresh grant, a fresh per-session credential, a fresh connection, so every attempt re-runs the same\nauthorization the first attach did. On success it prints `[cotal: reconnected]`, the manager repaints\nthe seat's current screen the way it does for any attach, and you carry on in the same terminal.\nRetries wait 1s, 2s, 5s, 10s, then 30s, for as long as the seat exists. The detach key is read the\nwhole time the loop runs, the waits and the attempts alike, so a reconnect never traps you: press it\nwhile a session is being established and the attach ends there, and a session that lands behind the\npress is handed back to the manager rather than left holding a slot. Everything else you type while\nthere is no session is dropped rather than queued, so keystrokes aimed at a terminal that turned out\nto be frozen, Ctrl-C included, are not delivered to the agent by a reconnect you did not know had\nhappened. That starts before the first session, not at the first reconnect: at a terminal, `attach`\nreads and drops what you type while it is still resolving the mesh, so a key struck at a prompt that\nhas not come up yet does not reach the agent when it does.\nThe terminal is in raw mode for the whole reconnect, including when the link died before the first\nsession finished opening, so the detach key works there too instead of echoing as `^]`.\n\nA **pipe** carries script input. For example, `printf 'ls\\n' | cotal attach --name web` is\nbuffered until the session opens. Buffering continues across reconnects, so\n`tail -f log | cotal attach --name web` does not lose the part of its feed written while the link was\ndown. Only a terminal gets the reader; `--no-reconnect` keeps the old behaviour on both.\n\nIt stops on its own when reconnecting cannot help, and says why: a manager that refuses the attach\nexits non-zero with the manager's own message, and a reconnect that finds the seat no longer there\n(despawned, or its agent exited while the link was down) exits cleanly with `seat <name> is gone`.\nA local connect refusal that retrying cannot fix, such as a static-auth mesh whose seed is now\nmissing, also exits non-zero with the refusal's own sentence. A broker that is still unreachable\nkeeps the loop trying in silence.\nA refusal that could still pass, such as a manager at its session ceiling, is relayed in the\nmanager's own words while the loop keeps trying, once per refusal rather than once per attempt.\nPressing the detach key, or the agent's process exiting while you are attached, ends the attach as\nit always did. `--no-reconnect` turns all of this off and restores the single-session behaviour,\nwhich is what a script wants.\n\nEach reconnect also hands the abandoned session back to the manager, over the first link that can\ncarry the message, so an attach that flaps does not eat the manager's session slots one outage at a\ntime. If that message never gets a link, the attach says so when it ends. The live-session ceiling\ndefaults to 64 concurrent sessions (`--max-sessions`); the browser console opens one session per\npane, so a dashboard over a large mesh should size for agents \xD7 panes. Hitting the ceiling refuses\nbefore a credential is minted and names `--max-sessions`.\n\nWhich mesh `attach` resolves also decides **how it redeems the grant**. On a registered open mesh\nthere is no local seed. The CLI connects bare, the same way other control commands already do, and\nthe session rail is the caller rail that a real open-mode connection already reaches. Telling the\noperator to re-register the root is false: the registered root is already the contract. On a\nstatic-auth mesh the grant is still redeemed by minting a short-lived\nsession-scoped credential from the seed at the root the mesh resolved to, never from a `.cotal`\nfound by walking up from whichever directory you happen to be standing in. The difference is not\nhypothetical: `~/.cotal` exists on every install because the mesh registry lives there, so a command\nrun anywhere under your home directory but outside a project used to mint from your home\ndirectory's trust and present it to a broker that trusts a different chain, which surfaced as a\nbare authorization failure that named nothing. A directory that does hold another chain for the\nsame space is now reported on the way past, and not obeyed:\n\n```text\n! this directory resolves to /Users/you, whose .cotal/auth holds a DIFFERENT trust chain for space \"team\".\n attach used /Users/you/projects/app, the root this mesh resolved to. The other one is not being used, and is worth a look.\n```\n\nWhen a **static-auth** mesh holds no seed at the resolved root, `attach` refuses and names what it\nresolved, the broker and the root, instead of describing a directory it did not use and instead of\ntaking the open-mode path. An authenticated registry entry with a missing seed is still\nauthenticated. On a USER-AUTH mesh `attach` reads no seed. It sends your login and the session grant to\nthe auth service, which issues a `session-caller` bearer only if your owner and actor hold that\nsession. The connection it opens expires with the session grant.\n\nTerminal bytes stream over the mesh; the manager's own HTTP/WS face serves the console. That endpoint binds\n**loopback by default**, so nothing is exposed by accident; `cotal up --host <addr>` passes its bind\naddress down, which is what lets you reach the browser console (`cotal console`) for an agent whose manager runs on another machine.\n`attach` does not use that face: it redeems a signed mesh session grant over the broker instead (see above), so it reaches a\nremote manager regardless of the bind address. A\nbare `cotal supervise` and an embedded manager stay machine-local. Set it directly with\n`supervise --console-host <host>`.\n\nThat address is **recorded on the mesh** and carried forward, because it is a decision rather than\nsomething later commands can work out for themselves (a broker dial address is not a manager bind\naddress). Every later manager launch for the same mesh reuses it, including a same-root `cotal up` repair,\nan adopted preserved or restored listener, and a `spawn -f` manifest deploy. A manager replacement\ndoes not quietly move a reachable attach face back to loopback. Passing `--host` again overrides it,\nso you can widen or narrow exposure whenever you like; a mesh that never asked stays loopback-only\nand records nothing.\n\nBecause that face mints terminal read and write authority for every managed agent's browser session, it is credentialed in two\ntiers. A mesh caller receives a **ticket** bound to the single agent the manager just authorized,\nsingle-use and short-lived, so one authorized attach can never be re-pointed at someone else's\nagent. The **console token** is the operator's own, reaches every agent, and is printed only to the\nmanager's output. The roster, the live feed, and the PTY stream all answer `401` without one; the\nstatic console shell is served openly, since it describes no agent.\n\n## input\n\n```bash\ncotal input --name <n> --text <text> [--no-enter] [--on <instance>] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | | Managed agent to type into (required) |\n| `--text <text>` | | The text to type, taken verbatim (required) |\n| `--no-enter` | off | Type the text and stop there, without pressing Enter |\n| `--on <instance>` | class anycast | Pin to one manager instance id using the same rules as [`attach`](#managed-seats) |\n\nTypes one line into a running agent's terminal, as if you had typed it there, and returns. This is\nthe half of [`attach`](#managed-seats) that a program wants: `attach` is a live stream that holds a\nsession open and expects a terminal on your side, so a script, a cron job or a web UI cannot use it\nto send a single line. `input` is one authorized call.\n\nWhat it is for is **harness commands**. A line beginning with `/` is not chat and not a message: it\nis something the agent's own harness handles, and the only way in is the keyboard.\n\n```bash\ncotal input --name reviewer --text \"/compact\" # ask the harness to compact its context\ncotal input --name reviewer --text \"/model opus\" # switch its model\ncotal input --name reviewer --text \"hold on that PR\" # ordinary typing works too\n```\n\n**Quoting.** `--text` takes a value, so a payload starting with `/` survives as written. A payload\nstarting with a dash needs the `=` form, because the shell-style `--text --foo` is ambiguous and is\nrefused rather than guessed:\n\n```bash\ncotal input --name reviewer --text=--verbose # dash-leading text: use --text=<value>\n```\n\nEnter is pressed by default, since a command typed but never submitted has not been delivered.\n`--no-enter` types the text and leaves it sitting at the prompt, which is how you stage a line and\nsend it later.\n\nNothing comes back but a delivery receipt (`\u2713 sent 9 bytes to reviewer`, counting the trailing\ncarriage return). Whatever the agent does next shows up where its output already goes: the mesh, its\ntranscript, or an `attach`.\n\n**This one is operator-only, and more narrowly than `stop` or `attach`.** Those two are granted to\nanything holding `spawn`, so an agent can stop and attach to seats under its own owner. `input` is\nnot: it is granted only to operator credentials, which on a user-auth mesh means your ledger row\nneeds the `admin` scope, the same scope [`ps`](#managed-seats) already needs there. The reason is\nthat a write into a terminal is control of whatever is running in it, and on a user-auth mesh the\nown-owner rule covers every seat under you, not only the ones you launched: a `spawn`-scoped agent\ncould otherwise type into a sibling it never started. Seat locality is still resolved for you.\n\nOnly the `pty` runtime can be typed into. The external terminal runtimes (`tmux`, `cmux`, `orca`,\n`herdr`) attach to a process they do not own, so they have no input stream for it and the command\nrefuses by name rather than dropping the keystroke.\n\n## personas\n\n```bash\ncotal personas list [-v] [--running]\ncotal personas show <name>\ncotal personas edit <name>\ncotal personas new <name> (--prompt <t> | --from <f>) [--role <r>] [--model <m>]\ncotal personas rm <name> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh's persona catalog |\n| `--role <r>` | none | `new`: the persona's role |\n| `--model <m>` | none | `new`: the persona's model |\n| `--prompt <t>` | none | `new`: the persona's prompt text |\n| `--from <f>` | none | `new`: seed the prompt from a file |\n| `--verbose`, `-v` | off | `list`: include role / model / description |\n| `--running` | off | `list`: mark personas live on the mesh |\n| `--force` | none | `rm`: required, delete without prompting |\n\nPersonas are the local agent files under the resolved mesh root's `.cotal/agents/`, the same catalog\n`cotal spawn` launches from. `--space` and `--server` therefore move every list, read, write, delete\nand completion operation to the selected mesh. An unresolved target refuses rather than falling back\nto the current directory. See [Agent files](agent-files.md) for the file format.\n\n## supervise\n\n```bash\ncotal supervise [--runtime <name>] [--space <s>] [--server <url>] [--spawn <names>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space to supervise |\n| `--server <url>` | hosting mesh, or matching registered mesh | Broker URL. A registered mesh supplies it when omitted; a different explicit value is refused before anything is dialed. |\n| `--runtime <name>` | `pty` | Agent runtime (`pty` built in; extension runtimes are explicit-only) |\n| `--console-port <n>` | none | Protocol-console port |\n| `--console-host <host>` | loopback | Bind host for the console endpoint. Loopback keeps it machine-local; `cotal up` passes the address it bound the broker to, which is what lets the browser console reach this manager from another machine. `cotal attach` does not use this face: it redeems a mesh session grant over the broker |\n| `--max-sessions <n>` | 64 | Live-session ceiling. Each console pane and each `cotal attach` is one session, so size for agents \xD7 panes, not agent count. A capacity refusal names this flag. `cotal up --max-sessions` records the same number on the mesh so a later `supervise` started by repair or `spawn -f` keeps it |\n| `--roster <file>` | none | Declarative roster to boot at startup. See [Roster files](define-a-team.md#roster-files) |\n| `--launch <spec>` | none | Resolved manifest launch spec (from `up -f` / `spawn -f`) |\n| `--spawn <names>` | none | Comma-separated personas to pre-spawn at startup |\n\nThe manager is the agent supervisor and control plane: it answers `spawn --detach`, `stop`, `ps`,\n`attach`, and the `cotal_*` manager tools. `cotal up --detach` starts one for you; run `supervise`\ndirectly to recover a dead manager or drive a custom runtime. Default runtime is `pty`; install an\noptional provider first (`cotal ext add @cotal-ai/orca`, `@cotal-ai/tmux`, `@cotal-ai/cmux`, or `@cotal-ai/herdr`) and\nselect it explicitly. A missing provider or app fails loudly; there is no fallback. See [Deploy](deploy.md).\nBoot inventory decides whether this process takes unpinned `spawn`/`launch` on the class rail:\nif every declared connector is unavailable, those commands stay on this instance rail only\n(`status` reports `classSpawn: false`). `describe` still answers on the class rail, so an\nunpinned spawn can bind-fence against a skip member; re-issue, or pin `--on`. A partial\ninventory keeps the class rail and names `--on` on a harness refusal, because sibling\ninventories are not readable from the serve credential. See [control surface](control-surface.md#instance-routing).\n\nOn a normal `SIGINT`/`SIGTERM`, the manager stops every seat and requires the selected runtime to\nprove the seat is gone before it releases the manager lease or service registration. A stop that\ncannot prove exit fails loud and keeps manager authority instead of reporting a clean shutdown while\nan orphan still holds broker rails. After an abrupt manager death, the same logical successor\nterminalizes only its own durable static slots, verify-evicts the predecessor's broker principal,\nrecords that result in the lifecycle's caller-readable audit detail, reaps the predecessor's seat\nprocess through the runtime's custody reference recorded on the slot (the pty runtime verifies the\nprocess start identity in its seat record, so a reused pid is never signalled), and only then\nretires the lifecycle and frees the alias. A runtime that custodies its seats reserves that\nreference before it launches one, and the manager records it on the slot's first durable row, so a\nmanager that dies part-way through a spawn also leaves a seat its successor can address. A\nsame-lifecycle restart or a resume records the new seat's reference on the slot the same way, and\nwhen the slot does not take it the restart or resume fails and stops any seat it started, so the\nslot never names a seat that has already exited while its replacement runs. A resumed seat keeps\nits retained credentials, so the resume frees it only once its exit is proved; a seat whose stop\ncannot be proved stays managed, and the resume's error says so. A spawn\nthat launched its seat and then failed is rolled back by the manager that launched it, and that\nrollback reaps the seat through the same reserved reference before the lifecycle retires. Missing or unverified broker evidence keeps the slot\nterminalizing, and so does a runtime that cannot reap by reference.\n\nA `meshes add --mode user` entry is a **participant** registration, not hosting authority. A\nparticipant may run `supervise` only when the host advertises the remote manager authority service\nand the signed-in actor has the dedicated `supervise` ledger scope. The CLI obtains the closed,\nloopback-only `manager-service` view; `spawn` and `admin` do not substitute for that scope. The\nhost issues the manager's public-nkey JWT material through its lifecycle-bound prepare \u2192 activate\n\u2192 renew protocol, never by handing the participant a signer or static provisioner credential.\nThe host also performs instance-scoped eviction and guarded gate reconciliation. A remote manager\nrefreshes its short-lived registration executor before clean deregistration, so a long-running\nprocess removes its service row on `SIGINT` or `SIGTERM`. After an unclean stop, the same instance\nverify-evicts its superseded family and advances the process epoch. If an abandoned frozen gate\nholds the manager governance slot, a different supervise-scoped manager asks the host to reconcile\nthat holder after a complete gone verdict, then retries its registration once.\n\nA remote supervise never uses local signing trust: with host-issued authority in hand, the\nmanager mints from that authority alone and consults local records only to refuse a conflict,\nnamely the supervised space's own trust records under the cwd root. A root that hosts another\nstatic space beside the sign-in is a normal configuration and is never read as this space's\ntrust.\n\nThe broker URL in the registry entry decides the transport. A remote broker is often published\nover a `wss://` edge rather than a raw `nats://` port, and `supervise` dials whichever scheme the\nrecord holds, starting with the manager-authority registration it runs before the manager exists.\nThe record also decides whether that registration requires TLS, so a participant never downgrades\nthe credential exchange to a plaintext connection the registry did not describe.\n\nWithout that advertised host service or scope, `supervise` refuses before it starts a manager.\nRun `cotal spawn` without `--detach` to launch a foreground agent, or ask the space host to enable\nthe authority service and grant `supervise` for detached agents. If a running remote manager loses\nrenewal, it reports degraded state and refuses unsafe new starts and restarts; live agents are not\nsilently replaced. Do not run `cotal down` or `cotal up` on a participant machine to repair this\ncondition.\n\n## service\n\n```bash\ncotal service install [--mesh <name>] [--linger]\ncotal service status [--mesh <name>] [--json]\ncotal service uninstall [--mesh <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--mesh <name>` | this folder's mesh | The mesh whose manager the service runs; one unit per mesh |\n| `--linger` | off | install: when lingering is off, ask logind to enable it so the user manager starts at boot and the service survives logout. Never enabled silently |\n| `--json` | off | status: machine-readable output |\n\nRuns the manager as a user service so it survives logout and reboot. On Linux this installs a\nsystemd user unit (`~/.config/systemd/user/cotal-manager@<key>.service`, where `<key>` is the\ncase-safe mesh key); on macOS a launchd agent plist under `~/Library/LaunchAgents/`. Any other\nplatform, or an absent systemd/launchd user session, fails with a message naming what is missing.\n\n`install` resolves the mesh from the registry and binds the unit to that entry's root and broker\naddress, so it can be run from any directory. The mesh must be registered (`cotal up` or\n`cotal meshes add`) before installing; an unregistered name refuses before anything is written.\n\nThe unit's `ExecStart` is the bare `supervise` command. The mesh facts travel in the unit's\nenvironment (`COTAL_SPACE`, `COTAL_SERVER` pinned to the registered broker URL, whatever port it\nlistens on) rather than the command line, because command lines are readable by every user on a\nmulti-user host. On Linux that environment is a `0600` `EnvironmentFile`; on macOS it is the\nplist's `EnvironmentVariables`. The same environment gives the service a private `COTAL_HOME` and\n`XDG_CONFIG_HOME` under the unit directory, so the service manager never touches the login\nuser's `~/.cotal`. First-run connector seeding runs synchronously inside `service install`,\nagainst that private config root; the unit itself starts with `COTAL_SKIP_CONNECTOR_SEED=1`\nso a manager is never interrupted mid-seed by a restart. An install whose pre-seed cannot\ncomplete (network unreachable, registry error) refuses instead of deferring.\n\nThe same environment pins `PATH` to the `PATH` of the shell that ran `install`.\nWithout it the unit inherits the service manager's own short\n`PATH`, which usually lacks `~/.local/bin` and Homebrew, so the manager's boot inventory would\nreport a harness unavailable that your shell resolves. Install from a shell that resolves every\nharness the service should launch, and reinstall after moving one. A relative entry, including\nan empty one, is resolved against the directory you ran `install` from, because the unit starts in\nthe mesh root where the same spelling names another directory. An entry with a `..` segment is\npinned as the directory your shell reaches through it, with symlinks followed, and refuses when it\nreaches none. A `PATH` set to the empty string is one empty entry, so it pins that directory. An\nunset `PATH` refuses.\n\nEvery value the unit derives from a path (`WorkingDirectory`, the `EnvironmentFile` path, the\n`ExecStart` tokens) is escaped for systemd specifiers (`%` becomes `%%`), so a mesh root that\ncontains `%` starts over its real path instead of a path systemd rewrote by expanding it. The\nprovenance comment records the root unescaped.\n\nOn Linux a user unit starts at boot and survives logout only while the user lingers. Without\nlingering, systemd starts no user manager at boot, so an enabled unit stays inert until the next\nlogin and stops at the last logout. `install` checks lingering before it writes anything, and when\nlingering is off it fails with the root command that turns it on (`sudo loginctl enable-linger\n<user>`). With `--linger` it first asks logind to enable lingering for the current user, and fails\nwith the same command when logind refuses (unprivileged users over SSH get `Access denied`).\n`service status` prints that command while lingering is off. A Linger query that does not answer\n`yes` or `no` (logind unreachable, no `loginctl`) is never read as off: `install` refuses with\nthe query's own error and enables nothing, and `service status` shows lingering as unknown with\nthat error (`--json` gives `\"linger\": { \"error\": ... }`).\n\n`service install` also refuses while a manager is already running for the mesh (`cotal down\nmanager` first). The restart policy is `Restart=always` with `RestartSec=20s`, chosen for\nmanager units in production: a manager exits for reasons that are not failures (broker\nrestarts, host suspend), where `on-failure` with a short interval thrashes.\n\nThe unit also sets a start limit (`StartLimitIntervalSec=30min`, `StartLimitBurst=20`). A manager\nthat keeps failing to start stops after 20 attempts, about seven minutes at 20 seconds apart, and\nthe unit is left `failed` instead of restarting forever. One such failure is deliberate. After an\nunclean stop, a manager that cannot verify eviction of its predecessor's credentials exits 1 and\nleaves the issuance gate frozen, because starting without that proof could let two incarnations\nserve at once (SPEC 13.1). It first waits up to 60 seconds for the delivery daemon to answer, so a\ndaemon that is still starting does not fail the start. The log names the cause. When the delivery\ndaemon is down, it says the daemon is not reachable on the `ctl.delivery-admin` rail. When the\ndaemon answers and refuses, for example because the space is missing a `$SYS` cred, it prints the\ndaemon's own reason and repair step. Fix that cause, then run `systemctl --user reset-failed\n<unit>` and `systemctl --user start <unit>`. The macOS agent has no start limit: launchd's\n`ThrottleInterval` only spaces restarts.\n\n`service status` reports the unit state from systemd/launchd, the manager's own health read from\nits pidfile at the unit's recorded root, and the machine facts a hosting side asks for:\narchitecture, OS (the platform, never the hostname), whether `/dev/kvm` is present and\naccessible, CPU count, and total memory. `--json` returns the same fields as one object. The\nmanager row names the recorded pid, and the command it runs when another program has reused that\npid. `--json` also gives the command of a live recorded pid whenever it can be read.\n\n`service uninstall` stops and disables the unit and removes it plus the private state directory.\nIt works from any directory: the unit's own records name the mesh and root it serves, and an\nexplicit `--mesh <name>` selects it. It refuses any unit that was not written by `service\ninstall` (the files carry a provenance comment), whose recorded mesh is missing, or that was\ninstalled for a different mesh, so operator-written units are never destroyed; `service status`\napplies the same rule and never reports a mesh a unit does not record.\n\nThis command installs only the manager. The per-space auth service and the delivery daemon are\nnot installed by it: on a shared broker an operator runs three units per space with `After=`\nedges (auth service, then manager, then delivery) and stops them in reverse. A broker-side `cotal\nup` unit is a separate unit documented in [Run a mesh](run-a-mesh.md).\n\n## reconcile-gate\n\n```bash\ncotal reconcile-gate [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the frozen gate lives in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint whose gate is frozen |\n| `--instance <id>` | this folder's persisted manager instance | Instance id |\n\n**When you need this.** A manager restart killed after deregistration begins but before the new\nincarnation finishes leaves the endpoint's issuance gate *frozen*, held by a\nprocess that no longer exists. The freeze is what stops two incarnations serving at once, which is\ncorrect. The successor manager now completes that dead registration itself on boot, including on\nthe remote user-auth path. A foreign remote manager blocked by this gate also asks the host to repair\nit before one registration retry. Both use the same guard this command uses: they act only when the freeze-holder is affirmatively gone under a complete\nCONNZ sweep (`gone` and `sweepComplete=true`). If that registration's spec write already committed,\nit finishes the same freeze at the committed registration revision. If the spec did not advance, it\nabort-reopens the gate at generation+1 with processEpoch unchanged and continues the normal takeover.\nLive, unknown, unestablishable, and\nwrong-op-kind still refuse; there is no TTL.\n\nUse this command when the automatic path cannot run: the delivery daemon is down, the repair targets a\nnon-manager endpoint, or you want to lift the freeze without starting a manager. It checks that the\nholder really is gone, prints what it found, and then finishes the dead operation the same way as the\ninterrupted restart would have: revoke the old credentials, evict their holders with verification,\nand reopen the gate.\n\nThe command revokes the old credentials 16 at a time. It then verifies the holders' eviction in\nshared sweeps of up to 256 holders on the delivery daemon. Each sweep scans the broker a fixed number\nof times and kicks live connections 16 at a time, so holders that are already gone add almost\nnothing and live ones add one broker round trip per 16 connections. The daemon must serve the\n`evictPrincipals` verb; an older daemon refuses it and the gate stays frozen.\n\nEach sweep durably records the holders it verified before the next sweep starts. If a holder is not\nverified gone, the command leaves the gate frozen with those records kept. An interrupted sweep\nrecords nothing, and the sweeps before it stay recorded. A retry still repeats the freeze-holder\nliveness check, then skips only progress bound to the same registration operation, frozen-gate\nrevision, and holder set.\nThe output reports holders completed before this attempt, completed now, and still remaining. A new\nfreeze or changed holder set starts from zero. Cursor cleanup happens only after reopen; a retained\ncursor is harmless because its old gate revision cannot authorize a later freeze.\n\n**It refuses far more often than it acts, on purpose**, and always says which check stopped it:\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `holder-alive` | The freeze-holder still has a live connection: a manager *is* running | Stop that process first. Reconciling would evict a live manager's credentials |\n| `holder-unknown` | The connection sweep could not prove the holder absent | Not safe to proceed: an unprovable holder is treated as a live one. Re-run once the broker answers completely |\n| `liveness-unestablishable` | The delivery daemon gave no verdict: it was unreachable, timed out, or refused | Act on the delivery lease line in the refusal (below). Silence is never read as death |\n| `not-frozen` / `no-gate` | The gate is open, or there is no gate at that coordinate | Nothing to repair: check `--endpoint` / `--instance` |\n| `wrong-op-kind` | Frozen under a takeover or retirement, not a registration | Out of scope for this command; it will not reinterpret another operation's intent |\n| `eviction-unverified` | The holder looked gone but eviction could not be verified | The gate is left frozen, unchanged. Investigate the broker before retrying |\n| `raced` | A newer manager moved the gate mid-repair | Re-run `cotal doctor` and look again |\n\nWhen the daemon gives no verdict, the refusal also reads the delivery lease (`lease.0`) and names\nwhat is blocking the rail:\n\n| Lease reading | What to do |\n|---|---|\n| absent | No daemon is running. Start it (`cotal up` runs it) and re-run |\n| unreadable | The daemon cannot be named, so do not assume none is running. Fix the lease read, then re-run |\n| held, not ready | That holder claimed the shard and has not bound its rails. Wait for it, or stop it so its lease lapses |\n| held, ready, no answer | The query may have gone to another daemon still subscribed to the rail, such as a stopped one whose lease lapsed. Re-run before stopping anything. If no run gets an answer, stop any other delivery daemon for the space, then stop or restart the holder |\n| changed hands | The holder took the shard after the query was sent, so it was never asked. Re-run before stopping anything |\n\nThe command reads the lease before it sends the query and again after the query fails. It names a\nholder as the blocker only when the same run of the same daemon held the lease both times, and two\nrows from a daemon too old to record its run never count as the same run. Even then a ready holder\nmay not have been asked: the rail is queue-grouped, so any daemon still subscribed to it can take\nthe query. A row whose times are not valid dates reads as unreadable.\n\nA daemon that answered and refused keeps its own reason, followed by the same lease line. The lease\nline names the holder, whether it is ready, the space account that holds the lease bucket, when that\nholder acquired the shard, and when the row was last written. A ready holder rewrites the row on\nevery renewal and keeps its acquisition time, which only a successful acquisition sets. A row\nwritten by a daemon that predates the acquisition time reports it as unknown. The lease reads never\nchange the outcome: the gate stays frozen and the command exits 2. A manager's boot self-heal uses\nthe same check and reports the same line.\n\nThere is no `--force`, and no path that discards gate state: the only way this reopens a gate is by\nproving the holder is gone and then completing the operation properly.\n\n**What reopening the gate does for the endpoint's governance slot.** A registration takes the\nendpoint-wide governance slot before it publishes its spec, and holds it until its gate reopens. An\ninstance that died between those two points leaves the slot held with no registration behind it.\nThis command does not write that slot and never has; the registration path is its only writer. What\nthe reopen does is advance the holder's gate past the generation the slot is stamped with, which is\nwhat marks the slot abandoned. The next registration for that endpoint then reclaims it as part of\nits ordinary start. So the repair here is still one command followed by starting the manager, and\nthe slot needs no separate step.\n\n## deregister-instance\n\n```bash\ncotal deregister-instance [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the instance is registered in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint the instance serves |\n| `--instance <id>` | this folder's persisted manager instance | Instance id, the whole id as `cotal ps` prints it |\n\n**When you need this.** The service registry records *registration*, not liveness, and nothing in\nthe model expires a row. A manager that stops cleanly removes its own registration. One whose host\ndied without writing anything cannot, so its record goes on claiming a live instance forever: every\nclass scatter in that space freezes the dead slot in, and `cotal ps`, `stop` and `attach` each pay\ntheir whole deadline waiting for a machine that is never coming back. A laptop that was reimaged, a\ncontainer that was deleted, a box that will not be back on the network: those registrations have no\nother exit.\n\nThis command is that exit. It asks the instance first, and it removes a record only when the broker\naffirms the instance's own rail is empty: nothing subscribed there. Then it deletes the\nregistration's two records keys, each pinned to the revision it read, and prints what it removed.\n\n**Silence alone never passes.** An unanswered describe is what a dead host, a wedged process and a\nslow one all look like, and a hung process still holds its subscriptions, so the broker sees\ninterest on its rail. That instance is refused and the observation is printed. A dead process holds\nno connection and therefore no subscription, so a real corpse is still removed.\n\n**Every refusal names the failed check:**\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `instance-answered` | The instance answered a pinned describe. It is alive | Nothing to repair. If it is wedged rather than gone, stop the process first; its own clean stop removes the record |\n| `instance-not-affirmed-gone` | It did not answer, and the broker did not report its rail empty, which is what a held subscription looks like: slow or hung, not affirmed gone | Nothing was removed. Stop the process; its record goes on its own clean stop, or re-run this once it is down |\n| `liveness-unestablishable` | The probe itself failed, so nothing was learned | Fix the probe's path (credential, broker) and re-run. A probe that could not run is never read as death |\n| `not-registered` | No registration at that coordinate | Check `--instance` and `--endpoint`. This takes the whole id, never a prefix |\n| `registration-in-flight` | The instance holds the endpoint governance slot at the live issuance-gate generation, so a registration is still completing | Nothing was removed. Wait for that registration to finish, then re-run |\n| `superseded` | The record moved between the read and the delete | Something is writing to it. Nothing was removed; re-observe before retrying |\n\nThere is no `--force` and no sweep: silence is not death, and a rule that removed rows on silence\nwould eventually remove a live instance that was merely slow. An operator names one instance, the\nbroker's verdict on its rail is what authorizes the removal, and the guard's job is to show them\nthey named a dead one. Removal is not a one way door either. The same instance re-registers over\nthe tombstone on its next start, under the same identity.\n\n## runtimes\n\n```bash\ncotal runtimes\n```\n\nLists every agent runtime the manager can spawn through: the built-in `pty`, the official providers\n(`orca`, `tmux`, `cmux`, `herdr`), and any custom provider installed via `cotal ext add`. Each installed\nprovider is probed so you can see what is actually reachable on this machine before selecting it:\n\n```\npty built in\norca installed \xB7 reachable @cotal-ai/orca\ntmux available \xB7 cotal ext add @cotal-ai/tmux\ncmux available \xB7 cotal ext add @cotal-ai/cmux\nherdr available \xB7 cotal ext add @cotal-ai/herdr\n```\n\n`installed \xB7 reachable` / `unreachable` is the provider's own `available()` probe; `available` means\nit is a known runtime you can add with the shown command. Selecting an unknown or uninstalled runtime\nvia `up`/`spawn --runtime <name>` fails loud and, for a known one, points at the exact `cotal ext add`\npackage. There is no silent fallback to `pty`.\n\n## seats\n\n```bash\ncotal seats [--drain]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--drain` | off | Retire every seat whose agent has exited. A seat whose agent still runs is kept |\n\nThe pty runtime used to start a detached custodian process for every Linux seat. It now spawns\nin-process, but custodians that an earlier manager started keep running, and one whose agent has\nexited stays resident while a manager still holds its connection. This command lists the custody\nrecords under `COTAL_SEAT_ROOT` (default `~/.cotal/seats`), one line per seat:\n\n| State | Meaning |\n|---|---|\n| `live-child` | The agent process still runs. The seat is never signalled, and a manager can still adopt it |\n| `childless` | The agent has exited, or the record comes from an earlier boot. `--drain` retires the seat |\n| `drained` | `--drain` proved the custodian and the agent gone and removed the record |\n| `refused` | The record cannot be read, carries no start or boot identity, this host publishes no boot identity, or the reap could not prove the processes gone. The record stays on disk |\n\nA drain signals only a custodian whose recorded start identity still matches the live process, so\na reused pid is never touched. No process outlives a reboot, so a record from an earlier boot is\nreported childless and `--drain` removes it without signalling anything. A record with no start or\nboot identity is refused with or without `--drain`, and is never reported as running or exited.\nOn a host that publishes no boot identity (`/proc/sys/kernel/random/boot_id`) every record is\nrefused the same way, because no record can be tied to this boot.\nThat refusal and an unreadable record signal nothing. A refusal from the reap itself can come after the drain already\nsent `SIGKILL` to the custodian. Its detail names the pid or process group the reap could not prove\ngone, so check those processes before you retry. The command exits non-zero when any record is\nrefused. It is Linux-only and throws on other platforms.\n\n## send\n\n```bash\ncotal send dm <agent> \"<text>\" [--space <s>] [--server <url>] [--creds <path>]\ncotal send msg <channel> \"<text>\"\ncotal send ask <role> \"<text>\"\n```\n\nA `send dm` prints one line naming three facts: `\u2192 <name> stored seq <N>; recipient <status>\nat send; delivery not confirmed <text>`. `stored seq N` is the JetStream sequence the broker\nassigned to the publish; `recipient <status> at send` is the roster status (`idle`, `working`,\nor `offline`) resolved right before the publish, which can change the instant after; the send\nnever prints `delivered`, because the sender's credential cannot read the recipient's durable\nto confirm it. Inspect what the broker actually holds for a recipient with\n[`cotal deliver pending`](#deliver).\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and (off-registry) which credential |\n\nOne-shot messaging: connect, send a single direct message (`dm`), channel post (`msg`), or role\nask/anycast (`ask`), then exit. For a running conversation, agents use the mesh tools instead\n([MCP tools](mcp-tools.md)).\n\n`cotal send` works from an operator shell or from a seat. Its display name is `<login>@<host>` of\nthe shell that ran it, so the recipient can tell one operator's send from another's; it is taken\nfrom the operating system, never from `COTAL_NAME`. The wire principal comes from the resolved\noperator credential or user bearer, not from `COTAL_NAME`, `COTAL_ID`, `COTAL_OWNER`, or\n`COTAL_ACTOR`. On an open mesh the transient endpoint self-mints its principal.\n\nThe transient endpoint never joins the roster and binds no inbox. A recipient can still answer a\n`send dm` or `send ask` with `cotal_dm`, by the sender's name or by the id on the message it\nholds: the reply is stored under the sender's id in the space's DM history, which an operator's DM\nview such as the dashboard's Direct messages lens shows. The `cotal send` that asked has already\nexited, so the reply never reaches that shell.\n\n## channels\n\n```bash\ncotal channels list\ncotal channels set <name> [--replay | --no-replay] [--window <n>] [--desc <s>] [--instructions <s>]\ncotal channels default --replay | --no-replay\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--replay` / `--no-replay` | none | `set`/`default`: replay history to new joiners, or not |\n| `--window <n>` | none | `set`: replay window size |\n| `--desc <s>` | none | `set`: one-line channel description |\n| `--instructions <s>` | none | `set`: instructions shown to joiners |\n\nInspects and edits the channel registry: replay policy, description, and joiner instructions. ACL\nsemantics (who may read or post) are set at mint / provision time, not here; see\n[Channels and permissions](channels-and-permissions.md). On a user-auth mesh, `list` rides your\nown login as is; `set` and `default` edit the registry over a short-lived\nchannel-writer view, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\nOn a remote user-auth mesh that view is served by the public exchange; space-history `purger`\nand the read-only admin view are not.\n\n\n## history\n\n```bash\ncotal history clear --force [--dms] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--dms` | off | Also clear DM history |\n| `--force` | none | Required: clear without prompting |\n\nPurges retained channel history; `--dms` extends it to direct-message history. An alias of\n[`clean history`](#clean). On a user-auth mesh the purge rides a short-lived purger view over\nyour login, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\n\n## console\n\n```bash\ncotal console [--plain] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to watch |\n| `--plain` | off | Line stream instead of the TUI |\n\nA live protocol view for a space: a lazygit-style TUI, or a plain line stream on `--plain`. On a\nuser-auth mesh it rides the read-only admin view over your login, which needs ledger scope\n`admin`. Inside the TUI, operator control (`D` kill, `:spawn`, `:status`, `:purge`) rides the\nsame per-action instrument path as `cotal stop` and `cotal ps`, never the observer; a raw\n`--creds` file cannot drive it. `a` (or `:attach <agent>`) runs\n[`cotal attach`](#managed-seats) in place and returns to the console on detach. See\n[Watch a mesh](watch-a-mesh.md).\n\n## web\n\n```bash\ncotal web [--detach] [--host <host>] [--port <n>] [--no-open] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to serve |\n| `--host <host>` | `127.0.0.1` | Concrete HTTP bind and browser host; wildcard addresses are refused |\n| `--port <n>` | `7799` | HTTP port, a decimal number from 1 to 65535 |\n| `--detach` | off | Run in the background; stop with `cotal down web` or bare `cotal down` |\n| `--no-open` | off | Don't open the browser |\n\nThe browser observability dashboard: presence, channels, and a live feed. It is **not** part of\n`cotal up`: it ships inside `cotal-ai` as the `@cotal-ai/web` extension, seeded automatically on first\nrun (like the built-in connectors) so it always matches your CLI version. It self-registers `cotal web`\ninto this surface and serves\n`http://cotal.localhost:7799` by default (loopback; `*.localhost` resolves in Chrome/Firefox/Edge; for Safari\nor a system resolver such as WSL2's, the launch link is also printed at `http://127.0.0.1:7799`).\nOn a user-auth mesh the dashboard rides the read-only admin view\nover your login, and a channel purge asks for its own channel-purger view per click; both need\nledger scope `admin`. The public exchange serves `channel-purger` for a remote owner; it still\nrefuses the startup admin view, so a remote `cotal web` is not a complete channel-management\nsurface. Detached mode re-execs the current Cotal installation, writes diagnostics to\nthe mesh root's `.cotal/web.log`, and reports success only after the HTTP server answers. It requires\na recorded mesh root, but can be launched from any directory once `cotal up` has recorded the mesh.\nSee [Watch a mesh](watch-a-mesh.md).\n\n## deliver\n\n```bash\ncotal deliver [--space <s>] [--server <url>] [--tls] [--creds <file>] [--root <dir>] [--shard <n>] [--shards <n>] [--dev-mint]\ncotal deliver pending <name> [--limit <n>] [--durable <name>] [--json]\n```\n\nWith no positional, `cotal deliver` runs the delivery daemon (see\n[the delivery daemon](delivery-daemon.md)). `deliver pending <name>` never starts the daemon: it\nis an operator-only read over one recipient's DM durable, for the moment after a send when the\nquestion is \"what does the broker actually hold for them.\" It resolves `<name>` against a short\npresence watch (an `offline` card still counts, since the recipient may be dead, that is what\nthe verb exists to inspect); when neither a card nor the durable can be found, it prints\n`\u2717 not-found: no agent \"<name>\" and no DM durable for it in space <s>` and exits non-zero, never\n`pending 0`. On a match it prints the durable name and one fact per line: `pending`,\n`ack-pending`, `delivered`, `ack-floor`, `created`, `frontier`, and the stream's `max_age` /\n`max_msgs_per_subject` / `discard` limits (`--json` prints the same facts as one object), followed\nby a bounded, unacked read of up to `--limit` (default 20) recent candidate message ids under the\nheading `recent candidate ids (from the ack floor; not proof of a hole)`, a list of what is\nthere, not proof that nothing was lost.\n\nThe verb needs the `admin` credential profile: it runs through the same static-mesh route as\n`cotal mint --profile admin`, and refuses a user-mode mesh, naming the retired static credential,\nbecause there is no user-mode inspection authority yet. Pass `--creds <file>` for an off-registry\nadmin credential. A same-name respawn never inherits a predecessor's held DMs (the durable is\nlifecycle-keyed); an old lifecycle's durable is reachable only by the name a live read printed\n(the `<durable>` line on the first line of this verb's output). Pass that name with `--durable\n<name>` to read it directly once the lifecycle's card is gone from the roster. This skips the\npresence watch on `<name>` entirely, so `<name>` is required but only echoed in error text.\n\n## mint\n\n```bash\ncotal mint <name> [--profile <agent|observer|admin>] [--out <path>] [--signer]\ncotal mint <name> --provision [--role <role>] [--space <s>] [--server <url>]\ncotal mint <name> --expires-in <seconds> | --expires-at <unix-seconds>\ncotal mint <name> --identity <creds> [--expires-in <seconds>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--profile <agent\\|observer\\|admin>` | `agent` | Credential profile |\n| `--out <path>` | `.cotal/auth/creds/space.<key>/<name>.creds` | Output path - the default sits under the resolved space's segment (`<key>` is that space's hex encoding, as in [Project files](config.md#project-files)) |\n| `--signer` | off | Emit a stripped account-signing file instead |\n| `--force` | off | With `--signer`: overwrite an existing file |\n| `--allow-subscribe <a,b>` | the agent file's, else subscribe | Read-ACL override, **agent profile only**: `observer` and `admin` carry a fixed read set, and `mint` refuses this flag there rather than narrowing nothing |\n| `--allow-publish <a,b>` | the agent file's, else deny | Post-ACL override, **agent profile only** |\n| `--role <role>` | the agent file's | Agent profile: the anycast task queue the identity pulls (`svc_<role>`) |\n| `--provision` | off | Agent profile: also pre-create the identity's bind-only DM/deliver durables (and its role's task queue) on the live mesh, so the credential can consume |\n| `--expires-in <seconds>` | unbounded | Bound the credential's lifetime: the JWT `exp` is `iat + <seconds>`. A positive integer; refused together with `--expires-at` |\n| `--expires-at <unix-seconds>` | unbounded | Bound the credential to an absolute `exp` (unix seconds). Refused together with `--expires-in` |\n| `--identity <creds>` | a fresh identity | Re-mint for the nkey carried by this creds file, keeping the principal and every durable keyed to it. The file is read by the same loader the endpoint uses; a file with no seed is refused by name |\n| `--space <s>`, `--server <url>` | the resolved mesh | Which root supplies the agent file, static trust and default credential storage; with `--provision`, also which live mesh receives the durables |\n\nMints a NATS creds file for a space in **static** auth mode, scoped to a profile and (optionally)\nexplicit read/post ACLs. `--signer` emits an account-signing file for delegating minting to another\nhost. A per-user-auth space refuses `mint`: agents there join under a logged-in user\n([`login`](#login) + [`actor grant`](#actor)), never via a handed-out creds file. See\n[Identity and auth](identity-and-auth.md).\n\nFor an agent profile, the resolved mesh root supplies the persona ACL, the signing material and the\ndefault credential destination as one authority. If the current folder also holds trust for a\ndifferent space or account, mint refuses before writing and names both roots. It never combines a\npersona from one root with credentials signed or stored under another.\n\nA plain mint is creds only: the identity can publish within its post ACL at once, but on an authed\nmesh its DM inbox and task queue are provisioner-pre-created and bind-only, so a **consuming**\nconnect fails until they exist. `--provision` performs that pre-create in the same command (a\nprovisioner cred is minted from the space's trust material, used, and dropped), so a long-running\nclient you start yourself can receive DMs and role anycasts like a spawned seat. The command prints\nthe identity's principal (its wire id) and lifecycle uid; a consuming client passes that uid as its\n`lifecycleUid`. Agent profile only; an open mesh needs none of this (peers self-create there). The\nsame resolved authority is used for both the credential and `--provision`, so the broker\nfootprint cannot be created under a different root's trust material.\n\nThe CLI-mintable profiles carry no default TTL: without a lifetime flag the credential is\nunbounded, and a standing-renewal consumer refuses it. `--expires-in <seconds>` (or\n`--expires-at`) is the door the renewal seam's own error names. `--identity <creds>` re-mints for\nthe nkey the file already carries, so the new credential presents the SAME principal and every\ndurable keyed to it survives; combine it with a lifetime flag to rotate an expiring credential\nwithout churning the identity.\n\n## Login\n\n```bash\ncotal login --idp <auth base URL> [--client-id <id>]\ncotal logout --idp <auth base URL>\n```\n\nSigns you in to a per-user-auth mesh's IdP (device code flow) and caches the session; run it\nonce per machine. The IdP URL must use `https://`. Plain `http://` is accepted only on a loopback IP\nliteral such as `127.0.0.1` or `::1`, for an IdP on this machine. `localhost` is refused because it\nis a name: a hosts entry would choose the IdP. It prints your IdP subject, the id the operator\ngrants against. When the trusted\n`/token` response advertises a same-origin space catalog, login validates and records that account's\nspaces immediately. After a\nlogin, every command on that mesh works under your identity: each connect takes a fresh IdP\nproof, exchanges it locally for a short-lived bearer, and is authorized against the actor\nledger at connect time. `logout` revokes the IdP session, clears its cache, and removes only that\naccount's discovered registry entries. See\n[identity & auth](identity-and-auth.md).\n\n## actor\n\n```bash\n# an upsert of the WHOLE row: name all three ACL flags, or pass --full for the wide defaults below\ncotal actor grant <actor> --sub <IdP subject> --scope a,b --allow-subscribe a,b --allow-publish a,b [--role <r>] [--label <l>]\ncotal actor grant <actor> --sub <IdP subject> --full [--scope a,b] [--allow-subscribe a,b] [--allow-publish a,b] [--role <r>] [--label <l>]\ncotal actor revoke <actor> (--sub <IdP subject> | --owner <u_\u2026>)\ncotal actor list\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | the folder's | Space whose ledger to manage |\n| `--sub <subject>` | none | The IdP subject (shown by `cotal login`) the actor belongs to |\n| `--owner <u_\u2026>` | none | The derived owner token (alternative to `--sub`) |\n| `--full` | off | Fill each ACL flag left off with its wide default; without it, `grant` refuses unless all three are named |\n| `--scope <a,b>` | `spawn,role:default` with `--full` | Capability scope (`''` = none; `spawn` = may run agents; `role:<r>` = may delegate role r; `admin` = cross-agent control; `supervise` = eligible for the closed remote manager-service view when the host enables it) |\n| `--allow-subscribe <a,b>` | `>` (all channels) with `--full` | Channel read ACL; the user's envelope, their agents can never read beyond it |\n| `--allow-publish <a,b>` | `>` (all channels) with `--full` | Channel post ACL; also the envelope for their agents' posting |\n| `--role <r>` | none | Role (scopes the task-queue consumer) |\n| `--label <l>` | none | Display label for `actor list` (never the IdP subject) |\n\nThe actor ledger is the single authorization source of a user-auth space: no row, no access.\n`grant --full` is the **full** envelope (all channels; scope `spawn,role:default`, so it may spawn and may delegate the default role). A\n`grant` that leaves off `--scope`, `--allow-subscribe` or `--allow-publish` without `--full` is\nrefused and writes nothing. A re-grant **replaces the whole row**, not the one field you name, so to add a capability spell\nevery field out: the new scope plus the row's current read set, post set, role and label\n(`cotal actor list` shows what a row holds). Under `--full`, a field left off does not stay as it\nwas: it reverts to the wide default in the table above. A re-grant retires the current interactive lifecycle through the running auth\nservice before it rotates the row, so copied bearers cannot cross an authorization update. If that\nretirement cannot be confirmed, the row is left unchanged and the command fails with the recovery\naction. `revoke` uses the same retirement before deleting the row, which lets a later grant create a\nreal successor instead of colliding with a live predecessor. `supervise` is separate from `spawn` and `admin`: it only makes a signed-in\nperson eligible for the host-provided closed remote manager-service view; it does not grant\nmanagement of another owner or a general host profile. `revoke` denies the next exchange and\nthe next connect with no restart, and evicts the principal's live connections. Managed-agent rows\n(written by the spawn path) live in a disjoint row space this command never touches. See\n[identity & auth](identity-and-auth.md).\n\n## doctor\n\n```bash\ncotal doctor auth [--fix]\n```\n\nCredential-health diagnosis and repair for this folder's mesh: renders every managed\ncredential as healthy / near-expiry / expired and ends in `healthy` or the exact next\ncommand; `--fix` applies the repairs it can. The one surface every stale-credential error\npoints at. `--fix` takes the mesh's renewal lease when the broker answers and refuses while\na manager or another doctor holds it; with no broker it repairs offline and says so.\n\n## join\n\n```bash\ncotal join --space <s> --name <n> [--role <r>] [--channel <c>]\ncotal join --link <url> | --token <t>\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and which credential |\n| `--name <n>` | none | Your presence name |\n| `--role <r>` | none | Your role |\n| `--channel <c>` | none | Channel to join |\n| `--kind <k>` | `agent` | Endpoint kind |\n| `--link <url>` | none | Join link (`cotal://\u2026`) |\n| `--token <t>` | none | Join token |\n| `--lifecycle-uid <uid>` | none | Required with `--creds`: the lifecycle UID minted alongside the credential (`COTAL_LIFECYCLE_UID` works too). A credential's durable grants name exact lifecycle-keyed resources, so `join` refuses to invent one |\n| `--tls` | off | Connect over TLS |\n\nAn interactive presence: join a space under your own name and role, without launching an agent\nharness. A `--link` or `--token` supplies the where and the auth in one value. See\n[Spaces](spaces.md) and [Identity and auth](identity-and-auth.md).\n\n## Manifest deploys\n\nA `cotal.yaml` manifest declares a whole mesh (channels, personas, roles, and ACLs) in one file.\nThree commands consume it, plus a read-only validator:\n\n```bash\ncotal up -f cotal.yaml # boot a fresh mesh from the manifest\ncotal spawn -f cotal.yaml # deploy the manifest additively onto a running mesh\ncotal down -f cotal.yaml # tear that deploy down (or --run <id> for one run)\ncotal topology view -f cotal.yaml # validate + view the access graph, change nothing\n```\n\n`up -f` and `spawn -f` differ in target: `up -f` brings up a new broker and applies the manifest;\n`spawn -f` requires an already-reachable mesh and applies additively (ownership-scoped). On a\nuser-auth mesh, `spawn -f` deploys over your own login (the deployer view, gated on ledger scope\n`spawn`): the manifest's agents land under your owner, a manifest claiming another owner is\nrefused, and seeding new channels additionally needs scope `admin`. Both take\n`--dry-run` to print the plan without mutating anything. `topology` validates the manifest and\nrenders its channel / role / ACL graph. See [Define a team](define-a-team.md) and the\n[manifest reference](manifest.md).\n\n## ext\n\n```bash\ncotal ext # same as `list`\ncotal ext add <npm-package>\ncotal ext remove <name>\ncotal ext list [--json]\ncotal ext root # print just the install prefix (scriptable)\ncotal ext seed [--repair|--reset|--force]\n```\n\nOperator-installed extensions: `add` installs an npm package into a cotal-owned prefix and records\nevery registry provider it contributes. Commands appear in help, completion, and dispatch; runtime\nproviders are lazy-loaded by commands such as `supervise`; local process providers participate in\n`status` and selective `down`. `remove` and `list` manage them. The `@cotal-ai/web` dashboard is the\ncanonical command/process example. Installed packages and their location are described in\n[config](config.md). When a package needs an export its linked `@cotal-ai/*` peer does not have,\n`add` rolls back and names which install is behind, as a later load of an installed one does.\n\nBare `cotal ext` lists the inventory, headed by the install prefix. That prefix is a cotal-owned npm\nroot kept **separate** from npm's own global tree. These packages never show up in `npm list -g`,\n`cotal ext` (or the Extensions section of `cotal status`) is the canonical inventory. `cotal ext root`\nprints only the path, for scripts. The versions shown are the manifest pin recorded at add time.\n\n`ext list --json` (or bare `ext --json`) prints one JSON object per installed extension per line:\n`pkg`, `version`, `spec` (what `ext add` was given), `seeded` (`true` for an entry the built-in seed\ninstalled) and `provides` (the `kind:name` refs the table shows). No header or footer reaches stdout,\nand an empty prefix prints nothing and exits 0. The table is presentation and is not a stable parsing\ntarget. The other `ext` subcommands refuse `--json`.\n\nRemoving an extension that owns a running local process is refused with the mesh root and its\n`cotal down <component>` command; stop it first so uninstalling the package never strands a process\nwhose lifecycle provider is gone.\n\n### Built-in connectors are seeded extensions\n\nThe first-party agent connectors (`claude`, `opencode`, `codex`, `hermes`, `jcode`, `pi`) are not compiled into\nthe binary. They are seeded on first run through the **same** `ext add` path a third party uses, and\nappear in `cotal ext list` like any other extension. So you can remove one you do not want\n(`cotal ext remove @cotal-ai/connector-hermes`), and a deliberately-removed connector STAYS removed\nacross upgrades. `cotal ext add <your-package>` adds a third-party connector the same way. The web\ndashboard (`@cotal-ai/web`, providing `command:web`) is the seventh built-in seeded on the same path.\n\n`cotal ext seed` is the maintenance entry for that seeding (it runs automatically on the first real\ncommand of each boot, so you rarely call it). Each seeded connector's `\u2713 added` line goes to stderr,\nso the command that triggered the seed keeps stdout to itself:\n\n| Flag | Meaning |\n|---|---|\n| (none) | Reconcile: seed any never-seeded built-in, refresh a seeded one whose version the binary bumped, leave a removed one removed. A no-op once current. |\n| `--repair` | Recover after an interrupted seed or a lost authority (rebuilds the interrupted connector; restores the removed-vs-never-seeded record from its durable backup). |\n| `--reset` | Discard the record and re-seed all seven built-ins (the six connectors plus the web dashboard). **Resurrects any you removed.** Rebuilds cleanly over corrupt seed state. |\n| `--force` | Re-seed the built-ins even when the version stamp is current or a downgrade. |\n\nWhen a newer `cotal` advances the operator-global seed store to its generation, it prints one\nmigration line naming the old and new generations, the exact CLI entry that wrote the store, the\ncommit timestamp, and `seed/stamp.json`. That writer and timestamp are kept in the stamp, so a later\nolder CLI refusal can say which executable wrote the generation it will not overwrite and when.\nLegacy generation-only stamps remain readable; their refusal simply has no writer provenance to add.\n\nAn older `cotal` refuses a seed store written by a newer version. When it can verify a sufficient\n`cotal` executable on PATH or at the installer's `~/.local/bin/cotal` location, the refusal names\nthat absolute path so a reduced service PATH does not select the older binary again. Otherwise it\nkeeps the generic newer-version instruction. `--force` rebuilds the store for the running older\nversion without discarding the ever-seeded authority. `--reset` still exists for corrupt state and\nresurrects deliberately-removed connectors.\n\nA source-checkout CLI (`pnpm cotal`, `tsx bin/cotal.ts`, `node bin/cotal.ts`, or a suite child of\nthose) refuses to write or garbage-collect that store. The refusal names the path, the generation\nit declined, and `COTAL_SKIP_CONNECTOR_SEED=1` as the way to run other commands from a checkout,\nbecause pointing `$XDG_CONFIG_HOME` at a scratch dir alone does not lift it. With the skip set, a fresh\nconfig gets no built-in connectors (`cotal ext list` shows none), so add one from the checkout with\n`cotal ext add <checkout>/extensions/connector-claude-code` against that scratch `$XDG_CONFIG_HOME`. `COTAL_HOME` does not relocate this\nstore. An entry that cannot be proven as a released install is refused the same way. Isolated\nrelease tests that must seed from a checkout-shaped `bin/` set `COTAL_ALLOW_CHECKOUT_SEED=1` after\npointing `$XDG_CONFIG_HOME` at a scratch dir; that override is documented here, not on the refusal\nline. An opt-in write still records the checkout path in `seed/stamp.json` as `writtenBy`.\n\nThe default connector for a bare `cotal spawn` (no `--agent`) is the persona's `agent:` pin if it\nhas one, else `claude`; set `COTAL_DEFAULT_AGENT` (e.g. `opencode`) to change the fallback. It is\na default, so a persona that pins its harness still wins over it. An `--agent` naming a removed\nconnector fails loud with the exact\n`cotal ext add` to restore it. Set `COTAL_SKIP_CONNECTOR_SEED=1` to turn off the automatic first-run\nseed/refresh entirely (for a controlled or offline setup that manages connectors by hand); `cotal ext\nseed` still runs on request. `cotal agent-bearer` never takes the seed at all: it is exec'd by\nspawned seats on every bearer refresh, so it neither reconciles nor is refused by the store's\ngeneration (see [Plumbing](#plumbing)).\n\n## completion\n\n```bash\ncotal completion <bash|zsh|fish|powershell> # print a stub to eval / source\ncotal completion install [shell] # install it persistently\n```\n\nPrints or installs shell completion. Completion candidates come from each command's declared flags\nand, where useful, live mesh state (spaces, personas, managed agents) resolved offline.\n\n## feedback\n\n```bash\ncotal feedback \"<summary>\" [--type <t>] [--email <e>] [--details <text>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--type <t>` | none | `bug` \\| `idea` \\| `friction` \\| `praise` \\| `other` |\n| `--details <text>` | none | Longer free-form details |\n| `--severity <s>` | none | `low` \\| `medium` \\| `high` |\n| `--area <a>` | none | The part of Cotal this concerns |\n| `--email <e>` | git email | Contact email (required on the keyless public path) |\n| `--name <n>` | none | Your name (optional) |\n| `--url <url>` | keyed / public intake | Intake URL override |\n| `--key <k>` | `COTAL_FEEDBACK_KEY` | Feedback key |\n\nSends feedback to the Cotal developers. With a key (`--key` / `COTAL_FEEDBACK_KEY`) it routes to the\nkeyed beta intake; without one it goes to the public `cotal.ai` intake and requires a contact email\n(`--email` / `COTAL_FEEDBACK_EMAIL`, else your git email). Run a self-hosted intake with\n[`feedback-intake`](#server-daemons).\n\n## run\n\nOperate durable workflow runs (cotal-lang programs) from the terminal.\n\n```bash\ncotal run start --file <program> [--timeout <dur>] [--local]\ncotal run resume <runId> [--local --file <program>]\ncotal run ps [--endpoint <ep>] [--json]\ncotal run journal <runId> [--endpoint <ep>] [--json]\ncotal run answer <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]\ncotal run amend <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]\ncotal run migrate <runId> --local --file <program> [--endpoint <ep>]\n```\n\n`start` hands the program to the mesh's manager, which validates it, mints the run id (the record\nnever takes a caller-supplied one), drives it in its own process, and answers with the id once the\nrun is recorded; a program that does not validate is refused with every problem listed. `resume`\nasks the manager to take an existing run back and continue it from its step journal; the source is\nthe recorded program, so no `--file` is taken. Neither takes `--endpoint`: the manager records\nits runs under its own endpoint, and naming another is refused. `ps` lists the run records and\n`journal` renders one run's durable records; both only inspect. An open pause prints its question.\nA pause settled with an accepted answer prints its value as JSON plus the recorded answerer,\nartifact when present, time, and answer id, then one `amended` line per later amendment, in the\norder the store committed them, so the last is the current position.\nExpired pauses and ordinary steps print no answer line.\n`--json` on `ps` or `journal` prints each row the manager answers with (or `--local` reads) as one\nJSON object per line. A `ps` row carries `runId`, `endpoint`, `state`, `holder`, `epoch`,\n`journalHigh`, `forkedFrom`, `startedAt` and `programHash` (the values the program's `run()`\nreports; `programHash` is absent for a run with no recorded program), and `revoked` or\n`revocationUnreadable` when the marker says so. A `journal` row is an `activation` or a `step`. A\nstep row carries its `step` key, the `effect` kind and its `name`, `state`, `outcome`, the recorded\n`status` and `errorCode` once settled, and `startedAt` and `endedAt` in epoch milliseconds. An open\npause adds its `asks`, its `deadlineAt`, and for a checkpoint the `onExpiry` it was armed with; a\nsettled pause adds its `answer` and its `amendments`, as the text view prints them. A field the\njournal does not record is absent: a checkpoint opened before `onExpiry` was recorded carries none.\nThe run header and errors go to stderr, so stdout carries only rows; an unreadable revocation marker\nprints its reason there and still exits 1. The text view is presentation and is not a stable\nparsing target. `--json` on any other verb is refused.\n`answer` resolves an open\ncheckpoint through the manager, presenting as the holder that armed it; the manager records the\nanswerer from your credential, so no `--by` is taken there. A settled step refuses a second\n`answer`. `amend` records a changed position on a settled checkpoint or `ask`: it files a new\nanswer beside the accepted one, naming it, and the journal lists it under the step. The pause stays\nsettled and the run keeps the answer it acted on. A step that is still open or settled without an\nanswer refuses an amend. A spawned seat may amend only an answer recorded under its own name. `migrate` runs the migrate check of an\nedited program against a run's journal, from this terminal under a read credential (`--local`\nonly; the manager serves no run-migrate command): it prints whether the migration is admissible,\nevery orphaned step with its verdict and code, and exits 0 on admissible and non-zero on not. It\nwrites nothing: the commit that would file the migration is not reachable yet, and the report\nsays so. `--timeout` sets the default\ncheckpoint timeout for a drive (default 1h). `--local` drives in this process instead, over one\nconnection per invocation under the run's own credential minted from the project folder's trust\nmaterial, and is the path on a bare broker with no manager or for a run with no recorded program\n(`cotal run resume <runId> --local --file <program>`); `answer --local` and `amend --local` take\n`--by <who>`. On a\nuser-auth mesh the host's own manager refuses the family by name, and `--local` has no credential\nthere. A participant's manager started with `cotal supervise` hosts a logged-in user's runs through\nits issuing host: the auth callout issues the user's manager connection, and every `run` verb rides\nthe versioned rail under that issuance.\n[User-auth run start](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/user-auth-run-start.md)\nrecords the path. The guide is [workflows](workflows.md).\n\n## Server daemons\n\nTwo long-lived infra roles ship with the CLI. They are not part of everyday operation; the delivery\ndaemon comes up automatically with `cotal up --detach` in auth mode.\n\n```bash\ncotal deliver --space <s> [--server <url>] [--creds <file>] [--root <dir>]\ncotal auth-service --space <s> --server <url> [--port <n>] [--exchange-public-port <n>] [--exchange-public-url <https://\u2026>] [--exchange-trusted-proxy]\ncotal feedback-intake --keys <keys.json> [--port <n>] [--creds <file>]\n```\n\n`auth-service` runs a user-auth space's identity plane: the NATS auth callout, the\ncapability-gated local exchange and JWKS, and, when `--exchange-public-port` is set, the closed public\nexchange/discovery face forwarded by an HTTPS reverse proxy. `--exchange-public-url` is the proxy URL\nadvertised to clients; `--exchange-trusted-proxy` opts into last-hop `X-Forwarded-For` attribution.\n`cotal up --user-auth` starts and supervises the service for you, so you run it directly only to\nrecover one by hand.\n\n`deliver` runs the server-side Plane-3 delivery daemon: the durable backstop and membership/ACL\nauthority. It is auth-mode-only and single-instance (`--shard`/`--shards` accept only `N=1`);\n`--dev-mint` mints a scoped cred from the local signer for standalone dev. `--creds` can start a\ndaemon that already looks healthy, but production renewal is not that file alone: the manager and\nthe daemon must address one credential store. On a stock split host with two project roots, a\ndirect `deliver` is not an independent repair; keep the daemon under `cotal up` on the broker\nhost, or inject the same store into both processes ([embedding](embedding.md#supervisor-signing-authority)).\nTyped by hand on the workstation, `deliver` dials the broker recorded for `--space` in the mesh\nregistry (a mismatching `--server` is refused before any dial, and a record for a different\nworkspace root is refused outright); with no record for the space it falls back to the local mesh.\nThe daemon serves the workspace root that `--root <dir>` names, which must hold `.cotal/`, or else\nthe nearest `.cotal/` above its working directory. With neither, it refuses at start and names the\ndirectory it searched from, before it reads a credential or dials a broker.\nSee the [delivery daemon](delivery-daemon.md). `feedback-intake` runs a self-hosted feedback server\n(requires `--keys` and a scoped `--creds`), announcing submissions into a space channel; flags\ninclude `--host`/`--port`, `--store`, `--space`/`--channel`, `--max-bytes`, and `--rate-limit`.\n\n## Plumbing\n\n`cotal __complete <words\u2026>` is the internal entry the shell-completion stubs call to emit candidates\nfor the current command line; you never run it directly. `cotal agent-bearer` is machine-facing\nplumbing on user-auth meshes: spawned agents exec it to print a fresh short-lived bearer from their\nspawn-time secret; you never run it directly either. Its local arm uses `--dir` to discover the\ncapability-gated loopback service. A remotely enrolled, already-granted agent instead receives\n`--exchange-url <https://base>` in its launch argv: that arm sends `{owner, actor, actorToken}` to the\npinned public exchange with no local capability, follows no redirects, and refuses every non-HTTPS\nURL because the actor token is the credential in the request body. Because a seat execs it on every\nbearer refresh, it skips the connector-seed boot gate entirely: it reads one 0600 token file,\nexchanges it and prints the bearer without consulting or writing the operator-global seed store, so\na newer store generation cannot refuse a live seat's refresh. `--manager-call` asks for the\ninstance-bound `manager-caller` view; `--manager-instance <id>` selects an explicit live candidate.\nThat mode still prints only the raw token and does not update `--health-file`. A spawn runs it once as the agent auth\npreflight. When it fails there without printing a sentence of its own, the refusal names the cause: the 30 second\ntimeout, the signal that killed it, or its exit code. (`cotal start` is a removed tombstone: it\nerrors and points you to `cotal spawn --detach`.)\n"
|
|
57782
57782
|
},
|
|
57783
57783
|
{
|
|
57784
57784
|
"slug": "config",
|
|
@@ -57792,7 +57792,7 @@ function loadDocsBundle() {
|
|
|
57792
57792
|
"title": "Connect Claude",
|
|
57793
57793
|
"kind": "Guide (informative)",
|
|
57794
57794
|
"summary": "The Claude Code connector turns a real claude session into a Cotal mesh peer.",
|
|
57795
|
-
"body": "# Connect Claude\n\n> **Guide** (informative) \xB7 **For:** operators \xB7 **Prereqs:** [Quickstart](getting-started.md)\n\nThe Claude Code connector turns a real `claude` session into a Cotal mesh peer. A bundled\nplugin inside the session joins NATS, maps lifecycle hooks to presence, and exposes the\nmesh tools. Nothing wraps Claude; it is an ordinary session that happens to be on the\nmesh.\n\nThe shared mesh runtime (agent, `cotal_*` tools, hook relay) lives in\n[`@cotal-ai/connector-core`](../extensions/connector-core); this connector is the thin\nClaude-specific adapter over it. Its lifecycle hook imports the relay from the\n`@cotal-ai/connector-core/relay` subpath, so each hook process loads the relay and its environment\nreaders and none of the NATS client, zod or yaml. Siblings: [OpenCode](connect-opencode.md) (beta),\n[Hermes](connect-hermes.md) (alpha), [pi](connect-pi.md) (alpha); the\n[Connectors](connectors.md) matrix compares them feature-by-feature.\n\n## Set up\n\n```bash\ncotal setup # one-time: installs the plugin, seeds one agent; launches nothing\ncotal up # brings up the mesh + delivery daemon + a detached manager\n```\n\n`cotal setup` installs the cotal plugin (so the repo's Claude sessions get the `cotal_*`\ntools), shares your own MCP servers with spawned sessions on its first run (see\n[Sharing your MCP servers](#sharing-your-mcp-servers)), and seeds one `default` persona; `cotal up` brings up the local stack so\n`cotal spawn --detach` / `cotal_spawn` work right away. Re-running either is idempotent.\nThe install mechanics and the invariants behind them are in\n[setup internals](setup-internals.md).\n\n`cotal setup` also installs Cotal's authored Agent Skills (`SKILL.md`, the agentskills.io format) for\ncoordinating agent teams (today `team-topology`), from one canonical source, on two channels:\n\n- **Claude Code** gets a second, skills-only plugin, `cotal-skills`, from the same `cotal-mesh`\n marketplace, at **user scope** (machine-wide). The Claude connector declares and implements this\n setup provider, including the marketplace assets and native plugin commands; the base CLI only passes\n the vendor-neutral Agent Skills directory. The plugin carries no code and no core dependency,\n and uninstalls on its own with `claude plugin uninstall cotal-skills --scope user`. Its plugin version\n is stamped from the running CLI release, so an upgrade + `cotal setup --skills` runs `claude plugin update` and\n the deployed install actually gets the new skill. `cotal setup` installs it on first run and on repeat\n runs, so upgraders are not left behind. The same provider reports the plugin and skills plugin rows\n in `cotal status`, which point a stale or missing skills plugin at `cotal setup --skills`.\n- **Every other harness** (Codex, Cursor, OpenCode, Gemini CLI, Windsurf/Devin) reads the cross-vendor\n `~/.agents/skills/` directory convention, which has no remote index, so `cotal setup` **reconciles** it\n (and `cotal setup --skills` does only that):\n it installs/updates each Cotal skill, backs up a copy you have edited to `SKILL.md.bak` before\n replacing it, and removes a Cotal skill that is no longer shipped. Only skills Cotal owns are touched;\n your own or third-party skills there are left alone. `cotal status` reports whether the drop is current,\n stale, missing, or has a retired skill to reconcile, and names `cotal setup --skills` as the remedy. This is the working cross-vendor path.\n\nCotal also generates an [Agent Skills discovery index](https://cotal.ai/.well-known/agent-skills/index.json)\non cotal.ai, but that RFC is still a draft with no harness consuming it yet, so it is a forward bet,\nnot a channel to rely on today.\n\n## Spawn a session\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn dave --detach # supervised: the manager runs it in a PTY\n```\n\nA spawn resolves a persona from `.cotal/agents/<name>.md` ([agent files](agent-files.md));\n`--model`, `--variant`, `--cwd`, `--prompt`, ACL overrides, and `--share-tools` apply to\nboth forms ([run a mesh](run-a-mesh.md) has the full resolution rules). The session joins\nwith identity from its environment and auto-registers presence by the time it is\ninteractive.\n\nInside the session, the agent orients with one read-only tool, `cotal_orientation`: its\nidentity, the channels it reads and may post to, its capabilities, the tools available,\nwho's present, and unread counts. The full tool surface is the\n[MCP tool catalog](mcp-tools.md). In auth mode the team-supervision tools\n(`cotal_spawn` / `cotal_persona` / `cotal_personas`) are injected **only** for personas declaring\n`capabilities: [spawn]` (the same grant that opens the privileged control subject), so an\nagent's toolset matches its declared capabilities. `cotal_run` is gated separately by\n`run`; use `capabilities: [spawn, run]` for both. Fresh setup defaults include both.\nSee [workflow tool setup](workflows.md#from-an-agent-session) for a first run and missing-tool checks.\nClearing retained history is\noperator-only ([run a mesh](run-a-mesh.md)), never an agent tool.\n\n## How it binds\n\nClaude Code exposes four integration surfaces, and three of them collapse into a single\ndual-purpose MCP server:\n\n| Surface | Mechanism |\n|---|---|\n| Outbound, ambient | `http` lifecycle hooks \u2192 POST to the connector (presence, activity) |\n| Outbound, deliberate | MCP tools `cotal_send` / `cotal_dm` / `cotal_anycast` (+ `cotal_feedback`) |\n| Inbound, pull | MCP tool `cotal_inbox` (same server) |\n| Inbound, push | Channel nudge + hook drain (below) |\n\nThe manager launches the *real* `claude` (no wrapper):\n\n```\nclaude --strict-mcp-config --mcp-config '{\"mcpServers\":{\"cotal\":{\u2026}}}' \\\n --dangerously-load-development-channels server:cotal\n# env: COTAL_SPACE, COTAL_NAME, COTAL_ROLE, COTAL_CHANNEL=1, plus claude's documented auth vars\n```\n\n- **Model auth.** Locally, `claude` still reads macOS Keychain / `~/.claude`. In a container or\n CI there is no Keychain, so the connector forwards the documented credential set:\n `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`), `ANTHROPIC_API_KEY` /\n `ANTHROPIC_AUTH_TOKEN`, and the cloud-provider flags plus their credential vars. Host-session\n markers (`CLAUDE_CODE_CHILD_SESSION`, `CLAUDECODE`) stay out so a nested seat still saves a\n transcript. See [Deploy](deploy.md).\n- **Persona privacy.** The persona body is written to a private file and Claude receives only\n `--append-system-prompt-file <path>`. The body never appears in the spawned process argv. The\n carrier is a 0600 file inside a 0700 directory on POSIX, with equivalent owner-only ACL hardening\n on Windows. That is OS-user isolation: any process running as your user can read it while it\n exists. The manager or the foreground `cotal spawn` removes it, and the shared-server MCP config\n file, once it has proved the `claude` process gone. If the launcher is killed first, a watcher\n started beside `claude` removes them when `claude` exits.\n- **MCP servers.** `--strict-mcp-config` ignores every ambient MCP source, so a spawned agent\n loads the cotal server plus the servers the cotal config shares. First-run `cotal setup`\n fills that list with your own user-scope servers, so a spawned session has the tools you know\n (see below).\n- **Installed plugin.** The plugin is installed once (`claude plugin install\n cotal@cotal-mesh --scope local`) because its hooks bind only to an *installed* plugin.\n The repo's `.claude-plugin/marketplace.json` lists the committed plugin tree under\n `claude-plugin/`, which each release regenerates with the built bundles, the skills and the\n release version, so an install from the repo or from a pinned commit runs without a build\n ([Release](release.md)). `cotal setup` (npx, no clone) materializes the same marketplace under\n `~/.cotal/claude-plugin/` from the installed CLI (each plugin dir is rebuilt from scratch and\n atomically replaced, never merged, so no stale file rides in). The\n `cotal-skills` plugin installs from that same marketplace at user scope (`claude plugin install\n cotal-skills@cotal-mesh --scope user`); its manifest and install behavior ship inside the Claude connector, and\n its version tracks the CLI release so updates land.\n- **Identity-gated.** Connector code requires `COTAL_NAME`, `COTAL_LINK` or `COTAL_AGENT_FILE`.\n A plain `claude` with none of them never joins, so your own sessions in a repo do not appear\n as stray peers. Its MCP server still answers `initialize` and lists one static tool,\n `cotal_how_to_join`, which explains how to launch a session on a mesh. It builds no mesh\n agent, opens no broker connection and binds no control socket.\n- **Hands-free.** The dev-channels flag prints a one-time confirm prompt. The runtime waits for the\n dialog title in normalized terminal text and presses Enter once when it appears, so startup speed\n does not affect a supervised launch. The PTY runtime reads the child's output, and the tmux, cmux,\n Orca and Herdr runtimes read the pane's screen. If the declared prompt never appears within 15\n seconds, the seat ends with a bounded error naming the unmatched prompt instead of hanging\n silently or answering another dialog. The PTY runtime writes that error to the seat's output, and\n the other runtimes write it to the manager's log.\n- **Trusted directory.** Claude opens a directory it has not trusted on its workspace-trust dialog,\n and the dialog's default answer exits. No one is at a supervised seat to answer it, so a launch\n whose directory the manager host's own Claude does not trust is refused before it starts, naming\n the directory and the dialog. Trust is read as Claude reads it: trust given to a parent directory\n counts up to the root of the directory's own Git repository, and a linked worktree shares the trust\n of its repository's main checkout. Open `claude` in that directory on the manager host once and\n trust it, then spawn again. A foreground `cotal spawn` shows the dialog in your own terminal instead.\n\nInbound mesh messages arrive in context as\n`<channel source=\"cotal\" from=\"bob\" kind=\"dm\" \u2026>\u2026</channel>`: each meta key a tag\nattribute the agent can read for routing.\n\n## How messages reach the session\n\nDurable deliveries land in the connector's inbox from JetStream consumers\n([SPEC \xA78](../SPEC.md#8-nats--jetstream-binding)); live channel traffic can instead arrive\nthrough an at-most-once core subscription. A durable message sent while the agent is busy\nor offline waits on the stream. Two things move a message from inbox to model; one\ndelivers, the other only wakes:\n\n- **Hook drain (delivery).** `SessionStart` / `UserPromptSubmit` hooks read automatic inbox items and\n inject them as `additionalContext`. This is the single authoritative path: deterministic and works\n on any Claude Code build. Quiet ambient is excluded and stays buffered for `cotal_inbox`.\n A message is **acked only once the hook reply carrying it has cleared both legs of its journey**:\n the connector's control socket to the hook process (which gives up after 2s), and the hook\n process's own stdout to Claude Code (which it force-exits 1s after starting to write). The relay\n sends a receipt back down the control socket from that stdout write's callback, and only on a\n clean write (a runtime whose pipe has gone away fails it), and the connector treats that receipt,\n not its own socket write, as delivery. So a large injection killed mid-flush, or one written to a\n broken pipe, leaves the message un-acked and JetStream redelivers it. What this does *not* prove is\n that Claude Code read or applied the reply: a payload small enough to fit the pipe buffer is\n reported written the moment the kernel takes it. That residual is why the path errs toward\n at-least-once rather than treating a confirmed write as a confirmed read. Acking when\n the reply was merely *formatted* meant a lost reply was a lost message: it was already marked\n handled, so its own redelivery was silently acked on arrival.\n A hook whose handler throws still returns an empty reply so the session is never blocked, and\n that reply carries nothing, so it commits nothing: the batch it had started to surface stays\n un-acked and goes out on a later frame. The seat also drops any `turn-pending` row that breaks\n the manager contract, such as one with no integer deadline, and says so once in its log. A reply\n with no `turns` array changes nothing: the seat keeps the turns it already holds.\n This errs toward **at-least-once**: if a reply lands but its confirmation does not, the batch is\n surfaced again and flagged as a possible repeat. A duplicate injection is noise; a buried DM stops\n the peer answering at all.\n- **Channel nudge (wake).** An arriving message fires a `notifications/claude/channel`\n event that wakes an *idle* session into a turn, so the drain runs *now* instead of at\n the next prompt. The nudge never acks anything. A nudge that the host rejects is retried with a\n bounded backoff while anything is still pending. For an idle session it is the only wake source,\n so dropping it means silence until someone types. When the channel becomes active, the connector\n first re-fires a focus mention remembered during startup, otherwise one buffered wake. A rejected\n push keeps its bounded retry, and JetStream redelivery remains the durable backstop for unacked\n inbox items. Neither a redelivery nor that retry repeats a nudge already pushed for that message,\n whether the message had its own nudge or was counted in a batch one, so a session held in a long\n tool call gets one nudge per message. Once a hook frame carries the message, or the push that\n announced it fails, its next redelivery nudges again, so a reply that never reached Claude Code\n still recovers. If the channel cannot run at all, delivery still waits for the next hook. Live-only\n traffic has no durable retry.\n\n**Two priority tiers.** A *directed* message (DM, anycast, or a channel message that\n`@mentions` us) always nudges. *Ambient* channel chatter does not nudge mid-turn; it\naccumulates, and the `Stop` \u2192 idle transition fires one batch nudge so the backlog drains\ntogether.\n\n**Constraints (accepted).** Channels are a Claude Code research preview (\u2265 v2.1.80;\npermission relay \u2265 v2.1.81): Anthropic auth only, admin-enabled on Team/Enterprise, and a\ncustom channel needs the `--dangerously-load-development-channels` launch flag. The hook\ndrain does not depend on any of that; the channel only adds \"wake me when idle.\"\n\nThe same channel also relays **tool-permission requests** onto the mesh, so a peer (a\nhuman at the CLI, a policy node) can approve or deny an agent's pending tool call through\nCotal rather than a per-terminal prompt.\n\n### Attention\n\nAn agent picks how aggressively peer traffic reaches it with\n`cotal_status({ attention })` (three modes, orthogonal to presence):\n\n| arrival | open (default) | dnd | focus |\n|---|---|---|---|\n| directed (dm / anycast) | wake + inject | wake + inject | wake + inject |\n| channel `@mention` | wake + inject | wake + inject | ack-drop; wake to *pull*; not injected |\n| ambient channel chatter | wake when idle; hold while working | never wakes; injects next turn | ack-drop; recall via `cotal_inbox` |\n\nPer-channel overrides refine this: **quiet** (delivered, never wakes; `@mention` still\nwakes) and **muted** (dropped on receive, mentions included; DMs/anycast unaffected), set\nwith `cotal_channel_mode` or as agent-file defaults (`quiet:` / `muted:`,\n[agent files](agent-files.md)). A per-channel override is the final word for that channel.\nQuiet ambient is pull-only: it never hitchhikes on a human prompt, DM, mention, or other\nconnector-driven turn. `cotal_inbox` explicitly surfaces and clears it. A quiet-channel\n`@mention` remains automatic and injects normally.\n\nA pull is bounded too, and clears only what it hands over. One `cotal_inbox` call carries at most a\nreceivable window (direct messages and role requests first, then channel traffic, replayed history\nlast); whatever does not fit stays buffered, is named in the reply, and comes back on the next call.\nA message too large for one whole response is delivered in parts: once no smaller mail is waiting,\neach call carries its next part, and it is cleared only after its last part goes out, because clearing\nwhat was not handed over is the loss this bound exists to stop.\nThat matters most on the path where it is easiest to lose mail: reconnecting brings a channel-history\nreplay with it, so the largest payload and the least expendable message arrive in the same read.\n\nThe local inbox is bounded. On pathological overflow it evicts pull-only items first, then other\nchannel traffic, and a direct message or role request only when the whole buffer is directed mail.\nAn evicted channel item is acknowledged. An evicted direct message or role request never is, and\nthe broker redelivers it after the ack wait until a redelivery finds room. A direct message stays\npending on the session's DM durable, where `cotal deliver pending <name>` counts it. A role request\nstays on its role's shared queue, which that command does not read. A full inbox therefore delays\ndirected mail until the session drains it.\nIf the bounded live/durable classification guard also fills, the connector fails closed:\notherwise-normal ambient becomes pull-only until restart. Muted hard-drop and normal focus recall\nstill take precedence. Focus also keeps a bounded exclusion list so mode toggles cannot recall\nquiet/muted traffic; if that safety bound fills, recall skips the affected channel and reports it\nas incomplete rather than risk resurfacing excluded content. Recall cannot tell one message with an\nempty id from an identical one with another disposition, so in focus such a message is held in the\nlocal inbox as pull-only instead of being dropped, and a mention of it still wakes the agent. When\nthe session settles an id-less copy, it reads the chat stream's last sequence. Identical copies\narrive in stream order, and those reads can answer out of order, so the first read that can see the\ncopies binds them in arrival order, latest first, each to the latest unbound stream copy at or below\nthe lowest sequence read for it or any later identical copy.\nWhile the connection stays up, every stream copy at or below that sequence reached the session\nfirst, so a later identical copy sent during a reconnect gap is above it and stays unbound. A copy\nthat arrives while a read runs may not be in that read, so it binds nothing there and recall reads\nthe channel again, up to three times; if it still could be a stream copy in the last read, recall\nleaves that stream copy in the stream and reports the channel as incomplete. A settled copy a\ncomplete read cannot bind is behind the focus start or out of retention, and is forgotten. Recall\nhands back into the inbox only the stream copies nothing is bound to, such as one sent during a\nreconnect gap or one the inbox evicted on overflow, and `cotal_inbox` hands each over once. Overflow\nfrees only the evicted copy, held or quiet, and a copy no read has bound yet keeps its place in\narrival order, so an identical muted copy stays out of recall. When\nthe inbox is full, recall leaves them in the stream for a later call and reports the channel as\nincomplete. A history read that fails, or a channel with replay off, settles nothing and is reported\nas incomplete, and recall calls run one at a time. If the sequence read for a settled copy fails,\nor answers only after the connection dropped, recall skips that channel for the rest of the focus\nperiod and reports it as incomplete. Recall cannot tell a late copy of a message it handed back from\na new identical message, so every copy takes its own disposition: a new identical quiet mention is\nstill delivered automatically, and a late copy can surface a second time.\nA recalled message that already went out in part is read to its last part, even if an exclusion\nlands after its first part. One session reads its inbox one call at a time: a `cotal_inbox` call\nthat overlaps another waits for it to finish, so neither decides from a view the other has already\nmoved past.\nIf the separate hard-drop disposition guard fills, channel traffic is dropped for the rest of the\nsession rather than risk a late copy bypassing an earlier muted/focus decision; DMs and anycast are\nunaffected.\n\nAttention is **advisory UX, not a boundary**: any peer can wake a dnd/focus agent by\nnaming it, and `muted` means \"I opted out of receiving\", not \"the channel is blocked\";\nthe broker still authorizes and delivers. Focus's real effect is shrinking the\nuntrusted-ambient injection surface (only subject-authenticated dm/anycast auto-inject).\nIt resets to **open** on `SessionStart`, so a restarted agent never stays silently deaf.\nYour attention is mirrored into presence so peers can see it.\n\nWhatever does reach a turn is framed so a peer cannot write the frame. A line that begins at column\nzero is written by the connector; one message is one line plus indented continuations, with the\nsender inside a single bracket pair. A message body, a sender name and role, and a service or\nchannel label are all peer-controlled, so each passes through the same neutralization the\n`cotal_inbox` reply uses: no line break a splitter may honour and no bracket survives into a\nrendered attribution. This matters more for an injected block than for a reply, because the agent\ndid not ask for it and so never had the chance to distrust it.\n\n## Presence mapping\n\nThe connector wires a small subset of Claude Code hooks to presence states; presence is\ncoarse, and \"what it is doing\" rides on activity updates. Presence is **advisory**: a presence\npublish that fails (the endpoint mid-reconnect, say) is swallowed and never prevents the same hook\nfrom delivering messages or flushing held ones.\nA `SessionStart` during an open turn, including compaction, preserves the current `working` or\n`waiting` status until `Stop`, `StopFailure`, or `SessionEnd` closes the turn.\n\n| Hook | \u2192 state |\n|---|---|\n| `SessionStart` | `idle` only when no turn is open (join; surfaces the inbox; captures the live model into `meta.model` when no pin) |\n| `UserPromptSubmit` | `working` (turn starts; surfaces the inbox) |\n| `PreToolUse` | no change; records *what* is about to run, so a permission wait can name it |\n| `Notification` (`permission_prompt` / `agent_needs_input`) | `waiting` with condition `approval` / `input` (activity leads with the pending tool, e.g. `Bash: git push \u2026`) |\n| `Stop` / `StopFailure` | `idle` (turn done / died on an API error; flushes anything held while busy). `StopFailure` also relays Claude Code's native error value as `condition.source` and maps it to the closed condition vocabulary. On the [event plane](#event-plane) it closes the run with `RUN_ERROR`. |\n| `SessionEnd` | `offline` (graceful leave) |\n\nThe connector also leaves gracefully when its stdin closes. An MCP client closes it to end the\nsession, and a killed `claude` closes it with no `SessionEnd`, so a dead session drops off the\nroster instead of staying on it as a live peer.\n\n`StopFailure` maps `rate_limit` and `overloaded` directly; auth and credential failures to\n`auth`; account and billing failures to `billing`; `invalid_request` to `request`;\n`model_not_found` to `model`; `server_error` to `server`; `max_output_tokens` to `context`; and\n`unknown` to `failed`. The native value remains in `condition.source`.\n\nHooks are relayed over the connector's **authenticated** local control endpoint (per-user\nsocket + per-launch token, constant-time checked), so a local process that finds the path\nstill can't drive presence or stop the agent. The full Claude Code hook-event list lives\nwith the adapter:\n[`extensions/connector-claude-code`](../extensions/connector-claude-code/README.md).\n\n## Event plane\n\nA spawned session publishes a **structured** account of what it\ndid: run boundaries per turn, assistant text, reasoning, and each tool call with its start\nand its end. Not prose about the work, the work itself, in a vocabulary a program can\nread. The launcher sets `COTAL_EVENTS` by default; pass `--no-events` to opt out on an unrestricted\nspace. A user-auth registration with `policy: { events: \"required\" }` carries `eventsRequired` in the\nprivate launch material, so the connector arms even without `COTAL_EVENTS`; `--no-events` is refused.\nA hand-driven user-mode session may carry the same decision as `COTAL_EVENTS_REQUIRED=1`. Its own\npublish grant must cover `events.<owner>.<actor>` or the connector refuses before joining. An unmanaged\nsession with no launch material and no required-policy fallback keeps the generic default behavior.\n\nIf the event plane stops for good, the space's policy decides what happens to the seat, on every\nconnector. On a space that requires events the seat stops and leaves the mesh. On any other space\nit keeps running without events, and the connector log records `AG-UI emitter stopped` with the\nreason. For Claude Code the connector is the MCP server: it leaves the mesh and exits with code 1,\nand its stderr carries that line.\n\nA new session includes its first run even when Claude writes a positional startup prompt before the\nconnector receives `SessionStart`. That from-zero read is keyed only to Claude's explicit\n`source: \"startup\"`; resumed, forked, cleared, and compacted sessions adopt at the transcript boundary\ncaptured at that adopt, before the mesh link connects, so nothing Claude appends while the connector\nis still starting up lands behind the cursor and is silently dropped. Crash recovery follows the\ncursor already stored in the event write-ahead log, regardless of the new process's startup label.\n\nClaude starts each hook in its own process, so a prompt or stop relay can reach Cotal before the\n`SessionStart` relay. The connector holds those event flushes and the terminal until `SessionStart`\nsupplies the source, then enqueues adopt, flush, and close in that order.\n\n`SessionStart` can also run before the connector process has bound its local control socket. The hook\nthe `SessionStart` relay retries only transient pre-connect listener errors, with capped backoff\ninside its existing two-second budget. Later hooks and permanent local faults still fail open\nimmediately. Once a socket has connected, a broken exchange is not retried: the connector may\nalready have handled the frame, so replaying it could apply one lifecycle event twice.\nThat retained `SessionStart` can itself arrive before Claude creates the transcript path. A genuinely\nnew startup waits up to five seconds for that file with capped backoff, and the same deadline bounds\none stalled file read; expiry fails loud instead of silently losing the first run. A forked session\ngets the same wait, because Claude copies the parent transcript into the fork's own file after the\nhook, and then adopts at the end of that copy. Resumed, cleared and compacted starts and recovered\ncursors still require their existing source at once.\n\nTool arguments (`TOOL_CALL_ARGS`) and tool results (`TOOL_CALL_RESULT`) are not republished\nonto this channel. The durable emitter drops those events before they are written to the\nwrite-ahead log, because this channel's read ACL is not the ACL the tool ran under. Content is\nmandatory on both kinds, so the event is suppressed rather than emptied or replaced with a\nplaceholder. Tool start and end still go out. A restart that finds a pending pre-fix frame\nstill carrying those kinds HALTS rather than republishing it.\n\nThe channel is **`events.<owner>.<actor>`**, named after the session's principal. What the actor\nhalf is depends on the mesh, and the difference matters when you go looking for it: on a static mesh\nit is a key the manager allocated, never the display name, so two live agents sharing a display name\ndo not share a stream; on a user-auth mesh it is the agent's own name, because that is what the\nledger row is keyed on. Spelled out again with both halves below. The launch grants publish rights\non that channel alone. A spawn\nthat asks for a *different* agent's event channel is refused at the door rather than granted, since\nthat channel is that session's event stream. The same rule runs on restart: a manager\nresume document that names another agent's event channel is refused rather than adopted, because the\nmanaged row is re-armed from that document and the credential is re-minted from the row.\n\nThe rule reads a **concrete** channel, two principal tokens and nothing else. A pattern such as\n`events.<owner>.>` is not an event channel to it and passes untouched, governed by ordinary ACL\nauthority: on a user mesh the delegation envelope, on a static mesh the spawning credential itself.\nThat is deliberate, because the pattern is the form an operator writes on purpose for an observer,\nand it is worth knowing rather than assuming the fence is total.\n\nTo let something else read a plane, grant it out of band. The refusal prints the command for the\nmesh it is running on, spelled out in full, and only that one.\n\nOn a **user-auth** mesh:\n\n```bash\ncotal actor grant <reader> --owner <owner> --scope '' --allow-subscribe 'events.<owner>.<actor>' --allow-publish ''\n```\n\nEvery field, deliberately. `actor grant` is an upsert of the whole row, so it refuses a grant that\nleaves off any of the three ACL flags. Only `--full` turns an omitted flag into the wide default\n(`>` read, `>` post, `spawn,role:default` scope), which is the opposite of what a scoped watcher is for.\n\nOn a **static** mesh there is no actor ledger for `actor grant` to write to, and the refusal says\nso; mint the reader instead:\n\n```bash\ncotal mint watcher --profile agent --allow-subscribe 'events.<owner>.<actor>' --provision\n```\n\nThe **agent** profile, not the observer one. `mint` reads `--allow-subscribe` only for that\nprofile, and refuses it anywhere else: `--profile observer --allow-subscribe <channel>` exits\nnon-zero and writes no creds file, because the observer profile carries a fixed read set over the\nwhole chat plane, which is the opposite of what a scoped watcher is for. The agent profile also prints the lifecycle uid the\nreader needs, since an authed consuming endpoint refuses to start without one.\n\nOn an **open** mesh there is nothing to grant: the mesh has no credentials and no ACLs, so any peer\nthat lists the channel reads it, and the refusal says so instead of naming a command. The\nown-channel rule still applies there, because a spawn is not the place to hand out a read on\nanother agent's tool inputs and outputs.\n\nTwo things a reader has to do that are not obvious, both on `CotalEndpoint`. It must pass the event\nchannel in `channels`: an endpoint reads the channels it lists, so one constructed without\nthe event channel joins nothing and the frames never arrive. And it reads history with `readHistory(channel)`, the delivery daemon's mediated read, not\n`channelHistory(channel)`: a scoped credential is denied the ad-hoc consumer the direct read\ncreates, by design. `cotal console` and the web console already do both.\n\nThe `<owner>.<actor>` pair is the session's principal. On a user-auth mesh the actor half is the\nagent's own name, so the channel is `events.<your-owner>.<agent-name>`. On a\nstatic mesh the owner half is the literal `local` and the actor is a key the manager allocated, so\nthe channel is `events.local.<key>`; the spawn reply carries that key as `id`. Note\nthat `cotal console` and the web console keep event channels out of their channel lists on purpose,\nsince a plane is a machine feed rather than a conversation; they draw the frames when you open the\nchannel by name.\n\nThe rule governs the manager's doors, which are the ones a caller other than you can reach. A\nforeground `cotal spawn` on your own machine mints from your own signing material, so it can still\ngrant any channel you name: that is the out-of-band grant, not a way around the rule.\n\n**Failed turns publish run errors.** Claude Code decides for itself\nwhether a turn finished or died and fires one of two hooks accordingly, so the connector relays that\ndecision rather than making one of its own: a turn that ended on an API error ends its run with\n`RUN_ERROR` carrying the fixed message `run failed` and no code. Neither the detail Claude Code\nreported nor its error kind is published there: both are upstream values that can echo your prompt or\ntool output, and the events channel has a different read ACL. The error kind still reaches presence\nas the agent's condition (`rate_limit`, `auth`, `billing` and the rest). A turn that ended normally still\nends with a run-finished event carrying no outcome, which says the turn ended and does not claim it\nsucceeded.\n\nEvents are written to a per-session write-ahead log before they are published, so a hook that fires\nafter a restart resumes at the cursor it left rather than replaying or skipping, and a run that was\nopen when the session stopped is closed rather than left dangling.\n\nOne channel carries **every session of one agent**, because it is named after the principal and not\nafter the session. Alongside the per-session logs the connector keeps one small record per principal,\nholding the last sequence the broker assigned on that channel, so a new session continues the stream\nits predecessor left instead of starting again from nothing. Both live under the events state root\n(`COTAL_WORKSPACE_ROOT`), and neither is something you edit by hand.\n\nA **missing** record is not a fault: the connector rebuilds it from the session logs beside it,\nwhich is how an agent that was already running before this record existed keeps its stream. That\nrebuild stops if any one of those session logs is damaged. Unreadable, not valid JSON, and written\nfor a different principal all count, and so does a session directory or a log that is a link rather\nthan the real file the connector wrote, or a log that has more than one name. A tip taken from the\nrest would be too low, and it would stop publication later with nothing left to point at the cause.\nThe connector names the file instead, and the only way past it is the directory removal described\nbelow, under the same condition. A record that **disagrees with the broker** is a fault, and the\nconnector stops publishing and says why rather than guessing. A record that **moved while a session\nwas writing to it** is refused the same way: it means something else wrote the principal's record,\nand the connector reports which value it held and which the file holds rather than writing over the\nlater one. There is no command to clear it. The state is the principal's directory under the events\nroot, and clearing it by hand means removing that directory whole: the sequence, the cursor and the\nper-session logs only mean anything together, so removing part of it leaves a state the next start\nrefuses. Removing it is only half a remedy, and the half that comes first is the channel. The\ndirectory is where the agent's memory of the tip lives, not the tip itself, so on a channel that\nstill holds frames the next session opens expecting an empty one and stops on the same\ndisagreement, with the logs a tip could have been rebuilt from now gone. Purge the channel first,\nthen remove the directory.\n\nReading it: `cotal console` and the web console draw event frames directly. A frame carries no text\npart by design, so a surface that renders a message as flat text shows a marker instead of prose.\n\n**On a per-user-auth mesh, the default event plane needs the spawner's grant to cover the channel.** The event\nchannel is added to the child's publish set, and delegation only narrows: an agent may hand down\na subset of what it holds and no more. So a peer-initiated spawn is refused unless the\nspawning identity's own grant already covers the child's event channel. The refusal prints the\nexact `cotal actor grant` command that widens it. An operator launch, whose chain reaches an\nadmin-scoped or roster row, is unaffected. Passing `events: false` is the explicit opt-out.\n\nArming the event plane through a typed spawn request (`manager.spawn` with `events`, including\nthe CLI's `cotal spawn --detach --events`) additionally requires the caller's admin tier on a\nuser mesh. A non-admin caller that asks for the plane is refused before anything is provisioned,\nand one that stays silent gets a spawn without it, with the reply saying so.\n\n## Resume a session\n\n`--resume <session-id>` pulls an existing Claude session, its context and transcript,\ninto the mesh. It **forks**: Claude mints a *new* session id from that transcript\n(`--resume <id> --fork-session`), so the meshed agent gets its own session and the\noriginal is untouched.\n\n- `cotal spawn --resume <id>` (foreground) is the primary surface: the transcript is on\n *your* machine, and errors are Claude's own stderr, inline.\n- `--detach --resume <id> --on <instance>` carries a session held on *your* machine to that\n manager instance, which may run on another host. The CLI finds the transcript under your\n Claude config (`~/.claude`, or `$CLAUDE_CONFIG_DIR`), sends it through a JetStream Object\n Store bucket only that instance reads, under a writer credential pinned to that one transcript,\n and prints `carried session <id> to <instance>:\n sha256:<hex>, <sent> of <size> bytes sent in <chunks> chunks`. A re-run of the same bytes\n sends nothing, and an interrupted carry continues where it stopped. The seat forks it in a\n private Claude home under the manager's `.cotal/seat-homes/`, which no other seat's Claude\n lists or finds, and which is removed when the seat stops. When Claude starts the fork, the seat\n records the SHA-256 of the transcript it read from its own project; the manager stops a seat\n whose record names other bytes than the carried ones, or that records none within the join\n timeout after it joins, an uncertain launch included, and otherwise shows that record as the\n seat's provenance. `cotal attach` to such a seat names its source after the seat\n name, as `(resumed from <host>:<id>)`. A remote manager receives a carry when its host issues it a\n transfer reader. On a user-auth mesh the CLI exchanges the operator's login for a one-object\n `transfer-writer` view, which needs scope `admin`.\n- A session name in place of an id is refused, listing each session on this host that carries\n that name with its id, SHA-256 and modification time. An id this host does not hold resolves\n against the **manager host's** `~/.claude`, as before.\n- A seat-private home holds no login. The manager host needs `CLAUDE_CODE_OAUTH_TOKEN` (from\n `claude setup-token`), `ANTHROPIC_AUTH_TOKEN`, or a cloud provider selection in its\n environment; `ANTHROPIC_API_KEY` alone is refused. The launch directory must already be\n trusted by the manager host's own Claude, and Claude must be 2.1.234 or later.\n- The manager waits for a real outcome: `\u2713 started` means the agent *joined the mesh*,\n `\u2717 exited on launch` carries Claude's last output, and an uncertain launch (~30 s) is\n reported without tearing the agent down.\n- Resume is an **operator surface only**, deliberately not exposed on MCP `cotal_spawn`\n (a mesh peer naming host-local transcripts would widen `spawn` into transcript\n disclosure). Only the Claude connector supports it today; OpenCode and Hermes fail loud.\n- Needs a `claude` new enough for `--resume \u2026 --fork-session` (verified on 2.1.197).\n\n## Sharing your MCP servers\n\nA spawned session keeps your own MCP servers by default. On its first run, `cotal setup` copies\nthe user-scope servers from your Claude Code config (`~/.claude.json`, or the one under\n`$CLAUDE_CONFIG_DIR`) into the cotal config file (`~/.config/cotal/config.json`) under\n`connectors.claude.mcpServers`, and names them in its output. With none to copy it writes an\nempty list. Each entry is the familiar `.mcp.json` shape ([full format](config.md)). A cotal\nconfig that already declares that list keeps it, and a later `cotal setup` never changes it.\n\nThe cotal config holds secrets only as `${VAR}` references. Setup cannot tell literal text from\na secret, so it leaves out a server with an `env` or `headers` value that is anything but `${VAR}`\nreferences (a `Bearer ${TOKEN}` header among them) and names it in its output. To share one,\nadd it to the cotal config with each secret written as a `${VAR}` reference, and export that\nvariable where you spawn. Setup also leaves out and names an entry no session can start, such as\none with a missing or empty `command` or `url`, or one whose `command` is not a string.\n\nAt launch the connector forwards *only* the named vars the chosen servers declare and\npasses the merged config as an owner-only temp file; `--strict-mcp-config` stays on, so\nonly cotal + the shared servers load.\n\nFor a lighter seat, share fewer. Remove an entry from the cotal config to drop it from every\nspawn, or scope one spawn with `--share-tools tavily,figma` (or `--share-tools none` for cotal\nalone). An empty list (`\"mcpServers\": {}`) in `~/.config/cotal/config.json` keeps every spawn\nisolated, and setup leaves it as it is.\n\nTwo caveats: sharing a server grants its credential to the agent (the var lives in the\nClaude process's environment, so share only when you're fine with that teammate holding\nthe key), and memory adds up, because a heavy server boots once per spawn, multiplied\nacross a team, and can starve a small machine.\n\n## Feedback\n\n`cotal_feedback` works out of the box: without a key it posts to the public intake at\n`https://cotal.ai/v1/feedback` (needs a contact email: `COTAL_FEEDBACK_EMAIL`, then\n`git config user.email`, else the agent asks). Set `COTAL_FEEDBACK_KEY=fbk_<key>` in a\nbeta tester's environment to route to the keyed intake (`Authorization: Bearer`, identity\nderived from the key); `COTAL_FEEDBACK_URL` overrides either endpoint. The CLI can send\ntoo: `cotal feedback \"<summary>\" [--type bug]`. Each submission carries\n`origin: human | agent`, whether the tester asked, or the agent auto-reported a major\nissue.\n"
|
|
57795
|
+
"body": "# Connect Claude\n\n> **Guide** (informative) \xB7 **For:** operators \xB7 **Prereqs:** [Quickstart](getting-started.md)\n\nThe Claude Code connector turns a real `claude` session into a Cotal mesh peer. A bundled\nplugin inside the session joins NATS, maps lifecycle hooks to presence, and exposes the\nmesh tools. Nothing wraps Claude; it is an ordinary session that happens to be on the\nmesh.\n\nThe shared mesh runtime (agent, `cotal_*` tools, hook relay) lives in\n[`@cotal-ai/connector-core`](../extensions/connector-core); this connector is the thin\nClaude-specific adapter over it. Its lifecycle hook imports the relay from the\n`@cotal-ai/connector-core/relay` subpath, so each hook process loads the relay and its environment\nreaders and none of the NATS client, zod or yaml. Siblings: [OpenCode](connect-opencode.md) (beta),\n[Hermes](connect-hermes.md) (alpha), [pi](connect-pi.md) (alpha); the\n[Connectors](connectors.md) matrix compares them feature-by-feature.\n\n## Set up\n\n```bash\ncotal setup # one-time: installs the plugin, seeds one agent; launches nothing\ncotal up # brings up the mesh + delivery daemon + a detached manager\n```\n\n`cotal setup` installs the cotal plugin (so the repo's Claude sessions get the `cotal_*`\ntools), shares your own MCP servers with spawned sessions on its first run (see\n[Sharing your MCP servers](#sharing-your-mcp-servers)), and seeds one `default` persona; `cotal up` brings up the local stack so\n`cotal spawn --detach` / `cotal_spawn` work right away. Re-running either is idempotent.\nThe install mechanics and the invariants behind them are in\n[setup internals](setup-internals.md).\n\n`cotal setup` also installs Cotal's authored Agent Skills (`SKILL.md`, the agentskills.io format) for\ncoordinating agent teams (today `team-topology`), from one canonical source, on two channels:\n\n- **Claude Code** gets a second, skills-only plugin, `cotal-skills`, from the same `cotal-mesh`\n marketplace, at **user scope** (machine-wide). The Claude connector declares and implements this\n setup provider, including the marketplace assets and native plugin commands; the base CLI only passes\n the vendor-neutral Agent Skills directory. The plugin carries no code and no core dependency,\n and uninstalls on its own with `claude plugin uninstall cotal-skills --scope user`. Its plugin version\n is stamped from the running CLI release, so an upgrade + `cotal setup --skills` runs `claude plugin update` and\n the deployed install actually gets the new skill. `cotal setup` installs it on first run and on repeat\n runs, so upgraders are not left behind. The same provider reports the plugin and skills plugin rows\n in `cotal status`, which point a stale or missing skills plugin at `cotal setup --skills`.\n- **Every other harness** (Codex, Cursor, OpenCode, Gemini CLI, Windsurf/Devin) reads the cross-vendor\n `~/.agents/skills/` directory convention, which has no remote index, so `cotal setup` **reconciles** it\n (and `cotal setup --skills` does only that):\n it installs/updates each Cotal skill, backs up a copy you have edited to `SKILL.md.bak` before\n replacing it, and removes a Cotal skill that is no longer shipped. Only skills Cotal owns are touched;\n your own or third-party skills there are left alone. `cotal status` reports whether the drop is current,\n stale, missing, or has a retired skill to reconcile, and names `cotal setup --skills` as the remedy. This is the working cross-vendor path.\n\nCotal also generates an [Agent Skills discovery index](https://cotal.ai/.well-known/agent-skills/index.json)\non cotal.ai, but that RFC is still a draft with no harness consuming it yet, so it is a forward bet,\nnot a channel to rely on today.\n\n## Spawn a session\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn dave --detach # supervised: the manager runs it in a PTY\n```\n\nA spawn resolves a persona from `.cotal/agents/<name>.md` ([agent files](agent-files.md));\n`--model`, `--variant`, `--cwd`, `--prompt`, ACL overrides, and `--share-tools` apply to\nboth forms ([run a mesh](run-a-mesh.md) has the full resolution rules). The session joins\nwith identity from its environment and auto-registers presence by the time it is\ninteractive.\n\nInside the session, the agent orients with one read-only tool, `cotal_orientation`: its\nidentity, the channels it reads and may post to, its capabilities, the tools available,\nwho's present, and unread counts. The full tool surface is the\n[MCP tool catalog](mcp-tools.md). In auth mode the team-supervision tools\n(`cotal_spawn` / `cotal_persona` / `cotal_personas`) are injected **only** for personas declaring\n`capabilities: [spawn]` (the same grant that opens the privileged control subject), so an\nagent's toolset matches its declared capabilities. `cotal_run` is gated separately by\n`run`; use `capabilities: [spawn, run]` for both. Fresh setup defaults include both.\nSee [workflow tool setup](workflows.md#from-an-agent-session) for a first run and missing-tool checks.\nClearing retained history is\noperator-only ([run a mesh](run-a-mesh.md)), never an agent tool.\n\n## How it binds\n\nClaude Code exposes four integration surfaces, and three of them collapse into a single\ndual-purpose MCP server:\n\n| Surface | Mechanism |\n|---|---|\n| Outbound, ambient | `http` lifecycle hooks \u2192 POST to the connector (presence, activity) |\n| Outbound, deliberate | MCP tools `cotal_send` / `cotal_dm` / `cotal_anycast` (+ `cotal_feedback`) |\n| Inbound, pull | MCP tool `cotal_inbox` (same server) |\n| Inbound, push | Channel nudge + hook drain (below) |\n\nThe manager launches the *real* `claude` (no wrapper):\n\n```\nclaude --strict-mcp-config --mcp-config '{\"mcpServers\":{\"cotal\":{\u2026}}}' \\\n --dangerously-load-development-channels server:cotal\n# env: COTAL_SPACE, COTAL_NAME, COTAL_ROLE, COTAL_CHANNEL=1, plus claude's documented auth vars\n```\n\n- **Model auth.** Locally, `claude` still reads macOS Keychain / `~/.claude`. In a container or\n CI there is no Keychain, so the connector forwards the documented credential set:\n `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`), `ANTHROPIC_API_KEY` /\n `ANTHROPIC_AUTH_TOKEN`, and the cloud-provider flags plus their credential vars. Host-session\n markers (`CLAUDE_CODE_CHILD_SESSION`, `CLAUDECODE`) stay out so a nested seat still saves a\n transcript. See [Deploy](deploy.md).\n- **Persona privacy.** The persona body is written to a private file and Claude receives only\n `--append-system-prompt-file <path>`. The body never appears in the spawned process argv. The\n carrier is a 0600 file inside a 0700 directory on POSIX, with equivalent owner-only ACL hardening\n on Windows. That is OS-user isolation: any process running as your user can read it while it\n exists. The manager or the foreground `cotal spawn` removes it, and the shared-server MCP config\n file, once it has proved the `claude` process gone. If the launcher is killed first, a watcher\n started beside `claude` removes them when `claude` exits.\n- **MCP servers.** `--strict-mcp-config` ignores every ambient MCP source, so a spawned agent\n loads the cotal server plus the servers the cotal config shares. First-run `cotal setup`\n fills that list with your own user-scope servers, so a spawned session has the tools you know\n (see below).\n- **Installed plugin.** The plugin is installed once (`claude plugin install\n cotal@cotal-mesh --scope local`) because its hooks bind only to an *installed* plugin.\n The repo's `.claude-plugin/marketplace.json` lists the committed plugin tree under\n `claude-plugin/`, which each release regenerates with the built bundles, the skills and the\n release version, so an install from the repo or from a pinned commit runs without a build\n ([Release](release.md)). `cotal setup` (npx, no clone) materializes the same marketplace under\n `~/.cotal/claude-plugin/` from the installed CLI (each plugin dir is rebuilt from scratch and\n atomically replaced, never merged, so no stale file rides in). The\n `cotal-skills` plugin installs from that same marketplace at user scope (`claude plugin install\n cotal-skills@cotal-mesh --scope user`); its manifest and install behavior ship inside the Claude connector, and\n its version tracks the CLI release so updates land.\n- **Identity-gated.** Connector code requires `COTAL_NAME`, `COTAL_LINK` or `COTAL_AGENT_FILE`.\n A plain `claude` with none of them never joins, so your own sessions in a repo do not appear\n as stray peers. Its MCP server still answers `initialize` and lists one static tool,\n `cotal_how_to_join`, which explains how to launch a session on a mesh. It builds no mesh\n agent, opens no broker connection and binds no control socket.\n- **Hands-free.** The dev-channels flag prints a one-time confirm prompt. The runtime waits for the\n dialog title in normalized terminal text and presses Enter once when it appears, so startup speed\n does not affect a supervised launch. The PTY runtime reads the child's output, and the tmux, cmux,\n Orca and Herdr runtimes read the pane's screen. If the declared prompt never appears within 15\n seconds, the seat ends with a bounded error naming the unmatched prompt instead of hanging\n silently or answering another dialog. The PTY runtime writes that error to the seat's output, and\n the other runtimes write it to the manager's log.\n- **Trusted directory.** Claude opens a directory it has not trusted on its workspace-trust dialog,\n and the dialog's default answer exits. No one is at a supervised seat to answer it, so a launch\n whose directory the manager host's own Claude does not trust is refused before it starts, naming\n the directory and the dialog. Trust is read as Claude reads it: trust given to a parent directory\n counts up to the root of the directory's own Git repository, and a linked worktree shares the trust\n of its repository's main checkout. Open `claude` in that directory on the manager host once and\n trust it, then spawn again. A foreground `cotal spawn` shows the dialog in your own terminal instead.\n\nInbound mesh messages arrive in context as\n`<channel source=\"cotal\" from=\"bob\" kind=\"dm\" \u2026>\u2026</channel>`: each meta key a tag\nattribute the agent can read for routing.\n\n## How messages reach the session\n\nDurable deliveries land in the connector's inbox from JetStream consumers\n([SPEC \xA78](../SPEC.md#8-nats--jetstream-binding)); live channel traffic can instead arrive\nthrough an at-most-once core subscription. A durable message sent while the agent is busy\nor offline waits on the stream. Two things move a message from inbox to model; one\ndelivers, the other only wakes:\n\n- **Hook drain (delivery).** `SessionStart` / `UserPromptSubmit` hooks read automatic inbox items and\n inject them as `additionalContext`. This is the single authoritative path: deterministic and works\n on any Claude Code build. Quiet ambient is excluded and stays buffered for `cotal_inbox`.\n A message is **acked only once the hook reply carrying it has cleared both legs of its journey**:\n the connector's control socket to the hook process (which gives up after 2s), and the hook\n process's own stdout to Claude Code (which it force-exits 1s after starting to write). The relay\n sends a receipt back down the control socket from that stdout write's callback, and only on a\n clean write (a runtime whose pipe has gone away fails it), and the connector treats that receipt,\n not its own socket write, as delivery. So a large injection killed mid-flush, or one written to a\n broken pipe, leaves the message un-acked and JetStream redelivers it. What this does *not* prove is\n that Claude Code read or applied the reply: a payload small enough to fit the pipe buffer is\n reported written the moment the kernel takes it. That residual is why the path errs toward\n at-least-once rather than treating a confirmed write as a confirmed read. Acking when\n the reply was merely *formatted* meant a lost reply was a lost message: it was already marked\n handled, so its own redelivery was silently acked on arrival.\n A hook whose handler throws still returns an empty reply so the session is never blocked, and\n that reply carries nothing, so it commits nothing: the batch it had started to surface stays\n un-acked and goes out on a later frame. The seat also drops any `turn-pending` row that breaks\n the manager contract, such as one with no integer deadline, and says so once in its log. A reply\n with no `turns` array changes nothing: the seat keeps the turns it already holds.\n This errs toward **at-least-once**: if a reply lands but its confirmation does not, the batch is\n surfaced again and flagged as a possible repeat. A duplicate injection is noise; a buried DM stops\n the peer answering at all.\n- **Channel nudge (wake).** An arriving message fires a `notifications/claude/channel`\n event that wakes an *idle* session into a turn, so the drain runs *now* instead of at\n the next prompt. The nudge never acks anything. A nudge that the host rejects is retried with a\n bounded backoff while anything is still pending. For an idle session it is the only wake source,\n so dropping it means silence until someone types. When the channel becomes active, the connector\n first re-fires a focus mention remembered during startup, otherwise one buffered wake. A rejected\n push keeps its bounded retry, and JetStream redelivery remains the durable backstop for unacked\n inbox items. Neither a redelivery nor that retry repeats a nudge already pushed for that message,\n whether the message had its own nudge or was counted in a batch one, so a session held in a long\n tool call gets one nudge per message. Once a hook frame carries the message, or the push that\n announced it fails, its next redelivery nudges again, so a reply that never reached Claude Code\n still recovers. If the channel cannot run at all, delivery still waits for the next hook. Live-only\n traffic has no durable retry.\n\n**Two priority tiers.** A *directed* message (DM, anycast, or a channel message that\n`@mentions` us) always nudges. *Ambient* channel chatter does not nudge mid-turn; it\naccumulates, and the `Stop` \u2192 idle transition fires one batch nudge so the backlog drains\ntogether.\n\n**Constraints (accepted).** Channels are a Claude Code research preview (\u2265 v2.1.80;\npermission relay \u2265 v2.1.81): Anthropic auth only, admin-enabled on Team/Enterprise, and a\ncustom channel needs the `--dangerously-load-development-channels` launch flag. The hook\ndrain does not depend on any of that; the channel only adds \"wake me when idle.\"\n\nThe same channel also relays **tool-permission requests** onto the mesh, so a peer (a\nhuman at the CLI, a policy node) can approve or deny an agent's pending tool call through\nCotal rather than a per-terminal prompt.\n\n### Attention\n\nAn agent picks how aggressively peer traffic reaches it with\n`cotal_status({ attention })` (three modes, orthogonal to presence):\n\n| arrival | open (default) | dnd | focus |\n|---|---|---|---|\n| directed (dm / anycast) | wake + inject | wake + inject | wake + inject |\n| channel `@mention` | wake + inject | wake + inject | ack-drop; wake to *pull*; not injected |\n| ambient channel chatter | wake when idle; hold while working | never wakes; injects next turn | ack-drop; recall via `cotal_inbox` |\n\nPer-channel overrides refine this: **quiet** (delivered, never wakes; `@mention` still\nwakes) and **muted** (dropped on receive, mentions included; DMs/anycast unaffected), set\nwith `cotal_channel_mode` or as agent-file defaults (`quiet:` / `muted:`,\n[agent files](agent-files.md)). A per-channel override is the final word for that channel.\nQuiet ambient is pull-only: it never hitchhikes on a human prompt, DM, mention, or other\nconnector-driven turn. `cotal_inbox` explicitly surfaces and clears it. A quiet-channel\n`@mention` remains automatic and injects normally.\n\nA pull is bounded too, and clears only what it hands over. One `cotal_inbox` call carries at most a\nreceivable window (direct messages and role requests first, then channel traffic, replayed history\nlast); whatever does not fit stays buffered, is named in the reply, and comes back on the next call.\nA message too large for one whole response is delivered in parts: once no smaller mail is waiting,\neach call carries its next part, and it is cleared only after its last part goes out, because clearing\nwhat was not handed over is the loss this bound exists to stop.\nThat matters most on the path where it is easiest to lose mail: reconnecting brings a channel-history\nreplay with it, so the largest payload and the least expendable message arrive in the same read.\n\nThe local inbox is bounded. On pathological overflow it evicts pull-only items first, then other\nchannel traffic, and a direct message or role request only when the whole buffer is directed mail.\nAn evicted channel item is acknowledged. An evicted direct message or role request never is, and\nthe broker redelivers it after the ack wait until a redelivery finds room. A direct message stays\npending on the session's DM durable, where `cotal deliver pending <name>` counts it. A role request\nstays on its role's shared queue, which that command does not read. A full inbox therefore delays\ndirected mail until the session drains it.\nIf the bounded live/durable classification guard also fills, the connector fails closed:\notherwise-normal ambient becomes pull-only until restart. Muted hard-drop and normal focus recall\nstill take precedence. Focus also keeps a bounded exclusion list so mode toggles cannot recall\nquiet/muted traffic; if that safety bound fills, recall skips the affected channel and reports it\nas incomplete rather than risk resurfacing excluded content. Recall cannot tell one message with an\nempty id from an identical one with another disposition, so in focus such a message is held in the\nlocal inbox as pull-only instead of being dropped, and a mention of it still wakes the agent. When\nthe session settles an id-less copy, it reads the chat stream's last sequence. Identical copies\narrive in stream order, and those reads can answer out of order, so the first read that can see the\ncopies binds them in arrival order, latest first, each to the latest unbound stream copy at or below\nthe lowest sequence read for it or any later identical copy.\nWhile the connection stays up, every stream copy at or below that sequence reached the session\nfirst, so a later identical copy sent during a reconnect gap is above it and stays unbound. A copy\nthat arrives while a read runs may not be in that read, so it binds nothing there and recall reads\nthe channel again, up to three times; if it still could be a stream copy in the last read, recall\nleaves that stream copy in the stream and reports the channel as incomplete. A settled copy a\ncomplete read cannot bind is behind the focus start or out of retention, and is forgotten. Recall\nhands back into the inbox only the stream copies nothing is bound to, such as one sent during a\nreconnect gap or one the inbox evicted on overflow, and `cotal_inbox` hands each over once. Overflow\nfrees only the evicted copy, held or quiet, and a copy no read has bound yet keeps its place in\narrival order, so an identical muted copy stays out of recall. When\nthe inbox is full, recall leaves them in the stream for a later call and reports the channel as\nincomplete. A history read that fails, or a channel with replay off, settles nothing and is reported\nas incomplete, and recall calls run one at a time. If the sequence read for a settled copy fails,\nor answers only after the connection dropped, recall skips that channel for the rest of the focus\nperiod and reports it as incomplete. Recall cannot tell a late copy of a message it handed back from\na new identical message, so every copy takes its own disposition: a new identical quiet mention is\nstill delivered automatically, and a late copy can surface a second time.\nA recalled message that already went out in part is read to its last part, even if an exclusion\nlands after its first part. One session reads its inbox one call at a time: a `cotal_inbox` call\nthat overlaps another waits for it to finish, so neither decides from a view the other has already\nmoved past.\nIf the separate hard-drop disposition guard fills, channel traffic is dropped for the rest of the\nsession rather than risk a late copy bypassing an earlier muted/focus decision; DMs and anycast are\nunaffected.\n\nAttention is **advisory UX, not a boundary**: any peer can wake a dnd/focus agent by\nnaming it, and `muted` means \"I opted out of receiving\", not \"the channel is blocked\";\nthe broker still authorizes and delivers. Focus's real effect is shrinking the\nuntrusted-ambient injection surface (only subject-authenticated dm/anycast auto-inject).\nIt resets to **open** on `SessionStart`, so a restarted agent never stays silently deaf.\nYour attention is mirrored into presence so peers can see it.\n\nWhatever does reach a turn is framed so a peer cannot write the frame. A line that begins at column\nzero is written by the connector; one message is one line plus indented continuations, with the\nsender inside a single bracket pair. A message body, a sender name and role, and a service or\nchannel label are all peer-controlled, so each passes through the same neutralization the\n`cotal_inbox` reply uses: no line break a splitter may honour and no bracket survives into a\nrendered attribution. This matters more for an injected block than for a reply, because the agent\ndid not ask for it and so never had the chance to distrust it.\n\n## Presence mapping\n\nThe connector wires a small subset of Claude Code hooks to presence states; presence is\ncoarse, and \"what it is doing\" rides on activity updates. Presence is **advisory**: a presence\npublish that fails (the endpoint mid-reconnect, say) is swallowed and never prevents the same hook\nfrom delivering messages or flushing held ones.\nA `SessionStart` during an open turn, including compaction, preserves the current `working` or\n`waiting` status until `Stop`, `StopFailure`, or `SessionEnd` closes the turn.\n\n| Hook | \u2192 state |\n|---|---|\n| `SessionStart` | `idle` only when no turn is open (join; surfaces the inbox; captures the live model into `meta.model` when no pin) |\n| `UserPromptSubmit` | `working` (turn starts; surfaces the inbox) |\n| `PreToolUse` | no change; records *what* is about to run, so a permission wait can name it |\n| `Notification` (`permission_prompt` / `agent_needs_input`) | `waiting` with condition `approval` / `input` (activity leads with the pending tool, e.g. `Bash: git push \u2026`) |\n| `Stop` / `StopFailure` | `idle` (turn done / died on an API error; flushes anything held while busy). `StopFailure` also relays Claude Code's native error value as `condition.source` and maps it to the closed condition vocabulary. On the [event plane](#event-plane) it closes the run with `RUN_ERROR`. |\n| `SessionEnd` | `offline` (graceful leave) |\n\nThe connector also leaves gracefully when its stdin closes. An MCP client closes it to end the\nsession, and a killed `claude` closes it with no `SessionEnd`, so a dead session drops off the\nroster instead of staying on it as a live peer.\n\n`StopFailure` maps `rate_limit` and `overloaded` directly; auth and credential failures to\n`auth`; account and billing failures to `billing`; `invalid_request` to `request`;\n`model_not_found` to `model`; `server_error` to `server`; `max_output_tokens` to `context`; and\n`unknown` to `failed`. The native value remains in `condition.source`.\n\nHooks are relayed over the connector's **authenticated** local control endpoint (per-user\nsocket + per-launch token, constant-time checked), so a local process that finds the path\nstill can't drive presence or stop the agent. The full Claude Code hook-event list lives\nwith the adapter:\n[`extensions/connector-claude-code`](../extensions/connector-claude-code/README.md).\n\n## Event plane\n\nA spawned session publishes a **structured** account of what it\ndid: run boundaries per turn, assistant text, reasoning, and each tool call with its start\nand its end. Not prose about the work, the work itself, in a vocabulary a program can\nread. The launcher sets `COTAL_EVENTS` by default; pass `--no-events` to opt out on an unrestricted\nspace. A user-auth registration with `policy: { events: \"required\" }` carries `eventsRequired` in the\nprivate launch material, so the connector arms even without `COTAL_EVENTS`; `--no-events` is refused.\nA hand-driven user-mode session may carry the same decision as `COTAL_EVENTS_REQUIRED=1`. Its own\npublish grant must cover `events.<owner>.<actor>` or the connector refuses before joining. An unmanaged\nsession with no launch material and no required-policy fallback keeps the generic default behavior.\n\nIf the event plane stops for good, the space's policy decides what happens to the seat, on every\nconnector. On a space that requires events the seat stops and leaves the mesh. On any other space\nit keeps running without events, and the connector log records `AG-UI emitter stopped` with the\nreason. For Claude Code the connector is the MCP server: it leaves the mesh and exits with code 1,\nand its stderr carries that line.\n\nA new session includes its first run even when Claude writes a positional startup prompt before the\nconnector receives `SessionStart`. That from-zero read is keyed only to Claude's explicit\n`source: \"startup\"`; resumed, forked, cleared, and compacted sessions adopt at the transcript boundary\ncaptured at that adopt, before the mesh link connects, so nothing Claude appends while the connector\nis still starting up lands behind the cursor and is silently dropped. Crash recovery follows the\ncursor already stored in the event write-ahead log, regardless of the new process's startup label.\n\nClaude starts each hook in its own process, so a prompt or stop relay can reach Cotal before the\n`SessionStart` relay. The connector holds those event flushes and the terminal until `SessionStart`\nsupplies the source, then enqueues adopt, flush, and close in that order.\n\n`SessionStart` can also run before the connector process has bound its local control socket. The hook\nthe `SessionStart` relay retries only transient pre-connect listener errors, with capped backoff\ninside its existing two-second budget. Later hooks and permanent local faults still fail open\nimmediately. Once a socket has connected, a broken exchange is not retried: the connector may\nalready have handled the frame, so replaying it could apply one lifecycle event twice.\nThat retained `SessionStart` can itself arrive before Claude creates the transcript path. A genuinely\nnew startup waits up to five seconds for that file with capped backoff, and the same deadline bounds\none stalled file read; expiry fails loud instead of silently losing the first run. A forked session\ngets the same wait, because Claude copies the parent transcript into the fork's own file after the\nhook, and then adopts at the end of that copy. Resumed, cleared and compacted starts and recovered\ncursors still require their existing source at once.\n\nTool arguments (`TOOL_CALL_ARGS`) and tool results (`TOOL_CALL_RESULT`) are not republished\nonto this channel. The durable emitter drops those events before they are written to the\nwrite-ahead log, because this channel's read ACL is not the ACL the tool ran under. Content is\nmandatory on both kinds, so the event is suppressed rather than emptied or replaced with a\nplaceholder. Tool start and end still go out. A restart that finds a pending pre-fix frame\nstill carrying those kinds HALTS rather than republishing it.\n\nThe channel is **`events.<owner>.<actor>`**, named after the session's principal. What the actor\nhalf is depends on the mesh, and the difference matters when you go looking for it: on a static mesh\nit is a key the manager allocated, never the display name, so two live agents sharing a display name\ndo not share a stream; on a user-auth mesh it is the agent's own name, because that is what the\nledger row is keyed on. Spelled out again with both halves below. The launch grants publish rights\non that channel alone. A spawn\nthat asks for a *different* agent's event channel is refused at the door rather than granted, since\nthat channel is that session's event stream. The same rule runs on restart: a manager\nresume document that names another agent's event channel is refused rather than adopted, because the\nmanaged row is re-armed from that document and the credential is re-minted from the row.\n\nThe rule reads a **concrete** channel, two principal tokens and nothing else. A pattern such as\n`events.<owner>.>` is not an event channel to it and passes untouched, governed by ordinary ACL\nauthority: on a user mesh the delegation envelope, on a static mesh the spawning credential itself.\nThat is deliberate, because the pattern is the form an operator writes on purpose for an observer,\nand it is worth knowing rather than assuming the fence is total.\n\nTo let something else read a plane, grant it out of band. The refusal prints the command for the\nmesh it is running on, spelled out in full, and only that one.\n\nOn a **user-auth** mesh:\n\n```bash\ncotal actor grant <reader> --owner <owner> --scope '' --allow-subscribe 'events.<owner>.<actor>' --allow-publish ''\n```\n\nEvery field, deliberately. `actor grant` is an upsert of the whole row, so it refuses a grant that\nleaves off any of the three ACL flags. Only `--full` turns an omitted flag into the wide default\n(`>` read, `>` post, `spawn,role:default` scope), which is the opposite of what a scoped watcher is for.\n\nOn a **static** mesh there is no actor ledger for `actor grant` to write to, and the refusal says\nso; mint the reader instead:\n\n```bash\ncotal mint watcher --profile agent --allow-subscribe 'events.<owner>.<actor>' --provision\n```\n\nThe **agent** profile, not the observer one. `mint` reads `--allow-subscribe` only for that\nprofile, and refuses it anywhere else: `--profile observer --allow-subscribe <channel>` exits\nnon-zero and writes no creds file, because the observer profile carries a fixed read set over the\nwhole chat plane, which is the opposite of what a scoped watcher is for. The agent profile also prints the lifecycle uid the\nreader needs, since an authed consuming endpoint refuses to start without one.\n\nOn an **open** mesh there is nothing to grant: the mesh has no credentials and no ACLs, so any peer\nthat lists the channel reads it, and the refusal says so instead of naming a command. The\nown-channel rule still applies there, because a spawn is not the place to hand out a read on\nanother agent's tool inputs and outputs.\n\nTwo things a reader has to do that are not obvious, both on `CotalEndpoint`. It must pass the event\nchannel in `channels`: an endpoint reads the channels it lists, so one constructed without\nthe event channel joins nothing and the frames never arrive. And it reads history with `readHistory(channel)`, the delivery daemon's mediated read, not\n`channelHistory(channel)`: a scoped credential is denied the ad-hoc consumer the direct read\ncreates, by design. `cotal console` and the web console already do both.\n\nThe `<owner>.<actor>` pair is the session's principal. On a user-auth mesh the actor half is the\nagent's own name, so the channel is `events.<your-owner>.<agent-name>`. On a\nstatic mesh the owner half is the literal `local` and the actor is a key the manager allocated, so\nthe channel is `events.local.<key>`; the spawn reply carries that key as `id`. Note\nthat `cotal console` and the web console keep event channels out of their channel lists on purpose,\nsince a plane is a machine feed rather than a conversation; they draw the frames when you open the\nchannel by name.\n\nThe rule governs the manager's doors, which are the ones a caller other than you can reach. A\nforeground `cotal spawn` on your own machine mints from your own signing material, so it can still\ngrant any channel you name: that is the out-of-band grant, not a way around the rule.\n\n**Failed turns publish run errors.** Claude Code decides for itself\nwhether a turn finished or died and fires one of two hooks accordingly, so the connector relays that\ndecision rather than making one of its own: a turn that ended on an API error ends its run with\n`RUN_ERROR` carrying the fixed message `run failed` and no code. Neither the detail Claude Code\nreported nor its error kind is published there: both are upstream values that can echo your prompt or\ntool output, and the events channel has a different read ACL. The error kind still reaches presence\nas the agent's condition (`rate_limit`, `auth`, `billing` and the rest). A turn that ended normally still\nends with a run-finished event carrying no outcome, which says the turn ended and does not claim it\nsucceeded.\n\nEvents are written to a per-session write-ahead log before they are published, so a hook that fires\nafter a restart resumes at the cursor it left rather than replaying or skipping, and a run that was\nopen when the session stopped is closed rather than left dangling.\n\nOne channel carries **every session of one agent**, because it is named after the principal and not\nafter the session. Alongside the per-session logs the connector keeps one small record per principal,\nholding the last sequence the broker assigned on that channel, so a new session continues the stream\nits predecessor left instead of starting again from nothing. Both live under the events state root\n(`COTAL_WORKSPACE_ROOT`), and neither is something you edit by hand.\n\nA **missing** record is not a fault: the connector rebuilds it from the session logs beside it,\nwhich is how an agent that was already running before this record existed keeps its stream. That\nrebuild stops if any one of those session logs is damaged. Unreadable, not valid JSON, and written\nfor a different principal all count, and so does a session directory or a log that is a link rather\nthan the real file the connector wrote, or a log that has more than one name. A tip taken from the\nrest would be too low, and it would stop publication later with nothing left to point at the cause.\nThe connector names the file instead, and the only way past it is the directory removal described\nbelow, under the same condition. A record that **disagrees with the broker** is a fault, and the\nconnector stops publishing and says why rather than guessing. A record that **moved while a session\nwas writing to it** is refused the same way: it means something else wrote the principal's record,\nand the connector reports which value it held and which the file holds rather than writing over the\nlater one. There is no command to clear it. The state is the principal's directory under the events\nroot, and clearing it by hand means removing that directory whole: the sequence, the cursor and the\nper-session logs only mean anything together, so removing part of it leaves a state the next start\nrefuses. Removing it is only half a remedy, and the half that comes first is the channel. The\ndirectory is where the agent's memory of the tip lives, not the tip itself, so on a channel that\nstill holds frames the next session opens expecting an empty one and stops on the same\ndisagreement, with the logs a tip could have been rebuilt from now gone. Purge the channel first,\nthen remove the directory.\n\nReading it: `cotal console` and the web console draw event frames directly. A frame carries no text\npart by design, so a surface that renders a message as flat text shows a marker instead of prose.\n\n**On a per-user-auth mesh, the default event plane needs the spawner's grant to cover the channel.** The event\nchannel is added to the child's publish set, and delegation only narrows: an agent may hand down\na subset of what it holds and no more. So a peer-initiated spawn is refused unless the\nspawning identity's own grant already covers the child's event channel. The refusal prints the\nexact `cotal actor grant` command that widens it. An operator launch, whose chain reaches an\nadmin-scoped or roster row, is unaffected. Passing `events: false` is the explicit opt-out.\n\nArming the event plane through a typed spawn request (`manager.spawn` with `events`, including\nthe CLI's `cotal spawn --detach --events`) is allowed on a user mesh when the child runs under\nthe caller's owner. This also satisfies a registration policy that requires the plane. A spawn\nunder another owner still needs the caller's admin tier. Without it, an explicit request or a\nrequired plane is refused before provisioning. A silent non-owner caller in a space without that\npolicy has the plane disarmed with a notice in a successful reply. Later provisioning can still\nrefuse the spawn. The child's own-channel rule and the ledger's delegation envelope still apply.\n\n## Resume a session\n\n`--resume <session-id>` pulls an existing Claude session, its context and transcript,\ninto the mesh. It **forks**: Claude mints a *new* session id from that transcript\n(`--resume <id> --fork-session`), so the meshed agent gets its own session and the\noriginal is untouched.\n\n- `cotal spawn --resume <id>` (foreground) is the primary surface: the transcript is on\n *your* machine, and errors are Claude's own stderr, inline.\n- `--detach --resume <id> --on <instance>` carries a session held on *your* machine to that\n manager instance, which may run on another host. The CLI finds the transcript under your\n Claude config (`~/.claude`, or `$CLAUDE_CONFIG_DIR`), sends it through a JetStream Object\n Store bucket only that instance reads, under a writer credential pinned to that one transcript,\n and prints `carried session <id> to <instance>:\n sha256:<hex>, <sent> of <size> bytes sent in <chunks> chunks`. A re-run of the same bytes\n sends nothing, and an interrupted carry continues where it stopped. The seat forks it in a\n private Claude home under the manager's `.cotal/seat-homes/`, which no other seat's Claude\n lists or finds, and which is removed when the seat stops. When Claude starts the fork, the seat\n records the SHA-256 of the transcript it read from its own project; the manager stops a seat\n whose record names other bytes than the carried ones, or that records none within the join\n timeout after it joins, an uncertain launch included, and otherwise shows that record as the\n seat's provenance. `cotal attach` to such a seat names its source after the seat\n name, as `(resumed from <host>:<id>)`. A remote manager receives a carry when its host issues it a\n transfer reader. On a user-auth mesh the CLI exchanges the operator's login for a one-object\n `transfer-writer` view, which needs scope `admin`.\n- A session name in place of an id is refused, listing each session on this host that carries\n that name with its id, SHA-256 and modification time. An id this host does not hold resolves\n against the **manager host's** `~/.claude`, as before.\n- A seat-private home holds no login. The manager host needs `CLAUDE_CODE_OAUTH_TOKEN` (from\n `claude setup-token`), `ANTHROPIC_AUTH_TOKEN`, or a cloud provider selection in its\n environment; `ANTHROPIC_API_KEY` alone is refused. The launch directory must already be\n trusted by the manager host's own Claude, and Claude must be 2.1.234 or later.\n- The manager waits for a real outcome: `\u2713 started` means the agent *joined the mesh*,\n `\u2717 exited on launch` carries Claude's last output, and an uncertain launch (~30 s) is\n reported without tearing the agent down.\n- Resume is an **operator surface only**, deliberately not exposed on MCP `cotal_spawn`\n (a mesh peer naming host-local transcripts would widen `spawn` into transcript\n disclosure). Only the Claude connector supports it today; OpenCode and Hermes fail loud.\n- Needs a `claude` new enough for `--resume \u2026 --fork-session` (verified on 2.1.197).\n\n## Sharing your MCP servers\n\nA spawned session keeps your own MCP servers by default. On its first run, `cotal setup` copies\nthe user-scope servers from your Claude Code config (`~/.claude.json`, or the one under\n`$CLAUDE_CONFIG_DIR`) into the cotal config file (`~/.config/cotal/config.json`) under\n`connectors.claude.mcpServers`, and names them in its output. With none to copy it writes an\nempty list. Each entry is the familiar `.mcp.json` shape ([full format](config.md)). A cotal\nconfig that already declares that list keeps it, and a later `cotal setup` never changes it.\n\nThe cotal config holds secrets only as `${VAR}` references. Setup cannot tell literal text from\na secret, so it leaves out a server with an `env` or `headers` value that is anything but `${VAR}`\nreferences (a `Bearer ${TOKEN}` header among them) and names it in its output. To share one,\nadd it to the cotal config with each secret written as a `${VAR}` reference, and export that\nvariable where you spawn. Setup also leaves out and names an entry no session can start, such as\none with a missing or empty `command` or `url`, or one whose `command` is not a string.\n\nAt launch the connector forwards *only* the named vars the chosen servers declare and\npasses the merged config as an owner-only temp file; `--strict-mcp-config` stays on, so\nonly cotal + the shared servers load.\n\nFor a lighter seat, share fewer. Remove an entry from the cotal config to drop it from every\nspawn, or scope one spawn with `--share-tools tavily,figma` (or `--share-tools none` for cotal\nalone). An empty list (`\"mcpServers\": {}`) in `~/.config/cotal/config.json` keeps every spawn\nisolated, and setup leaves it as it is.\n\nTwo caveats: sharing a server grants its credential to the agent (the var lives in the\nClaude process's environment, so share only when you're fine with that teammate holding\nthe key), and memory adds up, because a heavy server boots once per spawn, multiplied\nacross a team, and can starve a small machine.\n\n## Feedback\n\n`cotal_feedback` works out of the box: without a key it posts to the public intake at\n`https://cotal.ai/v1/feedback` (needs a contact email: `COTAL_FEEDBACK_EMAIL`, then\n`git config user.email`, else the agent asks). Set `COTAL_FEEDBACK_KEY=fbk_<key>` in a\nbeta tester's environment to route to the keyed intake (`Authorization: Bearer`, identity\nderived from the key); `COTAL_FEEDBACK_URL` overrides either endpoint. The CLI can send\ntoo: `cotal feedback \"<summary>\" [--type bug]`. Each submission carries\n`origin: human | agent`, whether the tester asked, or the agent auto-reported a major\nissue.\n"
|
|
57796
57796
|
},
|
|
57797
57797
|
{
|
|
57798
57798
|
"slug": "connect-codex",
|
|
@@ -57841,7 +57841,7 @@ function loadDocsBundle() {
|
|
|
57841
57841
|
"title": "The control surface",
|
|
57842
57842
|
"kind": "Concept (informative)",
|
|
57843
57843
|
"summary": "Cotal once had a privileged control rail: a fixed set of named service tiers (self / manager / admin / delivery) on their own ctl.",
|
|
57844
|
-
"body": "# The control surface\n\n> **Concept** (informative) \xB7 **For:** operators and client authors who want to know how the manager and other daemons are driven \xB7 **Normative:** [SPEC \xA713](../SPEC.md#13-endpoint-control-surface-v04)\n\nCotal once had a privileged control rail: a fixed set of named service tiers\n(`self` / `manager` / `admin` / `delivery`) on their own `ctl.*` subjects, with the manager\nas a special case the broker recognised by name. That rail is gone. Everything that serves\nstructured commands now, the manager, the delivery daemon, a wrapped MCP server, a\nthird-party service, is an ordinary **endpoint**: a daemon that registers a service\nidentity, publishes its contracts, and answers `describe`. `manager` is an endpoint name\nlike any other; no subject, envelope, or grant in this surface knows it specially. The\nmanager is a service on the mesh, not an authority over it: it holds only the capability\nrows its callers grant it, and serves over a scoped credential.\n\n## The `ep` rails\n\nOne kind, `ep`, carries every request under a mode token that says where the request\nroutes, never which verb it is (the verb rides the envelope): `one` (queue-group\nanycast, one and only one class member), `all` (scatter, every instance), and `inst` (one instance by its\nstable address). Replies come back on a `reply` rail keyed to the serving instance and its\nepoch. Around these sit the sibling planes the composites use: per-goal events, timers,\nsessions, and the journal that holds durable facts. Every request carries the caller as\nthree forge-locked tokens, `owner`, `actor`, and lifecycle `uid`, plus an unguessable\nnonce, so the broker polices who is calling in the subject grammar itself. See\n[SPEC \xA713.2](../SPEC.md#132-grammar) for the grammar and [\xA713.5](../SPEC.md#135-verbs) for\nthe verbs (`call`, `cast`, `watch`, `claim`, `scatter`).\n\n## Lifecycle identity\n\nA principal `owner.actor` is a reusable routing alias: a despawn frees the actor name and a\nlater spawn may legitimately reuse it, so the alias alone is never authority. Two further\ncoordinates make an identity durable: a **lifecycle uid**, an unguessable, never-reused id\nfor one managed lifecycle under a principal, and a **process epoch**, the fenced ownership\nepoch of the process currently animating it, advanced on every restart or takeover. At most\none live epoch owns an identity, and a superseded epoch must stop serving. Durables and\ncredentials key on the lifecycle uid, not the reusable name, which is what lets a\nsupervised restart recover the same lifecycle instead of minting a new one. See\n[SPEC \xA713.1](../SPEC.md#131-lifecycle-identity) and [identity & auth](identity-and-auth.md).\n\n## Service discovery\n\nNo client has compile-time knowledge of any endpoint's commands. `cotal describe\n<endpoint>` resolves a registered endpoint's command set off the wire: the reserved\n`describe` command answers the registered contract digests, the schemas are fetched from the\nspace's content-addressed contract store, recompiled, and verified against those digests.\nEach command prints with its capability class and targeting shape. `cotal invoke <endpoint>\n<command> --args '<json>'` then calls one command by name, validating the arguments as they\nwill be sent (JSON drops a key whose value is undefined) against the fetched input schema before\npublish. A refusal at that check means nothing was sent, and it is marked `not-executed`. A\nsigned-in user invokes the same surface through their bearer, and the broker enforces each\ncommand's existing capability grant. A manager alias supplied through `--name` resolves through\nits name-keyed `inspect` command, so an authorized targeted call does not need the manager-wide\n`ps` enumeration grant. Every built-in manager command uses this\nsame trust chain, so there is nothing the built-ins can reach that a described contract cannot. The registered\n`auth` endpoint is describable the same way `manager` is: `cotal describe auth` lists\n`retire-lifecycle` and its exact-mode target shape.\nSee [SPEC \xA713.7](../SPEC.md#137-contracts-and-discovery) and [cli.md](cli.md).\n\nThe manager's `resolve-cwd` command is in the `manager.spawn` capability class. It accepts an\nabsolute path on that manager's host and returns its canonical directory plus the host name. It\nrefuses a relative, missing or non-directory path with `failed-precondition`; it creates nothing.\n`spawn` applies the same check at admission, before any credentials or durables are minted.\n\n### Inspecting a managed name\n\nManager `inspect` keeps its successful response as the live managed-agent row. A live hit does\nnot read durable lifecycle state, so a temporary records-store failure cannot break inspection of\nan agent the manager currently holds.\n\nOn a live miss, a static manager point-reads its durable slot row. A name with no slot, or a slot\nwhose phase is `retired`, remains `not-found`. A nonterminal slot returns\n`failed-precondition` with `error.details[].kind =\nai.cotal.manager.static-slot-observation`. The detail carries the slot's `slotPhase`,\n`owner`, `actor`, `slotLifecycleUid`, `cleanupComplete` when recorded, and `slotRevision`. It\nthen carries the separate lifecycle head's `headState`, `headOp` when present,\n`headLifecycleUid`, and `headRevision`. Head fields are absent when provisioning has not written\nthe lifecycle head yet.\nThe error message carries the same diagnostic summary so string-only operator paths do not hide\nthe structured detail.\n\nA slot row records the manager instance that owns it. In a space with more than one manager, the\nclass queue can hand `inspect` to an instance that does not host the name. When the row names a\ndifferent instance and is not `retired`, the miss returns `failed-precondition` with the same\ndetail plus `ownerInstanceId`, which names the only manager that can act on it. The message names\nboth instances, so a caller that reads only the string can tell it from `not-found`. A sibling's\n`retired` row remains `not-found`. A named `cotal_despawn` resolves its target through this read\nand cannot address an instance, so it asks again until the owning instance answers, up to 16\ntimes.\n\nThe slot is read before the head. These records do not form one atomic snapshot, so the detail\nalso carries `readOrder: [\"slot\", \"head\"]` and `consistency: \"ordered-not-atomic\"`. A head can\nadvance between the reads. The issuance gate is not projected because the retirement operation\nneeded for this diagnosis is already recorded on the head, and reading a third record would add\nanother non-atomic edge without changing the per-name result.\n\nIf either durable read fails or exceeds its bound, the miss returns `unavailable` with\n`ai.cotal.manager.static-slot-read-failed` rather than claiming the name is absent. That detail\nnames the inspected `name`, the failed `record` (`slot`, `head`, or `slot-or-head` when the layer\ncannot distinguish them), and `operation: \"read\"`. User-auth managers do not own `mgrslot` rows,\nso their inspect misses remain live-map reads.\n\nFor Linux custodied seats, retirement requires the runtime's process-exit evidence before\nfreeing the alias or deleting its credentials and delivery state. Socket loss alone is not\nproof of exit. The runtime retains the record captured at launch or adoption so a clean\ncustodian exit can unlink its file without losing the recorded boot and process identities.\nIf the file is missing, reaping uses that retained record and the existing kernel identity\nchecks. An unknown reference without either record refuses cleanup. Reused process ids\nare never signalled on the strength of the old record.\n\n### Listing the durable slots\n\nManager `slots` (`manager.read`, untargeted) lists the durable static slot rows this manager\nowns. Only static managers hold these rows: a user-mode or open manager answers\n`failed-precondition`, and a manager whose durable store is not standing answers `unavailable`.\nEach row carries the same `readOrder` and `consistency` fields `inspect` uses, because the list\nis read the same way: torn across rows as well as within each row's slot/head pair. A `retired`\nrow is never listed. `live` reflects the manager's live roster at render time, not the durable\nrow. A slot row is never deleted, so a row whose latest operation is a DEL or PURGE marker is\ncorruption: the list answers `unavailable` naming that row, as `inspect` does for its name.\n\n## Spawn is a goal\n\nLong-running commands are **actions** ([SPEC \xA713.6](../SPEC.md#136-composites)): the caller\nsubmits with a client-generated `goalId` and a request fingerprint, the endpoint records a\ndurable accept or reject decision, progress rides per-goal events, and the work ends in one\nterminal outcome (`succeeded`, `failed`, `cancelled`, `expired`, or `uncertain`). Spawn is\nthe reference case. Rather than block the caller for up to 30 seconds while an agent comes\nup, the manager accepts the goal and returns the allocated identity at once:\n\n```json\n{\n \"name\": \"reviewer_2\",\n \"owner\": \"u_...\", \"actor\": \"reviewer_2\", \"uid\": \"...\",\n \"goalId\": \"...\", \"fingerprint\": \"...\",\n \"readinessDeadlineMs\": 30000,\n \"executor\": { \"lifecycleUid\": \"...\", \"epoch\": 3 }\n}\n```\n\nThe `uid` is the lifecycle the agent runs at. On a participant manager whose host enrolls its\nagents, the host picks that uid, so the manager accepts the goal only after the host has answered.\nA host refusal there refuses the spawn, and no goal is bound.\n\nThe name is the one actually allocated: a persona-derived collision is auto-numbered\n(`reviewer`, then `reviewer_2`), while a hard-pinned `--name` that collides with a live\nagent is refused at accept, before anything is minted. Auto-numbering never hands out a numbered\nname it has already issued in that manager process, even after the agent holding it is gone, so\na collision takes the next number. Only numbering consults that history: a hard-pinned `--name`,\nor a persona whose own name is a numbered string, takes that name whenever it is free, and\nnumbering does not skip a string such a spawn held before. The triple plus `goalId` let the\ncaller follow progress (connector handoff, process launched, presence join) and reconcile\nlater against the exact instance that accepted. Presence within the manager's default\n30-second readiness window, or a connector's declared bounded window, settles the goal\n`succeeded`; an early process exit is `failed`; the window passing with neither is `uncertain`,\na bounded, durable outcome that a later `ps` or status read settles against the live roster.\n`uncertain` is a real terminal outcome, not an absence and not a silent hang. It carries the\ndiagnosis of whoever owned the deadline: for a launch that\nnames the agent and says to inspect it rather than re-issue, since re-issuing after a launch\nthat in fact succeeded mints a duplicate. A follower keeps the acceptance as the data of any\nterminal other than `succeeded`, so `cotal_spawn` returns an uncertain launch as a pending result\ninstead of an error: it names the allocated agent, its id, and its manager, and tells the calling\nagent to watch the roster. A committer that supplies no diagnosis falls back to\n\"the success signal did not arrive within the readiness deadline\". The agent's own eventual\nstate is then observable on its presence record.\n\nThe acceptance carries that exact `readinessDeadlineMs`. A synchronous follower treats its own\nrequest deadline as a floor and waits through the accepted readiness budget plus delivery margin,\nso a connector-specific slow boot cannot be reported as a caller timeout while the manager is\nstill legitimately waiting for its terminal.\n\nA spawn that is **refused** because a lifecycle barrier already holds the actor (a frozen\nissuance gate, a retiring alias, a retired uid) is not a wait-timeout. The manager already\nknows the blocked op (`registration` / `retirement` / `activation` / `takeover`), the `opId`\nholding it, and the remedy when one exists (`retry`, `cotal reconcile-gate`). The detail\ncarries `headState` (`active` / `retiring` / `retired`) only when the refusing site read the\nlifecycle head, and `gateState` (`frozen` / `retired`) only when it read the issuance gate. A\ngate frozen by a takeover or a registration says nothing about the head, so that refusal\ncarries `gateState=frozen` and no `headState`. Those facts ride `error.details[]` as\n`kind = ai.cotal.ep.lifecycle-blocked` and are also appended to the error string, so a\ncaller that only prints `error.message` still sees them. The CLI and the connector tools hand a\nrefusal on in one shape, so `cotal spawn -f` keeps the same code, details, rendered facts and\nacceptance data as `cotal spawn --detach` and `cotal_spawn`. A connector that collapses the\nrefusal to \"startup failed (unknown)\" or a SPEC 13.6 wait-timeout is hiding a knowable\nstate, not reporting a missing one.\n\n## Instance routing\n\nA space can run more than one manager. Each manager persists a stable logical instance id\nacross restarts and advances its process epoch when it comes back, so callers address a\nspecific manager without caring which process currently serves it. A start serves only at the\nepoch its own registration committed, never at the epoch of a later start of the same instance.\nOn a static or open mesh,\nan untargeted spawn rides class anycast (any manager may accept, and the acceptance records which one did).\n`cotal spawn <persona> --detach --on <instance>` and `cotal_spawn(instance: \"<instance>\")`\npin one instance by its exact id. A foreground CLI spawn has no manager to pin and refuses the\nflag. An MCP pin that does not resolve is refused without falling back to class anycast. There are no ordinal\naliases and no short forms: wherever a display names an instance you can address, it prints\nthe whole id, because both surfaces take nothing else.\n\nOn a user-auth mesh, manager commands obtain a short-lived `manager-caller` view from the\nexchange. It authorizes one concrete manager instance using the caller's current actor grant and\nthe host's registered service records. Discovery and invocation both use that instance's `inst`\nroute. This view grants no registry scan, class queue, or additional command capability. An absent,\nambiguous or unauthorized selection refuses before the command is sent.\n\nManaged launches carry `COTAL_MANAGER_INSTANCE` so their tools address the manager that launched\nthem. Existing unbound sessions can use the exchange's unique authorized selection without\nreplacing their actor or conversation. The connector uses a separate control connection; the\nstanding message connection and its credential source are unchanged. Accepted spawn goals are\nfollowed on that renewing connection, using its existing caller-scoped progress grant, so a long\nreadiness budget does not depend on the short-lived control credential. The follower confirms its\nprogress subscription with the broker before submitting on the separate connection. A caller still\nchecks the resolved instance and epoch, and never retries an ambiguous mutation outcome.\n\nThe manager's `goal-result` command accepts `{goalId}` and returns `{goalId, result?}`. It reads\nonly the authenticated caller's owner, actor and lifecycle through the manager's separate trusted\ngoal-writer connection. The caller receives an attributed reply, never a raw JetStream reader\ngrant. Each read is admitted by the connection's broker-enforced command grant. A live user-auth\nconnection remains bounded by its bearer expiry after revocation; a renewed connection is checked\nagainst fresh authority. There is no separate per-read ledger check. An absent `result` means no\nterminal is recorded; it does not prove the goal is running or permit another submission. The\nexisting trusted goal-writer's leader-served EPF read is space-wide at the broker; the handler\nconfines it to this endpoint and caller triple.\n\nThe manager's reserved `cancel` command accepts `{goalId, mode?}` and returns `{goalId, state}`.\nIt is served for a turn the manager relays. The goal is the authenticated caller's own, so a caller\nwithdraws only a turn it submitted. The turn ends `cancelled`, its seat is not shown it again, and a\nlater yield of it is answered with that terminal. A goal that already ended is refused\n`failed-precondition` with its cached outcome attached, and a goal this manager does not relay is\nrefused without being changed. So is a second cancel that arrives while a first is still ending the\nturn; a first that fails leaves the turn pending unless something ended it meanwhile. A workflow\nrun sends it for the turn, ask attempt or escalation of a branch it cancelled.\n\nA followed mutation requires a manager whose attributed describe includes `goal-result`. Update\nthe manager, issuer and client together before using that recovery path. Reloading an issuer alone\ncannot change an already-running participant manager. Recovery re-resolves the accepting instance's\nepoch, preserves the caller lifecycle and validates the result against the accepted goal and any\nacceptance fingerprint. Stopping the caller ends its observation, not the already accepted goal.\n\nA followed call resolves the endpoint within its deadline before the submission starts, so a\nrefused or unanswered describe surfaces as its own error. A describe or command publish that the\nbroker refuses reports `not-executed`.\nCancellation before submission reports `not-executed`. Once submission starts, cancellation or a\nlost reply reports an unknown outcome unless an attributed refusal proves otherwise. A received\nrefusal remains a refusal even when stop races it. Local failures do not invent responder identities.\nThe follower owns its subscription, timers and read cancellation signal. Reconciliation begins\nbefore the wait deadline, and late read completions cannot settle an expired observation. Its read\ncallback receives the accepting caller triple, remaining budget and abort signal; borrowed bearer\ncommands and control connections use that signal. An in-flight dial that finishes after cancellation\ncloses without publishing. A local reply-subscription failure prevents publication and is observed\nby the same request promise, including when the transport is closing or draining. Request\ncancellation does not revoke or resubmit the accepted operation.\n\n\"Only one manager per space\" is not the current invariant. A split topology that keeps the\nbroker host manager-free is still a topology choice: `cotal up` on that host starts a\nmanager you then stop with `cotal down manager` after `\u2713 manager up` in\n`.cotal/manager.<spaceKey>.log` (detach stdout listing `manager` is pidfile liveness, not a\nteardown boundary), and `cotal supervise\n--server` runs the manager elsewhere ([Run a mesh](run-a-mesh.md)). Extra live managers\nare addressable, not an error.\n\nThe reserved `describe` bootstrap is the one request the resolver may repeat while waiting: it is\nread-only, it is re-published under the same request binding, and every attempt stays inside the\noriginal deadline. This covers the startup window where Core NATS discards the first request before\nthe manager has subscribed. If the connection closes while the resolver waits, the describe fails\ncleanly instead of throwing from the retry timer. The resolved command is never repeated by this\nreadiness behavior.\n\nThe resolve and the invoke are separate trips through the same anycast queue, so in a\nmulti-manager space an unpinned call can land on an instance the caller did not resolve. Every\ncall carries the incarnation it resolved against, and a manager that is not that incarnation\n**refuses before running the command**, so the failure an operator sees says the command did\nnot run, and re-issuing it cannot duplicate the effect. That is the difference that matters for\na mutation: the older behaviour detected the mismatch on the reply, after the manager had\nalready acted, and could only tell you to go and check. `--on` still matters for reaching a\nspecific manager (`ps`, `stop`, `attach`, `spawn --detach`), but it is no longer what stands\nbetween a split and a duplicated spawn. Against a manager older than this fence the refusal is\nstill after the fact, and its message says so. The re-issue is automatic only when the refusal\nstates `not-executed` in its `outcome` field; a refusal that omits the field, or states\n`unknown`, is surfaced to the caller instead of repaired, because neither proves the command did\nnot run. The CLI's manager commands, `cotal invoke`, the `cotal run` verbs and the manager row of\n`cotal status` re-describe and re-issue an unpinned call after each such refusal, up to 16 times,\nso a split reaches the operator only when every attempt split. A hosted run's own manager calls\nuse the same bound. A pinned call is never re-issued. An agent's own manager\ntools, such as `cotal_spawn` and `cotal_despawn`, re-describe and re-issue with the same bound,\nincluding the goal-result read that follows a spawn to its outcome.\n\nAn unpinned targeted call, such as `cotal_despawn` or a hosted run's turn relay, can also reach a\nmanager that does not host its target, because each manager resolves targets against the agents\nit runs. That manager refuses with `expired` and `not-executed` and says it holds no mapping for\nthe target, and the same re-issue repairs it within the same bound. An agent that no manager hosts\nstill ends in that refusal once the re-issues run out. A pinned call gets the refusal of the\ninstance it named.\n\nA manager whose boot inventory marked every declared connector unavailable does not subscribe\n`spawn` or `launch` on the class `one` rail. Those commands stay on scatter and on this\ninstance's `inst` rail, so a sibling that can launch them can take an unpinned spawn, and a\ncaller that pins this instance with `--on` still gets a named harness refusal. `describe`\nstill lists the commands: the instance rail serves them, and `describe` itself stays on the\nclass rail (SPEC 13.7). An unpinned `spawn` can therefore bind-fence: `describe` may land on\nthe skip member while `spawn` lands on a sibling, the command was not run, and the caller\nre-issues or pins `--on`. `status` reports `classSpawn: false` when that skip is in effect.\nA manager that can launch some connectors keeps the class rail. If the queue hands it a\nharness its inventory marked unavailable, the refusal names `--on` because the standing serve\ncredential cannot read sibling inventories. Pin the capable instance (the whole id, as `ps`\nprints it).\n\n`ps` and\n`status` become a **scatter** across every registered instance: the caller freezes the\nexpected set from the service registry, invokes each under a shared deadline, and merges the\nresults with per-instance attribution. A non-answering instance is labelled as registered\nwith no answer within the deadline, never silently omitted. See [SPEC \xA713.5](../SPEC.md#135-verbs) (scatter) and [cli.md](cli.md).\n\nThe expected set comes from the **registry**, which records registration rather than liveness.\nAn instance that crashes never deregisters, so it stays in the set and the gather has nothing\nleft to wait for but an answer that cannot come. It pays the whole deadline, on every scatter,\nindefinitely. A scatter can therefore be given a per-instance liveness probe: when the broker\nitself reports that an instance holds no subscription on its own instance rail, the gather stops\nwaiting for it. Only that affirmative report counts. A lapsed presence entry, a probe that timed\nout, and a probe that failed are all *absence of evidence*, and treating any of them as death\nwould turn a slow correct answer into a fast wrong one, so they leave the full deadline standing.\nNothing about the outcome changes either way: an instance that did not answer is still\nunreachable, still surfaced, and the scatter is still not complete.\n\nThe probe is supplied by the **caller**, not invented by the scatter. Asking about an instance is\na publish on that instance's rail, and a credential that holds no row for it is refused by the\nbroker asynchronously, while the publish itself returns normally. The probe verb watches for that\nrefusal and raises it as `permission-denied` naming the rail, so it is never mistaken for a quiet\ninstance, and it never burns the probe budget waiting out a refusal. Only the layer that\nminted the credential knows which ids it may ask about, so that layer asks about those and no\nothers. `cotal ps` freezes the class on its first connection, re-mints an instrument pinned only\nto the frozen ids, and scatters on a second; a refusal the broker raises anyway is printed and\nthe instance's row says the probe was refused, which is a fact about the credential, not about\nthe instance.\n\nThis does not help against an instance that is **connected but not answering**. A hung manager\nholds its subscriptions, so it is indistinguishable from a slow one, and it still costs the full\ndeadline. That is the correct result, not a gap in the probe.\n\n### Deregistration\n\nA probe makes a dead registration cheap to skip; it does not remove it. Removal is the\nregistration's own exit, and there are two explicit routes to it\n([SPEC \xA713.5](../SPEC.md#135-verbs): a deleted `svc` spec *is* the deregistration).\n\nA manager that stops cleanly removes its own registration, so an ordinary shutdown leaves no stale\nrow. The delete is pinned to the registration revision that process wrote. When a successor has\nregistered the same instance since then, the stop logs that and leaves the successor's registration\nalone. It refuses that delete while this instance holds the endpoint governance slot at the live\nissuance-gate generation (a registration still completing its reopen). A leftover slot whose\ngeneration is behind that live generation is not in-flight and does not block the stop. A manager\nthat cannot renew or read its lease keeps serving, stays registered, and retries. If another process\nholds the same instance key, that process has taken the instance over, so this one logs the conflict\nand exits without deregistering, leaving the successor's registration alone.\n\nA restart that died *mid-registration* is a different residue: the issuance gate stays frozen under\nthat op. The successor completes the dead registration on boot when the freeze-holder is\naffirmatively gone under a complete CONNZ sweep (the same composition as\n[`cotal reconcile-gate`](cli.md#reconcile-gate)). A committed spec write is finished under that\nsame freeze; only a definite no-commit abort-reopens and then runs the normal takeover.\nIt does not invent a TTL and it does not start a new freeze over a still-held one.\n\nThat residue has a second half, and it is the endpoint governance slot rather than the gate. Every\nregistration takes the endpoint-wide slot before it publishes its spec and holds it until its own\ngate reopens, which is what serializes registration for the endpoint. An instance that stopped\nbetween those two points leaves the slot held with no registration behind it, so the endpoint\nrefuses new registrations while nothing is actually in flight. The slot is stamped with the\ngeneration of the gate its holder had frozen when it took it, and a slot is promoted only at that\nsame generation. So once the holder's gate has reopened past the stamp, the slot can never be\npromoted by anyone, and the next registration for that endpoint replaces it. That reclaim is part of\nan ordinary start and needs no operator step.\n\nA slot whose holder's gate is still at the stamped generation is a registration that is genuinely in\nflight, and it keeps refusing. The two states read differently only in the holder's gate coordinate,\nso reopening that gate is what separates them: the holder's own restart heals it on boot, and\n[`cotal reconcile-gate`](cli.md#reconcile-gate) is the operator's route when the boot path cannot\nrun. The registration path is the slot's only writer, and neither repair command writes it.\nA registration that cannot read the holder's gate at all refuses, because an unreadable gate does\nnot distinguish the two states either. Each of these refusals carries\n`kind = ai.cotal.ep.foreign-slot-held` in `error.details[]` with the holder's instance id and the\n`condition` that refused: `in-flight` for a holder gate still at the stamp, or `no-seam`,\n`unreadable`, `garbled` or `behind` when the registration could not read that gate or read it below\nthe stamp. A remote manager asks its host to reconcile the holder only on `in-flight`, the one\ncondition a gate repair can clear.\n\nFor the instance that cannot cooperate, an operator names it:\n`cotal deregister-instance --instance <id>` ([cli.md](cli.md#deregister-instance)). It removes the\nrecord only on the same evidence `cotal ps` acts on: the broker reporting nothing subscribed on\nthat instance's own rail. It refuses if the instance answers a describe, refuses if the probe could\nnot run at all, and refuses if the instance is merely quiet, because a hung process still holds its\nsubscriptions and is therefore not affirmed gone. It also refuses while that instance holds the\nendpoint governance slot at the live issuance-gate generation (a registration still completing);\na leftover slot behind that generation is not in-flight and does not block. Nothing sweeps the\nregistry on an age threshold or on silence.\nAn instance that is deregistered while it is merely wedged re-registers over the tombstone on its\nnext start, which is what makes the operator's decision a recoverable one.\n\n## Attach sessions\n\n`cotal attach` no longer returns a `ws://127.0.0.1` URL. It creates a one-use, holder-bound\nsession offer: the manager mints a token bound to the caller, the target lifecycle, its own\ninstance id and epoch, and an expiry, and replies with a session id and expiry only, no URL\nand no secret in the reply. The CLI redeems the offer over the mesh (a second redeem is\nrefused). On a registered open mesh that redeem is a bare connection, the same path other\ncontrol commands already use; on a static-auth mesh it is still a session-caller credential\nminted from the resolved root's seed. On a user-auth mesh the CLI holds no seed: it exchanges its\nlogin and the grant for a `session-caller` view bearer, and the callout mints the same caller rails\nwith the grant's expiry. Terminal bytes then stream on core-NATS session subjects\nscoped to the two parties. Backpressure is a bounded in-flight window with an explicit drop notice, never\nsilent loss; a late attach still repaints the full screen from a replayed terminal\nsnapshot. Close, expiry, target despawn, and a manager restart are distinct, surfaced end\nstates: a restarted manager's successor refuses the old epoch's sessions and the client\nshows \"manager restarted; re-attach\".\n\n## Seat input\n\n`attach` is a stream, so it is the wrong shape for a program that wants to send one line: it\nholds a session open and expects a terminal at the caller's end. The `input` command is the\nother half. One authorized call writes text into a running seat's terminal as if it had been\ntyped there, and answers with the seat and the number of bytes delivered.\n\nIt exists for **harness commands**. A line beginning with `/` (`/compact`, `/clear`, `/model`)\nis neither chat nor an event: the agent's own harness handles it, and the keyboard is the only\nway in. An external control surface that can already read a seat's turns and talk to it still\ncannot drive it without this.\n\nThe op is targeted, rides the `manager.lifecycle` capability, and declares authz modes `owner`\nand `any`, the row shape `attach` and `despawn` already carry, checked by the same authorization.\nEnter is appended unless the caller suppresses it, and nothing is echoed back, since the resulting\nturns already have somewhere to go.\n\n**Who may call it is narrower than either of those**, and the reasoning is worth stating because\nthe natural assumption is wrong. `despawn` and `attach` are granted to anything holding `spawn`;\n`input` is granted only to operator credentials. The tempting argument for treating them alike is\nthat an attach session's `write` already reaches the same terminal, so `input` adds nothing. It\ndoes not reach it: an attach yields a signed session offer, and redeeming one needs a per-session\ncredential minted from the space signing seed, which no agent holds. So `input` would be new\nauthority, and the own-owner rule that bounds `despawn` covers every seat under an owner rather\nthan only the ones a caller launched. Killing a peer is denial; typing into a peer is control of\nit. The write therefore sits with the credential that is already the administrative authority for\nthe domain.\n\nOnly a runtime that owns the child's input stream can serve it. The `pty` runtime does; the\nexternal terminal runtimes attach to a process they do not own, and there the command refuses\nand names the runtime rather than dropping the keystroke. A seat that is not running refuses for\nits own reason, and the two are distinguishable, so a caller can tell \"this will never work\"\nfrom \"not right now\". See [cli.md](cli.md#input).\n\n## Grants\n\nThere is no broad control credential. A caller holds one capability row per command it is\nallowed to send, and minting maps each named capability to the request subjects it needs and no\nothers. The manager serves over a scoped serve credential that can answer and\nreply but cannot, for instance, write another endpoint's records or forge a goal terminal;\nthe goal-fact writer and the session writer are separate, narrowly scoped credentials the\nbroker fences by subject. Authorization is checked at the serving boundary, and for actions\nit linearises at acceptance: a spawn refused there mints no reservation and leaves no\nprocess. See [SPEC \xA713.9](../SPEC.md#139-authority-boundary) and\n[identity & auth](identity-and-auth.md).\n\nA carried resume transcript never rides the rails. The operator-only `transcript-receive` command\nanswers whether to upload and hands back a one-time claim for `spawn`, and the bytes travel through\nthe target instance's own transfer bucket under two one-shot credentials: a writer the operator\nmints for that one transcript, and a reader the target instance mints for its own bucket, or that\nthe host issues a remote manager through its `transferReader` authority operation.\n\n## See also\n\n- [Architecture](architecture.md), where the manager and the wire fit in the whole system.\n- [CLI](cli.md), for `describe`, `invoke`, `spawn`, `ps`, `status`, `attach`, and `input`.\n- [SPEC \xA713](../SPEC.md#13-endpoint-control-surface-v04), the normative contract.\n"
|
|
57844
|
+
"body": "# The control surface\n\n> **Concept** (informative) \xB7 **For:** operators and client authors who want to know how the manager and other daemons are driven \xB7 **Normative:** [SPEC \xA713](../SPEC.md#13-endpoint-control-surface-v04)\n\nCotal once had a privileged control rail: a fixed set of named service tiers\n(`self` / `manager` / `admin` / `delivery`) on their own `ctl.*` subjects, with the manager\nas a special case the broker recognised by name. That rail is gone. Everything that serves\nstructured commands now, the manager, the delivery daemon, a wrapped MCP server, a\nthird-party service, is an ordinary **endpoint**: a daemon that registers a service\nidentity, publishes its contracts, and answers `describe`. `manager` is an endpoint name\nlike any other; no subject, envelope, or grant in this surface knows it specially. The\nmanager is a service on the mesh, not an authority over it: it holds only the capability\nrows its callers grant it, and serves over a scoped credential.\n\n## The `ep` rails\n\nOne kind, `ep`, carries every request under a mode token that says where the request\nroutes, never which verb it is (the verb rides the envelope): `one` (queue-group\nanycast, one and only one class member), `all` (scatter, every instance), and `inst` (one instance by its\nstable address). Replies come back on a `reply` rail keyed to the serving instance and its\nepoch. Around these sit the sibling planes the composites use: per-goal events, timers,\nsessions, and the journal that holds durable facts. Every request carries the caller as\nthree forge-locked tokens, `owner`, `actor`, and lifecycle `uid`, plus an unguessable\nnonce, so the broker polices who is calling in the subject grammar itself. See\n[SPEC \xA713.2](../SPEC.md#132-grammar) for the grammar and [\xA713.5](../SPEC.md#135-verbs) for\nthe verbs (`call`, `cast`, `watch`, `claim`, `scatter`).\n\n## Lifecycle identity\n\nA principal `owner.actor` is a reusable routing alias: a despawn frees the actor name and a\nlater spawn may legitimately reuse it, so the alias alone is never authority. Two further\ncoordinates make an identity durable: a **lifecycle uid**, an unguessable, never-reused id\nfor one managed lifecycle under a principal, and a **process epoch**, the fenced ownership\nepoch of the process currently animating it, advanced on every restart or takeover. At most\none live epoch owns an identity, and a superseded epoch must stop serving. Durables and\ncredentials key on the lifecycle uid, not the reusable name, which is what lets a\nsupervised restart recover the same lifecycle instead of minting a new one. See\n[SPEC \xA713.1](../SPEC.md#131-lifecycle-identity) and [identity & auth](identity-and-auth.md).\n\n## Service discovery\n\nNo client has compile-time knowledge of any endpoint's commands. `cotal describe\n<endpoint>` resolves a registered endpoint's command set off the wire: the reserved\n`describe` command answers the registered contract digests, the schemas are fetched from the\nspace's content-addressed contract store, recompiled, and verified against those digests.\nEach command prints with its capability class and targeting shape. `cotal invoke <endpoint>\n<command> --args '<json>'` then calls one command by name, validating the arguments as they\nwill be sent (JSON drops a key whose value is undefined) against the fetched input schema before\npublish. A refusal at that check means nothing was sent, and it is marked `not-executed`. A\nsigned-in user invokes the same surface through their bearer, and the broker enforces each\ncommand's existing capability grant. A manager alias supplied through `--name` resolves through\nits name-keyed `inspect` command, so an authorized targeted call does not need the manager-wide\n`ps` enumeration grant. Every built-in manager command uses this\nsame trust chain, so there is nothing the built-ins can reach that a described contract cannot. The registered\n`auth` endpoint is describable the same way `manager` is: `cotal describe auth` lists\n`retire-lifecycle` and its exact-mode target shape.\nSee [SPEC \xA713.7](../SPEC.md#137-contracts-and-discovery) and [cli.md](cli.md).\n\nThe manager's `resolve-cwd` command is in the `manager.spawn` capability class. It accepts an\nabsolute path on that manager's host and returns its canonical directory plus the host name. It\nrefuses a relative, missing or non-directory path with `failed-precondition`; it creates nothing.\n`spawn` applies the same check at admission, before any credentials or durables are minted.\n\n`spawn` refuses an empty or whitespace-only value in any of its optional string fields, such as\n`role`, `agent` or `identity`, with `bad-request` naming the field. To take the persona file's value,\nomit the field.\n\n### Inspecting a managed name\n\nManager `inspect` keeps its successful response as the live managed-agent row. A live hit does\nnot read durable lifecycle state, so a temporary records-store failure cannot break inspection of\nan agent the manager currently holds.\n\nOn a live miss, a static manager point-reads its durable slot row. A name with no slot, or a slot\nwhose phase is `retired`, remains `not-found`. A nonterminal slot returns\n`failed-precondition` with `error.details[].kind =\nai.cotal.manager.static-slot-observation`. The detail carries the slot's `slotPhase`,\n`owner`, `actor`, `slotLifecycleUid`, `cleanupComplete` when recorded, and `slotRevision`. It\nthen carries the separate lifecycle head's `headState`, `headOp` when present,\n`headLifecycleUid`, and `headRevision`. Head fields are absent when provisioning has not written\nthe lifecycle head yet.\nThe error message carries the same diagnostic summary so string-only operator paths do not hide\nthe structured detail.\n\nA slot row records the manager instance that owns it. In a space with more than one manager, the\nclass queue can hand `inspect` to an instance that does not host the name. When the row names a\ndifferent instance and is not `retired`, the miss returns `failed-precondition` with the same\ndetail plus `ownerInstanceId`, which names the only manager that can act on it. The message names\nboth instances, so a caller that reads only the string can tell it from `not-found`. A sibling's\n`retired` row remains `not-found`. A named `cotal_despawn` resolves its target through this read\nand cannot address an instance, so it asks again until the owning instance answers, up to 16\ntimes.\n\nThe slot is read before the head. These records do not form one atomic snapshot, so the detail\nalso carries `readOrder: [\"slot\", \"head\"]` and `consistency: \"ordered-not-atomic\"`. A head can\nadvance between the reads. The issuance gate is not projected because the retirement operation\nneeded for this diagnosis is already recorded on the head, and reading a third record would add\nanother non-atomic edge without changing the per-name result.\n\nIf either durable read fails or exceeds its bound, the miss returns `unavailable` with\n`ai.cotal.manager.static-slot-read-failed` rather than claiming the name is absent. That detail\nnames the inspected `name`, the failed `record` (`slot`, `head`, or `slot-or-head` when the layer\ncannot distinguish them), and `operation: \"read\"`. User-auth managers do not own `mgrslot` rows,\nso their inspect misses remain live-map reads.\n\nFor Linux custodied seats, retirement requires the runtime's process-exit evidence before\nfreeing the alias or deleting its credentials and delivery state. Socket loss alone is not\nproof of exit. The runtime retains the record captured at launch or adoption so a clean\ncustodian exit can unlink its file without losing the recorded boot and process identities.\nIf the file is missing, reaping uses that retained record and the existing kernel identity\nchecks. An unknown reference without either record refuses cleanup. Reused process ids\nare never signalled on the strength of the old record.\n\n### Listing the durable slots\n\nManager `slots` (`manager.read`, untargeted) lists the durable static slot rows this manager\nowns. Only static managers hold these rows: a user-mode or open manager answers\n`failed-precondition`, and a manager whose durable store is not standing answers `unavailable`.\nEach row carries the same `readOrder` and `consistency` fields `inspect` uses, because the list\nis read the same way: torn across rows as well as within each row's slot/head pair. A `retired`\nrow is never listed. `live` reflects the manager's live roster at render time, not the durable\nrow. A slot row is never deleted, so a row whose latest operation is a DEL or PURGE marker is\ncorruption: the list answers `unavailable` naming that row, as `inspect` does for its name.\n\n## Spawn is a goal\n\nLong-running commands are **actions** ([SPEC \xA713.6](../SPEC.md#136-composites)): the caller\nsubmits with a client-generated `goalId` and a request fingerprint, the endpoint records a\ndurable accept or reject decision, progress rides per-goal events, and the work ends in one\nterminal outcome (`succeeded`, `failed`, `cancelled`, `expired`, or `uncertain`). Spawn is\nthe reference case. Rather than block the caller for up to 30 seconds while an agent comes\nup, the manager accepts the goal and returns the allocated identity at once:\n\n```json\n{\n \"name\": \"reviewer_2\",\n \"owner\": \"u_...\", \"actor\": \"reviewer_2\", \"uid\": \"...\",\n \"goalId\": \"...\", \"fingerprint\": \"...\",\n \"readinessDeadlineMs\": 30000,\n \"executor\": { \"lifecycleUid\": \"...\", \"epoch\": 3 }\n}\n```\n\nThe `uid` is the lifecycle the agent runs at. On a participant manager whose host enrolls its\nagents, the host picks that uid, so the manager accepts the goal only after the host has answered.\nA host refusal there refuses the spawn, and no goal is bound.\n\nThe name is the one actually allocated: a persona-derived collision is auto-numbered\n(`reviewer`, then `reviewer_2`), while a hard-pinned `--name` that collides with a live\nagent is refused at accept, before anything is minted. Auto-numbering never hands out a numbered\nname it has already issued in that manager process, even after the agent holding it is gone, so\na collision takes the next number. Only numbering consults that history: a hard-pinned `--name`,\nor a persona whose own name is a numbered string, takes that name whenever it is free, and\nnumbering does not skip a string such a spawn held before. The triple plus `goalId` let the\ncaller follow progress (connector handoff, process launched, presence join) and reconcile\nlater against the exact instance that accepted. Presence within the manager's default\n30-second readiness window, or a connector's declared bounded window, settles the goal\n`succeeded`; an early process exit is `failed`; the window passing with neither is `uncertain`,\na bounded, durable outcome that a later `ps` or status read settles against the live roster.\n`uncertain` is a real terminal outcome, not an absence and not a silent hang. It carries the\ndiagnosis of whoever owned the deadline: for a launch that\nnames the agent and says to inspect it rather than re-issue, since re-issuing after a launch\nthat in fact succeeded mints a duplicate. A follower keeps the acceptance as the data of any\nterminal other than `succeeded`, so `cotal_spawn` returns an uncertain launch as a pending result\ninstead of an error: it names the allocated agent, its id, and its manager, and tells the calling\nagent to watch the roster. A committer that supplies no diagnosis falls back to\n\"the success signal did not arrive within the readiness deadline\". The agent's own eventual\nstate is then observable on its presence record.\n\nThe acceptance carries that exact `readinessDeadlineMs`. A synchronous follower treats its own\nrequest deadline as a floor and waits through the accepted readiness budget plus delivery margin,\nso a connector-specific slow boot cannot be reported as a caller timeout while the manager is\nstill legitimately waiting for its terminal.\n\nA spawn that is **refused** because a lifecycle barrier already holds the actor (a frozen\nissuance gate, a retiring alias, a retired uid) is not a wait-timeout. The manager already\nknows the blocked op (`registration` / `retirement` / `activation` / `takeover`), the `opId`\nholding it, and the remedy when one exists (`retry`, `cotal reconcile-gate`). The detail\ncarries `headState` (`active` / `retiring` / `retired`) only when the refusing site read the\nlifecycle head, and `gateState` (`frozen` / `retired`) only when it read the issuance gate. A\ngate frozen by a takeover or a registration says nothing about the head, so that refusal\ncarries `gateState=frozen` and no `headState`. Those facts ride `error.details[]` as\n`kind = ai.cotal.ep.lifecycle-blocked` and are also appended to the error string, so a\ncaller that only prints `error.message` still sees them. The CLI and the connector tools hand a\nrefusal on in one shape, so `cotal spawn -f` keeps the same code, details, rendered facts and\nacceptance data as `cotal spawn --detach` and `cotal_spawn`. A connector that collapses the\nrefusal to \"startup failed (unknown)\" or a SPEC 13.6 wait-timeout is hiding a knowable\nstate, not reporting a missing one.\n\n## Instance routing\n\nA space can run more than one manager. Each manager persists a stable logical instance id\nacross restarts and advances its process epoch when it comes back, so callers address a\nspecific manager without caring which process currently serves it. A start serves only at the\nepoch its own registration committed, never at the epoch of a later start of the same instance.\nOn a static or open mesh,\nan untargeted spawn rides class anycast (any manager may accept, and the acceptance records which one did).\n`cotal spawn <persona> --detach --on <instance>` and `cotal_spawn(instance: \"<instance>\")`\npin one instance by its exact id. A foreground CLI spawn has no manager to pin and refuses the\nflag. An MCP pin that does not resolve is refused without falling back to class anycast. There are no ordinal\naliases and no short forms: wherever a display names an instance you can address, it prints\nthe whole id, because both surfaces take nothing else.\n\nOn a user-auth mesh, manager commands obtain a short-lived `manager-caller` view from the\nexchange. It authorizes one concrete manager instance using the caller's current actor grant and\nthe host's registered service records. Discovery and invocation both use that instance's `inst`\nroute. This view grants no registry scan, class queue, or additional command capability. An absent,\nambiguous or unauthorized selection refuses before the command is sent.\n\nManaged launches carry `COTAL_MANAGER_INSTANCE` so their tools address the manager that launched\nthem. Existing unbound sessions can use the exchange's unique authorized selection without\nreplacing their actor or conversation. The connector uses a separate control connection; the\nstanding message connection and its credential source are unchanged. Accepted spawn goals are\nfollowed on that renewing connection, using its existing caller-scoped progress grant, so a long\nreadiness budget does not depend on the short-lived control credential. The follower confirms its\nprogress subscription with the broker before submitting on the separate connection. A caller still\nchecks the resolved instance and epoch, and never retries an ambiguous mutation outcome.\n\nThe manager's `goal-result` command accepts `{goalId}` and returns `{goalId, result?}`. It reads\nonly the authenticated caller's owner, actor and lifecycle through the manager's separate trusted\ngoal-writer connection. The caller receives an attributed reply, never a raw JetStream reader\ngrant. Each read is admitted by the connection's broker-enforced command grant. A live user-auth\nconnection remains bounded by its bearer expiry after revocation; a renewed connection is checked\nagainst fresh authority. There is no separate per-read ledger check. An absent `result` means no\nterminal is recorded; it does not prove the goal is running or permit another submission. The\nexisting trusted goal-writer's leader-served EPF read is space-wide at the broker; the handler\nconfines it to this endpoint and caller triple.\n\nThe manager's reserved `cancel` command accepts `{goalId, mode?}` and returns `{goalId, state}`.\nIt is served for a turn the manager relays. The goal is the authenticated caller's own, so a caller\nwithdraws only a turn it submitted. The turn ends `cancelled`, its seat is not shown it again, and a\nlater yield of it is answered with that terminal. A goal that already ended is refused\n`failed-precondition` with its cached outcome attached, and a goal this manager does not relay is\nrefused without being changed. So is a second cancel that arrives while a first is still ending the\nturn; a first that fails leaves the turn pending unless something ended it meanwhile. A workflow\nrun sends it for the turn, ask attempt or escalation of a branch it cancelled.\n\nA followed mutation requires a manager whose attributed describe includes `goal-result`. Update\nthe manager, issuer and client together before using that recovery path. Reloading an issuer alone\ncannot change an already-running participant manager. Recovery re-resolves the accepting instance's\nepoch, preserves the caller lifecycle and validates the result against the accepted goal and any\nacceptance fingerprint. Stopping the caller ends its observation, not the already accepted goal.\n\nA followed call resolves the endpoint within its deadline before the submission starts, so a\nrefused or unanswered describe surfaces as its own error. A describe or command publish that the\nbroker refuses reports `not-executed`.\nCancellation before submission reports `not-executed`. Once submission starts, cancellation or a\nlost reply reports an unknown outcome unless an attributed refusal proves otherwise. A received\nrefusal remains a refusal even when stop races it. Local failures do not invent responder identities.\nThe follower owns its subscription, timers and read cancellation signal. Reconciliation begins\nbefore the wait deadline, and late read completions cannot settle an expired observation. Its read\ncallback receives the accepting caller triple, remaining budget and abort signal; borrowed bearer\ncommands and control connections use that signal. An in-flight dial that finishes after cancellation\ncloses without publishing. A local reply-subscription failure prevents publication and is observed\nby the same request promise, including when the transport is closing or draining. Request\ncancellation does not revoke or resubmit the accepted operation.\n\n\"Only one manager per space\" is not the current invariant. A split topology that keeps the\nbroker host manager-free is still a topology choice: `cotal up` on that host starts a\nmanager you then stop with `cotal down manager` after `\u2713 manager up` in\n`.cotal/manager.<spaceKey>.log` (detach stdout listing `manager` is pidfile liveness, not a\nteardown boundary), and `cotal supervise\n--server` runs the manager elsewhere ([Run a mesh](run-a-mesh.md)). Extra live managers\nare addressable, not an error.\n\nThe reserved `describe` bootstrap is the one request the resolver may repeat while waiting: it is\nread-only, it is re-published under the same request binding, and every attempt stays inside the\noriginal deadline. This covers the startup window where Core NATS discards the first request before\nthe manager has subscribed. If the connection closes while the resolver waits, the describe fails\ncleanly instead of throwing from the retry timer. The resolved command is never repeated by this\nreadiness behavior.\n\nThe resolve and the invoke are separate trips through the same anycast queue, so in a\nmulti-manager space an unpinned call can land on an instance the caller did not resolve. Every\ncall carries the incarnation it resolved against, and a manager that is not that incarnation\n**refuses before running the command**, so the failure an operator sees says the command did\nnot run, and re-issuing it cannot duplicate the effect. That is the difference that matters for\na mutation: the older behaviour detected the mismatch on the reply, after the manager had\nalready acted, and could only tell you to go and check. `--on` still matters for reaching a\nspecific manager (`ps`, `stop`, `attach`, `spawn --detach`), but it is no longer what stands\nbetween a split and a duplicated spawn. Against a manager older than this fence the refusal is\nstill after the fact, and its message says so. The re-issue is automatic only when the refusal\nstates `not-executed` in its `outcome` field; a refusal that omits the field, or states\n`unknown`, is surfaced to the caller instead of repaired, because neither proves the command did\nnot run. The CLI's manager commands, `cotal invoke`, the `cotal run` verbs and the manager row of\n`cotal status` re-describe and re-issue an unpinned call after each such refusal, up to 16 times,\nso a split reaches the operator only when every attempt split. A hosted run's own manager calls\nuse the same bound. A pinned call is never re-issued. An agent's own manager\ntools, such as `cotal_spawn` and `cotal_despawn`, re-describe and re-issue with the same bound,\nincluding the goal-result read that follows a spawn to its outcome.\n\nAn unpinned targeted call, such as `cotal_despawn` or a hosted run's turn relay, can also reach a\nmanager that does not host its target, because each manager resolves targets against the agents\nit runs. That manager refuses with `expired` and `not-executed` and says it holds no mapping for\nthe target, and the same re-issue repairs it within the same bound. An agent that no manager hosts\nstill ends in that refusal once the re-issues run out. A pinned call gets the refusal of the\ninstance it named.\n\nA manager whose boot inventory marked every declared connector unavailable does not subscribe\n`spawn` or `launch` on the class `one` rail. Those commands stay on scatter and on this\ninstance's `inst` rail, so a sibling that can launch them can take an unpinned spawn, and a\ncaller that pins this instance with `--on` still gets a named harness refusal. `describe`\nstill lists the commands: the instance rail serves them, and `describe` itself stays on the\nclass rail (SPEC 13.7). An unpinned `spawn` can therefore bind-fence: `describe` may land on\nthe skip member while `spawn` lands on a sibling, the command was not run, and the caller\nre-issues or pins `--on`. `status` reports `classSpawn: false` when that skip is in effect.\nA manager that can launch some connectors keeps the class rail. If the queue hands it a\nharness its inventory marked unavailable, the refusal names `--on` because the standing serve\ncredential cannot read sibling inventories. Pin the capable instance (the whole id, as `ps`\nprints it).\n\n`ps` and\n`status` become a **scatter** across every registered instance: the caller freezes the\nexpected set from the service registry, invokes each under a shared deadline, and merges the\nresults with per-instance attribution. A non-answering instance is labelled as registered\nwith no answer within the deadline, never silently omitted. See [SPEC \xA713.5](../SPEC.md#135-verbs) (scatter) and [cli.md](cli.md).\n\nThe expected set comes from the **registry**, which records registration rather than liveness.\nAn instance that crashes never deregisters, so it stays in the set and the gather has nothing\nleft to wait for but an answer that cannot come. It pays the whole deadline, on every scatter,\nindefinitely. A scatter can therefore be given a per-instance liveness probe: when the broker\nitself reports that an instance holds no subscription on its own instance rail, the gather stops\nwaiting for it. Only that affirmative report counts. A lapsed presence entry, a probe that timed\nout, and a probe that failed are all *absence of evidence*, and treating any of them as death\nwould turn a slow correct answer into a fast wrong one, so they leave the full deadline standing.\nNothing about the outcome changes either way: an instance that did not answer is still\nunreachable, still surfaced, and the scatter is still not complete.\n\nThe probe is supplied by the **caller**, not invented by the scatter. Asking about an instance is\na publish on that instance's rail, and a credential that holds no row for it is refused by the\nbroker asynchronously, while the publish itself returns normally. The probe verb watches for that\nrefusal and raises it as `permission-denied` naming the rail, so it is never mistaken for a quiet\ninstance, and it never burns the probe budget waiting out a refusal. Only the layer that\nminted the credential knows which ids it may ask about, so that layer asks about those and no\nothers. `cotal ps` freezes the class on its first connection, re-mints an instrument pinned only\nto the frozen ids, and scatters on a second; a refusal the broker raises anyway is printed and\nthe instance's row says the probe was refused, which is a fact about the credential, not about\nthe instance.\n\nThis does not help against an instance that is **connected but not answering**. A hung manager\nholds its subscriptions, so it is indistinguishable from a slow one, and it still costs the full\ndeadline. That is the correct result, not a gap in the probe.\n\n### Deregistration\n\nA probe makes a dead registration cheap to skip; it does not remove it. Removal is the\nregistration's own exit, and there are two explicit routes to it\n([SPEC \xA713.5](../SPEC.md#135-verbs): a deleted `svc` spec *is* the deregistration).\n\nA manager that stops cleanly removes its own registration, so an ordinary shutdown leaves no stale\nrow. The delete is pinned to the registration revision that process wrote. When a successor has\nregistered the same instance since then, the stop logs that and leaves the successor's registration\nalone. It refuses that delete while this instance holds the endpoint governance slot at the live\nissuance-gate generation (a registration still completing its reopen). A leftover slot whose\ngeneration is behind that live generation is not in-flight and does not block the stop. A manager\nthat cannot renew or read its lease keeps serving, stays registered, and retries. If another process\nholds the same instance key, that process has taken the instance over, so this one logs the conflict\nand exits without deregistering, leaving the successor's registration alone.\n\nA restart that died *mid-registration* is a different residue: the issuance gate stays frozen under\nthat op. The successor completes the dead registration on boot when the freeze-holder is\naffirmatively gone under a complete CONNZ sweep (the same composition as\n[`cotal reconcile-gate`](cli.md#reconcile-gate)). A committed spec write is finished under that\nsame freeze; only a definite no-commit abort-reopens and then runs the normal takeover.\nIt does not invent a TTL and it does not start a new freeze over a still-held one.\n\nThat residue has a second half, and it is the endpoint governance slot rather than the gate. Every\nregistration takes the endpoint-wide slot before it publishes its spec and holds it until its own\ngate reopens, which is what serializes registration for the endpoint. An instance that stopped\nbetween those two points leaves the slot held with no registration behind it, so the endpoint\nrefuses new registrations while nothing is actually in flight. The slot is stamped with the\ngeneration of the gate its holder had frozen when it took it, and a slot is promoted only at that\nsame generation. So once the holder's gate has reopened past the stamp, the slot can never be\npromoted by anyone, and the next registration for that endpoint replaces it. That reclaim is part of\nan ordinary start and needs no operator step.\n\nA slot whose holder's gate is still at the stamped generation is a registration that is genuinely in\nflight, and it keeps refusing. The two states read differently only in the holder's gate coordinate,\nso reopening that gate is what separates them: the holder's own restart heals it on boot, and\n[`cotal reconcile-gate`](cli.md#reconcile-gate) is the operator's route when the boot path cannot\nrun. The registration path is the slot's only writer, and neither repair command writes it.\nA registration that cannot read the holder's gate at all refuses, because an unreadable gate does\nnot distinguish the two states either. Each of these refusals carries\n`kind = ai.cotal.ep.foreign-slot-held` in `error.details[]` with the holder's instance id and the\n`condition` that refused: `in-flight` for a holder gate still at the stamp, or `no-seam`,\n`unreadable`, `garbled` or `behind` when the registration could not read that gate or read it below\nthe stamp. A remote manager asks its host to reconcile the holder only on `in-flight`, the one\ncondition a gate repair can clear.\n\nFor the instance that cannot cooperate, an operator names it:\n`cotal deregister-instance --instance <id>` ([cli.md](cli.md#deregister-instance)). It removes the\nrecord only on the same evidence `cotal ps` acts on: the broker reporting nothing subscribed on\nthat instance's own rail. It refuses if the instance answers a describe, refuses if the probe could\nnot run at all, and refuses if the instance is merely quiet, because a hung process still holds its\nsubscriptions and is therefore not affirmed gone. It also refuses while that instance holds the\nendpoint governance slot at the live issuance-gate generation (a registration still completing);\na leftover slot behind that generation is not in-flight and does not block. Nothing sweeps the\nregistry on an age threshold or on silence.\nAn instance that is deregistered while it is merely wedged re-registers over the tombstone on its\nnext start, which is what makes the operator's decision a recoverable one.\n\n## Attach sessions\n\n`cotal attach` no longer returns a `ws://127.0.0.1` URL. It creates a one-use, holder-bound\nsession offer: the manager mints a token bound to the caller, the target lifecycle, its own\ninstance id and epoch, and an expiry, and replies with a session id and expiry only, no URL\nand no secret in the reply. The CLI redeems the offer over the mesh (a second redeem is\nrefused). On a registered open mesh that redeem is a bare connection, the same path other\ncontrol commands already use; on a static-auth mesh it is still a session-caller credential\nminted from the resolved root's seed. On a user-auth mesh the CLI holds no seed: it exchanges its\nlogin and the grant for a `session-caller` view bearer, and the callout mints the same caller rails\nwith the grant's expiry. Terminal bytes then stream on core-NATS session subjects\nscoped to the two parties. Backpressure is a bounded in-flight window with an explicit drop notice, never\nsilent loss; a late attach still repaints the full screen from a replayed terminal\nsnapshot. Close, expiry, target despawn, and a manager restart are distinct, surfaced end\nstates: a restarted manager's successor refuses the old epoch's sessions and the client\nshows \"manager restarted; re-attach\".\n\n## Seat input\n\n`attach` is a stream, so it is the wrong shape for a program that wants to send one line: it\nholds a session open and expects a terminal at the caller's end. The `input` command is the\nother half. One authorized call writes text into a running seat's terminal as if it had been\ntyped there, and answers with the seat and the number of bytes delivered.\n\nIt exists for **harness commands**. A line beginning with `/` (`/compact`, `/clear`, `/model`)\nis neither chat nor an event: the agent's own harness handles it, and the keyboard is the only\nway in. An external control surface that can already read a seat's turns and talk to it still\ncannot drive it without this.\n\nThe op is targeted, rides the `manager.lifecycle` capability, and declares authz modes `owner`\nand `any`, the row shape `attach` and `despawn` already carry, checked by the same authorization.\nEnter is appended unless the caller suppresses it, and nothing is echoed back, since the resulting\nturns already have somewhere to go.\n\n**Who may call it is narrower than either of those**, and the reasoning is worth stating because\nthe natural assumption is wrong. `despawn` and `attach` are granted to anything holding `spawn`;\n`input` is granted only to operator credentials. The tempting argument for treating them alike is\nthat an attach session's `write` already reaches the same terminal, so `input` adds nothing. It\ndoes not reach it: an attach yields a signed session offer, and redeeming one needs a per-session\ncredential minted from the space signing seed, which no agent holds. So `input` would be new\nauthority, and the own-owner rule that bounds `despawn` covers every seat under an owner rather\nthan only the ones a caller launched. Killing a peer is denial; typing into a peer is control of\nit. The write therefore sits with the credential that is already the administrative authority for\nthe domain.\n\nOnly a runtime that owns the child's input stream can serve it. The `pty` runtime does; the\nexternal terminal runtimes attach to a process they do not own, and there the command refuses\nand names the runtime rather than dropping the keystroke. A seat that is not running refuses for\nits own reason, and the two are distinguishable, so a caller can tell \"this will never work\"\nfrom \"not right now\". See [cli.md](cli.md#input).\n\n## Grants\n\nThere is no broad control credential. A caller holds one capability row per command it is\nallowed to send, and minting maps each named capability to the request subjects it needs and no\nothers. The manager serves over a scoped serve credential that can answer and\nreply but cannot, for instance, write another endpoint's records or forge a goal terminal;\nthe goal-fact writer and the session writer are separate, narrowly scoped credentials the\nbroker fences by subject. Authorization is checked at the serving boundary, and for actions\nit linearises at acceptance: a spawn refused there mints no reservation and leaves no\nprocess. See [SPEC \xA713.9](../SPEC.md#139-authority-boundary) and\n[identity & auth](identity-and-auth.md).\n\nA carried resume transcript never rides the rails. The operator-only `transcript-receive` command\nanswers whether to upload and hands back a one-time claim for `spawn`, and the bytes travel through\nthe target instance's own transfer bucket under two one-shot credentials: a writer the operator\nmints for that one transcript, and a reader the target instance mints for its own bucket, or that\nthe host issues a remote manager through its `transferReader` authority operation.\n\n## See also\n\n- [Architecture](architecture.md), where the manager and the wire fit in the whole system.\n- [CLI](cli.md), for `describe`, `invoke`, `spawn`, `ps`, `status`, `attach`, and `input`.\n- [SPEC \xA713](../SPEC.md#13-endpoint-control-surface-v04), the normative contract.\n"
|
|
57845
57845
|
},
|
|
57846
57846
|
{
|
|
57847
57847
|
"slug": "define-a-team",
|
|
@@ -57974,7 +57974,7 @@ function loadDocsBundle() {
|
|
|
57974
57974
|
"title": "Upgrading a running deployment",
|
|
57975
57975
|
"kind": "Guide (informative)",
|
|
57976
57976
|
"summary": "Substrate stability tells you what the version numbers promise.",
|
|
57977
|
-
"body": "# Upgrading a running deployment\n\n> **Guide** (informative) \xB7 **For:** operators upgrading a mesh that already exists \xB7 **See also:** [Substrate stability](stability.md), [Run a mesh](run-a-mesh.md), [Identity and auth](identity-and-auth.md)\n\n[Substrate stability](stability.md) tells you what the version numbers promise. This page is the\nother half: what to actually do when the deployment already exists, has credentials in it, and\ncannot simply be recreated. Every release that breaks a running deployment gets a section here,\nnaming what migrates on its own, what does not, and the order to move the pieces in.\n\n## The pre-1.0 upgrade contract\n\nThe packages are pre-1.0, so a minor bump may break an API or an on-disk expectation. Four\ncommitments make that survivable for someone with a fleet:\n\n- **Pin an exact version.** `0.N.P`, never `^0.N.P`. A range can pull a breaking minor in during an\n unrelated reinstall.\n- **Every break that touches a running deployment gets a section on this page**, written in terms of\n what an operator does, not in terms of which module changed.\n- **Read the section before you start, not halfway through.** A section names the work up front\n precisely so the operation does not change shape once it is underway.\n- **A break that cannot be made automatic says so.** Where credentials or state must be recreated by\n hand, the section says which ones and when, rather than leaving you to discover it at the moment\n the first one stops working.\n- **A change to the shape of a credential, or to who may renew one, is breaking whatever the commit\n marker says.** This rule is stated because the marker is a judgement made while writing the code\n and the consequence is felt by someone running it a day later. A fleet that keeps authenticating\n looks compatible and is not, if nothing in it can renew. Any automated check of this rule would\n read commit markers, so a break recorded as a feature is the one case it could not see, which is\n why the rule is written for people first. **The marker held for this release: the 0.49.0 change\n that caused all of this, `36d177951 feat(core)!`, did carry its `!`.** The rule exists for the\n next one that does not.\n\nWhat this page does not promise is a rolling upgrade. Nothing in the current line dual-serves two\nauthority versions, so where broker and manager run separately there is a window in which the mesh\nis down. The sections below give that window's shape so it can be scheduled rather than endured.\n\n## Auth context closure in 0.71.0 (unreleased)\n\nExisting deployments need no credential migration or restart for these additive APIs. Embedded\nhosts can now inspect `handle.connections()` and await `handle.closed` after `close()` or `drain()`\nto prove every owned transport ended, including the callout, replaced readiness readers and\nshort-lived clients. The inventory is a detached snapshot.\n\nA transport close failure now rejects with its connection label. The terminal signal stays pending\nwhile any connection remains live. Repair the failure and retry `close()` before awaiting\n`handle.closed`. Closing one hosted context does not close another account's context.\n\nRead a space's claim with `readPlaneClaim(kv, space)` on that account's leader-only auth bucket.\nAn unclaimed space returns `undefined`; held and released rows retain their claim identity.\nDeleted, malformed and foreign-space rows refuse. `PlaneClaimRow` and `PLANE_CLAIM_KEY` are exported.\n\nUse `observeAccountLivenessWithCreds({ servers, observerCreds, accountId, options })` with the\naccount-scoped membership-observer credential to list that account's connections. It never widens\ncredentials or evicts connections. Zero rows prove absence only with a complete sweep and the\nsingle-server proof. An embedded endpoint's trusted composition can retain transport custody\nthrough `EndpointOptions.onConnection`.\n\n## Hermes model from the environment in 0.68.0\n\nA connector now launches on the model and variant its launcher resolved (the `--model` or\n`--variant` flag, else the agent file's `model:` or `variant:`) and no longer reads them again from\nthe agent file. The Hermes connector also no longer takes a model from `HERMES_MODEL` in the\nenvironment of the process that spawns the seat, including when `spawn.env` lists it.\n\n### What stops working\n\nA Hermes spawn whose only model was `HERMES_MODEL` in the spawning environment is refused at launch,\nand the refusal names both ways to set a model. Spawns that set `--model` or `model:` are unchanged,\non every connector.\n\nCode that calls a connector's `buildLaunch` directly with only `configPath` now gets no model or\nvariant from that file. Pass them as `model` and `variant`.\n\n### Before the upgrade\n\nMove each Hermes seat's model from `HERMES_MODEL` onto its spawn with `--model`, or into its\npersona's `model:`.\n\n## Run answers on a participant manager in 0.68.0\n\nA participant manager now asks its issuing host for an answering credential by naming the run and\nstep it answers. The host reads the pause's token off that run's journal and no longer accepts a\ntoken from the manager. Runs on a mesh with no participant manager are unaffected.\n\n### What stops working\n\nWhile a participant manager and its issuing host run different sides of this release, the host\nrefuses every `cotal run answer` and every amendment that manager serves, because each side refuses\nthe other's request shape. Starting, resuming and reading runs is unchanged. A pause stays waiting\nthrough the window, or follows its timeout if it has one.\n\n### Before the upgrade\n\nUpgrade the auth service and every participant manager registered with it in the same window, then\nanswer the pauses that waited.\n\n## Headless OpenCode handshake in 0.69.0\n\nWith `COTAL_SERVE_HEADLESS=1`, the OpenCode launcher's `[cotal-serve]` line on stdout now carries\nonly `port` and `session`. The server password no longer appears in it, and the 1.x TUI no longer\nreceives the password on its command line.\n\n### What stops working\n\nA headless host that read `password` from that line has no password, and the server refuses its\nrequests. Seats with a TUI, and headless seats that no host drives, are unaffected.\n\n### Before the upgrade\n\nHave each headless host mint a password and pass it to the launcher as `OPENCODE_SERVER_PASSWORD`,\nthen use it for basic auth as before. Without that variable the launcher mints its own.\n\n## Filesystem store identity in 0.69.0\n\nThe delivery daemon's answer to the manager's store check now names a filesystem store by its root\nand by a random `id` that the store records once in `store.id` inside its own directory:\n`.cotal/store.id` for a workspace root, or the directory of the file for `cotal deliver --creds\n<file>`. A manager no longer counts the daemon's store as its own because the two roots have the\nsame path. On a split whose broker host and manager host use one root path, the manager host now\nstays off the daemon-credential renewal lease, so `cotal doctor auth --fix` on the broker host can\nrenew the daemon credentials.\n\n### What stops working\n\nA manager and a delivery daemon on different sides of this release refuse each other's answer to\nthe store check. The manager then remints no daemon credential, and a manager that is booting does\nnot start. This is read from the code and was not measured across two releases. A\n`cotal deliver --creds <file>` whose directory is a read-only mount and holds no `store.id` stops at\nstart. So does a `--creds` file that is its directory's `store.id` under any name, and a `store.id`\nthat is a symbolic link or holds anything but a lowercase UUID.\n\n### Before the upgrade\n\nUpgrade the broker host and every manager host of a space in the same window. For a `--creds` file\non a read-only mount, add a regular `store.id` file beside it that holds a new lowercase UUID and no newline,\nas `node -e 'process.stdout.write(crypto.randomUUID())' > store.id` writes. Move a `--creds` file\nnamed or linked as `store.id` to a file of its own.\n\n## Detached spawns with `--share-tools` in 0.69.0\n\nThe manager's `spawn` operation now takes `shareTools` as a list of MCP server names. The CLI parses\n`--share-tools` into that list before it sends the request, and the manager cluster document moves\nto revision 22. A cut taken with `cotal down --preserve-state` before the upgrade still resumes: the\nmanager reads its `cotal-manager-resume/v1` inventory and writes new cuts as\n`cotal-manager-resume/v2`.\n\n### What stops working\n\nA CLI and a manager on different sides of this release refuse a detached spawn that passes\n`--share-tools`, because the CLI checks each request against the contract the manager serves. This\nis read from the code and was not measured across two releases. A detached spawn without the flag,\na foreground spawn and a roster entry are unaffected. A manager older than this release cannot\nresume a cut that this release took.\n\n### Before the upgrade\n\nUpgrade the CLI on every host that runs `cotal spawn --detach` in the same window as the managers\nit reaches.\n\n## Shared MCP server checks in 0.69.0\n\nThe cotal config reader now checks each server under `connectors.<name>.mcpServers` when it reads\nthe file, and refuses one that cannot launch as written, naming the file and the field. The rules\nare in [the config file](config.md#the-config-file).\n\n### What stops working\n\nA config file that holds such a server refuses every Claude spawn that reads it, including one with\n`--share-tools none`. Before, a field of the wrong type failed each Claude spawn that shared the\nserver with a `TypeError` that named neither the file nor the server, a spawn that did not share it\nlaunched, and a server with no `command` or `url` was passed to `claude`, which never started it.\nRead from the code and not measured: spawns on other connectors, a manager resume and the step of\n`cotal setup` that records the shared list read the same files, so each stops at the same refusal.\n\n### Before the upgrade\n\nCheck `connectors.<name>.mcpServers` in the operator-level config file and in each space's\n`.cotal/config.json`. Give each server a string `command`, or a `type` of `http`, `sse` or `ws` with\na string `url`. Write `args` as a list of strings and `env` and `headers` as objects of strings, or\nremove the server.\n\n## Remote manager family eviction in 0.69.0\n\nA remote manager registered through its host now asks the host to evict up to 256 holders of its\ncredential family in one maintenance request, and the host reads the family once for the whole set.\nBefore, a restart sent one request per holder and the host read the whole family for each one.\nMeshes with no remote manager are unaffected.\n\n### What stops working\n\nWhile a remote manager and its issuing host run different sides of this release, each side refuses\nthe other's eviction request shape. A restart whose credential family already has holders then fails\nat its eviction step and leaves the manager's registration gate frozen. A first start, a clean stop\nand the host's reconciliation of a foreign slot holder are unchanged.\n\n### Before the upgrade\n\nUpgrade the auth service and every remote manager registered with it in the same window. A manager\nthat restarted inside the window resumes its frozen registration on its next start once both sides\nrun this release.\n\n## AG-UI emitter holder hooks in 0.69.0\n\n`AguiEmitterHolder` from `@cotal-ai/connector-core` now takes its hooks as one named object after\nthe emitter factory: `new AguiEmitterHolder(startEmitter, { onError, onRunClosed, waitLive, runMeta })`.\nOnly `onError` is required. Nothing about a running mesh changes, and every shipped connector passes\nits hooks by name. Only a connector of your own that builds a holder is affected.\n\n### What stops working\n\nA holder built with positional hooks, such as `new AguiEmitterHolder(start, onError, onRunClosed)`,\nno longer compiles, because the constructor takes two arguments. Plain JavaScript that keeps the\npositional form still runs, but the holder calls none of its hooks, so a failure never reaches\n`onError`.\n\n### Before the upgrade\n\nPass each hook by name, for example `new AguiEmitterHolder(start, { onError, onRunClosed })`, and\ndrop any `undefined` that filled an earlier slot to reach a later hook.\n\n## Worker run failure type in 0.70.0\n\n`WorkerRunFailed`, the failed result of `runInWorker` in `@cotal-ai/lang`, is now a union on\n`class`: `released`, `held`, `effect`, `too-large`, `rejected` or `error`. A running mesh needs\nnothing, because the runtime host and the engine thread ship in the same install. A run on the\ncompiled engine whose program throws an object with `code: \"L5012\"` or `code: \"L5025\"` used to end\nreleased and now ends failed, as it does on the walker.\n\n### What stops working\n\nTypeScript code that reads `code`, `reason`, `step`, `pending`, `kind`, `detail` or `tooLarge` on a\n`WorkerRunFailed` it has not narrowed fails with TS2339. A released, held, too-large or rejected\nresult no longer carries `code`, so JavaScript that branched on `L5012`, `L5025`, `L5006` or\n`L5010` stops matching with no error. `tooLarge` is gone.\n\n### Before the upgrade\n\nBranch on `class` where such code read `code`: `released` for L5012, `held` for L5025, `too-large`\nfor L5006 and `rejected` for L5010. An `effect` or `error` result keeps its `code`. Once narrowed to\n`too-large`, a result carries the `stepKey`, `bytes` and `bound` that `tooLarge` held.\n\n## Remote manager request builder in 0.70.0\n\n`remoteManagerClient.remoteManagerAuthorityRequest` from `@cotal-ai/manager` now takes an\noperation's coordinates as one object, and `remoteManagerRegistrationProof` from `@cotal-ai/core`\ncomputes the proof from the manager's identity state instead of a request. Nothing about a running\nmesh changes: the proof digest and the request on the wire are the same, so a manager and a host on\ndifferent sides of this release still accept each other. Only code that builds remote manager\nrequests itself is affected, in TypeScript and in plain JavaScript.\n\n### What stops working\n\nA call that passes the registration proof, contract artifacts, session, retirement or transfer\nreader as positional arguments after the operation no longer compiles. A call that passes a request\nto `remoteManagerRegistrationProof` no longer compiles either, because the second argument now names\nthe lifecycle `lifecycleUid`, as the identity state does.\n\nPlain JavaScript runs both old calls without an error. The builder drops the positional coordinates,\nso the host refuses the request with `requires a sha256 registrationProof`. A proof computed from a\nrequest leaves out the lifecycle, so the host refuses a request that carries it as a proof mismatch.\n\n### Before the upgrade\n\nName the coordinates, for example\n`remoteManagerAuthorityRequest(state, \"cli\", \"retire\", { registrationProof, retirement })`.\nCompute the proof as `remoteManagerRegistrationProof(owner, state)`, adding the contract artifacts\nas a third argument for activation only. A host that recomputes the proof from a received request\npasses `{ space, instanceId, lifecycleUid: managerLifecycleUid, identities }` from that request.\n\n## Bearer validator lifetime cap in 0.70.0\n\n`validateUserToken` from `@cotal-ai/auth` no longer takes `maxTtlSec`. It caps a bearer's lifetime\nat the cap of the bearer's view, the same cap the issuer applies when it mints: 900 seconds, or 300\nfor a `transfer-writer` bearer. The auth callout never passed the option, so a running mesh behaves\nas before. Only code of your own that calls the validator with `maxTtlSec` is affected.\n\n### What stops working\n\nA call that passes `maxTtlSec` in an object literal no longer compiles. Plain JavaScript that keeps\nit still runs, and the value is ignored. A `NaN` value, such as `Number()` of an unset environment\nvariable, used to turn the lifetime check off and accept a bearer of any lifetime. That bearer is\nnow refused at its view's cap.\n\n### Before the upgrade\n\nRemove `maxTtlSec` from each call. A test that needs a bearer to expire sooner mints one with a\nshorter lifetime.\n\n## Persisted identity records in 0.70.0\n\nThe manager instance identity, the manager sibling identities, the auth plane instance identity and\na participant manager's remote authority state now share one reader and one first mint in\n`@cotal-ai/workspace`, exported as `claimIdentityRecord` with the nkey check `identityOf`. Each\nrecord is read as a regular file, must hold non-empty nkeys and is created exclusively, so\nconcurrent first starts of a participant manager on one root now settle on one identity where each\nused to keep its own. `saveManagerInstanceIdentity` and `saveAuthInstanceIdentity` are gone. A\nrunning mesh whose records are plain files needs nothing.\n\n### What stops working\n\nA manager instance, auth instance or remote authority record that is a symlink, a directory or any\nother non-regular entry is refused where it used to be followed. The manager, the auth plane and a\nparticipant manager fail to start on it, and `cotal reconcile-gate` and `cotal deregister-instance`\nrefuse it. Retirement already refused it. A remote authority record with an empty nkey id or seed\nis refused too. A first mint that loses its race and cannot read the winner now refuses with\n`identity-record-create-lost` in place of `manager-instance-identity-create-lost` or\n`auth-instance-identity-create-lost`. Code that imports either `save` function no longer compiles.\n\n### Before the upgrade\n\nReplace a symlinked identity record with a copy of the file it points to. Code that wrote a record\nwith a `save` function plants it with `createManagerInstanceIdentity` or\n`createAuthInstanceIdentity`, which create the record when it is absent and otherwise return the\nstored one unchanged. Nothing replaces an overwrite of a stored identity.\n\n## Manager instance in user credentials in 0.70.0\n\n`AuthProvider.userCredentials` from `@cotal-ai/core` no longer returns `managerInstanceId`. A\n`manager-caller` credential's manager instance is the signed `act.managerInstanceId` claim in its\nbearer, which the broker verifies and the CLI already used. The reference provider in\n`@cotal-ai/auth` stops copying the exchange response's field into its result, where nothing\ncompared it with the bearer. The exchange still answers with the field, so a running mesh behaves as\nbefore.\n\n### What stops working\n\nCode of your own that reads `managerInstanceId` from a `userCredentials` result no longer compiles,\nand plain JavaScript reads `undefined` there.\n\n### Before the upgrade\n\nRead the instance from the bearer's `act.managerInstanceId` claim.\n\n## Auth plane identity location in 0.70.0\n\nThe user-auth service keeps its instance identity in the root's `.cotal/space.<hex>/auth-instance.json`,\nbeside the manager's. It used to sit inside `.cotal/auth`, at\n`space.<hex>/.cotal/auth/auth-instance.<hex>.json`, so a copy of that folder carried it. The first\nstart of an upgraded root moves the record and keeps the instance. A hosted context started through\n`startAuthService` has its record moved the same way inside its `stateDir`.\n\n### What stops working\n\nCode that calls `openAuthAuthorityPlane` without the new `identityRoot` option no longer compiles. A\nstart that finds a record both in `.cotal/space.<hex>/` and at its older place refuses and names the\ntwo files. A start also refuses when the older place of the auth or manager identity holds a symlink,\na directory or anything else that is not a regular file. The manager used to skip a dangling symlink\nthere and mint a new identity.\n\n### Before the upgrade\n\nPass `identityRoot` to `openAuthAuthorityPlane`. When `dir` is a workspace root's user-auth state\ndir, `<root>/.cotal/auth/space.<hex>`, pass that root. A plane with no workspace root, as\n`startAuthService` runs, passes `dir` itself. Either keeps the identity the plane already has: on the\nfirst start it moves from `<dir>/.cotal/auth/` to `<identityRoot>/.cotal/space.<hex>/`. Never pass a\ndirectory inside `.cotal/auth`: the record would land in the folder an operator copies and travel\nwith it again.\n\nA copy of `.cotal/auth` taken from a root last run by an older Cotal carries that root's record.\nDelete `.cotal/auth/space.<hex>/.cotal/auth/auth-instance.<hex>.json` from the root you copied it to\nbefore the first `cotal up --user-auth` there.\n\n## Per-seat `COTAL_` names in `spawn.env` in 0.71.0\n\n`spawn.env` in the cotal config no longer forwards a `COTAL_` name the launcher sets for each seat,\nsuch as `COTAL_ROLE`, `COTAL_MODEL` or `COTAL_SUBSCRIBE`. Before, a seat launched with no value of\nits own took the spawning process's value and ran under that role, model or read set. The\nmachine-wide knobs a seat already receives, such as `COTAL_HOME`, may still be listed.\n\n### What stops working\n\nEvery spawn and resume under a config whose `spawn.env` lists such a name is refused before\nlaunch, and the refusal names the entry. Code that calls `launchEnv` from `@cotal-ai/connector-core`\nwith such a name in `envAllow` gets the same error.\n\n### Before the upgrade\n\nRemove those names from `spawn.env`. Give each seat its role, model and channels with `--role`,\n`--model` and `--subscribe`, or in its persona's `role:`, `model:` and `subscribe:`.\n\n## Role addresses in 0.71.0\n\nA role must be one `[A-Za-z0-9_-]` token. Before 0.71.0 any other spelling was rewritten into one:\n` probe ` reached the `probe` queue and `pro.be` reached `pro_be`, while the message kept the\nspelling sent. An anycast to `*` was accepted and stored where no holder reads it.\n\n### What stops working\n\nAn agent whose role is outside the token set no longer starts, however it is launched:\n`cotal join --role`, `cotal spawn --role`, an agent file's `role:`, `COTAL_ROLE` and an embedded\nendpoint's `card.role` are all refused before the agent joins.\n\nA send to such a role, or to `*`, through `cotal send ask`, `/anycast` or `cotal_anycast` is refused,\nand nothing is stored.\n\n`routeToken` is no longer exported from `@cotal-ai/core`. A role routes as spelled, so code that\nused it to name a role's queue uses the role itself, and `assertValidRole` checks one.\n\n### What migrates on its own\n\nEvery task queue. A `svc_<role>` durable was always named from the rewritten token, so its pending\nrequests and its holders carry over.\n\n### Before the upgrade\n\nRename each role outside the token set to the token it already routed to: remove the surrounding\nspaces and replace every other character outside the set with `_`. Rename it where the holder is\nlaunched and in every script or prompt that sends to it.\n\n## Carrying a resumed Claude session to another host in 0.67.0\n\n`cotal spawn --resume <id> --detach --on <instance>` now carries a Claude session held on the\noperator's host to the target manager instance. Both sides need this release: an older manager does\nnot serve `transcript-receive`, and the CLI then stops with that manager's refusal instead of\nlaunching. The manager cluster document moves to revision 21, and the `ps` row's `resume` object\ngains `host` and `transferredAt`.\n\nA manager host that runs carried seats needs `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_AUTH_TOKEN` or a\ncloud provider selection in its environment, because each carried seat runs in its own Claude home\nwith no stored login. On an authenticated mesh the CLI mints the transfer writer from the space's\nsigning seed, so the carrying host needs that seed, as for any other operator command. On a user-auth\nmesh it exchanges the operator's login for a `transfer-writer` view instead, so the operator's grant\nneeds scope `admin`, and the auth service must run this release. A remote manager receives a carry once\nits host serves the manager-service `transferReader` operation. A seat launched without carrying,\nincluding any `--resume` whose id this host does not hold, is unchanged.\n\n## Lifecycle head type in 0.67.0\n\n`LifecycleMapping`, the type `parseLifecycleHead` returns, is now a union on `state`. Nothing about\na running mesh changes: heads that parsed before parse the same way, and the refusals are\nunchanged. Only TypeScript code that compiles against `@cotal-ai/core` is affected.\n\n### What stops working\n\nAn `interface` that extends `LifecycleMapping` fails with TS2312, because an interface cannot extend\na union. Code that builds a head in memory no longer compiles when the head is `retiring` without\nits `op`, or `active` or `retired` with one. The parser already refused those heads.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type ActiveMapping = LifecycleMapping & { state: \"active\" }`. A reader that has checked\n`state === \"retiring\"` reads `op` without a guard.\n\n## Issuance gate types in 0.67.0\n\n`EpGateRow` and `EndpointGateRow`, which `parseIssuanceGate` and `parseEndpointGate` return, and\n`EpGateState`, which an `EpIssuanceGate` or `EpIssuanceBarrier` returns from `observe`, are now\nunions on `state`. Nothing about a running mesh changes: gates that parsed before parse the same\nway, and the refusals are unchanged. Only TypeScript code that compiles against `@cotal-ai/core`\nis affected.\n\n### What stops working\n\nAn `interface` that extends one of these types fails with TS2312, because an interface cannot\nextend a union. Code that builds a gate in memory, such as a custom barrier's `observe`, no longer\ncompiles when the gate is `frozen` or `retired` without its `op`, or `open` with one. The gate\nparsers already refused those rows.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type CustomGateRow = EpGateRow & { custom: string }`.\n\n## Lifecycle-blocked refusals in 0.66.0\n\nA refusal that carries `ai.cotal.ep.lifecycle-blocked` now reports only the lifecycle state it\nread. Nothing about a running mesh changes. A client that branches on the detail must read the new\nfield.\n\n### What stops working\n\nA refusal raised at the issuance gate used to carry `headState` without reading the head:\n`retiring` for a frozen gate and `retired` for a retired one. It now carries `gateState`\n(`frozen` or `retired`) and no `headState`. A client that treats `headState: \"retired\"` as a\nburned uid, or `headState: \"retiring\"` as a retirement in flight, no longer matches those\nrefusals, and the `[lifecycle ...]` suffix on the error string changes the same way. A custom\nissuance barrier whose `observe` returns a frozen gate without a valid `op` (a string `opId` and\none of the four op kinds) is now refused as `internal` by `registerServiceInstance`.\n\n### Before the upgrade\n\nUpdate such a client to read `gateState` for a gate refusal and `blockedOp` for the operation that\nholds the gate. `headState` is present only when the refusal read the head, for example an\nactivation refused because the head is still retiring.\n\n## Workflow programs that bind `once` in 0.65.0\n\n`once` is now a scope of the workflow language, so it is a reserved name. A program that declares\nits own `once` binding (`const once = ...`, a parameter or a function named `once`) is refused at\nvalidation with L2002. Nothing else about a running mesh changes.\n\n### What stops working\n\nA run whose recorded program binds `once` cannot be resumed after the upgrade, because a resume\nvalidates the recorded program again. A new `cotal run start` of such a program is refused before\nanything is recorded.\n\n### Before the upgrade\n\nList the runs with `cotal run ps` and check each program that is still running or held for a\nbinding named `once`. Let those runs finish on the old version before you upgrade the manager, and\nrename the binding in the program before you start it again.\n\n## From 0.58.0 to 0.59.0\n\nEvery connector now publishes a failed run's `RUN_ERROR` on `events.<owner>.<actor>` with the fixed\nmessage `run failed` and no `code` or `rawEvent`. The error text and error kind a harness reports\ncan echo a prompt, a peer message or tool output, and that channel has a different read ACL. A\nreader that showed the message or branched on `code` gets neither after the upgrade. Where a\nconnector reports the error kind as the agent's presence condition, that is unchanged.\n\n### Settle pending event frames before the upgrade\n\nEach session's events are frozen in its event write-ahead log before they are published. A session\nrestarted on 0.59.0 whose log still holds an unacknowledged frame with an older `RUN_ERROR` does not\nrepublish it: its event emitter halts with `egress-run-error` and publishes nothing further for that\nsession. The broker may or may not already hold that frame, so the halt cannot settle it.\n\n1. Stop the seats cleanly on 0.58.0, with the broker still up.\n2. List the logs that still hold a pending frame. The logs live under the events state root\n (`COTAL_WORKSPACE_ROOT`). Empty output means there is nothing to settle.\n\n ```sh\n find \"$COTAL_WORKSPACE_ROOT/.cotal/events\" -name wal.json \\\n -exec jq -r 'select(.pending != null) | input_filename' {} +\n ```\n\n3. For each session listed, start it again on 0.58.0 while the broker is reachable, let it recover,\n stop it, and run step 2 again. Recovery publishes the frame as 0.58.0 would have, error text\n included, so it only finishes what 0.58.0 had already started.\n\n If that start halts with `cas-loss` instead, the agent's subject is no longer at the sequence this\n log expects, and no restart settles that log, on 0.58.0 or later. A lost acknowledgement is one\n cause: the broker stored the frame, so it and its error text are already on the channel, and every\n retry halts the same way because the stream checks the frozen expectation before it deduplicates.\n The halt message names the other causes, such as a second emitter for the same agent under a\n different state root, a restored stream or frontier record, or a purged channel. With those the\n pending frame may never have reached the broker, so a `cas-loss` does not tell you whether it\n landed. Find and stop any second writer and rule out a restored state first. Clearing the halt\n then means purging the agent's event channel and removing the agent's directory under the events\n state root whole (see [Event plane](connect-claude.md#event-plane)). That abandons the pending\n frame whether or not the broker has it, and the purge also drops the earlier frames of every\n session of that agent.\n4. Upgrade once step 2 prints nothing.\n\nIf a session halts with `egress-run-error` after the upgrade, go back to step 3 for that session on\n0.58.0. Do not edit or delete `wal.json` on its own to get past either halt: clearing the pending\nframe abandons that epoch, an event the broker never received is lost, and removing part of the\ndirectory leaves a state the next start refuses.\n\n## Explicit actor grants in 0.59.0\n\n`cotal actor grant` no longer fills an omitted ACL flag with its wide default. A grant names\n`--scope`, `--allow-subscribe` and `--allow-publish`, or passes `--full` to give the ones it leaves\noff their wide defaults (`spawn,role:default`, `>` read, `>` post). Any other grant is refused. The\nbreak is in the CLI on the machine that holds the actor ledger, the one that ran\n`cotal up --user-auth --idp <url>`. No stored row, credential or wire message changes.\n\n### What keeps working\n\nExisting actor ledger rows keep the authority they were granted, and their users and agents connect\nas before. `actor revoke`, `actor list` and a `grant` that names all three ACL flags behave as they\ndid on 0.58.0. Nothing on disk is converted.\n\n### What stops working\n\nA grant that leaves off any of the three flags without `--full` exits 1 with\n`refusing to grant \"<actor>\" with --scope, --allow-subscribe, --allow-publish left off`, naming the\nflags it is missing, and then prints both accepted forms. It writes no row and does not retire the\nactor's current lifecycle. An existing row stays as it was, and an actor granted for the first time\nstays out until the grant is run again. This includes the bare grant printed on 0.58.0 by\n`cotal login`, `cotal status`, `actor list` and the not-granted refusal. Look for it in provisioning\nscripts, onboarding runbooks and anything that pastes those hints.\n\n### Upgrade order\n\nChange the scripts before the ledger machine is upgraded, and make each grant name all three flags.\n0.58.0 and 0.59.0 both accept that form. To keep a wide row, write its defaults out:\n\n```sh\ncotal actor grant <actor> --sub <IdP subject> \\\n --scope spawn,role:default --allow-subscribe '>' --allow-publish '>'\n```\n\nSwitch to `--full` only once the ledger machine runs 0.59.0. 0.58.0 refuses it with\n`Unknown option '--full'` before it reads the ledger. Brokers, managers and participant machines\nneed nothing for this break, so their order is the one the section above gives.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused grant changes nothing. The\nexposure is a grant script that runs against 0.59.0 before it was changed: it fails and grants\nnothing.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. On the ledger machine, save the output\nof `cotal actor list` to compare rows after the changed scripts run, and list the scripts that call\n`cotal actor grant`.\n\n### The upgrade end to end\n\n```sh\n# on the ledger machine, still on 0.58.0\ncotal actor list > actors-before.txt\ngrep -rn 'actor grant' <your provisioning scripts>\n# make every grant name --scope, --allow-subscribe and --allow-publish, run them, then upgrade\nnpm i -g cotal-ai@0.59.0\ncotal actor list | diff actors-before.txt -\n```\n\nBoth refusals quoted here were run on 0.58.0 and on the 0.59.0 code. That brokers, managers and stored\nrows need nothing is read from the change, which touches only the CLI and its hints, and was not run\non a live split deployment.\n\n## Repeated flags refused in 0.59.0\n\nA `cotal` flag given more than once is now a usage error unless the command declares it repeatable.\nOn 0.58.0 the last value won with no message, so `cotal down web --space a --space b` acted on `b`\nwhile a wrapper that checked the first `--space` verified `a`. The break is in the command-line\nparser on the machine that runs the command, including commands added with `cotal ext add`. No\nstored state, credential or wire message changes.\n\n### What keeps working\n\nA command line that gives each flag once parses as it did on 0.58.0, in any order and in the\n`--flag=value` form. Flags whose help says repeatable, such as `--opt` and `down --session-store`,\nstill collect every value. A flag-shaped word after `--` is still a positional. The daemons, units\nand agents that `cotal` starts for itself are given each flag once, so a fleet driven only by `cotal`\ncommands typed by hand needs no action.\n\n### What stops working\n\nA command line that repeats any other flag exits 1 before the command runs. It prints\n`Option '--space' cannot be repeated`, or `Option '-f, --file' cannot be repeated` for a flag with a\nshort form, followed by the command's help. `-f` and `--file` count as the same flag. Look for it in\nscripts, aliases and wrappers that append a flag to override one set earlier, such as a fixed\n`--space` followed by `\"$@\"`.\n\n### Upgrade order\n\nChange those scripts first so each flag is given once. 0.58.0 and 0.59.0 both accept that form.\nBrokers, managers and participant machines need nothing for this break, and each machine's CLI\napplies it when that machine is upgraded, so their order is the one the sections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused command does nothing. The\nexposure is a script that still repeats a flag when it runs on 0.59.0: it exits 1 instead of acting on\nthe last value.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the scripts, aliases and wrappers\nthat call `cotal` so each one can be checked.\n\n### The upgrade end to end\n\n```sh\n# still on 0.58.0\ngrep -rn 'cotal ' <your scripts and wrappers>\n# give each non-repeatable flag once, then upgrade\nnpm i -g cotal-ai@0.59.0\n# run each changed script; a repeat left behind exits 1 with the usage error and does nothing\n```\n\nThe refusal and its messages were run against the 0.59.0 parser and `cotal topology view`. That the\nargument lists `cotal` builds for its own processes give each flag once is read from the code, and\nwas not run on a live split deployment.\n\n## Detached spawns from a seat's shell in 0.62.0\n\nOn a static or open mesh, `cotal spawn --detach` run inside a managed seat's shell now launches as\nthat seat. On 0.61.0 it minted a one-shot operator instrument, so the manager recorded that\ninstrument as the spawner and the seat's own `cotal_despawn` of the child was refused with\n`not authorized: <seat> was not spawned by <caller> (admin tier required)`. The break is in the CLI\non the machine where the seats run. No stored state, credential or wire message changes.\n\n### What keeps working\n\n`cotal spawn --detach` from an operator terminal or from a script outside any seat launches as\nbefore, and so does any call with `--creds`, one aimed at a space other than the seat's own, or a raw\nopen target named with `--server` and an unregistered `--space`. A user-auth mesh is unchanged. A seat with\n`capabilities: [spawn]` still spawns from its shell, and can now stop that child with\n`cotal_despawn`. `--on <instance>` from a seat's shell still lands on that manager instance, now as\nthe seat.\n\n### What stops working\n\n- On a static mesh, a seat without `capabilities: [spawn]` can no longer spawn from its shell. Its\n own credential holds no spawn subject, so the broker refuses the request and the command exits 1.\n- A child launched from a seat's shell is now that seat's child, so the manager stops it when the\n seat exits, as it does for a `cotal_spawn` child. A child that has to outlive the seat that\n started it now goes with the seat.\n- A seat launched without `COTAL_SPACE` is placed by its static credential. Every connector sets\n that variable, so this only reaches a hand-built launch: from such a seat's shell, a spawn aimed at\n a static space that holds no credential for the seat is refused instead of running as the operator.\n\n### Upgrade order\n\nOnly the CLI that seats run from their shell changes, which is the one installed on the host where\nthe seats run. Brokers and managers need nothing for this break, so their order is the one the\nsections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it. A child already running when you upgrade\nkeeps the spawner the manager recorded at its launch.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the agent files whose seats run\n`cotal spawn --detach` from their shell, note which of them lack `capabilities: [spawn]`, and note\nwhich of their children must outlive the seat.\n\n### The upgrade end to end\n\n```sh\n# still on 0.61.0: find the seats that spawn from their shell\ngrep -rln 'cotal spawn' .cotal/agents\n# add `capabilities: [spawn]` to each of those agent files that lacks it, and launch any child\n# that must outlive its seat from an operator terminal instead\nnpm i -g cotal-ai@0.62.0\n```\n\nThe attribution, the despawn, the refusal of a seat without `spawn`, the stop on seat exit and a\nseat's `--on` spawn were run on a local static mesh, and the attribution and the despawn on a local\nopen mesh.\n\n## From 0.53.0 to 0.54.0\n\nManager calls now borrow an instance-bound `manager-caller` credential. Followed mutations require\n`manager.goal-result` on the selected manager, so a compatible issuer, manager and client must be\nloaded together. An older manager is refused before a followed mutation; upgrading an installed\nbinary alone does not replace code in a running manager, connector or embedded client.\n\n### Preserve state before changing processes\n\nSnapshot the broker's durable storage using its supported backup procedure, the host authority and\nactor ledgers, and each participant's manager identity, runtime custody records, credentials and\nsaved sessions. Include the embedding application's database and configuration under its supported\nbackup procedure. Record the loaded package versions and the CLI path used by bearer helpers.\nKeep these copies private. Do not change the IdP issuer, regenerate manager identities, rotate agent\ncredentials or recreate tenant storage to make the upgrade pass.\n\nNo ledger, goal-history or session conversion is required for this change. Existing ordinary\nmessaging credentials retain their normal expiry rules. New manager-caller credentials are obtained\non demand from the current grant; old manager-call credentials do not gain the new view automatically.\nExisting accepted goals remain durable and must not be submitted again merely because observation\nwas interrupted. Fresh remote registration publishes its service status at the current revision and\nepoch; do not seed that status manually.\n\n### Upgrade the split deployment\n\n1. Stage one pinned 0.54.0 package set for the host and participants, including the embedding SDKs.\n Pause new manager mutations and let accepted work settle where possible before reloading processes.\n2. Upgrade the host issuer and embedding first. Keep the broker, its account identities and durable\n storage in place. Then load the matching manager release on each participating machine.\n3. Preserve active seats through the runtime's supported update path. A Linux custodial runtime may\n release and re-adopt seats within its 600-second unattended window; verify the actual runtime,\n custody records and process identities before relying on it. A legacy PTY runtime without release\n support cannot preserve active seats through a generic manager restart. Drain it at an approved\n idle window instead of signalling the manager or replacing conversations.\n4. Reload the clients and connectors through their session-preserving host controls. Refresh any\n bearer helper captured from an older immutable CLI path. A transport-only reconnect does not reload\n JavaScript. Verify authenticated instance selection, a read-only manager command and canonical\n result recovery before allowing new followed mutations.\n\nTreat the interval from issuer reload through compatible manager/client reload as a manager-control\noutage. Mixed versions can refuse discovery or commands; there is no promised rolling transition.\nOrdinary agent sessions survive only where their runtime and credentials permit it. If verification\nfails, keep mutations paused and repair forward from the preserved state rather than resetting it.\nThis release does not add host-backed enrollment or terminal release for stock participant detached\nagents; see [Remote supervised agents](run-a-mesh.md#remote-supervised-agents).\n\n## From 0.48.2 to 0.49.0\n\n0.49.0 changes how a credential's authority is recorded. A credential is no longer only a signed\nfile: it is an *issuance*, with a generation the issuer chose and durable evidence of the ceiling it\nwas granted under. The important consequence for a running deployment is not at connect time. It is\nat renewal time.\n\n### What keeps working without any action\n\n- **Existing agent credentials keep authenticating.** A credential minted under 0.48.2 is not\n revoked and is not rejected at connect. Nothing needs to be re-issued to bring the fleet back up\n after the upgrade.\n- **The channel registry survives.** Channels, their replay settings, descriptions, and usage text\n are ordinary durable state and are not rewritten by the upgrade.\n- **`cotal deliver` is still a standalone command.** Running the delivery daemon as its own process\n remains supported; it is not restricted to being a child of `cotal up`.\n- **`cotal join` keeps its flags.** In particular `--lifecycle-uid` is not new in 0.49.0. It has\n been required alongside `--creds` since well before this release, and the pairing rule did not\n change here. A scripted external join that worked under 0.48.2 works unchanged.\n\n### What does not migrate\n\n**A credential minted before 0.49.0 cannot be renewed.** Managed agent credentials carry a\n24-hour lifetime and the manager re-signs one once it passes **75%** of its life, ticking every\nquarter of the TTL so a tick always lands inside that window. When the manager reaches a credential\nthat carries no issuance, it refuses to renew it and logs the agent by name:\n\n```\n! managed cred renewal <agent>: renewManagedStaticCred: <agent> carries no issuance;\n a static credential minted before SPEC 13.15 is not renewed under an unbound generation\n - respawn the agent\n - the agent dies loud at this cred's expiry unless it is reminted\n```\n\nSo the fleet comes up fine, runs normally, and then each agent stops at its own credential's\nexpiry, within roughly a day of the upgrade, one at a time rather than together. The refusal is\ndeliberate: the renewal would otherwise have to invent a generation nobody issued, which is the\nstate the release exists to remove.\n\n**Respawn the managed agents as the last step of the upgrade.** For this particular upgrade the\nrespawn is not optional: stopping a 0.48.2 manager ends its agent processes whichever CLI you use,\nfor the reason given under the outage window below. The respawn is how they come back, and it is\nalso what mints each credential as an issuance so it renews from then on. One planned pass over the\nfleet is the whole job. Skipping it leaves agents stopped and, for any credential that survived\ninto 0.49.0 unminted, brings the renewal cliff above a day later, one agent at a time.\n\n### Credentials you minted yourself\n\n**A credential you minted with `cotal mint` is a different case, and it very likely needs\nnothing.** The distinction that matters is not the word \"static\", which covers both. It is **what\nminted the credential and who owns its renewal**. A credential the **manager** minted for an agent\nit spawned carries a lifetime and is renewed by the manager, so it is the subject of everything\nabove. A credential **you** minted with `cotal mint` and handed to an external peer is issued with\n**no expiry at all**, and no manager renews it: it is not in the sweep, so there is no renewal to\nfail. It keeps working after the upgrade, and re-minting it would mean coordinating with a third\nparty for no gain.\n\nThe manager says which one it is holding. Where a credential has no expiry to reach, the sweep\nnames it and moves on rather than refusing:\n\n```\n! managed cred renewal <agent>: credential is unbounded - not renewed\n (a pre-TTL credential stays as minted until respawn)\n```\n\nRe-mint an external peer's credential only if you want it to carry a lifetime, and at a time you\nchoose.\n\n### How to read the boot log\n\nA 0.49.0 manager starting over an existing space may print lines like:\n\n```\n verified evicted: <holder-key> (3/12)\n already verified (durable): <holder-key>\n\u2713 boot self-heal: manager/<id> registration gate reopened at generation <n>\n```\n\nThese are **not** a credential migration, and reading them as one is the most likely way to\nconclude the fleet is fine when it is not. They come from the manager repairing **one** endpoint\nregistration gate that a previous restart left frozen, and they enumerate that single gate's\ncredential-family holders as it verifies each one evicted. `already verified (durable)` on a later\nstart is the repair cursor resuming, not a credential that became durable. The repair is real and\nuseful (it is what previously needed `cotal reconcile-gate` by hand), but it says nothing about\nwhether your agent credentials carry issuances. The renewal refusal above is the signal that does.\n\n### Which side to upgrade first in a split topology\n\nMove the manager first.\n\nThe stores 0.49.0 introduces are created by the **manager** at its own boot, not by the broker.\nThey are create-or-verify and idempotent, so a 0.49.0 manager brings the space's authority stores\nup to the new shape itself, and it does so against whichever broker is answering.\n\nBeing honest about the evidence behind each direction, because they are not equally established:\n\n- **Broker-first was measured on a live 30-agent deployment** (issue #1578). Upgrading the broker\n first locks the old manager out immediately: `cotal up` re-renders the broker's generated config\n from the trust record, and after the restart the still-0.48.2 manager is refused on every\n connection with an `authentication error` naming the Nkey, continuously. That text comes from the\n broker process, not from a Cotal command, so match on its shape rather than on an exact string.\n `cotal ps` reports zero agents while\n the agent processes are still alive, because the manager has lost its view of them, not because\n they died. Upgrading the manager clears it immediately.\n- **Manager-first is reasoned from where the new stores are provisioned**, not from a measured\n fleet upgrade. It is the recommended order because the manager is the component that creates what\n 0.49.0 adds, but it has not been run end to end on a production split topology at the time of\n writing. Treat it as the better-supported order rather than a guaranteed one, and keep the\n rollback below ready either way.\n\nWhichever order you pick, **this is not a rolling upgrade**. Between the two steps the mesh is down\nand the manager cannot see its agents. Go straight through rather than pausing between them, and\nschedule it as an outage window.\n\n### What the window looks like\n\n- **The managed agent processes do not survive step 1, in either order.** This is the one place\n where the obvious reordering does not rescue you, so it is worth understanding rather than\n working around. Sparing agents on a bare manager stop is a **handshake**: a 0.49.0 manager\n publishes a capability file proving it can release its agents, and a 0.49.0 `cotal down` refuses\n the stop unless it finds one. **A 0.48.2 manager never publishes that file**, because the\n mechanism ships in the release you are installing. So the old CLI against the old manager sends a\n plain stop and takes every seat with it, and the new CLI against the old manager either refuses\n (leaving `--with-agents`, which reaps deliberately) or falls to the legacy path, warns that it\n cannot verify the manager can spare its agents, and signals it anyway.\n- **You can confirm which side you are on in one command, without stopping anything.** The flag that\n marks the newer behaviour is absent from the older CLI, and its summary line makes the difference\n plain:\n\n ```\n $ cotal down --help # on 0.48.2\n cotal down - stop the whole local stack, or name only the components to stop\n\n $ cotal down --help # on 0.49.0\n cotal down - stop the whole local stack (managed agents stay running unless --with-agents), ...\n ```\n\n If your `cotal down --help` does not mention `--with-agents`, stopping the manager stops the\n agents with it.\n- **Therefore the respawn in step 5 is mandatory recovery for this upgrade, not an optional pass.**\n It is also the step that re-mints credentials as issuances, so it is the same action either way.\n Plan the window to include it rather than treating it as cleanup.\n- The **manager's view** of them is lost while the two sides disagree, so `cotal ps` reports zero\n and control commands do not reach seats.\n- **Messages are not delivered** while the mesh is down.\n- The window is as long as it takes to restart the second component, plus the manager's own start.\n It is minutes, not hours, provided you do not stop between the steps.\n- **Nothing self-heals if you stop halfway.** The refusal is continuous until both sides match.\n\n### Snapshot this before you start\n\nTake these while the deployment is still on 0.48.2. The two `cotal` reads are live reads and must\nhappen before anything stops.\n\n- **A filesystem or volume snapshot of both containers**, if your platform offers one. This is the\n only rollback that covers every case, and it is what the reporting deployment used.\n- **`cotal backup create <dir>`**, for the durable space state, **but read the next paragraph before\n you rely on it**: on a split broker and manager topology it is very likely unavailable to you, and\n the volume snapshot above is your actual rollback.\n- **The trust records and credential directory** under `.cotal/auth` on the manager host, including\n the per-space material directory. These are what a re-mint would otherwise have to replace.\n- **A copy of the channel registry**, so you can verify it came back rather than assuming it did:\n `cotal channels list` before and after.\n- **The output of `cotal ps`**, so you know how many seats you expect to see afterwards and can tell\n a lost view from a lost agent.\n\n#### `cotal backup` on a split topology\n\n**`cotal backup create` cannot read a running stack.** It requires a completed cut, and only\n`cotal down --preserve-state` publishes one:\n\n```\n$ cotal backup create ./backup.0482\n\u2717 backup requires a completed cut; run `cotal down --preserve-state` first\n```\n\n**And `cotal down --preserve-state` requires a manager alive on the host you run it from.** It uses\nthat manager to attest that every retained child stopped, and the check is deliberately fail-closed:\na manager that is dead or merely uncertain refuses rather than preserving an unproven cut. The check\nreads a local pidfile, so a **remote** manager does not satisfy it. On a split topology the broker\nhost has no local manager, which means the documented durable-backup path is not available there.\n\n**Measured rather than assumed, at 0.48.2**: the backup refusal above is executed output. The\npreservation requirement is read from `down.ts` at the same tag, where the preserve path asks a\nmanager to prepare an inventory and then requires that manager to be locally alive before it\ncommits. The part not executed end to end is a genuine two-host split, which needs two real hosts.\n\n**What to do instead.** Use the filesystem or volume snapshot of both containers. That is the\nrollback the reporting deployment actually used, it covers the broker's durable state and the\nmanager's credential material together, and it does not depend on either component being able to\nattest for the other. If you want `cotal backup` as well, take it from a host that does have a live\nlocal manager, and understand it is a second copy rather than the primary rollback.\n\n**This looks like a product limitation rather than a documentation gap**, and it is written here as\none so an operator is not left thinking they mis-typed a command. The upgrade path for the exact\ntopology this page is addressed to cannot use the documented backup command.\n\n### The upgrade end to end\n\n```bash\n# 0. on 0.48.2, STILL RUNNING: record what you expect to see afterwards.\n# These two are live reads, so they must happen before anything stops.\ncotal channels list > channels.before\ncotal ps > ps.before\n\n# 1. manager host. READ THE NOTE BELOW THE BLOCK FIRST: this step ends the\n# managed agent processes whichever order you choose, and the respawn in\n# step 5 is how they come back. It is recovery, not tidying.\n#\n# STOP THE MANAGER WITH THE 0.48.2 CLI, BEFORE INSTALLING 0.49.0. The\n# order matters and it is not recoverable once you install: a 0.49.0\n# `down manager` REFUSES to stop a 0.48.2 manager whose pid record carries\n# a start token, which is every manager on a platform that can read one\n# (Linux can):\n# refusing bare manager stop: ... does not prove this manager can detach\n# its agents; use --with-agents or stop the agents explicitly\n# The refusal names two remedies and NEITHER clears it for this case. The\n# check reads a capability file that only a 0.49.0 manager writes; it never\n# counts agents, so stopping them first changes nothing. And `--with-agents`\n# is whole-stack only, so `down manager --with-agents` is refused by its own\n# flag rule. See #1592.\ncotal down manager # the 0.48.2 CLI, still installed.\n # 0.48.2 has no --with-agents; this\n # is the whole route. On a host that\n # runs the whole stack, the 0.49.0\n # `cotal down --with-agents` after\n # installing is the alternative.\nnpm install -g cotal-ai@0.49.0 # ONLY after the stop above\n# `supervise` RUNS IN THE FOREGROUND and holds the terminal until you stop\n# it. There is no --detach on this command. Start it under whatever keeps\n# your manager alive normally (systemd unit, container entrypoint, or a\n# second terminal), and run the remaining steps from another shell.\ncotal supervise --space <space> --server nats://<broker>:4222\n\n# 2. broker host: stop the stack.\n# NOT `--preserve-state` on a split topology: it needs a manager alive on\n# THIS host to attest its children stopped, and yours is on the other one.\n# Your rollback is the volume snapshot from \"Snapshot this before you\n# start\", not `cotal backup`.\n# See \"cotal backup on a split topology\" above.\ncotal down\n\n# 3. broker host: install 0.49.0 and start it again\nnpm install -g cotal-ai@0.49.0\n# Record the manager log's size BEFORE starting, so step 3a can tell THIS\n# boot's output from every earlier one. It must be captured here, ahead of\n# the start: taken afterwards it sits past the new line and the wait hangs.\n# `<spaceKey>` is NOT the space name. It is lowercase hex of the name's\n# UTF-8 bytes, so space `prod` is `manager.70726f64.log`. Do not guess it:\n# `cotal up` prints the real path on its launch line. Substituting the\n# plain name points at a file that does not exist, and the wait below then\n# burns its full timeout before telling you.\nLOG=.cotal/manager.<spaceKey>.log\nOFF=$( [ -f \"$LOG\" ] && wc -c < \"$LOG\" || echo 0 )\ncotal up --detach --host 0.0.0.0 --space <space> --no-manager\n\n# 3a. SPLIT TOPOLOGY ONLY: `--no-manager` above boots the broker (and the\n# delivery daemon) with NO local manager on the broker host, so there is\n# no wait-and-stop step on a current cotal-ai. The rest of this step is\n# the OLDER-host recipe, kept because the flag is refused there and that\n# refusal is your signal you are on it: without the flag the `up` also\n# starts a local manager, and you must wait for the log to show it is up,\n# then stop it, or you finish the upgrade with two managers and the one\n# you did not intend is the one nobody is watching.\n# A bare `grep -q` does NOT wait: it reads once and exits 1 immediately\n# if the line has not been written yet. Bound the wait instead, so a\n# manager that never comes up fails loudly rather than reading as ready.\n# The log is opened APPEND-ONLY, so on any host that has run a manager\n# before, this file ALREADY carries a `manager up` line from an earlier\n# boot. Grepping the whole file therefore matches instantly and waits for\n# nothing. Read only what THIS boot appended, using the $OFF captured in\n# step 3 above (before the start, which is the only point it is correct):\ntimeout 60 bash -c \\\n \"until tail -c +$((OFF+1)) \\\"$LOG\\\" | grep -q '. manager up'; do sleep 1; done\"\n# exit 0 = THIS boot logged it; exit 124 = it never did, so STOP and look.\n# This manager is 0.49.0 and publishes its own spare-capability file, so\n# the bare stop below is NOT the refusal case from step 1.\ncotal down manager # broker + delivery remain\n# On a current cotal-ai the two commands above are unnecessary (nothing\n# to wait for, nothing to stop) and `cotal down manager` simply reports\n# no manager to stop.\n\n# 4. verify the mesh is whole again before touching the fleet.\n# Do NOT compare `cotal ps` against ps.before yet: step 1 ended the agent\n# processes, so at this point it is EXPECTED to be empty, and an empty\n# `ps` is also the signature of the broker/manager mismatch described\n# above. The two are indistinguishable here, so compare what the mesh\n# itself should have carried across instead:\ncotal channels list # compare against channels.before: this SHOULD match now\ncotal ps # expect it to be EMPTY here; ps.before is the target for\n # step 5, not for this step\n\n# 5. the step that is easy to skip: respawn the managed agents so their\n# credentials are re-minted as issuances and can renew. Persona is a\n# POSITIONAL argument here, unlike `cotal stop`, which requires --name.\n# One call per agent:\ncotal spawn <persona> --detach --name <n> --space <space>\n# then the comparison step 4 could not make:\ncotal ps # NOW compare against ps.before: seat count should match\n```\n\nThe mesh is down from step 2 until step 3 finishes. That is the window. On a split topology there is\nno cut and no backup inside it, so the window is the stop, the install and the restart, nothing more.\n\n## Adding a section for a future release\n\n**Every changeset marked breaking adds a section to this page.** A release that changes what an\noperator must do, in what order, or what stops working, is not finished until the section exists.\n`scripts/upgrade-section-gate.mjs` grades a commit range for this: run it as\n`pnpm upgrade-section-gate --base <ref>` and it reds when the range carries a breaking change and\nadds no new release section. CI runs its self-test and, as a step of the `attribution` job, grades\neach pull request's own range as `HEAD^1..HEAD` over the merge snapshot it checked out. That job is\nthe only context in the branch protection rule set, so a red gate FAILS A REQUIRED CHECK AND BLOCKS\nTHE MERGE. The section is not optional and a reviewer cannot wave it through without an\nadministrator overriding branch protection. Be precise about what the check proves either\nway, because one trusted past its evidence is worse than none. It proves a section for a release\n**was written here**. It cannot prove the section is **correct**, or that it describes the break\nthat actually landed, and it cannot see a breaking change that carries no marker at all. Reviewing\nthe words remains a person's job.\n\n**Mark the break, or the gate cannot see it.** Any one of these is enough, and they are the only\nthings it reads:\n\n- a `!` before the colon in the commit subject, as in `feat(core)!: bind hosted runs to the caller`\n- a `BREAKING CHANGE:` footer in the commit body\n- a changeset in `.changeset/` declaring a `major` bump for any package\n\nThe marker must survive the squash. A `!` that lives only in a commit you squash away is not in the\nrange the gate grades, so put it in the subject that lands on `main`.\n\n**The heading is a `##` and names the release**, like `## From 0.48.2 to 0.49.0`. Both matter, and\nneither is a style preference. Coverage is claimed by a heading, so a heading that names\nno release claims every release and distinguishes none: `## Notes` with a sentence under it would\notherwise satisfy the rule. Naming the release also makes the section the one an operator upgrading\nthat release will search for. Use `###` freely for detail inside a section. Subsections belong to\ntheir release rather than counting as separate coverage.\n\nName the release that first carries the change: the next version Changesets publishes, which\n`pnpm changeset status --verbose` lists. `bin/package.json` on `main` still reads the release already\npublished. If a release is cut while the change is open, the change ships in the release after it,\nso move the heading before merging. The gate accepts any version in a heading, so before merging a\nrelease pull request, check every heading added since the previous tag against the version it\npublishes.\n\nA section is written for the operator, not for the reviewer. It answers, in this order:\n\n1. What keeps working with no action at all.\n2. What does **not** migrate, and when that becomes visible. Name the log line if there is one.\n3. The order to move components in for a split topology, and why that order.\n4. What the outage window looks like, including what survives it.\n5. What to snapshot before starting.\n6. The commands, end to end.\n\n**Where an answer was not measured, say so in the document rather than guessing.** An operator who\nknows which half of a recommendation is reasoned and which is measured can plan around it; one who\nfinds out afterwards cannot.\n"
|
|
57977
|
+
"body": "# Upgrading a running deployment\n\n> **Guide** (informative) \xB7 **For:** operators upgrading a mesh that already exists \xB7 **See also:** [Substrate stability](stability.md), [Run a mesh](run-a-mesh.md), [Identity and auth](identity-and-auth.md)\n\n[Substrate stability](stability.md) tells you what the version numbers promise. This page is the\nother half: what to actually do when the deployment already exists, has credentials in it, and\ncannot simply be recreated. Every release that breaks a running deployment gets a section here,\nnaming what migrates on its own, what does not, and the order to move the pieces in.\n\n## The pre-1.0 upgrade contract\n\nThe packages are pre-1.0, so a minor bump may break an API or an on-disk expectation. Four\ncommitments make that survivable for someone with a fleet:\n\n- **Pin an exact version.** `0.N.P`, never `^0.N.P`. A range can pull a breaking minor in during an\n unrelated reinstall.\n- **Every break that touches a running deployment gets a section on this page**, written in terms of\n what an operator does, not in terms of which module changed.\n- **Read the section before you start, not halfway through.** A section names the work up front\n precisely so the operation does not change shape once it is underway.\n- **A break that cannot be made automatic says so.** Where credentials or state must be recreated by\n hand, the section says which ones and when, rather than leaving you to discover it at the moment\n the first one stops working.\n- **A change to the shape of a credential, or to who may renew one, is breaking whatever the commit\n marker says.** This rule is stated because the marker is a judgement made while writing the code\n and the consequence is felt by someone running it a day later. A fleet that keeps authenticating\n looks compatible and is not, if nothing in it can renew. Any automated check of this rule would\n read commit markers, so a break recorded as a feature is the one case it could not see, which is\n why the rule is written for people first. **The marker held for this release: the 0.49.0 change\n that caused all of this, `36d177951 feat(core)!`, did carry its `!`.** The rule exists for the\n next one that does not.\n\nWhat this page does not promise is a rolling upgrade. Nothing in the current line dual-serves two\nauthority versions, so where broker and manager run separately there is a window in which the mesh\nis down. The sections below give that window's shape so it can be scheduled rather than endured.\n\n## Auth context closure in 0.71.0 (unreleased)\n\nExisting deployments need no credential migration or restart for these additive APIs. Embedded\nhosts can now inspect `handle.connections()` and await `handle.closed` after `close()` or `drain()`\nto prove every owned transport ended, including the callout, replaced readiness readers and\nshort-lived clients. The inventory is a detached snapshot.\n\nA transport close failure now rejects with its connection label. The terminal signal stays pending\nwhile any connection remains live. Repair the failure and retry `close()` before awaiting\n`handle.closed`. Closing one hosted context does not close another account's context.\n\nRead a space's claim with `readPlaneClaim(kv, space)` on that account's leader-only auth bucket.\nAn unclaimed space returns `undefined`; held and released rows retain their claim identity.\nDeleted, malformed and foreign-space rows refuse. `PlaneClaimRow` and `PLANE_CLAIM_KEY` are exported.\n\nUse `observeAccountLivenessWithCreds({ servers, observerCreds, accountId, options })` with the\naccount-scoped membership-observer credential to list that account's connections. It never widens\ncredentials or evicts connections. Zero rows prove absence only with a complete sweep and the\nsingle-server proof. An embedded endpoint's trusted composition can retain transport custody\nthrough `EndpointOptions.onConnection`.\n\n## Unreleased\n\nOn a per-user-auth mesh, a spawn-scoped caller can arm the event plane of a child under its own\nowner without `admin`. This fixes owned spawns in spaces whose registration policy requires the\nplane. Upgrade the manager to pick up the admission change. No credential or state migration is\nneeded. Cross-owner arming still requires `admin`, and the child's own-channel rule and ledger\nenvelope are unchanged. A silent non-owner caller in a space without the policy still has the\nplane disarmed, with a notice if provisioning succeeds.\n\n## Hermes model from the environment in 0.68.0\n\nA connector now launches on the model and variant its launcher resolved (the `--model` or\n`--variant` flag, else the agent file's `model:` or `variant:`) and no longer reads them again from\nthe agent file. The Hermes connector also no longer takes a model from `HERMES_MODEL` in the\nenvironment of the process that spawns the seat, including when `spawn.env` lists it.\n\n### What stops working\n\nA Hermes spawn whose only model was `HERMES_MODEL` in the spawning environment is refused at launch,\nand the refusal names both ways to set a model. Spawns that set `--model` or `model:` are unchanged,\non every connector.\n\nCode that calls a connector's `buildLaunch` directly with only `configPath` now gets no model or\nvariant from that file. Pass them as `model` and `variant`.\n\n### Before the upgrade\n\nMove each Hermes seat's model from `HERMES_MODEL` onto its spawn with `--model`, or into its\npersona's `model:`.\n\n## Run answers on a participant manager in 0.68.0\n\nA participant manager now asks its issuing host for an answering credential by naming the run and\nstep it answers. The host reads the pause's token off that run's journal and no longer accepts a\ntoken from the manager. Runs on a mesh with no participant manager are unaffected.\n\n### What stops working\n\nWhile a participant manager and its issuing host run different sides of this release, the host\nrefuses every `cotal run answer` and every amendment that manager serves, because each side refuses\nthe other's request shape. Starting, resuming and reading runs is unchanged. A pause stays waiting\nthrough the window, or follows its timeout if it has one.\n\n### Before the upgrade\n\nUpgrade the auth service and every participant manager registered with it in the same window, then\nanswer the pauses that waited.\n\n## Headless OpenCode handshake in 0.69.0\n\nWith `COTAL_SERVE_HEADLESS=1`, the OpenCode launcher's `[cotal-serve]` line on stdout now carries\nonly `port` and `session`. The server password no longer appears in it, and the 1.x TUI no longer\nreceives the password on its command line.\n\n### What stops working\n\nA headless host that read `password` from that line has no password, and the server refuses its\nrequests. Seats with a TUI, and headless seats that no host drives, are unaffected.\n\n### Before the upgrade\n\nHave each headless host mint a password and pass it to the launcher as `OPENCODE_SERVER_PASSWORD`,\nthen use it for basic auth as before. Without that variable the launcher mints its own.\n\n## Filesystem store identity in 0.69.0\n\nThe delivery daemon's answer to the manager's store check now names a filesystem store by its root\nand by a random `id` that the store records once in `store.id` inside its own directory:\n`.cotal/store.id` for a workspace root, or the directory of the file for `cotal deliver --creds\n<file>`. A manager no longer counts the daemon's store as its own because the two roots have the\nsame path. On a split whose broker host and manager host use one root path, the manager host now\nstays off the daemon-credential renewal lease, so `cotal doctor auth --fix` on the broker host can\nrenew the daemon credentials.\n\n### What stops working\n\nA manager and a delivery daemon on different sides of this release refuse each other's answer to\nthe store check. The manager then remints no daemon credential, and a manager that is booting does\nnot start. This is read from the code and was not measured across two releases. A\n`cotal deliver --creds <file>` whose directory is a read-only mount and holds no `store.id` stops at\nstart. So does a `--creds` file that is its directory's `store.id` under any name, and a `store.id`\nthat is a symbolic link or holds anything but a lowercase UUID.\n\n### Before the upgrade\n\nUpgrade the broker host and every manager host of a space in the same window. For a `--creds` file\non a read-only mount, add a regular `store.id` file beside it that holds a new lowercase UUID and no newline,\nas `node -e 'process.stdout.write(crypto.randomUUID())' > store.id` writes. Move a `--creds` file\nnamed or linked as `store.id` to a file of its own.\n\n## Detached spawns with `--share-tools` in 0.69.0\n\nThe manager's `spawn` operation now takes `shareTools` as a list of MCP server names. The CLI parses\n`--share-tools` into that list before it sends the request, and the manager cluster document moves\nto revision 22. A cut taken with `cotal down --preserve-state` before the upgrade still resumes: the\nmanager reads its `cotal-manager-resume/v1` inventory and writes new cuts as\n`cotal-manager-resume/v2`.\n\n### What stops working\n\nA CLI and a manager on different sides of this release refuse a detached spawn that passes\n`--share-tools`, because the CLI checks each request against the contract the manager serves. This\nis read from the code and was not measured across two releases. A detached spawn without the flag,\na foreground spawn and a roster entry are unaffected. A manager older than this release cannot\nresume a cut that this release took.\n\n### Before the upgrade\n\nUpgrade the CLI on every host that runs `cotal spawn --detach` in the same window as the managers\nit reaches.\n\n## Shared MCP server checks in 0.69.0\n\nThe cotal config reader now checks each server under `connectors.<name>.mcpServers` when it reads\nthe file, and refuses one that cannot launch as written, naming the file and the field. The rules\nare in [the config file](config.md#the-config-file).\n\n### What stops working\n\nA config file that holds such a server refuses every Claude spawn that reads it, including one with\n`--share-tools none`. Before, a field of the wrong type failed each Claude spawn that shared the\nserver with a `TypeError` that named neither the file nor the server, a spawn that did not share it\nlaunched, and a server with no `command` or `url` was passed to `claude`, which never started it.\nRead from the code and not measured: spawns on other connectors, a manager resume and the step of\n`cotal setup` that records the shared list read the same files, so each stops at the same refusal.\n\n### Before the upgrade\n\nCheck `connectors.<name>.mcpServers` in the operator-level config file and in each space's\n`.cotal/config.json`. Give each server a string `command`, or a `type` of `http`, `sse` or `ws` with\na string `url`. Write `args` as a list of strings and `env` and `headers` as objects of strings, or\nremove the server.\n\n## Remote manager family eviction in 0.69.0\n\nA remote manager registered through its host now asks the host to evict up to 256 holders of its\ncredential family in one maintenance request, and the host reads the family once for the whole set.\nBefore, a restart sent one request per holder and the host read the whole family for each one.\nMeshes with no remote manager are unaffected.\n\n### What stops working\n\nWhile a remote manager and its issuing host run different sides of this release, each side refuses\nthe other's eviction request shape. A restart whose credential family already has holders then fails\nat its eviction step and leaves the manager's registration gate frozen. A first start, a clean stop\nand the host's reconciliation of a foreign slot holder are unchanged.\n\n### Before the upgrade\n\nUpgrade the auth service and every remote manager registered with it in the same window. A manager\nthat restarted inside the window resumes its frozen registration on its next start once both sides\nrun this release.\n\n## AG-UI emitter holder hooks in 0.69.0\n\n`AguiEmitterHolder` from `@cotal-ai/connector-core` now takes its hooks as one named object after\nthe emitter factory: `new AguiEmitterHolder(startEmitter, { onError, onRunClosed, waitLive, runMeta })`.\nOnly `onError` is required. Nothing about a running mesh changes, and every shipped connector passes\nits hooks by name. Only a connector of your own that builds a holder is affected.\n\n### What stops working\n\nA holder built with positional hooks, such as `new AguiEmitterHolder(start, onError, onRunClosed)`,\nno longer compiles, because the constructor takes two arguments. Plain JavaScript that keeps the\npositional form still runs, but the holder calls none of its hooks, so a failure never reaches\n`onError`.\n\n### Before the upgrade\n\nPass each hook by name, for example `new AguiEmitterHolder(start, { onError, onRunClosed })`, and\ndrop any `undefined` that filled an earlier slot to reach a later hook.\n\n## Worker run failure type in 0.70.0\n\n`WorkerRunFailed`, the failed result of `runInWorker` in `@cotal-ai/lang`, is now a union on\n`class`: `released`, `held`, `effect`, `too-large`, `rejected` or `error`. A running mesh needs\nnothing, because the runtime host and the engine thread ship in the same install. A run on the\ncompiled engine whose program throws an object with `code: \"L5012\"` or `code: \"L5025\"` used to end\nreleased and now ends failed, as it does on the walker.\n\n### What stops working\n\nTypeScript code that reads `code`, `reason`, `step`, `pending`, `kind`, `detail` or `tooLarge` on a\n`WorkerRunFailed` it has not narrowed fails with TS2339. A released, held, too-large or rejected\nresult no longer carries `code`, so JavaScript that branched on `L5012`, `L5025`, `L5006` or\n`L5010` stops matching with no error. `tooLarge` is gone.\n\n### Before the upgrade\n\nBranch on `class` where such code read `code`: `released` for L5012, `held` for L5025, `too-large`\nfor L5006 and `rejected` for L5010. An `effect` or `error` result keeps its `code`. Once narrowed to\n`too-large`, a result carries the `stepKey`, `bytes` and `bound` that `tooLarge` held.\n\n## Remote manager request builder in 0.70.0\n\n`remoteManagerClient.remoteManagerAuthorityRequest` from `@cotal-ai/manager` now takes an\noperation's coordinates as one object, and `remoteManagerRegistrationProof` from `@cotal-ai/core`\ncomputes the proof from the manager's identity state instead of a request. Nothing about a running\nmesh changes: the proof digest and the request on the wire are the same, so a manager and a host on\ndifferent sides of this release still accept each other. Only code that builds remote manager\nrequests itself is affected, in TypeScript and in plain JavaScript.\n\n### What stops working\n\nA call that passes the registration proof, contract artifacts, session, retirement or transfer\nreader as positional arguments after the operation no longer compiles. A call that passes a request\nto `remoteManagerRegistrationProof` no longer compiles either, because the second argument now names\nthe lifecycle `lifecycleUid`, as the identity state does.\n\nPlain JavaScript runs both old calls without an error. The builder drops the positional coordinates,\nso the host refuses the request with `requires a sha256 registrationProof`. A proof computed from a\nrequest leaves out the lifecycle, so the host refuses a request that carries it as a proof mismatch.\n\n### Before the upgrade\n\nName the coordinates, for example\n`remoteManagerAuthorityRequest(state, \"cli\", \"retire\", { registrationProof, retirement })`.\nCompute the proof as `remoteManagerRegistrationProof(owner, state)`, adding the contract artifacts\nas a third argument for activation only. A host that recomputes the proof from a received request\npasses `{ space, instanceId, lifecycleUid: managerLifecycleUid, identities }` from that request.\n\n## Bearer validator lifetime cap in 0.70.0\n\n`validateUserToken` from `@cotal-ai/auth` no longer takes `maxTtlSec`. It caps a bearer's lifetime\nat the cap of the bearer's view, the same cap the issuer applies when it mints: 900 seconds, or 300\nfor a `transfer-writer` bearer. The auth callout never passed the option, so a running mesh behaves\nas before. Only code of your own that calls the validator with `maxTtlSec` is affected.\n\n### What stops working\n\nA call that passes `maxTtlSec` in an object literal no longer compiles. Plain JavaScript that keeps\nit still runs, and the value is ignored. A `NaN` value, such as `Number()` of an unset environment\nvariable, used to turn the lifetime check off and accept a bearer of any lifetime. That bearer is\nnow refused at its view's cap.\n\n### Before the upgrade\n\nRemove `maxTtlSec` from each call. A test that needs a bearer to expire sooner mints one with a\nshorter lifetime.\n\n## Persisted identity records in 0.70.0\n\nThe manager instance identity, the manager sibling identities, the auth plane instance identity and\na participant manager's remote authority state now share one reader and one first mint in\n`@cotal-ai/workspace`, exported as `claimIdentityRecord` with the nkey check `identityOf`. Each\nrecord is read as a regular file, must hold non-empty nkeys and is created exclusively, so\nconcurrent first starts of a participant manager on one root now settle on one identity where each\nused to keep its own. `saveManagerInstanceIdentity` and `saveAuthInstanceIdentity` are gone. A\nrunning mesh whose records are plain files needs nothing.\n\n### What stops working\n\nA manager instance, auth instance or remote authority record that is a symlink, a directory or any\nother non-regular entry is refused where it used to be followed. The manager, the auth plane and a\nparticipant manager fail to start on it, and `cotal reconcile-gate` and `cotal deregister-instance`\nrefuse it. Retirement already refused it. A remote authority record with an empty nkey id or seed\nis refused too. A first mint that loses its race and cannot read the winner now refuses with\n`identity-record-create-lost` in place of `manager-instance-identity-create-lost` or\n`auth-instance-identity-create-lost`. Code that imports either `save` function no longer compiles.\n\n### Before the upgrade\n\nReplace a symlinked identity record with a copy of the file it points to. Code that wrote a record\nwith a `save` function plants it with `createManagerInstanceIdentity` or\n`createAuthInstanceIdentity`, which create the record when it is absent and otherwise return the\nstored one unchanged. Nothing replaces an overwrite of a stored identity.\n\n## Manager instance in user credentials in 0.70.0\n\n`AuthProvider.userCredentials` from `@cotal-ai/core` no longer returns `managerInstanceId`. A\n`manager-caller` credential's manager instance is the signed `act.managerInstanceId` claim in its\nbearer, which the broker verifies and the CLI already used. The reference provider in\n`@cotal-ai/auth` stops copying the exchange response's field into its result, where nothing\ncompared it with the bearer. The exchange still answers with the field, so a running mesh behaves as\nbefore.\n\n### What stops working\n\nCode of your own that reads `managerInstanceId` from a `userCredentials` result no longer compiles,\nand plain JavaScript reads `undefined` there.\n\n### Before the upgrade\n\nRead the instance from the bearer's `act.managerInstanceId` claim.\n\n## Auth plane identity location in 0.70.0\n\nThe user-auth service keeps its instance identity in the root's `.cotal/space.<hex>/auth-instance.json`,\nbeside the manager's. It used to sit inside `.cotal/auth`, at\n`space.<hex>/.cotal/auth/auth-instance.<hex>.json`, so a copy of that folder carried it. The first\nstart of an upgraded root moves the record and keeps the instance. A hosted context started through\n`startAuthService` has its record moved the same way inside its `stateDir`.\n\n### What stops working\n\nCode that calls `openAuthAuthorityPlane` without the new `identityRoot` option no longer compiles. A\nstart that finds a record both in `.cotal/space.<hex>/` and at its older place refuses and names the\ntwo files. A start also refuses when the older place of the auth or manager identity holds a symlink,\na directory or anything else that is not a regular file. The manager used to skip a dangling symlink\nthere and mint a new identity.\n\n### Before the upgrade\n\nPass `identityRoot` to `openAuthAuthorityPlane`. When `dir` is a workspace root's user-auth state\ndir, `<root>/.cotal/auth/space.<hex>`, pass that root. A plane with no workspace root, as\n`startAuthService` runs, passes `dir` itself. Either keeps the identity the plane already has: on the\nfirst start it moves from `<dir>/.cotal/auth/` to `<identityRoot>/.cotal/space.<hex>/`. Never pass a\ndirectory inside `.cotal/auth`: the record would land in the folder an operator copies and travel\nwith it again.\n\nA copy of `.cotal/auth` taken from a root last run by an older Cotal carries that root's record.\nDelete `.cotal/auth/space.<hex>/.cotal/auth/auth-instance.<hex>.json` from the root you copied it to\nbefore the first `cotal up --user-auth` there.\n\n## Per-seat `COTAL_` names in `spawn.env` in 0.71.0\n\n`spawn.env` in the cotal config no longer forwards a `COTAL_` name the launcher sets for each seat,\nsuch as `COTAL_ROLE`, `COTAL_MODEL` or `COTAL_SUBSCRIBE`. Before, a seat launched with no value of\nits own took the spawning process's value and ran under that role, model or read set. The\nmachine-wide knobs a seat already receives, such as `COTAL_HOME`, may still be listed.\n\n### What stops working\n\nEvery spawn and resume under a config whose `spawn.env` lists such a name is refused before\nlaunch, and the refusal names the entry. Code that calls `launchEnv` from `@cotal-ai/connector-core`\nwith such a name in `envAllow` gets the same error.\n\n### Before the upgrade\n\nRemove those names from `spawn.env`. Give each seat its role, model and channels with `--role`,\n`--model` and `--subscribe`, or in its persona's `role:`, `model:` and `subscribe:`.\n\n## Role addresses in 0.71.0\n\nA role must be one `[A-Za-z0-9_-]` token. Before 0.71.0 any other spelling was rewritten into one:\n` probe ` reached the `probe` queue and `pro.be` reached `pro_be`, while the message kept the\nspelling sent. An anycast to `*` was accepted and stored where no holder reads it.\n\n### What stops working\n\nAn agent whose role is outside the token set no longer starts, however it is launched:\n`cotal join --role`, `cotal spawn --role`, an agent file's `role:`, `COTAL_ROLE` and an embedded\nendpoint's `card.role` are all refused before the agent joins.\n\nA send to such a role, or to `*`, through `cotal send ask`, `/anycast` or `cotal_anycast` is refused,\nand nothing is stored.\n\n`routeToken` is no longer exported from `@cotal-ai/core`. A role routes as spelled, so code that\nused it to name a role's queue uses the role itself, and `assertValidRole` checks one.\n\n### What migrates on its own\n\nEvery task queue. A `svc_<role>` durable was always named from the rewritten token, so its pending\nrequests and its holders carry over.\n\n### Before the upgrade\n\nRename each role outside the token set to the token it already routed to: remove the surrounding\nspaces and replace every other character outside the set with `_`. Rename it where the holder is\nlaunched and in every script or prompt that sends to it.\n\n## IdP URLs on `localhost` in 0.73.0\n\nThe IdP URL that `cotal login --idp` and `cotal up --user-auth --idp` take, the JWKS URL the auth\nservice pins from it, the space-catalog link and `cotal sync` now share one rule: `https://`, or\n`http://` on a loopback IP literal. `localhost` is a name, so it no longer counts: a hosts entry\nwould choose the IdP, and with it the keys the callout trusts. Every loopback literal now passes, so\n`http://127.0.0.2/api/auth` and `http://[::ffff:127.0.0.1]/api/auth` are accepted where they were\nrefused before. A JWKS URL with any scheme other than `https:` or `http:` is refused. A mesh whose\nIdP uses HTTPS is unaffected.\n\n### What stops working\n\nOn a mesh whose IdP is pinned at `http://localhost:<port>/...`, the auth service builds its key\nresolver from the pinned JWKS URL at start, and that resolver now refuses it with\n`JWKS origin must be https (or http on a loopback IP literal for dev)`.\n\n`cotal login` and `cotal logout` with `--idp http://localhost:<port>/...` are refused with\n`idp url must be https (or http on a loopback IP literal such as 127.0.0.1 for local dev)`, and so is\nevery command that reads a session cached under that URL.\n\n### Before the upgrade\n\nUse `127.0.0.1` (or `::1`) in place of `localhost`. In the space's `idp.json` under the mesh's\n`.cotal/auth`, change `url` and `jwksUri` to the literal spelling and leave `issuer` and `audience`\nas they are. Owners derive from the issuer, so existing grants keep matching. Change a manifest's\n`broker.idp` the same way, because `cotal up` refuses an `--idp` that differs from the pin. Then\nhave each person run `cotal login --idp http://127.0.0.1:<port>/api/auth` again, because sessions\nare cached under the URL.\n\n## Carrying a resumed Claude session to another host in 0.67.0\n\n`cotal spawn --resume <id> --detach --on <instance>` now carries a Claude session held on the\noperator's host to the target manager instance. Both sides need this release: an older manager does\nnot serve `transcript-receive`, and the CLI then stops with that manager's refusal instead of\nlaunching. The manager cluster document moves to revision 21, and the `ps` row's `resume` object\ngains `host` and `transferredAt`.\n\nA manager host that runs carried seats needs `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_AUTH_TOKEN` or a\ncloud provider selection in its environment, because each carried seat runs in its own Claude home\nwith no stored login. On an authenticated mesh the CLI mints the transfer writer from the space's\nsigning seed, so the carrying host needs that seed, as for any other operator command. On a user-auth\nmesh it exchanges the operator's login for a `transfer-writer` view instead, so the operator's grant\nneeds scope `admin`, and the auth service must run this release. A remote manager receives a carry once\nits host serves the manager-service `transferReader` operation. A seat launched without carrying,\nincluding any `--resume` whose id this host does not hold, is unchanged.\n\n## Lifecycle head type in 0.67.0\n\n`LifecycleMapping`, the type `parseLifecycleHead` returns, is now a union on `state`. Nothing about\na running mesh changes: heads that parsed before parse the same way, and the refusals are\nunchanged. Only TypeScript code that compiles against `@cotal-ai/core` is affected.\n\n### What stops working\n\nAn `interface` that extends `LifecycleMapping` fails with TS2312, because an interface cannot extend\na union. Code that builds a head in memory no longer compiles when the head is `retiring` without\nits `op`, or `active` or `retired` with one. The parser already refused those heads.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type ActiveMapping = LifecycleMapping & { state: \"active\" }`. A reader that has checked\n`state === \"retiring\"` reads `op` without a guard.\n\n## Issuance gate types in 0.67.0\n\n`EpGateRow` and `EndpointGateRow`, which `parseIssuanceGate` and `parseEndpointGate` return, and\n`EpGateState`, which an `EpIssuanceGate` or `EpIssuanceBarrier` returns from `observe`, are now\nunions on `state`. Nothing about a running mesh changes: gates that parsed before parse the same\nway, and the refusals are unchanged. Only TypeScript code that compiles against `@cotal-ai/core`\nis affected.\n\n### What stops working\n\nAn `interface` that extends one of these types fails with TS2312, because an interface cannot\nextend a union. Code that builds a gate in memory, such as a custom barrier's `observe`, no longer\ncompiles when the gate is `frozen` or `retired` without its `op`, or `open` with one. The gate\nparsers already refused those rows.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type CustomGateRow = EpGateRow & { custom: string }`.\n\n## Lifecycle-blocked refusals in 0.66.0\n\nA refusal that carries `ai.cotal.ep.lifecycle-blocked` now reports only the lifecycle state it\nread. Nothing about a running mesh changes. A client that branches on the detail must read the new\nfield.\n\n### What stops working\n\nA refusal raised at the issuance gate used to carry `headState` without reading the head:\n`retiring` for a frozen gate and `retired` for a retired one. It now carries `gateState`\n(`frozen` or `retired`) and no `headState`. A client that treats `headState: \"retired\"` as a\nburned uid, or `headState: \"retiring\"` as a retirement in flight, no longer matches those\nrefusals, and the `[lifecycle ...]` suffix on the error string changes the same way. A custom\nissuance barrier whose `observe` returns a frozen gate without a valid `op` (a string `opId` and\none of the four op kinds) is now refused as `internal` by `registerServiceInstance`.\n\n### Before the upgrade\n\nUpdate such a client to read `gateState` for a gate refusal and `blockedOp` for the operation that\nholds the gate. `headState` is present only when the refusal read the head, for example an\nactivation refused because the head is still retiring.\n\n## Workflow programs that bind `once` in 0.65.0\n\n`once` is now a scope of the workflow language, so it is a reserved name. A program that declares\nits own `once` binding (`const once = ...`, a parameter or a function named `once`) is refused at\nvalidation with L2002. Nothing else about a running mesh changes.\n\n### What stops working\n\nA run whose recorded program binds `once` cannot be resumed after the upgrade, because a resume\nvalidates the recorded program again. A new `cotal run start` of such a program is refused before\nanything is recorded.\n\n### Before the upgrade\n\nList the runs with `cotal run ps` and check each program that is still running or held for a\nbinding named `once`. Let those runs finish on the old version before you upgrade the manager, and\nrename the binding in the program before you start it again.\n\n## From 0.58.0 to 0.59.0\n\nEvery connector now publishes a failed run's `RUN_ERROR` on `events.<owner>.<actor>` with the fixed\nmessage `run failed` and no `code` or `rawEvent`. The error text and error kind a harness reports\ncan echo a prompt, a peer message or tool output, and that channel has a different read ACL. A\nreader that showed the message or branched on `code` gets neither after the upgrade. Where a\nconnector reports the error kind as the agent's presence condition, that is unchanged.\n\n### Settle pending event frames before the upgrade\n\nEach session's events are frozen in its event write-ahead log before they are published. A session\nrestarted on 0.59.0 whose log still holds an unacknowledged frame with an older `RUN_ERROR` does not\nrepublish it: its event emitter halts with `egress-run-error` and publishes nothing further for that\nsession. The broker may or may not already hold that frame, so the halt cannot settle it.\n\n1. Stop the seats cleanly on 0.58.0, with the broker still up.\n2. List the logs that still hold a pending frame. The logs live under the events state root\n (`COTAL_WORKSPACE_ROOT`). Empty output means there is nothing to settle.\n\n ```sh\n find \"$COTAL_WORKSPACE_ROOT/.cotal/events\" -name wal.json \\\n -exec jq -r 'select(.pending != null) | input_filename' {} +\n ```\n\n3. For each session listed, start it again on 0.58.0 while the broker is reachable, let it recover,\n stop it, and run step 2 again. Recovery publishes the frame as 0.58.0 would have, error text\n included, so it only finishes what 0.58.0 had already started.\n\n If that start halts with `cas-loss` instead, the agent's subject is no longer at the sequence this\n log expects, and no restart settles that log, on 0.58.0 or later. A lost acknowledgement is one\n cause: the broker stored the frame, so it and its error text are already on the channel, and every\n retry halts the same way because the stream checks the frozen expectation before it deduplicates.\n The halt message names the other causes, such as a second emitter for the same agent under a\n different state root, a restored stream or frontier record, or a purged channel. With those the\n pending frame may never have reached the broker, so a `cas-loss` does not tell you whether it\n landed. Find and stop any second writer and rule out a restored state first. Clearing the halt\n then means purging the agent's event channel and removing the agent's directory under the events\n state root whole (see [Event plane](connect-claude.md#event-plane)). That abandons the pending\n frame whether or not the broker has it, and the purge also drops the earlier frames of every\n session of that agent.\n4. Upgrade once step 2 prints nothing.\n\nIf a session halts with `egress-run-error` after the upgrade, go back to step 3 for that session on\n0.58.0. Do not edit or delete `wal.json` on its own to get past either halt: clearing the pending\nframe abandons that epoch, an event the broker never received is lost, and removing part of the\ndirectory leaves a state the next start refuses.\n\n## Explicit actor grants in 0.59.0\n\n`cotal actor grant` no longer fills an omitted ACL flag with its wide default. A grant names\n`--scope`, `--allow-subscribe` and `--allow-publish`, or passes `--full` to give the ones it leaves\noff their wide defaults (`spawn,role:default`, `>` read, `>` post). Any other grant is refused. The\nbreak is in the CLI on the machine that holds the actor ledger, the one that ran\n`cotal up --user-auth --idp <url>`. No stored row, credential or wire message changes.\n\n### What keeps working\n\nExisting actor ledger rows keep the authority they were granted, and their users and agents connect\nas before. `actor revoke`, `actor list` and a `grant` that names all three ACL flags behave as they\ndid on 0.58.0. Nothing on disk is converted.\n\n### What stops working\n\nA grant that leaves off any of the three flags without `--full` exits 1 with\n`refusing to grant \"<actor>\" with --scope, --allow-subscribe, --allow-publish left off`, naming the\nflags it is missing, and then prints both accepted forms. It writes no row and does not retire the\nactor's current lifecycle. An existing row stays as it was, and an actor granted for the first time\nstays out until the grant is run again. This includes the bare grant printed on 0.58.0 by\n`cotal login`, `cotal status`, `actor list` and the not-granted refusal. Look for it in provisioning\nscripts, onboarding runbooks and anything that pastes those hints.\n\n### Upgrade order\n\nChange the scripts before the ledger machine is upgraded, and make each grant name all three flags.\n0.58.0 and 0.59.0 both accept that form. To keep a wide row, write its defaults out:\n\n```sh\ncotal actor grant <actor> --sub <IdP subject> \\\n --scope spawn,role:default --allow-subscribe '>' --allow-publish '>'\n```\n\nSwitch to `--full` only once the ledger machine runs 0.59.0. 0.58.0 refuses it with\n`Unknown option '--full'` before it reads the ledger. Brokers, managers and participant machines\nneed nothing for this break, so their order is the one the section above gives.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused grant changes nothing. The\nexposure is a grant script that runs against 0.59.0 before it was changed: it fails and grants\nnothing.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. On the ledger machine, save the output\nof `cotal actor list` to compare rows after the changed scripts run, and list the scripts that call\n`cotal actor grant`.\n\n### The upgrade end to end\n\n```sh\n# on the ledger machine, still on 0.58.0\ncotal actor list > actors-before.txt\ngrep -rn 'actor grant' <your provisioning scripts>\n# make every grant name --scope, --allow-subscribe and --allow-publish, run them, then upgrade\nnpm i -g cotal-ai@0.59.0\ncotal actor list | diff actors-before.txt -\n```\n\nBoth refusals quoted here were run on 0.58.0 and on the 0.59.0 code. That brokers, managers and stored\nrows need nothing is read from the change, which touches only the CLI and its hints, and was not run\non a live split deployment.\n\n## Repeated flags refused in 0.59.0\n\nA `cotal` flag given more than once is now a usage error unless the command declares it repeatable.\nOn 0.58.0 the last value won with no message, so `cotal down web --space a --space b` acted on `b`\nwhile a wrapper that checked the first `--space` verified `a`. The break is in the command-line\nparser on the machine that runs the command, including commands added with `cotal ext add`. No\nstored state, credential or wire message changes.\n\n### What keeps working\n\nA command line that gives each flag once parses as it did on 0.58.0, in any order and in the\n`--flag=value` form. Flags whose help says repeatable, such as `--opt` and `down --session-store`,\nstill collect every value. A flag-shaped word after `--` is still a positional. The daemons, units\nand agents that `cotal` starts for itself are given each flag once, so a fleet driven only by `cotal`\ncommands typed by hand needs no action.\n\n### What stops working\n\nA command line that repeats any other flag exits 1 before the command runs. It prints\n`Option '--space' cannot be repeated`, or `Option '-f, --file' cannot be repeated` for a flag with a\nshort form, followed by the command's help. `-f` and `--file` count as the same flag. Look for it in\nscripts, aliases and wrappers that append a flag to override one set earlier, such as a fixed\n`--space` followed by `\"$@\"`.\n\n### Upgrade order\n\nChange those scripts first so each flag is given once. 0.58.0 and 0.59.0 both accept that form.\nBrokers, managers and participant machines need nothing for this break, and each machine's CLI\napplies it when that machine is upgraded, so their order is the one the sections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused command does nothing. The\nexposure is a script that still repeats a flag when it runs on 0.59.0: it exits 1 instead of acting on\nthe last value.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the scripts, aliases and wrappers\nthat call `cotal` so each one can be checked.\n\n### The upgrade end to end\n\n```sh\n# still on 0.58.0\ngrep -rn 'cotal ' <your scripts and wrappers>\n# give each non-repeatable flag once, then upgrade\nnpm i -g cotal-ai@0.59.0\n# run each changed script; a repeat left behind exits 1 with the usage error and does nothing\n```\n\nThe refusal and its messages were run against the 0.59.0 parser and `cotal topology view`. That the\nargument lists `cotal` builds for its own processes give each flag once is read from the code, and\nwas not run on a live split deployment.\n\n## Detached spawns from a seat's shell in 0.62.0\n\nOn a static or open mesh, `cotal spawn --detach` run inside a managed seat's shell now launches as\nthat seat. On 0.61.0 it minted a one-shot operator instrument, so the manager recorded that\ninstrument as the spawner and the seat's own `cotal_despawn` of the child was refused with\n`not authorized: <seat> was not spawned by <caller> (admin tier required)`. The break is in the CLI\non the machine where the seats run. No stored state, credential or wire message changes.\n\n### What keeps working\n\n`cotal spawn --detach` from an operator terminal or from a script outside any seat launches as\nbefore, and so does any call with `--creds`, one aimed at a space other than the seat's own, or a raw\nopen target named with `--server` and an unregistered `--space`. A user-auth mesh is unchanged. A seat with\n`capabilities: [spawn]` still spawns from its shell, and can now stop that child with\n`cotal_despawn`. `--on <instance>` from a seat's shell still lands on that manager instance, now as\nthe seat.\n\n### What stops working\n\n- On a static mesh, a seat without `capabilities: [spawn]` can no longer spawn from its shell. Its\n own credential holds no spawn subject, so the broker refuses the request and the command exits 1.\n- A child launched from a seat's shell is now that seat's child, so the manager stops it when the\n seat exits, as it does for a `cotal_spawn` child. A child that has to outlive the seat that\n started it now goes with the seat.\n- A seat launched without `COTAL_SPACE` is placed by its static credential. Every connector sets\n that variable, so this only reaches a hand-built launch: from such a seat's shell, a spawn aimed at\n a static space that holds no credential for the seat is refused instead of running as the operator.\n\n### Upgrade order\n\nOnly the CLI that seats run from their shell changes, which is the one installed on the host where\nthe seats run. Brokers and managers need nothing for this break, so their order is the one the\nsections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it. A child already running when you upgrade\nkeeps the spawner the manager recorded at its launch.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the agent files whose seats run\n`cotal spawn --detach` from their shell, note which of them lack `capabilities: [spawn]`, and note\nwhich of their children must outlive the seat.\n\n### The upgrade end to end\n\n```sh\n# still on 0.61.0: find the seats that spawn from their shell\ngrep -rln 'cotal spawn' .cotal/agents\n# add `capabilities: [spawn]` to each of those agent files that lacks it, and launch any child\n# that must outlive its seat from an operator terminal instead\nnpm i -g cotal-ai@0.62.0\n```\n\nThe attribution, the despawn, the refusal of a seat without `spawn`, the stop on seat exit and a\nseat's `--on` spawn were run on a local static mesh, and the attribution and the despawn on a local\nopen mesh.\n\n## From 0.53.0 to 0.54.0\n\nManager calls now borrow an instance-bound `manager-caller` credential. Followed mutations require\n`manager.goal-result` on the selected manager, so a compatible issuer, manager and client must be\nloaded together. An older manager is refused before a followed mutation; upgrading an installed\nbinary alone does not replace code in a running manager, connector or embedded client.\n\n### Preserve state before changing processes\n\nSnapshot the broker's durable storage using its supported backup procedure, the host authority and\nactor ledgers, and each participant's manager identity, runtime custody records, credentials and\nsaved sessions. Include the embedding application's database and configuration under its supported\nbackup procedure. Record the loaded package versions and the CLI path used by bearer helpers.\nKeep these copies private. Do not change the IdP issuer, regenerate manager identities, rotate agent\ncredentials or recreate tenant storage to make the upgrade pass.\n\nNo ledger, goal-history or session conversion is required for this change. Existing ordinary\nmessaging credentials retain their normal expiry rules. New manager-caller credentials are obtained\non demand from the current grant; old manager-call credentials do not gain the new view automatically.\nExisting accepted goals remain durable and must not be submitted again merely because observation\nwas interrupted. Fresh remote registration publishes its service status at the current revision and\nepoch; do not seed that status manually.\n\n### Upgrade the split deployment\n\n1. Stage one pinned 0.54.0 package set for the host and participants, including the embedding SDKs.\n Pause new manager mutations and let accepted work settle where possible before reloading processes.\n2. Upgrade the host issuer and embedding first. Keep the broker, its account identities and durable\n storage in place. Then load the matching manager release on each participating machine.\n3. Preserve active seats through the runtime's supported update path. A Linux custodial runtime may\n release and re-adopt seats within its 600-second unattended window; verify the actual runtime,\n custody records and process identities before relying on it. A legacy PTY runtime without release\n support cannot preserve active seats through a generic manager restart. Drain it at an approved\n idle window instead of signalling the manager or replacing conversations.\n4. Reload the clients and connectors through their session-preserving host controls. Refresh any\n bearer helper captured from an older immutable CLI path. A transport-only reconnect does not reload\n JavaScript. Verify authenticated instance selection, a read-only manager command and canonical\n result recovery before allowing new followed mutations.\n\nTreat the interval from issuer reload through compatible manager/client reload as a manager-control\noutage. Mixed versions can refuse discovery or commands; there is no promised rolling transition.\nOrdinary agent sessions survive only where their runtime and credentials permit it. If verification\nfails, keep mutations paused and repair forward from the preserved state rather than resetting it.\nThis release does not add host-backed enrollment or terminal release for stock participant detached\nagents; see [Remote supervised agents](run-a-mesh.md#remote-supervised-agents).\n\n## From 0.48.2 to 0.49.0\n\n0.49.0 changes how a credential's authority is recorded. A credential is no longer only a signed\nfile: it is an *issuance*, with a generation the issuer chose and durable evidence of the ceiling it\nwas granted under. The important consequence for a running deployment is not at connect time. It is\nat renewal time.\n\n### What keeps working without any action\n\n- **Existing agent credentials keep authenticating.** A credential minted under 0.48.2 is not\n revoked and is not rejected at connect. Nothing needs to be re-issued to bring the fleet back up\n after the upgrade.\n- **The channel registry survives.** Channels, their replay settings, descriptions, and usage text\n are ordinary durable state and are not rewritten by the upgrade.\n- **`cotal deliver` is still a standalone command.** Running the delivery daemon as its own process\n remains supported; it is not restricted to being a child of `cotal up`.\n- **`cotal join` keeps its flags.** In particular `--lifecycle-uid` is not new in 0.49.0. It has\n been required alongside `--creds` since well before this release, and the pairing rule did not\n change here. A scripted external join that worked under 0.48.2 works unchanged.\n\n### What does not migrate\n\n**A credential minted before 0.49.0 cannot be renewed.** Managed agent credentials carry a\n24-hour lifetime and the manager re-signs one once it passes **75%** of its life, ticking every\nquarter of the TTL so a tick always lands inside that window. When the manager reaches a credential\nthat carries no issuance, it refuses to renew it and logs the agent by name:\n\n```\n! managed cred renewal <agent>: renewManagedStaticCred: <agent> carries no issuance;\n a static credential minted before SPEC 13.15 is not renewed under an unbound generation\n - respawn the agent\n - the agent dies loud at this cred's expiry unless it is reminted\n```\n\nSo the fleet comes up fine, runs normally, and then each agent stops at its own credential's\nexpiry, within roughly a day of the upgrade, one at a time rather than together. The refusal is\ndeliberate: the renewal would otherwise have to invent a generation nobody issued, which is the\nstate the release exists to remove.\n\n**Respawn the managed agents as the last step of the upgrade.** For this particular upgrade the\nrespawn is not optional: stopping a 0.48.2 manager ends its agent processes whichever CLI you use,\nfor the reason given under the outage window below. The respawn is how they come back, and it is\nalso what mints each credential as an issuance so it renews from then on. One planned pass over the\nfleet is the whole job. Skipping it leaves agents stopped and, for any credential that survived\ninto 0.49.0 unminted, brings the renewal cliff above a day later, one agent at a time.\n\n### Credentials you minted yourself\n\n**A credential you minted with `cotal mint` is a different case, and it very likely needs\nnothing.** The distinction that matters is not the word \"static\", which covers both. It is **what\nminted the credential and who owns its renewal**. A credential the **manager** minted for an agent\nit spawned carries a lifetime and is renewed by the manager, so it is the subject of everything\nabove. A credential **you** minted with `cotal mint` and handed to an external peer is issued with\n**no expiry at all**, and no manager renews it: it is not in the sweep, so there is no renewal to\nfail. It keeps working after the upgrade, and re-minting it would mean coordinating with a third\nparty for no gain.\n\nThe manager says which one it is holding. Where a credential has no expiry to reach, the sweep\nnames it and moves on rather than refusing:\n\n```\n! managed cred renewal <agent>: credential is unbounded - not renewed\n (a pre-TTL credential stays as minted until respawn)\n```\n\nRe-mint an external peer's credential only if you want it to carry a lifetime, and at a time you\nchoose.\n\n### How to read the boot log\n\nA 0.49.0 manager starting over an existing space may print lines like:\n\n```\n verified evicted: <holder-key> (3/12)\n already verified (durable): <holder-key>\n\u2713 boot self-heal: manager/<id> registration gate reopened at generation <n>\n```\n\nThese are **not** a credential migration, and reading them as one is the most likely way to\nconclude the fleet is fine when it is not. They come from the manager repairing **one** endpoint\nregistration gate that a previous restart left frozen, and they enumerate that single gate's\ncredential-family holders as it verifies each one evicted. `already verified (durable)` on a later\nstart is the repair cursor resuming, not a credential that became durable. The repair is real and\nuseful (it is what previously needed `cotal reconcile-gate` by hand), but it says nothing about\nwhether your agent credentials carry issuances. The renewal refusal above is the signal that does.\n\n### Which side to upgrade first in a split topology\n\nMove the manager first.\n\nThe stores 0.49.0 introduces are created by the **manager** at its own boot, not by the broker.\nThey are create-or-verify and idempotent, so a 0.49.0 manager brings the space's authority stores\nup to the new shape itself, and it does so against whichever broker is answering.\n\nBeing honest about the evidence behind each direction, because they are not equally established:\n\n- **Broker-first was measured on a live 30-agent deployment** (issue #1578). Upgrading the broker\n first locks the old manager out immediately: `cotal up` re-renders the broker's generated config\n from the trust record, and after the restart the still-0.48.2 manager is refused on every\n connection with an `authentication error` naming the Nkey, continuously. That text comes from the\n broker process, not from a Cotal command, so match on its shape rather than on an exact string.\n `cotal ps` reports zero agents while\n the agent processes are still alive, because the manager has lost its view of them, not because\n they died. Upgrading the manager clears it immediately.\n- **Manager-first is reasoned from where the new stores are provisioned**, not from a measured\n fleet upgrade. It is the recommended order because the manager is the component that creates what\n 0.49.0 adds, but it has not been run end to end on a production split topology at the time of\n writing. Treat it as the better-supported order rather than a guaranteed one, and keep the\n rollback below ready either way.\n\nWhichever order you pick, **this is not a rolling upgrade**. Between the two steps the mesh is down\nand the manager cannot see its agents. Go straight through rather than pausing between them, and\nschedule it as an outage window.\n\n### What the window looks like\n\n- **The managed agent processes do not survive step 1, in either order.** This is the one place\n where the obvious reordering does not rescue you, so it is worth understanding rather than\n working around. Sparing agents on a bare manager stop is a **handshake**: a 0.49.0 manager\n publishes a capability file proving it can release its agents, and a 0.49.0 `cotal down` refuses\n the stop unless it finds one. **A 0.48.2 manager never publishes that file**, because the\n mechanism ships in the release you are installing. So the old CLI against the old manager sends a\n plain stop and takes every seat with it, and the new CLI against the old manager either refuses\n (leaving `--with-agents`, which reaps deliberately) or falls to the legacy path, warns that it\n cannot verify the manager can spare its agents, and signals it anyway.\n- **You can confirm which side you are on in one command, without stopping anything.** The flag that\n marks the newer behaviour is absent from the older CLI, and its summary line makes the difference\n plain:\n\n ```\n $ cotal down --help # on 0.48.2\n cotal down - stop the whole local stack, or name only the components to stop\n\n $ cotal down --help # on 0.49.0\n cotal down - stop the whole local stack (managed agents stay running unless --with-agents), ...\n ```\n\n If your `cotal down --help` does not mention `--with-agents`, stopping the manager stops the\n agents with it.\n- **Therefore the respawn in step 5 is mandatory recovery for this upgrade, not an optional pass.**\n It is also the step that re-mints credentials as issuances, so it is the same action either way.\n Plan the window to include it rather than treating it as cleanup.\n- The **manager's view** of them is lost while the two sides disagree, so `cotal ps` reports zero\n and control commands do not reach seats.\n- **Messages are not delivered** while the mesh is down.\n- The window is as long as it takes to restart the second component, plus the manager's own start.\n It is minutes, not hours, provided you do not stop between the steps.\n- **Nothing self-heals if you stop halfway.** The refusal is continuous until both sides match.\n\n### Snapshot this before you start\n\nTake these while the deployment is still on 0.48.2. The two `cotal` reads are live reads and must\nhappen before anything stops.\n\n- **A filesystem or volume snapshot of both containers**, if your platform offers one. This is the\n only rollback that covers every case, and it is what the reporting deployment used.\n- **`cotal backup create <dir>`**, for the durable space state, **but read the next paragraph before\n you rely on it**: on a split broker and manager topology it is very likely unavailable to you, and\n the volume snapshot above is your actual rollback.\n- **The trust records and credential directory** under `.cotal/auth` on the manager host, including\n the per-space material directory. These are what a re-mint would otherwise have to replace.\n- **A copy of the channel registry**, so you can verify it came back rather than assuming it did:\n `cotal channels list` before and after.\n- **The output of `cotal ps`**, so you know how many seats you expect to see afterwards and can tell\n a lost view from a lost agent.\n\n#### `cotal backup` on a split topology\n\n**`cotal backup create` cannot read a running stack.** It requires a completed cut, and only\n`cotal down --preserve-state` publishes one:\n\n```\n$ cotal backup create ./backup.0482\n\u2717 backup requires a completed cut; run `cotal down --preserve-state` first\n```\n\n**And `cotal down --preserve-state` requires a manager alive on the host you run it from.** It uses\nthat manager to attest that every retained child stopped, and the check is deliberately fail-closed:\na manager that is dead or merely uncertain refuses rather than preserving an unproven cut. The check\nreads a local pidfile, so a **remote** manager does not satisfy it. On a split topology the broker\nhost has no local manager, which means the documented durable-backup path is not available there.\n\n**Measured rather than assumed, at 0.48.2**: the backup refusal above is executed output. The\npreservation requirement is read from `down.ts` at the same tag, where the preserve path asks a\nmanager to prepare an inventory and then requires that manager to be locally alive before it\ncommits. The part not executed end to end is a genuine two-host split, which needs two real hosts.\n\n**What to do instead.** Use the filesystem or volume snapshot of both containers. That is the\nrollback the reporting deployment actually used, it covers the broker's durable state and the\nmanager's credential material together, and it does not depend on either component being able to\nattest for the other. If you want `cotal backup` as well, take it from a host that does have a live\nlocal manager, and understand it is a second copy rather than the primary rollback.\n\n**This looks like a product limitation rather than a documentation gap**, and it is written here as\none so an operator is not left thinking they mis-typed a command. The upgrade path for the exact\ntopology this page is addressed to cannot use the documented backup command.\n\n### The upgrade end to end\n\n```bash\n# 0. on 0.48.2, STILL RUNNING: record what you expect to see afterwards.\n# These two are live reads, so they must happen before anything stops.\ncotal channels list > channels.before\ncotal ps > ps.before\n\n# 1. manager host. READ THE NOTE BELOW THE BLOCK FIRST: this step ends the\n# managed agent processes whichever order you choose, and the respawn in\n# step 5 is how they come back. It is recovery, not tidying.\n#\n# STOP THE MANAGER WITH THE 0.48.2 CLI, BEFORE INSTALLING 0.49.0. The\n# order matters and it is not recoverable once you install: a 0.49.0\n# `down manager` REFUSES to stop a 0.48.2 manager whose pid record carries\n# a start token, which is every manager on a platform that can read one\n# (Linux can):\n# refusing bare manager stop: ... does not prove this manager can detach\n# its agents; use --with-agents or stop the agents explicitly\n# The refusal names two remedies and NEITHER clears it for this case. The\n# check reads a capability file that only a 0.49.0 manager writes; it never\n# counts agents, so stopping them first changes nothing. And `--with-agents`\n# is whole-stack only, so `down manager --with-agents` is refused by its own\n# flag rule. See #1592.\ncotal down manager # the 0.48.2 CLI, still installed.\n # 0.48.2 has no --with-agents; this\n # is the whole route. On a host that\n # runs the whole stack, the 0.49.0\n # `cotal down --with-agents` after\n # installing is the alternative.\nnpm install -g cotal-ai@0.49.0 # ONLY after the stop above\n# `supervise` RUNS IN THE FOREGROUND and holds the terminal until you stop\n# it. There is no --detach on this command. Start it under whatever keeps\n# your manager alive normally (systemd unit, container entrypoint, or a\n# second terminal), and run the remaining steps from another shell.\ncotal supervise --space <space> --server nats://<broker>:4222\n\n# 2. broker host: stop the stack.\n# NOT `--preserve-state` on a split topology: it needs a manager alive on\n# THIS host to attest its children stopped, and yours is on the other one.\n# Your rollback is the volume snapshot from \"Snapshot this before you\n# start\", not `cotal backup`.\n# See \"cotal backup on a split topology\" above.\ncotal down\n\n# 3. broker host: install 0.49.0 and start it again\nnpm install -g cotal-ai@0.49.0\n# Record the manager log's size BEFORE starting, so step 3a can tell THIS\n# boot's output from every earlier one. It must be captured here, ahead of\n# the start: taken afterwards it sits past the new line and the wait hangs.\n# `<spaceKey>` is NOT the space name. It is lowercase hex of the name's\n# UTF-8 bytes, so space `prod` is `manager.70726f64.log`. Do not guess it:\n# `cotal up` prints the real path on its launch line. Substituting the\n# plain name points at a file that does not exist, and the wait below then\n# burns its full timeout before telling you.\nLOG=.cotal/manager.<spaceKey>.log\nOFF=$( [ -f \"$LOG\" ] && wc -c < \"$LOG\" || echo 0 )\ncotal up --detach --host 0.0.0.0 --space <space> --no-manager\n\n# 3a. SPLIT TOPOLOGY ONLY: `--no-manager` above boots the broker (and the\n# delivery daemon) with NO local manager on the broker host, so there is\n# no wait-and-stop step on a current cotal-ai. The rest of this step is\n# the OLDER-host recipe, kept because the flag is refused there and that\n# refusal is your signal you are on it: without the flag the `up` also\n# starts a local manager, and you must wait for the log to show it is up,\n# then stop it, or you finish the upgrade with two managers and the one\n# you did not intend is the one nobody is watching.\n# A bare `grep -q` does NOT wait: it reads once and exits 1 immediately\n# if the line has not been written yet. Bound the wait instead, so a\n# manager that never comes up fails loudly rather than reading as ready.\n# The log is opened APPEND-ONLY, so on any host that has run a manager\n# before, this file ALREADY carries a `manager up` line from an earlier\n# boot. Grepping the whole file therefore matches instantly and waits for\n# nothing. Read only what THIS boot appended, using the $OFF captured in\n# step 3 above (before the start, which is the only point it is correct):\ntimeout 60 bash -c \\\n \"until tail -c +$((OFF+1)) \\\"$LOG\\\" | grep -q '. manager up'; do sleep 1; done\"\n# exit 0 = THIS boot logged it; exit 124 = it never did, so STOP and look.\n# This manager is 0.49.0 and publishes its own spare-capability file, so\n# the bare stop below is NOT the refusal case from step 1.\ncotal down manager # broker + delivery remain\n# On a current cotal-ai the two commands above are unnecessary (nothing\n# to wait for, nothing to stop) and `cotal down manager` simply reports\n# no manager to stop.\n\n# 4. verify the mesh is whole again before touching the fleet.\n# Do NOT compare `cotal ps` against ps.before yet: step 1 ended the agent\n# processes, so at this point it is EXPECTED to be empty, and an empty\n# `ps` is also the signature of the broker/manager mismatch described\n# above. The two are indistinguishable here, so compare what the mesh\n# itself should have carried across instead:\ncotal channels list # compare against channels.before: this SHOULD match now\ncotal ps # expect it to be EMPTY here; ps.before is the target for\n # step 5, not for this step\n\n# 5. the step that is easy to skip: respawn the managed agents so their\n# credentials are re-minted as issuances and can renew. Persona is a\n# POSITIONAL argument here, unlike `cotal stop`, which requires --name.\n# One call per agent:\ncotal spawn <persona> --detach --name <n> --space <space>\n# then the comparison step 4 could not make:\ncotal ps # NOW compare against ps.before: seat count should match\n```\n\nThe mesh is down from step 2 until step 3 finishes. That is the window. On a split topology there is\nno cut and no backup inside it, so the window is the stop, the install and the restart, nothing more.\n\n## Adding a section for a future release\n\n**Every changeset marked breaking adds a section to this page.** A release that changes what an\noperator must do, in what order, or what stops working, is not finished until the section exists.\n`scripts/upgrade-section-gate.mjs` grades a commit range for this: run it as\n`pnpm upgrade-section-gate --base <ref>` and it reds when the range carries a breaking change and\nadds no new release section. CI runs its self-test and, as a step of the `attribution` job, grades\neach pull request's own range as `HEAD^1..HEAD` over the merge snapshot it checked out. That job is\nthe only context in the branch protection rule set, so a red gate FAILS A REQUIRED CHECK AND BLOCKS\nTHE MERGE. The section is not optional and a reviewer cannot wave it through without an\nadministrator overriding branch protection. Be precise about what the check proves either\nway, because one trusted past its evidence is worse than none. It proves a section for a release\n**was written here**. It cannot prove the section is **correct**, or that it describes the break\nthat actually landed, and it cannot see a breaking change that carries no marker at all. Reviewing\nthe words remains a person's job.\n\n**Mark the break, or the gate cannot see it.** Any one of these is enough, and they are the only\nthings it reads:\n\n- a `!` before the colon in the commit subject, as in `feat(core)!: bind hosted runs to the caller`\n- a `BREAKING CHANGE:` footer in the commit body\n- a changeset in `.changeset/` declaring a `major` bump for any package\n\nThe marker must survive the squash. A `!` that lives only in a commit you squash away is not in the\nrange the gate grades, so put it in the subject that lands on `main`.\n\n**The heading is a `##` and names the release**, like `## From 0.48.2 to 0.49.0`. Both matter, and\nneither is a style preference. Coverage is claimed by a heading, so a heading that names\nno release claims every release and distinguishes none: `## Notes` with a sentence under it would\notherwise satisfy the rule. Naming the release also makes the section the one an operator upgrading\nthat release will search for. Use `###` freely for detail inside a section. Subsections belong to\ntheir release rather than counting as separate coverage.\n\nName the release that first carries the change: the next version Changesets publishes, which\n`pnpm changeset status --verbose` lists. `bin/package.json` on `main` still reads the release already\npublished. If a release is cut while the change is open, the change ships in the release after it,\nso move the heading before merging. The gate accepts any version in a heading, so before merging a\nrelease pull request, check every heading added since the previous tag against the version it\npublishes.\n\nA section is written for the operator, not for the reviewer. It answers, in this order:\n\n1. What keeps working with no action at all.\n2. What does **not** migrate, and when that becomes visible. Name the log line if there is one.\n3. The order to move components in for a split topology, and why that order.\n4. What the outage window looks like, including what survives it.\n5. What to snapshot before starting.\n6. The commands, end to end.\n\n**Where an answer was not measured, say so in the document rather than guessing.** An operator who\nknows which half of a recommendation is reasoned and which is measured can plan around it; one who\nfinds out afterwards cannot.\n"
|
|
57978
57978
|
},
|
|
57979
57979
|
{
|
|
57980
57980
|
"slug": "watch-a-mesh",
|