@cotal-ai/connector-core 0.52.0 → 0.52.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,5 +1,5 @@
1
1
  export const DOCS_BUNDLE = {
2
- "version": "0.52.0",
2
+ "version": "0.52.1",
3
3
  "generatedFrom": "docs/*.md + SPEC.md + spec/cotal-lang.md + spec/cotal.schema.json",
4
4
  "pages": [
5
5
  {
@@ -42,7 +42,7 @@ export const DOCS_BUNDLE = {
42
42
  "title": "Identity",
43
43
  "kind": "Concept (informative)",
44
44
  "summary": "Who can do what on a mesh, and how it is enforced.",
45
- "body": "# Identity\n\n> **Concept** (informative) · **For:** operators and implementers · **Normative:** [SPEC §2](../SPEC.md#2-identity), [§9](../SPEC.md#9-nats--jetstream-security-and-authorization), [§10](../SPEC.md#10-connection-and-onboarding), [Appendix B](../SPEC.md#appendix-b-profile-acls)\n\nWho can do what on a mesh, and how it is enforced. The design goal: the mesh is a **real\nboundary against untrusted peers in a shared space**; an agent can only speak as itself\nand only where its declared permissions allow, enforced by the broker, not by agent\ngoodwill. What that boundary does and does not protect is the\n[security model](security.md); the exact ACLs are\n[SPEC Appendix B](../SPEC.md#appendix-b-profile-acls).\n\n## On by default\n\n`cotal up` provisions a JWT-authed space; `cotal up --open` runs an unauthenticated dev\nmesh instead. Both bind loopback by default. `--host 0.0.0.0` widens the bind\nindependently, so \"network-reachable\" never silently means \"unauthenticated\". Open mode\nis for quick local experiments and sits outside every security claim\n([SPEC §9](../SPEC.md#9-nats--jetstream-security-and-authorization)).\n\n## Shared identity\n\nAn agent's wire identity is a **principal**: an `owner.actor` pair, where the owner is\nthe account (a human, or an organization) the agent acts on behalf of, and the actor is\nthe agent's own handle under that owner ([SPEC §2](../SPEC.md#2-identity)). The same pair\nis the card id, the sender tokens in every subject it publishes, the presence key, and\nits durable-consumer names. On an open dev mesh the owner is the literal `local`; on a\nper-user-auth mesh it is a derived token (`u_` plus 26 characters, so no PII rides the\nwire). The connection still authenticates with an **nkey**, generated locally (the signer\nonly ever sees the public half), but the nkey is the transport credential, not the\nidentity: it scopes only the per-connection reply inbox.\n\n**The sender is encoded in the subject.** Every publish carries the sender's owner and\nactor in positions the broker's permissions pin to that connection, so an agent *cannot*\nemit as anyone else: not as another owner, and not as a sibling actor under its own\nowner. Receivers verify the payload's `from.id` against the subject sender and reject\nmismatches; sender authenticity is broker-enforced end to end\n([SPEC §3](../SPEC.md#3-subject-layout), [§5](../SPEC.md#5-envelopes)).\n\n**Account = space, user = agent.** A space is one NATS account, a server-enforced\nisolation boundary. An operator signs the account; an account **signing key** mints\nper-agent user JWTs.\n\n## Provisioner\n\nThe **provisioner** is whoever holds the account signing key. It mints profile-scoped\ncredentials and pre-creates the durables agents may only *bind* (their DM inbox, their\nrole's task queue). The manager hosts it today, but nothing is manager-special about it;\nprivilege attaches to the signer, and a space can run without a manager.\n`cotal mint <name> --profile <agent|observer|admin>` is the out-of-band path; spawn calls\nthe same library ([CLI](cli.md)). Minting static creds is a **static-auth** surface: a\nper-user-auth space refuses it, because agents there join under a logged-in user, never\nvia a handed-out file (see *Per-user auth* below).\n\nAgent-profile minting resolves one mesh root for the persona ACL, account signer and default\ncredential storage. If the current folder holds trust for a different space or account, mint\nrefuses and names both roots. It never signs one root's persona policy with another root's authority.\n\n## Profiles\n\nEvery credential is a profile: an explicit allow-list built from the same\nsubject/stream/durable builders as the wire layout, so ACLs cannot drift from it. The\nnormative shapes are [SPEC Appendix B](../SPEC.md#appendix-b-profile-acls); in brief:\n\n| Profile | Is |\n|---|---|\n| **agent** | The ordinary peer: publishes as itself to its declared channels, reads within its read ACL + its own DM/task inboxes. Its read-only presence and channel-registry watches may create, inspect, and delete only their own client-managed ordered consumers; those cleanup grants cannot delete KV records or streams. |\n| **observer** | Read-only chat + presence; DMs invisible. What `cotal console` runs. |\n| **admin** | Elevated *read-only* god-view: sees DMs and anycast live, still writes nothing. A deliberate opt-in (`cotal web`). |\n| operator-side | Narrow single-purpose creds for the machinery (supervising, provisioning, teardown, delivery); the reference implementation splits these so no one connection can read every DM *and* delete every stream ([security model](security.md)). |\n| **run-driver** | One workflow run and takeover attempt: its journal subject, replay durable and run-owned record writes. Store reads and effects go through the host. |\n| **run-mediator** | The trusted hosting process's separate connection for workflow effects and leader reads. It exposes journal-checked, run-bound operations and never hands its credential to the driver. |\n| **run-operator** | One served run read, or one half of an answer, minted per call: a read holds the records walk, one run's replay and the admission read; the answering half is minted for one checkpoint token and holds that pause's answer record and settle alone. |\n| **issuer** | One issuance window: the party holding the space signer mints it for a few minutes to stage and release a credential's evidence, retire the issuances a lifecycle terminal leaves behind, or resolve the evidence a request rides. |\n| **run-admitter** | One run's admission record or revocation marker, minted per run for a minute: two exact keys in the admission store and nothing else. |\n\n**An agent's channel scope is three verbs**: `subscribe` (reads at boot),\n`allowSubscribe` (read ACL), `allowPublish` (post ACL, default-deny), declared in its\n[agent file](agent-files.md) or [manifest](manifest.md), minted into its cred. One card\nwith the recipes: [Channels & permissions](channels-and-permissions.md).\n\n**DM confidentiality** holds against peers by construction: deliveries ride per-identity\ninbox prefixes, and the DM/task consumers are provisioner-pre-created and bind-only, so an\nagent cannot create a consumer filtered to someone else's inbox\n([SPEC §9](../SPEC.md#9-nats--jetstream-security-and-authorization) items 1–5).\n\n## Issued authority\n\nA credential says who is calling. It does not, by itself, say what the caller was granted, and a\nhost that acts on a caller's behalf (a workflow run, [workflows](workflows.md)) needs that from\nthe issuer, not from a ledger that may have changed since. So a static agent credential is an\n**issuance** ([SPEC §13.15](../SPEC.md#1315-issued-authority)): before the material is handed\nout, the issuer records the credential's final permission ceiling, as evidence keyed by a fresh\n**generation**, in `cotal_issued_<space>`. That store is append-only at the broker and the\nevidence is read as the first message on its key, so nothing written later on the same key, by\nanyone, changes what a resolver sees. The credential's endpoint rows then ride a versioned\nrail, `cotal.<space>.ep.v1.…`, with that generation pinned beside the caller triple, so the\nbroker binds every request to the ceiling the issuer accepted. The legacy `ep.` rail and the\n`ep.v1.` rail are disjoint subject spaces; a credential holds rows on one of them, and every\nendpoint serves both.\n\nA connected client learns its generation by reading one row in `cotal_accepted_<space>` under a\ntoken the launching party chose at mint time, through a per-key read grant its own ceiling carries. It never trusts\nwhat the file says. A renewal keeps the generation only while the ceiling is byte-identical to the\nevidence; a changed scope is a fresh issuance on a fresh generation, adopted by a new connection.\nA static agent's evidence names its credential ledger family as the source it depends on, and the\nlifecycle terminal that retires that family retires its issuances with it.\n\nThe manager, `cotal spawn`, and the CLI's control-caller instruments all mint through an `issuer`\nsession. The deployer instrument does not: it is the one instrument with no default expiry, and an\nissuance with no lifecycle gate must carry one, so a deploy rides the legacy rail under its own\nlifecycle uid. Only workflow `run-start` requires the binding today: a request for it on the legacy\nrail is refused with `permission-denied` and the detail `ai.cotal.ep.unbound-caller-authority`\nnaming the caller. Every other command serves both rails.\n\n## Declared capabilities\n\nControl-plane power is a **declared capability**, not a default. An agent file carrying\n`capabilities: [spawn]` gets the privileged control subject minted into its cred: spawn,\nplus stop/despawn of its *own* children, plus persona definition. Without it, an agent can\nonly self-despawn and pull or yield the run turns addressed to it. `capabilities: [run]` mints\nthe manager's workflow-run commands (start, resume, answer, status, list) together with the spawn\nset, since a program the agent starts may spawn; the manager drives the run under a per-run\n`run-driver` credential of its own, never the caller's. The tool surface mirrors the grant:\n`cotal_spawn` / `cotal_persona` / `cotal_personas` are injected only for `spawn`, and `cotal_run`\nonly for `run` ([agent files](agent-files.md)). Destructive\noperator ops (history purge, cross-agent stop) live on a third tier no agent credential\nreaches. Persona redefinition separates content from policy; the write path takes only\n`model`/`persona`, so a peer cannot grant itself a capability by redefining a file.\n\n## Per-user authentication\n\n`cotal up --user-auth --idp <auth base URL>` (or manifest `broker.auth: \"user\"`) puts a\n**human identity plane** above the per-agent one: people sign in to an external IdP once,\nand every connect is authorized live against the operator's **actor ledger**. No creds\nfiles to hand out, and revoking a grant actually bites.\n\n**The flow.** Each person runs `cotal login --idp <url>` once per machine. After that,\nany command works: cached IdP session → fresh IdP proof per connect (so IdP-side\nrevocation bites here too) → the configured exchange turns it into a short-lived Cotal bearer →\nthe broker's **auth callout** checks the bearer and the ledger at connect time and mints\na scoped credential on the spot. Every bearer also names a **root credential** row in the\nspace's credential ledger, proved live at each connect, so revoking that one credential\nbites at the very next connect. The operator grants access with\n`cotal actor grant <actor> --sub <their id>`; a bare grant is the full envelope (all\nchannels, may spawn), and `--allow-subscribe` / `--allow-publish` / `--scope` narrow it.\nNo ledger row, no access; there is no allow-by-default.\n\n**Space catalogs.** A successful authenticated `GET <idp>/token` may advertise one catalog with:\n\n```http\nLink: <https://idp.example/spaces>; rel=\"https://cotal.ai/relations/space-catalog\"\n```\n\nThe target must use HTTPS and the same origin as the normalized IdP URL. Loopback IP literals may\nuse HTTP for local development. A missing, foreign-origin, or insecure link records that this account\nhas no catalog. The client never guesses a path.\n\nThe catalog request carries the opaque cached session as its bearer and returns a complete snapshot:\n\n```json\n{\n \"v\": 1,\n \"account\": { \"idpUrl\": \"https://idp.example/api/auth\", \"issuer\": \"https://idp.example\", \"sub\": \"user-id\" },\n \"spaces\": [\n { \"id\": \"space-id\", \"slug\": \"shared_project\", \"name\": \"Shared project\", \"kind\": \"hosted\", \"role\": \"owner\", \"registration\": {} }\n ]\n}\n```\n\nThe client checks every `registration` with the same `checkUserBundle` validator used by `cotal\nmeshes add`. One invalid entry refuses the whole candidate snapshot. Conditional refresh uses the\ncatalog's `ETag`; a transport error, non-success response, or invalid candidate leaves the prior\nsnapshot intact and reports the failure. The registry is reconciled under the provider's catalog\nlock, and a snapshot whose reconciliation was interrupted is reconciled again by the next refresh\nbefore it counts as fresh or not modified.\n\nThe user-auth registration document may include one closed policy object:\n\n```json\n{ \"policy\": { \"events\": \"required\" } }\n```\n\nNo other key under `policy` and no other value for `policy.events` is accepted. The registry preserves\nthis field for manual, discovered, and enrollment-created entries. A pre-policy manual entry is\nrefreshed from its own pinned exchange origin by the command that consumes the policy, a spawn, a\njoin, or a manager start, after that command's own local refusals; the returned space, broker,\ntransport, IdP, issuer, audience, and exchange pins must all match before only the policy is added.\nA failed expired refresh refuses that operation. Read-only commands such as `status` and `meshes`\nnever refresh. The five-second warm window makes no request.\n\nEvery registration is also bound to the proved account. Its IdP URL must match the account and its\nissuer must match the exact JWT `iss` pin. Its exchange, provisioning, and manager-authority\nendpoints must be same-origin with that IdP. One foreign pin or endpoint refuses the whole\ncandidate snapshot.\n\nThe `slug` is the space identity resolved by `--space`, `use`, registry roots, and collisions. The\n`name` is a display label only.\n\nDiscovered registry entries are owned by the normalized IdP origin plus the proved `sub`. That key is\nstored as an opaque digest, so accounts on one machine never union their spaces and the registry does\nnot persist the subject. A manual or locally started record with the same name is never overwritten.\nLogout removes only the entries owned by the account whose session was revoked. Local teardown,\ncleanup, and liveness pruning do not remove discovered entries.\n\nAn account that previously advertised no catalog is checked again by explicit `cotal sync`. The\nordinary lazy path checks again after its five-second capability window, so an IdP can enable the\nLink for an existing login without making the person sign in again.\n\n**One auth service per space** hosts both halves: the NATS auth callout and the token\nexchange. Its default HTTP listener remains loopback-only and requires the per-start capability\nstored in the owner-only `auth-service.json` file. An operator may add a second listener with\n`cotal up --user-auth ... --exchange-public-port <port> --exchange-public-url https://auth.example`.\nThat listener still binds `127.0.0.1`; put a reverse proxy in front of it and terminate TLS there.\nIn-process TLS is deliberately not another deployment mode: it would duplicate certificate renewal\nand fork proxy-based deployments.\n\nThe public listener has a closed surface: `GET /health`, `GET /jwks`, `POST /exchange`, and\n`GET /.well-known/cotal-mesh`; every other path is 404. It does **not** require the loopback\ncapability. That capability proves same-uid access to a 0600 local file and has no remote meaning;\non the public face the credential is the proof. A human presents an EdDSA IdP JWT checked against\nthe pinned JWKS, issuer, and audience. An agent presents its spawn-time actor token, whose hash must\nmatch a fresh managed-ledger row. The public face mints only two elevated views, both still\ngated on ledger scope `admin`: `channel-writer` (`cotal channels set/default`) and\n`channel-purger` (the dashboard's per-click channel delete). God-view (`admin`), space-history\n`purger`, `deployer`, and `manager-service` stay loopback-only. A managed-agent secret exchange\nnever mints a view on either face. This is not full remote channel management: `cotal web` still\nmints the read-only admin view at startup, so a remote dashboard that needs that god-view still\nfails even when a later delete would mint `channel-purger`.\n\nThe well-known response contains the IdP pins and the actual deny-all sentinel credential remote\nagents need before the bearer-driven auth callout. The pins ride a `userAuth` arm that names the\nauth provider, and that name is the same one the local arm registers under. A document naming a\ndifferent provider than the one serving it would register an entry nothing can resolve, so both read\none constant. Treat it as bootstrap material: the sentinel cannot publish or subscribe, but\nconsumers must still take the bundle only from the intended HTTPS origin and must verify TLS.\n`--exchange-trusted-proxy` opts into peer attribution by the **last** `X-Forwarded-For` hop; use it\nonly when the listener is reachable solely through a proxy you control.\nWithout it, forwarded headers are ignored and the socket address is the peer key. Public failure\nbuckets are per-source and separate from loopback exchange budgets. The in-process LRU retains at\nmost 1024 peer buckets: that bounds memory and isolates ordinary sources, but an attacker cycling\nmore than 1024 trusted-proxy last hops can evict earlier 429 state. It is not a mint bypass; a valid\ncredential is still required, so use upstream reverse-proxy rate limiting when that throttle-escape\nmatters to the deployment.\n\n### Enrollment redeem\n\nA remote owner may pre-mint a one-time enrollment for a seat that has no browser, TTY, or cached\nIdP login. The enrollment is a secret-bearing URL. The client performs one request:\n\n```http\nGET <enrollment URL>\n```\n\nIt sends no `Authorization` header and no request body. The URL must be HTTPS, except for plain HTTP\nto a loopback IP literal. The client redeems only an enrollment URL that is already in canonical\nform and contains none of `\\ @ ? #`. That is checked on the raw string before parsing, so every\nrewrite a URL parser would perform, backslash folding, userinfo erasure, scheme or host case\nfolding, default-port removal, dot-segment resolution, and short-host canonicalization, is a refusal\nrather than a redeem of a URL the owner never minted. Redirects are refused. The client never\nretries because a successful claim deletes the server-side token row. The token expires five minutes\nafter mint.\n\nSuccess is `200` with this JSON object:\n\n```text\nspace\nbrokerAccess { kind, ... }\nowner\nactor\nlifecycleUid\nactorToken\nsentinelCreds\nauthServiceUrl\nidp { url, issuer, audience }\nsubscribe[]\nallowSubscribe[]\nallowPublish[]\n```\n\nThe grant arrays are informational; the broker row remains authoritative. A stock-dialable\ndeployment also includes `server`, `tlsRequired`, `userAuth`, and optional `policy`, forming the same user-bundle\nsuperset that `cotal meshes add --user-auth-file` accepts. That lets a bare seat register the mesh\nfrom the enrollment response before launch. For `brokerAccess.kind: \"direct\"`, the stock `server`\nmust equal `brokerAccess.url` byte for byte or the client refuses the bundle before registration. A\ntunnel kind carries no dial address, so its `brokerAccess` is not compared to the operator-asserted\nstock `server` face.\n\nUnknown, expired, revoked, and already-used enrollments are intentionally indistinguishable. They\nall return `404 {\"error\":\"unknown, expired, or already-used enrollment\"}`. The client reports only\n`enrollment refused: unknown, expired, or already-used; ask the owner for a fresh one`. It does not\nguess which case occurred.\n\nAfter redeem, the seat stores only the normal remote user-mesh and agent material. The actor token\nis exchanged at `authServiceUrl` through the existing `agent-bearer --exchange-url` path. The\nenrollment URL is not logged, persisted, or forwarded into any child process, including the bearer\npreflight and harness.\n\nThe service starts with the broker, is torn down by `cotal down`, and holds the\ndata-account signing key for the callout (a running manager is the other standing holder, for\nthe creds it mints); the operator seed never enters it. It also owns the space's two authority\nstores (lifecycle records and the credential ledger), provisions them at boot, and refuses\nconnects it cannot credential-check against them; there is no fallback path. If it\ndies while the broker lives, re-running `cotal up` heals it, and a boot whose auth\nservice never became ready exits non-zero, so automation never reads a dead identity\nplane as success. Changing any public-listener flag requires `cotal down` followed by `cotal up`\nwith the new values; a refresh adopts an already-running auth service rather than silently replacing\nits listener policy. \"One per space\" is enforced, not assumed (SPEC §13.13): at boot the\nservice takes a broker-backed ownership claim, so a second same-space auth process refuses\nwith instructions instead of silently splitting the plane, and a crashed one's claim is\nreclaimed only once the broker confirms its connections are gone. That verdict is trusted only\non a standalone broker (a clustered one refuses the reclaim, since a partitioned member\ncould still hold them). If the claim's connections die mid-run, the service downs itself\nloudly instead of serving from a half-dead plane.\n\n**Your agents are yours.** `cotal spawn` on a user mesh grants a managed actor under the\n*spawning operator's* owner and launches the agent with a bearer command instead of a\ncreds file. The agent exchanges its spawn-time secret for short bearers (five minutes or\nless) and refreshes ahead of each expiry. Rows are runtime grants: every start rotates\nthe secret, every stop or despawn revokes the row, so a non-running agent holds no\nstanding authority. Manifest deploys (`up -f`) stamp the logged-in owner into the launch,\nso those agents are yours too.\n\n**Despawn tears the lifecycle down, then frees the name.** When you despawn an agent, the manager\ndrives the *full* teardown of that lifecycle: it shreds the local credential files, revokes the\nagent's standing mint authority (its ledger row, so a copied token can no longer mint a fresh\ncredential), deletes its broker footprint (the lifecycle-keyed durables + read-ACL row), and asks\nthe auth service to *retire* the lifecycle (settle in-flight work, evict the departed credentials,\nrecord it retired). The name is held *reserved pending retirement* until **all** of that completes,\nthe broker-footprint cleanup, the standing-authority revoke, **and** the lifecycle retirement, not the\nretirement alone, so a same-name respawn in the gap is refused with\na plain reason and a retry hint rather than quietly handing the alias to a new agent while\nthe old lifecycle's teardown is still running. Only once the broker footprint is gone, the standing\nauthority is revoked, and the retirement is confirmed does the name free, and `cotal spawn <same-name>`\ngives you a fresh agent cleanly. This is what makes reusing an agent's name safe: the old lifecycle is\nfully torn down before the new one takes the alias. If the auth service is unreachable or the\nstanding-authority revoke fails, the despawn still stops the agent and *holds* the name. **A\nsame-name `cotal spawn` re-drives the whole teardown** and finishes it. Retrying the despawn has no\neffect because the agent is already stopped. The operator copy tells you to recover the stack\n(`cotal supervise`) rather than reusing the name over an unretired predecessor.\n\n**A crash mid-retirement resumes at the next boot.** The retirement's last two steps (recording the\nissuance gate terminal, then the lifecycle head terminal) are separate durable writes, and a crash\nbetween them leaves the gate retired while the head is still `retiring`: an alias that can neither\nmint nor be replaced. The auth service's boot crash-resume decides what it owes across *both*\nobjects (the gate *and* the alias head), so the next boot finishes that tail from the durable\noperation intent: nothing is re-revoked or re-drained, and a completed retirement (its head terminal\nlanded, or a successor already took the alias) is left skipped. A retry of the despawn converges on\nthe same recovery.\n\n**Delegation only narrows (the envelope rule).** A user's grant is their envelope:\neverything under their owner (their CLI, every agent they spawn, every agent those\nspawn) stays within its channel lists and its capability scope. Handing a role to a\nspawned agent needs the matching `role:<r>` capability in the spawner's scope. The whole\ndelegation chain is checked, not just the last link, and re-checked at every bearer\nexchange, so narrowing a user's grant reaches their agents within minutes, and revoking\nthe user revokes everything under them, grandchildren included. A spawn beyond the\nenvelope is refused with the exact widening re-grant to ask the operator for.\n\n**Control ops ride your own login**, gated by ledger scope. `spawn` covers launching,\n`ps`, and stop/attach of the agents under **your own owner**: the owner is the\nadministrative boundary of its own subtree, so you (and your agents) manage what you own\nwithout any extra grant. `admin` is the explicit opt-in for touching **other owners'**\nagents; it is never part of a default grant and never accepted from a manifest.\n\n**Elevated operator surfaces ride the same login** through a short-lived *view*: the\nexchange stamps a server-authored view claim into the bearer, and the callout mints that\nconnection as the matching non-agent profile instead of `agent`. `cotal web` and\n`cotal console` ask for the read-only admin view, `clean history` for the purger,\n`channels set/default` for the channel-writer (all gated on ledger scope `admin`);\n`up -f` deploys over the deployer view, gated on `spawn`, because deploying your own team\nis spawn-grade (the manager still refuses a manifest claiming another owner). Views exist\nonly on a signed-in human exchange (an agent's managed exchange never mints one), are\nauthorized against the fresh ledger row at every connect, and expire with the bearer, so\nnarrowing or revoking a grant bites within minutes here too. On the public exchange face only\n`channel-writer` and `channel-purger` are served; `admin`, `purger`, `deployer`, and\n`manager-service` remain loopback-only.\n\n### Remote manager authority\n\nA registered user remains an ordinary `agent` bearer by default. Running a detached manager\non a remote user-auth mesh needs the closed server-authored **`manager-service`** view, which\nis distinct from every general-purpose profile. The operator grants it only by adding\n`supervise` to that user's actor-ledger scope. `supervise` is deliberately distinct from\n`spawn` and `admin`: spawn controls your agents, admin permits the separate cross-owner\noperations, and neither grants persistent manager registration authority.\n\nOnly a signed-in human may request this view from the loopback/operator exchange. The public\nexchange and every managed-agent secret exchange refuse it. At exchange and each connection,\nthe auth service re-reads the actor row; revoking or removing `supervise` therefore denies the\nnext view exchange and connection. A grant must carry the whole requested row just like every\nother actor update, so re-grant its channel envelope, role, and all wanted scope tokens, not\nonly `supervise`.\n\nThe service is one opaque manager instance for the user's derived owner and a fixed\nserver-selected manager actor. Its authority is limited to that instance's manager\nregistration, contracts, status, endpoint rails, gate and credential family; it cannot read or\nwrite another owner or instance. It never exposes a signer, static provisioner credential, owner\nsecret, raw stream/KV/consumer authority, or a generic credential-mint API. The host creates the\npublic-nkey JWT material through the typed lifecycle-bound protocol: **prepare → activate →\nrenew**, plus a one-shot **retire** phase for one exact managed lifecycle. Each request is replay-safe and idempotent at its lifecycle/instance operation\ncoordinate; the host writes its credential ledger row and finalizes the gate before it releases\nusable material. The retire phase fresh-checks the current manager instance, server-derived serve\nprincipal, serve epoch, same-owner target and lifecycle UID. It returns only a short-lived requester\ncredential pinned to that target. The manager sends it on the existing auth retirement rail with the\noperation id derived from the target lifecycle UID. The terminal rail recomputes it from the\nbroker-pinned target before any durable access. A caller cannot substitute another valid operation\nidentity for the same target, and retries plus auth-service boot recovery finish the same terminal\nbarrier. It never exposes the barrier executor or a general mint surface.\n\nA remote manager can provision only descendants of the same derived owner, and the host\nvalidates that relation and the current manager grant for every provision. It cannot broaden the\nuser's envelope or provision a sibling owner's agent. Renewals are bounded. If login, the\n`supervise` grant, or the host manager authority service is unavailable, the manager reports a\ndegraded state and refuses new agents, restarts, or replacement credentials rather than\nsubstituting local/static authority. Existing live agents remain running only while their own\nvalid authority permits it; recovery requires the host service and a fresh successful renewal.\n\n**User authentication has one path.** On a user-auth space, commands never fall back to\nstatic minting or credless connects: a missing login or a down auth service is one\nsentence naming the exact recovery, and static agent/observer/admin minting is refused\noutright. The refusal is deny-new: a static cred signed before the space flipped stays\nbroker-valid until the signing key is rotated ([security model](security.md)).\n\n## The IdP callout contract\n\nAny OIDC identity provider that issues **EdDSA/Ed25519** JWTs plugs in here directly; a provider that\nissues RS256 or ES256 tokens (many managed OIDC services do) needs a host-side normalization or\nre-issuance adapter first, because the reference bridge pins the token algorithm to EdDSA. The\nreference implementation ships **Better Auth** as a\ndev and test fixture only (it is a `devDependency` of `@cotal-ai/auth`; the only code that imports\nit is the `dev-idp.ts` harness and the smoke tests, never the runtime `src`). The one runtime\ncoupling to an IdP is the `idp.ts` bridge plus the `auth-provider` extension. The bridge core\n(`createIdpBridge`) is IdP-generic for **EdDSA** tokens (issuer, audience, JWKS as configuration).\nThe stock end-to-end flow around it, though, is **Better-Auth-shaped**: `cotalAuthProvider` pins\n`<base>/jwks` and issuer/audience to the IdP origin, and the login client speaks Better Auth's\ndevice-code endpoints (`/device/code`, `/device/token`, `/token`) with an opaque revocable session.\nSo a Better-Auth-shaped EdDSA IdP uses the stock flow directly; **any other production IdP is a\nhosted-composability gap, not a configuration change**. A host integrates it by building its own\nlogin and provider wiring on the low-level primitives (`createIdpBridge`, `createUserTokenIssuer`),\nnot by reusing the stock provider. Note that importing `@cotal-ai/auth` self-registers\n`cotalAuthProvider`, and `resolveAuthProvider()` throws when two providers are registered, so a host\non the registry-resolution path must not also register its own. Whatever the path, never loosen the\nissuer/audience/JWKS pins to force-fit an IdP.\n\nThe bridge (`createIdpBridge`) exchanges a verified IdP token for a Cotal bearer in three steps:\n\n1. **Bearer validation.** Verify the IdP's JWT offline against its **pinned JWKS**, with the token\n algorithm pinned to EdDSA. Keys resolve only through the pinned JWKS: a token carrying embedded\n key material (`jku`/`jwk`/`x5u`/`x5c`) is rejected, so the token can never influence key\n resolution. Issuer and audience are checked, and the minted Cotal bearer is capped to the\n upstream proof's remaining lifetime.\n2. **Owner derivation.** The opaque per-space owner derives deterministically from the JSON-array\n encoding of `[idp issuer, sub]`, namespaced by issuer so no issuer/sub pair can straddle a\n delimiter, and re-login re-lands the same person in the same lanes. The owner-token *format*\n (`u_` followed by 26 base32-lower characters) is normative\n ([SPEC section 2](../SPEC.md#2-identity)). At the contract level the *derivation* from an\n identity is a pluggable edge, but the reference `createIdpBridge` fixes it\n (`deriveOwnerForIdpSubject`) and takes no derivation callback, so what a host configures is the\n IdP, not the derivation. **The encoding is frozen:** changing it, or changing the IdP issuer\n string, re-keys every owner in the space, which is a migration on the order of rotating the space\n secret.\n3. **Actor authorization and mint.** The operator's ledger hook authorizes the `(owner, actor)` pair\n and is the only source of the bearer's `scope`/`parent`; the issuer then mints the Cotal bearer,\n re-asserting every claim shape.\n\nA host wires this with the IdP's own coordinates and nothing from `@cotal-ai/auth` changes:\n\n```ts\nimport { createIdpBridge, pinnedJwksResolver, createUserTokenIssuer } from \"@cotal-ai/auth\";\nconst bridge = createIdpBridge({\n idp: { issuer: idpIssuer, audience, key: pinnedJwksResolver(jwksUri) }, // your production IdP\n space,\n spaceSecret, // identity-plane owner-derivation secret (>=32 bytes), held by the auth service at runtime\n issuer: createUserTokenIssuer({ issuer: cotalIssuer, key: signingKey }), // mints the Cotal bearer\n authorizeActor: (owner, actor) => grantFromLedger(owner, actor), // your ledger, returns an ActorGrant\n});\n```\n\n## Joining\n\nA single **join link** carries server, auth, and space\n([SPEC §10](../SPEC.md#10-connection-and-onboarding)):\n\n```\ncotals://<token>@host:4222/<space>?channel=general # cotals:// = TLS required; cotal:// = TLS not required (downgrade-tolerant)\n```\n\nHumans: `cotal join --link …`. Agents: `COTAL_LINK=… ` in the environment. The connector\nexpands it and auto-joins. Token/user-pass links are the open-mode path; the default\nauthed path threads a minted creds file, and the endpoint adopts the credential's identity\nas its card id. A seat the manager spawned reaches that file through its **launch\nmaterial** rather than through `COTAL_CREDS` in an environment every descendant process\ninherits (see [Configuration](config.md#launch-material)); a session you drive by hand\nstill sets `COTAL_CREDS` itself.\n\n## Honest limitations (v0)\n\n- **The signing key is hot** on the mint/manager box of a static-auth mesh; the \"real\n boundary\" holds given operator-controlled cred distribution. On a per-user-auth mesh\n the data-account signing key is held by the auth service (the callout stage) and by any\n running manager, which loads the trust bundle and self-mints its supervisor cred and\n renewals from it; a copied signing *seed* still stays valid for its identity until the\n signing key is rotated. Rotation remains the revocation lever for trust material.\n- **The two `$SYS` creds renew through rotation.** `membership-observer` and\n `connection-evictor` are signed by the system-account seed, which is never persisted, so no\n running process re-signs them: they carry a 30-day expiry and are renewed by issuing a new\n system account (`cotal down` then `cotal up --rotate-sys`), which leaves the data account,\n every agent cred and the store untouched but does invalidate earlier full backups (they bind to\n the operator JWT and system account they were taken under, so re-run `cotal backup` after). Past that horizon the mesh keeps delivering, but the\n membership feed and live eviction stop; `cotal doctor auth` and the manager warn from the 75%\n point onward.\n- **Static agent creds are long-lived; the machinery's are not.** One-shot command creds\n expire in minutes and the standing daemon creds in 24h with the manager renewing them\n (`cotal doctor auth` is the one diagnosis and repair surface). But a static *agent*\n cred has no TTL yet: `cotal_despawn` cuts a session, not a credential, and a\n compromised agent that copied its creds can reconnect until the signing key is\n rotated. Per-user-auth spaces close this: bearers live minutes, `cotal actor revoke`\n denies the next exchange and the next connect and evicts the principal's live\n connections immediately.\n- **Not non-repudiation.** Authenticity is broker-enforced, not portable proof; it does\n not survive an untrusted relay. Signed envelopes are reserved\n ([SPEC §11](../SPEC.md#11-versioning-and-extensibility)).\n- **Chat metadata leaks in-space.** Content reads are ACL-bounded; stream metadata\n (channel names, per-subject counts) is not yet ([security model](security.md)).\n\n**Denials are loud, never silent.** A publish outside an ACL surfaces as a logged denial\n(\"denied, not absent\") on the endpoint's error path; an over-tight ACL never looks like a\nmissing peer ([run a mesh](run-a-mesh.md)).\n"
45
+ "body": "# Identity\n\n> **Concept** (informative) · **For:** operators and implementers · **Normative:** [SPEC §2](../SPEC.md#2-identity), [§9](../SPEC.md#9-nats--jetstream-security-and-authorization), [§10](../SPEC.md#10-connection-and-onboarding), [Appendix B](../SPEC.md#appendix-b-profile-acls)\n\nWho can do what on a mesh, and how it is enforced. The design goal: the mesh is a **real\nboundary against untrusted peers in a shared space**; an agent can only speak as itself\nand only where its declared permissions allow, enforced by the broker, not by agent\ngoodwill. What that boundary does and does not protect is the\n[security model](security.md); the exact ACLs are\n[SPEC Appendix B](../SPEC.md#appendix-b-profile-acls).\n\n## On by default\n\n`cotal up` provisions a JWT-authed space; `cotal up --open` runs an unauthenticated dev\nmesh instead. Both bind loopback by default. `--host 0.0.0.0` widens the bind\nindependently, so \"network-reachable\" never silently means \"unauthenticated\". Open mode\nis for quick local experiments and sits outside every security claim\n([SPEC §9](../SPEC.md#9-nats--jetstream-security-and-authorization)).\n\n## Shared identity\n\nAn agent's wire identity is a **principal**: an `owner.actor` pair, where the owner is\nthe account (a human, or an organization) the agent acts on behalf of, and the actor is\nthe agent's own handle under that owner ([SPEC §2](../SPEC.md#2-identity)). The same pair\nis the card id, the sender tokens in every subject it publishes, the presence key, and\nits durable-consumer names. On an open dev mesh the owner is the literal `local`; on a\nper-user-auth mesh it is a derived token (`u_` plus 26 characters, so no PII rides the\nwire). The connection still authenticates with an **nkey**, generated locally (the signer\nonly ever sees the public half), but the nkey is the transport credential, not the\nidentity: it scopes only the per-connection reply inbox.\n\n**The sender is encoded in the subject.** Every publish carries the sender's owner and\nactor in positions the broker's permissions pin to that connection, so an agent *cannot*\nemit as anyone else: not as another owner, and not as a sibling actor under its own\nowner. Receivers verify the payload's `from.id` against the subject sender and reject\nmismatches; sender authenticity is broker-enforced end to end\n([SPEC §3](../SPEC.md#3-subject-layout), [§5](../SPEC.md#5-envelopes)).\n\n**Account = space, user = agent.** A space is one NATS account, a server-enforced\nisolation boundary. An operator signs the account; an account **signing key** mints\nper-agent user JWTs.\n\n## Provisioner\n\nThe **provisioner** is whoever holds the account signing key. It mints profile-scoped\ncredentials and pre-creates the durables agents may only *bind* (their DM inbox, their\nrole's task queue). The manager hosts it today, but nothing is manager-special about it;\nprivilege attaches to the signer, and a space can run without a manager.\n`cotal mint <name> --profile <agent|observer|admin>` is the out-of-band path; spawn calls\nthe same library ([CLI](cli.md)). Minting static creds is a **static-auth** surface: a\nper-user-auth space refuses it, because agents there join under a logged-in user, never\nvia a handed-out file (see *Per-user auth* below).\n\nAgent-profile minting resolves one mesh root for the persona ACL, account signer and default\ncredential storage. If the current folder holds trust for a different space or account, mint\nrefuses and names both roots. It never signs one root's persona policy with another root's authority.\n\n## Profiles\n\nEvery credential is a profile: an explicit allow-list built from the same\nsubject/stream/durable builders as the wire layout, so ACLs cannot drift from it. The\nnormative shapes are [SPEC Appendix B](../SPEC.md#appendix-b-profile-acls); in brief:\n\n| Profile | Is |\n|---|---|\n| **agent** | The ordinary peer: publishes as itself to its declared channels, reads within its read ACL + its own DM/task inboxes. Its read-only presence and channel-registry watches may create, inspect, and delete only their own client-managed ordered consumers; those cleanup grants cannot delete KV records or streams. |\n| **observer** | Read-only chat + presence; DMs invisible. What `cotal console` runs. |\n| **admin** | Elevated *read-only* god-view: sees DMs and anycast live, still writes nothing. A deliberate opt-in (`cotal web`). |\n| operator-side | Narrow single-purpose creds for the machinery (supervising, provisioning, teardown, delivery); the reference implementation splits these so no one connection can read every DM *and* delete every stream ([security model](security.md)). |\n| **run-driver** | One workflow run and takeover attempt: its journal subject, replay durable and run-owned record writes. Store reads and effects go through the host. |\n| **run-mediator** | The trusted hosting process's separate connection for workflow effects and leader reads. It exposes journal-checked, run-bound operations and never hands its credential to the driver. |\n| **run-operator** | One served run read, or one half of an answer, minted per call: a read holds the records walk, one run's replay and the admission read; the answering half is minted for one checkpoint token and holds that pause's answer record and settle alone. |\n| **issuer** | One issuance window: the party holding the space signer mints it for a few minutes to stage and release a credential's evidence, retire the issuances a lifecycle terminal leaves behind, or resolve the evidence a request rides. |\n| **run-admitter** | One run's admission record or revocation marker, minted per run for a minute: two exact keys in the admission store and nothing else. |\n\n**An agent's channel scope is three verbs**: `subscribe` (reads at boot),\n`allowSubscribe` (read ACL), `allowPublish` (post ACL, default-deny), declared in its\n[agent file](agent-files.md) or [manifest](manifest.md), minted into its cred. One card\nwith the recipes: [Channels & permissions](channels-and-permissions.md).\n\n**DM confidentiality** holds against peers by construction: deliveries ride per-identity\ninbox prefixes, and the DM/task consumers are provisioner-pre-created and bind-only, so an\nagent cannot create a consumer filtered to someone else's inbox\n([SPEC §9](../SPEC.md#9-nats--jetstream-security-and-authorization) items 1–5).\n\n## Issued authority\n\nA credential says who is calling. It does not, by itself, say what the caller was granted, and a\nhost that acts on a caller's behalf (a workflow run, [workflows](workflows.md)) needs that from\nthe issuer, not from a ledger that may have changed since. So a static agent credential is an\n**issuance** ([SPEC §13.15](../SPEC.md#1315-issued-authority)): before the material is handed\nout, the issuer records the credential's final permission ceiling, as evidence keyed by a fresh\n**generation**, in `cotal_issued_<space>`. That store is append-only at the broker and the\nevidence is read as the first message on its key, so nothing written later on the same key, by\nanyone, changes what a resolver sees. The credential's endpoint rows then ride a versioned\nrail, `cotal.<space>.ep.v1.…`, with that generation pinned beside the caller triple, so the\nbroker binds every request to the ceiling the issuer accepted. The legacy `ep.` rail and the\n`ep.v1.` rail are disjoint subject spaces; a credential holds rows on one of them, and every\nendpoint serves both.\n\nA connected client learns its generation by reading one row in `cotal_accepted_<space>` under a\ntoken the launching party chose at mint time, through a per-key read grant its own ceiling carries. It never trusts\nwhat the file says. A renewal keeps the generation only while the ceiling is byte-identical to the\nevidence; a changed scope is a fresh issuance on a fresh generation, adopted by a new connection.\nA static agent's evidence names its credential ledger family as the source it depends on, and the\nlifecycle terminal that retires that family retires its issuances with it.\n\nThe manager, `cotal spawn`, and the CLI's control-caller instruments all mint through an `issuer`\nsession. The deployer instrument does not: it is the one instrument with no default expiry, and an\nissuance with no lifecycle gate must carry one, so a deploy rides the legacy rail under its own\nlifecycle uid. Only workflow `run-start` requires the binding today: a request for it on the legacy\nrail is refused with `permission-denied` and the detail `ai.cotal.ep.unbound-caller-authority`\nnaming the caller. Every other command serves both rails.\n\n## Declared capabilities\n\nControl-plane power is a **declared capability**, not a default. An agent file carrying\n`capabilities: [spawn]` gets the privileged control subject minted into its cred: spawn,\nplus stop/despawn of its *own* children, plus persona definition. Without it, an agent can\nonly self-despawn and pull or yield the run turns addressed to it. `capabilities: [run]` mints\nthe manager's workflow-run commands (start, resume, answer, status, list) together with the spawn\nset, since a program the agent starts may spawn; the manager drives the run under a per-run\n`run-driver` credential of its own, never the caller's. The tool surface mirrors the grant:\n`cotal_spawn` / `cotal_persona` / `cotal_personas` are injected only for `spawn`, and `cotal_run`\nonly for `run` ([agent files](agent-files.md)). Destructive\noperator ops (history purge, cross-agent stop) live on a third tier no agent credential\nreaches. Persona redefinition separates content from policy; the write path takes only\n`model`/`persona`, so a peer cannot grant itself a capability by redefining a file.\n\n## Per-user authentication\n\n`cotal up --user-auth --idp <auth base URL>` (or manifest `broker.auth: \"user\"`) puts a\n**human identity plane** above the per-agent one: people sign in to an external IdP once,\nand every connect is authorized live against the operator's **actor ledger**. No creds\nfiles to hand out, and revoking a grant actually bites.\n\n**The flow.** Each person runs `cotal login --idp <url>` once per machine. After that,\nany command works: cached IdP session → fresh IdP proof per connect (so IdP-side\nrevocation bites here too) → the configured exchange turns it into a short-lived Cotal bearer →\nthe broker's **auth callout** checks the bearer and the ledger at connect time and mints\na scoped credential on the spot. Every bearer also names a **root credential** row in the\nspace's credential ledger, proved live at each connect, so revoking that one credential\nbites at the very next connect. The operator grants access with\n`cotal actor grant <actor> --sub <their id>`; a bare grant is the full envelope (all\nchannels, may spawn), and `--allow-subscribe` / `--allow-publish` / `--scope` narrow it.\nNo ledger row, no access; there is no allow-by-default.\n\n**Space catalogs.** A successful authenticated `GET <idp>/token` may advertise one catalog with:\n\n```http\nLink: <https://idp.example/spaces>; rel=\"https://cotal.ai/relations/space-catalog\"\n```\n\nThe target must use HTTPS and the same origin as the normalized IdP URL. Loopback IP literals may\nuse HTTP for local development. A missing, foreign-origin, or insecure link records that this account\nhas no catalog. The client never guesses a path.\n\nThe catalog request carries the opaque cached session as its bearer and returns a complete snapshot:\n\n```json\n{\n \"v\": 1,\n \"account\": { \"idpUrl\": \"https://idp.example/api/auth\", \"issuer\": \"https://idp.example\", \"sub\": \"user-id\" },\n \"spaces\": [\n { \"id\": \"space-id\", \"slug\": \"shared_project\", \"name\": \"Shared project\", \"kind\": \"hosted\", \"role\": \"owner\", \"registration\": {} }\n ]\n}\n```\n\nThe client checks every `registration` with the same `checkUserBundle` validator used by `cotal\nmeshes add`. One invalid entry refuses the whole candidate snapshot. Conditional refresh uses the\ncatalog's `ETag`; a transport error, non-success response, or invalid candidate leaves the prior\nsnapshot intact and reports the failure. Only a 401 response for the saved session recommends signing\nin again. Transport errors and server failures report that account's refresh error without discarding\nthe session. The registry is reconciled under the provider's catalog lock, and a snapshot whose\nreconciliation was interrupted is reconciled again by the next refresh before it counts as fresh or\nnot modified.\n\nThe user-auth registration document may include one closed policy object:\n\n```json\n{ \"policy\": { \"events\": \"required\" } }\n```\n\nNo other key under `policy` and no other value for `policy.events` is accepted. The registry preserves\nthis field for manual, discovered, and enrollment-created entries. A pre-policy manual entry is\nrefreshed from its own pinned exchange origin by the command that consumes the policy, a spawn, a\njoin, or a manager start, after that command's own local refusals; the returned space, broker,\ntransport, IdP, issuer, audience, and exchange pins must all match before only the policy is added.\nA failed expired refresh refuses that operation. Read-only commands such as `status` and `meshes`\nnever refresh. The five-second warm window makes no request.\n\nEvery registration is also bound to the proved account. Its IdP URL must match the account and its\nissuer must match the exact JWT `iss` pin. Its exchange, provisioning, and manager-authority\nendpoints must be same-origin with that IdP. One foreign pin or endpoint refuses the whole\ncandidate snapshot.\n\nThe `slug` is the space identity resolved by `--space`, `use`, registry roots, and collisions. The\n`name` is a display label only.\n\nDiscovered registry entries are owned by the normalized IdP origin plus the proved `sub`. That key is\nstored as an opaque digest, so accounts on one machine never union their spaces and the registry does\nnot persist the subject. A manual or locally started record with the same name is never overwritten.\nLogout removes only the entries owned by the account whose session was revoked. Local teardown,\ncleanup, and liveness pruning do not remove discovered entries.\n\nAn account that previously advertised no catalog is checked again by explicit `cotal sync`. The\nordinary lazy path checks again after its five-second capability window, so an IdP can enable the\nLink for an existing login without making the person sign in again.\n\n**One auth service per space** hosts both halves: the NATS auth callout and the token\nexchange. Its default HTTP listener remains loopback-only and requires the per-start capability\nstored in the owner-only `auth-service.json` file. An operator may add a second listener with\n`cotal up --user-auth ... --exchange-public-port <port> --exchange-public-url https://auth.example`.\nThat listener still binds `127.0.0.1`; put a reverse proxy in front of it and terminate TLS there.\nIn-process TLS is deliberately not another deployment mode: it would duplicate certificate renewal\nand fork proxy-based deployments.\n\nThe public listener has a closed surface: `GET /health`, `GET /jwks`, `POST /exchange`, and\n`GET /.well-known/cotal-mesh`; every other path is 404. It does **not** require the loopback\ncapability. That capability proves same-uid access to a 0600 local file and has no remote meaning;\non the public face the credential is the proof. A human presents an EdDSA IdP JWT checked against\nthe pinned JWKS, issuer, and audience. An agent presents its spawn-time actor token, whose hash must\nmatch a fresh managed-ledger row. The public face mints only two elevated views, both still\ngated on ledger scope `admin`: `channel-writer` (`cotal channels set/default`) and\n`channel-purger` (the dashboard's per-click channel delete). God-view (`admin`), space-history\n`purger`, `deployer`, and `manager-service` stay loopback-only. A managed-agent secret exchange\nnever mints a view on either face. This is not full remote channel management: `cotal web` still\nmints the read-only admin view at startup, so a remote dashboard that needs that god-view still\nfails even when a later delete would mint `channel-purger`.\n\nThe well-known response contains the IdP pins and the actual deny-all sentinel credential remote\nagents need before the bearer-driven auth callout. The pins ride a `userAuth` arm that names the\nauth provider, and that name is the same one the local arm registers under. A document naming a\ndifferent provider than the one serving it would register an entry nothing can resolve, so both read\none constant. Treat it as bootstrap material: the sentinel cannot publish or subscribe, but\nconsumers must still take the bundle only from the intended HTTPS origin and must verify TLS.\n`--exchange-trusted-proxy` opts into peer attribution by the **last** `X-Forwarded-For` hop; use it\nonly when the listener is reachable solely through a proxy you control.\nWithout it, forwarded headers are ignored and the socket address is the peer key. Public failure\nbuckets are per-source and separate from loopback exchange budgets. The in-process LRU retains at\nmost 1024 peer buckets: that bounds memory and isolates ordinary sources, but an attacker cycling\nmore than 1024 trusted-proxy last hops can evict earlier 429 state. It is not a mint bypass; a valid\ncredential is still required, so use upstream reverse-proxy rate limiting when that throttle-escape\nmatters to the deployment.\n\n### Enrollment redeem\n\nA remote owner may pre-mint a one-time enrollment for a seat that has no browser, TTY, or cached\nIdP login. The enrollment is a secret-bearing URL. The client performs one request:\n\n```http\nGET <enrollment URL>\n```\n\nIt sends no `Authorization` header and no request body. The URL must be HTTPS, except for plain HTTP\nto a loopback IP literal. The client redeems only an enrollment URL that is already in canonical\nform and contains none of `\\ @ ? #`. That is checked on the raw string before parsing, so every\nrewrite a URL parser would perform, backslash folding, userinfo erasure, scheme or host case\nfolding, default-port removal, dot-segment resolution, and short-host canonicalization, is a refusal\nrather than a redeem of a URL the owner never minted. Redirects are refused. The client never\nretries because a successful claim deletes the server-side token row. The token expires five minutes\nafter mint.\n\nSuccess is `200` with this JSON object:\n\n```text\nspace\nbrokerAccess { kind, ... }\nowner\nactor\nlifecycleUid\nactorToken\nsentinelCreds\nauthServiceUrl\nidp { url, issuer, audience }\nsubscribe[]\nallowSubscribe[]\nallowPublish[]\n```\n\nThe grant arrays are informational; the broker row remains authoritative. A stock-dialable\ndeployment also includes `server`, `tlsRequired`, `userAuth`, and optional `policy`, forming the same user-bundle\nsuperset that `cotal meshes add --user-auth-file` accepts. That lets a bare seat register the mesh\nfrom the enrollment response before launch. For `brokerAccess.kind: \"direct\"`, the stock `server`\nmust equal `brokerAccess.url` byte for byte or the client refuses the bundle before registration. A\ntunnel kind carries no dial address, so its `brokerAccess` is not compared to the operator-asserted\nstock `server` face.\n\nUnknown, expired, revoked, and already-used enrollments are intentionally indistinguishable. They\nall return `404 {\"error\":\"unknown, expired, or already-used enrollment\"}`. The client reports only\n`enrollment refused: unknown, expired, or already-used; ask the owner for a fresh one`. It does not\nguess which case occurred.\n\nAfter redeem, the seat stores only the normal remote user-mesh and agent material. The actor token\nis exchanged at `authServiceUrl` through the existing `agent-bearer --exchange-url` path. The\nenrollment URL is not logged, persisted, or forwarded into any child process, including the bearer\npreflight and harness.\n\nThe service starts with the broker, is torn down by `cotal down`, and holds the\ndata-account signing key for the callout (a running manager is the other standing holder, for\nthe creds it mints); the operator seed never enters it. It also owns the space's two authority\nstores (lifecycle records and the credential ledger), provisions them at boot, and refuses\nconnects it cannot credential-check against them; there is no fallback path. If it\ndies while the broker lives, re-running `cotal up` heals it, and a boot whose auth\nservice never became ready exits non-zero, so automation never reads a dead identity\nplane as success. Changing any public-listener flag requires `cotal down` followed by `cotal up`\nwith the new values; a refresh adopts an already-running auth service rather than silently replacing\nits listener policy. \"One per space\" is enforced, not assumed (SPEC §13.13): at boot the\nservice takes a broker-backed ownership claim, so a second same-space auth process refuses\nwith instructions instead of silently splitting the plane, and a crashed one's claim is\nreclaimed only once the broker confirms its connections are gone. That verdict is trusted only\non a standalone broker (a clustered one refuses the reclaim, since a partitioned member\ncould still hold them). If the claim's connections die mid-run, the service downs itself\nloudly instead of serving from a half-dead plane.\n\n**Your agents are yours.** `cotal spawn` on a user mesh grants a managed actor under the\n*spawning operator's* owner and launches the agent with a bearer command instead of a\ncreds file. The agent exchanges its spawn-time secret for short bearers (five minutes or\nless) and refreshes ahead of each expiry. Rows are runtime grants: every start rotates\nthe secret, every stop or despawn revokes the row, so a non-running agent holds no\nstanding authority. Manifest deploys (`up -f`) stamp the logged-in owner into the launch,\nso those agents are yours too.\n\n**Despawn tears the lifecycle down, then frees the name.** When you despawn an agent, the manager\ndrives the *full* teardown of that lifecycle: it shreds the local credential files, revokes the\nagent's standing mint authority (its ledger row, so a copied token can no longer mint a fresh\ncredential), deletes its broker footprint (the lifecycle-keyed durables + read-ACL row), and asks\nthe auth service to *retire* the lifecycle (settle in-flight work, evict the departed credentials,\nrecord it retired). The name is held *reserved pending retirement* until **all** of that completes,\nthe broker-footprint cleanup, the standing-authority revoke, **and** the lifecycle retirement, not the\nretirement alone, so a same-name respawn in the gap is refused with\na plain reason and a retry hint rather than quietly handing the alias to a new agent while\nthe old lifecycle's teardown is still running. Only once the broker footprint is gone, the standing\nauthority is revoked, and the retirement is confirmed does the name free, and `cotal spawn <same-name>`\ngives you a fresh agent cleanly. This is what makes reusing an agent's name safe: the old lifecycle is\nfully torn down before the new one takes the alias. If the auth service is unreachable or the\nstanding-authority revoke fails, the despawn still stops the agent and *holds* the name. **A\nsame-name `cotal spawn` re-drives the whole teardown** and finishes it. Retrying the despawn has no\neffect because the agent is already stopped. The operator copy tells you to recover the stack\n(`cotal supervise`) rather than reusing the name over an unretired predecessor.\n\n**A crash mid-retirement resumes at the next boot.** The retirement's last two steps (recording the\nissuance gate terminal, then the lifecycle head terminal) are separate durable writes, and a crash\nbetween them leaves the gate retired while the head is still `retiring`: an alias that can neither\nmint nor be replaced. The auth service's boot crash-resume decides what it owes across *both*\nobjects (the gate *and* the alias head), so the next boot finishes that tail from the durable\noperation intent: nothing is re-revoked or re-drained, and a completed retirement (its head terminal\nlanded, or a successor already took the alias) is left skipped. A retry of the despawn converges on\nthe same recovery.\n\n**Delegation only narrows (the envelope rule).** A user's grant is their envelope:\neverything under their owner (their CLI, every agent they spawn, every agent those\nspawn) stays within its channel lists and its capability scope. Handing a role to a\nspawned agent needs the matching `role:<r>` capability in the spawner's scope. The whole\ndelegation chain is checked, not just the last link, and re-checked at every bearer\nexchange, so narrowing a user's grant reaches their agents within minutes, and revoking\nthe user revokes everything under them, grandchildren included. A spawn beyond the\nenvelope is refused with the exact widening re-grant to ask the operator for.\n\n**Control ops ride your own login**, gated by ledger scope. `spawn` covers launching,\n`ps`, and stop/attach of the agents under **your own owner**: the owner is the\nadministrative boundary of its own subtree, so you (and your agents) manage what you own\nwithout any extra grant. `admin` is the explicit opt-in for touching **other owners'**\nagents; it is never part of a default grant and never accepted from a manifest.\n\n**Elevated operator surfaces ride the same login** through a short-lived *view*: the\nexchange stamps a server-authored view claim into the bearer, and the callout mints that\nconnection as the matching non-agent profile instead of `agent`. `cotal web` and\n`cotal console` ask for the read-only admin view, `clean history` for the purger,\n`channels set/default` for the channel-writer (all gated on ledger scope `admin`);\n`up -f` deploys over the deployer view, gated on `spawn`, because deploying your own team\nis spawn-grade (the manager still refuses a manifest claiming another owner). Views exist\nonly on a signed-in human exchange (an agent's managed exchange never mints one), are\nauthorized against the fresh ledger row at every connect, and expire with the bearer, so\nnarrowing or revoking a grant bites within minutes here too. On the public exchange face only\n`channel-writer` and `channel-purger` are served; `admin`, `purger`, `deployer`, and\n`manager-service` remain loopback-only.\n\n### Remote manager authority\n\nA registered user remains an ordinary `agent` bearer by default. Running a detached manager\non a remote user-auth mesh needs the closed server-authored **`manager-service`** view, which\nis distinct from every general-purpose profile. The operator grants it only by adding\n`supervise` to that user's actor-ledger scope. `supervise` is deliberately distinct from\n`spawn` and `admin`: spawn controls your agents, admin permits the separate cross-owner\noperations, and neither grants persistent manager registration authority.\n\nOnly a signed-in human may request this view from the loopback/operator exchange. The public\nexchange and every managed-agent secret exchange refuse it. At exchange and each connection,\nthe auth service re-reads the actor row; revoking or removing `supervise` therefore denies the\nnext view exchange and connection. A grant must carry the whole requested row just like every\nother actor update, so re-grant its channel envelope, role, and all wanted scope tokens, not\nonly `supervise`.\n\nThe service is one opaque manager instance for the user's derived owner and a fixed\nserver-selected manager actor. Its authority is limited to that instance's manager\nregistration, contracts, status, endpoint rails, gate and credential family; it cannot read or\nwrite another owner or instance. It never exposes a signer, static provisioner credential, owner\nsecret, raw stream/KV/consumer authority, or a generic credential-mint API. The host creates the\npublic-nkey JWT material through the typed lifecycle-bound protocol: **prepare → activate →\nrenew**, plus a one-shot **retire** phase for one exact managed lifecycle. Each request is replay-safe and idempotent at its lifecycle/instance operation\ncoordinate; the host writes its credential ledger row and finalizes the gate before it releases\nusable material. The retire phase fresh-checks the current manager instance, server-derived serve\nprincipal, serve epoch, same-owner target and lifecycle UID. It returns only a short-lived requester\ncredential pinned to that target. The manager sends it on the existing auth retirement rail with the\noperation id derived from the target lifecycle UID. The terminal rail recomputes it from the\nbroker-pinned target before any durable access. A caller cannot substitute another valid operation\nidentity for the same target, and retries plus auth-service boot recovery finish the same terminal\nbarrier. It never exposes the barrier executor or a general mint surface.\n\nA remote manager can provision only descendants of the same derived owner, and the host\nvalidates that relation and the current manager grant for every provision. It cannot broaden the\nuser's envelope or provision a sibling owner's agent. Renewals are bounded. If login, the\n`supervise` grant, or the host manager authority service is unavailable, the manager reports a\ndegraded state and refuses new agents, restarts, or replacement credentials rather than\nsubstituting local/static authority. Existing live agents remain running only while their own\nvalid authority permits it; recovery requires the host service and a fresh successful renewal.\n\n**User authentication has one path.** On a user-auth space, commands never fall back to\nstatic minting or credless connects: a missing login or a down auth service is one\nsentence naming the exact recovery, and static agent/observer/admin minting is refused\noutright. The refusal is deny-new: a static cred signed before the space flipped stays\nbroker-valid until the signing key is rotated ([security model](security.md)).\n\n## The IdP callout contract\n\nAny OIDC identity provider that issues **EdDSA/Ed25519** JWTs plugs in here directly; a provider that\nissues RS256 or ES256 tokens (many managed OIDC services do) needs a host-side normalization or\nre-issuance adapter first, because the reference bridge pins the token algorithm to EdDSA. The\nreference implementation ships **Better Auth** as a\ndev and test fixture only (it is a `devDependency` of `@cotal-ai/auth`; the only code that imports\nit is the `dev-idp.ts` harness and the smoke tests, never the runtime `src`). The one runtime\ncoupling to an IdP is the `idp.ts` bridge plus the `auth-provider` extension. The bridge core\n(`createIdpBridge`) is IdP-generic for **EdDSA** tokens (issuer, audience, JWKS as configuration).\nThe stock end-to-end flow around it, though, is **Better-Auth-shaped**: `cotalAuthProvider` pins\n`<base>/jwks` and issuer/audience to the IdP origin, and the login client speaks Better Auth's\ndevice-code endpoints (`/device/code`, `/device/token`, `/token`) with an opaque revocable session.\nSo a Better-Auth-shaped EdDSA IdP uses the stock flow directly; **any other production IdP is a\nhosted-composability gap, not a configuration change**. A host integrates it by building its own\nlogin and provider wiring on the low-level primitives (`createIdpBridge`, `createUserTokenIssuer`),\nnot by reusing the stock provider. Note that importing `@cotal-ai/auth` self-registers\n`cotalAuthProvider`, and `resolveAuthProvider()` throws when two providers are registered, so a host\non the registry-resolution path must not also register its own. Whatever the path, never loosen the\nissuer/audience/JWKS pins to force-fit an IdP.\n\nThe bridge (`createIdpBridge`) exchanges a verified IdP token for a Cotal bearer in three steps:\n\n1. **Bearer validation.** Verify the IdP's JWT offline against its **pinned JWKS**, with the token\n algorithm pinned to EdDSA. Keys resolve only through the pinned JWKS: a token carrying embedded\n key material (`jku`/`jwk`/`x5u`/`x5c`) is rejected, so the token can never influence key\n resolution. Issuer and audience are checked, and the minted Cotal bearer is capped to the\n upstream proof's remaining lifetime.\n2. **Owner derivation.** The opaque per-space owner derives deterministically from the JSON-array\n encoding of `[idp issuer, sub]`, namespaced by issuer so no issuer/sub pair can straddle a\n delimiter, and re-login re-lands the same person in the same lanes. The owner-token *format*\n (`u_` followed by 26 base32-lower characters) is normative\n ([SPEC section 2](../SPEC.md#2-identity)). At the contract level the *derivation* from an\n identity is a pluggable edge, but the reference `createIdpBridge` fixes it\n (`deriveOwnerForIdpSubject`) and takes no derivation callback, so what a host configures is the\n IdP, not the derivation. **The encoding is frozen:** changing it, or changing the IdP issuer\n string, re-keys every owner in the space, which is a migration on the order of rotating the space\n secret.\n3. **Actor authorization and mint.** The operator's ledger hook authorizes the `(owner, actor)` pair\n and is the only source of the bearer's `scope`/`parent`; the issuer then mints the Cotal bearer,\n re-asserting every claim shape.\n\nA host wires this with the IdP's own coordinates and nothing from `@cotal-ai/auth` changes:\n\n```ts\nimport { createIdpBridge, pinnedJwksResolver, createUserTokenIssuer } from \"@cotal-ai/auth\";\nconst bridge = createIdpBridge({\n idp: { issuer: idpIssuer, audience, key: pinnedJwksResolver(jwksUri) }, // your production IdP\n space,\n spaceSecret, // identity-plane owner-derivation secret (>=32 bytes), held by the auth service at runtime\n issuer: createUserTokenIssuer({ issuer: cotalIssuer, key: signingKey }), // mints the Cotal bearer\n authorizeActor: (owner, actor) => grantFromLedger(owner, actor), // your ledger, returns an ActorGrant\n});\n```\n\n## Joining\n\nA single **join link** carries server, auth, and space\n([SPEC §10](../SPEC.md#10-connection-and-onboarding)):\n\n```\ncotals://<token>@host:4222/<space>?channel=general # cotals:// = TLS required; cotal:// = TLS not required (downgrade-tolerant)\n```\n\nHumans: `cotal join --link …`. Agents: `COTAL_LINK=… ` in the environment. The connector\nexpands it and auto-joins. Token/user-pass links are the open-mode path; the default\nauthed path threads a minted creds file, and the endpoint adopts the credential's identity\nas its card id. A seat the manager spawned reaches that file through its **launch\nmaterial** rather than through `COTAL_CREDS` in an environment every descendant process\ninherits (see [Configuration](config.md#launch-material)); a session you drive by hand\nstill sets `COTAL_CREDS` itself.\n\n## Honest limitations (v0)\n\n- **The signing key is hot** on the mint/manager box of a static-auth mesh; the \"real\n boundary\" holds given operator-controlled cred distribution. On a per-user-auth mesh\n the data-account signing key is held by the auth service (the callout stage) and by any\n running manager, which loads the trust bundle and self-mints its supervisor cred and\n renewals from it; a copied signing *seed* still stays valid for its identity until the\n signing key is rotated. Rotation remains the revocation lever for trust material.\n- **The two `$SYS` creds renew through rotation.** `membership-observer` and\n `connection-evictor` are signed by the system-account seed, which is never persisted, so no\n running process re-signs them: they carry a 30-day expiry and are renewed by issuing a new\n system account (`cotal down` then `cotal up --rotate-sys`), which leaves the data account,\n every agent cred and the store untouched but does invalidate earlier full backups (they bind to\n the operator JWT and system account they were taken under, so re-run `cotal backup` after). Past that horizon the mesh keeps delivering, but the\n membership feed and live eviction stop; `cotal doctor auth` and the manager warn from the 75%\n point onward.\n- **Static agent creds are long-lived; the machinery's are not.** One-shot command creds\n expire in minutes and the standing daemon creds in 24h with the manager renewing them\n (`cotal doctor auth` is the one diagnosis and repair surface). But a static *agent*\n cred has no TTL yet: `cotal_despawn` cuts a session, not a credential, and a\n compromised agent that copied its creds can reconnect until the signing key is\n rotated. Per-user-auth spaces close this: bearers live minutes, `cotal actor revoke`\n denies the next exchange and the next connect and evicts the principal's live\n connections immediately.\n- **Not non-repudiation.** Authenticity is broker-enforced, not portable proof; it does\n not survive an untrusted relay. Signed envelopes are reserved\n ([SPEC §11](../SPEC.md#11-versioning-and-extensibility)).\n- **Chat metadata leaks in-space.** Content reads are ACL-bounded; stream metadata\n (channel names, per-subject counts) is not yet ([security model](security.md)).\n\n**Denials are loud, never silent.** A publish outside an ACL surfaces as a logged denial\n(\"denied, not absent\") on the endpoint's error path; an over-tight ACL never looks like a\nmissing peer ([run a mesh](run-a-mesh.md)).\n"
46
46
  },
47
47
  {
48
48
  "slug": "agent-files",
@@ -70,7 +70,7 @@ export const DOCS_BUNDLE = {
70
70
  "title": "`cotal` CLI reference",
71
71
  "kind": "Reference: describes the TypeScript reference implementation (the `cotal` CLI), not the wire contract.",
72
72
  "summary": "cotal is the operator command line for the reference implementation: bring a mesh up, mint identities, launch agents, watch what they do, and tear it all down.",
73
- "body": "# `cotal` CLI reference\n\n> **Reference**: describes the TypeScript reference implementation (the `cotal` CLI), not the wire contract. · **For:** operators · **Wire contract:** [SPEC](../SPEC.md)\n\n`cotal` is the operator command line for the reference implementation: bring a mesh up, mint\nidentities, launch agents, watch what they do, and tear it all down. It is a thin client over the\nwire contract: the normative subjects and schemas live in the [SPEC](../SPEC.md); this page is\nlookup material for the commands, not a walkthrough; if you are new, start with\n[Getting started](getting-started.md).\n\n## Running it\n\n```bash\nnpm install -g cotal-ai # puts `cotal` on your PATH (needs Node 22+)\ncotal --help # every command, grouped\ncotal --version # cotal-ai version + each installed extension's (also `cotal -v`)\ncotal <command> --help # one command's flags and usage\n```\n\n`npx cotal-ai <command>` runs it without a global install; in a dev clone, `pnpm cotal <command>`\nruns it through `tsx` with no build step. Bare `cotal` prints help. Every command generates its own\n`--help`, usage, and shell completion from its declared flags.\n\nCommands come from the surfaces the binary composes: the base mesh CLI, the manager\n(`supervise`), and the delivery daemon (`deliver`), plus any operator-installed extensions.\n`cotal ext add <npm-package>` installs any registry providers a package contributes: commands,\nruntimes, and local process lifecycle descriptors. The `web` dashboard and optional manager\nruntimes ship this way.\n\n## Commands\n\n| Area | Command | Purpose |\n|---|---|---|\n| Set up & lifecycle | [`setup`](#setup) | Guided, configure-only setup (installs, seeds personas; launches nothing) |\n| Set up & lifecycle | [`update`](#update) | Reconcile first-party extensions and check or opt into a coherent CLI upgrade |\n| Set up & lifecycle | [`up`](#up) | Start a local mesh (nats-server + JetStream), or boot a whole manifest with `-f` |\n| Set up & lifecycle | [`down`](#down) | Stop the whole stack, selected registered components, or a manifest deploy |\n| Set up & lifecycle | [`backup`](#backups) | Create an offline full-space or registry-only artifact from a preserved cut |\n| Set up & lifecycle | [`clean`](#clean) | Configurable cleanup: purge history (live), or wipe the local store / identity (stopped) |\n| Set up & lifecycle | [`meshes`](#mesh-registry) | List the running meshes on this machine |\n| Set up & lifecycle | [`sync`](#mesh-registry) | Refresh the signed-in account's advertised spaces |\n| Set up & lifecycle | [`use`](#mesh-registry) | Set the default mesh a bare `cotal spawn` joins |\n| Set up & lifecycle | [`status`](#mesh-registry) | Read-only diagnostics for setup, processes, and the selected mesh |\n| Agents & personas | [`spawn`](#spawn) | Launch an agent from a persona (foreground, or `--detach` via the manager) |\n| Agents & personas | [`models`](#models) | List connector model catalogs and variants from the manager |\n| Agents & personas | [`ps`](#managed-seats) | List managed agents and their mesh status |\n| Agents & personas | [`stop`](#managed-seats) | Ask the manager to stop a managed agent |\n| Agents & personas | [`attach`](#managed-seats) | Stream and drive a managed agent's terminal (pty runtime) |\n| Agents & personas | [`input`](#input) | Type one line into a managed agent's terminal without attaching |\n| Agents & personas | [`personas`](#personas) | List, show, edit, create, or remove local personas |\n| Agents & personas | [`supervise`](#supervise) | Run a manager daemon (the agent supervisor / control plane) |\n| Agents & personas | [`service`](#service) | Run the manager as a user service (survives logout and reboot) |\n| Agents & personas | [`runtimes`](#runtimes) | List the agent runtimes the manager can spawn through and whether each is reachable |\n| Agents & personas | [`reconcile-gate`](#reconcile-gate) | Unfreeze an issuance gate left frozen by a crashed restart when the successor cannot boot-heal it (holder gone, complete CONNZ sweep) |\n| Messaging & watching | [`endpoints`](#endpoints) | List every endpoint in the live presence roster, including infrastructure |\n| Messaging & watching | [`describe` / `invoke`](#endpoint-control) | Resolve a v0.4 service's command surface off the wire; invoke one command by name |\n| Messaging & watching | [`send`](#send) | Send one message, then exit: DM a peer, post a channel, or ask a role |\n| Messaging & watching | [`channels`](#channels) | Inspect or set the channel registry |\n| Messaging & watching | [`history`](#history) | Clear retained message history |\n| Messaging & watching | [`console`](#console) | Live protocol view for a space (TUI, or `--plain` line stream) |\n| Messaging & watching | [`web`](#web) | Browser dashboard (installed as the `@cotal-ai/web` extension) |\n| Auth & meshes | [`mint`](#mint) | Mint a creds file for a space (static auth mode) |\n| Auth & meshes | [`login`](#login) | Sign in to a per-user-auth mesh's IdP (once per machine) |\n| Auth & meshes | [`logout`](#login) | Revoke the IdP session and clear the cached login |\n| Auth & meshes | [`actor`](#actor) | Manage a user-auth space's actor ledger (grant / revoke / list) |\n| Auth & meshes | [`doctor`](#doctor) | Credential-health diagnosis and repair (`doctor auth`) |\n| Auth & meshes | [`join`](#join) | Join a space as your own presence (interactive) |\n| Manifest | [`topology`](#manifest-deploys) | Validate and view a mesh manifest's access graph (read-only) |\n| Extensions & misc | [`ext`](#ext) | Install / remove operator CLI extensions |\n| Extensions & misc | [`completion`](#completion) | Print or install shell completion |\n| Extensions & misc | [`feedback`](#feedback) | Send feedback to the Cotal developers |\n| Extensions & misc | [`deliver`](#server-daemons) | Run the server-side Plane-3 delivery daemon |\n| Workflow runs | [`run`](#run) | Operate durable workflow runs: start, resume, list, inspect, answer a checkpoint, check an edited program with migrate |\n| Extensions & misc | [`feedback-intake`](#server-daemons) | Run a self-hosted feedback intake server |\n\nThe manifest modes of `up`, `spawn`, and `down` (`-f <cotal.yaml>`) plus `topology` are covered\ntogether under [Manifest deploys](#manifest-deploys).\n\n## setup\n\n```bash\ncotal setup [--full] [--demo] [--yes] [--skills]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--full` | off | Redo the full guided flow (implies `--demo`) |\n| `--demo` | off | Also seed the guided expert team (`david`, `sven`, `me`) |\n| `--yes`, `-y` | off | Non-interactive accept-all (for agents / CI) |\n| `--skills` | off | Reconcile Cotal skills only through installed connector providers, plus `~/.agents/skills`. Refused with `--full` or `--demo`. |\n\nGuided setup is **configure-only**: it checks prerequisites, invokes installed connectors' declared setup providers, and\nseeds persona files, and it launches nothing (no mesh, no web, no manager). First run gets the\nnarrated flow; later runs print a status card. By default it seeds one `default` persona; the\n`david`/`sven`/`me` team is opt-in via `--demo`. `cotal status` points stale Claude skills and\nout-of-date `.agents` skills at `cotal setup --skills`, not unscoped `setup`. See [Getting started](getting-started.md) and, for\nmaintainers, [setup internals](setup-internals.md).\n\nWhen a mesh resolves, setup seeds that mesh's recorded `.cotal/agents` catalog, the same catalog a\nfollowing `cotal spawn` reads. It prints the absolute destination. On a fresh machine with no mesh it\nuses this folder and says why; when several meshes are available and none is selected, it refuses\nrather than choosing a catalog.\n\n## update\n\n```bash\ncotal update [--self] [--space <s>] [--server <url>] [--creds <path>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--self` | off | If a newer release exists, install that exact validated `cotal-ai` version globally and reconcile through the newly installed binary |\n| `--space`, `--server`, `--creds` | resolved mesh | Select the running manager whose continuity state is reported |\n\nWithout `--self`, `update` keeps the installed first-party surfaces coherent with the running\nbinary: it force-reconciles the four built-in connectors, then reinstalls other `@cotal-ai/*`\noperator extensions at the binary's exact version. Each extension runs in an isolated child, so one\nfailure cannot poison later replays. It then checks npm; a newer binary is an informational notice\nwith `cotal update --self` as the next command, not an automatic install.\n\nAfter disk reconciliation, `update` reads the selected running manager. A machine with no recorded\nmesh has no running manager to observe, so that read is skipped and the command completes. The same\nholds when every recorded mesh is down and none is selected. A recorded mesh that is down is still a\nrefusal when the command selects it, with `--space` or by running inside its project, and so is a\nnamed space that is not running. With several meshes running and no `--space`, `--server` or\n`--creds`, the install is machine-wide, so every running manager is reported in turn, each under its\nspace name, before anything is written; a `legacy` verdict on any of them makes the whole run not a\nhot update. A selector flag still reports one manager. A manager without a\ncustody generation is reported as `legacy`: it cannot preserve its manager-owned PTYs, so the\ncommand says that this is not a hot update and prints `exact`, `fork`, `fresh`, or `drain-only`\nfor every seat. This report sends no stop, preservation-commit, or replacement command.\nIt does not preserve a running PTY on a legacy manager. On Linux a detached custodian\nowns each PTY, so a manager-worker death no longer closes the seat and `status` reports\n`custodied`. Other platforms still spawn in-process and report `legacy`. An incompatible native\n`@lydell/node-pty` or ConPTY ABI break remains an explicit per-seat maintenance cut.\n\nWith `--self`, the selected running manager is reported before any global install. When a newer\nrelease exists, Cotal then installs the exact version it validated, resolves and verifies that\npackage in npm's global root, then launches that binary with the same `--space` / `--server` /\n`--creds` selection to reconcile connectors and first-party extensions to the new generation. An npx\nor dev-clone invocation therefore installs and continues through a separate global copy; it never\nclaims the already-running process changed. If the binary is current, `--self` performs the normal\nlocal reconcile without reinstalling it.\n\nThird-party extensions are listed with their installed version and recorded spec but are not\nauto-updated in v1. Floating third-party updates require `@cotal-ai/*` peer-range validation and are\na future follow-up. A failed connector/extension install, npm metadata check, or requested global\ninstall is reported and makes the command exit nonzero. Independent extension attempts continue so\nthe output includes every failure; an unavailable npm registry does not undo a completed local\nreconcile, but the command still exits nonzero because it could not establish that the install is\ncurrent.\n\n## up\n\n```bash\ncotal up [--detach] [--open] [--space <s>] [--server <url>] [--channels <path>] [--runtime <name>]\ncotal up --user-auth --idp <url> [--exchange-public-port <n> --exchange-public-url <https://…> [--exchange-trusted-proxy]]\ncotal up --tls-cert <cert.pem> --tls-key <key.pem> # serve broker TLS (both, or neither)\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\ncotal up -f <cotal.yaml> [--dry-run] [--runtime <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--server <url>` | auto (free local port) | Listen URL override |\n| `--host <host>` | none | Bind host override for a **fresh** broker boot: an IP or hostname only, never a URL (that is `--server`) and never `host:port` (the port comes from `--server` or its default); a URL or port-bearing value is refused pointing at the right flag. With no `--server`, the broker URL is derived from it, so `--host <addr>` alone is enough to make a mesh reachable at that address; a `--host`/`--server` pair naming different addresses is refused. A wildcard bind (`0.0.0.0`, `::`) keeps a dialable loopback URL. Recorded on the mesh and reused by every later manager launch, so a repair or resume keeps remote [`attach`](#managed-seats) working. A live refresh (`✓ mesh already running`) does not rewrite `.cotal/auth/server.conf` or rebind nats; stop the broker, then re-run `up --host` |\n| `--space <s>` | the folder's name | Space name |\n| `--store-dir <dir>` | none | JetStream store directory |\n| `--max-file-store <bytes>` | nats-server's dynamic cap | JetStream file storage cap in bytes (a positive integer, no unit suffix). Without it nats-server sizes the store at start as three quarters of the free space on its filesystem. The cap is fixed at broker start: a running broker cannot change it (`cotal down` first), `down --preserve-state` keeps it for the resume, and a resume with a different value is refused. Not accepted with `-f` |\n| `--channels <path>` | `.cotal/channels.json` if present | Channel-registry seed file (JSON). An explicit path that is missing is an error |\n| `--restore <dir>` | none | Restore a completed offline backup before exposing the normal listener |\n| `--restore-only registry` | artifact selection | Restore only the registry component |\n| `--accept-missing-source` | off | Explicit disaster consent when the inode-bound preserved source is absent |\n| `--accept-stale-checkpoint` | off | Explicit consent to resume a seat whose checkpoint was captured outside its recorded recency horizon |\n| `--open` | off (auth) | Unauthenticated dev mesh: no JWT, no ACLs |\n| `--user-auth` | off | Per-user auth: people `cotal login`; connects are authorized against the actor ledger |\n| `--idp <url>` | none | With `--user-auth`: the IdP auth base URL to pin on first enable |\n| `--exchange-public-port <n>` | none | With `--user-auth`: add the public exchange face on this loopback port, for an HTTPS reverse proxy to forward to |\n| `--exchange-public-url <https://…>` | none | With `--exchange-public-port`: advertise the reverse proxy's HTTPS URL in discovery |\n| `--exchange-trusted-proxy` | off | With `--exchange-public-port`: attribute public failure buckets to the last `X-Forwarded-For` hop. Enable only when the listener is reachable solely through a trusted proxy; otherwise the socket address is used |\n| `--detach` | off | Run in the background (stop with `cotal down`) |\n| `--tls-cert <path>` | none | PEM certificate to serve TLS with. Must be given together with `--tls-key`. Before starting the broker, Cotal checks readability, private-key mode, key/certificate match, the validity window, and host coverage. `nats-server` accepts an expired certificate and leaves the failure to clients, so Cotal performs these checks first. The decision is recorded; a later bare `cotal up` keeps serving TLS |\n| `--tls-key <path>` | none | PEM private key for `--tls-cert`. Refused if group- or other-readable (tighten to `600`) |\n| `--file <cotal.yaml>`, `-f` | none | Launch a whole mesh from a manifest |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--runtime <name>` | `pty` (or the manifest's, with `-f`) | Agent runtime for the mesh manager (`pty` built in; others are installed extensions, explicit-only). Resolved + probed before the broker starts; an uninstalled/unreachable runtime fails loud. With `-f`, overrides the manifest's runtime |\n| `--max-sessions <n>` | 64 | Live-session ceiling for the mesh manager. Each console pane and each `cotal attach` is one session, so size for agents × panes, not agent count. Recorded on the mesh and reused by every later manager launch, so a repair or resume does not silently drop back to 64. A running manager cannot change it: `cotal down` first, then `cotal up --max-sessions <n>` |\n| `--no-manager` | off | Broker-only boot: start the broker and, in auth mode, the delivery daemon, and no local manager. A refresh under the flag of a mesh whose manager is live refuses rather than keeping or stopping it (`cotal down manager` first). Cannot be combined with `--runtime`, `--max-sessions`, or an agent-declaring manifest |\n| `--rotate-sys` | off | Rotate the space's system account and re-mint its two `$SYS` creds. Needs a stopped mesh; refused with `--open` |\n\n`cotal up` boots a local nats-server with JetStream and, in auth mode (the default), JWT auth and\nper-agent ACLs; `--detach` records the mesh so `cotal spawn` from any directory can find it. With no\n`--server`, it auto-selects a free port if the default address is taken; an explicit `--server`\nstays fail-loud on collision. `--detach` also brings up the control plane (delivery daemon in auth\nmode, then the manager). `--no-manager` is the broker-only mode: it boots\nthe broker (and the delivery daemon in auth mode) and starts no manager, so there is no manager\npidfile to leave stale. A refresh under the flag of a mesh whose manager is live refuses rather\nthan keeping or stopping it: `cotal down manager` first. For a split topology with a manager, wait for `.cotal/manager.<spaceKey>.log` to contain `✓ manager up`, then `cotal down manager` on that\nhost and run [`supervise`](#supervise) against the remote broker; see\n[Run a mesh](run-a-mesh.md). `cotal up --detach` prints `✓ running in the background:` with\n`manager` listed (pidfile liveness, not a teardown boundary); with `--no-manager` the line lists\nonly what actually started. The `-f` form is a [manifest deploy](#manifest-deploys).\n\nThe generated `.cotal/auth/server.conf` is written on a real broker boot and is not an\noperator-owned config. `--host` changes that file only when nats is actually started. A unit\nrestart that leaves an answering listener in place is a refresh, not a rebind.\n\nOn an existing mesh, `cotal up` reconciles the presence and lease bucket TTLs. It writes a reserved\ncanary and waits for the bucket to expire it before reporting success. If the broker accepts the\nstream update but the backing store does not persist or enforce it, `up` exits nonzero with a TTL\npersistence error instead of trusting the value returned by stream info.\n\n`--user-auth --idp <url>` starts the space's auth service alongside the broker: the NATS\nauth callout plus its capability-gated local exchange, and optionally the closed public exchange\nface configured by the three `--exchange-*` flags above. The service is torn down with `cotal down`,\nand a re-run of `cotal up` heals a dead service on a running broker. `--user-auth` and `--open`\ncontradict each other and are refused loudly; a running broker cannot change auth mode\nwithout a `cotal down` first. See [identity & auth](identity-and-auth.md).\n\n`--rotate-sys` renews the two `$SYS` credentials (`membership-observer`, `connection-evictor`).\nThey carry a 30-day expiry and nothing re-signs them in place, because the system-account seed is\nnever persisted, so they are renewed by issuing a **new system account** under the same broker\noperator and minting fresh creds against it. A plain re-`up` does **not** do this: it reuses the\nexisting trust record, and its `$SYS` creds along with it.\n\nThe rotation is safe to run on a real space, with one operational cost. The data account, the account\nsigning key, every agent credential minted from it, and the JetStream store are all untouched; what\ndies is the retired system account, and with it any out-of-band copy of the old `$SYS` creds, on every\nbroker that loads the rotated config. The cost is that **earlier full backups stop being restorable**\n(see below), so this is not a no-consequence operation. It needs the broker to restart on the rewritten\nconfig, so it runs as part of a boot:\n\n```bash\ncotal down\ncotal up --rotate-sys --detach # agents reconnect; nothing is re-provisioned\ncotal doctor auth # both $SYS creds healthy again, 30 days out\n```\n\nA rotation is a stopped, fresh boot, and anything that is not one refuses it, all for the same reason\n(the on-disk material and the broker it runs on must never end up on different generations):\n\n- a live mesh, because the running broker would keep serving the retired account;\n- an open mesh, whether that comes from `--open` or from `broker.auth: false` in a manifest, which\n has no system account at all;\n- `--restore`, because reinstating a trust root and superseding it in one command leaves no way to\n say which authority the mesh came up on;\n- an unfinished restore or resume attempt on this root, including one `cotal up` would recover on\n its own, because those paths can adopt a live listener and return without booting a broker;\n- a root that hosts more than one space, because the system account lives in the shared broker\n record and a rotation would retire every tenant's, while the root holds one `$SYS` cred pair\n pinned to one data account.\n\nTwo things to know before you run it:\n\n- **The retirement is config-load-bound.** Old `$SYS` creds are refused by any broker that loads the\n rotated config. A stale `nats-server` still running the *previous* config in memory would keep\n honouring them, so stop every broker for this root first. `--rotate-sys` refuses if this root's\n mesh is recorded as running, if anything unidentified is answering at the address it was given, or\n if the root's pid file names a live (or unreadable) process. Those are Cotal's own ownership\n records, not a scan of the process table: a `nats-server` you started by hand against this root's\n `server.conf` on some other port writes none of them and will not be seen. Do not run one.\n- **It invalidates earlier full backups.** A full artifact binds to the trust chain it was taken\n against, and that commitment covers the operator JWT and the system account. Every full backup\n taken before a rotation refuses to restore afterwards, so take a fresh `cotal backup` once the\n rotated mesh is up. `cotal up --restore` names this case when the data account still matches.\n\nThe commit is not atomic (a trust-record write plus two credential writes), so an interrupted\nrotation leaves the record ahead of the creds. That split is detected rather than silent: every\n`cotal up` on an auth mesh, and every `cotal doctor auth`, compares each `$SYS` cred's issuer against\nthe persisted record and names the retired account. `up` warns rather than refusing, because these\ncreds power the membership graph and live eviction, both of which degrade fail-soft; the mesh is not\nworth taking down over them. Re-running the rotation heals it, at the cost of one generation.\n\nWhile those creds are expired the mesh keeps delivering messages, but the\n[membership feed](delivery-daemon.md) and live connection eviction stay down; `cotal doctor auth`\nand the manager's log both name the credential and this repair.\n\n## down\n\n```bash\ncotal down\ncotal down --with-agents\ncotal down --preserve-state [--store-dir <dir>] [--session-store <dir> …]\ncotal down manager [delivery auth web nats ...]\ncotal down web [--space <name>]\ncotal down -f <cotal.yaml> | --run <id> [--dry-run]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--file <cotal.yaml>`, `-f` | none | Tear down this manifest's deploy |\n| `--run <id>` | none | Tear down one `spawn -f` run by id |\n| `--space <name>` | current mesh | With components: the mesh whose target-addressed components (e.g. `web`) to stop |\n| `--dry-run` | off | Print the manifest teardown or selected components, mutate nothing |\n| `--with-agents` | off | Bare whole stack only: also stop and deprovision every managed agent |\n| `--preserve-state` | off | Bare whole stack only: fence the manager, retain principals and durable state, stop and prove the stack down, then publish `ready` |\n| `--store-dir <dir>` | `.cotal/nats` | With `--preserve-state`: the actual store path (required for a custom store) |\n| `--session-store <dir>` | none | With `--preserve-state`: a harness transcript store directory to capture with every continuation-capable retained seat. Repeatable. No default and never inferred from a connector name; a path that does not exist or is not a directory is refused before anything stops |\n\nBare `cotal down` stops the whole local stack in dependency order and leaves managed agents running.\nBefore signalling the manager it verifies that the exact recorded manager supports releasing its\nlocal custody, and it reports the agents left behind plus `cotal down --with-agents` as the explicit\nreap. `--with-agents` is a one-shot destructive policy bound to the exact verified manager process\nand the exact live `down` stop reservation; a stale, malformed, crashed, or different stop attempt\ncannot turn a later bare shutdown destructive. Positional component names stop\nonly those self-registered local processes; for example, `cotal down manager` leaves delivery and\nthe broker running, and `cotal down web` is available when the web extension is installed. A\ncomponent that starts target-resolved (the web dashboard) is stopped the same way: `cotal down web`\nresolves the mesh the same way as `cotal web` (registry current mesh first, `--space` to name one), so\nit works from any directory; the other components always stop under the folder you run it in. The\n`-f` / `--run` forms tear down a [manifest deploy](#manifest-deploys) without stopping the whole mesh\nand cannot be combined with component names. Stopping `nats` alone is refused while an unselected\nregistered daemon is still live; include those components or use bare `cotal down`.\n\nBare `cotal down` inventories by pidfile. When this folder's registered broker answers and no\n`nats.pid` records it, the command does not say nothing is running. It names the space and the\nbroker address, says no pidfile records that process, says it will not stop a process it did not\nstart, and exits 1. Stop that broker with whatever started it (an init unit, a container, or the\nhand-run process). `cotal meshes rm <space>` only drops the registration. The probe runs whether or\nnot other owned components were running: they stop and clear their artifacts first, then the broker\nis named. A component stop and `--dry-run` stay pidfile-only and do not probe.\n\n**Teardown verifies pinned process identity before signalling.** PIDs are recycled by every OS,\nso a recorded pid alone is not a durable target identity. `up` records each stack process's\ncreation identity in a sibling `<pidfile>.identity` pin, which holds the pid and the process start\nreported by the OS. Every stop path, including `down` for the broker, web and extension components,\nand the manager, delivery and auth-service stops, applies the same rule. A pin that names a different\nstart means the pid was reused, so teardown refuses and preserves it. A torn or unreadable pin also\nrefuses.\n\nThe pidfile pid and the pin pid are two coordinates. Automatic cleanup follows **proven death of\nthe pidfile target** (ESRCH on that pid): a torn sibling pin does not wedge a dead pidfile pid.\nA torn pairing where the pin names another pid, while the pidfile pid is still live or not proven\ndead, still refuses. Inspect both pids with `ps`. Do not delete `<pidfile>.identity` to force a\nstop; that weakens target-identity protection. Once the pidfile process is dead, rerunning\nteardown clears the stale record automatically.\n\nThe first teardown after upgrading a running pre-pin stack has a narrower guarantee. A live record\nwith no identity pin is signalled after a loud warning that it predates identity pinning. Restarting\nthe component writes the pin, so later teardowns receive full match and mismatch protection. The\nsame warning applies on platforms where no stable start token is available. For a legacy manager,\nbare `cotal down` also warns that agent sparing cannot be verified before it signals. Because the\nCLI cannot establish which SIGTERM handler that already-running binary carries, it never presents\nthe pre-signal seat inventory as confirmed spared; a genuinely older destructive handler may still\nreap those agents. `--with-agents` publishes a one-shot reduced-guarantee handoff bound to the\nrecorded manager pid and the live `.stopping` reservation's inode, then signals unconditionally.\nThat handoff cannot be replayed by a later stop attempt. A pin that exists and does not match the\nlive process still refuses before signal.\n\n`--with-agents` performs the old destructive logical teardown: managed processes stop and their\ncredentials, ACL rows, and delivery footprints are deprovisioned. `--preserve-state` is a different\nmaintenance transition: it stops retained processes while suppressing leave/deprovision cleanup, persists the manager's\nsame-principal resume inventory, stops the entire stack without removing run/auth artifacts, and\npublishes a stable inode-bound cut only after every recorded process is proven stopped and the exact\nrecorded NATS endpoint is unreachable. A missing or stale broker pidfile never counts as stopped. The\nattempt is bound durably before the manager is fenced, the resume document and attempt-bound\n`cut-intent` are fsynced before manager commit, and the manager's commitment itself is journaled\n(`cut-committed`) before any process stops. A retry after a crash at any of those boundaries reuses\nthe exact recorded attempt and finishes the remaining stop and endpoint proofs idempotently, without\nneeding the (by then intentionally dead) manager. A partial cut never publishes `ready`. It cannot\nbe combined with component names, manifest teardown, or `--dry-run`.\n\n**Seat checkpoints.** After the stack is proven down, the cut writes one checkpoint per retained\nseat under `.cotal/maintenance/v1/checkpoints/<attempt>/<seat>/`, and prints the path, the\ncontinuity class and the generation for each. The path carries the preservation attempt because a\ncheckpoint is immutable once sealed: a shared directory would make the second cut in a root refuse\non the first cut's leftovers, and clearing it would destroy an artifact a rollback still needs. The\ncapture happens only at that point because anything earlier races a harness that is still writing\nits transcript and its working tree.\n\nEach checkpoint directory is created 0700, refuses a destination that already exists, and holds:\n\n- `repo.bundle`, the seat `cwd`'s reachable history, anchored on the base commit the record names\n by full object id;\n- `repo.index.diff` and `repo.worktree.diff`, the staging state as two diffs, base to index and\n index to worktree. Two rather than one because a single combined diff restores a mixed tree with\n the right bytes and the wrong index: a source reporting `MM README` would come back as ` M README`;\n- `repo.untracked.tar`, the untracked files in scope;\n- the harness session pointer, when the seat's connector declares one, and the transcript store\n files the operator named with `--session-store`. Each records where the destination puts it back\n as an anchor (the workspace root, the account home, or the seat's `cwd`) plus a relative path,\n because the destination's root and home are its own and the source host's absolute spelling would\n either miss them or write outside them;\n- `checkpoint.json`, written last, after every digest is computed over the bytes that landed.\n\nThe record carries the manager's resume entry unchanged as its first field, then the space, the seat\nname, the recovered `lifecycleUid`, the writer generation the cut was taken at, `capturedAt`, the\nrecency horizon, the applied profile revision, the seat's `git status --porcelain` as the cut read\nit, and the continuity class. Every captured file is\nlisted with its byte size and sha256, so an operator verifies the whole artifact with `sha256sum`\nand `git bundle verify`. No secret values, no operator keys and no source-host launch material\nenter it.\n\nThe continuity class is what the connector declares, capped by what the checkpoint carries. A\nconnector declaring session continuation classifies as `exact`, but reopening a session takes both\nhalves, the pointer that names it and the store that holds its transcript. A checkpoint missing\neither one cannot reopen that session, so it is recorded as `fresh` when the connector declares a\nfresh start and `drain-only` otherwise. A pointer with no store is capped the same way as a cut\ncarrying neither, because it names a session whose bytes the artifact does not contain. A class is a promise the destination is entitled\nto act on, so it never describes bytes the artifact does not contain. The transcript store stays an\noperator input: this repository does not know where a harness keeps its transcript, so `exact`\nrequires `--session-store` to name one.\n\nThe recorded status is read under the same selection rule as the untracked set, so it describes the\nstate the captured bytes can reproduce. The destination re-reads it in the promoted tree and refuses\na difference.\n\nThe untracked selection rule is recorded in the record and is\n`git ls-files --others --exclude-standard -z, excluding .cotal/`. It honors `.gitignore`, so an\nignored file the seat needs does not travel and has to be moved separately. The `.cotal/` exclusion\nis a secrecy boundary rather than a size one: when a seat's `cwd` is also the mesh root, the control\ndirectory is untracked, and without the exclusion the broker trust material, the space account, the\nmanager instance identity's private seed and the seat's own credentials would land inside the\nartifact. A checkpoint carries credential references only; the destination resolves that material\nitself.\n\nA seat whose launch options could not be resolved is refused rather than checkpointed, with the\nmanager's own wording: `imperative launch options have no non-secret durable source (<keys>)`. The\nrefusal arrives at prepare time, so the cut stops before any child does.\n\n## clean\n\n```bash\ncotal clean <history|store|all> --force\ncotal clean restore-attempt --attempt <id> --force\ncotal clean restore-fallback --attempt <id> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | `history`: target mesh |\n| `--dms` | off | `history`: also clear DM history |\n| `--store-dir <dir>` | `.cotal/nats` | `store`/`all`: JetStream store directory |\n| `--force` | none | Required: destructive, no prompting |\n| `--attempt <id>` | none | `restore-attempt`: exact stale pre-commit attempt; `restore-fallback`: matching healthy committed restore |\n\nOne configurable cleanup verb; every target requires `--force`.\n\n- `history` purges the retained message backlog on the **running** broker (channels, plus DMs\n with `--dms`). The same operation as [`history clear`](#history), which stays as an alias.\n- `store` deletes the **stopped** mesh's JetStream store (`.cotal/nats`): streams, durable\n consumers, and messages. This is the reset for stale on-disk broker state, e.g. durables\n minted by an older, incompatible Cotal generation surviving a `down`/`up` cycle.\n- `all` is `store` plus the space identity (`.cotal/auth`), the local creds and markers tied to\n it, any crash residue a normal `down` would have swept (stale pidfiles, `run/`), and the mesh's\n registry entry; the next `cotal up` mints a fresh identity.\n\n`history` needs the mesh up; `store` and `all` refuse while any recorded mesh process is still\nalive or any same-root recorded broker endpoint remains reachable (run `cotal down` first). They\nalso refuse outright on a root that holds accounts for several spaces: the store and the broker\ntrust record are shared by every space on the broker, so both targets would take out all of them\nand no `--space` can narrow that. `down`, `backup` and `up --restore` refuse there for the same\nreason. `cotal status` lists the tenants on such a root. Personas\n(`.cotal/agents`) and logs are never touched. A custom\nstore location is not recorded anywhere, so `--store-dir` must repeat whatever the mesh was\nlaunched with. Custom cleanup targets must contain either the Cotal store-generation marker or a\nreal `jetstream/` store directory; filesystem roots, project roots, and Cotal auth/maintenance trees\nare always refused.\n\n`store` and `all` also refuse every maintenance journal state. After a healthy committed restore,\n`restore-fallback` is the only supported way to remove the recorded unchanged old-store inode; it\nnever deletes the active target, requires both the exact attempt id and `--force`, and retires the\ncompleted restore journal so a later `down --preserve-state` can start a new backup cycle.\n\n## Backups\n\n```bash\ncotal down --preserve-state [--store-dir <dir>]\ncotal backup create <dir> [--only full|registry] [--store-dir <dir>]\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\n```\n\nBackup is offline-only. It requires the stable `ready` record from `down --preserve-state`, an exact\nstore match, no live recorded process, and an unreachable exact endpoint from the recorded cut.\nThat endpoint is probed immediately before cloning, so a live broker with a missing or stale pidfile\nis still refused. It claims the cut, reflink/copies the stopped source to a\nprivate attempt clone, and opens only that clone on a random loopback bootstrap broker with an\nindependent parent/deadline watchdog. It validates the canonical stream and pull-consumer inventory,\nwrites native snapshots with consumers excluded, and stores conservative contiguous ACK-floor\ncheckpoints separately. The original store is never opened by the backup broker, and the stack is\nnot restarted implicitly. Artifact destinations must not overlap the preserved source or maintenance\nattempt tree. Restore artifacts and targets likewise cannot nest inside or contain each other, the\npreserved source, or the maintenance attempt tree.\n\nStopped client-managed KV ordered consumers are ephemeral read residue, not backup state. Backup\nignores only the pinned client's exact stopped shapes: ordinary last-value watchers and the\nwhole-bucket scanner that uses all-history delivery to collapse concurrent tombstones. A bound\nconsumer or any lookalike with a different filter, inbox, lifetime, or other config is still refused.\n\n`full` is the default and indivisible: channel registry, CHAT/DM/TASK/INBOX/DLV, ACL, MEMBERS, and\nvalidated durable checkpoints. `registry` is the sole partial artifact. Presence, derived membership\nfeed, leases, native ephemeral/history consumers, credentials, keys, tokens, owner secrets, and actor\nledger files are excluded. `full` means every transferable message and registry stream, not every\nJetStream resource: endpoint submissions/facts/events/timers/workflow state, contract artifacts, and\nthe records/auth/session stores are nonportable control state. Restore recreates those streams empty\nwith their canonical configs before exposing the normal listener, so active endpoint runs,\nlifecycles, and sessions do not cross a backup. Artifacts are exclusively created `0700`;\nsnapshot/checkpoint files and\nthe manifest are `0600`; `manifest.json` is written last with exact sizes and SHA-256 values. The\ndirectory is trusted operator input: hashes detect corruption, not malicious rewriting.\n\nRestore validates and stages the exact allowlisted artifact bytes before moving or creating a store.\nIt requires the same space and existing trust state. The whole pre-commit window holds a journaled\nliveness claim (coordinator, watchdogs, brokers, absolute deadline): ordinary `up` and a repeated\n`up --restore` refuse while the claim is live, and a stale attempt is recovered only after the\ndeadline has elapsed and every recorded owner is proven dead. A retried `up --restore` handles this\nautomatically; an operator can also recover it explicitly with `cotal clean restore-attempt --attempt <id> --force`. Nothing\never rolls back a live attempt. A registry-only artifact restores as registry-only whether or not\n`--restore-only registry` is passed; omitted infrastructure is always created and the exact\npost-restore stream inventory is asserted before commit intent. Ordinary `up` from a preserved cut\nresumes only the exact recorded source store and runtime; a contradicting `--store-dir` or\n`--runtime` fails in preflight.\n\n**Admitting a seat checkpoint.** An ordinary `up` from a preserved cut admits that cut's seat\ncheckpoints before it journals the resume attempt and before any process starts, so a refusal costs\nnothing. Three gates run in order, each naming what it saw.\n\n1. *Integrity.* Every file the record names must be present, a regular non-symlink file, the\n recorded byte size and the recorded sha256, re-stat'd after the read so a file that moved is a\n refusal. Failure here consults no other gate.\n2. *Identity.* The recorded space must match, the recorded `lifecycleUid` must not belong to a live\n incarnation, and the profile revision must match this host's or be resumed under deliberately\n this host's. A differing revision is refused with both digests and the remedy, and there is no\n override: the checkpoint carries the recorded digest and not the config bytes, so nothing could\n run the seat under the recorded revision, and the manager re-digests the same file and refuses\n drift on its own. This gate has no blanket override, which is the only reason the next one may\n have one.\n3. *Recency.* `capturedAt` is compared to this host's clock against the horizon the record carries.\n Inside it, the seat resumes. Outside it, `up` refuses and prints the capture instant, the clock\n reading and the horizon; `--accept-stale-checkpoint` admits it anyway and the exercised consent\n is printed with the actual age. An unreadable `capturedAt` is refused with no override, because a\n freshness gate that fails open is not a gate.\n\nCustody transfers only after all three pass. The destination claims the recorded generation plus one\nby exclusive create, before it launches anything. A lost create means another destination is already\nclaiming that seat, and it refuses with `seat-writer-generation-create-lost` rather than adopting\nthe winner and becoming a second writer. The recorded `lifecycleUid` is reused and never minted, so\nthe resumed seat binds the same lifecycle-keyed durables.\n\nAdmission is reconciled against the inventory the resume is about to hand the manager, and that\nreconciliation finishes before the restore moves a single tree. A checkpoint whose recorded\n`lifecycleUid` is not the one the retained inventory carries describes a different incarnation of\nthat seat, and it refuses with both uids while every live working tree is still untouched and no\ngeneration is claimed. A retained\nseat with no admitted checkpoint refuses the resume by name: an absent checkpoint directory and an\nabsent record are indistinguishable from a seat that was never checkpointed, and a seat that starts\nwithout passing the gates has claimed no generation. `--accept-stale-checkpoint` is recorded in the\nresume journal with the seat, the capture instant, the admitted age and the horizon, so the consent\nsurvives the terminal it was typed into.\n\nThe whole admission is all or nothing. Coverage is settled first, then every gate runs over every\ncheckpoint, and only then is any generation claimed. A refusal at any point leaves every generation\nunclaimed, including a lost exclusive create during the claim itself: the claims that attempt made\nare removed before the refusal is raised, by the exact paths it wrote, so a generation another\ndestination holds is never touched. A claim is a create that can never be made again, so a refusal\nthat left one behind would consume the retry over the same checkpoint set.\n\n**Restoring a seat checkpoint.** Once every gate has passed over every checkpoint, and before a\nsingle generation is claimed, `up` puts each admitted seat's captured bytes back. A refusal here\ncosts nothing for the same reason a gate failure does: no claim has been made and nothing has\nstarted.\n\nA restore never moves or replaces the destination's own control directory. A checkpoint excludes\n`.cotal/` by design, so a seat whose `cwd` holds one, which is the layout an operator gets by\nrunning `up` and `spawn` in a single directory, is refused before anything is staged: promoting a\ntree that cannot contain `.cotal/` over that `cwd` would carry this host's live trust material and\nmaintenance state away with the superseded tree. The refusal names the control directory it found\nand the remedy, which is to give the seat a working tree that is not a workspace root.\n\nEach seat is staged beside its own `cwd`, in `<cwd>.incoming`:\n\n1. every recorded digest is verified again over the files as they are now;\n2. the bundle is cloned into `<cwd>.incoming`, which is refused when that path already exists;\n3. the recorded base commit is verified in the clone and checked out detached, so a bundle that does\n not contain it stops the resume instead of continuing against a different history;\n4. the index diff is applied with `--index` and the worktree diff without it, both `--binary\n --allow-empty`. That order is what puts staged content back in the index rather than only in the\n worktree, and `--allow-empty` is why a seat with a clean tree is still restorable;\n5. the untracked archive is extracted.\n\nEvery seat stages before any seat is promoted. Promotion moves an existing `cwd` aside to\n`<cwd>.superseded.<timestamp>` and renames the staging directory into place, then puts the session\npointer and store files where the destination's connector reads them, then re-reads\n`git status --porcelain` in the promoted tree and compares it to the status the checkpoint recorded.\nA restore that applied without error and produced a different index is a refusal, not a warning. The\ntwo renames are the only steps that touch the path the seat will use, so a failure anywhere leaves\nevery seat's live `cwd` as it was.\n\nThe rename itself claims the superseded name, and a taken name gets a numeric suffix. The timestamp\nhas one-second resolution, so two promotions of the same seat within one second compute the same\npath; a rename onto a name that already holds a tree fails on every platform, and that failure is\nread as taken. Nothing creates the name ahead of the move, because Windows refuses to rename onto an\nexisting directory at all. A superseded tree is the thing that rename exists to keep.\n\n`git` and `tar` run as child processes with argument arrays, never a shell string.\n\nA leftover `<cwd>.incoming` refuses the resume by name. A staging directory from a failed run is the\nonly record of what failed, so nothing removes one automatically: inspect it, remove it by hand, and\nresume. A pre-existing `cwd` is renamed rather than deleted, so a wrong checkpoint costs a rename\ninstead of a tree. When a promotion fails, the renames that attempt made are undone and the staging\ntree is left where it is, as the evidence for what did not verify.\n\nA session pointer whose recorded `sessionId` is not the one the retained inventory reopens is\nrefused before anything is cloned. A session file already present at its destination is judged by\ncontent: bytes equal to the recorded digest are already restored, and different bytes under the path\nthe connector is about to read are refused with both digests rather than clobbered.\n\n`up --restore <dir>` reaches the same admission and the same restore, after the store is restored\nand validated and before commit intent is journaled. A registry-only restore resumes no seat, so it\nadmits and restores nothing.\n\nOne limit is worth stating plainly. The writer generation is claimed by exclusive create inside one\nworkspace root, so it fences two resumes on the same host and does not fence two independent\ndestinations: copy a checkpoint to two roots and both claim the same successor. A real cross-host\nfence needs a coordinate neither root owns.\n\nAuthenticated restores validate the complete\nspace trust bundle before staging, including nkeys, seed matches, JWTs, signers, and space binding;\nfull restores commit to the validated operator, system-account, data-account, and active-signer root\nchain in addition to the static/user authority fingerprint. Because the system account is part of that\ncommitment, a [`cotal up --rotate-sys`](#up) makes every full artifact taken before it unrestorable\nagainst this root: take a fresh full backup after each rotation. The composed commitment is revalidated\nimmediately before store mutation and never includes secret seeds. Restore never creates fresh auth.\nSame-path restores atomically retain the old\nsource at the journaled fallback path; alternate targets retain it in place; a missing canonical\nsource needs explicit `--accept-missing-source`. Quarantine and target restores use current canonical\nconfigs on isolated random-loopback brokers, never expose native snapshot consumers, and publish a\ncommit-intent immediately before the normal listener starts. Archive bytes never instantiate the real\ntarget: after quarantine validation, every stream is re-snapshotted from the validated quarantine\nstate into attempt-owned sanitized files, and the target is restored solely from those. Before that boundary, failure rolls back\nthe attempt-owned target; after it, ambiguity preserves both stores and records forward-repair\nrecourse. The cooperative maintenance lock excludes Cotal commands, not arbitrary raw NATS processes.\n\nBootstrap brokers in every auth mode, including open, mount the store under a local account with\nrandom operation-specific logins only, each carrying the exact per-phase subject permission matrix;\nnormal static credentials and user-auth sentinel/bearer connections are rejected, and no auth\nservice or callout starts. Open mode differs only in its account label, never in authority. Inventory, each stream snapshot,\nrestore initiation, exact upload id, validation, and each checkpoint recreation use separate exact\nauthorities. Every checkpoint carries the source stream's message/first/last sequence state and must\nmatch its snapshot record before mutation; core then derives and validates the only allowed start\npolicy. TASK is not a CLI exception: the same core checkpoint API recreates its canonical `DeliverAll`\nWorkQueue durable because acknowledged tasks are absent from retention and NATS forbids a\nstart-sequence policy there. Registry-only restore creates every omitted canonical stream and transient\nbucket on the isolated target before the normal listener is exposed. It deliberately does not resume\nretained agents or recreate their DM/DLV/TASK/ACL state; their identity material stays retained and\nstopped rather than being reprovisioned into a partial restore.\n\nAfter listener readiness, the manager starts attempt-bound, validates retained credentials/tokens\nwithout granting or reprovisioning, and resumes the exact persisted principals under cleanup\nsuppression. Registry-only restore uses the same flow with an empty agent set. `commitResume` is an\nidempotent validation barrier only: success must be `awaitingFinalize` with an attempt-bound 64-hex\ncommit token and does not release suppression. Under the workspace lock, the CLI first fsyncs that\nexact evidence as `manager-committed` (restore) or `resume-committed` (ordinary resume), then calls\ntoken-bound `finalizeResume`; only an `active` response for the exact token releases suppression. The\nCLI records the same token in finalization evidence before a restore becomes `active`, or before an\nordinary resume retires and consumes the marker. Re-entry from either committed state skips the prior\nidempotent activation/commit phases, retries finalization with the durable token, and finishes the\nworkspace transition. Failure before finalization preserves the committed state and cleanup\nsuppression; it is not rewritten through a degraded transition. Re-entry between any two earlier\nboundaries reuses the same attempt and may retry the idempotent phases without deleting retained state. A missing or\nchanged per-agent dependency is a named fail-closed result; the journal becomes degraded and remains\navailable for forward repair. A retry from `resume-intent`,\n`resume-active`, or `resume-degraded` reuses the same attempt and inventory after the prior listener is\nproven stopped. Every normal restore listener has an unguessable attempt-bound NATS server name. The\nCLI fsyncs its exact name/nonce, canonical endpoint, process owner, and generation-bound target identity\nimmediately after spawn. Re-entry accepts a surviving listener only when its INFO server name, live PID\nrecord, endpoint, and target identity all match that proof; degraded restore repair then moves through\nthe guarded workspace transition only after manager commit. If an uncommitted bound owner is provably\ndead, recovery retires that exact proof under the maintenance lock and binds a fresh listener for the\nsame attempt, endpoint, and target with a new nonce and server name. A live foreign/mismatched listener\nor ambiguous owner is preserved and refused, never adopted by reachability alone. A reconstructed\ncommit/degraded attempt without either the exact bound proof or a durable dead-listener replacement\nrecord fails closed even when the recorded port is free. A later ordinary startup may pass an `active`\nrestore only when its details prove manager commit and its exact recorded listener is dead.\n\n## Mesh registry\n\n```bash\ncotal meshes\ncotal meshes add # guided, on a terminal\ncotal meshes add <space> --server <url> [--root <dir>] [--mode auth|open|user] [--tls] [--force]\ncotal meshes add <space> --mode user (--user-auth-file <bundle.json> | --from <https url>)\ncotal meshes rm <space> [<space> …] [--force]\ncotal sync [--idp <auth base URL>]\ncotal use <space>\ncotal status [--space <s>] [--server <url>] [--components]\n```\n\n`meshes` lists the meshes this machine knows; a `*` marks the `current` default a bare\n`cotal spawn` joins. Entries learned from a signed-in account are marked `discovered`. Their\nregistration trust is stored under the account's private auth state, and the registry contains no\nsession token or sentinel credential bytes. Commands resolve the catalog `slug`; a different human\n`name` is rendered only as a label.\n\nA registry record this build cannot use is refused by name, never rendered and never skipped. One\nthat does not parse, or is missing a field every consumer reads (`server`, `mode`, `root`, `ts`,\n`space`), makes every registry command exit 1 with the file's path and what is wrong with it.\nRemove the file or restore the record; nothing repairs or invents a field for you.\n\nAn IdP may advertise a same-origin space catalog during login. Cotal reads the complete snapshot and\nadds every valid registration without a separate `meshes add`. A snapshot younger than five seconds\nis used without a request. After that, commands that select a discovered space require one successful\nconditional refresh before target resolution. A failed refresh refuses the operation. `cotal sync`\nbypasses freshness and reports added, changed, removed, unchanged, and name collisions. `--idp`\nlimits it to one signed-in account. It never connects to a broker.\n\nThe registry is updated under the same lock that guards the catalog cache, so a command never lists\na discovered space set that another command is still writing. The cache records a fetched snapshot\nas not yet applied before the first registry write and as applied after the last. If a command dies\nor is stopped in between, the next command applies that snapshot again before it can use it, with\nno request inside the freshness window.\n\nThe shared dispatcher applies this preparation to every command that declares both `--space` and\n`--server` as mesh-target flags, including commands registered by other packages and commands that\ndeclare their own equivalent flag objects. Daemon and startup commands that use those names only as\nconfiguration explicitly opt out. Registry-local `meshes add` and `meshes rm` never refresh a catalog.\n\nRun on a terminal with the space or `--server` missing, **`meshes add` is guided**: it asks for the\none thing that cannot be derived (the broker URL), probes it, and tells you what answered - open or\nrequiring credentials. It then offers the spaces your `--root` already holds credentials for, states\nthe mode as a fact about that broker rather than asking, and shows the exact record before writing\nanything. A broker that does not answer, or a space name already registered, becomes a choice rather\nthan an error. Anything you pass on the command line is taken as given and not asked again. Without\na terminal - a script, an agent, CI - nothing prompts and the flag form's errors stand\n(`COTAL_NO_PROMPT=1` forces that too).\n\n`cotal up` and `cotal down` maintain their own records. `meshes add` registers a mesh they cannot\nspeak for: one running on another machine, a shared broker, a hosted space. `--root` is the folder\nwhose `.cotal/auth` holds that mesh's credentials and whose `.cotal/agents` holds its personas.\nThe default is the project you run it in. The registry stores that path, never a secret. `--mode`\ndefaults to `auth` when the root holds the space's account record and to `open` otherwise. The\nbroker is probed before anything is recorded, so a wrong address, or credentials that mesh will\nnot accept, fails here instead of at the first `spawn`; `--force` records without verifying (and\nreplaces an existing record).\n\nA hostname or public address is registrable only when the connection will **require TLS**. Pass\n`--tls`, or use a `tls://` URL. The scheme is recorded as enforced intent, so every later dial\nthrough the record demands the handshake (and `meshes add tls://…` against a plaintext broker is\nrefused at registration). Without required TLS the fence admits loopback and private-overlay\nliterals only. RFC1918 addresses are refused in both modes because a cafe LAN is private but does not belong to you.\n\nA **user-auth** mesh registers from supplied pinned trust, never guessed: `--user-auth-file`\ntakes the bundle exported where the mesh runs; `--from` asks before it dials the address at all,\nthen fetches its `/.well-known/cotal-mesh` discovery document (HTTPS only), displays the pins, and\nasks again before adopting them. Neither fetch follows redirects: a 302 can move a pinned fetch\nonto plaintext or onto another host, so it is refused rather than followed, and the pinned\nexchange must itself be an `https://` URL, except for an exchange on this machine, where plain\n`http://` is accepted for a loopback *literal* (`127.0.0.1`, `::1`, any spelling of them) but not\nfor `localhost`, which is a name rather than an address. Registration verifies that the exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also verifies that the broker refuses a bare\nconnect; that auth-required refusal is the pass. The sentinel credentials land in a 0600 file under\nthe entry's root; the registry records only the path.\n\n`meshes rm` drops records. It never stops a mesh. For a mesh running on this machine `cotal down`\nis the right verb, and `rm` says so unless you pass `--force`. A hand-added record is removed by\n`meshes rm`, by an `add --force` replacement, or by a `cotal up` that actually starts the broker for that same space, server and root, which becomes that\nmesh and so takes the record over (a `cotal up` for that space anywhere else refuses instead).\nNothing that merely *infers* a record is stale from a dead broker touches it: an\nunreachable broker is listed `offline` and stays, whether `cotal up` or `cotal meshes add`\nwrote the record. A bare command does not treat that offline record as a running mesh;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the project\nthey tear down; a hand-added one they leave alone even when it shares a root, because nothing\non this machine could write it back.\n\nA discovered entry belongs to the normalized IdP origin and proved subject that supplied it. Local\nteardown, cleanup, and liveness pruning do not remove it. A manual or locally started entry with the\nsame name wins and remains untouched; that discovered name is reported as a collision. Logging out\nremoves only the discovered entries owned by that account.\n\n`cotal meshes` and `cotal status` print `events: required` for a registration carrying\n`policy: { events: \"required\" }`. On that space, foreground spawn, detached spawn, manager starts,\nand interactive `join` cannot opt out or join without an event plane. `--no-events` is refused with\nthe space named. A connector without an event plane is refused with both the space and connector\nnamed. A session whose own grant omits `events.<owner>.<actor>` is refused before joining and the\nmessage names a full-row `actor grant` repair.\n\n`use <space>` sets that default; the selection applies from every directory,\nincluding inside another mesh's project. `status` is a read-only report: machine prerequisites\n(starting with the installed `cotal-ai` version), the installed extensions and their versions, this\nfolder's `.cotal/`, the recorded meshes, and a live snapshot of the selected mesh (roster, channels,\nmembership feed). Stale Claude skills and out-of-date `.agents` skills recommend `cotal setup --skills`,\nnot unscoped `cotal setup`. `status` takes `--space` / `--server` to pick the mesh to inspect; it starts\nnothing.\n\nIf a refresh fails, `status` may still show the kept catalog bytes for diagnosis. It labels them\nstale with the last successful snapshot timestamp and the refresh error. It never calls that state\nsynchronized or online. If a selected discovered space vanishes from a successful snapshot, the\nselection is cleared and the command reports that no default is selected.\n\nFor a user-auth mesh the selected-mesh section reports the login `status` works as: the signed-in\nsubject when this machine holds a cached session for the entry's pinned IdP, or the exact `cotal\nlogin --idp <url>` line when it does not, with no network round trip either way. A locally\nprovisioned space also shows the actor grant row; a discovered or registered remote entry reports\nthe grant as not checkable on this machine, because the ledger runs where the space was\nprovisioned. `--components` on a user-mode target probes as that same signed-in login (`ps`'s\ncredential), never a static mint; when the login cannot supply a credential, the row says why\ninstead of printing the broker's refusal of an unauthenticated probe.\n\nPersona rows name the catalog they describe. If this folder and the selected mesh use different\ncatalogs, status names both and marks which one spawn launches from. A green `default` means the file\npasses the same agent-file loader spawn uses; a present but invalid file is reported as invalid.\n\n`cotal status --components` adds a fail-loud per-component health pass. It reads **each\ncomponent's own control surface**, rather than treating a PID, a lease, or a successful probe of a\nsibling as proof that the component serves. It prints one of `serving`, `absent`, `not-serving`, or\n`refused` for each component and exits `0`, `1`, `2`, or `3` respectively (the highest observed\nstate wins):\n\n- **manager**: local PID record, its liveness-lease holder and PID, then the manager's own typed\n `status` service reachability from this host. Manager builds that do not report static\n reconciliation say `static reconciliation not reported by this manager build`; the line stays\n visible even when the manager is otherwise `serving`.\n- **delivery**: local PID record, its ready lease (`ready` is the daemon's own bound-control\n signal), and the latest `renewal.<spaceKey>.json` adoption verdict, the record of the space the\n command was asked about, keyed per space the way the pidfiles are. A re-signed credential and a\n broker-accepted adoption stay distinct facts. A root-only `renewal.json` left by an older build\n names no space and is never read as any space's verdict (`doctor auth` names it as a leftover).\n- **web**: local PID record and the dashboard's own loopback `/api/meta` response, which must name\n the same PID and its requested port. A different process on the port, an unreadable PID command,\n or an unrecognizable process record is `refused`, not a green default-port guess.\n- **broker**: the registered mesh URL dialed from this host with its recorded TLS requirement.\n\n`absent` means Cotal has no live local component record (or has a stale record); `not-serving`\nmeans the component record is live but its service/readiness surface did not answer or is not ready.\nThose are intentionally separate exit cases. A failed or unreadable probe is `refused`, never an\nabsent component or a clean zero.\n\n## spawn\n\n```bash\ncotal spawn [<persona>] [--detach] [--name <n>] [--agent <a>] [--model <m>] [--variant <v>] [--prompt <text>] [--cwd <dir>]\ncotal spawn -f <cotal.yaml> [--dry-run]\n```\n\nFor a foreground spawn onto a remote user-auth mesh, a launcher may supply a one-time enrollment\ninstead of a cached human login. Prefer a private file:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./seat.md --space main\n```\n\nThe file contains only the enrollment URL, ending with at most one line terminator, and must be\nmode `0600` on POSIX. An orchestrator that cannot mount a file may set `COTAL_ENROLLMENT_URL`\ninstead; that value is redeemed byte for byte, so a trailing newline in it is refused. Setting both\nis refused. Enrollment input\nrequires `--space` and applies only to a foreground persona spawn. If the mesh is not registered yet,\nthe enrollment response must carry the stock user-bundle fields and the command needs\n`--config <persona-file>` because there is no local remote-mesh persona catalog to read. The client\nredeems the URL once, registers the returned mesh material, exchanges the returned actor token at the\npinned auth service, and removes both enrollment variables before starting any child process.\n\nA cached login for the same IdP and an enrollment are conflicting proofs, so the command refuses\nrather than choosing one. An invalid enrollment never falls back to login provisioning. Unknown,\nexpired, revoked, and already-used enrollments all produce one response: ask the owner for a fresh\none. See [Enrollment redeem](identity-and-auth.md#enrollment-redeem) for the HTTP contract.\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | resolved mesh | Target space |\n| `--server <url>` | registry entry | Broker URL override |\n| `--creds <path>` | none | Control-caller creds for an off-registry manager (`--detach` only) |\n| `--name <n>` | persona's `name:` | Presence-name override (does not choose the persona) |\n| `--config <persona-or-path>` | none | Persona catalog name or file path; wins over the positional |\n| `--agent <a>` | persona's `agent:`, else `COTAL_DEFAULT_AGENT`, else `claude` | Connector type (`claude`, `opencode`, `jcode`, `hermes`, and so on) |\n| `--role <r>` | persona's `role:` | Role override |\n| `--model <m>` | persona's `model:` | Model override |\n| `--variant <v>` | persona's `variant:` | Model variant override (connector-defined; e.g. OpenCode reasoning tiers) |\n| `--cwd <dir>` | this cwd | Working directory to root the agent at |\n| `--prompt <text>` | none | Initial prompt auto-submitted at start |\n| `--resume <id>` | none | Fork an existing session id into the mesh; only connectors that declare resume support accept it (see [the matrix](connectors.md)) |\n| `--no-events` | event plane on where supported | Opt out of the session's structured event plane (`--events` only restates the default) |\n| `--share-tools <sel>` | none | Share named operator MCP servers with the agent |\n| `--subscribe <a,b>` | persona's | Channel read-set override |\n| `--allow-subscribe <a,b>` | = subscribe | Read-ACL override |\n| `--allow-publish <a,b>` | deny | Post-ACL override |\n| `--detach`, `-d` | off | Launch via the manager into a detached PTY (reattach with `cotal attach`) |\n| `--on <instance>` | class anycast | With `--detach` only: pin the launch to one manager instance id (the whole id, as `ps` prints it). Refused on a foreground spawn (no manager to pin), with `-f` (a manifest deploy launches through the manager class queue), and when empty |\n| `--file <cotal.yaml>`, `-f` | none | Deploy a manifest onto the running mesh |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--allow-stale <a,b>` | none | With `-f`: waive named stale agents (apply-only) |\n| `--runtime <name>` | manifest's | With `-f`: override the manifest's runtime |\n\nEach session uses its connector's **event plane** by default: a stream of structured events\ndescribing what the agent did, rather than the prose it wrote, on a channel of its own. The channel is named after\nthe agent's principal, `events.<owner>.<actor>`, never after its display name, because two live\nagents are allowed to share a display name and would then share a stream. The launch grants publish\nrights on that channel alone, foreground and detached alike. `--no-events` is the explicit opt-out\nunless the selected registration says `policy: { events: \"required\" }`. Required policy makes the\nevent arm and grant mandatory, so `--no-events` and connectors without an event plane are refused.\n\nThe launch decision and the grant are separate on purpose. Holding publish rights on a channel is\nnot a request to publish to it, so writing an event channel into an agent file's `allowPublish`\ndoes not override `--no-events`.\n\nThe persona (`--config` > positional > `COTAL_DEFAULT_PERSONA` > `default`) is loaded from the\ntarget mesh's `.cotal/agents/`; the launch flags override the file. Foreground runs the agent\nattached to your terminal; `--detach` hands the launch to the running manager. Both modes get the\ndurable backstop on a mesh that runs the delivery daemon; `--live-only` skips it for a foreground\nspawn (messages posted while it is disconnected are then not replayed). A foreground exit retires\nthe agent's creds and broker footprint, like a manager despawn. On a user-auth mesh the two arms\ndiffer: a spawn against a mesh this machine provisioned revokes the actor row on exit, while a\nremote spawn (an enrollment or the advertised provisioning endpoint) removes only this machine's\ncredential files; its grant stays until the mesh operator revokes it, and the launch line says\nwhich arm you are on. A `--detach` spawn is an\n**action**: the manager accepts it and returns the allocated identity at once, then the launch\nfollows to a terminal outcome rather than blocking (see [the control surface](control-surface.md)).\nSee [Connect Claude Code](connect-claude.md) and [Agent files](agent-files.md); `-f` is a\n[manifest deploy](#manifest-deploys). (`cotal start` was merged into `cotal spawn --detach`.)\n\n## models\n\n```bash\ncotal models [--agent <connector>] [--refresh]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--agent <connector>` | all registered connectors | Connector whose catalog to list |\n| `--refresh` | off | Ask the connector to refresh its provider cache |\n\nAsks the running manager for each connector's model catalog (model ids plus their variants)\nfor connectors that expose one. OpenCode and Codex query harness/provider surfaces; Jcode reads\nproviders that enable `model_catalog = true` in the operator Jcode `config.toml`. Jcode's listed\neffort tiers render as `variants (declared, not provider-verified)`, and launch can still refuse one.\nA connector without a catalog says so. Pick a result with `cotal spawn --model <id> --variant <v>`,\nwhere `<id>` is the model id as the catalog printed it. OpenCode and Codex ids are the full\n`provider/model`; Jcode ids are bare (`opus-5`, not `cliproxy/opus-5`), because the provider is\nselected by the operator's Jcode config and a prefixed id is refused at launch with the bare form\nnamed.\n\n## endpoints\n\n```bash\ncotal endpoints [--space <s>] [--server <url>] [--creds <path>]\n```\n\nLists the mesh presence roster: agents, the manager, and any other protocol endpoint, with each\nendpoint's role, kind, status, and current activity. Unlike `ps`, this is a read-only presence view;\nit is not limited to child processes owned by the manager.\n\n## Endpoint control\n\n```bash\ncotal describe <endpoint> [--space <s>]\ncotal invoke <endpoint> <command> [--args '<json>'] [--space <s>]\ncotal invoke <endpoint> <command> --name <agent> [--admin] [--space <s>]\n```\n\nThe generic v0.4 service surface. `describe` resolves a registered endpoint's command set off the\nwire - the reserved `describe` command answers the registered contract digests, the schemas are\nfetched from the space's content-addressed contract store, recompiled, and verified against those\ndigests - and prints each command with its capability class and targeting shape. `invoke` calls one\ncommand by name: `--args` is a JSON object validated against the fetched input schema *before*\npublish; a targeted command takes `--name <agent>` (resolved to the agent's current principal through\n`inspect`) or `--self`. `--admin` uses the admin instrument credential, whose cross-agent reach rides\nthe operator-only `any` authorization mode. Neither command has compile-time knowledge of any\nendpoint's schemas - this is the same trust chain every built-in control command now uses. Needs an\nauth mesh: the manager registers its service on both static and per-user meshes (a signed-in user\nrides their bearer; each visible or invoked command still requires its existing grant, and cross-agent\nreach needs the `admin` scope). An open mesh has no service registry.\n\n## Managed seats\n\n```bash\ncotal ps [--on <instance>] [--wide | --json] [--space <s>]\ncotal stop --name <n> [--on <instance>] [--space <s>]\ncotal attach --name <n> [--on <instance>] [--no-reconnect] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | none | Managed agent to stop / attach (required) |\n| `--on <instance>` | class anycast (`ps`: class scatter) | Pin to one manager instance id (multi-manager space); takes the whole id as `ps` prints it, not a prefix. An empty value (`--on \"\"`, an unset shell variable) is refused, never treated as absent |\n| `--wide` (`ps`) | off | After each seat's compact row, print extra operational facts the manager records: `cwd`, `pid`, spawner, lifecycle uid, and the owning manager's instance id and host. Model and requested variant stay in the identity row rather than printing twice. A fact the manager did not record (for example a runtime with no real process) prints nothing, never a placeholder |\n| `--json` (`ps`) | off | Machine-readable: one JSON object per seat per line, copied unchanged from the manager row. Instance headers and errors go to stderr, so stdout contains only rows. Mutually exclusive with `--wide` |\n| `--no-reconnect` (`attach`) | off | End the attach when its session ends, instead of re-establishing it. For scripts that want one run and one exit code |\n\nThe human `ps` row is presentation text and is not a stable parsing target. Scripts use `--json`,\nwhich is the machine-readable row contract.\n\nThese are operator clients over the running manager's control plane. The default row includes the\nconnector, model pin, optional requested variant, and runtime as operational descriptors for the\nmanaged row. They do not make a shared display name a unique protocol identity; use `--json` when\nunambiguous owner+actor attribution is required. An omitted variant means no override was requested;\nCotal does not invent an effective provider default it cannot observe. `ps` also prints two state\nfacts per managed agent, because they answer different questions: the process fact from the manager's\nown runtime handle (`running` with its uptime, or `exited` with how long it ran), and the mesh fact\nfrom the roster (`idle` / `working` / `waiting` / `mesh offline`, or `not in roster` when the seat has\nno presence row at all: a seat that has not joined yet, or one that never did). A seat can be\n`running` and `mesh offline` at once: the process is alive and its presence has lapsed. The mesh fact\nis only a verdict while the manager's own presence watch is fresh: when that watch has been silent past\nthe liveness window, or has not replayed the bucket yet, every row prints `mesh unknown` with the reason\ninstead (`--json` carries it as `meshView: stale | unpopulated`), because `offline` and `not in roster`\nwould then describe the manager's watch rather than the seat. The manager rebinds a watch that goes\nsilent under a live connection on its own, so `mesh unknown` normally clears within a liveness window.\nOn a user-auth mesh `ps` also renders each managed agent's last credential-refresh outcome, fail-closed.\n\n**Mode split (chosen up front, never try-scatter-then-degrade):**\n\n- **Static / open mesh.** Bare `ps` is a **class scatter**: it freezes the live manager class from\n the records registry, merges every registered instance's agents grouped and attributed per\n instance, and a non-answering instance is shown as `registered, no answer within the deadline`\n (never silently omitted). A refused list or a missing answer makes the census incomplete: rows\n from other instances remain visible, but `ps` prints an incomplete-census warning on stderr and\n exits non-zero, including with `--json`. Those rows are not a complete seat count. A contract\n mismatch prints one plain comparison of the requested and served input/output digest pairs and\n advises aligning manager versions. The no-answer label means only that the instance is registered\n and did not answer. It does not say the host is down, because a dead host never deregisters itself and a\n live one can be slow; if it is gone, deregister it.\n `--on <instance>` pins the read to one exact instance id instead. A wrong pin fails loud\n rather than falling through: a well-formed id that no live manager carries is reported as\n `manager instance <id> did not answer` (nothing else is asked), and a credential without that\n instance's rail is reported as refused by the broker, not as an unresponsive manager. A manager\n that answers with a refusal is shown with its own cause; \"no manager reachable\" is said only when\n nothing answered at all. If the scatter's own registry read fails (the freeze or the reconcile),\n `ps` says the manager registry could not be read rather than pronouncing on the managers, which\n may all be up.\n\n**The verdict is scoped to the endpoint rail the request rode.** An issued caller rides the\nversioned `ep.v1` rail, a separate subject space from the legacy `ep` rail, and an endpoint serves\nboth (SPEC 13.15). A manager older than the versioned rail serves `ep` alone, so it can be running,\nregistered and answering while an issued caller's request reaches nobody. Silence on `ep.v1` is\nreported as `no manager answered on the ep.v1 rail` and names both causes it is consistent with:\nno manager running, or one older than the rail. The CLI cannot tell them apart, because the service\nregistry records no package version, so check whether a manager is running and, if it is, its\nversion. The same scoping applies to `cotal run`'s hosted verbs, which drop the `--local`\nsuggestion there, since `--local` drives the run from the calling process and names the caller as\nits answerer.\n\n**`stop` and `attach` route by seat locality.** A seat can only be stopped or attached by the\nmanager actually running it, and the class queue does not know which one that is. So on a\nstatic/open mesh both verbs first ask every registered instance which one hosts the named seat, then\naddress that instance directly. This happens by default; you do not need `--on`.\n\n`--on <instance>` remains the override, for when you already know where the seat lives or the\nlookup itself is degraded. It is also the **only** route on a **user-auth mesh**: a ledger-scoped\nbearer does not hold the registry-read rows the lookup needs, so there the verbs stay on the class\nqueue unless you pin them yourself.\n\nA seat is reported as **not found** only when every reachable instance answered for itself. An\ninstance that stayed silent past the deadline, or that refused the read rather than answering, said\nnothing about which seats it hosts, so the seat may be running on it. That case reports that the\nlocation could not be established, names the instances that did not answer, and states outright\nthat it is not a report that the seat is gone. Read it as unknown and retry with\n`--on <instance>`; a retry loop that treats it as \"already gone\" stops looking for a seat that is\nstill running. A single manager cannot tell \"hosted elsewhere\" from \"does not exist\": it answers\n`not-found` for both, which is why the search asks all of them and why an incomplete search\nconcludes nothing.\n- **User-auth mesh.** `cotal ps` reports what **one** manager knows about your agents (an `ep.one`\n read against the manager's in-memory roster, owner-filtered). It does **not** report other\n manager instances. It cannot tell you that one is down: an unreachable manager is absent\n from the list. Completeness across a multi-manager user-auth space is not claimed.\n A manager that does not answer fails the command outright (exit non-zero), rather than printing\n an empty list that could be read as \"no agents\". Your ledger row needs the `admin` scope to\n reach `ps` at all; `spawn` alone is refused by the broker (the ep tier boundary).\n\n`attach` streams and drives an agent's terminal on the `pty` runtime; detach with the escape key\n(Ctrl-] by default; see [`COTAL_DETACH_KEY`](config.md)). It does so over a one-use, holder-bound\nmesh session ([SPEC](../SPEC.md) §13.6): the manager replies with a signed session grant (never a\n`127.0.0.1` URL), the CLI redeems it once over the broker, and the browser console (`cotal console`)\ndrives the same session. `stop` and `attach` need a running manager to talk to. On a static mesh\nthey are cross-agent admin operations. On a user-auth mesh, your own agents (any agent under your\nowner) need only the `spawn` scope; another owner's agent needs `admin` on your ledger row\n([identity & auth](identity-and-auth.md)). Launch detached agents with [`spawn --detach`](#spawn).\n\n**`attach` reconnects when the link dies.** A session lives on a network link, and a laptop that\nsleeps, a VPN that drops or a wifi handover kills it. When that happens `attach` prints\n`[cotal: connection lost, reconnecting]` on stderr and starts asking the manager for a new session:\na fresh grant, a fresh per-session credential, a fresh connection, so every attempt re-runs the same\nauthorization the first attach did. On success it prints `[cotal: reconnected]`, the manager repaints\nthe seat's current screen the way it does for any attach, and you carry on in the same terminal.\nRetries wait 1s, 2s, 5s, 10s, then 30s, for as long as the seat exists. The detach key is read the\nwhole time the loop runs, the waits and the attempts alike, so a reconnect never traps you: press it\nwhile a session is being established and the attach ends there, and a session that lands behind the\npress is handed back to the manager rather than left holding a slot. Everything else you type while\nthere is no session is dropped rather than queued, so keystrokes aimed at a terminal that turned out\nto be frozen, Ctrl-C included, are not delivered to the agent by a reconnect you did not know had\nhappened. That starts before the first session, not at the first reconnect: at a terminal, `attach`\nreads and drops what you type while it is still resolving the mesh, so a key struck at a prompt that\nhas not come up yet does not reach the agent when it does.\n\nA **pipe** carries script input. For example, `printf 'ls\\n' | cotal attach --name web` is\nbuffered until the session opens. Buffering continues across reconnects, so\n`tail -f log | cotal attach --name web` does not lose the part of its feed written while the link was\ndown. Only a terminal gets the reader; `--no-reconnect` keeps the old behaviour on both.\n\nIt stops on its own when reconnecting cannot help, and says why: a manager that refuses the attach\nexits non-zero with the manager's own message, and a reconnect that finds the seat no longer there\n(despawned, or its agent exited while the link was down) exits cleanly with `seat <name> is gone`.\nA refusal that could still pass, such as a manager at its session ceiling, is relayed in the\nmanager's own words while the loop keeps trying, once per refusal rather than once per attempt.\nPressing the detach key, or the agent's process exiting while you are attached, ends the attach as\nit always did. `--no-reconnect` turns all of this off and restores the single-session behaviour,\nwhich is what a script wants.\n\nEach reconnect also hands the abandoned session back to the manager, over the first link that can\ncarry the message, so an attach that flaps does not eat the manager's session slots one outage at a\ntime. If that message never gets a link, the attach says so when it ends. The live-session ceiling\ndefaults to 64 concurrent sessions (`--max-sessions`); the browser console opens one session per\npane, so a dashboard over a large mesh should size for agents × panes. Hitting the ceiling refuses\nbefore a credential is minted and names `--max-sessions`.\n\nWhich mesh `attach` resolves also decides **how it redeems the grant**. On a registered open mesh\nthere is no local seed. The CLI connects bare, the same way other control commands already do, and\nthe session rail is the caller rail that a real open-mode connection already reaches. Telling the\noperator to re-register the root is false: the registered root is already the contract. On a\nstatic-auth mesh the grant is still redeemed by minting a short-lived\nsession-scoped credential from the seed at the root the mesh resolved to, never from a `.cotal`\nfound by walking up from whichever directory you happen to be standing in. The difference is not\nhypothetical: `~/.cotal` exists on every install because the mesh registry lives there, so a command\nrun anywhere under your home directory but outside a project used to mint from your home\ndirectory's trust and present it to a broker that trusts a different chain, which surfaced as a\nbare authorization failure that named nothing. A directory that does hold another chain for the\nsame space is now reported on the way past, and not obeyed:\n\n```text\n! this directory resolves to /Users/you, whose .cotal/auth holds a DIFFERENT trust chain for space \"team\".\n attach used /Users/you/projects/app, the root this mesh resolved to. The other one is not being used, and is worth a look.\n```\n\nWhen a **static-auth** mesh holds no seed at the resolved root, `attach` refuses and names what it\nresolved, the broker and the root, instead of describing a directory it did not use and instead of\ntaking the open-mode path. An authenticated registry entry with a missing seed is still\nauthenticated. A USER-AUTH mesh still refuses loud: two-step user-mode redemption is not wired.\n\nTerminal bytes stream over the mesh; the manager's own HTTP/WS face serves the console. That endpoint binds\n**loopback by default**, so nothing is exposed by accident; `cotal up --host <addr>` passes its bind\naddress down, which is what lets you attach to an agent whose manager runs on another machine. A\nbare `cotal supervise` and an embedded manager stay machine-local. Set it directly with\n`supervise --console-host <host>`.\n\nThat address is **recorded on the mesh** and carried forward, because it is a decision rather than\nsomething later commands can work out for themselves (a broker dial address is not a manager bind\naddress). Every later manager launch for the same mesh reuses it, including a same-root `cotal up` repair,\nan adopted preserved or restored listener, and a `spawn -f` manifest deploy. A manager replacement\ndoes not quietly move a reachable attach face back to loopback. Passing `--host` again overrides it,\nso you can widen or narrow exposure whenever you like; a mesh that never asked stays loopback-only\nand records nothing.\n\nBecause that face carries terminal read and write for every managed agent, it is credentialed in two\ntiers. A mesh caller receives a **ticket** bound to the single agent the manager just authorized,\nsingle-use and short-lived, so one authorized attach can never be re-pointed at someone else's\nagent. The **console token** is the operator's own, reaches every agent, and is printed only to the\nmanager's output. The roster, the live feed, and the PTY stream all answer `401` without one; the\nstatic console shell is served openly, since it describes no agent.\n\n## input\n\n```bash\ncotal input --name <n> --text <text> [--no-enter] [--on <instance>] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | | Managed agent to type into (required) |\n| `--text <text>` | | The text to type, taken verbatim (required) |\n| `--no-enter` | off | Type the text and stop there, without pressing Enter |\n| `--on <instance>` | class anycast | Pin to one manager instance id using the same rules as [`attach`](#managed-seats) |\n\nTypes one line into a running agent's terminal, as if you had typed it there, and returns. This is\nthe half of [`attach`](#managed-seats) that a program wants: `attach` is a live stream that holds a\nsession open and expects a terminal on your side, so a script, a cron job or a web UI cannot use it\nto send a single line. `input` is one authorized call.\n\nWhat it is for is **harness commands**. A line beginning with `/` is not chat and not a message: it\nis something the agent's own harness handles, and the only way in is the keyboard.\n\n```bash\ncotal input --name reviewer --text \"/compact\" # ask the harness to compact its context\ncotal input --name reviewer --text \"/model opus\" # switch its model\ncotal input --name reviewer --text \"hold on that PR\" # ordinary typing works too\n```\n\n**Quoting.** `--text` takes a value, so a payload starting with `/` survives as written. A payload\nstarting with a dash needs the `=` form, because the shell-style `--text --foo` is ambiguous and is\nrefused rather than guessed:\n\n```bash\ncotal input --name reviewer --text=--verbose # dash-leading text: use --text=<value>\n```\n\nEnter is pressed by default, since a command typed but never submitted has not been delivered.\n`--no-enter` types the text and leaves it sitting at the prompt, which is how you stage a line and\nsend it later.\n\nNothing comes back but a delivery receipt (`✓ sent 9 bytes to reviewer`, counting the trailing\ncarriage return). Whatever the agent does next shows up where its output already goes: the mesh, its\ntranscript, or an `attach`.\n\n**This one is operator-only, and more narrowly than `stop` or `attach`.** Those two are granted to\nanything holding `spawn`, so an agent can stop and attach to seats under its own owner. `input` is\nnot: it is granted only to operator credentials, which on a user-auth mesh means your ledger row\nneeds the `admin` scope, the same scope [`ps`](#managed-seats) already needs there. The reason is\nthat a write into a terminal is control of whatever is running in it, and on a user-auth mesh the\nown-owner rule covers every seat under you, not only the ones you launched: a `spawn`-scoped agent\ncould otherwise type into a sibling it never started. Seat locality is still resolved for you.\n\nOnly the `pty` runtime can be typed into. The external terminal runtimes (`tmux`, `cmux`, `orca`,\n`herdr`) attach to a process they do not own, so they have no input stream for it and the command\nrefuses by name rather than dropping the keystroke.\n\n## personas\n\n```bash\ncotal personas list [-v] [--running]\ncotal personas show <name>\ncotal personas edit <name>\ncotal personas new <name> (--prompt <t> | --from <f>) [--role <r>] [--model <m>]\ncotal personas rm <name> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh's persona catalog |\n| `--role <r>` | none | `new`: the persona's role |\n| `--model <m>` | none | `new`: the persona's model |\n| `--prompt <t>` | none | `new`: the persona's prompt text |\n| `--from <f>` | none | `new`: seed the prompt from a file |\n| `--verbose`, `-v` | off | `list`: include role / model / description |\n| `--running` | off | `list`: mark personas live on the mesh |\n| `--force` | none | `rm`: required, delete without prompting |\n\nPersonas are the local agent files under the resolved mesh root's `.cotal/agents/`, the same catalog\n`cotal spawn` launches from. `--space` and `--server` therefore move every list, read, write, delete\nand completion operation to the selected mesh. An unresolved target refuses rather than falling back\nto the current directory. See [Agent files](agent-files.md) for the file format.\n\n## supervise\n\n```bash\ncotal supervise [--runtime <name>] [--space <s>] [--server <url>] [--spawn <names>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space to supervise |\n| `--server <url>` | hosting mesh, or matching registered mesh | Broker URL. A registered mesh supplies it when omitted; a different explicit value is refused before anything is dialed. |\n| `--runtime <name>` | `pty` | Agent runtime (`pty` built in; extension runtimes are explicit-only) |\n| `--console-port <n>` | none | Protocol-console port |\n| `--console-host <host>` | loopback | Bind host for the console + attach endpoint. Loopback keeps it machine-local; `cotal up` passes the address it bound the broker to, which is what lets `cotal attach` reach this manager from another machine |\n| `--max-sessions <n>` | 64 | Live-session ceiling. Each console pane and each `cotal attach` is one session, so size for agents × panes, not agent count. A capacity refusal names this flag. `cotal up --max-sessions` records the same number on the mesh so a later `supervise` started by repair or `spawn -f` keeps it |\n| `--roster <file>` | none | Declarative roster to boot at startup |\n| `--launch <spec>` | none | Resolved manifest launch spec (from `up -f` / `spawn -f`) |\n| `--spawn <names>` | none | Comma-separated personas to pre-spawn at startup |\n\nThe manager is the agent supervisor and control plane: it answers `spawn --detach`, `stop`, `ps`,\n`attach`, and the `cotal_*` manager tools. `cotal up --detach` starts one for you; run `supervise`\ndirectly to recover a dead manager or drive a custom runtime. Default runtime is `pty`; install an\noptional provider first (`cotal ext add @cotal-ai/orca`, `@cotal-ai/tmux`, `@cotal-ai/cmux`, or `@cotal-ai/herdr`) and\nselect it explicitly. A missing provider or app fails loudly; there is no fallback. See [Deploy](deploy.md).\nBoot inventory decides whether this process takes unpinned `spawn`/`launch` on the class rail:\nif every declared connector is unavailable, those commands stay on this instance rail only\n(`status` reports `classSpawn: false`). `describe` still answers on the class rail, so an\nunpinned spawn can bind-fence against a skip member; re-issue, or pin `--on`. A partial\ninventory keeps the class rail and names `--on` on a harness refusal, because sibling\ninventories are not readable from the serve credential. See [control surface](control-surface.md#instance-routing).\n\nOn a normal `SIGINT`/`SIGTERM`, the manager stops every seat and requires the selected runtime to\nprove the seat is gone before it releases the manager lease or service registration. A stop that\ncannot prove exit fails loud and keeps manager authority instead of reporting a clean shutdown while\nan orphan still holds broker rails. After an abrupt manager death, the same logical successor\nterminalizes only its own durable static slots, verify-evicts the predecessor's broker principal,\nrecords that result in the lifecycle's caller-readable audit detail, reaps the predecessor's seat\nprocess through the runtime's custody reference recorded on the slot (the pty runtime verifies the\nprocess start identity in its seat record, so a reused pid is never signalled), and only then\nretires the lifecycle and frees the alias. A runtime that custodies its seats reserves that\nreference before it launches one, and the manager records it on the slot's first durable row, so a\nmanager that dies part-way through a spawn also leaves a seat its successor can address. A spawn\nthat launched its seat and then failed is rolled back by the manager that launched it, and that\nrollback reaps the seat through the same reserved reference before the lifecycle retires. Missing or unverified broker evidence keeps the slot\nterminalizing, and so does a runtime that cannot reap by reference.\n\nA `meshes add --mode user` entry is a **participant** registration, not hosting authority. A\nparticipant may run `supervise` only when the host advertises the remote manager authority service\nand the signed-in actor has the dedicated `supervise` ledger scope. The CLI obtains the closed,\nloopback-only `manager-service` view; `spawn` and `admin` do not substitute for that scope. The\nhost issues the manager's public-nkey JWT material through its lifecycle-bound prepare → activate\n→ renew protocol, never by handing the participant a signer or static provisioner credential.\n\nThe broker URL in the registry entry decides the transport. A remote broker is often published\nover a `wss://` edge rather than a raw `nats://` port, and `supervise` dials whichever scheme the\nrecord holds, starting with the manager-authority registration it runs before the manager exists.\nThe record also decides whether that registration requires TLS, so a participant never downgrades\nthe credential exchange to a plaintext connection the registry did not describe.\n\nWithout that advertised host service or scope, `supervise` refuses before it starts a manager.\nRun `cotal spawn` without `--detach` to launch a foreground agent, or ask the space host to enable\nthe authority service and grant `supervise` for detached agents. If a running remote manager loses\nrenewal, it reports degraded state and refuses unsafe new starts and restarts; live agents are not\nsilently replaced. Do not run `cotal down` or `cotal up` on a participant machine to repair this\ncondition.\n\n## service\n\n```bash\ncotal service install [--mesh <name>] [--linger]\ncotal service status [--mesh <name>] [--json]\ncotal service uninstall [--mesh <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--mesh <name>` | this folder's mesh | The mesh whose manager the service runs; one unit per mesh |\n| `--linger` | off | install: also enable user lingering so the user manager starts at boot and the service survives logout. Never enabled silently |\n| `--json` | off | status: machine-readable output |\n\nRuns the manager as a user service so it survives logout and reboot. On Linux this installs a\nsystemd user unit (`~/.config/systemd/user/cotal-manager@<key>.service`, where `<key>` is the\ncase-safe mesh key); on macOS a launchd agent plist under `~/Library/LaunchAgents/`. Any other\nplatform, or an absent systemd/launchd user session, fails with a message naming what is missing.\n\n`install` resolves the mesh from the registry and binds the unit to that entry's root and broker\naddress, so it can be run from any directory. The mesh must be registered (`cotal up` or\n`cotal meshes add`) before installing; an unregistered name refuses before anything is written.\n\nThe unit's `ExecStart` is the bare `supervise` command. The mesh facts travel in a `0600`\n`EnvironmentFile` (`COTAL_SPACE`, `COTAL_SERVER` pinned to the registered broker URL, whatever\nport it listens on) rather than the command line, because command lines are readable by every\nuser on a multi-user host. The same file gives the service a private `COTAL_HOME` and\n`XDG_CONFIG_HOME` under the unit directory, so the service manager never touches the login\nuser's `~/.cotal`. First-run connector seeding runs synchronously inside `service install`,\nagainst that private config root; the unit itself starts with `COTAL_SKIP_CONNECTOR_SEED=1`\nso a manager is never interrupted mid-seed by a restart. An install whose pre-seed cannot\ncomplete (network unreachable, registry error) refuses instead of deferring.\n\nEvery value the unit derives from a path (`WorkingDirectory`, the `EnvironmentFile` path, the\n`ExecStart` tokens) is escaped for systemd specifiers (`%` becomes `%%`), so a mesh root that\ncontains `%` starts over its real path instead of a path systemd rewrote by expanding it. The\nprovenance comment records the root unescaped.\n\n`service install` also refuses while a manager is already running for the mesh (`cotal down\nmanager` first). The restart policy is `Restart=always` with `RestartSec=20s`, chosen for\nmanager units in production: a manager exits for reasons that are not failures (broker\nrestarts, host suspend), where `on-failure` with a short interval thrashes.\n\n`service status` reports the unit state from systemd/launchd, the manager's own health read from\nits pidfile at the unit's recorded root, and the machine facts a hosting side asks for:\narchitecture, OS (the platform, never the hostname), whether `/dev/kvm` is present and\naccessible, CPU count, and total memory. `--json` returns the same fields as one object.\n\n`service uninstall` stops and disables the unit and removes it plus the private state directory.\nIt works from any directory: the unit's own records name the mesh and root it serves, and an\nexplicit `--mesh <name>` selects it. It refuses any unit that was not written by `service\ninstall` (the files carry a provenance comment), whose recorded mesh is missing, or that was\ninstalled for a different mesh, so operator-written units are never destroyed; `service status`\napplies the same rule and never reports a mesh a unit does not record.\n\nThis command installs only the manager. The per-space auth service and the delivery daemon are\nnot installed by it: on a shared broker an operator runs three units per space with `After=`\nedges (auth service, then manager, then delivery) and stops them in reverse. A broker-side `cotal\nup` unit is a separate unit documented in [Run a mesh](run-a-mesh.md).\n\n## reconcile-gate\n\n```bash\ncotal reconcile-gate [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the frozen gate lives in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint whose gate is frozen |\n| `--instance <id>` | this folder's persisted manager instance | Instance id |\n\n**When you need this.** A manager restart killed after deregistration begins but before the new\nincarnation finishes leaves the endpoint's issuance gate *frozen*, held by a\nprocess that no longer exists. The freeze is what stops two incarnations serving at once, which is\ncorrect. The successor manager now completes that dead registration itself on boot, using the same\nguard this command uses: it acts only when the freeze-holder is affirmatively gone under a complete\nCONNZ sweep (`gone` and `sweepComplete=true`). If that registration's spec write already committed,\nit finishes the same freeze at the committed registration revision. If the spec did not advance, it\nabort-reopens the gate at generation+1 with processEpoch unchanged and continues the normal takeover.\nLive, unknown, unestablishable, and\nwrong-op-kind still refuse; there is no TTL.\n\nUse this command when the boot path cannot run: the delivery daemon is down, the repair targets a\nnon-manager endpoint, or you want to lift the freeze without starting a manager. It checks that the\nholder really is gone, prints what it found, and then finishes the dead operation the same way as the\ninterrupted restart would have: revoke the old credentials, evict their holders with verification,\nand reopen the gate.\n\nIf verification is interrupted, the command leaves the gate frozen and durably records each holder\nwhose eviction was already verified. A retry still repeats the freeze-holder liveness check, then\nskips only progress bound to the same registration operation, frozen-gate revision, and holder set.\nThe output reports holders completed before this attempt, completed now, and still remaining. A new\nfreeze or changed holder set starts from zero. Cursor cleanup happens only after reopen; a retained\ncursor is harmless because its old gate revision cannot authorize a later freeze.\n\n**It refuses far more often than it acts, on purpose**, and always says which check stopped it:\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `holder-alive` | The freeze-holder still has a live connection: a manager *is* running | Stop that process first. Reconciling would evict a live manager's credentials |\n| `holder-unknown` | The connection sweep could not prove the holder absent | Not safe to proceed: an unprovable holder is treated as a live one. Re-run once the broker answers completely |\n| `liveness-unestablishable` | The delivery daemon could not be asked at all | Start it (`cotal up` runs it) and re-run. Silence is never read as death |\n| `not-frozen` / `no-gate` | The gate is open, or there is no gate at that coordinate | Nothing to repair: check `--endpoint` / `--instance` |\n| `wrong-op-kind` | Frozen under a takeover or retirement, not a registration | Out of scope for this command; it will not reinterpret another operation's intent |\n| `eviction-unverified` | The holder looked gone but eviction could not be verified | The gate is left frozen, unchanged. Investigate the broker before retrying |\n| `raced` | A newer manager moved the gate mid-repair | Re-run `cotal doctor` and look again |\n\nThere is no `--force`, and no path that discards gate state: the only way this reopens a gate is by\nproving the holder is gone and then completing the operation properly.\n\n**What reopening the gate does for the endpoint's governance slot.** A registration takes the\nendpoint-wide governance slot before it publishes its spec, and holds it until its gate reopens. An\ninstance that died between those two points leaves the slot held with no registration behind it.\nThis command does not write that slot and never has; the registration path is its only writer. What\nthe reopen does is advance the holder's gate past the generation the slot is stamped with, which is\nwhat marks the slot abandoned. The next registration for that endpoint then reclaims it as part of\nits ordinary start. So the repair here is still one command followed by starting the manager, and\nthe slot needs no separate step.\n\n## deregister-instance\n\n```bash\ncotal deregister-instance [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the instance is registered in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint the instance serves |\n| `--instance <id>` | this folder's persisted manager instance | Instance id, the whole id as `cotal ps` prints it |\n\n**When you need this.** The service registry records *registration*, not liveness, and nothing in\nthe model expires a row. A manager that stops cleanly removes its own registration. One whose host\ndied without writing anything cannot, so its record goes on claiming a live instance forever: every\nclass scatter in that space freezes the dead slot in, and `cotal ps`, `stop` and `attach` each pay\ntheir whole deadline waiting for a machine that is never coming back. A laptop that was reimaged, a\ncontainer that was deleted, a box that will not be back on the network: those registrations have no\nother exit.\n\nThis command is that exit. It asks the instance first, and it removes a record only when the broker\naffirms the instance's own rail is empty: nothing subscribed there. Then it deletes the\nregistration's two records keys, each pinned to the revision it read, and prints what it removed.\n\n**Silence alone never passes.** An unanswered describe is what a dead host, a wedged process and a\nslow one all look like, and a hung process still holds its subscriptions, so the broker sees\ninterest on its rail. That instance is refused and the observation is printed. A dead process holds\nno connection and therefore no subscription, so a real corpse is still removed.\n\n**Every refusal names the failed check:**\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `instance-answered` | The instance answered a pinned describe. It is alive | Nothing to repair. If it is wedged rather than gone, stop the process first; its own clean stop removes the record |\n| `instance-not-affirmed-gone` | It did not answer, and the broker did not report its rail empty, which is what a held subscription looks like: slow or hung, not affirmed gone | Nothing was removed. Stop the process; its record goes on its own clean stop, or re-run this once it is down |\n| `liveness-unestablishable` | The probe itself failed, so nothing was learned | Fix the probe's path (credential, broker) and re-run. A probe that could not run is never read as death |\n| `not-registered` | No registration at that coordinate | Check `--instance` and `--endpoint`. This takes the whole id, never a prefix |\n| `registration-in-flight` | The instance holds the endpoint governance slot at the live issuance-gate generation, so a registration is still completing | Nothing was removed. Wait for that registration to finish, then re-run |\n| `superseded` | The record moved between the read and the delete | Something is writing to it. Nothing was removed; re-observe before retrying |\n\nThere is no `--force` and no sweep: silence is not death, and a rule that removed rows on silence\nwould eventually remove a live instance that was merely slow. An operator names one instance, the\nbroker's verdict on its rail is what authorizes the removal, and the guard's job is to show them\nthey named a dead one. Removal is not a one way door either. The same instance re-registers over\nthe tombstone on its next start, under the same identity.\n\n## runtimes\n\n```bash\ncotal runtimes\n```\n\nLists every agent runtime the manager can spawn through: the built-in `pty`, the official providers\n(`orca`, `tmux`, `cmux`, `herdr`), and any custom provider installed via `cotal ext add`. Each installed\nprovider is probed so you can see what is actually reachable on this machine before selecting it:\n\n```\npty built in\norca installed · reachable @cotal-ai/orca\ntmux available · cotal ext add @cotal-ai/tmux\ncmux available · cotal ext add @cotal-ai/cmux\nherdr available · cotal ext add @cotal-ai/herdr\n```\n\n`installed · reachable` / `unreachable` is the provider's own `available()` probe; `available` means\nit is a known runtime you can add with the shown command. Selecting an unknown or uninstalled runtime\nvia `up`/`spawn --runtime <name>` fails loud and, for a known one, points at the exact `cotal ext add`\npackage. There is no silent fallback to `pty`.\n\n## send\n\n```bash\ncotal send dm <agent> \"<text>\" [--space <s>] [--server <url>] [--creds <path>]\ncotal send msg <channel> \"<text>\"\ncotal send ask <role> \"<text>\"\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and (off-registry) which credential |\n\nOne-shot messaging: connect, send a single direct message (`dm`), channel post (`msg`), or role\nask/anycast (`ask`), then exit. For a running conversation, agents use the mesh tools instead\n([MCP tools](mcp-tools.md)).\n\n`cotal send` works from an operator shell or from a seat. It uses `cotal-send` as the advisory\ndisplay name. The wire principal comes from the resolved operator credential or user bearer, not\nfrom `COTAL_NAME`, `COTAL_ID`, `COTAL_OWNER`, or `COTAL_ACTOR`. On an open mesh the transient\nendpoint self-mints its principal.\n\n## channels\n\n```bash\ncotal channels list\ncotal channels set <name> [--replay | --no-replay] [--window <n>] [--desc <s>] [--instructions <s>]\ncotal channels default --replay | --no-replay\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--replay` / `--no-replay` | none | `set`/`default`: replay history to new joiners, or not |\n| `--window <n>` | none | `set`: replay window size |\n| `--desc <s>` | none | `set`: one-line channel description |\n| `--instructions <s>` | none | `set`: instructions shown to joiners |\n\nInspects and edits the channel registry: replay policy, description, and joiner instructions. ACL\nsemantics (who may read or post) are set at mint / provision time, not here; see\n[Channels and permissions](channels-and-permissions.md). On a user-auth mesh, `list` rides your\nown login as is; `set` and `default` edit the registry over a short-lived\nchannel-writer view, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\nOn a remote user-auth mesh that view is served by the public exchange; space-history `purger`\nand the read-only admin view are not.\n\n\n## history\n\n```bash\ncotal history clear --force [--dms] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--dms` | off | Also clear DM history |\n| `--force` | none | Required: clear without prompting |\n\nPurges retained channel history; `--dms` extends it to direct-message history. An alias of\n[`clean history`](#clean). On a user-auth mesh the purge rides a short-lived purger view over\nyour login, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\n\n## console\n\n```bash\ncotal console [--plain] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to watch |\n| `--plain` | off | Line stream instead of the TUI |\n\nA live protocol view for a space: a lazygit-style TUI, or a plain line stream on `--plain`. On a\nuser-auth mesh it rides the read-only admin view over your login, which needs ledger scope\n`admin`. Inside the TUI, operator control (`D` kill, `:spawn`, `:status`, `:purge`) rides the\nsame per-action instrument path as `cotal stop` and `cotal ps`, never the observer; a raw\n`--creds` file cannot drive it. `a` (or `:attach <agent>`) runs\n[`cotal attach`](#managed-seats) in place and returns to the console on detach. See\n[Watch a mesh](watch-a-mesh.md).\n\n## web\n\n```bash\ncotal web [--detach] [--host <host>] [--port <n>] [--no-open] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to serve |\n| `--host <host>` | `127.0.0.1` | Concrete HTTP bind and browser host; wildcard addresses are refused |\n| `--port <n>` | `7799` | HTTP port |\n| `--detach` | off | Run in the background; stop with `cotal down web` or bare `cotal down` |\n| `--no-open` | off | Don't open the browser |\n\nThe browser observability dashboard: presence, channels, and a live feed. It is **not** part of\n`cotal up`: it ships inside `cotal-ai` as the `@cotal-ai/web` extension, seeded automatically on first\nrun (like the built-in connectors) so it always matches your CLI version. It self-registers `cotal web`\ninto this surface and serves\n`http://cotal.localhost:7799` by default (loopback; `*.localhost` resolves in Chrome/Firefox/Edge; Safari may\nneed `http://127.0.0.1:7799`). On a user-auth mesh the dashboard rides the read-only admin view\nover your login, and a channel purge asks for its own channel-purger view per click; both need\nledger scope `admin`. The public exchange serves `channel-purger` for a remote owner; it still\nrefuses the startup admin view, so a remote `cotal web` is not a complete channel-management\nsurface. Detached mode re-execs the current Cotal installation, writes diagnostics to\nthe mesh root's `.cotal/web.log`, and reports success only after the HTTP server answers. It requires\na recorded mesh root, but can be launched from any directory once `cotal up` has recorded the mesh.\nSee [Watch a mesh](watch-a-mesh.md).\n\n## mint\n\n```bash\ncotal mint <name> [--profile <agent|observer|admin>] [--out <path>] [--signer]\ncotal mint <name> --provision [--role <role>] [--space <s>] [--server <url>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--profile <agent\\|observer\\|admin>` | `agent` | Credential profile |\n| `--out <path>` | `.cotal/auth/creds/space.<key>/<name>.creds` | Output path - the default sits under the resolved space's segment (`<key>` is that space's hex encoding, as in [Project files](config.md#project-files)) |\n| `--signer` | off | Emit a stripped account-signing file instead |\n| `--force` | off | With `--signer`: overwrite an existing file |\n| `--allow-subscribe <a,b>` | the agent file's, else subscribe | Read-ACL override, **agent profile only**: `observer` and `admin` carry a fixed read set, and `mint` refuses this flag there rather than narrowing nothing |\n| `--allow-publish <a,b>` | the agent file's, else deny | Post-ACL override, **agent profile only** |\n| `--role <role>` | the agent file's | Agent profile: the anycast task queue the identity pulls (`svc_<role>`) |\n| `--provision` | off | Agent profile: also pre-create the identity's bind-only DM/deliver durables (and its role's task queue) on the live mesh, so the credential can consume |\n| `--space <s>`, `--server <url>` | the resolved mesh | Which root supplies the agent file, static trust and default credential storage; with `--provision`, also which live mesh receives the durables |\n\nMints a NATS creds file for a space in **static** auth mode, scoped to a profile and (optionally)\nexplicit read/post ACLs. `--signer` emits an account-signing file for delegating minting to another\nhost. A per-user-auth space refuses `mint`: agents there join under a logged-in user\n([`login`](#login) + [`actor grant`](#actor)), never via a handed-out creds file. See\n[Identity and auth](identity-and-auth.md).\n\nFor an agent profile, the resolved mesh root supplies the persona ACL, the signing material and the\ndefault credential destination as one authority. If the current folder also holds trust for a\ndifferent space or account, mint refuses before writing and names both roots. It never combines a\npersona from one root with credentials signed or stored under another.\n\nA plain mint is creds only: the identity can publish within its post ACL at once, but on an authed\nmesh its DM inbox and task queue are provisioner-pre-created and bind-only, so a **consuming**\nconnect fails until they exist. `--provision` performs that pre-create in the same command (a\nprovisioner cred is minted from the space's trust material, used, and dropped), so a long-running\nclient you start yourself can receive DMs and role anycasts like a spawned seat. The command prints\nthe identity's principal (its wire id) and lifecycle uid; a consuming client passes that uid as its\n`lifecycleUid`. Agent profile only; an open mesh needs none of this (peers self-create there). The\nThe same resolved authority is used for both the credential and `--provision`, so the broker\nfootprint cannot be created under a different root's trust material.\n\n## Login\n\n```bash\ncotal login --idp <auth base URL> [--client-id <id>]\ncotal logout --idp <auth base URL>\n```\n\nSigns you in to a per-user-auth mesh's IdP (device code flow) and caches the session; run it\nonce per machine. It prints your IdP subject, the id the operator grants against. When the trusted\n`/token` response advertises a same-origin space catalog, login validates and records that account's\nspaces immediately. After a\nlogin, every command on that mesh works under your identity: each connect takes a fresh IdP\nproof, exchanges it locally for a short-lived bearer, and is authorized against the actor\nledger at connect time. `logout` revokes the IdP session, clears its cache, and removes only that\naccount's discovered registry entries. See\n[identity & auth](identity-and-auth.md).\n\n## actor\n\n```bash\n# an upsert of the WHOLE row: a flag left off is the WIDE default below, not \"unchanged\"\ncotal actor grant <actor> --sub <IdP subject> [--scope a,b] [--allow-subscribe a,b] [--allow-publish a,b] [--role <r>] [--label <l>]\ncotal actor revoke <actor> (--sub <IdP subject> | --owner <u_…>)\ncotal actor list\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | the folder's | Space whose ledger to manage |\n| `--sub <subject>` | none | The IdP subject (shown by `cotal login`) the actor belongs to |\n| `--owner <u_…>` | none | The derived owner token (alternative to `--sub`) |\n| `--scope <a,b>` | `spawn,role:default` | Capability scope (`''` = none; `spawn` = may run agents; `role:<r>` = may delegate role r; `admin` = cross-agent control; `supervise` = eligible for the closed remote manager-service view when the host enables it) |\n| `--allow-subscribe <a,b>` | `>` (all channels) | Channel read ACL; the user's envelope, their agents can never read beyond it |\n| `--allow-publish <a,b>` | `>` (all channels) | Channel post ACL; also the envelope for their agents' posting |\n| `--role <r>` | none | Role (scopes the task-queue consumer) |\n| `--label <l>` | none | Display label for `actor list` (never the IdP subject) |\n\nThe actor ledger is the single authorization source of a user-auth space: no row, no access.\nA bare `grant` is the **full** envelope (all channels, may spawn); the flags narrow it. A\nre-grant **replaces the whole row**, not the one field you name, so to add a capability spell\nevery field out: the new scope plus the row's current read set, post set, role and label\n(`cotal actor list` shows what a row holds). A field left off does not stay as it was, it\nreverts to the wide default in the table above, which is how a narrow reader becomes a reader\nof every channel. A re-grant retires the current interactive lifecycle through the running auth\nservice before it rotates the row, so copied bearers cannot cross an authorization update. If that\nretirement cannot be confirmed, the row is left unchanged and the command fails with the recovery\naction. `revoke` uses the same retirement before deleting the row, which lets a later grant create a\nreal successor instead of colliding with a live predecessor. `supervise` is separate from `spawn` and `admin`: it only makes a signed-in\nperson eligible for the host-provided closed remote manager-service view; it does not grant\nmanagement of another owner or a general host profile. `revoke` denies the next exchange and\nthe next connect with no restart, and evicts the principal's live connections. Managed-agent rows\n(written by the spawn path) live in a disjoint row space this command never touches. See\n[identity & auth](identity-and-auth.md).\n\n## doctor\n\n```bash\ncotal doctor auth [--fix]\n```\n\nCredential-health diagnosis and repair for this folder's mesh: renders every managed\ncredential as healthy / near-expiry / expired and ends in `healthy` or the exact next\ncommand; `--fix` applies the repairs it can. The one surface every stale-credential error\npoints at.\n\n## join\n\n```bash\ncotal join --space <s> --name <n> [--role <r>] [--channel <c>]\ncotal join --link <url> | --token <t>\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and which credential |\n| `--name <n>` | none | Your presence name |\n| `--role <r>` | none | Your role |\n| `--channel <c>` | none | Channel to join |\n| `--kind <k>` | `agent` | Endpoint kind |\n| `--link <url>` | none | Join link (`cotal://…`) |\n| `--token <t>` | none | Join token |\n| `--lifecycle-uid <uid>` | none | Required with `--creds`: the lifecycle UID minted alongside the credential (`COTAL_LIFECYCLE_UID` works too). A credential's durable grants name exact lifecycle-keyed resources, so `join` refuses to invent one |\n| `--tls` | off | Connect over TLS |\n\nAn interactive presence: join a space under your own name and role, without launching an agent\nharness. A `--link` or `--token` supplies the where and the auth in one value. See\n[Spaces](spaces.md) and [Identity and auth](identity-and-auth.md).\n\n## Manifest deploys\n\nA `cotal.yaml` manifest declares a whole mesh (channels, personas, roles, and ACLs) in one file.\nThree commands consume it, plus a read-only validator:\n\n```bash\ncotal up -f cotal.yaml # boot a fresh mesh from the manifest\ncotal spawn -f cotal.yaml # deploy the manifest additively onto a running mesh\ncotal down -f cotal.yaml # tear that deploy down (or --run <id> for one run)\ncotal topology view -f cotal.yaml # validate + view the access graph, change nothing\n```\n\n`up -f` and `spawn -f` differ in target: `up -f` brings up a new broker and applies the manifest;\n`spawn -f` requires an already-reachable mesh and applies additively (ownership-scoped). On a\nuser-auth mesh, `spawn -f` deploys over your own login (the deployer view, gated on ledger scope\n`spawn`): the manifest's agents land under your owner, a manifest claiming another owner is\nrefused, and seeding new channels additionally needs scope `admin`. Both take\n`--dry-run` to print the plan without mutating anything. `topology` validates the manifest and\nrenders its channel / role / ACL graph. See [Define a team](define-a-team.md) and the\n[manifest reference](manifest.md).\n\n## ext\n\n```bash\ncotal ext # same as `list`\ncotal ext add <npm-package>\ncotal ext remove <name>\ncotal ext list\ncotal ext root # print just the install prefix (scriptable)\ncotal ext seed [--repair|--reset|--force]\n```\n\nOperator-installed extensions: `add` installs an npm package into a cotal-owned prefix and records\nevery registry provider it contributes. Commands appear in help, completion, and dispatch; runtime\nproviders are lazy-loaded by commands such as `supervise`; local process providers participate in\n`status` and selective `down`. `remove` and `list` manage them. The `@cotal-ai/web` dashboard is the\ncanonical command/process example. Installed packages and their location are described in\n[config](config.md).\n\nBare `cotal ext` lists the inventory, headed by the install prefix. That prefix is a cotal-owned npm\nroot kept **separate** from npm's own global tree. These packages never show up in `npm list -g`,\n`cotal ext` (or the Extensions section of `cotal status`) is the canonical inventory. `cotal ext root`\nprints only the path, for scripts. The versions shown are the manifest pin recorded at add time.\n\nRemoving an extension that owns a running local process is refused with the mesh root and its\n`cotal down <component>` command; stop it first so uninstalling the package never strands a process\nwhose lifecycle provider is gone.\n\n### Built-in connectors are seeded extensions\n\nThe first-party agent connectors (`claude`, `opencode`, `codex`, `hermes`, `jcode`, `pi`) are not compiled into\nthe binary. They are seeded on first run through the **same** `ext add` path a third party uses, and\nappear in `cotal ext list` like any other extension. So you can remove one you do not want\n(`cotal ext remove @cotal-ai/connector-hermes`), and a deliberately-removed connector STAYS removed\nacross upgrades. `cotal ext add <your-package>` adds a third-party connector the same way. The web\ndashboard (`@cotal-ai/web`, providing `command:web`) is the seventh built-in seeded on the same path.\n\n`cotal ext seed` is the maintenance entry for that seeding (it runs automatically on the first real\ncommand of each boot, so you rarely call it):\n\n| Flag | Meaning |\n|---|---|\n| (none) | Reconcile: seed any never-seeded built-in, refresh a seeded one whose version the binary bumped, leave a removed one removed. A no-op once current. |\n| `--repair` | Recover after an interrupted seed or a lost authority (rebuilds the interrupted connector; restores the removed-vs-never-seeded record from its durable backup). |\n| `--reset` | Discard the record and re-seed all seven built-ins (the six connectors plus the web dashboard). **Resurrects any you removed.** Rebuilds cleanly over corrupt seed state. |\n| `--force` | Re-seed the built-ins even when the version stamp is current or a downgrade. |\n\nWhen a newer `cotal` advances the operator-global seed store to its generation, it prints one\nmigration line naming the old and new generations, the exact CLI entry that wrote the store, the\ncommit timestamp, and `seed/stamp.json`. That writer and timestamp are kept in the stamp, so a later\nolder CLI refusal can say which executable wrote the generation it will not overwrite and when.\nLegacy generation-only stamps remain readable; their refusal simply has no writer provenance to add.\n\nAn older `cotal` refuses a seed store written by a newer version. When it can verify a sufficient\n`cotal` executable on PATH or at the installer's `~/.local/bin/cotal` location, the refusal names\nthat absolute path so a reduced service PATH does not select the older binary again. Otherwise it\nkeeps the generic newer-version instruction. `--force` rebuilds the store for the running older\nversion without discarding the ever-seeded authority. `--reset` still exists for corrupt state and\nresurrects deliberately-removed connectors.\n\nA source-checkout CLI (`pnpm cotal`, `tsx bin/cotal.ts`, `node bin/cotal.ts`, or a suite child of\nthose) refuses to write or garbage-collect that store. The refusal names the path, the generation\nit declined, and `$XDG_CONFIG_HOME` as the isolation remedy. `COTAL_HOME` does not relocate this\nstore. An entry that cannot be proven as a released install is refused the same way. Isolated\nrelease tests that must seed from a checkout-shaped `bin/` set `COTAL_ALLOW_CHECKOUT_SEED=1` after\npointing `$XDG_CONFIG_HOME` at a scratch dir; that override is documented here, not on the refusal\nline. An opt-in write still records the checkout path in `seed/stamp.json` as `writtenBy`.\n\nThe default connector for a bare `cotal spawn` (no `--agent`) is the persona's `agent:` pin if it\nhas one, else `claude`; set `COTAL_DEFAULT_AGENT` (e.g. `opencode`) to change the fallback. It is\na default, so a persona that pins its harness still wins over it. An `--agent` naming a removed\nconnector fails loud with the exact\n`cotal ext add` to restore it. Set `COTAL_SKIP_CONNECTOR_SEED=1` to turn off the automatic first-run\nseed/refresh entirely (for a controlled or offline setup that manages connectors by hand); `cotal ext\nseed` still runs on request. `cotal agent-bearer` never takes the seed at all: it is exec'd by\nspawned seats on every bearer refresh, so it neither reconciles nor is refused by the store's\ngeneration (see [Plumbing](#plumbing)).\n\n## completion\n\n```bash\ncotal completion <bash|zsh|fish|powershell> # print a stub to eval / source\ncotal completion install [shell] # install it persistently\n```\n\nPrints or installs shell completion. Completion candidates come from each command's declared flags\nand, where useful, live mesh state (spaces, personas, managed agents) resolved offline.\n\n## feedback\n\n```bash\ncotal feedback \"<summary>\" [--type <t>] [--email <e>] [--details <text>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--type <t>` | none | `bug` \\| `idea` \\| `friction` \\| `praise` \\| `other` |\n| `--details <text>` | none | Longer free-form details |\n| `--severity <s>` | none | `low` \\| `medium` \\| `high` |\n| `--area <a>` | none | The part of Cotal this concerns |\n| `--email <e>` | git email | Contact email (required on the keyless public path) |\n| `--name <n>` | none | Your name (optional) |\n| `--url <url>` | keyed / public intake | Intake URL override |\n| `--key <k>` | `COTAL_FEEDBACK_KEY` | Feedback key |\n\nSends feedback to the Cotal developers. With a key (`--key` / `COTAL_FEEDBACK_KEY`) it routes to the\nkeyed beta intake; without one it goes to the public `cotal.ai` intake and requires a contact email\n(`--email` / `COTAL_FEEDBACK_EMAIL`, else your git email). Run a self-hosted intake with\n[`feedback-intake`](#server-daemons).\n\n## run\n\nOperate durable workflow runs (cotal-lang programs) from the terminal.\n\n```bash\ncotal run start --file <program> [--timeout <dur>] [--local]\ncotal run resume <runId> [--local --file <program>]\ncotal run ps [--endpoint <ep>]\ncotal run journal <runId> [--endpoint <ep>]\ncotal run answer <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]\ncotal run migrate <runId> --local --file <program> [--endpoint <ep>]\n```\n\n`start` hands the program to the mesh's manager, which validates it, mints the run id (the record\nnever takes a caller-supplied one), drives it in its own process, and answers with the id once the\nrun is recorded; a program that does not validate is refused with every problem listed. `resume`\nasks the manager to take an existing run back and continue it from its step journal; the source is\nthe recorded program, so no `--file` is taken. Neither takes `--endpoint`: the manager records\nits runs under its own endpoint, and naming another is refused. `ps` lists the run records and\n`journal` renders one run's durable records; both only inspect. An open pause prints its question.\nA pause settled with an accepted answer prints its value as JSON plus the recorded answerer,\nartifact when present, time, and answer id. Expired pauses and ordinary steps print no answer line.\n`answer` resolves an open\ncheckpoint through the manager, presenting as the holder that armed it; the manager records the\nanswerer from your credential, so no `--by` is taken there. `migrate` runs the migrate check of an\nedited program against a run's journal, from this terminal under a read credential (`--local`\nonly; the manager serves no run-migrate command): it prints whether the migration is admissible,\nevery orphaned step with its verdict and code, and exits 0 on admissible and non-zero on not. It\nwrites nothing: the commit that would file the migration is not reachable yet, and the report\nsays so. `--timeout` sets the default\ncheckpoint timeout for a drive (default 1h). `--local` drives in this process instead, over one\nconnection per invocation under the run's own credential minted from the project folder's trust\nmaterial, and is the path on a bare broker with no manager or for a run with no recorded program\n(`cotal run resume <runId> --local --file <program>`); `answer --local` takes `--by <who>`. A\nuser-auth mesh runs no programs yet: the manager refuses the family by name, and `--local` has no\ncredential there. The guide is [workflows](workflows.md).\n\n## Server daemons\n\nTwo long-lived infra roles ship with the CLI. They are not part of everyday operation; the delivery\ndaemon comes up automatically with `cotal up --detach` in auth mode.\n\n```bash\ncotal deliver --space <s> [--server <url>] [--creds <file>]\ncotal auth-service --space <s> --server <url> [--port <n>] [--exchange-public-port <n>] [--exchange-public-url <https://…>] [--exchange-trusted-proxy]\ncotal feedback-intake --keys <keys.json> [--port <n>] [--creds <file>]\n```\n\n`auth-service` runs a user-auth space's identity plane: the NATS auth callout, the\ncapability-gated local exchange and JWKS, and, when `--exchange-public-port` is set, the closed public\nexchange/discovery face forwarded by an HTTPS reverse proxy. `--exchange-public-url` is the proxy URL\nadvertised to clients; `--exchange-trusted-proxy` opts into last-hop `X-Forwarded-For` attribution.\n`cotal up --user-auth` starts and supervises the service for you, so you run it directly only to\nrecover one by hand.\n\n`deliver` runs the server-side Plane-3 delivery daemon: the durable backstop and membership/ACL\nauthority. It is auth-mode-only and single-instance (`--shard`/`--shards` accept only `N=1`);\n`--dev-mint` mints a scoped cred from the local signer for standalone dev. `--creds` can start a\ndaemon that already looks healthy, but production renewal is not that file alone: the manager and\nthe daemon must address one credential store. On a stock split host with two project roots, a\ndirect `deliver` is not an independent repair; keep the daemon under `cotal up` on the broker\nhost, or inject the same store into both processes ([embedding](embedding.md#supervisor-signing-authority)).\nSee the [delivery daemon](delivery-daemon.md). `feedback-intake` runs a self-hosted feedback server\n(requires `--keys` and a scoped `--creds`), announcing submissions into a space channel; flags\ninclude `--host`/`--port`, `--store`, `--space`/`--channel`, `--max-bytes`, and `--rate-limit`.\n\n## Plumbing\n\n`cotal __complete <words…>` is the internal entry the shell-completion stubs call to emit candidates\nfor the current command line; you never run it directly. `cotal agent-bearer` is machine-facing\nplumbing on user-auth meshes: spawned agents exec it to print a fresh short-lived bearer from their\nspawn-time secret; you never run it directly either. Its local arm uses `--dir` to discover the\ncapability-gated loopback service. A remotely enrolled, already-granted agent instead receives\n`--exchange-url <https://base>` in its launch argv: that arm sends `{owner, actor, actorToken}` to the\npinned public exchange with no local capability, follows no redirects, and refuses every non-HTTPS\nURL because the actor token is the credential in the request body. Because a seat execs it on every\nbearer refresh, it skips the connector-seed boot gate entirely: it reads one 0600 token file,\nexchanges it and prints the bearer without consulting or writing the operator-global seed store, so\na newer store generation cannot refuse a live seat's refresh. (`cotal start` is a removed tombstone: it\nerrors and points you to `cotal spawn --detach`.)\n"
73
+ "body": "# `cotal` CLI reference\n\n> **Reference**: describes the TypeScript reference implementation (the `cotal` CLI), not the wire contract. · **For:** operators · **Wire contract:** [SPEC](../SPEC.md)\n\n`cotal` is the operator command line for the reference implementation: bring a mesh up, mint\nidentities, launch agents, watch what they do, and tear it all down. It is a thin client over the\nwire contract: the normative subjects and schemas live in the [SPEC](../SPEC.md); this page is\nlookup material for the commands, not a walkthrough; if you are new, start with\n[Getting started](getting-started.md).\n\n## Running it\n\n```bash\nnpm install -g cotal-ai # puts `cotal` on your PATH (needs Node 22+)\ncotal --help # every command, grouped\ncotal --version # cotal-ai version + each installed extension's (also `cotal -v`)\ncotal <command> --help # one command's flags and usage\n```\n\n`npx cotal-ai <command>` runs it without a global install; in a dev clone, `pnpm cotal <command>`\nruns it through `tsx` with no build step. Bare `cotal` prints help. Every command generates its own\n`--help`, usage, and shell completion from its declared flags.\n\nCommands come from the surfaces the binary composes: the base mesh CLI, the manager\n(`supervise`), and the delivery daemon (`deliver`), plus any operator-installed extensions.\n`cotal ext add <npm-package>` installs any registry providers a package contributes: commands,\nruntimes, and local process lifecycle descriptors. The `web` dashboard and optional manager\nruntimes ship this way.\n\n## Commands\n\n| Area | Command | Purpose |\n|---|---|---|\n| Set up & lifecycle | [`setup`](#setup) | Guided, configure-only setup (installs, seeds personas; launches nothing) |\n| Set up & lifecycle | [`update`](#update) | Reconcile first-party extensions and check or opt into a coherent CLI upgrade |\n| Set up & lifecycle | [`up`](#up) | Start a local mesh (nats-server + JetStream), or boot a whole manifest with `-f` |\n| Set up & lifecycle | [`down`](#down) | Stop the whole stack, selected registered components, or a manifest deploy |\n| Set up & lifecycle | [`backup`](#backups) | Create an offline full-space or registry-only artifact from a preserved cut |\n| Set up & lifecycle | [`clean`](#clean) | Configurable cleanup: purge history (live), or wipe the local store / identity (stopped) |\n| Set up & lifecycle | [`meshes`](#mesh-registry) | List the running meshes on this machine |\n| Set up & lifecycle | [`sync`](#mesh-registry) | Refresh the signed-in account's advertised spaces |\n| Set up & lifecycle | [`use`](#mesh-registry) | Set the default mesh a bare `cotal spawn` joins |\n| Set up & lifecycle | [`status`](#mesh-registry) | Read-only diagnostics for setup, processes, and the selected mesh |\n| Agents & personas | [`spawn`](#spawn) | Launch an agent from a persona (foreground, or `--detach` via the manager) |\n| Agents & personas | [`models`](#models) | List connector model catalogs and variants from the manager |\n| Agents & personas | [`ps`](#managed-seats) | List managed agents and their mesh status |\n| Agents & personas | [`stop`](#managed-seats) | Ask the manager to stop a managed agent |\n| Agents & personas | [`attach`](#managed-seats) | Stream and drive a managed agent's terminal (pty runtime) |\n| Agents & personas | [`input`](#input) | Type one line into a managed agent's terminal without attaching |\n| Agents & personas | [`personas`](#personas) | List, show, edit, create, or remove local personas |\n| Agents & personas | [`supervise`](#supervise) | Run a manager daemon (the agent supervisor / control plane) |\n| Agents & personas | [`service`](#service) | Run the manager as a user service (survives logout and reboot) |\n| Agents & personas | [`runtimes`](#runtimes) | List the agent runtimes the manager can spawn through and whether each is reachable |\n| Agents & personas | [`reconcile-gate`](#reconcile-gate) | Unfreeze an issuance gate left frozen by a crashed restart when the successor cannot boot-heal it (holder gone, complete CONNZ sweep) |\n| Messaging & watching | [`endpoints`](#endpoints) | List every endpoint in the live presence roster, including infrastructure |\n| Messaging & watching | [`describe` / `invoke`](#endpoint-control) | Resolve a v0.4 service's command surface off the wire; invoke one command by name |\n| Messaging & watching | [`send`](#send) | Send one message, then exit: DM a peer, post a channel, or ask a role |\n| Messaging & watching | [`channels`](#channels) | Inspect or set the channel registry |\n| Messaging & watching | [`history`](#history) | Clear retained message history |\n| Messaging & watching | [`console`](#console) | Live protocol view for a space (TUI, or `--plain` line stream) |\n| Messaging & watching | [`web`](#web) | Browser dashboard (installed as the `@cotal-ai/web` extension) |\n| Auth & meshes | [`mint`](#mint) | Mint a creds file for a space (static auth mode) |\n| Auth & meshes | [`login`](#login) | Sign in to a per-user-auth mesh's IdP (once per machine) |\n| Auth & meshes | [`logout`](#login) | Revoke the IdP session and clear the cached login |\n| Auth & meshes | [`actor`](#actor) | Manage a user-auth space's actor ledger (grant / revoke / list) |\n| Auth & meshes | [`doctor`](#doctor) | Credential-health diagnosis and repair (`doctor auth`) |\n| Auth & meshes | [`join`](#join) | Join a space as your own presence (interactive) |\n| Manifest | [`topology`](#manifest-deploys) | Validate and view a mesh manifest's access graph (read-only) |\n| Extensions & misc | [`ext`](#ext) | Install / remove operator CLI extensions |\n| Extensions & misc | [`completion`](#completion) | Print or install shell completion |\n| Extensions & misc | [`feedback`](#feedback) | Send feedback to the Cotal developers |\n| Extensions & misc | [`deliver`](#server-daemons) | Run the server-side Plane-3 delivery daemon |\n| Workflow runs | [`run`](#run) | Operate durable workflow runs: start, resume, list, inspect, answer a checkpoint, check an edited program with migrate |\n| Extensions & misc | [`feedback-intake`](#server-daemons) | Run a self-hosted feedback intake server |\n\nThe manifest modes of `up`, `spawn`, and `down` (`-f <cotal.yaml>`) plus `topology` are covered\ntogether under [Manifest deploys](#manifest-deploys).\n\n## setup\n\n```bash\ncotal setup [--full] [--demo] [--yes] [--skills]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--full` | off | Redo the full guided flow (implies `--demo`) |\n| `--demo` | off | Also seed the guided expert team (`david`, `sven`, `me`) |\n| `--yes`, `-y` | off | Non-interactive accept-all (for agents / CI) |\n| `--skills` | off | Reconcile Cotal skills only through installed connector providers, plus `~/.agents/skills`. Refused with `--full` or `--demo`. |\n\nGuided setup is **configure-only**: it checks prerequisites, invokes installed connectors' declared setup providers, and\nseeds persona files, and it launches nothing (no mesh, no web, no manager). First run gets the\nnarrated flow; later runs print a status card. By default it seeds one `default` persona; the\n`david`/`sven`/`me` team is opt-in via `--demo`. `cotal status` points stale Claude skills and\nout-of-date `.agents` skills at `cotal setup --skills`, not unscoped `setup`. See [Getting started](getting-started.md) and, for\nmaintainers, [setup internals](setup-internals.md).\n\nWhen a mesh resolves, setup seeds that mesh's recorded `.cotal/agents` catalog, the same catalog a\nfollowing `cotal spawn` reads. It prints the absolute destination. On a fresh machine with no mesh it\nuses this folder and says why; when several meshes are available and none is selected, it refuses\nrather than choosing a catalog.\n\n## update\n\n```bash\ncotal update [--self] [--space <s>] [--server <url>] [--creds <path>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--self` | off | If a newer release exists, install that exact validated `cotal-ai` version globally and reconcile through the newly installed binary |\n| `--space`, `--server`, `--creds` | resolved mesh | Select the running manager whose continuity state is reported |\n\nWithout `--self`, `update` keeps the installed first-party surfaces coherent with the running\nbinary: it force-reconciles the four built-in connectors, then reinstalls other `@cotal-ai/*`\noperator extensions at the binary's exact version. Each extension runs in an isolated child, so one\nfailure cannot poison later replays. It then checks npm; a newer binary is an informational notice\nwith `cotal update --self` as the next command, not an automatic install.\n\nAfter disk reconciliation, `update` reads the selected running manager. A machine with no recorded\nmesh has no running manager to observe, so that read is skipped and the command completes. The same\nholds when every recorded mesh is down and none is selected. A recorded mesh that is down is still a\nrefusal when the command selects it, with `--space` or by running inside its project, and so is a\nnamed space that is not running. With several meshes running and no `--space`, `--server` or\n`--creds`, the install is machine-wide, so every running manager is reported in turn, each under its\nspace name, before anything is written; a `legacy` verdict on any of them makes the whole run not a\nhot update. A selector flag still reports one manager. A manager without a\ncustody generation is reported as `legacy`: it cannot preserve its manager-owned PTYs, so the\ncommand says that this is not a hot update and prints `exact`, `fork`, `fresh`, or `drain-only`\nfor every seat. This report sends no stop, preservation-commit, or replacement command.\nIt does not preserve a running PTY on a legacy manager. On Linux a detached custodian\nowns each PTY, so a manager-worker death no longer closes the seat and `status` reports\n`custodied`. Other platforms still spawn in-process and report `legacy`. An incompatible native\n`@lydell/node-pty` or ConPTY ABI break remains an explicit per-seat maintenance cut.\n\nWith `--self`, the selected running manager is reported before any global install. When a newer\nrelease exists, Cotal then installs the exact version it validated, resolves and verifies that\npackage in npm's global root, then launches that binary with the same `--space` / `--server` /\n`--creds` selection to reconcile connectors and first-party extensions to the new generation. An npx\nor dev-clone invocation therefore installs and continues through a separate global copy; it never\nclaims the already-running process changed. If the binary is current, `--self` performs the normal\nlocal reconcile without reinstalling it.\n\nThird-party extensions are listed with their installed version and recorded spec but are not\nauto-updated in v1. Floating third-party updates require `@cotal-ai/*` peer-range validation and are\na future follow-up. A failed connector/extension install, npm metadata check, or requested global\ninstall is reported and makes the command exit nonzero. Independent extension attempts continue so\nthe output includes every failure; an unavailable npm registry does not undo a completed local\nreconcile, but the command still exits nonzero because it could not establish that the install is\ncurrent.\n\n## up\n\n```bash\ncotal up [--detach] [--open] [--space <s>] [--server <url>] [--channels <path>] [--runtime <name>]\ncotal up --user-auth --idp <url> [--exchange-public-port <n> --exchange-public-url <https://…> [--exchange-trusted-proxy]]\ncotal up --tls-cert <cert.pem> --tls-key <key.pem> # serve broker TLS (both, or neither)\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\ncotal up -f <cotal.yaml> [--dry-run] [--runtime <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--server <url>` | auto (free local port) | Listen URL override |\n| `--host <host>` | none | Bind host override for a **fresh** broker boot: an IP or hostname only, never a URL (that is `--server`) and never `host:port` (the port comes from `--server` or its default); a URL or port-bearing value is refused pointing at the right flag. With no `--server`, the broker URL is derived from it, so `--host <addr>` alone is enough to make a mesh reachable at that address; a `--host`/`--server` pair naming different addresses is refused. A wildcard bind (`0.0.0.0`, `::`) keeps a dialable loopback URL. Recorded on the mesh and reused by every later manager launch, so a repair or resume keeps remote [`attach`](#managed-seats) working. A live refresh (`✓ mesh already running`) does not rewrite `.cotal/auth/server.conf` or rebind nats; stop the broker, then re-run `up --host` |\n| `--space <s>` | the folder's name | Space name |\n| `--store-dir <dir>` | none | JetStream store directory |\n| `--max-file-store <bytes>` | nats-server's dynamic cap | JetStream file storage cap in bytes (a positive integer, no unit suffix). Without it nats-server sizes the store at start as three quarters of the free space on its filesystem. The cap is fixed at broker start: a running broker cannot change it (`cotal down` first), `down --preserve-state` keeps it for the resume, and a resume with a different value is refused. Not accepted with `-f` |\n| `--channels <path>` | `.cotal/channels.json` if present | Channel-registry seed file (JSON). An explicit path that is missing is an error |\n| `--restore <dir>` | none | Restore a completed offline backup before exposing the normal listener |\n| `--restore-only registry` | artifact selection | Restore only the registry component |\n| `--accept-missing-source` | off | Explicit disaster consent when the inode-bound preserved source is absent |\n| `--accept-stale-checkpoint` | off | Explicit consent to resume a seat whose checkpoint was captured outside its recorded recency horizon |\n| `--open` | off (auth) | Unauthenticated dev mesh: no JWT, no ACLs |\n| `--user-auth` | off | Per-user auth: people `cotal login`; connects are authorized against the actor ledger |\n| `--idp <url>` | none | With `--user-auth`: the IdP auth base URL to pin on first enable |\n| `--exchange-public-port <n>` | none | With `--user-auth`: add the public exchange face on this loopback port, for an HTTPS reverse proxy to forward to |\n| `--exchange-public-url <https://…>` | none | With `--exchange-public-port`: advertise the reverse proxy's HTTPS URL in discovery |\n| `--exchange-trusted-proxy` | off | With `--exchange-public-port`: attribute public failure buckets to the last `X-Forwarded-For` hop. Enable only when the listener is reachable solely through a trusted proxy; otherwise the socket address is used |\n| `--detach` | off | Run in the background (stop with `cotal down`) |\n| `--tls-cert <path>` | none | PEM certificate to serve TLS with. Must be given together with `--tls-key`. Before starting the broker, Cotal checks readability, private-key mode, key/certificate match, the validity window, and host coverage. `nats-server` accepts an expired certificate and leaves the failure to clients, so Cotal performs these checks first. The decision is recorded; a later bare `cotal up` keeps serving TLS |\n| `--tls-key <path>` | none | PEM private key for `--tls-cert`. Refused if group- or other-readable (tighten to `600`) |\n| `--file <cotal.yaml>`, `-f` | none | Launch a whole mesh from a manifest |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--runtime <name>` | `pty` (or the manifest's, with `-f`) | Agent runtime for the mesh manager (`pty` built in; others are installed extensions, explicit-only). Resolved + probed before the broker starts; an uninstalled/unreachable runtime fails loud. With `-f`, overrides the manifest's runtime |\n| `--max-sessions <n>` | 64 | Live-session ceiling for the mesh manager. Each console pane and each `cotal attach` is one session, so size for agents × panes, not agent count. Recorded on the mesh and reused by every later manager launch, so a repair or resume does not silently drop back to 64. A running manager cannot change it: `cotal down` first, then `cotal up --max-sessions <n>` |\n| `--no-manager` | off | Broker-only boot: start the broker and, in auth mode, the delivery daemon, and no local manager. A refresh under the flag of a mesh whose manager is live refuses rather than keeping or stopping it (`cotal down manager` first). Cannot be combined with `--runtime`, `--max-sessions`, or an agent-declaring manifest |\n| `--rotate-sys` | off | Rotate the space's system account and re-mint its two `$SYS` creds. Needs a stopped mesh; refused with `--open` |\n\n`cotal up` boots a local nats-server with JetStream and, in auth mode (the default), JWT auth and\nper-agent ACLs; `--detach` records the mesh so `cotal spawn` from any directory can find it. With no\n`--server`, it auto-selects a free port if the default address is taken; an explicit `--server`\nstays fail-loud on collision. `--detach` also brings up the control plane (delivery daemon in auth\nmode, then the manager). `--no-manager` is the broker-only mode: it boots\nthe broker (and the delivery daemon in auth mode) and starts no manager, so there is no manager\npidfile to leave stale. A refresh under the flag of a mesh whose manager is live refuses rather\nthan keeping or stopping it: `cotal down manager` first. For a split topology with a manager, wait for `.cotal/manager.<spaceKey>.log` to contain `✓ manager up`, then `cotal down manager` on that\nhost and run [`supervise`](#supervise) against the remote broker; see\n[Run a mesh](run-a-mesh.md). `cotal up --detach` prints `✓ running in the background:` with\n`manager` listed (pidfile liveness, not a teardown boundary); with `--no-manager` the line lists\nonly what actually started. Ctrl-C on a foreground `up` spares managed agents and reports them\nunder the same rule as bare `cotal down` (see [`down`](#down)): when the manager cannot prove it can\nspare, Ctrl-C refuses the teardown, prints the refusal with the reap route, and leaves the stack\nrunning. The `-f` form is a\n[manifest deploy](#manifest-deploys).\n\nThe generated `.cotal/auth/server.conf` is written on a real broker boot and is not an\noperator-owned config. `--host` changes that file only when nats is actually started. A unit\nrestart that leaves an answering listener in place is a refresh, not a rebind.\n\nOn an existing mesh, `cotal up` reconciles the presence and lease bucket TTLs. It writes a reserved\ncanary and waits for the bucket to expire it before reporting success. If the broker accepts the\nstream update but the backing store does not persist or enforce it, `up` exits nonzero with a TTL\npersistence error instead of trusting the value returned by stream info.\n\n`--user-auth --idp <url>` starts the space's auth service alongside the broker: the NATS\nauth callout plus its capability-gated local exchange, and optionally the closed public exchange\nface configured by the three `--exchange-*` flags above. The service is torn down with `cotal down`,\nand a re-run of `cotal up` heals a dead service on a running broker. `up` waits for the service to\nfinish binding: while the daemon it launched (or found running) stays alive, the wait extends past\nthe base 15s up to 60s; a daemon that exits is refused at once with \"exited before becoming ready\",\nand one alive past 60s is refused as \"alive and still starting\" (wedged), naming the pid record and\nthe service log. `--user-auth` and `--open`\ncontradict each other and are refused loudly; a running broker cannot change auth mode\nwithout a `cotal down` first. See [identity & auth](identity-and-auth.md).\n\n`--rotate-sys` renews the two `$SYS` credentials (`membership-observer`, `connection-evictor`).\nThey carry a 30-day expiry and nothing re-signs them in place, because the system-account seed is\nnever persisted, so they are renewed by issuing a **new system account** under the same broker\noperator and minting fresh creds against it. A plain re-`up` does **not** do this: it reuses the\nexisting trust record, and its `$SYS` creds along with it.\n\nThe rotation is safe to run on a real space, with one operational cost. The data account, the account\nsigning key, every agent credential minted from it, and the JetStream store are all untouched; what\ndies is the retired system account, and with it any out-of-band copy of the old `$SYS` creds, on every\nbroker that loads the rotated config. The cost is that **earlier full backups stop being restorable**\n(see below), so this is not a no-consequence operation. It needs the broker to restart on the rewritten\nconfig, so it runs as part of a boot:\n\n```bash\ncotal down\ncotal up --rotate-sys --detach # agents reconnect; nothing is re-provisioned\ncotal doctor auth # both $SYS creds healthy again, 30 days out\n```\n\nA rotation is a stopped, fresh boot, and anything that is not one refuses it, all for the same reason\n(the on-disk material and the broker it runs on must never end up on different generations):\n\n- a live mesh, because the running broker would keep serving the retired account;\n- an open mesh, whether that comes from `--open` or from `broker.auth: false` in a manifest, which\n has no system account at all;\n- `--restore`, because reinstating a trust root and superseding it in one command leaves no way to\n say which authority the mesh came up on;\n- an unfinished restore or resume attempt on this root, including one `cotal up` would recover on\n its own, because those paths can adopt a live listener and return without booting a broker;\n- a root that hosts more than one space, because the system account lives in the shared broker\n record and a rotation would retire every tenant's, while the root holds one `$SYS` cred pair\n pinned to one data account.\n\nTwo things to know before you run it:\n\n- **The retirement is config-load-bound.** Old `$SYS` creds are refused by any broker that loads the\n rotated config. A stale `nats-server` still running the *previous* config in memory would keep\n honouring them, so stop every broker for this root first. `--rotate-sys` refuses if this root's\n mesh is recorded as running, if anything unidentified is answering at the address it was given, or\n if the root's pid file names a live (or unreadable) process. Those are Cotal's own ownership\n records, not a scan of the process table: a `nats-server` you started by hand against this root's\n `server.conf` on some other port writes none of them and will not be seen. Do not run one.\n- **It invalidates earlier full backups.** A full artifact binds to the trust chain it was taken\n against, and that commitment covers the operator JWT and the system account. Every full backup\n taken before a rotation refuses to restore afterwards, so take a fresh `cotal backup` once the\n rotated mesh is up. `cotal up --restore` names this case when the data account still matches.\n\nThe commit is not atomic (a trust-record write plus two credential writes), so an interrupted\nrotation leaves the record ahead of the creds. That split is detected rather than silent: every\n`cotal up` on an auth mesh, and every `cotal doctor auth`, compares each `$SYS` cred's issuer against\nthe persisted record and names the retired account. `up` warns rather than refusing, because these\ncreds power the membership graph and live eviction, both of which degrade fail-soft; the mesh is not\nworth taking down over them. Re-running the rotation heals it, at the cost of one generation.\n\nWhile those creds are expired the mesh keeps delivering messages, but the\n[membership feed](delivery-daemon.md) and live connection eviction stay down; `cotal doctor auth`\nand the manager's log both name the credential and this repair.\n\n## down\n\n```bash\ncotal down\ncotal down --with-agents\ncotal down --preserve-state [--store-dir <dir>] [--session-store <dir> …]\ncotal down manager [delivery auth web nats ...]\ncotal down web [--space <name>]\ncotal down -f <cotal.yaml> | --run <id> [--dry-run]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--file <cotal.yaml>`, `-f` | none | Tear down this manifest's deploy |\n| `--run <id>` | none | Tear down one `spawn -f` run by id |\n| `--space <name>` | current mesh | With components: the mesh whose target-addressed components (e.g. `web`) to stop |\n| `--dry-run` | off | Print the manifest teardown or selected components, mutate nothing |\n| `--with-agents` | off | Bare whole stack only: also stop and deprovision every managed agent |\n| `--preserve-state` | off | Bare whole stack only: fence the manager, retain principals and durable state, stop and prove the stack down, then publish `ready` |\n| `--store-dir <dir>` | `.cotal/nats` | With `--preserve-state`: the actual store path (required for a custom store) |\n| `--session-store <dir>` | none | With `--preserve-state`: a harness transcript store directory to capture with every continuation-capable retained seat. Repeatable. No default and never inferred from a connector name; a path that does not exist or is not a directory is refused before anything stops |\n\nBare `cotal down` stops the whole local stack in dependency order and leaves managed agents running.\nBefore signalling the manager it verifies that the exact recorded manager supports releasing its\nlocal custody, and it reports the agents left behind plus `cotal down --with-agents` as the explicit\nreap. Ctrl-C on a foreground `cotal up` follows the same rule: it verifies the spare capability,\nspares the agents, and prints the same report. When the capability cannot be verified, Ctrl-C\nrefuses the teardown and leaves the stack running; end it with `cotal down --with-agents`.\n`--with-agents` is a one-shot destructive policy bound to the exact verified manager process\nand the exact live `down` stop reservation; a stale, malformed, crashed, or different stop attempt\ncannot turn a later bare shutdown destructive. Positional component names stop\nonly those self-registered local processes; for example, `cotal down manager` leaves delivery and\nthe broker running, and `cotal down web` is available when the web extension is installed. A\ncomponent that starts target-resolved (the web dashboard) is stopped the same way: `cotal down web`\nresolves the mesh the same way as `cotal web` (registry current mesh first, `--space` to name one), so\nit works from any directory; the other components always stop under the folder you run it in. The\n`-f` / `--run` forms tear down a [manifest deploy](#manifest-deploys) without stopping the whole mesh\nand cannot be combined with component names. Stopping `nats` alone is refused while an unselected\nregistered daemon is still live; include those components or use bare `cotal down`.\n\nA pinned manager with no spare-capability record is not signalled by bare `cotal down` or `cotal\ndown manager`. The missing record can mean either that the manager predates capability reporting or\nthat a current manager cannot detach its agents. Stop each managed agent explicitly, then run `cotal\ndown --with-agents` from the mesh root to stop the whole stack. An older manager does not understand\nthe reap request, which is why the agents must already be stopped.\n\nBare `cotal down` inventories by pidfile. When this folder's registered broker answers and no\n`nats.pid` records it, the command does not say nothing is running. It names the space and the\nbroker address, says no pidfile records that process, says it will not stop a process it did not\nstart, and exits 1. Stop that broker with whatever started it (an init unit, a container, or the\nhand-run process). `cotal meshes rm <space>` only drops the registration. The probe runs whether or\nnot other owned components were running: they stop and clear their artifacts first, then the broker\nis named. A component stop and `--dry-run` stay pidfile-only and do not probe.\n\n**Teardown verifies pinned process identity before signalling.** PIDs are recycled by every OS,\nso a recorded pid alone is not a durable target identity. `up` records each stack process's\ncreation identity in a sibling `<pidfile>.identity` pin, which holds the pid and the process start\nreported by the OS. Every stop path, including `down` for the broker, web and extension components,\nand the manager, delivery and auth-service stops, applies the same rule. A pin that names a different\nstart means the pid was reused, so teardown refuses and preserves it. A torn or unreadable pin also\nrefuses.\n\nThe pidfile pid and the pin pid are two coordinates. Automatic cleanup follows **proven death of\nthe pidfile target** (ESRCH on that pid): a torn sibling pin does not wedge a dead pidfile pid.\nA torn pairing where the pin names another pid, while the pidfile pid is still live or not proven\ndead, still refuses. Inspect both pids with `ps`. Do not delete `<pidfile>.identity` to force a\nstop; that weakens target-identity protection. Once the pidfile process is dead, rerunning\nteardown clears the stale record automatically.\n\nThe first teardown after upgrading a running pre-pin stack has a narrower guarantee. A live record\nwith no identity pin is signalled after a loud warning that it predates identity pinning. Restarting\nthe component writes the pin, so later teardowns receive full match and mismatch protection. The\nsame warning applies on platforms where no stable start token is available. For a legacy manager,\nbare `cotal down` also warns that agent sparing cannot be verified before it signals. Because the\nCLI cannot establish which SIGTERM handler that already-running binary carries, it never presents\nthe pre-signal seat inventory as confirmed spared; a genuinely older destructive handler may still\nreap those agents. `--with-agents` publishes a one-shot reduced-guarantee handoff bound to the\nrecorded manager pid and the live `.stopping` reservation's inode, then signals unconditionally.\nThat handoff cannot be replayed by a later stop attempt. A pin that exists and does not match the\nlive process still refuses before signal.\n\n`--with-agents` performs the old destructive logical teardown: managed processes stop and their\ncredentials, ACL rows, and delivery footprints are deprovisioned. `--preserve-state` is a different\nmaintenance transition: it stops retained processes while suppressing leave/deprovision cleanup, persists the manager's\nsame-principal resume inventory, stops the entire stack without removing run/auth artifacts, and\npublishes a stable inode-bound cut only after every recorded process is proven stopped and the exact\nrecorded NATS endpoint is unreachable. A missing or stale broker pidfile never counts as stopped. The\nattempt is bound durably before the manager is fenced, the resume document and attempt-bound\n`cut-intent` are fsynced before manager commit, and the manager's commitment itself is journaled\n(`cut-committed`) before any process stops. A retry after a crash at any of those boundaries reuses\nthe exact recorded attempt and finishes the remaining stop and endpoint proofs idempotently, without\nneeding the (by then intentionally dead) manager. A partial cut never publishes `ready`. It cannot\nbe combined with component names, manifest teardown, or `--dry-run`.\n\n**Seat checkpoints.** After the stack is proven down, the cut writes one checkpoint per retained\nseat under `.cotal/maintenance/v1/checkpoints/<attempt>/<seat>/`, and prints the path, the\ncontinuity class and the generation for each. The path carries the preservation attempt because a\ncheckpoint is immutable once sealed: a shared directory would make the second cut in a root refuse\non the first cut's leftovers, and clearing it would destroy an artifact a rollback still needs. The\ncapture happens only at that point because anything earlier races a harness that is still writing\nits transcript and its working tree.\n\nEach checkpoint directory is created 0700, refuses a destination that already exists, and holds:\n\n- `repo.bundle`, the seat `cwd`'s reachable history, anchored on the base commit the record names\n by full object id;\n- `repo.index.diff` and `repo.worktree.diff`, the staging state as two diffs, base to index and\n index to worktree. Two rather than one because a single combined diff restores a mixed tree with\n the right bytes and the wrong index: a source reporting `MM README` would come back as ` M README`;\n- `repo.untracked.tar`, the untracked files in scope;\n- the harness session pointer, when the seat's connector declares one, and the transcript store\n files the operator named with `--session-store`. Each records where the destination puts it back\n as an anchor (the workspace root, the account home, or the seat's `cwd`) plus a relative path,\n because the destination's root and home are its own and the source host's absolute spelling would\n either miss them or write outside them;\n- `checkpoint.json`, written last, after every digest is computed over the bytes that landed.\n\nThe record carries the manager's resume entry unchanged as its first field, then the space, the seat\nname, the recovered `lifecycleUid`, the writer generation the cut was taken at, `capturedAt`, the\nrecency horizon, the applied profile revision, the seat's `git status --porcelain` as the cut read\nit, and the continuity class. Every captured file is\nlisted with its byte size and sha256, so an operator verifies the whole artifact with `sha256sum`\nand `git bundle verify`. No secret values, no operator keys and no source-host launch material\nenter it.\n\nThe continuity class is what the connector declares, capped by what the checkpoint carries. A\nconnector declaring session continuation classifies as `exact`, but reopening a session takes both\nhalves, the pointer that names it and the store that holds its transcript. A checkpoint missing\neither one cannot reopen that session, so it is recorded as `fresh` when the connector declares a\nfresh start and `drain-only` otherwise. A pointer with no store is capped the same way as a cut\ncarrying neither, because it names a session whose bytes the artifact does not contain. A class is a promise the destination is entitled\nto act on, so it never describes bytes the artifact does not contain. The transcript store stays an\noperator input: this repository does not know where a harness keeps its transcript, so `exact`\nrequires `--session-store` to name one.\n\nThe recorded status is read under the same selection rule as the untracked set, so it describes the\nstate the captured bytes can reproduce. The destination re-reads it in the promoted tree and refuses\na difference.\n\nThe untracked selection rule is recorded in the record and is\n`git ls-files --others --exclude-standard -z, excluding .cotal/`. It honors `.gitignore`, so an\nignored file the seat needs does not travel and has to be moved separately. The `.cotal/` exclusion\nis a secrecy boundary rather than a size one: when a seat's `cwd` is also the mesh root, the control\ndirectory is untracked, and without the exclusion the broker trust material, the space account, the\nmanager instance identity's private seed and the seat's own credentials would land inside the\nartifact. A checkpoint carries credential references only; the destination resolves that material\nitself.\n\nA seat whose launch options could not be resolved is refused rather than checkpointed, with the\nmanager's own wording: `imperative launch options have no non-secret durable source (<keys>)`. The\nrefusal arrives at prepare time, so the cut stops before any child does.\n\n## clean\n\n```bash\ncotal clean <history|store|all> --force\ncotal clean restore-attempt --attempt <id> --force\ncotal clean restore-fallback --attempt <id> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | `history`: target mesh |\n| `--dms` | off | `history`: also clear DM history |\n| `--store-dir <dir>` | `.cotal/nats` | `store`/`all`: JetStream store directory |\n| `--force` | none | Required: destructive, no prompting |\n| `--attempt <id>` | none | `restore-attempt`: exact stale pre-commit attempt; `restore-fallback`: matching healthy committed restore |\n\nOne configurable cleanup verb; every target requires `--force`.\n\n- `history` purges the retained message backlog on the **running** broker (channels, plus DMs\n with `--dms`). The same operation as [`history clear`](#history), which stays as an alias.\n- `store` deletes the **stopped** mesh's JetStream store (`.cotal/nats`): streams, durable\n consumers, and messages. This is the reset for stale on-disk broker state, e.g. durables\n minted by an older, incompatible Cotal generation surviving a `down`/`up` cycle.\n- `all` is `store` plus the space identity (`.cotal/auth`), the local creds and markers tied to\n it, any crash residue a normal `down` would have swept (stale pidfiles, `run/`), and the mesh's\n registry entry; the next `cotal up` mints a fresh identity.\n\n`history` needs the mesh up; `store` and `all` refuse while any recorded mesh process is still\nalive or any same-root recorded broker endpoint remains reachable (run `cotal down` first). They\nalso refuse outright on a root that holds accounts for several spaces: the store and the broker\ntrust record are shared by every space on the broker, so both targets would take out all of them\nand no `--space` can narrow that. `down`, `backup` and `up --restore` refuse there for the same\nreason. `cotal status` lists the tenants on such a root. Personas\n(`.cotal/agents`) and logs are never touched. A custom\nstore location is not recorded anywhere, so `--store-dir` must repeat whatever the mesh was\nlaunched with. Custom cleanup targets must contain either the Cotal store-generation marker or a\nreal `jetstream/` store directory; filesystem roots, project roots, and Cotal auth/maintenance trees\nare always refused.\n\n`store` and `all` also refuse every maintenance journal state. After a healthy committed restore,\n`restore-fallback` is the only supported way to remove the recorded unchanged old-store inode; it\nnever deletes the active target, requires both the exact attempt id and `--force`, and retires the\ncompleted restore journal so a later `down --preserve-state` can start a new backup cycle.\n\n## Backups\n\n```bash\ncotal down --preserve-state [--store-dir <dir>]\ncotal backup create <dir> [--only full|registry] [--store-dir <dir>]\ncotal up --restore <dir> [--restore-only registry] [--accept-missing-source]\n```\n\nBackup is offline-only. It requires the stable `ready` record from `down --preserve-state`, an exact\nstore match, no live recorded process, and an unreachable exact endpoint from the recorded cut.\nThat endpoint is probed immediately before cloning, so a live broker with a missing or stale pidfile\nis still refused. It claims the cut, reflink/copies the stopped source to a\nprivate attempt clone, and opens only that clone on a random loopback bootstrap broker with an\nindependent parent/deadline watchdog. It validates the canonical stream and pull-consumer inventory,\nwrites native snapshots with consumers excluded, and stores conservative contiguous ACK-floor\ncheckpoints separately. The original store is never opened by the backup broker, and the stack is\nnot restarted implicitly. Artifact destinations must not overlap the preserved source or maintenance\nattempt tree. Restore artifacts and targets likewise cannot nest inside or contain each other, the\npreserved source, or the maintenance attempt tree.\n\nStopped client-managed KV ordered consumers are ephemeral read residue, not backup state. Backup\nignores only the pinned client's exact stopped shapes: ordinary last-value watchers and the\nwhole-bucket scanner that uses all-history delivery to collapse concurrent tombstones. A bound\nconsumer or any lookalike with a different filter, inbox, lifetime, or other config is still refused.\n\n`full` is the default and indivisible: channel registry, CHAT/DM/TASK/INBOX/DLV, ACL, MEMBERS, and\nvalidated durable checkpoints. `registry` is the sole partial artifact. Presence, derived membership\nfeed, leases, native ephemeral/history consumers, credentials, keys, tokens, owner secrets, and actor\nledger files are excluded. `full` means every transferable message and registry stream, not every\nJetStream resource: endpoint submissions/facts/events/timers/workflow state, contract artifacts, and\nthe records/auth/session stores are nonportable control state. Restore recreates those streams empty\nwith their canonical configs before exposing the normal listener, so active endpoint runs,\nlifecycles, and sessions do not cross a backup. Artifacts are exclusively created `0700`;\nsnapshot/checkpoint files and\nthe manifest are `0600`; `manifest.json` is written last with exact sizes and SHA-256 values. The\ndirectory is trusted operator input: hashes detect corruption, not malicious rewriting.\n\nRestore validates and stages the exact allowlisted artifact bytes before moving or creating a store.\nIt requires the same space and existing trust state. The whole pre-commit window holds a journaled\nliveness claim (coordinator, watchdogs, brokers, absolute deadline): ordinary `up` and a repeated\n`up --restore` refuse while the claim is live, and a stale attempt is recovered only after the\ndeadline has elapsed and every recorded owner is proven dead. A retried `up --restore` handles this\nautomatically; an operator can also recover it explicitly with `cotal clean restore-attempt --attempt <id> --force`. Nothing\never rolls back a live attempt. A registry-only artifact restores as registry-only whether or not\n`--restore-only registry` is passed; omitted infrastructure is always created and the exact\npost-restore stream inventory is asserted before commit intent. Ordinary `up` from a preserved cut\nresumes only the exact recorded source store and runtime; a contradicting `--store-dir` or\n`--runtime` fails in preflight.\n\n**Admitting a seat checkpoint.** An ordinary `up` from a preserved cut admits that cut's seat\ncheckpoints before it journals the resume attempt and before any process starts, so a refusal costs\nnothing. Three gates run in order, each naming what it saw.\n\n1. *Integrity.* Every file the record names must be present, a regular non-symlink file, the\n recorded byte size and the recorded sha256, re-stat'd after the read so a file that moved is a\n refusal. Failure here consults no other gate.\n2. *Identity.* The recorded space must match, the recorded `lifecycleUid` must not belong to a live\n incarnation, and the profile revision must match this host's or be resumed under deliberately\n this host's. A differing revision is refused with both digests and the remedy, and there is no\n override: the checkpoint carries the recorded digest and not the config bytes, so nothing could\n run the seat under the recorded revision, and the manager re-digests the same file and refuses\n drift on its own. This gate has no blanket override, which is the only reason the next one may\n have one.\n3. *Recency.* `capturedAt` is compared to this host's clock against the horizon the record carries.\n Inside it, the seat resumes. Outside it, `up` refuses and prints the capture instant, the clock\n reading and the horizon; `--accept-stale-checkpoint` admits it anyway and the exercised consent\n is printed with the actual age. An unreadable `capturedAt` is refused with no override, because a\n freshness gate that fails open is not a gate.\n\nCustody transfers only after all three pass. The destination claims the recorded generation plus one\nby exclusive create, before it launches anything. A lost create means another destination is already\nclaiming that seat, and it refuses with `seat-writer-generation-create-lost` rather than adopting\nthe winner and becoming a second writer. The recorded `lifecycleUid` is reused and never minted, so\nthe resumed seat binds the same lifecycle-keyed durables.\n\nAdmission is reconciled against the inventory the resume is about to hand the manager, and that\nreconciliation finishes before the restore moves a single tree. A checkpoint whose recorded\n`lifecycleUid` is not the one the retained inventory carries describes a different incarnation of\nthat seat, and it refuses with both uids while every live working tree is still untouched and no\ngeneration is claimed. A retained\nseat with no admitted checkpoint refuses the resume by name: an absent checkpoint directory and an\nabsent record are indistinguishable from a seat that was never checkpointed, and a seat that starts\nwithout passing the gates has claimed no generation. `--accept-stale-checkpoint` is recorded in the\nresume journal with the seat, the capture instant, the admitted age and the horizon, so the consent\nsurvives the terminal it was typed into.\n\nThe whole admission is all or nothing. Coverage is settled first, then every gate runs over every\ncheckpoint, and only then is any generation claimed. A refusal at any point leaves every generation\nunclaimed, including a lost exclusive create during the claim itself: the claims that attempt made\nare removed before the refusal is raised, by the exact paths it wrote, so a generation another\ndestination holds is never touched. A claim is a create that can never be made again, so a refusal\nthat left one behind would consume the retry over the same checkpoint set.\n\n**Restoring a seat checkpoint.** Once every gate has passed over every checkpoint, and before a\nsingle generation is claimed, `up` puts each admitted seat's captured bytes back. A refusal here\ncosts nothing for the same reason a gate failure does: no claim has been made and nothing has\nstarted.\n\nA restore never moves or replaces the destination's own control directory. A checkpoint excludes\n`.cotal/` by design, so a seat whose `cwd` holds one, which is the layout an operator gets by\nrunning `up` and `spawn` in a single directory, is refused before anything is staged: promoting a\ntree that cannot contain `.cotal/` over that `cwd` would carry this host's live trust material and\nmaintenance state away with the superseded tree. The refusal names the control directory it found\nand the remedy, which is to give the seat a working tree that is not a workspace root.\n\nEach seat is staged beside its own `cwd`, in `<cwd>.incoming`:\n\n1. every recorded digest is verified again over the files as they are now;\n2. the bundle is cloned into `<cwd>.incoming`, which is refused when that path already exists;\n3. the recorded base commit is verified in the clone and checked out detached, so a bundle that does\n not contain it stops the resume instead of continuing against a different history;\n4. the index diff is applied with `--index` and the worktree diff without it, both `--binary\n --allow-empty`. That order is what puts staged content back in the index rather than only in the\n worktree, and `--allow-empty` is why a seat with a clean tree is still restorable;\n5. the untracked archive is extracted.\n\nEvery seat stages before any seat is promoted. Promotion moves an existing `cwd` aside to\n`<cwd>.superseded.<timestamp>` and renames the staging directory into place, then puts the session\npointer and store files where the destination's connector reads them, then re-reads\n`git status --porcelain` in the promoted tree and compares it to the status the checkpoint recorded.\nA restore that applied without error and produced a different index is a refusal, not a warning. The\ntwo renames are the only steps that touch the path the seat will use, so a failure anywhere leaves\nevery seat's live `cwd` as it was.\n\nThe rename itself claims the superseded name, and a taken name gets a numeric suffix. The timestamp\nhas one-second resolution, so two promotions of the same seat within one second compute the same\npath; a rename onto a name that already holds a tree fails on every platform, and that failure is\nread as taken. Nothing creates the name ahead of the move, because Windows refuses to rename onto an\nexisting directory at all. A superseded tree is the thing that rename exists to keep.\n\n`git` and `tar` run as child processes with argument arrays, never a shell string.\n\nA leftover `<cwd>.incoming` refuses the resume by name. A staging directory from a failed run is the\nonly record of what failed, so nothing removes one automatically: inspect it, remove it by hand, and\nresume. A pre-existing `cwd` is renamed rather than deleted, so a wrong checkpoint costs a rename\ninstead of a tree. When a promotion fails, the renames that attempt made are undone and the staging\ntree is left where it is, as the evidence for what did not verify.\n\nA session pointer whose recorded `sessionId` is not the one the retained inventory reopens is\nrefused before anything is cloned. A session file already present at its destination is judged by\ncontent: bytes equal to the recorded digest are already restored, and different bytes under the path\nthe connector is about to read are refused with both digests rather than clobbered.\n\n`up --restore <dir>` reaches the same admission and the same restore, after the store is restored\nand validated and before commit intent is journaled. A registry-only restore resumes no seat, so it\nadmits and restores nothing.\n\nOne limit is worth stating plainly. The writer generation is claimed by exclusive create inside one\nworkspace root, so it fences two resumes on the same host and does not fence two independent\ndestinations: copy a checkpoint to two roots and both claim the same successor. A real cross-host\nfence needs a coordinate neither root owns.\n\nAuthenticated restores validate the complete\nspace trust bundle before staging, including nkeys, seed matches, JWTs, signers, and space binding;\nfull restores commit to the validated operator, system-account, data-account, and active-signer root\nchain in addition to the static/user authority fingerprint. Because the system account is part of that\ncommitment, a [`cotal up --rotate-sys`](#up) makes every full artifact taken before it unrestorable\nagainst this root: take a fresh full backup after each rotation. The composed commitment is revalidated\nimmediately before store mutation and never includes secret seeds. Restore never creates fresh auth.\nSame-path restores atomically retain the old\nsource at the journaled fallback path; alternate targets retain it in place; a missing canonical\nsource needs explicit `--accept-missing-source`. Quarantine and target restores use current canonical\nconfigs on isolated random-loopback brokers, never expose native snapshot consumers, and publish a\ncommit-intent immediately before the normal listener starts. Archive bytes never instantiate the real\ntarget: after quarantine validation, every stream is re-snapshotted from the validated quarantine\nstate into attempt-owned sanitized files, and the target is restored solely from those. Before that boundary, failure rolls back\nthe attempt-owned target; after it, ambiguity preserves both stores and records forward-repair\nrecourse. The cooperative maintenance lock excludes Cotal commands, not arbitrary raw NATS processes.\n\nBootstrap brokers in every auth mode, including open, mount the store under a local account with\nrandom operation-specific logins only, each carrying the exact per-phase subject permission matrix;\nnormal static credentials and user-auth sentinel/bearer connections are rejected, and no auth\nservice or callout starts. Open mode differs only in its account label, never in authority. Inventory, each stream snapshot,\nrestore initiation, exact upload id, validation, and each checkpoint recreation use separate exact\nauthorities. Every checkpoint carries the source stream's message/first/last sequence state and must\nmatch its snapshot record before mutation; core then derives and validates the only allowed start\npolicy. TASK is not a CLI exception: the same core checkpoint API recreates its canonical `DeliverAll`\nWorkQueue durable because acknowledged tasks are absent from retention and NATS forbids a\nstart-sequence policy there. Registry-only restore creates every omitted canonical stream and transient\nbucket on the isolated target before the normal listener is exposed. It deliberately does not resume\nretained agents or recreate their DM/DLV/TASK/ACL state; their identity material stays retained and\nstopped rather than being reprovisioned into a partial restore.\n\nAfter listener readiness, the manager starts attempt-bound, validates retained credentials/tokens\nwithout granting or reprovisioning, and resumes the exact persisted principals under cleanup\nsuppression. Registry-only restore uses the same flow with an empty agent set. `commitResume` is an\nidempotent validation barrier only: success must be `awaitingFinalize` with an attempt-bound 64-hex\ncommit token and does not release suppression. Under the workspace lock, the CLI first fsyncs that\nexact evidence as `manager-committed` (restore) or `resume-committed` (ordinary resume), then calls\ntoken-bound `finalizeResume`; only an `active` response for the exact token releases suppression. The\nCLI records the same token in finalization evidence before a restore becomes `active`, or before an\nordinary resume retires and consumes the marker. Re-entry from either committed state skips the prior\nidempotent activation/commit phases, retries finalization with the durable token, and finishes the\nworkspace transition. Failure before finalization preserves the committed state and cleanup\nsuppression; it is not rewritten through a degraded transition. Re-entry between any two earlier\nboundaries reuses the same attempt and may retry the idempotent phases without deleting retained state. A missing or\nchanged per-agent dependency is a named fail-closed result; the journal becomes degraded and remains\navailable for forward repair. A retry from `resume-intent`,\n`resume-active`, or `resume-degraded` reuses the same attempt and inventory after the prior listener is\nproven stopped. Every normal restore listener has an unguessable attempt-bound NATS server name. The\nCLI fsyncs its exact name/nonce, canonical endpoint, process owner, and generation-bound target identity\nimmediately after spawn. Re-entry accepts a surviving listener only when its INFO server name, live PID\nrecord, endpoint, and target identity all match that proof; degraded restore repair then moves through\nthe guarded workspace transition only after manager commit. If an uncommitted bound owner is provably\ndead, recovery retires that exact proof under the maintenance lock and binds a fresh listener for the\nsame attempt, endpoint, and target with a new nonce and server name. A live foreign/mismatched listener\nor ambiguous owner is preserved and refused, never adopted by reachability alone. A reconstructed\ncommit/degraded attempt without either the exact bound proof or a durable dead-listener replacement\nrecord fails closed even when the recorded port is free. A later ordinary startup may pass an `active`\nrestore only when its details prove manager commit and its exact recorded listener is dead.\n\n## Mesh registry\n\n```bash\ncotal meshes\ncotal meshes add # guided, on a terminal\ncotal meshes add <space> --server <url> [--root <dir>] [--mode auth|open|user] [--tls] [--force]\ncotal meshes add <space> --mode user (--user-auth-file <bundle.json> | --from <https url>)\ncotal meshes rm <space> [<space> …] [--force]\ncotal sync [--idp <auth base URL>]\ncotal use <space>\ncotal status [--space <s>] [--server <url>] [--components]\n```\n\n`meshes` lists the meshes this machine knows; a `*` marks the `current` default a bare\n`cotal spawn` joins. Entries learned from a signed-in account are marked `discovered`. Their\nregistration trust is stored under the account's private auth state, and the registry contains no\nsession token or sentinel credential bytes. Commands resolve the catalog `slug`; a different human\n`name` is rendered only as a label.\n\nA registry record this build cannot use is refused by name, never rendered and never skipped. One\nthat does not parse, or is missing a field every consumer reads (`server`, `mode`, `root`, `ts`,\n`space`), makes every registry command exit 1 with the file's path and what is wrong with it.\nRemove the file or restore the record; nothing repairs or invents a field for you.\n\nAn IdP may advertise a same-origin space catalog during login. Cotal reads the complete snapshot and\nadds every valid registration without a separate `meshes add`. A snapshot younger than five seconds\nis used without a request. After that, commands that resolve a mesh target conditionally refresh the\nsaved catalogs. An operation targeting a discovered space refreshes only that space's account and\nrefuses if that account fails. Operations targeting local or manually registered meshes refresh every\naccount, print one warning for each failure, and continue. `cotal status` refreshes every account,\nnever refuses on a refresh failure, and lists each account as `fresh`, `updated`, `not-modified`,\n`no-catalog`, or `failed` with its error. `cotal sync` bypasses freshness and reports added, changed,\nremoved, unchanged, and name collisions. `--idp` limits it to one signed-in account. It never connects\nto a broker.\n\nThe registry is updated under the same lock that guards the catalog cache, so a command never lists\na discovered space set that another command is still writing. The cache records a fetched snapshot\nas not yet applied before the first registry write and as applied after the last. If a command dies\nor is stopped in between, the next command applies that snapshot again before it can use it, with\nno request inside the freshness window.\n\nThe shared dispatcher applies this preparation to every command that declares both `--space` and\n`--server` as mesh-target flags, including commands registered by other packages and commands that\ndeclare their own equivalent flag objects. Daemon and startup commands that use those names only as\nconfiguration explicitly opt out. Registry-local `meshes add` and `meshes rm` never refresh a catalog.\n\nRun on a terminal with the space or `--server` missing, **`meshes add` is guided**: it asks for the\none thing that cannot be derived (the broker URL), probes it, and tells you what answered - open or\nrequiring credentials. It then offers the spaces your `--root` already holds credentials for, states\nthe mode as a fact about that broker rather than asking, and shows the exact record before writing\nanything. A broker that does not answer, or a space name already registered, becomes a choice rather\nthan an error. Anything you pass on the command line is taken as given and not asked again. Without\na terminal - a script, an agent, CI - nothing prompts and the flag form's errors stand\n(`COTAL_NO_PROMPT=1` forces that too).\n\n`cotal up` and `cotal down` maintain their own records. `meshes add` registers a mesh they cannot\nspeak for: one running on another machine, a shared broker, a hosted space. `--root` is the folder\nwhose `.cotal/auth` holds that mesh's credentials and whose `.cotal/agents` holds its personas.\nThe default is the project you run it in. The registry stores that path, never a secret. `--mode`\ndefaults to `auth` when the root holds the space's account record and to `open` otherwise. The\nbroker is probed before anything is recorded, so a wrong address, or credentials that mesh will\nnot accept, fails here instead of at the first `spawn`; `--force` records without verifying (and\nreplaces an existing record).\n\nA hostname or public address is registrable only when the connection will **require TLS**. Pass\n`--tls`, or use a `tls://` URL. The scheme is recorded as enforced intent, so every later dial\nthrough the record demands the handshake (and `meshes add tls://…` against a plaintext broker is\nrefused at registration). Without required TLS the fence admits loopback and private-overlay\nliterals only. RFC1918 addresses are refused in both modes because a cafe LAN is private but does not belong to you.\n\nA **user-auth** mesh registers from supplied pinned trust, never guessed: `--user-auth-file`\ntakes the bundle exported where the mesh runs; `--from` asks before it dials the address at all,\nthen fetches its `/.well-known/cotal-mesh` discovery document (HTTPS only), displays the pins, and\nasks again before adopting them. Neither fetch follows redirects: a 302 can move a pinned fetch\nonto plaintext or onto another host, so it is refused rather than followed, and the pinned\nexchange must itself be an `https://` URL, except for an exchange on this machine, where plain\n`http://` is accepted for a loopback *literal* (`127.0.0.1`, `::1`, any spelling of them) but not\nfor `localhost`, which is a name rather than an address. Registration verifies that the exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also verifies that the broker refuses a bare\nconnect; that auth-required refusal is the pass. The sentinel credentials land in a 0600 file under\nthe entry's root; the registry records only the path.\n\n`meshes rm` drops records. It never stops a mesh. For a mesh running on this machine `cotal down`\nis the right verb, and `rm` says so unless you pass `--force`. A hand-added record is removed by\n`meshes rm`, by an `add --force` replacement, or by a `cotal up` that actually starts the broker for that same space, server and root, which becomes that\nmesh and so takes the record over (a `cotal up` for that space anywhere else refuses instead).\nNothing that merely *infers* a record is stale from a dead broker touches it: an\nunreachable broker is listed `offline` and stays, whether `cotal up` or `cotal meshes add`\nwrote the record. A bare command does not treat that offline record as a running mesh;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the project\nthey tear down; a hand-added one they leave alone even when it shares a root, because nothing\non this machine could write it back.\n\nA discovered entry belongs to the normalized IdP origin and proved subject that supplied it. Local\nteardown, cleanup, and liveness pruning do not remove it. A manual or locally started entry with the\nsame name wins and remains untouched; that discovered name is reported as a collision. Logging out\nremoves only the discovered entries owned by that account.\n\n`cotal meshes` and `cotal status` print `events: required` for a registration carrying\n`policy: { events: \"required\" }`. On that space, foreground spawn, detached spawn, manager starts,\nand interactive `join` cannot opt out or join without an event plane. `--no-events` is refused with\nthe space named. A connector without an event plane is refused with both the space and connector\nnamed. A session whose own grant omits `events.<owner>.<actor>` is refused before joining and the\nmessage names a full-row `actor grant` repair.\n\n`use <space>` sets that default; the selection applies from every directory,\nincluding inside another mesh's project. `status` is a read-only report: machine prerequisites\n(starting with the installed `cotal-ai` version), the installed extensions and their versions, this\nfolder's `.cotal/`, the recorded meshes, and a live snapshot of the selected mesh (roster, channels,\nmembership feed). Stale Claude skills and out-of-date `.agents` skills recommend `cotal setup --skills`,\nnot unscoped `cotal setup`. `status` takes `--space` / `--server` to pick the mesh to inspect; it starts\nnothing.\n\nIf a refresh fails, `status` may still show the kept catalog bytes for diagnosis. It labels them\nstale with the last successful snapshot timestamp and the refresh error. It never calls that state\nsynchronized or online. If a selected discovered space vanishes from a successful snapshot, the\nselection is cleared and the command reports that no default is selected.\n\nFor a user-auth mesh the selected-mesh section reports the login `status` works as: the signed-in\nsubject when this machine holds a cached session for the entry's pinned IdP, or the exact `cotal\nlogin --idp <url>` line when it does not, with no network round trip either way. A locally\nprovisioned space also shows the actor grant row; a discovered or registered remote entry reports\nthe grant as not checkable on this machine, because the ledger runs where the space was\nprovisioned. `--components` on a user-mode target probes as that same signed-in login (`ps`'s\ncredential), never a static mint; when the login cannot supply a credential, the row says why\ninstead of printing the broker's refusal of an unauthenticated probe.\n\nPersona rows name the catalog they describe. If this folder and the selected mesh use different\ncatalogs, status names both and marks which one spawn launches from. A green `default` means the file\npasses the same agent-file loader spawn uses; a present but invalid file is reported as invalid.\n\n`cotal status --components` adds a fail-loud per-component health pass. It reads **each\ncomponent's own control surface**, rather than treating a PID, a lease, or a successful probe of a\nsibling as proof that the component serves. It prints one of `serving`, `absent`, `not-serving`, or\n`refused` for each component and exits `0`, `1`, `2`, or `3` respectively (the highest observed\nstate wins):\n\n- **manager**: local PID record, its liveness-lease holder and PID, then the manager's own typed\n `status` service reachability from this host. Manager builds that do not report static\n reconciliation say `static reconciliation not reported by this manager build`; the line stays\n visible even when the manager is otherwise `serving`.\n- **delivery**: local PID record, its ready lease (`ready` is the daemon's own bound-control\n signal), and the latest `renewal.<spaceKey>.json` adoption verdict, the record of the space the\n command was asked about, keyed per space the way the pidfiles are. A re-signed credential and a\n broker-accepted adoption stay distinct facts. A root-only `renewal.json` left by an older build\n names no space and is never read as any space's verdict (`doctor auth` names it as a leftover).\n- **web**: local PID record and the dashboard's own loopback `/api/meta` response, which must name\n the same PID and its requested port. A different process on the port, an unreadable PID command,\n or an unrecognizable process record is `refused`, not a green default-port guess.\n- **broker**: the registered mesh URL dialed from this host with its recorded TLS requirement.\n\n`absent` means Cotal has no live local component record (or has a stale record); `not-serving`\nmeans the component record is live but its service/readiness surface did not answer or is not ready.\nThose are intentionally separate exit cases. A failed or unreadable probe is `refused`, never an\nabsent component or a clean zero.\n\n## spawn\n\n```bash\ncotal spawn [<persona>] [--detach] [--name <n>] [--agent <a>] [--model <m>] [--variant <v>] [--prompt <text>] [--cwd <dir>]\ncotal spawn -f <cotal.yaml> [--dry-run]\n```\n\nFor a foreground spawn onto a remote user-auth mesh, a launcher may supply a one-time enrollment\ninstead of a cached human login. Prefer a private file:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./seat.md --space main\n```\n\nThe file contains only the enrollment URL, ending with at most one line terminator, and must be\nmode `0600` on POSIX. An orchestrator that cannot mount a file may set `COTAL_ENROLLMENT_URL`\ninstead; that value is redeemed byte for byte, so a trailing newline in it is refused. Setting both\nis refused. Enrollment input\nrequires `--space` and applies only to a foreground persona spawn. If the mesh is not registered yet,\nthe enrollment response must carry the stock user-bundle fields and the command needs\n`--config <persona-file>` because there is no local remote-mesh persona catalog to read. The client\nredeems the URL once, registers the returned mesh material, exchanges the returned actor token at the\npinned auth service, and removes both enrollment variables before starting any child process.\n\nA cached login for the same IdP and an enrollment are conflicting proofs, so the command refuses\nrather than choosing one. An invalid enrollment never falls back to login provisioning. Unknown,\nexpired, revoked, and already-used enrollments all produce one response: ask the owner for a fresh\none. See [Enrollment redeem](identity-and-auth.md#enrollment-redeem) for the HTTP contract.\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | resolved mesh | Target space |\n| `--server <url>` | registry entry | Broker URL override |\n| `--creds <path>` | none | Control-caller creds for an off-registry manager (`--detach` only) |\n| `--name <n>` | persona's `name:` | Presence-name override (does not choose the persona) |\n| `--config <persona-or-path>` | none | Persona catalog name or file path; wins over the positional |\n| `--agent <a>` | persona's `agent:`, else `COTAL_DEFAULT_AGENT`, else `claude` | Connector type (`claude`, `opencode`, `jcode`, `hermes`, and so on) |\n| `--role <r>` | persona's `role:` | Role override |\n| `--model <m>` | persona's `model:` | Model override |\n| `--variant <v>` | persona's `variant:` | Model variant override (connector-defined; e.g. OpenCode reasoning tiers) |\n| `--cwd <dir>` | this cwd | Working directory to root the agent at |\n| `--prompt <text>` | none | Initial prompt auto-submitted at start |\n| `--resume <id>` | none | Fork an existing session id into the mesh; only connectors that declare resume support accept it (see [the matrix](connectors.md)) |\n| `--no-events` | event plane on where supported | Opt out of the session's structured event plane (`--events` only restates the default) |\n| `--share-tools <sel>` | none | Share named operator MCP servers with the agent |\n| `--subscribe <a,b>` | persona's | Channel read-set override |\n| `--allow-subscribe <a,b>` | = subscribe | Read-ACL override |\n| `--allow-publish <a,b>` | deny | Post-ACL override |\n| `--detach`, `-d` | off | Launch via the manager into a detached PTY (reattach with `cotal attach`) |\n| `--on <instance>` | class anycast | With `--detach` only: pin the launch to one manager instance id (the whole id, as `ps` prints it). Refused on a foreground spawn (no manager to pin), with `-f` (a manifest deploy launches through the manager class queue), and when empty |\n| `--file <cotal.yaml>`, `-f` | none | Deploy a manifest onto the running mesh |\n| `--dry-run` | off | With `-f`: print the plan, mutate nothing |\n| `--allow-stale <a,b>` | none | With `-f`: waive named stale agents (apply-only) |\n| `--runtime <name>` | manifest's | With `-f`: override the manifest's runtime |\n\nEach session uses its connector's **event plane** by default: a stream of structured events\ndescribing what the agent did, rather than the prose it wrote, on a channel of its own. The channel is named after\nthe agent's principal, `events.<owner>.<actor>`, never after its display name, because two live\nagents are allowed to share a display name and would then share a stream. The launch grants publish\nrights on that channel alone, foreground and detached alike. `--no-events` is the explicit opt-out\nunless the selected registration says `policy: { events: \"required\" }`. Required policy makes the\nevent arm and grant mandatory, so `--no-events` and connectors without an event plane are refused.\n\nThe launch decision and the grant are separate on purpose. Holding publish rights on a channel is\nnot a request to publish to it, so writing an event channel into an agent file's `allowPublish`\ndoes not override `--no-events`.\n\nThe persona (`--config` > positional > `COTAL_DEFAULT_PERSONA` > `default`) is loaded from the\ntarget mesh's `.cotal/agents/`; the launch flags override the file. Foreground runs the agent\nattached to your terminal; `--detach` hands the launch to the running manager. Both modes get the\ndurable backstop on a mesh that runs the delivery daemon; `--live-only` skips it for a foreground\nspawn (messages posted while it is disconnected are then not replayed). A foreground exit retires\nthe agent's creds and broker footprint, like a manager despawn. On a user-auth mesh the two arms\ndiffer: a spawn against a mesh this machine provisioned revokes the actor row on exit, while a\nremote spawn (an enrollment or the advertised provisioning endpoint) removes only this machine's\ncredential files; its grant stays until the mesh operator revokes it, and the launch line says\nwhich arm you are on. A `--detach` spawn is an\n**action**: the manager accepts it and returns the allocated identity at once, then the launch\nfollows to a terminal outcome rather than blocking (see [the control surface](control-surface.md)).\nSee [Connect Claude Code](connect-claude.md) and [Agent files](agent-files.md); `-f` is a\n[manifest deploy](#manifest-deploys). (`cotal start` was merged into `cotal spawn --detach`.)\n\n## models\n\n```bash\ncotal models [--agent <connector>] [--refresh]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--agent <connector>` | all registered connectors | Connector whose catalog to list |\n| `--refresh` | off | Ask the connector to refresh its provider cache |\n\nAsks the running manager for each connector's model catalog (model ids plus their variants)\nfor connectors that expose one. OpenCode and Codex query harness/provider surfaces; Jcode reads\nproviders that enable `model_catalog = true` in the operator Jcode `config.toml`. Jcode's listed\neffort tiers render as `variants (declared, not provider-verified)`, and launch can still refuse one.\nA connector without a catalog says so. Pick a result with `cotal spawn --model <id> --variant <v>`,\nwhere `<id>` is the model id as the catalog printed it. OpenCode and Codex ids are the full\n`provider/model`; Jcode ids are bare (`opus-5`, not `cliproxy/opus-5`), because the provider is\nselected by the operator's Jcode config and a prefixed id is refused at launch with the bare form\nnamed.\n\n## endpoints\n\n```bash\ncotal endpoints [--space <s>] [--server <url>] [--creds <path>]\n```\n\nLists the mesh presence roster: agents, the manager, and any other protocol endpoint, with each\nendpoint's role, kind, status, and current activity. Unlike `ps`, this is a read-only presence view;\nit is not limited to child processes owned by the manager.\n\n## Endpoint control\n\n```bash\ncotal describe <endpoint> [--space <s>]\ncotal invoke <endpoint> <command> [--args '<json>'] [--space <s>]\ncotal invoke <endpoint> <command> --name <agent> [--admin] [--space <s>]\n```\n\nThe generic v0.4 service surface. `describe` resolves a registered endpoint's command set off the\nwire - the reserved `describe` command answers the registered contract digests, the schemas are\nfetched from the space's content-addressed contract store, recompiled, and verified against those\ndigests - and prints each command with its capability class and targeting shape. `invoke` calls one\ncommand by name: `--args` is a JSON object validated against the fetched input schema *before*\npublish; a targeted command takes `--name <agent>` (resolved to the agent's current principal through\n`inspect`) or `--self`. `--admin` uses the admin instrument credential, whose cross-agent reach rides\nthe operator-only `any` authorization mode. Neither command has compile-time knowledge of any\nendpoint's schemas - this is the same trust chain every built-in control command now uses. Needs an\nauth mesh: the manager registers its service on both static and per-user meshes (a signed-in user\nrides their bearer; each visible or invoked command still requires its existing grant, and cross-agent\nreach needs the `admin` scope). An open mesh has no service registry.\n\n## Managed seats\n\n```bash\ncotal ps [--on <instance>] [--wide | --json] [--space <s>]\ncotal stop --name <n> [--on <instance>] [--space <s>]\ncotal attach --name <n> [--on <instance>] [--no-reconnect] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | none | Managed agent to stop / attach (required) |\n| `--on <instance>` | class anycast (`ps`: class scatter) | Pin to one manager instance id (multi-manager space); takes the whole id as `ps` prints it, not a prefix. An empty value (`--on \"\"`, an unset shell variable) is refused, never treated as absent |\n| `--wide` (`ps`) | off | After each seat's compact row, print extra operational facts the manager records: `cwd`, `pid`, spawner, lifecycle uid, and the owning manager's instance id and host. Model and requested variant stay in the identity row rather than printing twice. A fact the manager did not record (for example a runtime with no real process) prints nothing, never a placeholder |\n| `--json` (`ps`) | off | Machine-readable: one JSON object per seat per line, copied unchanged from the manager row. Instance headers and errors go to stderr, so stdout contains only rows. Mutually exclusive with `--wide` |\n| `--no-reconnect` (`attach`) | off | End the attach when its session ends, instead of re-establishing it. For scripts that want one run and one exit code |\n\nThe human `ps` row is presentation text and is not a stable parsing target. Scripts use `--json`,\nwhich is the machine-readable row contract.\n\nThese are operator clients over the running manager's control plane. The default row includes the\nconnector, model pin, optional requested variant, and runtime as operational descriptors for the\nmanaged row. They do not make a shared display name a unique protocol identity; use `--json` when\nunambiguous owner+actor attribution is required. An omitted variant means no override was requested;\nCotal does not invent an effective provider default it cannot observe. `ps` also prints two state\nfacts per managed agent, because they answer different questions: the process fact from the manager's\nown runtime handle (`running` with its uptime, or `exited` with how long it ran), and the mesh fact\nfrom the roster (`idle` / `working` / `waiting` / `mesh offline`, or `not in roster` when the seat has\nno presence row at all: a seat that has not joined yet, or one that never did). A seat can be\n`running` and `mesh offline` at once: the process is alive and its presence has lapsed. The mesh fact\nis only a verdict while the manager's own presence watch is fresh: when that watch has been silent past\nthe liveness window, or has not replayed the bucket yet, every row prints `mesh unknown` with the reason\ninstead (`--json` carries it as `meshView: stale | unpopulated`), because `offline` and `not in roster`\nwould then describe the manager's watch rather than the seat. The manager rebinds a watch that goes\nsilent under a live connection on its own, so `mesh unknown` normally clears within a liveness window.\nOn a user-auth mesh `ps` also renders each managed agent's last credential-refresh outcome, fail-closed.\n\n**Mode split (chosen up front, never try-scatter-then-degrade):**\n\n- **Static / open mesh.** Bare `ps` is a **class scatter**: it freezes the live manager class from\n the records registry, merges every registered instance's agents grouped and attributed per\n instance, and a non-answering instance is shown as `registered, no answer within the deadline`\n (never silently omitted). A refused list or a missing answer makes the census incomplete: rows\n from other instances remain visible, but `ps` prints an incomplete-census warning on stderr and\n exits non-zero, including with `--json`. Those rows are not a complete seat count. A contract\n mismatch prints one plain comparison of the requested and served input/output digest pairs and\n advises aligning manager versions. The no-answer label means only that the instance is registered\n and did not answer. It does not say the host is down, because a dead host never deregisters itself and a\n live one can be slow; if it is gone, deregister it.\n `--on <instance>` pins the read to one exact instance id instead. A wrong pin fails loud\n rather than falling through: a well-formed id that no live manager carries is reported as\n `manager instance <id> did not answer` (nothing else is asked), and a credential without that\n instance's rail is reported as refused by the broker, not as an unresponsive manager. A manager\n that answers with a refusal is shown with its own cause; \"no manager reachable\" is said only when\n nothing answered at all. If the scatter's own registry read fails (the freeze or the reconcile),\n `ps` says the manager registry could not be read rather than pronouncing on the managers, which\n may all be up.\n\n**The verdict is scoped to the endpoint rail the request rode.** An issued caller rides the\nversioned `ep.v1` rail, a separate subject space from the legacy `ep` rail, and an endpoint serves\nboth (SPEC 13.15). A manager older than the versioned rail serves `ep` alone, so it can be running,\nregistered and answering while an issued caller's request reaches nobody. Silence on `ep.v1` is\nreported as `no manager answered on the ep.v1 rail` and names both causes it is consistent with:\nno manager running, or one older than the rail. The CLI cannot tell them apart, because the service\nregistry records no package version, so check whether a manager is running and, if it is, its\nversion. The same scoping applies to `cotal run`'s hosted verbs, which drop the `--local`\nsuggestion there, since `--local` drives the run from the calling process and names the caller as\nits answerer.\n\n**`stop` and `attach` route by seat locality.** A seat can only be stopped or attached by the\nmanager actually running it, and the class queue does not know which one that is. So on a\nstatic/open mesh both verbs first ask every registered instance which one hosts the named seat, then\naddress that instance directly. This happens by default; you do not need `--on`.\n\n`--on <instance>` remains the override, for when you already know where the seat lives or the\nlookup itself is degraded. It is also the **only** route on a **user-auth mesh**: a ledger-scoped\nbearer does not hold the registry-read rows the lookup needs, so there the verbs stay on the class\nqueue unless you pin them yourself.\n\nA seat is reported as **not found** only when every reachable instance answered for itself. An\ninstance that stayed silent past the deadline, or that refused the read rather than answering, said\nnothing about which seats it hosts, so the seat may be running on it. That case reports that the\nlocation could not be established, names the instances that did not answer, and states outright\nthat it is not a report that the seat is gone. Read it as unknown and retry with\n`--on <instance>`; a retry loop that treats it as \"already gone\" stops looking for a seat that is\nstill running. A single manager cannot tell \"hosted elsewhere\" from \"does not exist\": it answers\n`not-found` for both, which is why the search asks all of them and why an incomplete search\nconcludes nothing.\n- **User-auth mesh.** `cotal ps` reports what **one** manager knows about your agents (an `ep.one`\n read against the manager's in-memory roster, owner-filtered). It does **not** report other\n manager instances. It cannot tell you that one is down: an unreachable manager is absent\n from the list. Completeness across a multi-manager user-auth space is not claimed.\n A manager that does not answer fails the command outright (exit non-zero), rather than printing\n an empty list that could be read as \"no agents\". Your ledger row needs the `admin` scope to\n reach `ps` at all; `spawn` alone is refused by the broker (the ep tier boundary).\n\n`attach` streams and drives an agent's terminal on the `pty` runtime; detach with the escape key\n(Ctrl-] by default; see [`COTAL_DETACH_KEY`](config.md)). It does so over a one-use, holder-bound\nmesh session ([SPEC](../SPEC.md) §13.6): the manager replies with a signed session grant (never a\n`127.0.0.1` URL), the CLI redeems it once over the broker, and the browser console (`cotal console`)\ndrives the same session. `stop` and `attach` need a running manager to talk to. On a static mesh\nthey are cross-agent admin operations. On a user-auth mesh, your own agents (any agent under your\nowner) need only the `spawn` scope; another owner's agent needs `admin` on your ledger row\n([identity & auth](identity-and-auth.md)). Launch detached agents with [`spawn --detach`](#spawn).\n\n**`attach` reconnects when the link dies.** A session lives on a network link, and a laptop that\nsleeps, a VPN that drops or a wifi handover kills it. When that happens `attach` prints\n`[cotal: connection lost, reconnecting]` on stderr and starts asking the manager for a new session:\na fresh grant, a fresh per-session credential, a fresh connection, so every attempt re-runs the same\nauthorization the first attach did. On success it prints `[cotal: reconnected]`, the manager repaints\nthe seat's current screen the way it does for any attach, and you carry on in the same terminal.\nRetries wait 1s, 2s, 5s, 10s, then 30s, for as long as the seat exists. The detach key is read the\nwhole time the loop runs, the waits and the attempts alike, so a reconnect never traps you: press it\nwhile a session is being established and the attach ends there, and a session that lands behind the\npress is handed back to the manager rather than left holding a slot. Everything else you type while\nthere is no session is dropped rather than queued, so keystrokes aimed at a terminal that turned out\nto be frozen, Ctrl-C included, are not delivered to the agent by a reconnect you did not know had\nhappened. That starts before the first session, not at the first reconnect: at a terminal, `attach`\nreads and drops what you type while it is still resolving the mesh, so a key struck at a prompt that\nhas not come up yet does not reach the agent when it does.\n\nA **pipe** carries script input. For example, `printf 'ls\\n' | cotal attach --name web` is\nbuffered until the session opens. Buffering continues across reconnects, so\n`tail -f log | cotal attach --name web` does not lose the part of its feed written while the link was\ndown. Only a terminal gets the reader; `--no-reconnect` keeps the old behaviour on both.\n\nIt stops on its own when reconnecting cannot help, and says why: a manager that refuses the attach\nexits non-zero with the manager's own message, and a reconnect that finds the seat no longer there\n(despawned, or its agent exited while the link was down) exits cleanly with `seat <name> is gone`.\nA refusal that could still pass, such as a manager at its session ceiling, is relayed in the\nmanager's own words while the loop keeps trying, once per refusal rather than once per attempt.\nPressing the detach key, or the agent's process exiting while you are attached, ends the attach as\nit always did. `--no-reconnect` turns all of this off and restores the single-session behaviour,\nwhich is what a script wants.\n\nEach reconnect also hands the abandoned session back to the manager, over the first link that can\ncarry the message, so an attach that flaps does not eat the manager's session slots one outage at a\ntime. If that message never gets a link, the attach says so when it ends. The live-session ceiling\ndefaults to 64 concurrent sessions (`--max-sessions`); the browser console opens one session per\npane, so a dashboard over a large mesh should size for agents × panes. Hitting the ceiling refuses\nbefore a credential is minted and names `--max-sessions`.\n\nWhich mesh `attach` resolves also decides **how it redeems the grant**. On a registered open mesh\nthere is no local seed. The CLI connects bare, the same way other control commands already do, and\nthe session rail is the caller rail that a real open-mode connection already reaches. Telling the\noperator to re-register the root is false: the registered root is already the contract. On a\nstatic-auth mesh the grant is still redeemed by minting a short-lived\nsession-scoped credential from the seed at the root the mesh resolved to, never from a `.cotal`\nfound by walking up from whichever directory you happen to be standing in. The difference is not\nhypothetical: `~/.cotal` exists on every install because the mesh registry lives there, so a command\nrun anywhere under your home directory but outside a project used to mint from your home\ndirectory's trust and present it to a broker that trusts a different chain, which surfaced as a\nbare authorization failure that named nothing. A directory that does hold another chain for the\nsame space is now reported on the way past, and not obeyed:\n\n```text\n! this directory resolves to /Users/you, whose .cotal/auth holds a DIFFERENT trust chain for space \"team\".\n attach used /Users/you/projects/app, the root this mesh resolved to. The other one is not being used, and is worth a look.\n```\n\nWhen a **static-auth** mesh holds no seed at the resolved root, `attach` refuses and names what it\nresolved, the broker and the root, instead of describing a directory it did not use and instead of\ntaking the open-mode path. An authenticated registry entry with a missing seed is still\nauthenticated. A USER-AUTH mesh still refuses loud: two-step user-mode redemption is not wired.\n\nTerminal bytes stream over the mesh; the manager's own HTTP/WS face serves the console. That endpoint binds\n**loopback by default**, so nothing is exposed by accident; `cotal up --host <addr>` passes its bind\naddress down, which is what lets you attach to an agent whose manager runs on another machine. A\nbare `cotal supervise` and an embedded manager stay machine-local. Set it directly with\n`supervise --console-host <host>`.\n\nThat address is **recorded on the mesh** and carried forward, because it is a decision rather than\nsomething later commands can work out for themselves (a broker dial address is not a manager bind\naddress). Every later manager launch for the same mesh reuses it, including a same-root `cotal up` repair,\nan adopted preserved or restored listener, and a `spawn -f` manifest deploy. A manager replacement\ndoes not quietly move a reachable attach face back to loopback. Passing `--host` again overrides it,\nso you can widen or narrow exposure whenever you like; a mesh that never asked stays loopback-only\nand records nothing.\n\nBecause that face carries terminal read and write for every managed agent, it is credentialed in two\ntiers. A mesh caller receives a **ticket** bound to the single agent the manager just authorized,\nsingle-use and short-lived, so one authorized attach can never be re-pointed at someone else's\nagent. The **console token** is the operator's own, reaches every agent, and is printed only to the\nmanager's output. The roster, the live feed, and the PTY stream all answer `401` without one; the\nstatic console shell is served openly, since it describes no agent.\n\n## input\n\n```bash\ncotal input --name <n> --text <text> [--no-enter] [--on <instance>] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which manager to reach |\n| `--name <n>` | | Managed agent to type into (required) |\n| `--text <text>` | | The text to type, taken verbatim (required) |\n| `--no-enter` | off | Type the text and stop there, without pressing Enter |\n| `--on <instance>` | class anycast | Pin to one manager instance id using the same rules as [`attach`](#managed-seats) |\n\nTypes one line into a running agent's terminal, as if you had typed it there, and returns. This is\nthe half of [`attach`](#managed-seats) that a program wants: `attach` is a live stream that holds a\nsession open and expects a terminal on your side, so a script, a cron job or a web UI cannot use it\nto send a single line. `input` is one authorized call.\n\nWhat it is for is **harness commands**. A line beginning with `/` is not chat and not a message: it\nis something the agent's own harness handles, and the only way in is the keyboard.\n\n```bash\ncotal input --name reviewer --text \"/compact\" # ask the harness to compact its context\ncotal input --name reviewer --text \"/model opus\" # switch its model\ncotal input --name reviewer --text \"hold on that PR\" # ordinary typing works too\n```\n\n**Quoting.** `--text` takes a value, so a payload starting with `/` survives as written. A payload\nstarting with a dash needs the `=` form, because the shell-style `--text --foo` is ambiguous and is\nrefused rather than guessed:\n\n```bash\ncotal input --name reviewer --text=--verbose # dash-leading text: use --text=<value>\n```\n\nEnter is pressed by default, since a command typed but never submitted has not been delivered.\n`--no-enter` types the text and leaves it sitting at the prompt, which is how you stage a line and\nsend it later.\n\nNothing comes back but a delivery receipt (`✓ sent 9 bytes to reviewer`, counting the trailing\ncarriage return). Whatever the agent does next shows up where its output already goes: the mesh, its\ntranscript, or an `attach`.\n\n**This one is operator-only, and more narrowly than `stop` or `attach`.** Those two are granted to\nanything holding `spawn`, so an agent can stop and attach to seats under its own owner. `input` is\nnot: it is granted only to operator credentials, which on a user-auth mesh means your ledger row\nneeds the `admin` scope, the same scope [`ps`](#managed-seats) already needs there. The reason is\nthat a write into a terminal is control of whatever is running in it, and on a user-auth mesh the\nown-owner rule covers every seat under you, not only the ones you launched: a `spawn`-scoped agent\ncould otherwise type into a sibling it never started. Seat locality is still resolved for you.\n\nOnly the `pty` runtime can be typed into. The external terminal runtimes (`tmux`, `cmux`, `orca`,\n`herdr`) attach to a process they do not own, so they have no input stream for it and the command\nrefuses by name rather than dropping the keystroke.\n\n## personas\n\n```bash\ncotal personas list [-v] [--running]\ncotal personas show <name>\ncotal personas edit <name>\ncotal personas new <name> (--prompt <t> | --from <f>) [--role <r>] [--model <m>]\ncotal personas rm <name> --force\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh's persona catalog |\n| `--role <r>` | none | `new`: the persona's role |\n| `--model <m>` | none | `new`: the persona's model |\n| `--prompt <t>` | none | `new`: the persona's prompt text |\n| `--from <f>` | none | `new`: seed the prompt from a file |\n| `--verbose`, `-v` | off | `list`: include role / model / description |\n| `--running` | off | `list`: mark personas live on the mesh |\n| `--force` | none | `rm`: required, delete without prompting |\n\nPersonas are the local agent files under the resolved mesh root's `.cotal/agents/`, the same catalog\n`cotal spawn` launches from. `--space` and `--server` therefore move every list, read, write, delete\nand completion operation to the selected mesh. An unresolved target refuses rather than falling back\nto the current directory. See [Agent files](agent-files.md) for the file format.\n\n## supervise\n\n```bash\ncotal supervise [--runtime <name>] [--space <s>] [--server <url>] [--spawn <names>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space to supervise |\n| `--server <url>` | hosting mesh, or matching registered mesh | Broker URL. A registered mesh supplies it when omitted; a different explicit value is refused before anything is dialed. |\n| `--runtime <name>` | `pty` | Agent runtime (`pty` built in; extension runtimes are explicit-only) |\n| `--console-port <n>` | none | Protocol-console port |\n| `--console-host <host>` | loopback | Bind host for the console + attach endpoint. Loopback keeps it machine-local; `cotal up` passes the address it bound the broker to, which is what lets `cotal attach` reach this manager from another machine |\n| `--max-sessions <n>` | 64 | Live-session ceiling. Each console pane and each `cotal attach` is one session, so size for agents × panes, not agent count. A capacity refusal names this flag. `cotal up --max-sessions` records the same number on the mesh so a later `supervise` started by repair or `spawn -f` keeps it |\n| `--roster <file>` | none | Declarative roster to boot at startup |\n| `--launch <spec>` | none | Resolved manifest launch spec (from `up -f` / `spawn -f`) |\n| `--spawn <names>` | none | Comma-separated personas to pre-spawn at startup |\n\nThe manager is the agent supervisor and control plane: it answers `spawn --detach`, `stop`, `ps`,\n`attach`, and the `cotal_*` manager tools. `cotal up --detach` starts one for you; run `supervise`\ndirectly to recover a dead manager or drive a custom runtime. Default runtime is `pty`; install an\noptional provider first (`cotal ext add @cotal-ai/orca`, `@cotal-ai/tmux`, `@cotal-ai/cmux`, or `@cotal-ai/herdr`) and\nselect it explicitly. A missing provider or app fails loudly; there is no fallback. See [Deploy](deploy.md).\nBoot inventory decides whether this process takes unpinned `spawn`/`launch` on the class rail:\nif every declared connector is unavailable, those commands stay on this instance rail only\n(`status` reports `classSpawn: false`). `describe` still answers on the class rail, so an\nunpinned spawn can bind-fence against a skip member; re-issue, or pin `--on`. A partial\ninventory keeps the class rail and names `--on` on a harness refusal, because sibling\ninventories are not readable from the serve credential. See [control surface](control-surface.md#instance-routing).\n\nOn a normal `SIGINT`/`SIGTERM`, the manager stops every seat and requires the selected runtime to\nprove the seat is gone before it releases the manager lease or service registration. A stop that\ncannot prove exit fails loud and keeps manager authority instead of reporting a clean shutdown while\nan orphan still holds broker rails. After an abrupt manager death, the same logical successor\nterminalizes only its own durable static slots, verify-evicts the predecessor's broker principal,\nrecords that result in the lifecycle's caller-readable audit detail, reaps the predecessor's seat\nprocess through the runtime's custody reference recorded on the slot (the pty runtime verifies the\nprocess start identity in its seat record, so a reused pid is never signalled), and only then\nretires the lifecycle and frees the alias. A runtime that custodies its seats reserves that\nreference before it launches one, and the manager records it on the slot's first durable row, so a\nmanager that dies part-way through a spawn also leaves a seat its successor can address. A spawn\nthat launched its seat and then failed is rolled back by the manager that launched it, and that\nrollback reaps the seat through the same reserved reference before the lifecycle retires. Missing or unverified broker evidence keeps the slot\nterminalizing, and so does a runtime that cannot reap by reference.\n\nA `meshes add --mode user` entry is a **participant** registration, not hosting authority. A\nparticipant may run `supervise` only when the host advertises the remote manager authority service\nand the signed-in actor has the dedicated `supervise` ledger scope. The CLI obtains the closed,\nloopback-only `manager-service` view; `spawn` and `admin` do not substitute for that scope. The\nhost issues the manager's public-nkey JWT material through its lifecycle-bound prepare → activate\n→ renew protocol, never by handing the participant a signer or static provisioner credential.\n\nThe broker URL in the registry entry decides the transport. A remote broker is often published\nover a `wss://` edge rather than a raw `nats://` port, and `supervise` dials whichever scheme the\nrecord holds, starting with the manager-authority registration it runs before the manager exists.\nThe record also decides whether that registration requires TLS, so a participant never downgrades\nthe credential exchange to a plaintext connection the registry did not describe.\n\nWithout that advertised host service or scope, `supervise` refuses before it starts a manager.\nRun `cotal spawn` without `--detach` to launch a foreground agent, or ask the space host to enable\nthe authority service and grant `supervise` for detached agents. If a running remote manager loses\nrenewal, it reports degraded state and refuses unsafe new starts and restarts; live agents are not\nsilently replaced. Do not run `cotal down` or `cotal up` on a participant machine to repair this\ncondition.\n\n## service\n\n```bash\ncotal service install [--mesh <name>] [--linger]\ncotal service status [--mesh <name>] [--json]\ncotal service uninstall [--mesh <name>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--mesh <name>` | this folder's mesh | The mesh whose manager the service runs; one unit per mesh |\n| `--linger` | off | install: also enable user lingering so the user manager starts at boot and the service survives logout. Never enabled silently |\n| `--json` | off | status: machine-readable output |\n\nRuns the manager as a user service so it survives logout and reboot. On Linux this installs a\nsystemd user unit (`~/.config/systemd/user/cotal-manager@<key>.service`, where `<key>` is the\ncase-safe mesh key); on macOS a launchd agent plist under `~/Library/LaunchAgents/`. Any other\nplatform, or an absent systemd/launchd user session, fails with a message naming what is missing.\n\n`install` resolves the mesh from the registry and binds the unit to that entry's root and broker\naddress, so it can be run from any directory. The mesh must be registered (`cotal up` or\n`cotal meshes add`) before installing; an unregistered name refuses before anything is written.\n\nThe unit's `ExecStart` is the bare `supervise` command. The mesh facts travel in a `0600`\n`EnvironmentFile` (`COTAL_SPACE`, `COTAL_SERVER` pinned to the registered broker URL, whatever\nport it listens on) rather than the command line, because command lines are readable by every\nuser on a multi-user host. The same file gives the service a private `COTAL_HOME` and\n`XDG_CONFIG_HOME` under the unit directory, so the service manager never touches the login\nuser's `~/.cotal`. First-run connector seeding runs synchronously inside `service install`,\nagainst that private config root; the unit itself starts with `COTAL_SKIP_CONNECTOR_SEED=1`\nso a manager is never interrupted mid-seed by a restart. An install whose pre-seed cannot\ncomplete (network unreachable, registry error) refuses instead of deferring.\n\nEvery value the unit derives from a path (`WorkingDirectory`, the `EnvironmentFile` path, the\n`ExecStart` tokens) is escaped for systemd specifiers (`%` becomes `%%`), so a mesh root that\ncontains `%` starts over its real path instead of a path systemd rewrote by expanding it. The\nprovenance comment records the root unescaped.\n\n`service install` also refuses while a manager is already running for the mesh (`cotal down\nmanager` first). The restart policy is `Restart=always` with `RestartSec=20s`, chosen for\nmanager units in production: a manager exits for reasons that are not failures (broker\nrestarts, host suspend), where `on-failure` with a short interval thrashes.\n\n`service status` reports the unit state from systemd/launchd, the manager's own health read from\nits pidfile at the unit's recorded root, and the machine facts a hosting side asks for:\narchitecture, OS (the platform, never the hostname), whether `/dev/kvm` is present and\naccessible, CPU count, and total memory. `--json` returns the same fields as one object.\n\n`service uninstall` stops and disables the unit and removes it plus the private state directory.\nIt works from any directory: the unit's own records name the mesh and root it serves, and an\nexplicit `--mesh <name>` selects it. It refuses any unit that was not written by `service\ninstall` (the files carry a provenance comment), whose recorded mesh is missing, or that was\ninstalled for a different mesh, so operator-written units are never destroyed; `service status`\napplies the same rule and never reports a mesh a unit does not record.\n\nThis command installs only the manager. The per-space auth service and the delivery daemon are\nnot installed by it: on a shared broker an operator runs three units per space with `After=`\nedges (auth service, then manager, then delivery) and stops them in reverse. A broker-side `cotal\nup` unit is a separate unit documented in [Run a mesh](run-a-mesh.md).\n\n## reconcile-gate\n\n```bash\ncotal reconcile-gate [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the frozen gate lives in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint whose gate is frozen |\n| `--instance <id>` | this folder's persisted manager instance | Instance id |\n\n**When you need this.** A manager restart killed after deregistration begins but before the new\nincarnation finishes leaves the endpoint's issuance gate *frozen*, held by a\nprocess that no longer exists. The freeze is what stops two incarnations serving at once, which is\ncorrect. The successor manager now completes that dead registration itself on boot, using the same\nguard this command uses: it acts only when the freeze-holder is affirmatively gone under a complete\nCONNZ sweep (`gone` and `sweepComplete=true`). If that registration's spec write already committed,\nit finishes the same freeze at the committed registration revision. If the spec did not advance, it\nabort-reopens the gate at generation+1 with processEpoch unchanged and continues the normal takeover.\nLive, unknown, unestablishable, and\nwrong-op-kind still refuse; there is no TTL.\n\nUse this command when the boot path cannot run: the delivery daemon is down, the repair targets a\nnon-manager endpoint, or you want to lift the freeze without starting a manager. It checks that the\nholder really is gone, prints what it found, and then finishes the dead operation the same way as the\ninterrupted restart would have: revoke the old credentials, evict their holders with verification,\nand reopen the gate.\n\nIf verification is interrupted, the command leaves the gate frozen and durably records each holder\nwhose eviction was already verified. A retry still repeats the freeze-holder liveness check, then\nskips only progress bound to the same registration operation, frozen-gate revision, and holder set.\nThe output reports holders completed before this attempt, completed now, and still remaining. A new\nfreeze or changed holder set starts from zero. Cursor cleanup happens only after reopen; a retained\ncursor is harmless because its old gate revision cannot authorize a later freeze.\n\n**It refuses far more often than it acts, on purpose**, and always says which check stopped it:\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `holder-alive` | The freeze-holder still has a live connection: a manager *is* running | Stop that process first. Reconciling would evict a live manager's credentials |\n| `holder-unknown` | The connection sweep could not prove the holder absent | Not safe to proceed: an unprovable holder is treated as a live one. Re-run once the broker answers completely |\n| `liveness-unestablishable` | The delivery daemon could not be asked at all | Start it (`cotal up` runs it) and re-run. Silence is never read as death |\n| `not-frozen` / `no-gate` | The gate is open, or there is no gate at that coordinate | Nothing to repair: check `--endpoint` / `--instance` |\n| `wrong-op-kind` | Frozen under a takeover or retirement, not a registration | Out of scope for this command; it will not reinterpret another operation's intent |\n| `eviction-unverified` | The holder looked gone but eviction could not be verified | The gate is left frozen, unchanged. Investigate the broker before retrying |\n| `raced` | A newer manager moved the gate mid-repair | Re-run `cotal doctor` and look again |\n\nThere is no `--force`, and no path that discards gate state: the only way this reopens a gate is by\nproving the holder is gone and then completing the operation properly.\n\n**What reopening the gate does for the endpoint's governance slot.** A registration takes the\nendpoint-wide governance slot before it publishes its spec, and holds it until its gate reopens. An\ninstance that died between those two points leaves the slot held with no registration behind it.\nThis command does not write that slot and never has; the registration path is its only writer. What\nthe reopen does is advance the holder's gate past the generation the slot is stamped with, which is\nwhat marks the slot abandoned. The next registration for that endpoint then reclaims it as part of\nits ordinary start. So the repair here is still one command followed by starting the manager, and\nthe slot needs no separate step.\n\n## deregister-instance\n\n```bash\ncotal deregister-instance [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | this folder's auth space | Space the instance is registered in |\n| `--server <url>` | the local mesh | Broker URL |\n| `--endpoint <e>` | `manager` | Endpoint the instance serves |\n| `--instance <id>` | this folder's persisted manager instance | Instance id, the whole id as `cotal ps` prints it |\n\n**When you need this.** The service registry records *registration*, not liveness, and nothing in\nthe model expires a row. A manager that stops cleanly removes its own registration. One whose host\ndied without writing anything cannot, so its record goes on claiming a live instance forever: every\nclass scatter in that space freezes the dead slot in, and `cotal ps`, `stop` and `attach` each pay\ntheir whole deadline waiting for a machine that is never coming back. A laptop that was reimaged, a\ncontainer that was deleted, a box that will not be back on the network: those registrations have no\nother exit.\n\nThis command is that exit. It asks the instance first, and it removes a record only when the broker\naffirms the instance's own rail is empty: nothing subscribed there. Then it deletes the\nregistration's two records keys, each pinned to the revision it read, and prints what it removed.\n\n**Silence alone never passes.** An unanswered describe is what a dead host, a wedged process and a\nslow one all look like, and a hung process still holds its subscriptions, so the broker sees\ninterest on its rail. That instance is refused and the observation is printed. A dead process holds\nno connection and therefore no subscription, so a real corpse is still removed.\n\n**Every refusal names the failed check:**\n\n| Refusal | What it means | What to do |\n|---|---|---|\n| `instance-answered` | The instance answered a pinned describe. It is alive | Nothing to repair. If it is wedged rather than gone, stop the process first; its own clean stop removes the record |\n| `instance-not-affirmed-gone` | It did not answer, and the broker did not report its rail empty, which is what a held subscription looks like: slow or hung, not affirmed gone | Nothing was removed. Stop the process; its record goes on its own clean stop, or re-run this once it is down |\n| `liveness-unestablishable` | The probe itself failed, so nothing was learned | Fix the probe's path (credential, broker) and re-run. A probe that could not run is never read as death |\n| `not-registered` | No registration at that coordinate | Check `--instance` and `--endpoint`. This takes the whole id, never a prefix |\n| `registration-in-flight` | The instance holds the endpoint governance slot at the live issuance-gate generation, so a registration is still completing | Nothing was removed. Wait for that registration to finish, then re-run |\n| `superseded` | The record moved between the read and the delete | Something is writing to it. Nothing was removed; re-observe before retrying |\n\nThere is no `--force` and no sweep: silence is not death, and a rule that removed rows on silence\nwould eventually remove a live instance that was merely slow. An operator names one instance, the\nbroker's verdict on its rail is what authorizes the removal, and the guard's job is to show them\nthey named a dead one. Removal is not a one way door either. The same instance re-registers over\nthe tombstone on its next start, under the same identity.\n\n## runtimes\n\n```bash\ncotal runtimes\n```\n\nLists every agent runtime the manager can spawn through: the built-in `pty`, the official providers\n(`orca`, `tmux`, `cmux`, `herdr`), and any custom provider installed via `cotal ext add`. Each installed\nprovider is probed so you can see what is actually reachable on this machine before selecting it:\n\n```\npty built in\norca installed · reachable @cotal-ai/orca\ntmux available · cotal ext add @cotal-ai/tmux\ncmux available · cotal ext add @cotal-ai/cmux\nherdr available · cotal ext add @cotal-ai/herdr\n```\n\n`installed · reachable` / `unreachable` is the provider's own `available()` probe; `available` means\nit is a known runtime you can add with the shown command. Selecting an unknown or uninstalled runtime\nvia `up`/`spawn --runtime <name>` fails loud and, for a known one, points at the exact `cotal ext add`\npackage. There is no silent fallback to `pty`.\n\n## send\n\n```bash\ncotal send dm <agent> \"<text>\" [--space <s>] [--server <url>] [--creds <path>]\ncotal send msg <channel> \"<text>\"\ncotal send ask <role> \"<text>\"\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and (off-registry) which credential |\n\nOne-shot messaging: connect, send a single direct message (`dm`), channel post (`msg`), or role\nask/anycast (`ask`), then exit. For a running conversation, agents use the mesh tools instead\n([MCP tools](mcp-tools.md)).\n\n`cotal send` works from an operator shell or from a seat. It uses `cotal-send` as the advisory\ndisplay name. The wire principal comes from the resolved operator credential or user bearer, not\nfrom `COTAL_NAME`, `COTAL_ID`, `COTAL_OWNER`, or `COTAL_ACTOR`. On an open mesh the transient\nendpoint self-mints its principal.\n\n## channels\n\n```bash\ncotal channels list\ncotal channels set <name> [--replay | --no-replay] [--window <n>] [--desc <s>] [--instructions <s>]\ncotal channels default --replay | --no-replay\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--replay` / `--no-replay` | none | `set`/`default`: replay history to new joiners, or not |\n| `--window <n>` | none | `set`: replay window size |\n| `--desc <s>` | none | `set`: one-line channel description |\n| `--instructions <s>` | none | `set`: instructions shown to joiners |\n\nInspects and edits the channel registry: replay policy, description, and joiner instructions. ACL\nsemantics (who may read or post) are set at mint / provision time, not here; see\n[Channels and permissions](channels-and-permissions.md). On a user-auth mesh, `list` rides your\nown login as is; `set` and `default` edit the registry over a short-lived\nchannel-writer view, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\nOn a remote user-auth mesh that view is served by the public exchange; space-history `purger`\nand the read-only admin view are not.\n\n\n## history\n\n```bash\ncotal history clear --force [--dms] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Target mesh |\n| `--dms` | off | Also clear DM history |\n| `--force` | none | Required: clear without prompting |\n\nPurges retained channel history; `--dms` extends it to direct-message history. An alias of\n[`clean history`](#clean). On a user-auth mesh the purge rides a short-lived purger view over\nyour login, which needs ledger scope `admin` ([Identity & auth](identity-and-auth.md)).\n\n## console\n\n```bash\ncotal console [--plain] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to watch |\n| `--plain` | off | Line stream instead of the TUI |\n\nA live protocol view for a space: a lazygit-style TUI, or a plain line stream on `--plain`. On a\nuser-auth mesh it rides the read-only admin view over your login, which needs ledger scope\n`admin`. Inside the TUI, operator control (`D` kill, `:spawn`, `:status`, `:purge`) rides the\nsame per-action instrument path as `cotal stop` and `cotal ps`, never the observer; a raw\n`--creds` file cannot drive it. `a` (or `:attach <agent>`) runs\n[`cotal attach`](#managed-seats) in place and returns to the console on detach. See\n[Watch a mesh](watch-a-mesh.md).\n\n## web\n\n```bash\ncotal web [--detach] [--host <host>] [--port <n>] [--no-open] [--space <s>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Space to serve |\n| `--host <host>` | `127.0.0.1` | Concrete HTTP bind and browser host; wildcard addresses are refused |\n| `--port <n>` | `7799` | HTTP port |\n| `--detach` | off | Run in the background; stop with `cotal down web` or bare `cotal down` |\n| `--no-open` | off | Don't open the browser |\n\nThe browser observability dashboard: presence, channels, and a live feed. It is **not** part of\n`cotal up`: it ships inside `cotal-ai` as the `@cotal-ai/web` extension, seeded automatically on first\nrun (like the built-in connectors) so it always matches your CLI version. It self-registers `cotal web`\ninto this surface and serves\n`http://cotal.localhost:7799` by default (loopback; `*.localhost` resolves in Chrome/Firefox/Edge; Safari may\nneed `http://127.0.0.1:7799`). On a user-auth mesh the dashboard rides the read-only admin view\nover your login, and a channel purge asks for its own channel-purger view per click; both need\nledger scope `admin`. The public exchange serves `channel-purger` for a remote owner; it still\nrefuses the startup admin view, so a remote `cotal web` is not a complete channel-management\nsurface. Detached mode re-execs the current Cotal installation, writes diagnostics to\nthe mesh root's `.cotal/web.log`, and reports success only after the HTTP server answers. It requires\na recorded mesh root, but can be launched from any directory once `cotal up` has recorded the mesh.\nSee [Watch a mesh](watch-a-mesh.md).\n\n## mint\n\n```bash\ncotal mint <name> [--profile <agent|observer|admin>] [--out <path>] [--signer]\ncotal mint <name> --provision [--role <role>] [--space <s>] [--server <url>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--profile <agent\\|observer\\|admin>` | `agent` | Credential profile |\n| `--out <path>` | `.cotal/auth/creds/space.<key>/<name>.creds` | Output path - the default sits under the resolved space's segment (`<key>` is that space's hex encoding, as in [Project files](config.md#project-files)) |\n| `--signer` | off | Emit a stripped account-signing file instead |\n| `--force` | off | With `--signer`: overwrite an existing file |\n| `--allow-subscribe <a,b>` | the agent file's, else subscribe | Read-ACL override, **agent profile only**: `observer` and `admin` carry a fixed read set, and `mint` refuses this flag there rather than narrowing nothing |\n| `--allow-publish <a,b>` | the agent file's, else deny | Post-ACL override, **agent profile only** |\n| `--role <role>` | the agent file's | Agent profile: the anycast task queue the identity pulls (`svc_<role>`) |\n| `--provision` | off | Agent profile: also pre-create the identity's bind-only DM/deliver durables (and its role's task queue) on the live mesh, so the credential can consume |\n| `--space <s>`, `--server <url>` | the resolved mesh | Which root supplies the agent file, static trust and default credential storage; with `--provision`, also which live mesh receives the durables |\n\nMints a NATS creds file for a space in **static** auth mode, scoped to a profile and (optionally)\nexplicit read/post ACLs. `--signer` emits an account-signing file for delegating minting to another\nhost. A per-user-auth space refuses `mint`: agents there join under a logged-in user\n([`login`](#login) + [`actor grant`](#actor)), never via a handed-out creds file. See\n[Identity and auth](identity-and-auth.md).\n\nFor an agent profile, the resolved mesh root supplies the persona ACL, the signing material and the\ndefault credential destination as one authority. If the current folder also holds trust for a\ndifferent space or account, mint refuses before writing and names both roots. It never combines a\npersona from one root with credentials signed or stored under another.\n\nA plain mint is creds only: the identity can publish within its post ACL at once, but on an authed\nmesh its DM inbox and task queue are provisioner-pre-created and bind-only, so a **consuming**\nconnect fails until they exist. `--provision` performs that pre-create in the same command (a\nprovisioner cred is minted from the space's trust material, used, and dropped), so a long-running\nclient you start yourself can receive DMs and role anycasts like a spawned seat. The command prints\nthe identity's principal (its wire id) and lifecycle uid; a consuming client passes that uid as its\n`lifecycleUid`. Agent profile only; an open mesh needs none of this (peers self-create there). The\nThe same resolved authority is used for both the credential and `--provision`, so the broker\nfootprint cannot be created under a different root's trust material.\n\n## Login\n\n```bash\ncotal login --idp <auth base URL> [--client-id <id>]\ncotal logout --idp <auth base URL>\n```\n\nSigns you in to a per-user-auth mesh's IdP (device code flow) and caches the session; run it\nonce per machine. It prints your IdP subject, the id the operator grants against. When the trusted\n`/token` response advertises a same-origin space catalog, login validates and records that account's\nspaces immediately. After a\nlogin, every command on that mesh works under your identity: each connect takes a fresh IdP\nproof, exchanges it locally for a short-lived bearer, and is authorized against the actor\nledger at connect time. `logout` revokes the IdP session, clears its cache, and removes only that\naccount's discovered registry entries. See\n[identity & auth](identity-and-auth.md).\n\n## actor\n\n```bash\n# an upsert of the WHOLE row: a flag left off is the WIDE default below, not \"unchanged\"\ncotal actor grant <actor> --sub <IdP subject> [--scope a,b] [--allow-subscribe a,b] [--allow-publish a,b] [--role <r>] [--label <l>]\ncotal actor revoke <actor> (--sub <IdP subject> | --owner <u_…>)\ncotal actor list\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` | the folder's | Space whose ledger to manage |\n| `--sub <subject>` | none | The IdP subject (shown by `cotal login`) the actor belongs to |\n| `--owner <u_…>` | none | The derived owner token (alternative to `--sub`) |\n| `--scope <a,b>` | `spawn,role:default` | Capability scope (`''` = none; `spawn` = may run agents; `role:<r>` = may delegate role r; `admin` = cross-agent control; `supervise` = eligible for the closed remote manager-service view when the host enables it) |\n| `--allow-subscribe <a,b>` | `>` (all channels) | Channel read ACL; the user's envelope, their agents can never read beyond it |\n| `--allow-publish <a,b>` | `>` (all channels) | Channel post ACL; also the envelope for their agents' posting |\n| `--role <r>` | none | Role (scopes the task-queue consumer) |\n| `--label <l>` | none | Display label for `actor list` (never the IdP subject) |\n\nThe actor ledger is the single authorization source of a user-auth space: no row, no access.\nA bare `grant` is the **full** envelope (all channels, may spawn); the flags narrow it. A\nre-grant **replaces the whole row**, not the one field you name, so to add a capability spell\nevery field out: the new scope plus the row's current read set, post set, role and label\n(`cotal actor list` shows what a row holds). A field left off does not stay as it was, it\nreverts to the wide default in the table above, which is how a narrow reader becomes a reader\nof every channel. A re-grant retires the current interactive lifecycle through the running auth\nservice before it rotates the row, so copied bearers cannot cross an authorization update. If that\nretirement cannot be confirmed, the row is left unchanged and the command fails with the recovery\naction. `revoke` uses the same retirement before deleting the row, which lets a later grant create a\nreal successor instead of colliding with a live predecessor. `supervise` is separate from `spawn` and `admin`: it only makes a signed-in\nperson eligible for the host-provided closed remote manager-service view; it does not grant\nmanagement of another owner or a general host profile. `revoke` denies the next exchange and\nthe next connect with no restart, and evicts the principal's live connections. Managed-agent rows\n(written by the spawn path) live in a disjoint row space this command never touches. See\n[identity & auth](identity-and-auth.md).\n\n## doctor\n\n```bash\ncotal doctor auth [--fix]\n```\n\nCredential-health diagnosis and repair for this folder's mesh: renders every managed\ncredential as healthy / near-expiry / expired and ends in `healthy` or the exact next\ncommand; `--fix` applies the repairs it can. The one surface every stale-credential error\npoints at.\n\n## join\n\n```bash\ncotal join --space <s> --name <n> [--role <r>] [--channel <c>]\ncotal join --link <url> | --token <t>\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--space <s>` / `--server <url>` / `--creds <path>` | resolved mesh | Which mesh, and which credential |\n| `--name <n>` | none | Your presence name |\n| `--role <r>` | none | Your role |\n| `--channel <c>` | none | Channel to join |\n| `--kind <k>` | `agent` | Endpoint kind |\n| `--link <url>` | none | Join link (`cotal://…`) |\n| `--token <t>` | none | Join token |\n| `--lifecycle-uid <uid>` | none | Required with `--creds`: the lifecycle UID minted alongside the credential (`COTAL_LIFECYCLE_UID` works too). A credential's durable grants name exact lifecycle-keyed resources, so `join` refuses to invent one |\n| `--tls` | off | Connect over TLS |\n\nAn interactive presence: join a space under your own name and role, without launching an agent\nharness. A `--link` or `--token` supplies the where and the auth in one value. See\n[Spaces](spaces.md) and [Identity and auth](identity-and-auth.md).\n\n## Manifest deploys\n\nA `cotal.yaml` manifest declares a whole mesh (channels, personas, roles, and ACLs) in one file.\nThree commands consume it, plus a read-only validator:\n\n```bash\ncotal up -f cotal.yaml # boot a fresh mesh from the manifest\ncotal spawn -f cotal.yaml # deploy the manifest additively onto a running mesh\ncotal down -f cotal.yaml # tear that deploy down (or --run <id> for one run)\ncotal topology view -f cotal.yaml # validate + view the access graph, change nothing\n```\n\n`up -f` and `spawn -f` differ in target: `up -f` brings up a new broker and applies the manifest;\n`spawn -f` requires an already-reachable mesh and applies additively (ownership-scoped). On a\nuser-auth mesh, `spawn -f` deploys over your own login (the deployer view, gated on ledger scope\n`spawn`): the manifest's agents land under your owner, a manifest claiming another owner is\nrefused, and seeding new channels additionally needs scope `admin`. Both take\n`--dry-run` to print the plan without mutating anything. `topology` validates the manifest and\nrenders its channel / role / ACL graph. See [Define a team](define-a-team.md) and the\n[manifest reference](manifest.md).\n\n## ext\n\n```bash\ncotal ext # same as `list`\ncotal ext add <npm-package>\ncotal ext remove <name>\ncotal ext list\ncotal ext root # print just the install prefix (scriptable)\ncotal ext seed [--repair|--reset|--force]\n```\n\nOperator-installed extensions: `add` installs an npm package into a cotal-owned prefix and records\nevery registry provider it contributes. Commands appear in help, completion, and dispatch; runtime\nproviders are lazy-loaded by commands such as `supervise`; local process providers participate in\n`status` and selective `down`. `remove` and `list` manage them. The `@cotal-ai/web` dashboard is the\ncanonical command/process example. Installed packages and their location are described in\n[config](config.md).\n\nBare `cotal ext` lists the inventory, headed by the install prefix. That prefix is a cotal-owned npm\nroot kept **separate** from npm's own global tree. These packages never show up in `npm list -g`,\n`cotal ext` (or the Extensions section of `cotal status`) is the canonical inventory. `cotal ext root`\nprints only the path, for scripts. The versions shown are the manifest pin recorded at add time.\n\nRemoving an extension that owns a running local process is refused with the mesh root and its\n`cotal down <component>` command; stop it first so uninstalling the package never strands a process\nwhose lifecycle provider is gone.\n\n### Built-in connectors are seeded extensions\n\nThe first-party agent connectors (`claude`, `opencode`, `codex`, `hermes`, `jcode`, `pi`) are not compiled into\nthe binary. They are seeded on first run through the **same** `ext add` path a third party uses, and\nappear in `cotal ext list` like any other extension. So you can remove one you do not want\n(`cotal ext remove @cotal-ai/connector-hermes`), and a deliberately-removed connector STAYS removed\nacross upgrades. `cotal ext add <your-package>` adds a third-party connector the same way. The web\ndashboard (`@cotal-ai/web`, providing `command:web`) is the seventh built-in seeded on the same path.\n\n`cotal ext seed` is the maintenance entry for that seeding (it runs automatically on the first real\ncommand of each boot, so you rarely call it):\n\n| Flag | Meaning |\n|---|---|\n| (none) | Reconcile: seed any never-seeded built-in, refresh a seeded one whose version the binary bumped, leave a removed one removed. A no-op once current. |\n| `--repair` | Recover after an interrupted seed or a lost authority (rebuilds the interrupted connector; restores the removed-vs-never-seeded record from its durable backup). |\n| `--reset` | Discard the record and re-seed all seven built-ins (the six connectors plus the web dashboard). **Resurrects any you removed.** Rebuilds cleanly over corrupt seed state. |\n| `--force` | Re-seed the built-ins even when the version stamp is current or a downgrade. |\n\nWhen a newer `cotal` advances the operator-global seed store to its generation, it prints one\nmigration line naming the old and new generations, the exact CLI entry that wrote the store, the\ncommit timestamp, and `seed/stamp.json`. That writer and timestamp are kept in the stamp, so a later\nolder CLI refusal can say which executable wrote the generation it will not overwrite and when.\nLegacy generation-only stamps remain readable; their refusal simply has no writer provenance to add.\n\nAn older `cotal` refuses a seed store written by a newer version. When it can verify a sufficient\n`cotal` executable on PATH or at the installer's `~/.local/bin/cotal` location, the refusal names\nthat absolute path so a reduced service PATH does not select the older binary again. Otherwise it\nkeeps the generic newer-version instruction. `--force` rebuilds the store for the running older\nversion without discarding the ever-seeded authority. `--reset` still exists for corrupt state and\nresurrects deliberately-removed connectors.\n\nA source-checkout CLI (`pnpm cotal`, `tsx bin/cotal.ts`, `node bin/cotal.ts`, or a suite child of\nthose) refuses to write or garbage-collect that store. The refusal names the path, the generation\nit declined, and `$XDG_CONFIG_HOME` as the isolation remedy. `COTAL_HOME` does not relocate this\nstore. An entry that cannot be proven as a released install is refused the same way. Isolated\nrelease tests that must seed from a checkout-shaped `bin/` set `COTAL_ALLOW_CHECKOUT_SEED=1` after\npointing `$XDG_CONFIG_HOME` at a scratch dir; that override is documented here, not on the refusal\nline. An opt-in write still records the checkout path in `seed/stamp.json` as `writtenBy`.\n\nThe default connector for a bare `cotal spawn` (no `--agent`) is the persona's `agent:` pin if it\nhas one, else `claude`; set `COTAL_DEFAULT_AGENT` (e.g. `opencode`) to change the fallback. It is\na default, so a persona that pins its harness still wins over it. An `--agent` naming a removed\nconnector fails loud with the exact\n`cotal ext add` to restore it. Set `COTAL_SKIP_CONNECTOR_SEED=1` to turn off the automatic first-run\nseed/refresh entirely (for a controlled or offline setup that manages connectors by hand); `cotal ext\nseed` still runs on request. `cotal agent-bearer` never takes the seed at all: it is exec'd by\nspawned seats on every bearer refresh, so it neither reconciles nor is refused by the store's\ngeneration (see [Plumbing](#plumbing)).\n\n## completion\n\n```bash\ncotal completion <bash|zsh|fish|powershell> # print a stub to eval / source\ncotal completion install [shell] # install it persistently\n```\n\nPrints or installs shell completion. Completion candidates come from each command's declared flags\nand, where useful, live mesh state (spaces, personas, managed agents) resolved offline.\n\n## feedback\n\n```bash\ncotal feedback \"<summary>\" [--type <t>] [--email <e>] [--details <text>]\n```\n\n| Flag | Default | Meaning |\n|---|---|---|\n| `--type <t>` | none | `bug` \\| `idea` \\| `friction` \\| `praise` \\| `other` |\n| `--details <text>` | none | Longer free-form details |\n| `--severity <s>` | none | `low` \\| `medium` \\| `high` |\n| `--area <a>` | none | The part of Cotal this concerns |\n| `--email <e>` | git email | Contact email (required on the keyless public path) |\n| `--name <n>` | none | Your name (optional) |\n| `--url <url>` | keyed / public intake | Intake URL override |\n| `--key <k>` | `COTAL_FEEDBACK_KEY` | Feedback key |\n\nSends feedback to the Cotal developers. With a key (`--key` / `COTAL_FEEDBACK_KEY`) it routes to the\nkeyed beta intake; without one it goes to the public `cotal.ai` intake and requires a contact email\n(`--email` / `COTAL_FEEDBACK_EMAIL`, else your git email). Run a self-hosted intake with\n[`feedback-intake`](#server-daemons).\n\n## run\n\nOperate durable workflow runs (cotal-lang programs) from the terminal.\n\n```bash\ncotal run start --file <program> [--timeout <dur>] [--local]\ncotal run resume <runId> [--local --file <program>]\ncotal run ps [--endpoint <ep>]\ncotal run journal <runId> [--endpoint <ep>]\ncotal run answer <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]\ncotal run migrate <runId> --local --file <program> [--endpoint <ep>]\n```\n\n`start` hands the program to the mesh's manager, which validates it, mints the run id (the record\nnever takes a caller-supplied one), drives it in its own process, and answers with the id once the\nrun is recorded; a program that does not validate is refused with every problem listed. `resume`\nasks the manager to take an existing run back and continue it from its step journal; the source is\nthe recorded program, so no `--file` is taken. Neither takes `--endpoint`: the manager records\nits runs under its own endpoint, and naming another is refused. `ps` lists the run records and\n`journal` renders one run's durable records; both only inspect. An open pause prints its question.\nA pause settled with an accepted answer prints its value as JSON plus the recorded answerer,\nartifact when present, time, and answer id. Expired pauses and ordinary steps print no answer line.\n`answer` resolves an open\ncheckpoint through the manager, presenting as the holder that armed it; the manager records the\nanswerer from your credential, so no `--by` is taken there. `migrate` runs the migrate check of an\nedited program against a run's journal, from this terminal under a read credential (`--local`\nonly; the manager serves no run-migrate command): it prints whether the migration is admissible,\nevery orphaned step with its verdict and code, and exits 0 on admissible and non-zero on not. It\nwrites nothing: the commit that would file the migration is not reachable yet, and the report\nsays so. `--timeout` sets the default\ncheckpoint timeout for a drive (default 1h). `--local` drives in this process instead, over one\nconnection per invocation under the run's own credential minted from the project folder's trust\nmaterial, and is the path on a bare broker with no manager or for a run with no recorded program\n(`cotal run resume <runId> --local --file <program>`); `answer --local` takes `--by <who>`. A\nuser-auth mesh runs no programs yet: the manager refuses the family by name, and `--local` has no\ncredential there. The guide is [workflows](workflows.md).\n\n## Server daemons\n\nTwo long-lived infra roles ship with the CLI. They are not part of everyday operation; the delivery\ndaemon comes up automatically with `cotal up --detach` in auth mode.\n\n```bash\ncotal deliver --space <s> [--server <url>] [--creds <file>]\ncotal auth-service --space <s> --server <url> [--port <n>] [--exchange-public-port <n>] [--exchange-public-url <https://…>] [--exchange-trusted-proxy]\ncotal feedback-intake --keys <keys.json> [--port <n>] [--creds <file>]\n```\n\n`auth-service` runs a user-auth space's identity plane: the NATS auth callout, the\ncapability-gated local exchange and JWKS, and, when `--exchange-public-port` is set, the closed public\nexchange/discovery face forwarded by an HTTPS reverse proxy. `--exchange-public-url` is the proxy URL\nadvertised to clients; `--exchange-trusted-proxy` opts into last-hop `X-Forwarded-For` attribution.\n`cotal up --user-auth` starts and supervises the service for you, so you run it directly only to\nrecover one by hand.\n\n`deliver` runs the server-side Plane-3 delivery daemon: the durable backstop and membership/ACL\nauthority. It is auth-mode-only and single-instance (`--shard`/`--shards` accept only `N=1`);\n`--dev-mint` mints a scoped cred from the local signer for standalone dev. `--creds` can start a\ndaemon that already looks healthy, but production renewal is not that file alone: the manager and\nthe daemon must address one credential store. On a stock split host with two project roots, a\ndirect `deliver` is not an independent repair; keep the daemon under `cotal up` on the broker\nhost, or inject the same store into both processes ([embedding](embedding.md#supervisor-signing-authority)).\nSee the [delivery daemon](delivery-daemon.md). `feedback-intake` runs a self-hosted feedback server\n(requires `--keys` and a scoped `--creds`), announcing submissions into a space channel; flags\ninclude `--host`/`--port`, `--store`, `--space`/`--channel`, `--max-bytes`, and `--rate-limit`.\n\n## Plumbing\n\n`cotal __complete <words…>` is the internal entry the shell-completion stubs call to emit candidates\nfor the current command line; you never run it directly. `cotal agent-bearer` is machine-facing\nplumbing on user-auth meshes: spawned agents exec it to print a fresh short-lived bearer from their\nspawn-time secret; you never run it directly either. Its local arm uses `--dir` to discover the\ncapability-gated loopback service. A remotely enrolled, already-granted agent instead receives\n`--exchange-url <https://base>` in its launch argv: that arm sends `{owner, actor, actorToken}` to the\npinned public exchange with no local capability, follows no redirects, and refuses every non-HTTPS\nURL because the actor token is the credential in the request body. Because a seat execs it on every\nbearer refresh, it skips the connector-seed boot gate entirely: it reads one 0600 token file,\nexchanges it and prints the bearer without consulting or writing the operator-global seed store, so\na newer store generation cannot refuse a live seat's refresh. (`cotal start` is a removed tombstone: it\nerrors and points you to `cotal spawn --detach`.)\n"
74
74
  },
75
75
  {
76
76
  "slug": "config",
@@ -84,7 +84,7 @@ export const DOCS_BUNDLE = {
84
84
  "title": "Connect Claude",
85
85
  "kind": "Guide (informative)",
86
86
  "summary": "The Claude Code connector turns a real claude session into a Cotal mesh peer.",
87
- "body": "# Connect Claude\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nThe Claude Code connector turns a real `claude` session into a Cotal mesh peer. A bundled\nplugin inside the session joins NATS, maps lifecycle hooks to presence, and exposes the\nmesh tools. Nothing wraps Claude; it is an ordinary session that happens to be on the\nmesh.\n\nThe shared mesh runtime (agent, `cotal_*` tools, hook relay) lives in\n[`@cotal-ai/connector-core`](../extensions/connector-core); this connector is the thin\nClaude-specific adapter over it. Siblings: [OpenCode](connect-opencode.md) (beta),\n[Hermes](connect-hermes.md) (alpha), [pi](connect-pi.md) (alpha); the\n[Connectors](connectors.md) matrix compares them feature-by-feature.\n\n## Set up\n\n```bash\ncotal setup # one-time: installs the plugin, seeds one agent; launches nothing\ncotal up # brings up the mesh + delivery daemon + a detached manager\n```\n\n`cotal setup` installs the cotal plugin (so the repo's Claude sessions get the `cotal_*`\ntools) and seeds one `default` persona; `cotal up` brings up the local stack so\n`cotal spawn --detach` / `cotal_spawn` work right away. Re-running either is idempotent.\nThe install mechanics and the invariants behind them are in\n[setup internals](setup-internals.md).\n\n`cotal setup` also installs Cotal's authored Agent Skills (`SKILL.md`, the agentskills.io format) for\ncoordinating agent teams (today `team-topology`), from one canonical source, on two channels:\n\n- **Claude Code** gets a second, skills-only plugin, `cotal-skills`, from the same `cotal-mesh`\n marketplace, at **user scope** (machine-wide). The Claude connector declares and implements this\n setup provider, including the marketplace assets and native plugin commands; the base CLI only passes\n the vendor-neutral Agent Skills directory. The plugin carries no code and no core dependency,\n and uninstalls on its own with `claude plugin uninstall cotal-skills --scope user`. Its plugin version\n is stamped from the running CLI release, so an upgrade + `cotal setup --skills` runs `claude plugin update` and\n the deployed install actually gets the new skill. `cotal setup` installs it on first run and on repeat\n runs, so upgraders are not left behind. `cotal status` points a stale or missing skills plugin at\n `cotal setup --skills`.\n- **Every other harness** (Codex, Cursor, OpenCode, Gemini CLI, Windsurf/Devin) reads the cross-vendor\n `~/.agents/skills/` directory convention, which has no remote index, so `cotal setup` **reconciles** it\n (and `cotal setup --skills` does only that):\n it installs/updates each Cotal skill, backs up a copy you have edited to `SKILL.md.bak` before\n replacing it, and removes a Cotal skill that is no longer shipped. Only skills Cotal owns are touched;\n your own or third-party skills there are left alone. `cotal status` reports whether the drop is current,\n stale, missing, or has a retired skill to reconcile, and names `cotal setup --skills` as the remedy. This is the working cross-vendor path.\n\nCotal also generates an [Agent Skills discovery index](https://cotal.ai/.well-known/agent-skills/index.json)\non cotal.ai, but that RFC is still a draft with no harness consuming it yet, so it is a forward bet,\nnot a channel to rely on today.\n\n## Spawn a session\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn dave --detach # supervised: the manager runs it in a PTY\n```\n\nA spawn resolves a persona from `.cotal/agents/<name>.md` ([agent files](agent-files.md));\n`--model`, `--variant`, `--cwd`, `--prompt`, ACL overrides, and `--share-tools` apply to\nboth forms ([run a mesh](run-a-mesh.md) has the full resolution rules). The session joins\nwith identity from its environment and auto-registers presence by the time it is\ninteractive.\n\nInside the session, the agent orients with one read-only tool, `cotal_orientation`: its\nidentity, the channels it reads and may post to, its capabilities, the tools available,\nwho's present, and unread counts. The full tool surface is the\n[MCP tool catalog](mcp-tools.md). In auth mode the team-supervision tools\n(`cotal_spawn` / `cotal_persona` / `cotal_personas`) are injected **only** for personas declaring\n`capabilities: [spawn]` (the same grant that opens the privileged control subject), so an\nagent's toolset matches its declared capabilities. `cotal_run` is gated separately by\n`run`; use `capabilities: [spawn, run]` for both. Fresh setup defaults include both.\nSee [workflow tool setup](workflows.md#from-an-agent-session) for a first run and missing-tool checks.\nClearing retained history is\noperator-only ([run a mesh](run-a-mesh.md)), never an agent tool.\n\n## How it binds\n\nClaude Code exposes four integration surfaces, and three of them collapse into a single\ndual-purpose MCP server:\n\n| Surface | Mechanism |\n|---|---|\n| Outbound, ambient | `http` lifecycle hooks → POST to the connector (presence, activity) |\n| Outbound, deliberate | MCP tools `cotal_send` / `cotal_dm` / `cotal_anycast` (+ `cotal_feedback`) |\n| Inbound, pull | MCP tool `cotal_inbox` (same server) |\n| Inbound, push | Channel nudge + hook drain (below) |\n\nThe manager launches the *real* `claude` (no wrapper):\n\n```\nclaude --strict-mcp-config --mcp-config '{\"mcpServers\":{\"cotal\":{…}}}' \\\n --dangerously-load-development-channels server:cotal\n# env: COTAL_SPACE, COTAL_NAME, COTAL_ROLE, COTAL_CHANNEL=1, plus claude's documented auth vars\n```\n\n- **Model auth.** Locally, `claude` still reads macOS Keychain / `~/.claude`. In a container or\n CI there is no Keychain, so the connector forwards the documented credential set:\n `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`), `ANTHROPIC_API_KEY` /\n `ANTHROPIC_AUTH_TOKEN`, and the cloud-provider flags plus their credential vars. Host-session\n markers (`CLAUDE_CODE_CHILD_SESSION`, `CLAUDECODE`) stay out so a nested seat still saves a\n transcript. See [Deploy](deploy.md).\n- **Persona privacy.** The persona body is written to a private file and Claude receives only\n `--append-system-prompt-file <path>`. The body never appears in the spawned process argv. The\n carrier is a 0600 file inside a 0700 directory on POSIX, with equivalent owner-only ACL hardening\n on Windows.\n- **MCP isolation.** A spawned agent runs with **only** the cotal MCP server:\n `--strict-mcp-config` ignores every other MCP source, crucially the operator's personal\n `~/.claude.json` servers (several spawns each booting a heavy helper would starve\n memory). Share your own servers deliberately (see below).\n- **Installed plugin.** The plugin is installed once (`claude plugin install\n cotal@cotal-mesh --scope local`) because its hooks bind only to an *installed* plugin.\n In a clone the marketplace is the repo's `.claude-plugin/marketplace.json`; `cotal setup`\n (npx, no clone) materializes the same marketplace under `~/.cotal/claude-plugin/` (each plugin dir is\n rebuilt from scratch and atomically replaced, never merged, so no stale file rides in). The\n `cotal-skills` plugin installs from that same marketplace at user scope (`claude plugin install\n cotal-skills@cotal-mesh --scope user`); its manifest and install behavior ship inside the Claude connector, and\n its version tracks the CLI release so updates land.\n- **Identity-gated.** Connector code requires `COTAL_NAME` *or* `COTAL_LINK`. A plain\n `claude` with no `COTAL_*` env stays inert and never joins, so your own sessions in a\n repo do not appear as stray peers.\n- **Hands-free.** The dev-channels flag prints a one-time confirm prompt. The PTY runtime waits for\n the dialog title in normalized terminal output and presses Enter once when it appears, so startup\n speed does not affect a supervised launch. If the declared prompt never appears, the seat exits\n with a bounded error naming the unmatched prompt instead of hanging silently.\n\nInbound mesh messages arrive in context as\n`<channel source=\"cotal\" from=\"bob\" kind=\"dm\" …>…</channel>`: each meta key a tag\nattribute the agent can read for routing.\n\n## How messages reach the session\n\nDurable deliveries land in the connector's inbox from JetStream consumers\n([SPEC §8](../SPEC.md#8-nats--jetstream-binding)); live channel traffic can instead arrive\nthrough an at-most-once core subscription. A durable message sent while the agent is busy\nor offline waits on the stream. Two things move a message from inbox to model; one\ndelivers, the other only wakes:\n\n- **Hook drain (delivery).** `SessionStart` / `UserPromptSubmit` hooks read automatic inbox items and\n inject them as `additionalContext`. This is the single authoritative path: deterministic and works\n on any Claude Code build. Quiet ambient is excluded and stays buffered for `cotal_inbox`.\n A message is **acked only once the hook reply carrying it has cleared both legs of its journey**:\n the connector's control socket to the hook process (which gives up after 2s), and the hook\n process's own stdout to Claude Code (which it force-exits 1s after starting to write). The relay\n sends a receipt back down the control socket from that stdout write's callback, and only on a\n clean write (a runtime whose pipe has gone away fails it), and the connector treats that receipt,\n not its own socket write, as delivery. So a large injection killed mid-flush, or one written to a\n broken pipe, leaves the message un-acked and JetStream redelivers it. What this does *not* prove is\n that Claude Code read or applied the reply: a payload small enough to fit the pipe buffer is\n reported written the moment the kernel takes it. That residual is why the path errs toward\n at-least-once rather than treating a confirmed write as a confirmed read. Acking when\n the reply was merely *formatted* meant a lost reply was a lost message: it was already marked\n handled, so its own redelivery was silently acked on arrival.\n This errs toward **at-least-once**: if a reply lands but its confirmation does not, the batch is\n surfaced again and flagged as a possible repeat. A duplicate injection is noise; a buried DM stops\n the peer answering at all.\n- **Channel nudge (wake).** An arriving message fires a `notifications/claude/channel`\n event that wakes an *idle* session into a turn, so the drain runs *now* instead of at\n the next prompt. The nudge never acks anything. A nudge that the host rejects is retried with a\n bounded backoff while anything is still pending. For an idle session it is the only wake source,\n so dropping it means silence until someone types. When the channel becomes active, the connector\n first re-fires a focus mention remembered during startup, otherwise one buffered wake. A rejected\n push keeps its bounded retry, and JetStream redelivery remains the durable backstop for unacked\n inbox items. If the channel cannot run at all, delivery still waits for the next hook. Live-only\n traffic has no durable retry.\n\n**Two priority tiers.** A *directed* message (DM, anycast, or a channel message that\n`@mentions` us) always nudges. *Ambient* channel chatter does not nudge mid-turn; it\naccumulates, and the `Stop` → idle transition fires one batch nudge so the backlog drains\ntogether.\n\n**Constraints (accepted).** Channels are a Claude Code research preview (≥ v2.1.80;\npermission relay ≥ v2.1.81): Anthropic auth only, admin-enabled on Team/Enterprise, and a\ncustom channel needs the `--dangerously-load-development-channels` launch flag. The hook\ndrain does not depend on any of that; the channel only adds \"wake me when idle.\"\n\nThe same channel also relays **tool-permission requests** onto the mesh, so a peer (a\nhuman at the CLI, a policy node) can approve or deny an agent's pending tool call through\nCotal rather than a per-terminal prompt.\n\n### Attention\n\nAn agent picks how aggressively peer traffic reaches it with\n`cotal_status({ attention })` (three modes, orthogonal to presence):\n\n| arrival | open (default) | dnd | focus |\n|---|---|---|---|\n| directed (dm / anycast) | wake + inject | wake + inject | wake + inject |\n| channel `@mention` | wake + inject | wake + inject | ack-drop; wake to *pull*; not injected |\n| ambient channel chatter | wake when idle; hold while working | never wakes; injects next turn | ack-drop; recall via `cotal_inbox` |\n\nPer-channel overrides refine this: **quiet** (delivered, never wakes; `@mention` still\nwakes) and **muted** (dropped on receive, mentions included; DMs/anycast unaffected), set\nwith `cotal_channel_mode` or as agent-file defaults (`quiet:` / `muted:`,\n[agent files](agent-files.md)). A per-channel override is the final word for that channel.\nQuiet ambient is pull-only: it never hitchhikes on a human prompt, DM, mention, or other\nconnector-driven turn. `cotal_inbox` explicitly surfaces and clears it. A quiet-channel\n`@mention` remains automatic and injects normally.\n\nA pull is bounded too, and clears only what it hands over. One `cotal_inbox` call carries at most a\nreceivable window (direct messages and role requests first, then channel traffic, replayed history\nlast); whatever does not fit stays buffered, is named in the reply, and comes back on the next call.\nA message too large for one whole response is never consumed at all: it is named with its sender and\nsize and left buffered, because clearing what cannot be delivered is the loss this bound exists to stop.\nThat matters most on the path where it is easiest to lose mail: reconnecting brings a channel-history\nreplay with it, so the largest payload and the least expendable message arrive in the same read.\n\nThe local inbox is bounded. On pathological overflow it evicts pull-only items before automatic\ntraffic. If the bounded live/durable classification guard also fills, the connector fails closed:\notherwise-normal ambient becomes pull-only until restart. Muted hard-drop and normal focus recall\nstill take precedence. Focus also keeps a bounded exclusion list so mode toggles cannot recall\nquiet/muted traffic; if that safety bound fills, recall skips the affected channel and reports it\nas incomplete rather than risk resurfacing excluded content.\nIf the separate hard-drop disposition guard fills, channel traffic is dropped for the rest of the\nsession rather than risk a late copy bypassing an earlier muted/focus decision; DMs and anycast are\nunaffected.\n\nAttention is **advisory UX, not a boundary**: any peer can wake a dnd/focus agent by\nnaming it, and `muted` means \"I opted out of receiving\", not \"the channel is blocked\";\nthe broker still authorizes and delivers. Focus's real effect is shrinking the\nuntrusted-ambient injection surface (only subject-authenticated dm/anycast auto-inject).\nIt resets to **open** on `SessionStart`, so a restarted agent never stays silently deaf.\nYour attention is mirrored into presence so peers can see it.\n\nWhatever does reach a turn is framed so a peer cannot write the frame. A line that begins at column\nzero is written by the connector; one message is one line plus indented continuations, with the\nsender inside a single bracket pair. A message body, a sender name and role, and a service or\nchannel label are all peer-controlled, so each passes through the same neutralization the\n`cotal_inbox` reply uses: no line break a splitter may honour and no bracket survives into a\nrendered attribution. This matters more for an injected block than for a reply, because the agent\ndid not ask for it and so never had the chance to distrust it.\n\n## Presence mapping\n\nThe connector wires a small subset of Claude Code hooks to presence states; presence is\ncoarse, and \"what it is doing\" rides on activity updates. Presence is **advisory**: a presence\npublish that fails (the endpoint mid-reconnect, say) is swallowed and never prevents the same hook\nfrom delivering messages or flushing held ones.\n\n| Hook | → state |\n|---|---|\n| `SessionStart` | `idle` (join; surfaces the inbox; captures the live model into `meta.model` when no pin) |\n| `UserPromptSubmit` | `working` (turn starts; surfaces the inbox) |\n| `PreToolUse` | no change; records *what* is about to run, so a permission wait can name it |\n| `Notification` (`permission_prompt` / `agent_needs_input`) | `waiting` with condition `approval` / `input` (activity leads with the pending tool, e.g. `Bash: git push …`) |\n| `Stop` / `StopFailure` | `idle` (turn done / died on an API error; flushes anything held while busy). `StopFailure` also relays Claude Code's native error value as `condition.source` and maps it to the closed condition vocabulary. On the [event plane](#event-plane) it closes the run with `RUN_ERROR`. |\n| `SessionEnd` | `offline` (graceful leave) |\n\n`StopFailure` maps `rate_limit` and `overloaded` directly; auth and credential failures to\n`auth`; account and billing failures to `billing`; `invalid_request` to `request`;\n`model_not_found` to `model`; `server_error` to `server`; `max_output_tokens` to `context`; and\n`unknown` to `failed`. The native value remains in `condition.source`.\n\nHooks are relayed over the connector's **authenticated** local control endpoint (per-user\nsocket + per-launch token, constant-time checked), so a local process that finds the path\nstill can't drive presence or stop the agent. The full Claude Code hook-event list lives\nwith the adapter:\n[`extensions/connector-claude-code`](../extensions/connector-claude-code/README.md).\n\n## Event plane\n\nA spawned session publishes a **structured** account of what it\ndid: run boundaries per turn, assistant text, reasoning, and each tool call with its start\nand its end. Not prose about the work, the work itself, in a vocabulary a program can\nread. The launcher sets `COTAL_EVENTS` by default; pass `--no-events` to opt out on an unrestricted\nspace. A user-auth registration with `policy: { events: \"required\" }` carries `eventsRequired` in the\nprivate launch material, so the connector arms even without `COTAL_EVENTS`; `--no-events` is refused.\nA hand-driven user-mode session may carry the same decision as `COTAL_EVENTS_REQUIRED=1`. Its own\npublish grant must cover `events.<owner>.<actor>` or the connector refuses before joining. An unmanaged\nsession with no launch material and no required-policy fallback keeps the generic default behavior.\n\nA new session includes its first run even when Claude writes a positional startup prompt before the\nconnector receives `SessionStart`. That from-zero read is keyed only to Claude's explicit\n`source: \"startup\"`; resumed, forked, cleared, and compacted sessions adopt at the current transcript\nboundary and do not republish retained history. Crash recovery follows the cursor already stored in\nthe event write-ahead log, regardless of the new process's startup label.\n\nClaude starts each hook in its own process, so a prompt or stop relay can reach Cotal before the\n`SessionStart` relay. The connector holds those event flushes and the terminal until `SessionStart`\nsupplies the source, then enqueues adopt, flush, and close in that order.\n\n`SessionStart` can also run before the connector process has bound its local control socket. The hook\nthe `SessionStart` relay retries only transient pre-connect listener errors, with capped backoff\ninside its existing two-second budget. Later hooks and permanent local faults still fail open\nimmediately. Once a socket has connected, a broken exchange is not retried: the connector may\nalready have handled the frame, so replaying it could apply one lifecycle event twice.\nThat retained `SessionStart` can itself arrive before Claude creates the transcript path. A genuinely\nnew startup waits up to five seconds for that file with capped backoff, and the same deadline bounds\none stalled file read; expiry fails loud instead of silently losing the first run. A forked session\ngets the same wait, because Claude copies the parent transcript into the fork's own file after the\nhook, and then adopts at the end of that copy. Resumed, cleared and compacted starts and recovered\ncursors still require their existing source at once.\n\nTool arguments (`TOOL_CALL_ARGS`) and tool results (`TOOL_CALL_RESULT`) are not republished\nonto this channel. The durable emitter drops those events before they are written to the\nwrite-ahead log, because this channel's read ACL is not the ACL the tool ran under. Content is\nmandatory on both kinds, so the event is suppressed rather than emptied or replaced with a\nplaceholder. Tool start and end still go out. A restart that finds a pending pre-fix frame\nstill carrying those kinds HALTS rather than republishing it.\n\nThe channel is **`events.<owner>.<actor>`**, named after the session's principal. What the actor\nhalf is depends on the mesh, and the difference matters when you go looking for it: on a static mesh\nit is a key the manager allocated, never the display name, so two live agents sharing a display name\ndo not share a stream; on a user-auth mesh it is the agent's own name, because that is what the\nledger row is keyed on. Spelled out again with both halves below. The launch grants publish rights\non that channel alone. A spawn\nthat asks for a *different* agent's event channel is refused at the door rather than granted, since\nthat channel is that session's event stream. The same rule runs on restart: a manager\nresume document that names another agent's event channel is refused rather than adopted, because the\nmanaged row is re-armed from that document and the credential is re-minted from the row.\n\nThe rule reads a **concrete** channel, two principal tokens and nothing else. A pattern such as\n`events.<owner>.>` is not an event channel to it and passes untouched, governed by ordinary ACL\nauthority: on a user mesh the delegation envelope, on a static mesh the spawning credential itself.\nThat is deliberate, because the pattern is the form an operator writes on purpose for an observer,\nand it is worth knowing rather than assuming the fence is total.\n\nTo let something else read a plane, grant it out of band. The refusal prints the command for the\nmesh it is running on, spelled out in full, and only that one.\n\nOn a **user-auth** mesh:\n\n```bash\ncotal actor grant <reader> --owner <owner> --scope '' --allow-subscribe 'events.<owner>.<actor>' --allow-publish ''\n```\n\nEvery field, deliberately. `actor grant` is an upsert of the whole row, and an omitted flag is not\n\"leave it alone\": it is the wide default, `>` read, `>` post, and `spawn,role:default` scope. A bare\n`cotal actor grant <reader>` therefore grants a reader of every channel in the space, which is the\nopposite of what a scoped watcher is for.\n\nOn a **static** mesh there is no actor ledger for `actor grant` to write to, and the refusal says\nso; mint the reader instead:\n\n```bash\ncotal mint watcher --profile agent --allow-subscribe 'events.<owner>.<actor>' --provision\n```\n\nThe **agent** profile, not the observer one. `mint` reads `--allow-subscribe` only for that\nprofile, and refuses it anywhere else: `--profile observer --allow-subscribe <channel>` exits\nnon-zero and writes no creds file, because the observer profile carries a fixed read set over the\nwhole chat plane, which is the opposite of what a scoped watcher is for. The agent profile also prints the lifecycle uid the\nreader needs, since an authed consuming endpoint refuses to start without one.\n\nTwo things a reader has to do that are not obvious, both on `CotalEndpoint`. It must pass the event\nchannel in `channels`: an endpoint reads the channels it lists, so one constructed without\nthe event channel joins nothing and the frames never arrive. And it reads history with `readHistory(channel)`, the delivery daemon's mediated read, not\n`channelHistory(channel)`: a scoped credential is denied the ad-hoc consumer the direct read\ncreates, by design. `cotal console` and the web console already do both.\n\nThe `<owner>.<actor>` pair is the session's principal, not its display name. On a user-auth mesh\nthe actor half **is** the agent's name, so the channel is `events.<your-owner>.<agent-name>`. On a\nstatic mesh the owner half is the literal `local` and the actor is a key the manager allocated, so\nthe channel is `events.local.<key>`; the spawn reply carries that key as `id`. Note\nthat `cotal console` and the web console keep event channels out of their channel lists on purpose,\nsince a plane is a machine feed rather than a conversation; they draw the frames when you open the\nchannel by name.\n\nThe rule governs the manager's doors, which are the ones a caller other than you can reach. A\nforeground `cotal spawn` on your own machine mints from your own signing material, so it can still\ngrant any channel you name: that is the out-of-band grant, not a way around the rule.\n\n**Failed turns publish run errors.** Claude Code decides for itself\nwhether a turn finished or died and fires one of two hooks accordingly, so the connector relays that\ndecision rather than making one of its own: a turn that ended on an API error ends its run with\n`RUN_ERROR` carrying the harness's own error kind (`rate_limit`, `billing_error`, `server_error`,\n`max_output_tokens` and the rest) as the code, and whatever detail it reported as the message. If that\ndetail cannot fit in the one closing frame, the shared close still publishes one `RUN_ERROR`\nthat does fit: it keeps the code and says the original detail was omitted or shortened because of the\nbound, so a reader is never shown a truncated message as complete. A turn that ended normally still\nends with a run-finished event carrying no outcome, which says the turn ended and does not claim it\nsucceeded.\n\nEvents are written to a per-session write-ahead log before they are published, so a hook that fires\nafter a restart resumes at the cursor it left rather than replaying or skipping, and a run that was\nopen when the session stopped is closed rather than left dangling.\n\nOne channel carries **every session of one agent**, because it is named after the principal and not\nafter the session. Alongside the per-session logs the connector keeps one small record per principal,\nholding the last sequence the broker assigned on that channel, so a new session continues the stream\nits predecessor left instead of starting again from nothing. Both live under the events state root\n(`COTAL_WORKSPACE_ROOT`), and neither is something you edit by hand.\n\nA **missing** record is not a fault: the connector rebuilds it from the session logs beside it,\nwhich is how an agent that was already running before this record existed keeps its stream. That\nrebuild stops if any one of those session logs is damaged. Unreadable, not valid JSON, and written\nfor a different principal all count, and so does a session directory or a log that is a link rather\nthan the real file the connector wrote, or a log that has more than one name. A tip taken from the\nrest would be too low, and it would stop publication later with nothing left to point at the cause.\nThe connector names the file instead, and the only way past it is the directory removal described\nbelow, under the same condition. A record that **disagrees with the broker** is a fault, and the\nconnector stops publishing and says why rather than guessing. A record that **moved while a session\nwas writing to it** is refused the same way: it means something else wrote the principal's record,\nand the connector reports which value it held and which the file holds rather than writing over the\nlater one. There is no command to clear it. The state is the principal's directory under the events\nroot, and clearing it by hand means removing that directory whole: the sequence, the cursor and the\nper-session logs only mean anything together, so removing part of it leaves a state the next start\nrefuses. Removing it is only half a remedy, and the half that comes first is the channel. The\ndirectory is where the agent's memory of the tip lives, not the tip itself, so on a channel that\nstill holds frames the next session opens expecting an empty one and stops on the same\ndisagreement, with the logs a tip could have been rebuilt from now gone. Purge the channel first,\nthen remove the directory.\n\nReading it: `cotal console` and the web console draw event frames directly. A frame carries no text\npart by design, so a surface that renders a message as flat text shows a marker instead of prose.\n\n**On a per-user-auth mesh, the default event plane needs the spawner's grant to cover the channel.** The event\nchannel is added to the child's publish set, and delegation only narrows: an agent may hand down\na subset of what it holds and no more. So a peer-initiated spawn is refused unless the\nspawning identity's own grant already covers the child's event channel. The refusal prints the\nexact `cotal actor grant` command that widens it. An operator launch, whose chain reaches an\nadmin-scoped or roster row, is unaffected. Passing `events: false` is the explicit opt-out.\n\n## Resume a session\n\n`--resume <session-id>` pulls an existing Claude session, its context and transcript,\ninto the mesh. It **forks**: Claude mints a *new* session id from that transcript\n(`--resume <id> --fork-session`), so the meshed agent gets its own session and the\noriginal is untouched.\n\n- `cotal spawn --resume <id>` (foreground) is the primary surface: the transcript is on\n *your* machine, and errors are Claude's own stderr, inline.\n- `--detach --resume <id>` works, with two differences: the id resolves against the\n **manager host's** `~/.claude` (you practically need `--cwd`), and the manager waits for\n a real outcome; `✓ started` means the agent *joined the mesh*, `✗ exited on launch`\n carries Claude's last output, and an uncertain launch (~30 s) is reported without\n tearing the agent down.\n- Resume is an **operator surface only**, deliberately not exposed on MCP `cotal_spawn`\n (a mesh peer naming host-local transcripts would widen `spawn` into transcript\n disclosure). Only the Claude connector supports it today; OpenCode and Hermes fail loud.\n- Needs a `claude` new enough for `--resume … --fork-session` (verified on 2.1.197).\n\n## Sharing your MCP servers\n\nIsolation is the default, but a meshed teammate sometimes genuinely needs one of your own\ntools (say, web search). The opt-in is the cotal config file\n(`~/.config/cotal/config.json`, or a space-local `.cotal/config.json` layered on top):\neach entry the familiar `.mcp.json` shape, secrets written as `${VAR}` references, never\nliterals ([full format](config.md)).\n\nAt launch the connector forwards *only* the named vars the chosen servers declare and\npasses the merged config as an owner-only temp file; `--strict-mcp-config` stays on, so\nonly cotal + the explicitly shared servers load. Scope per spawn with\n`--share-tools tavily,figma` (or `--share-tools none`).\n\nTwo caveats: sharing a server grants its credential to the agent (the var lives in the\nClaude process's environment, so share only when you're fine with that teammate holding\nthe key), and memory adds up, because a heavy server boots once per spawn, multiplied\nacross a team.\n\n## Feedback\n\n`cotal_feedback` works out of the box: without a key it posts to the public intake at\n`https://cotal.ai/v1/feedback` (needs a contact email: `COTAL_FEEDBACK_EMAIL`, then\n`git config user.email`, else the agent asks). Set `COTAL_FEEDBACK_KEY=fbk_<key>` in a\nbeta tester's environment to route to the keyed intake (`Authorization: Bearer`, identity\nderived from the key); `COTAL_FEEDBACK_URL` overrides either endpoint. The CLI can send\ntoo: `cotal feedback \"<summary>\" [--type bug]`. Each submission carries\n`origin: human | agent`, whether the tester asked, or the agent auto-reported a major\nissue.\n"
87
+ "body": "# Connect Claude\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nThe Claude Code connector turns a real `claude` session into a Cotal mesh peer. A bundled\nplugin inside the session joins NATS, maps lifecycle hooks to presence, and exposes the\nmesh tools. Nothing wraps Claude; it is an ordinary session that happens to be on the\nmesh.\n\nThe shared mesh runtime (agent, `cotal_*` tools, hook relay) lives in\n[`@cotal-ai/connector-core`](../extensions/connector-core); this connector is the thin\nClaude-specific adapter over it. Siblings: [OpenCode](connect-opencode.md) (beta),\n[Hermes](connect-hermes.md) (alpha), [pi](connect-pi.md) (alpha); the\n[Connectors](connectors.md) matrix compares them feature-by-feature.\n\n## Set up\n\n```bash\ncotal setup # one-time: installs the plugin, seeds one agent; launches nothing\ncotal up # brings up the mesh + delivery daemon + a detached manager\n```\n\n`cotal setup` installs the cotal plugin (so the repo's Claude sessions get the `cotal_*`\ntools) and seeds one `default` persona; `cotal up` brings up the local stack so\n`cotal spawn --detach` / `cotal_spawn` work right away. Re-running either is idempotent.\nThe install mechanics and the invariants behind them are in\n[setup internals](setup-internals.md).\n\n`cotal setup` also installs Cotal's authored Agent Skills (`SKILL.md`, the agentskills.io format) for\ncoordinating agent teams (today `team-topology`), from one canonical source, on two channels:\n\n- **Claude Code** gets a second, skills-only plugin, `cotal-skills`, from the same `cotal-mesh`\n marketplace, at **user scope** (machine-wide). The Claude connector declares and implements this\n setup provider, including the marketplace assets and native plugin commands; the base CLI only passes\n the vendor-neutral Agent Skills directory. The plugin carries no code and no core dependency,\n and uninstalls on its own with `claude plugin uninstall cotal-skills --scope user`. Its plugin version\n is stamped from the running CLI release, so an upgrade + `cotal setup --skills` runs `claude plugin update` and\n the deployed install actually gets the new skill. `cotal setup` installs it on first run and on repeat\n runs, so upgraders are not left behind. `cotal status` points a stale or missing skills plugin at\n `cotal setup --skills`.\n- **Every other harness** (Codex, Cursor, OpenCode, Gemini CLI, Windsurf/Devin) reads the cross-vendor\n `~/.agents/skills/` directory convention, which has no remote index, so `cotal setup` **reconciles** it\n (and `cotal setup --skills` does only that):\n it installs/updates each Cotal skill, backs up a copy you have edited to `SKILL.md.bak` before\n replacing it, and removes a Cotal skill that is no longer shipped. Only skills Cotal owns are touched;\n your own or third-party skills there are left alone. `cotal status` reports whether the drop is current,\n stale, missing, or has a retired skill to reconcile, and names `cotal setup --skills` as the remedy. This is the working cross-vendor path.\n\nCotal also generates an [Agent Skills discovery index](https://cotal.ai/.well-known/agent-skills/index.json)\non cotal.ai, but that RFC is still a draft with no harness consuming it yet, so it is a forward bet,\nnot a channel to rely on today.\n\n## Spawn a session\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn dave --detach # supervised: the manager runs it in a PTY\n```\n\nA spawn resolves a persona from `.cotal/agents/<name>.md` ([agent files](agent-files.md));\n`--model`, `--variant`, `--cwd`, `--prompt`, ACL overrides, and `--share-tools` apply to\nboth forms ([run a mesh](run-a-mesh.md) has the full resolution rules). The session joins\nwith identity from its environment and auto-registers presence by the time it is\ninteractive.\n\nInside the session, the agent orients with one read-only tool, `cotal_orientation`: its\nidentity, the channels it reads and may post to, its capabilities, the tools available,\nwho's present, and unread counts. The full tool surface is the\n[MCP tool catalog](mcp-tools.md). In auth mode the team-supervision tools\n(`cotal_spawn` / `cotal_persona` / `cotal_personas`) are injected **only** for personas declaring\n`capabilities: [spawn]` (the same grant that opens the privileged control subject), so an\nagent's toolset matches its declared capabilities. `cotal_run` is gated separately by\n`run`; use `capabilities: [spawn, run]` for both. Fresh setup defaults include both.\nSee [workflow tool setup](workflows.md#from-an-agent-session) for a first run and missing-tool checks.\nClearing retained history is\noperator-only ([run a mesh](run-a-mesh.md)), never an agent tool.\n\n## How it binds\n\nClaude Code exposes four integration surfaces, and three of them collapse into a single\ndual-purpose MCP server:\n\n| Surface | Mechanism |\n|---|---|\n| Outbound, ambient | `http` lifecycle hooks → POST to the connector (presence, activity) |\n| Outbound, deliberate | MCP tools `cotal_send` / `cotal_dm` / `cotal_anycast` (+ `cotal_feedback`) |\n| Inbound, pull | MCP tool `cotal_inbox` (same server) |\n| Inbound, push | Channel nudge + hook drain (below) |\n\nThe manager launches the *real* `claude` (no wrapper):\n\n```\nclaude --strict-mcp-config --mcp-config '{\"mcpServers\":{\"cotal\":{…}}}' \\\n --dangerously-load-development-channels server:cotal\n# env: COTAL_SPACE, COTAL_NAME, COTAL_ROLE, COTAL_CHANNEL=1, plus claude's documented auth vars\n```\n\n- **Model auth.** Locally, `claude` still reads macOS Keychain / `~/.claude`. In a container or\n CI there is no Keychain, so the connector forwards the documented credential set:\n `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`), `ANTHROPIC_API_KEY` /\n `ANTHROPIC_AUTH_TOKEN`, and the cloud-provider flags plus their credential vars. Host-session\n markers (`CLAUDE_CODE_CHILD_SESSION`, `CLAUDECODE`) stay out so a nested seat still saves a\n transcript. See [Deploy](deploy.md).\n- **Persona privacy.** The persona body is written to a private file and Claude receives only\n `--append-system-prompt-file <path>`. The body never appears in the spawned process argv. The\n carrier is a 0600 file inside a 0700 directory on POSIX, with equivalent owner-only ACL hardening\n on Windows.\n- **MCP isolation.** A spawned agent runs with **only** the cotal MCP server:\n `--strict-mcp-config` ignores every other MCP source, crucially the operator's personal\n `~/.claude.json` servers (several spawns each booting a heavy helper would starve\n memory). Share your own servers deliberately (see below).\n- **Installed plugin.** The plugin is installed once (`claude plugin install\n cotal@cotal-mesh --scope local`) because its hooks bind only to an *installed* plugin.\n In a clone the marketplace is the repo's `.claude-plugin/marketplace.json`; `cotal setup`\n (npx, no clone) materializes the same marketplace under `~/.cotal/claude-plugin/` (each plugin dir is\n rebuilt from scratch and atomically replaced, never merged, so no stale file rides in). The\n `cotal-skills` plugin installs from that same marketplace at user scope (`claude plugin install\n cotal-skills@cotal-mesh --scope user`); its manifest and install behavior ship inside the Claude connector, and\n its version tracks the CLI release so updates land.\n- **Identity-gated.** Connector code requires `COTAL_NAME` *or* `COTAL_LINK`. A plain\n `claude` with no `COTAL_*` env stays inert and never joins, so your own sessions in a\n repo do not appear as stray peers.\n- **Hands-free.** The dev-channels flag prints a one-time confirm prompt. The PTY runtime waits for\n the dialog title in normalized terminal output and presses Enter once when it appears, so startup\n speed does not affect a supervised launch. If the declared prompt never appears, the seat exits\n with a bounded error naming the unmatched prompt instead of hanging silently.\n\nInbound mesh messages arrive in context as\n`<channel source=\"cotal\" from=\"bob\" kind=\"dm\" …>…</channel>`: each meta key a tag\nattribute the agent can read for routing.\n\n## How messages reach the session\n\nDurable deliveries land in the connector's inbox from JetStream consumers\n([SPEC §8](../SPEC.md#8-nats--jetstream-binding)); live channel traffic can instead arrive\nthrough an at-most-once core subscription. A durable message sent while the agent is busy\nor offline waits on the stream. Two things move a message from inbox to model; one\ndelivers, the other only wakes:\n\n- **Hook drain (delivery).** `SessionStart` / `UserPromptSubmit` hooks read automatic inbox items and\n inject them as `additionalContext`. This is the single authoritative path: deterministic and works\n on any Claude Code build. Quiet ambient is excluded and stays buffered for `cotal_inbox`.\n A message is **acked only once the hook reply carrying it has cleared both legs of its journey**:\n the connector's control socket to the hook process (which gives up after 2s), and the hook\n process's own stdout to Claude Code (which it force-exits 1s after starting to write). The relay\n sends a receipt back down the control socket from that stdout write's callback, and only on a\n clean write (a runtime whose pipe has gone away fails it), and the connector treats that receipt,\n not its own socket write, as delivery. So a large injection killed mid-flush, or one written to a\n broken pipe, leaves the message un-acked and JetStream redelivers it. What this does *not* prove is\n that Claude Code read or applied the reply: a payload small enough to fit the pipe buffer is\n reported written the moment the kernel takes it. That residual is why the path errs toward\n at-least-once rather than treating a confirmed write as a confirmed read. Acking when\n the reply was merely *formatted* meant a lost reply was a lost message: it was already marked\n handled, so its own redelivery was silently acked on arrival.\n This errs toward **at-least-once**: if a reply lands but its confirmation does not, the batch is\n surfaced again and flagged as a possible repeat. A duplicate injection is noise; a buried DM stops\n the peer answering at all.\n- **Channel nudge (wake).** An arriving message fires a `notifications/claude/channel`\n event that wakes an *idle* session into a turn, so the drain runs *now* instead of at\n the next prompt. The nudge never acks anything. A nudge that the host rejects is retried with a\n bounded backoff while anything is still pending. For an idle session it is the only wake source,\n so dropping it means silence until someone types. When the channel becomes active, the connector\n first re-fires a focus mention remembered during startup, otherwise one buffered wake. A rejected\n push keeps its bounded retry, and JetStream redelivery remains the durable backstop for unacked\n inbox items. If the channel cannot run at all, delivery still waits for the next hook. Live-only\n traffic has no durable retry.\n\n**Two priority tiers.** A *directed* message (DM, anycast, or a channel message that\n`@mentions` us) always nudges. *Ambient* channel chatter does not nudge mid-turn; it\naccumulates, and the `Stop` → idle transition fires one batch nudge so the backlog drains\ntogether.\n\n**Constraints (accepted).** Channels are a Claude Code research preview (≥ v2.1.80;\npermission relay ≥ v2.1.81): Anthropic auth only, admin-enabled on Team/Enterprise, and a\ncustom channel needs the `--dangerously-load-development-channels` launch flag. The hook\ndrain does not depend on any of that; the channel only adds \"wake me when idle.\"\n\nThe same channel also relays **tool-permission requests** onto the mesh, so a peer (a\nhuman at the CLI, a policy node) can approve or deny an agent's pending tool call through\nCotal rather than a per-terminal prompt.\n\n### Attention\n\nAn agent picks how aggressively peer traffic reaches it with\n`cotal_status({ attention })` (three modes, orthogonal to presence):\n\n| arrival | open (default) | dnd | focus |\n|---|---|---|---|\n| directed (dm / anycast) | wake + inject | wake + inject | wake + inject |\n| channel `@mention` | wake + inject | wake + inject | ack-drop; wake to *pull*; not injected |\n| ambient channel chatter | wake when idle; hold while working | never wakes; injects next turn | ack-drop; recall via `cotal_inbox` |\n\nPer-channel overrides refine this: **quiet** (delivered, never wakes; `@mention` still\nwakes) and **muted** (dropped on receive, mentions included; DMs/anycast unaffected), set\nwith `cotal_channel_mode` or as agent-file defaults (`quiet:` / `muted:`,\n[agent files](agent-files.md)). A per-channel override is the final word for that channel.\nQuiet ambient is pull-only: it never hitchhikes on a human prompt, DM, mention, or other\nconnector-driven turn. `cotal_inbox` explicitly surfaces and clears it. A quiet-channel\n`@mention` remains automatic and injects normally.\n\nA pull is bounded too, and clears only what it hands over. One `cotal_inbox` call carries at most a\nreceivable window (direct messages and role requests first, then channel traffic, replayed history\nlast); whatever does not fit stays buffered, is named in the reply, and comes back on the next call.\nA message too large for one whole response is never consumed at all: it is named with its sender and\nsize and left buffered, because clearing what cannot be delivered is the loss this bound exists to stop.\nThat matters most on the path where it is easiest to lose mail: reconnecting brings a channel-history\nreplay with it, so the largest payload and the least expendable message arrive in the same read.\n\nThe local inbox is bounded. On pathological overflow it evicts pull-only items before automatic\ntraffic. If the bounded live/durable classification guard also fills, the connector fails closed:\notherwise-normal ambient becomes pull-only until restart. Muted hard-drop and normal focus recall\nstill take precedence. Focus also keeps a bounded exclusion list so mode toggles cannot recall\nquiet/muted traffic; if that safety bound fills, recall skips the affected channel and reports it\nas incomplete rather than risk resurfacing excluded content.\nIf the separate hard-drop disposition guard fills, channel traffic is dropped for the rest of the\nsession rather than risk a late copy bypassing an earlier muted/focus decision; DMs and anycast are\nunaffected.\n\nAttention is **advisory UX, not a boundary**: any peer can wake a dnd/focus agent by\nnaming it, and `muted` means \"I opted out of receiving\", not \"the channel is blocked\";\nthe broker still authorizes and delivers. Focus's real effect is shrinking the\nuntrusted-ambient injection surface (only subject-authenticated dm/anycast auto-inject).\nIt resets to **open** on `SessionStart`, so a restarted agent never stays silently deaf.\nYour attention is mirrored into presence so peers can see it.\n\nWhatever does reach a turn is framed so a peer cannot write the frame. A line that begins at column\nzero is written by the connector; one message is one line plus indented continuations, with the\nsender inside a single bracket pair. A message body, a sender name and role, and a service or\nchannel label are all peer-controlled, so each passes through the same neutralization the\n`cotal_inbox` reply uses: no line break a splitter may honour and no bracket survives into a\nrendered attribution. This matters more for an injected block than for a reply, because the agent\ndid not ask for it and so never had the chance to distrust it.\n\n## Presence mapping\n\nThe connector wires a small subset of Claude Code hooks to presence states; presence is\ncoarse, and \"what it is doing\" rides on activity updates. Presence is **advisory**: a presence\npublish that fails (the endpoint mid-reconnect, say) is swallowed and never prevents the same hook\nfrom delivering messages or flushing held ones.\nA `SessionStart` during an open turn, including compaction, preserves the current `working` or\n`waiting` status until `Stop`, `StopFailure`, or `SessionEnd` closes the turn.\n\n| Hook | → state |\n|---|---|\n| `SessionStart` | `idle` only when no turn is open (join; surfaces the inbox; captures the live model into `meta.model` when no pin) |\n| `UserPromptSubmit` | `working` (turn starts; surfaces the inbox) |\n| `PreToolUse` | no change; records *what* is about to run, so a permission wait can name it |\n| `Notification` (`permission_prompt` / `agent_needs_input`) | `waiting` with condition `approval` / `input` (activity leads with the pending tool, e.g. `Bash: git push …`) |\n| `Stop` / `StopFailure` | `idle` (turn done / died on an API error; flushes anything held while busy). `StopFailure` also relays Claude Code's native error value as `condition.source` and maps it to the closed condition vocabulary. On the [event plane](#event-plane) it closes the run with `RUN_ERROR`. |\n| `SessionEnd` | `offline` (graceful leave) |\n\n`StopFailure` maps `rate_limit` and `overloaded` directly; auth and credential failures to\n`auth`; account and billing failures to `billing`; `invalid_request` to `request`;\n`model_not_found` to `model`; `server_error` to `server`; `max_output_tokens` to `context`; and\n`unknown` to `failed`. The native value remains in `condition.source`.\n\nHooks are relayed over the connector's **authenticated** local control endpoint (per-user\nsocket + per-launch token, constant-time checked), so a local process that finds the path\nstill can't drive presence or stop the agent. The full Claude Code hook-event list lives\nwith the adapter:\n[`extensions/connector-claude-code`](../extensions/connector-claude-code/README.md).\n\n## Event plane\n\nA spawned session publishes a **structured** account of what it\ndid: run boundaries per turn, assistant text, reasoning, and each tool call with its start\nand its end. Not prose about the work, the work itself, in a vocabulary a program can\nread. The launcher sets `COTAL_EVENTS` by default; pass `--no-events` to opt out on an unrestricted\nspace. A user-auth registration with `policy: { events: \"required\" }` carries `eventsRequired` in the\nprivate launch material, so the connector arms even without `COTAL_EVENTS`; `--no-events` is refused.\nA hand-driven user-mode session may carry the same decision as `COTAL_EVENTS_REQUIRED=1`. Its own\npublish grant must cover `events.<owner>.<actor>` or the connector refuses before joining. An unmanaged\nsession with no launch material and no required-policy fallback keeps the generic default behavior.\n\nA new session includes its first run even when Claude writes a positional startup prompt before the\nconnector receives `SessionStart`. That from-zero read is keyed only to Claude's explicit\n`source: \"startup\"`; resumed, forked, cleared, and compacted sessions adopt at the current transcript\nboundary and do not republish retained history. Crash recovery follows the cursor already stored in\nthe event write-ahead log, regardless of the new process's startup label.\n\nClaude starts each hook in its own process, so a prompt or stop relay can reach Cotal before the\n`SessionStart` relay. The connector holds those event flushes and the terminal until `SessionStart`\nsupplies the source, then enqueues adopt, flush, and close in that order.\n\n`SessionStart` can also run before the connector process has bound its local control socket. The hook\nthe `SessionStart` relay retries only transient pre-connect listener errors, with capped backoff\ninside its existing two-second budget. Later hooks and permanent local faults still fail open\nimmediately. Once a socket has connected, a broken exchange is not retried: the connector may\nalready have handled the frame, so replaying it could apply one lifecycle event twice.\nThat retained `SessionStart` can itself arrive before Claude creates the transcript path. A genuinely\nnew startup waits up to five seconds for that file with capped backoff, and the same deadline bounds\none stalled file read; expiry fails loud instead of silently losing the first run. A forked session\ngets the same wait, because Claude copies the parent transcript into the fork's own file after the\nhook, and then adopts at the end of that copy. Resumed, cleared and compacted starts and recovered\ncursors still require their existing source at once.\n\nTool arguments (`TOOL_CALL_ARGS`) and tool results (`TOOL_CALL_RESULT`) are not republished\nonto this channel. The durable emitter drops those events before they are written to the\nwrite-ahead log, because this channel's read ACL is not the ACL the tool ran under. Content is\nmandatory on both kinds, so the event is suppressed rather than emptied or replaced with a\nplaceholder. Tool start and end still go out. A restart that finds a pending pre-fix frame\nstill carrying those kinds HALTS rather than republishing it.\n\nThe channel is **`events.<owner>.<actor>`**, named after the session's principal. What the actor\nhalf is depends on the mesh, and the difference matters when you go looking for it: on a static mesh\nit is a key the manager allocated, never the display name, so two live agents sharing a display name\ndo not share a stream; on a user-auth mesh it is the agent's own name, because that is what the\nledger row is keyed on. Spelled out again with both halves below. The launch grants publish rights\non that channel alone. A spawn\nthat asks for a *different* agent's event channel is refused at the door rather than granted, since\nthat channel is that session's event stream. The same rule runs on restart: a manager\nresume document that names another agent's event channel is refused rather than adopted, because the\nmanaged row is re-armed from that document and the credential is re-minted from the row.\n\nThe rule reads a **concrete** channel, two principal tokens and nothing else. A pattern such as\n`events.<owner>.>` is not an event channel to it and passes untouched, governed by ordinary ACL\nauthority: on a user mesh the delegation envelope, on a static mesh the spawning credential itself.\nThat is deliberate, because the pattern is the form an operator writes on purpose for an observer,\nand it is worth knowing rather than assuming the fence is total.\n\nTo let something else read a plane, grant it out of band. The refusal prints the command for the\nmesh it is running on, spelled out in full, and only that one.\n\nOn a **user-auth** mesh:\n\n```bash\ncotal actor grant <reader> --owner <owner> --scope '' --allow-subscribe 'events.<owner>.<actor>' --allow-publish ''\n```\n\nEvery field, deliberately. `actor grant` is an upsert of the whole row, and an omitted flag is not\n\"leave it alone\": it is the wide default, `>` read, `>` post, and `spawn,role:default` scope. A bare\n`cotal actor grant <reader>` therefore grants a reader of every channel in the space, which is the\nopposite of what a scoped watcher is for.\n\nOn a **static** mesh there is no actor ledger for `actor grant` to write to, and the refusal says\nso; mint the reader instead:\n\n```bash\ncotal mint watcher --profile agent --allow-subscribe 'events.<owner>.<actor>' --provision\n```\n\nThe **agent** profile, not the observer one. `mint` reads `--allow-subscribe` only for that\nprofile, and refuses it anywhere else: `--profile observer --allow-subscribe <channel>` exits\nnon-zero and writes no creds file, because the observer profile carries a fixed read set over the\nwhole chat plane, which is the opposite of what a scoped watcher is for. The agent profile also prints the lifecycle uid the\nreader needs, since an authed consuming endpoint refuses to start without one.\n\nTwo things a reader has to do that are not obvious, both on `CotalEndpoint`. It must pass the event\nchannel in `channels`: an endpoint reads the channels it lists, so one constructed without\nthe event channel joins nothing and the frames never arrive. And it reads history with `readHistory(channel)`, the delivery daemon's mediated read, not\n`channelHistory(channel)`: a scoped credential is denied the ad-hoc consumer the direct read\ncreates, by design. `cotal console` and the web console already do both.\n\nThe `<owner>.<actor>` pair is the session's principal, not its display name. On a user-auth mesh\nthe actor half **is** the agent's name, so the channel is `events.<your-owner>.<agent-name>`. On a\nstatic mesh the owner half is the literal `local` and the actor is a key the manager allocated, so\nthe channel is `events.local.<key>`; the spawn reply carries that key as `id`. Note\nthat `cotal console` and the web console keep event channels out of their channel lists on purpose,\nsince a plane is a machine feed rather than a conversation; they draw the frames when you open the\nchannel by name.\n\nThe rule governs the manager's doors, which are the ones a caller other than you can reach. A\nforeground `cotal spawn` on your own machine mints from your own signing material, so it can still\ngrant any channel you name: that is the out-of-band grant, not a way around the rule.\n\n**Failed turns publish run errors.** Claude Code decides for itself\nwhether a turn finished or died and fires one of two hooks accordingly, so the connector relays that\ndecision rather than making one of its own: a turn that ended on an API error ends its run with\n`RUN_ERROR` carrying the harness's own error kind (`rate_limit`, `billing_error`, `server_error`,\n`max_output_tokens` and the rest) as the code, and whatever detail it reported as the message. If that\ndetail cannot fit in the one closing frame, the shared close still publishes one `RUN_ERROR`\nthat does fit: it keeps the code and says the original detail was omitted or shortened because of the\nbound, so a reader is never shown a truncated message as complete. A turn that ended normally still\nends with a run-finished event carrying no outcome, which says the turn ended and does not claim it\nsucceeded.\n\nEvents are written to a per-session write-ahead log before they are published, so a hook that fires\nafter a restart resumes at the cursor it left rather than replaying or skipping, and a run that was\nopen when the session stopped is closed rather than left dangling.\n\nOne channel carries **every session of one agent**, because it is named after the principal and not\nafter the session. Alongside the per-session logs the connector keeps one small record per principal,\nholding the last sequence the broker assigned on that channel, so a new session continues the stream\nits predecessor left instead of starting again from nothing. Both live under the events state root\n(`COTAL_WORKSPACE_ROOT`), and neither is something you edit by hand.\n\nA **missing** record is not a fault: the connector rebuilds it from the session logs beside it,\nwhich is how an agent that was already running before this record existed keeps its stream. That\nrebuild stops if any one of those session logs is damaged. Unreadable, not valid JSON, and written\nfor a different principal all count, and so does a session directory or a log that is a link rather\nthan the real file the connector wrote, or a log that has more than one name. A tip taken from the\nrest would be too low, and it would stop publication later with nothing left to point at the cause.\nThe connector names the file instead, and the only way past it is the directory removal described\nbelow, under the same condition. A record that **disagrees with the broker** is a fault, and the\nconnector stops publishing and says why rather than guessing. A record that **moved while a session\nwas writing to it** is refused the same way: it means something else wrote the principal's record,\nand the connector reports which value it held and which the file holds rather than writing over the\nlater one. There is no command to clear it. The state is the principal's directory under the events\nroot, and clearing it by hand means removing that directory whole: the sequence, the cursor and the\nper-session logs only mean anything together, so removing part of it leaves a state the next start\nrefuses. Removing it is only half a remedy, and the half that comes first is the channel. The\ndirectory is where the agent's memory of the tip lives, not the tip itself, so on a channel that\nstill holds frames the next session opens expecting an empty one and stops on the same\ndisagreement, with the logs a tip could have been rebuilt from now gone. Purge the channel first,\nthen remove the directory.\n\nReading it: `cotal console` and the web console draw event frames directly. A frame carries no text\npart by design, so a surface that renders a message as flat text shows a marker instead of prose.\n\n**On a per-user-auth mesh, the default event plane needs the spawner's grant to cover the channel.** The event\nchannel is added to the child's publish set, and delegation only narrows: an agent may hand down\na subset of what it holds and no more. So a peer-initiated spawn is refused unless the\nspawning identity's own grant already covers the child's event channel. The refusal prints the\nexact `cotal actor grant` command that widens it. An operator launch, whose chain reaches an\nadmin-scoped or roster row, is unaffected. Passing `events: false` is the explicit opt-out.\n\n## Resume a session\n\n`--resume <session-id>` pulls an existing Claude session, its context and transcript,\ninto the mesh. It **forks**: Claude mints a *new* session id from that transcript\n(`--resume <id> --fork-session`), so the meshed agent gets its own session and the\noriginal is untouched.\n\n- `cotal spawn --resume <id>` (foreground) is the primary surface: the transcript is on\n *your* machine, and errors are Claude's own stderr, inline.\n- `--detach --resume <id>` works, with two differences: the id resolves against the\n **manager host's** `~/.claude` (you practically need `--cwd`), and the manager waits for\n a real outcome; `✓ started` means the agent *joined the mesh*, `✗ exited on launch`\n carries Claude's last output, and an uncertain launch (~30 s) is reported without\n tearing the agent down.\n- Resume is an **operator surface only**, deliberately not exposed on MCP `cotal_spawn`\n (a mesh peer naming host-local transcripts would widen `spawn` into transcript\n disclosure). Only the Claude connector supports it today; OpenCode and Hermes fail loud.\n- Needs a `claude` new enough for `--resume … --fork-session` (verified on 2.1.197).\n\n## Sharing your MCP servers\n\nIsolation is the default, but a meshed teammate sometimes genuinely needs one of your own\ntools (say, web search). The opt-in is the cotal config file\n(`~/.config/cotal/config.json`, or a space-local `.cotal/config.json` layered on top):\neach entry the familiar `.mcp.json` shape, secrets written as `${VAR}` references, never\nliterals ([full format](config.md)).\n\nAt launch the connector forwards *only* the named vars the chosen servers declare and\npasses the merged config as an owner-only temp file; `--strict-mcp-config` stays on, so\nonly cotal + the explicitly shared servers load. Scope per spawn with\n`--share-tools tavily,figma` (or `--share-tools none`).\n\nTwo caveats: sharing a server grants its credential to the agent (the var lives in the\nClaude process's environment, so share only when you're fine with that teammate holding\nthe key), and memory adds up, because a heavy server boots once per spawn, multiplied\nacross a team.\n\n## Feedback\n\n`cotal_feedback` works out of the box: without a key it posts to the public intake at\n`https://cotal.ai/v1/feedback` (needs a contact email: `COTAL_FEEDBACK_EMAIL`, then\n`git config user.email`, else the agent asks). Set `COTAL_FEEDBACK_KEY=fbk_<key>` in a\nbeta tester's environment to route to the keyed intake (`Authorization: Bearer`, identity\nderived from the key); `COTAL_FEEDBACK_URL` overrides either endpoint. The CLI can send\ntoo: `cotal feedback \"<summary>\" [--type bug]`. Each submission carries\n`origin: human | agent`, whether the tester asked, or the agent auto-reported a major\nissue.\n"
88
88
  },
89
89
  {
90
90
  "slug": "connect-codex",
@@ -161,7 +161,7 @@ export const DOCS_BUNDLE = {
161
161
  "title": "Embedding Cotal",
162
162
  "kind": "Guide (informative)",
163
163
  "summary": "The cotal binary in this repo is one composition root: an operator CLI.",
164
- "body": "# Embedding Cotal\n\n> **Guide** (informative) · **For:** implementers building a service on top of Cotal · **Prereqs:** [Architecture](architecture.md), [Identity and auth](identity-and-auth.md), [Delivery daemon](delivery-daemon.md)\n\nThe `cotal` binary in this repo is one composition root: an operator CLI. A separate service\n(for example a hosted, multi-tenant Cotal) does not fork this repo. It writes its **own**\ncomposition root that depends on the published `@cotal-ai/*` packages and imports the surfaces it\nwants. `bin/cotal.ts` uses the same composition pattern. This page is the contract for that: what is a real library\nexport you can build against, how to boot the server-side daemons from those exports, and where the\ncurrent export surface stops short of a fully hosted composition.\n\nThis is the \"guarded substrate\" boundary in practice. Nothing here reveals or assumes a specific\nhost; it documents the public seams any embedder composes.\n\n## What you embed\n\nThe supported reference shape here is **one broker operator serving one space** (one tenant: a\ndedicated data account, under an operator that also holds the system account and a quarantined\nauth-callout account) plus three standalone processes. The trust layer itself composes many spaces\nunder one broker operator today (`createBrokerAuth` + `createSpaceAccountAuth` + N-space\n`serverConfig`); what does not exist yet is the per-space **lifecycle** on a shared broker (see\n[Known gaps](#hosted-composition-gaps)). The three processes:\n\n| daemon | package | what it is |\n|---|---|---|\n| auth-service | `@cotal-ai/auth` | the NATS auth callout, the IdP token exchange, and JWKS. Plane 1 to Plane 2. |\n| delivery | `@cotal-ai/delivery` | the Plane-3 durable backstop: fan-out writer plus trusted reader, per space. |\n| supervise | `@cotal-ai/manager` | the per-machine agent lifecycle (spawn/despawn/attach), per space. |\n\n`mint`, `deliver`, and `auth-service` expose their behavior as direct library primitives, and the\nsupported one-space bootstrap below re-composes from exported low-level primitives. `supervise` and\nthe full `up` orchestration are **not** public runners: `up` also does broker bring-up, restore,\nprocess and registry management, and lifecycle work, and `supervise`'s orchestration is private (see\n[Supervisor signing authority](#supervisor-signing-authority)).\n\n## The export surface\n\nEverything below is a real export of a published package, reachable from the package root (each\npackage publishes only `.` via `dist/index.{js,d.ts}` and ships `files: [\"dist\"]`). Type-only names\nare marked; import them with `import type`.\n\n**Daemon runners and lifecycle**\n\n| symbol | package | purpose |\n|---|---|---|\n| `runAuthService(args, store?)` | `@cotal-ai/auth` | boot the auth-service daemon; `store` injects the secret material. |\n| `runDelivery(args, store?)` | `@cotal-ai/delivery` | boot the delivery daemon; `store` injects the scoped `delivery` cred. |\n| `deliveryCredsKey(space, composition)`, `membershipRwCredsKey(space, composition)` | `@cotal-ai/workspace` | build the secret-store keys the delivery cred and the membership feed's rw cred are read/re-signed under. Keys are **per-space**: `space.<hex>/<kind>`. A hosted composition passes `{ injected: true }`. |\n| `DELIVERY_CREDS_KIND`, `MEMBERSHIP_RW_CREDS_KIND` | `@cotal-ai/workspace` | the operator-facing KIND names (`delivery.creds`, `membership-rw.creds`) those keys are built from, and what renewal results report. A kind is **not** a key: putting a cred under the bare kind writes the pre-0.4 flat location, which nothing reads. |\n| `Manager`, `ManagerOptions` *(type)* | `@cotal-ai/manager` | construct and run a supervisor in-process; `ManagerOptions.secretStore` injects the one store it reads/writes every secret through. `ManagerOptions.remoteAuthority` is the hosted manager-service authority bundle, including host-owned release, retained-validation, goal-index, and serve-time admin-authorization callbacks. |\n| `createRuntime`, `Runtime` *(type)* | `@cotal-ai/manager` | resolve the spawn backend (pty built in). |\n\n**Provisioning and minting** (all `@cotal-ai/core`)\n\n| symbol | purpose |\n|---|---|\n| `createBrokerAuth(label)` | mint BROKER trust: the operator and system account one nats-server trusts. One per broker, shared by every space on it. |\n| `createSpaceAccountAuth(broker, space)` | mint one space's own data account, signed by that broker's operator: the add-a-tenant primitive. |\n| `createSpaceAuth(space)` | the one-space convenience: broker trust + one account in a single composed bundle. |\n| `setupSpaceStreams({ servers, space, creds })` | create the space's JetStream streams. |\n| `ensureDefaultDeliveryClass({ servers, space, creds?, deliveryClass })` | write the space's default delivery class at creation so it is wire-discoverable (SPEC section 4). |\n| `serverConfig(broker, spaces, { storeDir, maxFileStore?, extraAccounts?, port?, host? })` | render the broker config: one operator, N space accounts. `storeDir` is required, `maxFileStore` caps JetStream file storage in bytes (omitted, nats-server's dynamic default applies), and `extraAccounts` preloads the auth-callout account. |\n| `mintCreds(auth, identity, profile, opts?)` | mint a scoped cred for any `Profile`. |\n| `mintMembershipObserverCreds`, `mintConnectionEvictorCreds` | mint the membership/eviction scoped creds. |\n| `provisionAgent`, `provisionAgentDurables` | create a principal's bind-only durables. |\n| `newIdentity`, `stripSpaceAuth` | a fresh nkey identity; a stripped signer bundle (data signing seed only). |\n| `Profile`, `CredentialKind`, `MintOpts`, `SpaceAuth` *(types)*, `CREDENTIAL_LIFETIMES` | the profile matrix and cred lifetime policy. |\n\n**Auth building blocks** (all `@cotal-ai/auth`)\n\n| symbol | purpose |\n|---|---|\n| `createCalloutAuth`, `startAuthCallout` | the NATS auth-callout responder. |\n| `createUserTokenIssuer`, `pinnedJwksResolver` | mint and verify the Cotal user bearer. |\n| `createIdpBridge` | exchange a verified IdP JWT for a Cotal bearer (see [the callout contract](identity-and-auth.md#the-idp-callout-contract)). |\n| `deriveOwnerToken`, `validateUserToken` | owner derivation; strict bearer validation. |\n| `cotalAuthProvider` | the self-registering `auth-provider` extension. |\n| `ensureCalloutAuth`/`loadCalloutAuth`, `ensureIssuer`/`loadIssuer`, `ensureOwnerSecret`/`loadOwnerSecret` | read/write the auth secret kinds through a `SecretStore`. |\n\n**Seams and the wire** (all `@cotal-ai/core` unless noted)\n\n| symbol | purpose |\n|---|---|\n| `SecretStore` *(type)* | the durable hosted-secret seam (get/put/delete); `get()` returns raw seeds/keys into process memory, so it is a blob seam, not HSM/KMS signing. |\n| `FsSecretStore`, `workspaceSecretStore(root)` | the filesystem default. **These live in `@cotal-ai/workspace`, not core.** |\n| `AuthProvider` *(type)*, `Connector` *(type)*, `Runtime` *(type)*, `Command` *(type)* | the extension contracts; implementations self-register on import. |\n| `registry` | the shared registry a composition root pulls surfaces into. |\n| `CotalEndpoint`, subjects, message types | the wire client and shapes. |\n| `ParsedArgs` *(type)* | the shape the daemon runners take (see below). |\n\nThe runners take a CLI-shaped `ParsedArgs`, not a typed options object, so a host fabricates one:\n\n```ts\nconst args: ParsedArgs = { values: { space, server, port: \"0\" }, positionals: [], raw: [] };\n```\n\n### Long-lived endpoints take a bearer function\n\n`EndpointOptions.bearer` accepts either a string or a function, and the difference is not stylistic.\nA string is minted once, so when it expires (which it will: callout bearers live minutes) the\nendpoint has nothing to renew with. It will not present the dead token to the broker, since that is\na guaranteed denial that still costs a full auth-callout round trip. It refuses to reconnect, emits\n`warning` saying which case it is in, and retries on a widening backoff until the process\nre-authenticates and rebuilds it. Retry notices use `warning` rather than `error` because Node\nrethrows an unhandled `error` event and would kill a host the endpoint is still trying to recover.\n\nPass a **function** for anything that outlives one bearer. That is a renewal source: it is called\nahead of each expiry and again whenever a reconnect finds the cached bearer dead, and it requires\nexplicit `card.owner` and `card.actor`. The first-party surfaces already do this\n(`UserViewAuth.source`, the connector's `agentBearerCommand`). A string bearer is for a short\none-shot connection.\n\nLong-lived hosts must also subscribe to the endpoint's `warning` event. It carries conditions the\nendpoint is surviving, including failed credential renewal and reconnect retries. A host may choose\nto ignore warnings for a one-shot endpoint whose awaited operation owns the verdict, but that choice\nshould be explicit. An unhandled warning is nonfatal and silent.\n\n## Booting the daemons\n\n### auth-service\n\n`runAuthService(args, store?)` reads its provisioned long-lived secret kinds (service keys, callout\naccount, issuer keys, owner secret) through the injected `SecretStore`; a host provisions those into\nthe store first. It is a **signer and identity authority**, not a scoped daemon: at runtime it holds\nthe data-account and callout-account signing seeds, the issuer's private JWKs, and the\nowner-derivation secret in process memory (`SecretStore.get` exports raw values). The IdP pin and the\nactor ledger are **not** store-injected: `runAuthService` resolves them under\n`userAuthStateDir(findCotalRoot(), space)`, a path relative to the process working directory, so a\nhost provisions those into that exact directory (neither `store` nor `COTAL_HOME` selects it). It\nalso writes an ephemeral `auth-service.json` discovery file there that carries the live exchange\ncapability.\n\n```ts\nimport { runAuthService } from \"@cotal-ai/auth\";\n// store implements SecretStore over your secret backend; get() returns raw seeds into memory.\n// Provision the auth secret kinds into the store, AND the IdP pin + actor ledger under\n// userAuthStateDir(findCotalRoot(), space), before this call.\nawait runAuthService(\n { values: { space, server: brokerUrl, port: \"8081\" }, positionals: [], raw: [] },\n store,\n);\n```\n\n### delivery\n\n`runDelivery(args, store?)` runs from a **pre-minted scoped `delivery` cred** and never loads the\nsigner. Provide the cred either through the injected store (under\n`deliveryCredsKey(space, { injected: true })`) or with a\n`--creds` file; the two are mutually exclusive. The daemon re-fetches the cred from the store at 75%\nof its JWT lifetime and fails loud rather than riding to expiry, so **something must re-sign a fresh\ncred into that same store**. When that read finds the previous generation still there, the daemon\nreports the missed remint and retries in 60 seconds. The current cred stays live until its expiry,\nand the store is read once per retry rather than once per second.\n\n```ts\nimport { runDelivery } from \"@cotal-ai/delivery\";\nawait runDelivery({ values: { space, server: brokerUrl }, positionals: [], raw: [] }, store);\n```\n\nThat renewal is a **signer** operation, not the delivery daemon's:\n`remintDaemonCreds(root, space, store?, { preflight? })` (`@cotal-ai/workspace`) reads the `SpaceAuth`\nsigner **through the same resolved `store`** (`getSpaceAuth(store ?? workspaceSecretStore(root), space)`,\nkeys `auth/broker.json` + `auth/account.<key>.json`; the pre-split `auth/auth.json` monolith is\nmigration input and the container signer mount only) and re-signs the daemon creds (`delivery.creds` and the membership feed's\n`membership-rw.creds`) back into that store. The injected `store` is both the signer source and the\ncredential destination, never a split. `space` is **required** and validated against the store's signer, so a\nstore swapped to a different space cannot re-sign over the wrong broker's creds. `preflight` is a\ncaller-supplied proof that the broker accepts the credential. The reference `Manager` passes a\n`probeConnect` over its `servers`. It gates **every** candidate before overwriting the last-good,\nwhether the signer is a full bundle or a stripped projection: a bundle's JWT chain proves only that\nit is self-consistent and\nnamed the space, NOT that its account is the broker's *current* account for that space (two\n`createSpaceAuth(space)` calls yield same-named, different-account chains), so a same-label alternate\nsigner would otherwise mint a broker-dead cred and clobber the good one. The offline local repair (`doctor auth --fix`) has no preflight. It permits the overwrite only\nunder **authority continuity**: the candidate must be signed by the same account signing key (`iss`) as the current\n(already broker-accepted) cred. A same-label alternate account breaks continuity and is refused, full or\nstripped; a legitimate local re-sign is continuous and proceeds without a network. The reference\n`Manager` runs it on a schedule against its **own**\n`secretStore` (see below), so passing the manager and the delivery daemon the *same* store closes the\nrenewal loop end-to-end on an injected backend: the manager reads the signer from the store, re-signs\ninto it, and the daemon adopts each generation on a preflight-proven 75% timer. The stock\ncross-host composition cannot satisfy that by writing one filesystem and fingerprinting another:\n`Manager.start()` and every later remint challenge the daemon's `reloadStoreIdentity` and a\ndivergent pair is refused naming both stores. The identity is the store the daemon actually\nreloads: an injected coordinate, the workstation root only when `--creds` is\n`<root>/.cotal/<spaceSegment(space)>/delivery.creds` (matching the canonical arm), the\nfile's own directory for any other `--creds` path, or the workstation root. Uninjected\n`--creds` that names one real workstation while process cwd resolves another is refused\nat start, naming both, because membership-rw still uses `findCotalRoot`. A `--creds`\npath that is not under any `.cotal` tree is not that case and is not refused here. It never\nwalks ancestors with `findCotalRoot`. No bound daemon is not a named\nstore, so start proceeds; a later daemon on a foreign store is refused on the next remint.\nThe first-party filesystem adapter declares its workspace-root identity on the store itself. Other\ninjected adapters declare their stable coordinate on `SecretStore.identity`, or name it in\n`COTAL_SECRET_STORE` on both processes. It never throws: it\nreturns per-file results (`skipped: \"no-auth\"` when the store holds no signer records),\nso the caller must check them or the cred still rides to expiry. A composition whose signer lives in\nKMS/Vault simply injects that store; no bespoke renewal is needed. A `--creds` file path must be\nreplaced atomically before the 75% read. The signer can now be injected behind the store seam, which\nresolves custody. The remaining hosted gap is signer **isolation**. The seed is still decrypted\nin-process at the manager's uid, so it needs an OS sandbox or remote signer.\n\n### Supervisor signing authority\n\n`@cotal-ai/manager` exports the `Manager` class; there is **no** `runSupervise(opts)` runner. The\nprivate CLI `runManager` also does broker-reachability checks, space/default resolution,\nroster/launch parsing and materialization, installed-extension resolution, signal handling, staged\npre-spawn, and the forever wait. A host composes that lifecycle itself around `Manager`:\n\n```ts\nimport { Manager } from \"@cotal-ai/manager\";\nconst mgr = new Manager({ space, servers: brokerUrl, workspaceRoot });\nawait mgr.start(); // then wire your own SIGINT/SIGTERM -> mgr.stop()\n```\n\nUnlike delivery, the manager is **not** a pre-minted-scoped-cred daemon (auth-service is also a\nsigner: it holds fewer artifacts than the full trust bundle, but its data-account signing seed still\ngrants complete data-account mint authority on compromise, so this is not least-privilege). On `start()`\nthe manager reads its space's full trust chain **through its `secretStore`** (`getSpaceAuth(this.secrets,\nthis.space)`, composed from `auth/broker.json` + `auth/account.<key>.json`; a container may instead\nmount a stripped signer bundle at the legacy `auth/auth.json` key) and **self-mints** its supervisor cred and renewals from the\ndata-account signing seed. In static mode it also mints every per-agent cred from that seed; in user\nmode agents instead receive callout-minted bearers, but the manager still holds the signing seed for\nits own creds and renewal. So a hosted supervisor is a **trusted per-tenant account-signer process**,\nnot a least-privilege connect client. It additionally requires a `~/.cotal/meshes/space.<key>.json`\nregistry record and the workspace user-auth marker to start in user mode. `ManagerOptions.secretStore`\ninjects the one `SecretStore` the manager uses for **the signer itself (the split trust\nrecords)**, daemon-credential renewal (`remintDaemonCreds`), and per-agent secrets,\ndefaulting to the workspace filesystem store; pass the delivery daemon the *same* store for end-to-end\nhosted renewal. The store declares the same identity on both processes, or both set\n`COTAL_SECRET_STORE` to the same coordinate. The manager\nremints no daemon credential when the daemon names a different store, including a daemon that binds\nafter start; it keeps running and serving its own agents, so one space can carry a manager on more\nthan one workspace root. Pointing several managers at one coordinate is safe: the store identity\nalone cannot pick an owner (it carries no holder and no tiebreak, so every manager sharing the store\nmatches), so the manager that also holds the space's renewal lease is the one that remints and the\nrest skip it. Without that lease two owners would remint on independent timers with no ordering\nbetween them, and one write would land between the other's re-sign and its fingerprint-only\n`reloadCreds`. The signer IS now injectable: a hosted composition injects a KMS/Vault store and no\nsigning seed lands on the hosted disk. What remains is signer **isolation**. The seed is decrypted\nin-process at the manager's uid. That issue needs an OS sandbox or remote signer; it is no longer a\ncustody problem. The other knobs are `workspaceRoot` and the process-global `COTAL_HOME`.\n\n> Scope note: the **static-auth** operator paths (`cotal spawn`/`join`/`status`/`web`, via\n> `mesh-target` → `connect`/`preflight`) still read the signer from the local split records (sync\n> `loadSpaceAuth`). That is the single-machine composition, where the signer is on local disk by the\n> static-auth model; multi-tenant hosting runs **user mode**, which never mints from on-disk trust. The\n> store-injectable signer path is the hosted-server set: the manager, `remintDaemonCreds`, and delivery.\n\nThe typed remote-manager authority contract includes a one-shot terminal phase. A host implements\n`remoteAuthority.prepareAgentRetirement` to revoke the managed grant and finish its resumable\nrelease while preserving the UID, then `remoteAuthority.mintRetirementRequester` returns the\nhost-signed JWT for a fresh participant-owned nkey. The credential is pinned to the authenticated\nowner, server-derived manager serve principal, current instance epoch, and exact target lifecycle.\nThe manager then uses the existing auth `retireLifecycle` rail with the operation id derived by\n`managedRetirementOpId(target.lifecycleUid)`. This derivation is the reference remote-Manager\ncomposition's closed contract, not a rule for every retirement entry point; interactive retirement\nkeeps its existing operation identity and remains compatible. The `retireLifecycle` rail independently\nrecomputes the managed id from its broker-pinned target before any gate, head, intent, or barrier\naccess, so mint-time validation is not the terminal boundary. A failure keeps\nthe alias held. This does not expose the auth barrier or give the participant signer authority.\n\nA host that resumes retained managed actors also implements\n`remoteAuthority.validateRetainedAgent`. The participant sends back the actor token and sentinel it\nalready holds, plus the `nextRegistrationProof` returned by the activation response. That proof is\nhost-issued after registration and binds the manager owner, actor, lifecycle, identity nkeys, current\nregistration revision, and serving epoch. The host checks it against the current open manager gate,\nvalidates the retained secrets against its current managed row, and returns only the non-secret\nauthority shape. The manager binds every result coordinate and the returned authority back to its\ninventory before use. Do not copy the provider's `issuer.json` or `callout.json` into the participant\nstore. Both contain private signing or exchange authority.\n\nThe same composition supplies `remoteAuthority.agentBearerExchangeUrl`, the pinned public auth-service\nbase used by retained children. Remote adoption launches `agent-bearer --exchange-url <base>`; it must\nnot select the local `--dir` arm, which depends on a host-only auth-service process record.\n\nRemote user-mode managers must also supply `remoteAuthority.authorizeAdmin`. The manager builds each\nrequest only from the caller tuple parsed from the broker-authenticated endpoint subject, then relays\nthat tuple over the current registered manager lifecycle. HTTPS does not separately authenticate the\nrelayed caller. The host authenticates the manager operator, binds the request to the current open\nmanager gate, registration proof, serving epoch, and identity nkeys, then reads the caller's unified\nauthoritative row fresh. It returns only the manager owner and `authorized: boolean`, with every request\ncoordinate echoed. Missing, revoked, narrowed, foreign-owner, and stale-lifecycle callers all return\n`false`; malformed coordinates or corrupt and unavailable authority state fail the operation. The\nparticipant never reads or mirrors the host ledger, and the remote branch has no local fallback. The\nsame callback gates all `manager.admin` handlers, any-mode cross-owner control, and `ps` or `inspect`\ncross-owner visibility. Launch keeps its owner-equality policy.\n\nThe remote authority's instance executor remains the scoped maintenance credential for clean service\nderegistration and exact instance registration operations. It carries no records-stream consumer\nlifecycle authority. The manager's boot `goalidx` sweep uses the authenticated host operation, which\nreturns parsed `goalidx.manager.<owner>.>` entries for that owner only. The host keeps the sealed\nconsumer connection and its create/delete rights. The executor remains a five-minute credential for\nregistration operations and clean deregistration. A hosted process expected to\nrun beyond that window cannot yet renew it in place. Clean deregistration then fails loud and the\noperator removes the stale instance with `cotal deregister-instance`. Wiring the typed `renew` phase\ninto the running manager remains required for unattended long-lived hosting.\n\n**Signer isolation needs an OS sandbox.** The default pty runtime\nruns agent children under the *same* OS uid and the *same* `workspaceRoot`, so mode-0600 on\nthe trust records does not stop a hostile same-uid agent from reading their absolute paths. The reference\n[deploy](deploy.md) tree does not solve this: it mounts the signer into the agent's own container, so\nits phase-1 boundary isolates agents from each other, not the signer from the agent. A hosted\ncomposition must run the manager/minter that holds the signer in a different uid, container, or mount\nnamespace from the agent children, which mount no signer at all; that split is future\nhosted-composition work, so until it (or a remote/injected minter) exists, do not run untrusted\nagents under this manager.\n\n## Provisioning a space (one-space reference shape)\n\n```ts\nimport { createSpaceAuth, setupSpaceStreams, ensureDefaultDeliveryClass, mintCreds, newIdentity } from \"@cotal-ai/core\";\nconst auth = await createSpaceAuth(space); // trust bundle (in-memory seeds)\nconst provisionerCreds = await mintCreds(auth, newIdentity(), \"provisioner\");\nawait setupSpaceStreams({ servers: brokerUrl, space, creds: provisionerCreds });\n// SPEC section 4: write the default delivery class at space creation so it is wire-discoverable,\n// never inferred from the resolution fallback. A daemon-backed space is \"durable\".\nawait ensureDefaultDeliveryClass({ servers: brokerUrl, space, creds: provisionerCreds, deliveryClass: \"durable\" });\nconst deliveryCreds = await mintCreds(auth, newIdentity(), \"delivery\");\n// put deliveryCreds into your SecretStore under deliveryCredsKey(space, { injected: true })\n// (@cotal-ai/workspace) before booting delivery — the key is per-space, not the bare kind.\n```\n\nRendering the broker config for a user-auth space is `serverConfig(broker, spaces, { storeDir,\nmaxFileStore?, extraAccounts })`, where `extraAccounts` must include the callout account from\n`createCalloutAuth` so the auth-service has a broker account to answer on. That account never shares\nthe data account. `maxFileStore` is an optional positive integer byte cap; any other value throws.\n\nBroker trust and space accounts are separate authorities: `createBrokerAuth` mints the one\noperator + system account a broker trusts, `createSpaceAccountAuth(broker, space)` signs each\ntenant's data account under it, and `serverConfig(broker, spaces, opts)` renders them all into one\nconfig. A host composition can therefore provision several spaces on one broker today. `cotal up`\nrenders that config from every tenant the root's auth directory holds, so booting one space keeps\nthe broker trusting its siblings, and it refuses to render at all while any account record is\nunreadable. The rest of the CLI lifecycle is still broker-wide: `down`, `clean` and `backup` refuse\non a multi-space root rather than scoping to one tenant, and the per-space lifecycle is the\nremaining multi-space operator layer. See\n[Known gaps](#hosted-composition-gaps).\n\n## Hazardous provisioning primitives\n\n`mintCreds`, the full `Profile`/`CredentialKind` matrix, `createSpaceAuth`, and `stripSpaceAuth` are\nlow-level operator primitives. Handle them as account-authority material:\n\n- A holder of a `SpaceAuth` (or a `stripSpaceAuth` bundle, which **keeps** the data signing seed) is\n a fully-trusted tenant-account authority: it can mint `admin`, `provisioner`, and destructive\n profiles, not merely `supervisor`, and mint a DM-reading identity. `createSpaceAuth`'s full result\n holds operator, system, and account seeds in memory.\n- Choose `profile` and `MintOpts` from **server-side constants**, never from tenant input. `MintOpts`\n can widen the bounded TTL defaults; cap it at your boundary. `CREDENTIAL_LIFETIMES` is a policy\n record, not an authorization boundary.\n- Never log signer material or export it into env. Do not co-locate signer access with an untrusted\n connector/runtime process at the same OS uid (file permissions do not contain a same-uid reader;\n see the manager's isolation note). Segregate per tenant; rotate on compromise\n (`rotateDataAccountSigningKey`).\n\n## Hosted composition gaps\n\nThe primitives above are present as exports, but three capabilities are **not** cleanly composable\nfrom the public contract today. Each is tied to work in flight; a host either waits for the seam or\nscopes the capability out. None is a wire concern.\n\n1. **Delivery immediate live eviction and a fully-hosted membership feed.** The renewable\n `membership-rw.creds` is now a `SecretStore` kind. `startMembership` reads it through the\n injected store, and the manager re-signs it there. The graph-feed writer therefore renews on a hosted\n backend (its data connection adopts each generation on a preflight-proven 75% timer). What still\n reads from a fixed on-disk path are the *static* `membership-observer.creds` and\n `connection-evictor.creds` ($SYS creds, minted at the `up` that provisions the account and renewed by `up --rotate-sys`) and `membership.json`\n (`{accountId}`, non-secret config); those, plus the private provisioning wrapper, keep immediate\n live eviction and a fully-hosted feed a partial gap. Missing files degrade membership to\n traffic-only and make live eviction refuse (loudly). The supported delivery contract here is the\n Plane-3 durable backstop.\n2. **Supervisor signer isolation.** `ManagerOptions.secretStore` now injects the one `SecretStore` the\n manager reads/writes every secret through, including the composed `SpaceAuth`\n signer (the split trust records), its daemon-cred renewal, and its per-agent kinds. What remains is process\n isolation: the manager still decrypts the signer in-process at its uid, so untrusted agent children\n must run under a different uid/container/mount namespace or behind a future remote signer.\n3. **Per-space lifecycle on a shared broker.** The trust layer is multi-space\n (`createBrokerAuth` + `createSpaceAccountAuth` + N-space `serverConfig`, persisted as\n `broker.json` + `account.<key>.json`) and `cotal up` renders the whole tenant list, but there is\n no per-space provisioning verb and no per-space teardown/backup/restore: the CLI's broker-wide\n lifecycle verbs refuse on a multi-space root, naming the tenants.\n This is the remaining multi-space operator layer.\n4. **A non-Better-Auth production IdP.** The exchange core (`createIdpBridge`) is EdDSA-generic, but\n the stock provider and login client are Better-Auth-endpoint-shaped, `cotalAuthProvider`\n self-registers on import (colliding with a host-owned provider under `resolveAuthProvider`), and\n the login flow speaks Better Auth's device-code endpoints. A different IdP is a host-built auth\n composition on the low-level primitives, not a configuration change (see\n [the IdP callout contract](identity-and-auth.md#the-idp-callout-contract)).\n\n## Hosted durability\n\nSpace-durable **coordination** state (chat/DM/task history, live presence, membership runtime, the\ndurable ACL registry, leases) lives in **JetStream**, written by the delivery daemon and the\nendpoints. It is broker-resident and needs no host-side durable path.\n\nWhat is **not** in JetStream, and is hosting-critical, is trust and authorization state a host must\nplace and keep:\n\n| state | class | where today | hosted injection |\n|---|---|---|---|\n| full `SpaceAuth` trust chain (`auth/broker.json` + `auth/account.<key>.json`, composed; a stripped signer bundle may instead be mounted at the legacy `auth/auth.json` key) | signing authority | `SecretStore` | `SecretStore` (manager + renewal) |\n| auth kinds: callout account/creds/xkey, issuer private keys, owner-derivation secret, data-signer projection | signing/identity authority | four `SecretStore` kinds | `SecretStore` (auth-service) |\n| `delivery.creds` | standing scoped cred | `SecretStore` or `--creds` | `SecretStore` (delivery) |\n| actor ledger, IdP pin | authorization + trust config | ambient `userAuthStateDir(findCotalRoot(), space)` | none (root-relative; not `store`/`COTAL_HOME`) |\n| `membership-rw.creds` | standing scoped cred | `SecretStore` | `SecretStore` (delivery + manager renewal) |\n| membership-observer / connection-evictor creds + `membership.json` | scoped $SYS creds / config | workspace filesystem | none (see gap 1) |\n| manager agent creds, actor tokens, sentinel creds | lifecycle authority | `SecretStore` | `SecretStore` (manager `secretStore`) |\n| `~/.cotal/meshes/space.<key>.json` record (holds IdP trust pins/root pointers) | non-secret, integrity-critical | machine home | process-global `COTAL_HOME` only |\n| auth-health, renewal records | non-secret diagnostics | workspace filesystem | `workspaceRoot` |\n\nThe `SpaceAuth` trust chain and the auth-service store kinds are **separate** identities/projections,\nnever parts of one document. `auth-service.json` (the live exchange capability) is ephemeral runtime\nstate, not durable, but is sensitive while the daemon runs. `@cotal-ai/workspace` is machine-local\noperator tooling by design; personas, PID files, and the `current-mesh` pointer are truly local and\nmust **not** sit on a hosted durable path. Everything classed above as an authority is what a hosted\ncomposition must provision and persist: signer-bearing server secrets now have `SecretStore` seams;\nthe remaining non-injectable rows are the explicit ambient `workspaceRoot`/cwd paths above.\n\n## See also\n\n- [Substrate stability](stability.md): what v0.3 and the 0.x packages guarantee, and the projected v0.4 break.\n- [Identity and auth](identity-and-auth.md): the profile matrix, the signer, and the IdP callout contract.\n- [Delivery daemon](delivery-daemon.md): the Plane-3 durable backstop.\n- [Deploy](deploy.md): the reference container against an external broker.\n"
164
+ "body": "# Embedding Cotal\n\n> **Guide** (informative) · **For:** implementers building a service on top of Cotal · **Prereqs:** [Architecture](architecture.md), [Identity and auth](identity-and-auth.md), [Delivery daemon](delivery-daemon.md)\n\nThe `cotal` binary in this repo is one composition root: an operator CLI. A separate service\n(for example a hosted, multi-tenant Cotal) does not fork this repo. It writes its **own**\ncomposition root that depends on the published `@cotal-ai/*` packages and imports the surfaces it\nwants. `bin/cotal.ts` uses the same composition pattern. This page is the contract for that: what is a real library\nexport you can build against, how to boot the server-side daemons from those exports, and where the\ncurrent export surface stops short of a fully hosted composition.\n\nThis is the \"guarded substrate\" boundary in practice. Nothing here reveals or assumes a specific\nhost; it documents the public seams any embedder composes.\n\n## What you embed\n\nThe supported reference shape here is **one broker operator serving one space** (one tenant: a\ndedicated data account, under an operator that also holds the system account and a quarantined\nauth-callout account) plus three standalone processes. The trust layer itself composes many spaces\nunder one broker operator today (`createBrokerAuth` + `createSpaceAccountAuth` + N-space\n`serverConfig`); what does not exist yet is the per-space **lifecycle** on a shared broker (see\n[Known gaps](#hosted-composition-gaps)). The three processes:\n\n| daemon | package | what it is |\n|---|---|---|\n| auth-service | `@cotal-ai/auth` | the NATS auth callout, the IdP token exchange, and JWKS. Plane 1 to Plane 2. |\n| delivery | `@cotal-ai/delivery` | the Plane-3 durable backstop: fan-out writer plus trusted reader, per space. |\n| supervise | `@cotal-ai/manager` | the per-machine agent lifecycle (spawn/despawn/attach), per space. |\n\n`mint`, `deliver`, and `auth-service` expose their behavior as direct library primitives, and the\nsupported one-space bootstrap below re-composes from exported low-level primitives. `supervise` and\nthe full `up` orchestration are **not** public runners: `up` also does broker bring-up, restore,\nprocess and registry management, and lifecycle work, and `supervise`'s orchestration is private (see\n[Supervisor signing authority](#supervisor-signing-authority)).\n\n## The export surface\n\nEverything below is a real export of a published package, reachable from the package root (each\npackage publishes only `.` via `dist/index.{js,d.ts}` and ships `files: [\"dist\"]`). Type-only names\nare marked; import them with `import type`.\n\n**Daemon runners and lifecycle**\n\n| symbol | package | purpose |\n|---|---|---|\n| `runAuthService(args, store?)` | `@cotal-ai/auth` | boot the auth-service daemon; `store` injects the secret material. |\n| `runDelivery(args, store?)` | `@cotal-ai/delivery` | boot the delivery daemon; `store` injects the scoped `delivery` cred. |\n| `deliveryCredsKey(space, composition)`, `membershipRwCredsKey(space, composition)` | `@cotal-ai/workspace` | build the secret-store keys the delivery cred and the membership feed's rw cred are read/re-signed under. Keys are **per-space**: `space.<hex>/<kind>`. A hosted composition passes `{ injected: true }`. |\n| `DELIVERY_CREDS_KIND`, `MEMBERSHIP_RW_CREDS_KIND` | `@cotal-ai/workspace` | the operator-facing KIND names (`delivery.creds`, `membership-rw.creds`) those keys are built from, and what renewal results report. A kind is **not** a key: putting a cred under the bare kind writes the pre-0.4 flat location, which nothing reads. |\n| `Manager`, `ManagerOptions` *(type)* | `@cotal-ai/manager` | construct and run a supervisor in-process; `ManagerOptions.secretStore` injects the one store it reads/writes every secret through. `ManagerOptions.remoteAuthority` is the hosted manager-service authority bundle, including host-owned release, retained-validation, goal-index, and serve-time admin-authorization callbacks. |\n| `createRuntime`, `Runtime` *(type)* | `@cotal-ai/manager` | resolve the spawn backend (pty built in). |\n\n**Provisioning and minting** (all `@cotal-ai/core`)\n\n| symbol | purpose |\n|---|---|\n| `createBrokerAuth(label)` | mint BROKER trust: the operator and system account one nats-server trusts. One per broker, shared by every space on it. |\n| `createSpaceAccountAuth(broker, space)` | mint one space's own data account, signed by that broker's operator: the add-a-tenant primitive. |\n| `createSpaceAuth(space)` | the one-space convenience: broker trust + one account in a single composed bundle. |\n| `setupSpaceStreams({ servers, space, creds })` | create the space's JetStream streams. |\n| `ensureDefaultDeliveryClass({ servers, space, creds?, deliveryClass })` | write the space's default delivery class at creation so it is wire-discoverable (SPEC section 4). |\n| `serverConfig(broker, spaces, { storeDir, maxFileStore?, extraAccounts?, port?, host? })` | render the broker config: one operator, N space accounts. `storeDir` is required, `maxFileStore` caps JetStream file storage in bytes (omitted, nats-server's dynamic default applies), and `extraAccounts` preloads the auth-callout account. |\n| `mintCreds(auth, identity, profile, opts?)` | mint a scoped cred for any `Profile`. |\n| `mintMembershipObserverCreds`, `mintConnectionEvictorCreds` | mint the membership/eviction scoped creds. |\n| `provisionAgent`, `provisionAgentDurables` | create a principal's bind-only durables. |\n| `newIdentity`, `stripSpaceAuth` | a fresh nkey identity; a stripped signer bundle (data signing seed only). |\n| `Profile`, `CredentialKind`, `MintOpts`, `SpaceAuth` *(types)*, `CREDENTIAL_LIFETIMES` | the profile matrix and cred lifetime policy. |\n\n**Auth building blocks** (all `@cotal-ai/auth`)\n\n| symbol | purpose |\n|---|---|\n| `createCalloutAuth`, `startAuthCallout` | the NATS auth-callout responder. |\n| `createUserTokenIssuer`, `pinnedJwksResolver` | mint and verify the Cotal user bearer. |\n| `createIdpBridge` | exchange a verified IdP JWT for a Cotal bearer (see [the callout contract](identity-and-auth.md#the-idp-callout-contract)). |\n| `deriveOwnerToken`, `validateUserToken` | owner derivation; strict bearer validation. |\n| `cotalAuthProvider` | the self-registering `auth-provider` extension. |\n| `ensureCalloutAuth`/`loadCalloutAuth`, `ensureIssuer`/`loadIssuer`, `ensureOwnerSecret`/`loadOwnerSecret` | read/write the auth secret kinds through a `SecretStore`. |\n\n**Seams and the wire** (all `@cotal-ai/core` unless noted)\n\n| symbol | purpose |\n|---|---|\n| `SecretStore` *(type)* | the durable hosted-secret seam (get/put/delete); `get()` returns raw seeds/keys into process memory, so it is a blob seam, not HSM/KMS signing. |\n| `FsSecretStore`, `workspaceSecretStore(root)` | the filesystem default. **These live in `@cotal-ai/workspace`, not core.** |\n| `AuthProvider` *(type)*, `Connector` *(type)*, `Runtime` *(type)*, `Command` *(type)* | the extension contracts; implementations self-register on import. |\n| `registry` | the shared registry a composition root pulls surfaces into. |\n| `CotalEndpoint`, subjects, message types | the wire client and shapes. |\n| `ParsedArgs` *(type)* | the shape the daemon runners take (see below). |\n\nThe runners take a CLI-shaped `ParsedArgs`, not a typed options object, so a host fabricates one:\n\n```ts\nconst args: ParsedArgs = { values: { space, server, port: \"0\" }, positionals: [], raw: [] };\n```\n\n### Long-lived endpoints take a bearer function\n\n`EndpointOptions.bearer` accepts either a string or a function, and the difference is not stylistic.\nA string is minted once, so when it expires (which it will: callout bearers live minutes) the\nendpoint has nothing to renew with. It will not present the dead token to the broker, since that is\na guaranteed denial that still costs a full auth-callout round trip. It refuses to reconnect, emits\n`warning` saying which case it is in, and retries on a widening backoff until the process\nre-authenticates and rebuilds it. Retry notices use `warning` rather than `error` because Node\nrethrows an unhandled `error` event and would kill a host the endpoint is still trying to recover.\n\nPass a **function** for anything that outlives one bearer. That is a renewal source: it is called\nahead of each expiry and again whenever a reconnect finds the cached bearer dead, and it requires\nexplicit `card.owner` and `card.actor`. The first-party surfaces already do this\n(`UserViewAuth.source`, the connector's `agentBearerCommand`). A string bearer is for a short\none-shot connection.\n\nLong-lived hosts must also subscribe to the endpoint's `warning` event. It carries conditions the\nendpoint is surviving, including failed credential renewal and reconnect retries. A host may choose\nto ignore warnings for a one-shot endpoint whose awaited operation owns the verdict, but that choice\nshould be explicit. An unhandled warning is nonfatal and silent.\n\n## Booting the daemons\n\n### auth-service\n\n`runAuthService(args, store?)` reads its provisioned long-lived secret kinds (service keys, callout\naccount, issuer keys, owner secret) through the injected `SecretStore`; a host provisions those into\nthe store first. It is a **signer and identity authority**, not a scoped daemon: at runtime it holds\nthe data-account and callout-account signing seeds, the issuer's private JWKs, and the\nowner-derivation secret in process memory (`SecretStore.get` exports raw values). The IdP pin and the\nactor ledger are **not** store-injected: `runAuthService` resolves them under\n`userAuthStateDir(findCotalRoot(), space)`, a path relative to the process working directory, so a\nhost provisions those into that exact directory (neither `store` nor `COTAL_HOME` selects it). It\nalso writes an ephemeral `auth-service.json` discovery file there that carries the live exchange\ncapability. That file appears only after every plane is bound, so waiting on it is the readiness\nsignal: a host that also passes the daemon's pid to the provider's `ready()` gets a process-bound\nwait. The wait extends past the base timeout while that pid is alive, up to a fixed bound, and it\nends at once when the pid exits.\n\n```ts\nimport { runAuthService } from \"@cotal-ai/auth\";\n// store implements SecretStore over your secret backend; get() returns raw seeds into memory.\n// Provision the auth secret kinds into the store, AND the IdP pin + actor ledger under\n// userAuthStateDir(findCotalRoot(), space), before this call.\nawait runAuthService(\n { values: { space, server: brokerUrl, port: \"8081\" }, positionals: [], raw: [] },\n store,\n);\n```\n\n### delivery\n\n`runDelivery(args, store?)` runs from a **pre-minted scoped `delivery` cred** and never loads the\nsigner. Provide the cred either through the injected store (under\n`deliveryCredsKey(space, { injected: true })`) or with a\n`--creds` file; the two are mutually exclusive. The daemon re-fetches the cred from the store at 75%\nof its JWT lifetime and fails loud rather than riding to expiry, so **something must re-sign a fresh\ncred into that same store**. When that read finds the previous generation still there, the daemon\nreports the missed remint and retries in 60 seconds. The current cred stays live until its expiry,\nand the store is read once per retry rather than once per second.\n\n```ts\nimport { runDelivery } from \"@cotal-ai/delivery\";\nawait runDelivery({ values: { space, server: brokerUrl }, positionals: [], raw: [] }, store);\n```\n\nThat renewal is a **signer** operation, not the delivery daemon's:\n`remintDaemonCreds(root, space, store?, { preflight? })` (`@cotal-ai/workspace`) reads the `SpaceAuth`\nsigner **through the same resolved `store`** (`getSpaceAuth(store ?? workspaceSecretStore(root), space)`,\nkeys `auth/broker.json` + `auth/account.<key>.json`; the pre-split `auth/auth.json` monolith is\nmigration input and the container signer mount only) and re-signs the daemon creds (`delivery.creds` and the membership feed's\n`membership-rw.creds`) back into that store. The injected `store` is both the signer source and the\ncredential destination, never a split. `space` is **required** and validated against the store's signer, so a\nstore swapped to a different space cannot re-sign over the wrong broker's creds. `preflight` is a\ncaller-supplied proof that the broker accepts the credential. The reference `Manager` passes a\n`probeConnect` over its `servers`. It gates **every** candidate before overwriting the last-good,\nwhether the signer is a full bundle or a stripped projection: a bundle's JWT chain proves only that\nit is self-consistent and\nnamed the space, NOT that its account is the broker's *current* account for that space (two\n`createSpaceAuth(space)` calls yield same-named, different-account chains), so a same-label alternate\nsigner would otherwise mint a broker-dead cred and clobber the good one. The offline local repair (`doctor auth --fix`) has no preflight. It permits the overwrite only\nunder **authority continuity**: the candidate must be signed by the same account signing key (`iss`) as the current\n(already broker-accepted) cred. A same-label alternate account breaks continuity and is refused, full or\nstripped; a legitimate local re-sign is continuous and proceeds without a network. The reference\n`Manager` runs it on a schedule against its **own**\n`secretStore` (see below), so passing the manager and the delivery daemon the *same* store closes the\nrenewal loop end-to-end on an injected backend: the manager reads the signer from the store, re-signs\ninto it, and the daemon adopts each generation on a preflight-proven 75% timer. The stock\ncross-host composition cannot satisfy that by writing one filesystem and fingerprinting another:\n`Manager.start()` and every later remint challenge the daemon's `reloadStoreIdentity` and a\ndivergent pair is refused naming both stores. The identity is the store the daemon actually\nreloads: an injected coordinate, the workstation root only when `--creds` is\n`<root>/.cotal/<spaceSegment(space)>/delivery.creds` (matching the canonical arm), the\nfile's own directory for any other `--creds` path, or the workstation root. Uninjected\n`--creds` that names one real workstation while process cwd resolves another is refused\nat start, naming both, because membership-rw still uses `findCotalRoot`. A `--creds`\npath that is not under any `.cotal` tree is not that case and is not refused here. It never\nwalks ancestors with `findCotalRoot`. No bound daemon is not a named\nstore, so start proceeds; a later daemon on a foreign store is refused on the next remint.\nThe first-party filesystem adapter declares its workspace-root identity on the store itself. Other\ninjected adapters declare their stable coordinate on `SecretStore.identity`, or name it in\n`COTAL_SECRET_STORE` on both processes. It never throws: it\nreturns per-file results (`skipped: \"no-auth\"` when the store holds no signer records),\nso the caller must check them or the cred still rides to expiry. A composition whose signer lives in\nKMS/Vault simply injects that store; no bespoke renewal is needed. A `--creds` file path must be\nreplaced atomically before the 75% read. The signer can now be injected behind the store seam, which\nresolves custody. The remaining hosted gap is signer **isolation**. The seed is still decrypted\nin-process at the manager's uid, so it needs an OS sandbox or remote signer.\n\n### Supervisor signing authority\n\n`@cotal-ai/manager` exports the `Manager` class; there is **no** `runSupervise(opts)` runner. The\nprivate CLI `runManager` also does broker-reachability checks, space/default resolution,\nroster/launch parsing and materialization, installed-extension resolution, signal handling, staged\npre-spawn, and the forever wait. A host composes that lifecycle itself around `Manager`:\n\n```ts\nimport { Manager } from \"@cotal-ai/manager\";\nconst mgr = new Manager({ space, servers: brokerUrl, workspaceRoot });\nawait mgr.start(); // then wire your own SIGINT/SIGTERM -> mgr.stop()\n```\n\nUnlike delivery, the manager is **not** a pre-minted-scoped-cred daemon (auth-service is also a\nsigner: it holds fewer artifacts than the full trust bundle, but its data-account signing seed still\ngrants complete data-account mint authority on compromise, so this is not least-privilege). On `start()`\nthe manager reads its space's full trust chain **through its `secretStore`** (`getSpaceAuth(this.secrets,\nthis.space)`, composed from `auth/broker.json` + `auth/account.<key>.json`; a container may instead\nmount a stripped signer bundle at the legacy `auth/auth.json` key) and **self-mints** its supervisor cred and renewals from the\ndata-account signing seed. In static mode it also mints every per-agent cred from that seed; in user\nmode agents instead receive callout-minted bearers, but the manager still holds the signing seed for\nits own creds and renewal. So a hosted supervisor is a **trusted per-tenant account-signer process**,\nnot a least-privilege connect client. It additionally requires a `~/.cotal/meshes/space.<key>.json`\nregistry record and the workspace user-auth marker to start in user mode. `ManagerOptions.secretStore`\ninjects the one `SecretStore` the manager uses for **the signer itself (the split trust\nrecords)**, daemon-credential renewal (`remintDaemonCreds`), and per-agent secrets,\ndefaulting to the workspace filesystem store; pass the delivery daemon the *same* store for end-to-end\nhosted renewal. The store declares the same identity on both processes, or both set\n`COTAL_SECRET_STORE` to the same coordinate. The manager\nremints no daemon credential when the daemon names a different store, including a daemon that binds\nafter start; it keeps running and serving its own agents, so one space can carry a manager on more\nthan one workspace root. Pointing several managers at one coordinate is safe: the store identity\nalone cannot pick an owner (it carries no holder and no tiebreak, so every manager sharing the store\nmatches), so the manager that also holds the space's renewal lease is the one that remints and the\nrest skip it. Without that lease two owners would remint on independent timers with no ordering\nbetween them, and one write would land between the other's re-sign and its fingerprint-only\n`reloadCreds`. The signer IS now injectable: a hosted composition injects a KMS/Vault store and no\nsigning seed lands on the hosted disk. What remains is signer **isolation**. The seed is decrypted\nin-process at the manager's uid. That issue needs an OS sandbox or remote signer; it is no longer a\ncustody problem. The other knobs are `workspaceRoot` and the process-global `COTAL_HOME`.\n\n> Scope note: the **static-auth** operator paths (`cotal spawn`/`join`/`status`/`web`, via\n> `mesh-target` → `connect`/`preflight`) still read the signer from the local split records (sync\n> `loadSpaceAuth`). That is the single-machine composition, where the signer is on local disk by the\n> static-auth model; multi-tenant hosting runs **user mode**, which never mints from on-disk trust. The\n> store-injectable signer path is the hosted-server set: the manager, `remintDaemonCreds`, and delivery.\n\nThe typed remote-manager authority contract includes a one-shot terminal phase. A host implements\n`remoteAuthority.prepareAgentRetirement` to revoke the managed grant and finish its resumable\nrelease while preserving the UID, then `remoteAuthority.mintRetirementRequester` returns the\nhost-signed JWT for a fresh participant-owned nkey. The credential is pinned to the authenticated\nowner, server-derived manager serve principal, current instance epoch, and exact target lifecycle.\nThe manager then uses the existing auth `retireLifecycle` rail with the operation id derived by\n`managedRetirementOpId(target.lifecycleUid)`. This derivation is the reference remote-Manager\ncomposition's closed contract, not a rule for every retirement entry point; interactive retirement\nkeeps its existing operation identity and remains compatible. The `retireLifecycle` rail independently\nrecomputes the managed id from its broker-pinned target before any gate, head, intent, or barrier\naccess, so mint-time validation is not the terminal boundary. A failure keeps\nthe alias held. This does not expose the auth barrier or give the participant signer authority.\n\nA host that resumes retained managed actors also implements\n`remoteAuthority.validateRetainedAgent`. The participant sends back the actor token and sentinel it\nalready holds, plus the `nextRegistrationProof` returned by the activation response. That proof is\nhost-issued after registration and binds the manager owner, actor, lifecycle, identity nkeys, current\nregistration revision, and serving epoch. The host checks it against the current open manager gate,\nvalidates the retained secrets against its current managed row, and returns only the non-secret\nauthority shape. The manager binds every result coordinate and the returned authority back to its\ninventory before use. Do not copy the provider's `issuer.json` or `callout.json` into the participant\nstore. Both contain private signing or exchange authority.\n\nThe same composition supplies `remoteAuthority.agentBearerExchangeUrl`, the pinned public auth-service\nbase used by retained children. Remote adoption launches `agent-bearer --exchange-url <base>`; it must\nnot select the local `--dir` arm, which depends on a host-only auth-service process record.\n\nRemote user-mode managers must also supply `remoteAuthority.authorizeAdmin`. The manager builds each\nrequest only from the caller tuple parsed from the broker-authenticated endpoint subject, then relays\nthat tuple over the current registered manager lifecycle. HTTPS does not separately authenticate the\nrelayed caller. The host authenticates the manager operator, binds the request to the current open\nmanager gate, registration proof, serving epoch, and identity nkeys, then reads the caller's unified\nauthoritative row fresh. It returns only the manager owner and `authorized: boolean`, with every request\ncoordinate echoed. Missing, revoked, narrowed, foreign-owner, and stale-lifecycle callers all return\n`false`; malformed coordinates or corrupt and unavailable authority state fail the operation. The\nparticipant never reads or mirrors the host ledger, and the remote branch has no local fallback. The\nsame callback gates all `manager.admin` handlers, any-mode cross-owner control, and `ps` or `inspect`\ncross-owner visibility. Launch keeps its owner-equality policy.\n\nThe remote authority's instance executor remains the scoped maintenance credential for clean service\nderegistration and exact instance registration operations. It carries no records-stream consumer\nlifecycle authority. The manager's boot `goalidx` sweep uses the authenticated host operation, which\nreturns parsed `goalidx.manager.<owner>.>` entries for that owner only. The host keeps the sealed\nconsumer connection and its create/delete rights. The executor remains a five-minute credential for\nregistration operations and clean deregistration. A hosted process expected to\nrun beyond that window cannot yet renew it in place. Clean deregistration then fails loud and the\noperator removes the stale instance with `cotal deregister-instance`. Wiring the typed `renew` phase\ninto the running manager remains required for unattended long-lived hosting.\n\n**Signer isolation needs an OS sandbox.** The default pty runtime\nruns agent children under the *same* OS uid and the *same* `workspaceRoot`, so mode-0600 on\nthe trust records does not stop a hostile same-uid agent from reading their absolute paths. The reference\n[deploy](deploy.md) tree does not solve this: it mounts the signer into the agent's own container, so\nits phase-1 boundary isolates agents from each other, not the signer from the agent. A hosted\ncomposition must run the manager/minter that holds the signer in a different uid, container, or mount\nnamespace from the agent children, which mount no signer at all; that split is future\nhosted-composition work, so until it (or a remote/injected minter) exists, do not run untrusted\nagents under this manager.\n\n## Provisioning a space (one-space reference shape)\n\n```ts\nimport { createSpaceAuth, setupSpaceStreams, ensureDefaultDeliveryClass, mintCreds, newIdentity } from \"@cotal-ai/core\";\nconst auth = await createSpaceAuth(space); // trust bundle (in-memory seeds)\nconst provisionerCreds = await mintCreds(auth, newIdentity(), \"provisioner\");\nawait setupSpaceStreams({ servers: brokerUrl, space, creds: provisionerCreds });\n// SPEC section 4: write the default delivery class at space creation so it is wire-discoverable,\n// never inferred from the resolution fallback. A daemon-backed space is \"durable\".\nawait ensureDefaultDeliveryClass({ servers: brokerUrl, space, creds: provisionerCreds, deliveryClass: \"durable\" });\nconst deliveryCreds = await mintCreds(auth, newIdentity(), \"delivery\");\n// put deliveryCreds into your SecretStore under deliveryCredsKey(space, { injected: true })\n// (@cotal-ai/workspace) before booting delivery — the key is per-space, not the bare kind.\n```\n\nRendering the broker config for a user-auth space is `serverConfig(broker, spaces, { storeDir,\nmaxFileStore?, extraAccounts })`, where `extraAccounts` must include the callout account from\n`createCalloutAuth` so the auth-service has a broker account to answer on. That account never shares\nthe data account. `maxFileStore` is an optional positive integer byte cap; any other value throws.\n\nBroker trust and space accounts are separate authorities: `createBrokerAuth` mints the one\noperator + system account a broker trusts, `createSpaceAccountAuth(broker, space)` signs each\ntenant's data account under it, and `serverConfig(broker, spaces, opts)` renders them all into one\nconfig. A host composition can therefore provision several spaces on one broker today. `cotal up`\nrenders that config from every tenant the root's auth directory holds, so booting one space keeps\nthe broker trusting its siblings, and it refuses to render at all while any account record is\nunreadable. The rest of the CLI lifecycle is still broker-wide: `down`, `clean` and `backup` refuse\non a multi-space root rather than scoping to one tenant, and the per-space lifecycle is the\nremaining multi-space operator layer. See\n[Known gaps](#hosted-composition-gaps).\n\n## Hazardous provisioning primitives\n\n`mintCreds`, the full `Profile`/`CredentialKind` matrix, `createSpaceAuth`, and `stripSpaceAuth` are\nlow-level operator primitives. Handle them as account-authority material:\n\n- A holder of a `SpaceAuth` (or a `stripSpaceAuth` bundle, which **keeps** the data signing seed) is\n a fully-trusted tenant-account authority: it can mint `admin`, `provisioner`, and destructive\n profiles, not merely `supervisor`, and mint a DM-reading identity. `createSpaceAuth`'s full result\n holds operator, system, and account seeds in memory.\n- Choose `profile` and `MintOpts` from **server-side constants**, never from tenant input. `MintOpts`\n can widen the bounded TTL defaults; cap it at your boundary. `CREDENTIAL_LIFETIMES` is a policy\n record, not an authorization boundary.\n- Never log signer material or export it into env. Do not co-locate signer access with an untrusted\n connector/runtime process at the same OS uid (file permissions do not contain a same-uid reader;\n see the manager's isolation note). Segregate per tenant; rotate on compromise\n (`rotateDataAccountSigningKey`).\n\n## Hosted composition gaps\n\nThe primitives above are present as exports, but three capabilities are **not** cleanly composable\nfrom the public contract today. Each is tied to work in flight; a host either waits for the seam or\nscopes the capability out. None is a wire concern.\n\n1. **Delivery immediate live eviction and a fully-hosted membership feed.** The renewable\n `membership-rw.creds` is now a `SecretStore` kind. `startMembership` reads it through the\n injected store, and the manager re-signs it there. The graph-feed writer therefore renews on a hosted\n backend (its data connection adopts each generation on a preflight-proven 75% timer). What still\n reads from a fixed on-disk path are the *static* `membership-observer.creds` and\n `connection-evictor.creds` ($SYS creds, minted at the `up` that provisions the account and renewed by `up --rotate-sys`) and `membership.json`\n (`{accountId}`, non-secret config); those, plus the private provisioning wrapper, keep immediate\n live eviction and a fully-hosted feed a partial gap. Missing files degrade membership to\n traffic-only and make live eviction refuse (loudly). The supported delivery contract here is the\n Plane-3 durable backstop.\n2. **Supervisor signer isolation.** `ManagerOptions.secretStore` now injects the one `SecretStore` the\n manager reads/writes every secret through, including the composed `SpaceAuth`\n signer (the split trust records), its daemon-cred renewal, and its per-agent kinds. What remains is process\n isolation: the manager still decrypts the signer in-process at its uid, so untrusted agent children\n must run under a different uid/container/mount namespace or behind a future remote signer.\n3. **Per-space lifecycle on a shared broker.** The trust layer is multi-space\n (`createBrokerAuth` + `createSpaceAccountAuth` + N-space `serverConfig`, persisted as\n `broker.json` + `account.<key>.json`) and `cotal up` renders the whole tenant list, but there is\n no per-space provisioning verb and no per-space teardown/backup/restore: the CLI's broker-wide\n lifecycle verbs refuse on a multi-space root, naming the tenants.\n This is the remaining multi-space operator layer.\n4. **A non-Better-Auth production IdP.** The exchange core (`createIdpBridge`) is EdDSA-generic, but\n the stock provider and login client are Better-Auth-endpoint-shaped, `cotalAuthProvider`\n self-registers on import (colliding with a host-owned provider under `resolveAuthProvider`), and\n the login flow speaks Better Auth's device-code endpoints. A different IdP is a host-built auth\n composition on the low-level primitives, not a configuration change (see\n [the IdP callout contract](identity-and-auth.md#the-idp-callout-contract)).\n\n## Hosted durability\n\nSpace-durable **coordination** state (chat/DM/task history, live presence, membership runtime, the\ndurable ACL registry, leases) lives in **JetStream**, written by the delivery daemon and the\nendpoints. It is broker-resident and needs no host-side durable path.\n\nWhat is **not** in JetStream, and is hosting-critical, is trust and authorization state a host must\nplace and keep:\n\n| state | class | where today | hosted injection |\n|---|---|---|---|\n| full `SpaceAuth` trust chain (`auth/broker.json` + `auth/account.<key>.json`, composed; a stripped signer bundle may instead be mounted at the legacy `auth/auth.json` key) | signing authority | `SecretStore` | `SecretStore` (manager + renewal) |\n| auth kinds: callout account/creds/xkey, issuer private keys, owner-derivation secret, data-signer projection | signing/identity authority | four `SecretStore` kinds | `SecretStore` (auth-service) |\n| `delivery.creds` | standing scoped cred | `SecretStore` or `--creds` | `SecretStore` (delivery) |\n| actor ledger, IdP pin | authorization + trust config | ambient `userAuthStateDir(findCotalRoot(), space)` | none (root-relative; not `store`/`COTAL_HOME`) |\n| `membership-rw.creds` | standing scoped cred | `SecretStore` | `SecretStore` (delivery + manager renewal) |\n| membership-observer / connection-evictor creds + `membership.json` | scoped $SYS creds / config | workspace filesystem | none (see gap 1) |\n| manager agent creds, actor tokens, sentinel creds | lifecycle authority | `SecretStore` | `SecretStore` (manager `secretStore`) |\n| `~/.cotal/meshes/space.<key>.json` record (holds IdP trust pins/root pointers) | non-secret, integrity-critical | machine home | process-global `COTAL_HOME` only |\n| auth-health, renewal records | non-secret diagnostics | workspace filesystem | `workspaceRoot` |\n\nThe `SpaceAuth` trust chain and the auth-service store kinds are **separate** identities/projections,\nnever parts of one document. `auth-service.json` (the live exchange capability) is ephemeral runtime\nstate, not durable, but is sensitive while the daemon runs. `@cotal-ai/workspace` is machine-local\noperator tooling by design; personas, PID files, and the `current-mesh` pointer are truly local and\nmust **not** sit on a hosted durable path. Everything classed above as an authority is what a hosted\ncomposition must provision and persist: signer-bearing server secrets now have `SecretStore` seams;\nthe remaining non-injectable rows are the explicit ambient `workspaceRoot`/cwd paths above.\n\n## See also\n\n- [Substrate stability](stability.md): what v0.3 and the 0.x packages guarantee, and the projected v0.4 break.\n- [Identity and auth](identity-and-auth.md): the profile matrix, the signer, and the IdP callout contract.\n- [Delivery daemon](delivery-daemon.md): the Plane-3 durable backstop.\n- [Deploy](deploy.md): the reference container against an external broker.\n"
165
165
  },
166
166
  {
167
167
  "slug": "examples",
@@ -210,7 +210,7 @@ export const DOCS_BUNDLE = {
210
210
  "title": "Publishing a release",
211
211
  "kind": "Project (non-normative maintainer notes)",
212
212
  "summary": "Cotal uses Changesets to version and publish the workspace packages under packages/, extensions/, and implementations/ to npm.",
213
- "body": "# Publishing a release\n\n> **Project** (non-normative maintainer notes) · **For:** maintainers shipping Cotal\n\nCotal uses [Changesets](https://github.com/changesets/changesets) to version and publish the\nworkspace packages under `packages/*`, `extensions/*`, and `implementations/*` to npm.\n`examples/**` is ignored, since it is not published.\n\n## 0.11 runtime migration\n\nThe published binary no longer bundles the optional tmux and cmux runtimes. Existing operators\nmust run `cotal ext add @cotal-ai/tmux` or `cotal ext add @cotal-ai/cmux` once after upgrading,\nbefore using `runtime: tmux|cmux` in a manifest or passing `--runtime tmux|cmux`. Missing runtimes\nfail loudly with the matching install command; they never fall back to pty.\n\n## Trusted publishing\n\nTrusted publishing replaces the long-lived `NPM_TOKEN` secret with short-lived OIDC tokens\nissued by GitHub Actions. Each published package must be configured once on npmjs.com.\n\nThe `fixed` group in [`.changeset/config.json`](../.changeset/config.json) is the list that\ngets versioned and published. Derive the package list from it instead of\nmaintaining it by hand. It had drifted by six packages before this was last reconciled.\n\n### Deployment Environment setup\n\nBoth publishing jobs (`version` and `snapshot`) reference a GitHub Environment named\n`npm-publish`. The OIDC assertion rejects tokens that do not carry this environment claim,\nso the Environment must exist and must protect the release ref before the first publish.\n\n1. Go to **Settings > Environments** in the repository.\n2. Create a new Environment named **`npm-publish`**.\n3. Under **Deployment branches and tags**, select **Selected branches and tags** and add\n `main` as the only allowed branch. This restricts OIDC token issuance to runs on `main`.\n4. Optionally add required reviewers if your team wants a manual gate before each release.\n\nThe Environment name is embedded in the workflow, in the OIDC identity assertion, and in\nevery npm trusted-publisher record. All three must use the exact string `npm-publish`.\n\n> **Snapshot releases:** the snapshot job is also bound to the `npm-publish` Environment.\n> If the deployment branch policy allows only `main`, snapshot releases from other branches\n> are refused by the Environment gate before the OIDC exchange. To allow snapshots from\n> additional branches, add those branches to the Environment's deployment policy.\n\n### Per-package trusted publisher configuration\n\nFor **every** published package, `cotal-ai` (the binary), `@cotal-ai/core`,\n`@cotal-ai/workspace`, `@cotal-ai/cli`, `@cotal-ai/manager`, `@cotal-ai/delivery`,\n`@cotal-ai/web`, `@cotal-ai/cmux`, `@cotal-ai/orca`, `@cotal-ai/tmux`, `@cotal-ai/herdr`,\n`@cotal-ai/connector-core`, `@cotal-ai/connector-claude-code`, `@cotal-ai/connector-hermes`,\n`@cotal-ai/connector-opencode`, `@cotal-ai/connector-codex`, `@cotal-ai/pi`, `@cotal-ai/auth`:\n\n1. Go to `https://www.npmjs.com/package/<name>/access` (e.g.\n `https://www.npmjs.com/package/@cotal-ai/core/access`).\n2. Scroll to **Trusted publishing** > **Add a trusted publisher**.\n3. Pick **GitHub Actions**.\n4. Fill in:\n - **Organization or user:** the GitHub owner (your org or user).\n - **Repository:** `Cotal`.\n - **Workflow filename:** `changesets.yml`.\n - **Environment name:** `npm-publish`.\n5. Save. Repeat for every package.\n\n> The first time, you may need to publish a version manually (with a classic token) so the\n> package exists on npm. After that, OIDC takes over.\n\n> **Migration from blank Environment:** if packages were previously configured with a blank\n> Environment name, each must be updated to `npm-publish`. Delete the old trusted publisher\n> record and re-create it with the Environment name filled in. The preflight will refuse any\n> package whose trusted-publisher record does not carry the `npm-publish` environment.\n\n## Day-to-day flow\n\n1. Open a PR that changes code in a publishable package.\n2. Add a changeset describing the change:\n\n ```bash\n pnpm changeset\n ```\n\n Pick the affected packages plus the semver bump (patch / minor / major), and write a\n one-line summary. Commit the generated `.changeset/<name>.md` file alongside your code\n change.\n3. Merge to `main`.\n4. The `Changesets` workflow runs:\n - If there are pending changesets, it opens (or updates) a PR titled `chore(release):\n version packages` that bumps versions and updates `CHANGELOG.md` files.\n - When **that** PR is merged, the same workflow detects the bumped versions, runs `pnpm\n build`, and `pnpm publish`es each changed package to npm with provenance.\n\n## Publication workflow\n\n`ci:publish` in the root `package.json` is:\n\n- an exact-version census of every package in the Changesets fixed group against the registry;\n- a check that the public recursive workspace set is the complete Changesets fixed group;\n- one GitHub OIDC exchange per package when the release job exposes the OIDC requester;\n- a GET of each package's trusted-publisher document with that exchanged token, which must list\n a direct `npm publish` Allowed action on THIS repository's `changesets.yml` publisher;\n- only after those checks, the workspace build, native assembly, and recursive publish.\n\nThe census prints every package, version, OIDC result, and direct-publish result before it refuses.\nIf every exact version already exists, the preflight reports a no-op and exits successfully before\ncredential checks. A mixed census, incomplete fixed group, failed OIDC exchange, or stage-only\npackage exits before `pnpm publish`.\n\nThe publish job refuses an npm access token in its environment and publishes through OIDC only.\nThis prevents pnpm from falling back to a classic token when an OIDC exchange fails.\n\nHTTP 201 from the OIDC exchange is identity only. npm's trusted-publisher Allowed actions always\npermit `npm stage publish`; configurations created after 2026-09-03 default to stage and may omit\ndirect `npm publish`. Both paths use the same successful exchange, so the preflight never treats\nthat 201 as proof that sequential `pnpm publish -r` can write. Binding those Allowed actions to a\nGitHub Environment is done: the `version` and `snapshot` jobs reference `environment:\nnpm-publish`, and the OIDC identity assertion rejects tokens without the matching\nenvironment claim.\n\npnpm's `--batch` option was evaluated. It exists from pnpm 11.7 and is all-or-nothing only on a\nregistry implementing `PUT /-/pnpm/v1/publish` (pnpr does). npm's registry returns 404 for read-only\n`GET` and `OPTIONS` probes of that endpoint, and its published Registry API does not document it.\npnpm batch publishing also rejects provenance and requires one shared credential for the batch\ninstead of the per-package OIDC exchanges used here. The repository stays on the normal npm publish\nprotocol and treats the preflight as the fail-before-first-write control.\n\n```bash\nnode scripts/preflight-npm-publish.mjs && pnpm build && node scripts/seat-assemble-natives.mjs && pnpm publish -r --provenance --access=public --no-git-checks\n```\n\n- `preflight-npm-publish.mjs`: derive and print the full fixed-group package/version census. In\n GitHub Actions it exchanges a package-specific OIDC token, then GETs `/-/package/<name>/trust`\n and refuses unless THIS repository's `changesets.yml` publisher lists a direct-publish Allowed\n action. npm documents that identity on GET `/-/package/<name>/trust` as `claims.repository` and\n `claims.workflow_ref.file` with a `permissions` array. Other GitHub publishers on the same package\n are not proof that this job can publish. It refuses npm access-token environment variables before\n the census or OIDC exchange, so `ci:publish` cannot be used with a classic token.\n- `pnpm build`: build every workspace package first, supplying local workspace dependency outputs\n when a partial retry publishes only the packages still missing.\n- `seat-assemble-natives.mjs`: assemble the downloaded native seat artifacts before publication.\n- Seat's pack and publish hooks assert both native artifacts and compile its JavaScript and type\n entrypoints without rebuilding the native helpers.\n- `-r`: recursively publish all workspace packages.\n- `--provenance`: emit SLSA provenance attestations (a no-op without OIDC, automatic with it).\n- `--access=public`: required for scoped packages on first publish.\n- `--no-git-checks`: skip pnpm's branch / clean-tree guard, since CI does not need it.\n"
213
+ "body": "# Publishing a release\n\n> **Project** (non-normative maintainer notes) · **For:** maintainers shipping Cotal\n\nCotal uses [Changesets](https://github.com/changesets/changesets) to version and publish the\nworkspace packages under `packages/*`, `extensions/*`, and `implementations/*` to npm.\n`examples/**` is ignored, since it is not published.\n\n## 0.11 runtime migration\n\nThe published binary no longer bundles the optional tmux and cmux runtimes. Existing operators\nmust run `cotal ext add @cotal-ai/tmux` or `cotal ext add @cotal-ai/cmux` once after upgrading,\nbefore using `runtime: tmux|cmux` in a manifest or passing `--runtime tmux|cmux`. Missing runtimes\nfail loudly with the matching install command; they never fall back to pty.\n\n## Trusted publishing\n\nTrusted publishing replaces the long-lived `NPM_TOKEN` secret with short-lived OIDC tokens\nissued by GitHub Actions. Each published package must be configured once on npmjs.com.\n\nThe `fixed` group in [`.changeset/config.json`](../.changeset/config.json) is the list that\ngets versioned and published. Derive the package list from it instead of\nmaintaining it by hand. It had drifted by six packages before this was last reconciled.\n\n### Deployment Environment setup\n\nBoth publishing jobs (`version` and `snapshot`) reference a GitHub Environment named\n`npm-publish`. The OIDC assertion rejects tokens that do not carry this environment claim,\nso the Environment must exist and must protect the release ref before the first publish.\n\n1. Go to **Settings > Environments** in the repository.\n2. Create a new Environment named **`npm-publish`**.\n3. Under **Deployment branches and tags**, select **Selected branches and tags** and add\n `main` as the only allowed branch. This restricts OIDC token issuance to runs on `main`.\n4. Optionally add required reviewers if your team wants a manual gate before each release.\n\nThe Environment name is embedded in the workflow, in the OIDC identity assertion, and in\nevery npm trusted-publisher record. All three must use the exact string `npm-publish`.\n\n> **Snapshot releases:** the snapshot job is also bound to the `npm-publish` Environment.\n> If the deployment branch policy allows only `main`, snapshot releases from other branches\n> are refused by the Environment gate before the OIDC exchange. To allow snapshots from\n> additional branches, add those branches to the Environment's deployment policy.\n\n### Per-package trusted publisher configuration\n\nFor **every** published package, `cotal-ai` (the binary), `@cotal-ai/core`,\n`@cotal-ai/workspace`, `@cotal-ai/cli`, `@cotal-ai/manager`, `@cotal-ai/delivery`,\n`@cotal-ai/web`, `@cotal-ai/cmux`, `@cotal-ai/orca`, `@cotal-ai/tmux`, `@cotal-ai/herdr`,\n`@cotal-ai/connector-core`, `@cotal-ai/connector-claude-code`, `@cotal-ai/connector-hermes`,\n`@cotal-ai/connector-opencode`, `@cotal-ai/connector-codex`, `@cotal-ai/pi`, `@cotal-ai/auth`:\n\n1. Go to `https://www.npmjs.com/package/<name>/access` (e.g.\n `https://www.npmjs.com/package/@cotal-ai/core/access`).\n2. Scroll to **Trusted publishing** > **Add a trusted publisher**.\n3. Pick **GitHub Actions**.\n4. Fill in:\n - **Organization or user:** the GitHub owner (your org or user).\n - **Repository:** `Cotal`.\n - **Workflow filename:** `changesets.yml`.\n - **Environment name:** `npm-publish`.\n5. Save. Repeat for every package.\n\n> The first time, you may need to publish a version manually (with a classic token) so the\n> package exists on npm. After that, OIDC takes over.\n\n> **Migration from blank Environment:** if packages were previously configured with a blank\n> Environment name, each must be updated to `npm-publish`. Delete the old trusted publisher\n> record and re-create it with the Environment name filled in. The preflight will refuse any\n> package whose trusted-publisher record does not carry the `npm-publish` environment.\n\n## Day-to-day flow\n\n1. Open a PR that changes code in a publishable package.\n2. Add a changeset describing the change:\n\n ```bash\n pnpm changeset\n ```\n\n Pick the affected packages plus the semver bump (patch / minor / major), and write a\n one-line summary. Commit the generated `.changeset/<name>.md` file alongside your code\n change.\n3. Merge to `main`.\n4. The `Changesets` workflow runs:\n - If there are pending changesets, it opens (or updates) a PR titled `chore(release):\n version packages` that bumps versions and updates `CHANGELOG.md` files.\n - When **that** PR is merged, the same workflow detects the bumped versions, runs `pnpm\n build`, and `pnpm publish`es each changed package to npm with provenance.\n\n## Publication workflow\n\n`ci:publish` in the root `package.json` is:\n\n- an exact-version census of every package in the Changesets fixed group against the registry;\n- a check that the public recursive workspace set is the complete Changesets fixed group;\n- one GitHub OIDC exchange per package when the release job exposes the OIDC requester;\n- a GET of each package's trusted-publisher document with that exchanged token, which must list\n a direct `npm publish` Allowed action on THIS repository's `changesets.yml` publisher;\n- only after those checks, the workspace build, native assembly, and recursive publish.\n\nThe census prints every package, version, OIDC result, and direct-publish result before it refuses.\nIf every exact version already exists, the preflight reports a no-op and exits successfully before\ncredential checks. A mixed census, incomplete fixed group, failed OIDC exchange, or stage-only\npackage exits before `pnpm publish`.\n\nThe post-publish closure gate checks every package in the fixed group. Registry observations cannot\ndistinguish a partial publish from slow propagation: clean 404s and repeated non-404 failures both\nlack evidence that a package will never appear. The census therefore reports an incomplete or\nerrored closure as `UNSETTLED` and never fails the job on its own. Exit 1 remains reserved for future\npositive publisher evidence.\n\nRe-check a version that already shipped without publishing, tagging, or changing git:\n\n```bash\nnode scripts/verify-publish-closure.mjs 0.52.0 --recheck\n```\n\nThe publish job refuses an npm access token in its environment and publishes through OIDC only.\nThis prevents pnpm from falling back to a classic token when an OIDC exchange fails.\n\nHTTP 201 from the OIDC exchange is identity only. npm's trusted-publisher Allowed actions always\npermit `npm stage publish`; configurations created after 2026-09-03 default to stage and may omit\ndirect `npm publish`. Both paths use the same successful exchange, so the preflight never treats\nthat 201 as proof that sequential `pnpm publish -r` can write. Binding those Allowed actions to a\nGitHub Environment is done: the `version` and `snapshot` jobs reference `environment:\nnpm-publish`, and the OIDC identity assertion rejects tokens without the matching\nenvironment claim.\n\npnpm's `--batch` option was evaluated. It exists from pnpm 11.7 and is all-or-nothing only on a\nregistry implementing `PUT /-/pnpm/v1/publish` (pnpr does). npm's registry returns 404 for read-only\n`GET` and `OPTIONS` probes of that endpoint, and its published Registry API does not document it.\npnpm batch publishing also rejects provenance and requires one shared credential for the batch\ninstead of the per-package OIDC exchanges used here. The repository stays on the normal npm publish\nprotocol and treats the preflight as the fail-before-first-write control.\n\n```bash\nnode scripts/preflight-npm-publish.mjs && pnpm build && node scripts/seat-assemble-natives.mjs && pnpm publish -r --provenance --access=public --no-git-checks\n```\n\n- `preflight-npm-publish.mjs`: derive and print the full fixed-group package/version census. In\n GitHub Actions it exchanges a package-specific OIDC token, then GETs `/-/package/<name>/trust`\n and refuses unless THIS repository's `changesets.yml` publisher lists a direct-publish Allowed\n action. npm documents that identity on GET `/-/package/<name>/trust` as `claims.repository` and\n `claims.workflow_ref.file` with a `permissions` array. Other GitHub publishers on the same package\n are not proof that this job can publish. It refuses npm access-token environment variables before\n the census or OIDC exchange, so `ci:publish` cannot be used with a classic token.\n- `pnpm build`: build every workspace package first, supplying local workspace dependency outputs\n when a partial retry publishes only the packages still missing.\n- `seat-assemble-natives.mjs`: assemble the downloaded native seat artifacts before publication.\n- Seat's pack and publish hooks assert both native artifacts and compile its JavaScript and type\n entrypoints without rebuilding the native helpers.\n- `-r`: recursively publish all workspace packages.\n- `--provenance`: emit SLSA provenance attestations (a no-op without OIDC, automatic with it).\n- `--access=public`: required for scoped packages on first publish.\n- `--no-git-checks`: skip pnpm's branch / clean-tree guard, since CI does not need it.\n"
214
214
  },
215
215
  {
216
216
  "slug": "roadmap",
@@ -224,7 +224,7 @@ export const DOCS_BUNDLE = {
224
224
  "title": "Run a mesh",
225
225
  "kind": "Guide (informative)",
226
226
  "summary": "Day-to-day operation of a local mesh: what cotal up actually runs, how spawning resolves personas, harnesses, and models, how to reach a mesh from any directory, and the operator-only maintenance v…",
227
- "body": "# Run a mesh\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nDay-to-day operation of a local mesh: what `cotal up` actually runs, how spawning\nresolves personas, harnesses, and models, how to reach a mesh from any directory, and the\noperator-only maintenance verbs. Every command's full flag set is in the\n[CLI reference](cli.md).\n\n## The stack\n\n`cotal up` brings up the whole local stack and bare `cotal down` stops it. Managed\nagents stay running as unmanaged OS processes; pass `--with-agents` to take them\nwith the stack. A current manager proves that it can detach local PTY custody before\nbare down signals it. A pre-pin legacy manager instead receives a reduced-guarantee\nwarning and is signalled according to the documented upgrade contract. Its running binary\nmay still carry the older destructive SIGTERM handler, so the CLI does not claim its\npre-signal agent inventory was spared; those agents may have been reaped.\n\n- **Broker**: a local `nats-server` (logs to `.cotal/nats.log`).\n- **Delivery daemon**: the durable backstop, auth mode only\n ([what it does](delivery-daemon.md)).\n- **Manager**: a detached supervisor answering the control plane, so\n `cotal spawn --detach` and the `cotal_spawn` tool work right after `up`.\n\nThree modes:\n\n- **Default (static auth).** JWT-authed, on by default: sender authenticity and per-agent\n ACLs, enforced by the broker ([how](identity-and-auth.md)).\n- **`--user-auth --idp <url>`.** Per-user auth: people `cotal login` once, the operator\n grants their agents on the actor ledger, and every connect is authorized live against\n that grant. Starts the space's auth service alongside the broker\n ([how](identity-and-auth.md)).\n- **`--open`.** An unauthenticated, live-only dev mesh (no auth, no delivery daemon). For\n quick local experiments.\n\nThe broker and local services bind **loopback** by default. `--host 0.0.0.0` widens the broker\nbind independently of the auth mode, so \"network-reachable\" never silently means\n\"unauthenticated\". With no explicit `--server`, `cotal up` auto-selects a free local port when\nthe default address is already held by another project; an explicit `--server` fails loud on\ncollision.\n\n`--host` is a boot flag, not a live rebind. A fresh `cotal up` writes the generated\n`.cotal/auth/server.conf` (project-local, not `~/.cotal`) with that bind and starts nats against\nit. If anything is already answering at the mesh URL, `up` refreshes the recorded mesh and\nleaves the running nats listener alone, so passing `--host 0.0.0.0` on a live or orphaned\nbroker does not change who can connect. To change the bind: `cotal down`, then `cotal up --host\n<addr>` against a stopped broker so the generated file is rewritten. Do not edit `server.conf`\nby hand; the next real boot overwrites it.\n\nOn a stopped shared broker, `up` renders every persisted space account and every enabled\nspace's auth-callout account into the resolver preload, regardless of which space starts\nthe broker. A missing callout account for an enabled space stops the boot rather than\nstarting with a reduced resolver. An already-running broker is refreshed without rewriting\nits config.\n\nA broker-only host is a first-class `up` mode. `cotal up --no-manager` boots the broker and, in\nauth mode, the delivery daemon, and no local manager, so the broker host never has a manager to\nstop and never leaves a manager slot stale. A refresh under the flag of a mesh whose manager is\nlive refuses rather than keeping or stopping it: `cotal down manager` first. Without the flag,\nauth-mode `up` still starts nats, the delivery daemon, and a\nlocal manager. A space may run more than one manager, addressed by instance id\n([control surface](control-surface.md#instance-routing)); putting no manager on the broker host\nis a topology choice, not a singleton invariant. A manager whose boot inventory has no\navailable connector does not take unpinned `spawn`/`launch` on the class rail, so a sibling\nthat can launch the harness can. `describe` still rides the class rail, so an unpinned spawn\ncan bind-fence when that skip member answered describe; re-issue, or pin `--on`. Pin one\ninstance with `--on` when a partial inventory still answers with a harness refusal. The\nsupported split is:\n\n```bash\n# broker host (project root that owns the generated conf, pidfiles, and logs)\ncotal up --detach --host 0.0.0.0 --space main --no-manager\n# no local manager starts: the summary lists nats-server + delivery daemon, and there is no\n# `.cotal/manager.<spaceKey>.log` to wait for on this host\n\n# manager host (registered remote mesh, same space)\ncotal meshes add --server nats://broker.example:4222 --root ~/meshes/main\ncotal supervise --space main --server nats://broker.example:4222\n```\n\nWait for `✓ manager up` in `.cotal/manager.<spaceKey>.log` on the manager host before spawning\nagents. On a broker host started without `--no-manager`, `cotal up --detach` prints `✓ running in\nthe background:` with `manager` listed once the manager pidfile is live; stop that local manager\nonly after the `✓ manager up` line. A host started WITH `--no-manager` never runs one, so neither\nthe wait nor the stop applies there. That detach stdout is not a safe teardown boundary: it is\npidfile liveness, not `✓ manager up`. `✓ manager up` is supervise's post-start line after\n`await mgr.start()`. `cotal down manager` after only the detach line can still default-terminate\nthe child during registration after it has taken the governance slot. Stopping before that\npost-start log line can leave the endpoint governance slot held until the holder's gate\nreopens past the stamp (the successor's boot heal, or\n[`cotal reconcile-gate`](cli.md#reconcile-gate) when that boot cannot run). See\n[Gate recovery](#gate-recovery).\n\nStandalone `cotal deliver --creds` is not a repair for that split. Production renewal needs\nthe manager and the daemon to address one credential store. Separate host filesystems still\nleave manager root A writing and the daemon reloading root B; that composition is refused\nwhile the daemon stays up. Keep delivery on the broker host under `up`, and share one store\nonly when you are composing a hosted pair ([embedding](embedding.md#supervisor-signing-authority)).\n\n### Split host bind\n\nA remote manager cannot reach a loopback broker. After changing `--host`, confirm the\ngenerated `host:` in `.cotal/auth/server.conf` and that nats is listening on that address\nbefore registering the mesh on the manager host. Detached child logs stay under the **project**\n`.cotal/` that `up` ran in (see [When something looks absent](#when-something-looks-absent));\nthey are not `~/.cotal` unless that directory is the mesh root.\n\nA user-auth mesh can expose only its credential exchange through an operator-owned HTTPS reverse\nproxy while leaving the existing local exchange untouched:\n\n```bash\ncotal up --user-auth --idp https://idp.example/api/auth \\\n --exchange-public-port 7443 \\\n --exchange-public-url https://auth.example\n```\n\nThe public listener itself still binds `127.0.0.1:7443`; configure the proxy to terminate TLS and\nforward to it. It serves only `/health`, `/jwks`, `/exchange`, and `/.well-known/cotal-mesh` with\nthe documented methods. It needs no local file capability: the signed IdP JWT or managed-agent\nactor token is the proof, while the original loopback listener remains capability-gated. Add\n`--exchange-trusted-proxy` only when that listener is reachable exclusively through your trusted\nproxy; it keys failure throttling by the last `X-Forwarded-For` hop instead of the socket address.\nThe well-known bundle includes IdP pins and a deny-all sentinel credential, so fetch it only from\nthe configured HTTPS origin. To change these listener flags, stop and restart the mesh; a refresh\nof an already-running service does not replace its bind or proxy policy. See\n[Identity & auth](identity-and-auth.md#per-user-authentication) for the trust boundary.\n\n### Remote supervised seats by enrollment\n\nA remote seat does not need to run `cotal login` when the mesh owner pre-mints a single-use\nenrollment for it. Mount the enrollment URL as a private file, place the seat persona on the remote\nmachine, and launch the foreground seat:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./worker.md --space main\n```\n\nThe URL is redeemed once with an unauthenticated GET. Redirects, off-machine plain HTTP, retries,\nand login fallback are refused. If the seat has no mesh record yet, the enrollment response's stock\nuser-bundle fields register it before the launch. The returned actor token then uses the same remote\nauth-service exchange as a login-provisioned agent. The enrollment URL and file path do not enter the\npreflight or harness environment. A failed or reused enrollment leaves no actor material on disk; ask the owner\nfor a fresh enrollment. When the foreground seat exits, this machine's credential files are removed\nand the mesh-side grant stays until the mesh operator revokes it; the launch line says so. The exact\nserver contract is in\n[Enrollment redeem](identity-and-auth.md#enrollment-redeem).\n\n`cotal status` prints the detailed setup, process, registry, and live mesh status. Its Machine\nsection names the running CLI's source checkout, installed package root, or npx package root beside\nthe version. A stale Claude skills row names the installed and CLI versions it compared. `cotal\nsetup` (after the first run) prints the compact card.\n\nBefore reporting ready, the manager resolves every installed connector's declared harness\nbinaries against its own environment. A missing binary does not stop unrelated manager work: boot\ncontinues, but prints a named `connector <name> unavailable` line and records that reason in the\nmanager's `status` response. Available connector rows record the absolute paths boot resolved.\nSpawn keeps the same pre-mint check as a backstop for connectors registered after boot.\n\nOn an authenticated manager start, unfinished static lifecycle rows reconcile while the control\nendpoint is already serving. The manager `status` response reports\nthe `staticReconciliation` state, the last sweep counts, and each failed alias with its durable\nphase and literal disposition. `cotal status --components` reports the state and per-alias failure\ndetails. A failed exact terminal is retried in the same process after 1, 5,\nand 30 seconds. Each attempt re-reads the durable slot and re-enters the same deterministic terminal\noperation; the delays only schedule work and never release the lifecycle fence.\n\nOn shutdown, the manager fences new reconciliation work and waits for an exact terminal that already\nstarted. The current serial sweep stops before its next alias, and startup cannot publish the manager\nservice after `stop()` completes.\n\nThe four-attempt budget is per manager process. An exhausted row stays held and reports\n`retry-exhausted` with the remedy to restart the manager. The next process derives a fresh budget\nfrom the still-authoritative durable row. A `recovered` row remains visible until the next static\nreconciliation sweep, then clears. This component reports reconciliation outcomes. It does not say\nwhether footprint cleanup completed independently of the terminal result; that separate durable\nprojection remains tracked by #1274.\n\n`cotal service install` is the supported way to run the manager as a user service\n([CLI reference](cli.md#service)): a systemd user unit on Linux, a launchd agent on macOS, one\nper mesh, surviving logout and reboot when lingering is enabled with `--linger`. It installs only\nthe manager; the units below remain the process models for every other component, and they are\nstill **examples of process models** for those: copy them only after you decide which processes\nthe unit should own.\n\n### Supervising the detached stack\n\n`cotal up --detach` is a launcher: it starts the broker, delivery daemon, and manager, reports what\nstarted, then exits. Do not wrap it in a systemd service with `Type=oneshot` and\n`RemainAfterExit=yes` and treat `systemctl is-active` as stack health. That unit becomes `active\n(exited)` when the launcher exits successfully and stays active even if every detached process dies.\nWhen `up --detach` can identify that exact unit shape, it prints a warning but keeps the requested\nstartup behavior.\n\nFor a single-host stack, keep `cotal up` itself in the foreground so systemd tracks a long-running\nprocess and restarts the stack if that process fails:\n\n```ini\n[Service]\nType=simple\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal up --space main --host 0.0.0.0\nRestart=on-failure\nRestartSec=5s\n```\n\nAn active unit then proves the foreground launcher and broker are still running, but it still does\nnot prove that every child component serves. Pair it with the component check below. Also remember\nthat `cotal up` starts a local manager as well as the broker and delivery daemon; run\n`cotal up --no-manager` (add the flag to the unit's `ExecStart` too) on a host intended to be\nbroker-only, so the unit and the host agree.\n\nThat `Type=simple` shape puts nats in the unit's cgroup with the foreground `up` process. A\n`Restart=always` (or `on-failure`) of **this** unit therefore restarts nats as well, so remote\nmanagers drop for the time it takes the broker to come back. Wrapping `cotal up --detach` in\n`Type=oneshot` with `RemainAfterExit=yes` does not move nats out of that cgroup. Detached\nspawn starts a new process group, not a new systemd cgroup, and the default\n`KillMode=control-group` still signals every process left in the service cgroup on stop or\nrestart, including the nats PID. Escaping that cgroup needs an explicit unit setting such as\n`KillMode=process`, or a separate nats unit; this CLI does not ship that escape. The\n`Type=oneshot` unit below is a `cotal status --components` liveness check, not a\n`--detach` launcher. Neither trade is universal from\n`Type=simple` alone; it follows from which processes the unit actually owns. `cotal service\ninstall` covers only the manager, so for the broker and its siblings pick the example that\nmatches the ownership you want, and treat\n`systemctl is-active` as unit health, not mesh health.\n\nIf the deployment deliberately uses `cotal up --detach` as a boot action, monitor observed state\ninstead of the launcher's exit:\n\n```ini\n[Unit]\nDescription=Check Cotal component liveness\n\n[Service]\nType=oneshot\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal status --components --space main\n```\n\nRun that check from a systemd timer or another monitor and alert on a nonzero exit. The command\ndistinguishes `absent`, `not-serving`, and `refused` components and never treats a sibling's health as\nproof. Its delivery-process check is local to the broker host, so run it there. On a split topology,\nalso probe the broker URL from the manager host and monitor the manager's own service there. A remote\nmanager cannot observe the broker host's delivery PID, and an `active` unit on either host says\nnothing about the other host.\n\nStop one part without tearing down the mesh by naming its registered component: `cotal down\nmanager`, `cotal down delivery`, or `cotal down web`. Component names from installed extensions\njoin the same surface; `cotal down` with no names retains whole-stack behavior and\nleaves managed agents running as unmanaged OS processes. `cotal down --with-agents`\nis the previous reap.\n\n## Remote supervised agents\n\nOn a remote user-auth mesh, foreground `cotal spawn` remains the default participant path. A\nparticipant can run detached agents only after the host advertises and operates the remote manager\nauthority service, and the participant's actor-ledger row includes `supervise`. This is not implied\nby `spawn` or `admin`.\n\nThe participant's loopback/operator exchange obtains one closed `manager-service` view for its\nordinary derived owner, a fixed server-selected manager actor, and one opaque manager instance.\nThe host, not the participant, issues the public-nkey JWT material via the replay-safe,\nlifecycle-bound prepare → activate → renew exchange, plus a one-shot target-pinned retirement\nrequest for a host-managed terminal. It never exports the space signer, a static\nprovisioner credential, or generic storage authority. The manager may provision only descendants\nof that same owner, with host validation at each provision.\n\nThe registry entry decides the broker URL `supervise` dials, so a mesh published over `wss://` is\ndialed as a websocket. The manager-authority registration it runs first also takes its TLS\nrequirement from that entry, so the prepare credential is not exchanged over a plaintext\nconnection the record did not describe. `cotal meshes add` records both.\n\nWhen the authority service, login, or renewal is unavailable, the remote manager degrades\nfail-closed: it refuses new agents, restarts, and credential replacement rather than pretending\nlocal authority exists. Existing agents remain live only while their independent credentials are\nvalid. A hosted composition must revoke the managed grant and finish its resumable release before it\nrequests terminal retirement. Deleting DM or delivery consumers is not retirement and must not reset\na resumable lifecycle's frontier or pending state. The alias remains held until the terminal barrier\nconfirms. Restore service and renew successfully before asking it to recover an agent. See\n[Identity & auth](identity-and-auth.md#remote-manager-authority) and the [CLI\nreference](cli.md#supervise).\n\n## Spawning agents\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn reviewer --detach # supervised: the manager runs it in a PTY\ncotal attach --name reviewer # watch/type into a detached agent (Ctrl-] detaches)\ncotal ps # what the manager is running\ncotal stop --name reviewer # stop one\n```\n\nHow a spawn resolves:\n\n- **Persona.** A bare `cotal spawn` uses `.cotal/agents/default.md`; a positional name\n picks `.cotal/agents/<name>.md`; `--config` takes an explicit ref or path. Set\n `COTAL_DEFAULT_PERSONA=<name-or-path>` to change the fallback. Fields and format:\n [agent files](agent-files.md).\n- **Harness.** Resolution order is an explicit `--agent` or `cotal_spawn` `agent` argument,\n then the persona file's `agent:` pin, then the invoking caller's `COTAL_DEFAULT_AGENT`,\n then the manager's `COTAL_DEFAULT_AGENT`, then the product default (Claude). Compared in\n [Connectors](connectors.md); per-connector guides:\n [Claude](connect-claude.md) · [OpenCode](connect-opencode.md) ·\n [Hermes](connect-hermes.md) · [pi](connect-pi.md).\n- **Model.** `--model` overrides the persona file's `model:` (Claude: `opus` / `sonnet` or\n a full id; OpenCode: `provider/model`). Connectors that expose a catalog report it via\n `cotal models --agent opencode`: model ids plus available variants; pick one with\n `--model provider/model --variant high`.\n- **Tools.** A spawned agent gets only the cotal tools by default; share your own MCP\n servers deliberately with `--share-tools` ([config](config.md)).\n- **Launch options.** `--opt key=value` (repeatable) passes a native harness flag straight\n through; a persona or manifest `launchOptions:` mapping does the same declaratively (a\n `--opt` wins per key). It is a **raw passthrough**, with no allow/deny list: Claude renders\n each as `--key value` (a bare `--key` for an empty value), OpenCode merges them into its\n agent config, and Hermes has no option surface so it fails loud. The trust boundary is the\n `spawn` capability itself, not the flag set, so granting `spawn` is host-launch authority\n ([security](security.md)). A key must be a plain flag name; malformed or prototype-polluting\n keys are refused.\n\nDetach from an attached PTY with **Ctrl-]** (the agent keeps running); rebind it with\n`COTAL_DETACH_KEY=ctrl-<char>` when it clashes with a keybinding inside the agent's TUI.\n\n**Runtimes.** The manager spawns into a **pty** by default. On Linux a detached per-seat\ncustodian owns that PTY, so replacing the manager worker does not close the seat. A custodian\nwhose agent has exited exits a few seconds later on its own, and a manager that boots after a\ncrash reaps the seats its predecessor left running. Other\nplatforms still spawn the PTY in-process; `adopt` throws until their transport lands. Optional runtimes are installed\nthrough the extension surface, for example `cotal ext add @cotal-ai/orca`, then selected with\n`--runtime orca` (similarly `@cotal-ai/tmux`, `@cotal-ai/cmux`, and `@cotal-ai/herdr`). They put teammates in native\nterminal surfaces rather than manager-owned PTYs. Runtime names are open-ended and resolved from\nthe registry; a missing provider or app throws, never silently falls back\n([architecture](architecture.md)).\n\n## Mesh registry\n\n`cotal up` records each running mesh in a machine-local registry\n(`~/.cotal/meshes/space.<key>.json`, named by a case-safe hex encoding of the space: broker URL, the project root holding its creds and\npersonas, and its mode). So a bare `cotal spawn <persona>` from *any* directory joins the\nrunning mesh with the right credentials instead of mistaking the cwd for a space:\n\n- `cotal use <name>` sets the default from every directory, including inside another mesh's\n project. `--space <name>` overrides it for one command.\n- When one broker has records for several spaces, `cotal up --space <name>` refreshes that named\n space.\n- With no live selected default, a project with its own `.cotal/` resolves to that project's\n mesh; otherwise one running mesh is used automatically and several are an error.\n- `cotal meshes` lists them (a `*` marks the default); `cotal down` removes the entry.\n\nThe registry stores a *path*, never a secret; trust material stays in each project's\n`.cotal/auth`. If the mesh is down or won't take your creds, spawn fails with one\nsentence, never a raw NATS trace.\n\n### Meshes you did not start here\n\nA mesh running on another machine has no `cotal up` on this one, so register it by hand:\n\n```bash\ncotal meshes add # guided: asks for the broker, probes it, offers what it finds\ncotal meshes add optiplex --server nats://100.90.12.34:4222 --root ~/meshes/optiplex \\\n --allow-unencrypted-overlay # see below: an overlay address needs this\ncotal meshes rm optiplex\n```\n\nOn a terminal, a bare `cotal meshes add` walks you through it: it probes the broker you name and\nreports whether it is open or requires credentials, offers the spaces the folder already holds\ncredentials for, and shows the record before writing it. Scripts and agents keep the flag form -\nwithout a terminal nothing prompts.\n\n`--root` is the local folder holding that mesh's `.cotal/auth` and `.cotal/agents` (its personas);\nthe mode is inferred from what that folder holds.\n\n**Know what you are copying.** For an authenticated mesh that folder carries the space's account\n**signing seed**, which is the authority to mint any identity in the space. A machine holding it\nis a certificate authority for the mesh rather than a client of it: anyone who reads it can\nimpersonate any agent, read every retained channel and DM, change ACLs, and keep issuing\nthemselves credentials. There is no per-machine revocation; undoing it means rotating the signing\nkey and re-minting every credential in the space. Copy it only to machines you would trust with\nthe whole mesh. `cotal mint` on its own does not substitute here: registering an `auth` mesh needs\nsigning material that composes, which a minted user credential is not. The\nbroker is probed before the record is written, so a bad address or a credential that mesh will not\naccept fails at registration rather than at your first `spawn` (`--force` records it without verifying,\nuseful when the mesh is simply down right now).\n\n#### Which addresses you may register\n\nRegistering a mesh is how this machine starts sending agent credentials to a broker it does not\nrun. NATS announces itself in plaintext before anyone authenticates, so an attacker on the path\ncan pose as the broker and read the credential out of the connect unless the connection\n**requires TLS**, which is recorded on the entry and enforced on every dial through it.\n\nWhat the record will require decides what you may register:\n\n- **Without required TLS**, the address is the gate: **loopback** (`127.0.0.0/8`, `::1`), or\n **your private overlay** (`100.64.0.0/10`, `fd7a:115c:a1e0::/48`) with\n `--allow-unencrypted-overlay`. The tunnel provides the protection, and this command cannot check\n its state. Hostnames are refused because the lookup would choose which machine receives your\n credentials.\n- **With required TLS**, set `--tls` or use a `tls://` URL. The recorded scheme enforces the TLS\n requirement. A **hostname or public address** is accepted because the certificate chain and\n hostname check identify the peer. A registration whose broker cannot complete the handshake\n fails unless you pass `--force`, which records the entry without verification.\n\nOrdinary private ranges like `10.x` and `192.168.x` are refused in **both** modes. A café's wifi\nis private but does not belong to you, and no public CA issues certificates for those ranges. An\naddress spelling changes nothing: `[::ffff:192.168.1.10]`, `3232235786`, `0300.0250.01.012`, and\n`192.168.257` all resolve to private addresses and receive the same refusal as the dotted form.\n`--force` exists for a mesh that is down. It never permits an unsafe credential destination.\n\n#### Registering a hosted user-auth mesh\n\nA user-auth space's IdP pins are established where the mesh runs and are never guessed. Register\none from **supplied** trust: `--user-auth-file bundle.json` (exported on the mesh's machine), or\n`--from https://…/.well-known/cotal-mesh`, which asks before it contacts the address at all,\nfetches the discovery document over HTTPS, shows you the pins, and asks again before adopting\nthem. Redirects are refused because a 302 can walk a pinned fetch down to\nplaintext or onto another host, and the pinned exchange must be an `https://` URL too. The one\nexception is an exchange on **this machine**, where nothing leaves the box: plain `http://` is\naccepted for a loopback *literal* (`127.0.0.1`, `::1`, and any spelling of them), but **not** for\n`localhost`, which a hosts entry or poisoned lookup could point elsewhere. Use the\nliteral. Registration checks that the pinned exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also checks that the broker refuses a\nbare connect; that refusal is the pass. The bundle's sentinel credentials are written to a private (0600) file\nunder the entry's root; the registry itself never carries the secret.\n\n**Without required TLS**, an overlay address is **refused unless you accept the dependency\nexplicitly**, with `--allow-unencrypted-overlay`. The address is not the guarantee: it is protected\nwhile the tunnel is up, and if the tunnel is down that range is ordinary carrier-grade NAT and\nwhoever answers the dial receives your credentials. Only you can know which it is, so the command\nasks you to say so. Your acceptance is recorded on the mesh entry rather than printed and\nforgotten, and the guided form asks the same question instead of taking the flag.\n\n**With required TLS** (`--tls`, or a `tls://` URL) that consent is no longer asked for, and the\nflag is not needed: the handshake is what protects the connection, so the acceptance it stood in\nfor has been replaced by proof rather than promise. `cotal meshes add <space> --server\nnats://100.64.0.1 --tls` registers an overlay address with no prompt, no flag and no recorded\nacceptance. This is the \"the flag disappears once the broker can be served over TLS\" case, and it\nhas now arrived.\n\nThis gate is on **registration**. `cotal join --creds --server <url>` deliberately takes an\nexplicit connection at face value and does not consult the registry, so it is not covered. Join\nthat way only to an address you would have registered.\n\nRecords added this way are removed only by something that names them. A failed liveness probe\ndoes not delete any record: an unreachable broker, local or registered by hand, is shown as\n`offline` in `cotal meshes`. A bare command does not count that offline record as running;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the\nproject they are tearing down, and they leave a hand-registered one alone even when `--root`\npointed at that project. A `cotal up` for that space refuses outright (naming `cotal meshes rm`) unless it is\nthat same endpoint: finding a broker already answering there is a refresh that starts nothing and\nleaves the record's provenance alone, while actually starting the broker for that space, server and\nroot makes this machine the one running it, so the record becomes an ordinary local one that\n`cotal down` clears. `cotal meshes rm` drops it and re-registering with `--force` replaces it. `rm`\nonly forgets a mesh. To stop one running here, use `cotal down`.\n\n## Watching\n\n`cotal console` is the terminal view (TUI on a real terminal, plain line stream when\npiped); `cotal web` is the browser dashboard. Both are read-only observers; the\nwalkthrough is [Watch a mesh](watch-a-mesh.md).\n\n## History\n\nRetained history is operator-owned. `cotal clean history --force` purges a space's\nretained channel history; `--dms` also purges DMs (`cotal history clear` is an alias).\nIt is deliberately **not** an agent tool: agents cannot wipe the record\n([identity & auth](identity-and-auth.md)). For a **stopped** mesh, `cotal clean store\n--force` deletes the on-disk JetStream store outright, and `cotal clean all --force`\nalso resets the space identity ([CLI reference](cli.md#clean)).\n\n## Offline backup\n\nFor a coherent durable cut, preserve the whole stack first, then create the artifact while it stays\ndown:\n\n```bash\ncotal down --preserve-state\ncotal backup create ./space-backup # full by default\n# later: deliberately resume the unchanged source\ncotal up --detach\n# or, from another preserved cut, restore before the normal listener opens\ncotal up --restore ./space-backup --detach\n```\n\nUse `--store-dir` on both preservation and backup for a custom JetStream store. A store cap set\nwith `cotal up --max-file-store <bytes>` travels with the preserved state, and the resume renders it\nagain. nats-server reads the cap once at start and refuses a config reload that changes it, so a new\ncap always needs a restart. `registry` is the\nonly partial selection (`backup create ... --only registry`; `up --restore ... --restore-only\nregistry`). Backup never stops or restarts a mesh implicitly, never opens the original store, and\ndoes not contain credentials or trust secrets. Backup/restore in every auth mode, open included,\nuses isolated, operation-specific maintenance logins; normal agent credentials cannot enter that\nlistener. Full\nrestore requires the same space and exact current local trust continuity, recreates conservative\nconsumer checkpoints bound to their snapshot stream sequence state, and resumes retained agents under\ntheir original principals. The trust commitment includes the cryptographically validated full\noperator/system/data-account root chain as well as static/user authority state. A registry-only\nrestore completes canonical empty infrastructure but leaves retained agents stopped because their\nDM/DLV/TASK/ACL state is outside that selection. Authenticated restore validates the complete space\ntrust bundle before staging or changing the preserved store. Interrupted ordinary resume retries the\nsame durable attempt after its prior listener is stopped. Restore re-entry can recover a surviving normal listener\nonly when its attempt nonce, NATS server name, process owner, endpoint, and target-store identity all\nmatch the fsynced proof. A provably dead uncommitted owner is retired under lock and replaced with a\nfresh attempt-bound listener; an occupied foreign listener or ambiguous owner is never adopted. The\nmanager commit validates while retained cleanup is still suppressed; the CLI durably records its\nattempt-bound 64-hex token in `manager-committed` / `resume-committed` before `finalizeResume` can\nrelease suppression. A retry from either committed state goes straight to exact-token finalization;\nfailure preserves the committed gate and retained cleanup suppression. Missing commit evidence,\ninterrupted finalization, a live recorded endpoint despite missing pidfiles, or ambiguous proof fails closed. See the [CLI\nbackup and restore contract](cli.md#backups) for artifact, checkpoint, fallback,\ndisaster-consent, and degraded-recovery details.\n\n## Personas from the CLI\n\n`cotal personas` manages the local catalog offline: `list` (`--running` overlays live\nmarkers), `show <name>`, `edit <name>` (re-validates on save), `new <name>`, `rm <name>\n--force`. The runtime write is `cotal_persona`; the runtime read is `cotal_personas`\n(list / show), both over the wire with the manager's ownership checks. Fields: [agent files](agent-files.md).\n\n## Gate recovery\n\nA manager that dies mid-registration leaves its issuance gate *frozen* under that registration\nop. The freeze is correct: it stops two incarnations serving at once. The successor now completes\nthat dead op on boot, using the same guard as [`cotal reconcile-gate`](cli.md#reconcile-gate): it\nacts only when the freeze-holder is affirmatively gone under a complete CONNZ sweep (`gone` and\n`sweepComplete=true`). If the dead op's spec write committed, it finishes that same freeze\n(promote and reopen at the committed registration revision). If the spec did not advance, it\nabort-reopens the gate (generation+1, processEpoch unchanged) and continues the normal takeover.\nBoot heal and the following re-registration use separate one-shot executor windows, so a large\npredecessor family cannot spend the takeover's credential lifetime. If that later registration\nstill crosses a connection lifetime, it retries the same frozen operation with fresh authority\nand resumes verified-holder progress instead of freezing a new generation.\nA live holder, an incomplete sweep, or an unreachable delivery daemon still\nrefuses. Silence is never evidence of death, and there is no TTL. If holder verification is\ninterrupted, the frozen operation resumes from its durable, operation-and-gate-revision-bound\nprogress after liveness is checked again. A later freeze cannot reuse that progress: the cursor\nbinds the exact op, gate revision, and holder set. Use `cotal reconcile-gate` when the boot path cannot run\n(daemon down, a non-manager endpoint, or you want to lift the freeze without starting a manager). A spawn that hits the same frozen gate names that verb in the refusal\n(`blockedOp=registration`, the holding `opId`, `remedy=cotal reconcile-gate`) instead of a\nwait-timeout: the facts were always in the manager log; they now reach the spawn caller too.\n\nGive reconciliation a **quiet manager**. Suspend systemd restart policies, watchdogs, health-check\nrestart loops, and any other automation that can start or kill `cotal supervise` while boot healing or\n`cotal reconcile-gate` is running. Leave one recovery attempt in control until it finishes.\nRestarting the manager during the walk interrupts the current authority window. Durable progress makes\nthat interruption resumable, but a quiet manager is still the fastest and safest incident procedure.\n\n### Last-resort JetStream store replacement\n\nStore replacement is not normal gate recovery, is never automatic, and is destructive to mesh history.\nUse it only after the retained store cannot be reconciled and after deciding that losing its durable\ncontents is acceptable.\n\n1. Stop every actor touching the space: supervisor, watchdog, manager, delivery daemon, and broker.\n Confirm that no Cotal or NATS process still has the store open.\n2. Preserve the stopped store before changing anything. Move `.cotal/nats` aside to a dated backup and\n archive both `nats` and `auth`. Do not delete the only copy.\n3. Understand the loss: replacing the store removes JetStream message and control history and durable\n consumer state. Agent session files stored outside JetStream remain, but the mesh history they\n referenced does not.\n4. Start the broker against a new empty store, then start one manager. Wait until it reports\n serving successfully.\n5. Repopulate the mesh only after that manager is healthy. Re-enable supervisors, watchdogs, and other\n restart automation last.\n\nKeep the preserved store until the incident is reviewed and any required forensic or manual recovery is\ncomplete. Restoring it later restores the old durable state, including the fault that led to this last\nresort, so do not swap it back into a live mesh casually.\n\n## When something looks absent\n\nPermission denials are **loud, never silent**: an over-tight ACL rejects the endpoint call and\nalso shows up as a logged denial, instead of returning an empty or incomplete result that looks\nsuccessful. Check\n`.cotal/manager.<key>.log`, `.cotal/delivery.<key>.log` (one pair per space, keyed as\n[Config](config.md#project-files) describes), and `.cotal/nats.log`; `cotal status` shows\nwhat is actually running. Those files live under the **project** `.cotal/`, not `~/.cotal`,\nunless the mesh root is the home directory. `cotal up --detach` redirects delivery and manager\nstdio onto those files, so an operator-created systemd unit around that launcher does not put\nthe child logs in that unit's journal. `journalctl -u <unit>` can be empty while the crash\nreason is already in the project log. The access rules are collected in\n[Channels & permissions](channels-and-permissions.md).\n"
227
+ "body": "# Run a mesh\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nDay-to-day operation of a local mesh: what `cotal up` actually runs, how spawning\nresolves personas, harnesses, and models, how to reach a mesh from any directory, and the\noperator-only maintenance verbs. Every command's full flag set is in the\n[CLI reference](cli.md).\n\n## The stack\n\n`cotal up` brings up the whole local stack and bare `cotal down` stops it. Managed\nagents stay running as unmanaged OS processes; pass `--with-agents` to take them\nwith the stack. Ctrl-C on a foreground `up` follows the same sparing rule and\nprints the same report as bare down; when the manager cannot prove it can spare,\nCtrl-C refuses the teardown and leaves the stack running, and you end it with\n`cotal down --with-agents`. A current manager proves that it can detach local PTY custody before\nbare down signals it. A pre-pin legacy manager instead receives a reduced-guarantee\nwarning and is signalled according to the documented upgrade contract. Its running binary\nmay still carry the older destructive SIGTERM handler, so the CLI does not claim its\npre-signal agent inventory was spared; those agents may have been reaped.\n\n- **Broker**: a local `nats-server` (logs to `.cotal/nats.log`).\n- **Delivery daemon**: the durable backstop, auth mode only\n ([what it does](delivery-daemon.md)).\n- **Manager**: a detached supervisor answering the control plane, so\n `cotal spawn --detach` and the `cotal_spawn` tool work right after `up`.\n\nThree modes:\n\n- **Default (static auth).** JWT-authed, on by default: sender authenticity and per-agent\n ACLs, enforced by the broker ([how](identity-and-auth.md)).\n- **`--user-auth --idp <url>`.** Per-user auth: people `cotal login` once, the operator\n grants their agents on the actor ledger, and every connect is authorized live against\n that grant. Starts the space's auth service alongside the broker\n ([how](identity-and-auth.md)).\n- **`--open`.** An unauthenticated, live-only dev mesh (no auth, no delivery daemon). For\n quick local experiments.\n\nThe broker and local services bind **loopback** by default. `--host 0.0.0.0` widens the broker\nbind independently of the auth mode, so \"network-reachable\" never silently means\n\"unauthenticated\". With no explicit `--server`, `cotal up` auto-selects a free local port when\nthe default address is already held by another project; an explicit `--server` fails loud on\ncollision.\n\n`--host` is a boot flag, not a live rebind. A fresh `cotal up` writes the generated\n`.cotal/auth/server.conf` (project-local, not `~/.cotal`) with that bind and starts nats against\nit. If anything is already answering at the mesh URL, `up` refreshes the recorded mesh and\nleaves the running nats listener alone, so passing `--host 0.0.0.0` on a live or orphaned\nbroker does not change who can connect. To change the bind: `cotal down`, then `cotal up --host\n<addr>` against a stopped broker so the generated file is rewritten. Do not edit `server.conf`\nby hand; the next real boot overwrites it.\n\nOn a stopped shared broker, `up` renders every persisted space account and every enabled\nspace's auth-callout account into the resolver preload, regardless of which space starts\nthe broker. A missing callout account for an enabled space stops the boot rather than\nstarting with a reduced resolver. An already-running broker is refreshed without rewriting\nits config.\n\nA broker-only host is a first-class `up` mode. `cotal up --no-manager` boots the broker and, in\nauth mode, the delivery daemon, and no local manager, so the broker host never has a manager to\nstop and never leaves a manager slot stale. A refresh under the flag of a mesh whose manager is\nlive refuses rather than keeping or stopping it: `cotal down manager` first. Without the flag,\nauth-mode `up` still starts nats, the delivery daemon, and a\nlocal manager. A space may run more than one manager, addressed by instance id\n([control surface](control-surface.md#instance-routing)); putting no manager on the broker host\nis a topology choice, not a singleton invariant. A manager whose boot inventory has no\navailable connector does not take unpinned `spawn`/`launch` on the class rail, so a sibling\nthat can launch the harness can. `describe` still rides the class rail, so an unpinned spawn\ncan bind-fence when that skip member answered describe; re-issue, or pin `--on`. Pin one\ninstance with `--on` when a partial inventory still answers with a harness refusal. The\nsupported split is:\n\n```bash\n# broker host (project root that owns the generated conf, pidfiles, and logs)\ncotal up --detach --host 0.0.0.0 --space main --no-manager\n# no local manager starts: the summary lists nats-server + delivery daemon, and there is no\n# `.cotal/manager.<spaceKey>.log` to wait for on this host\n\n# manager host (registered remote mesh, same space)\ncotal meshes add --server nats://broker.example:4222 --root ~/meshes/main\ncotal supervise --space main --server nats://broker.example:4222\n```\n\nWait for `✓ manager up` in `.cotal/manager.<spaceKey>.log` on the manager host before spawning\nagents. On a broker host started without `--no-manager`, `cotal up --detach` prints `✓ running in\nthe background:` with `manager` listed once the manager pidfile is live; stop that local manager\nonly after the `✓ manager up` line. A host started WITH `--no-manager` never runs one, so neither\nthe wait nor the stop applies there. That detach stdout is not a safe teardown boundary: it is\npidfile liveness, not `✓ manager up`. `✓ manager up` is supervise's post-start line after\n`await mgr.start()`. `cotal down manager` after only the detach line can still default-terminate\nthe child during registration after it has taken the governance slot. Stopping before that\npost-start log line can leave the endpoint governance slot held until the holder's gate\nreopens past the stamp (the successor's boot heal, or\n[`cotal reconcile-gate`](cli.md#reconcile-gate) when that boot cannot run). See\n[Gate recovery](#gate-recovery).\n\nStandalone `cotal deliver --creds` is not a repair for that split. Production renewal needs\nthe manager and the daemon to address one credential store. Separate host filesystems still\nleave manager root A writing and the daemon reloading root B; that composition is refused\nwhile the daemon stays up. Keep delivery on the broker host under `up`, and share one store\nonly when you are composing a hosted pair ([embedding](embedding.md#supervisor-signing-authority)).\n\n### Split host bind\n\nA remote manager cannot reach a loopback broker. After changing `--host`, confirm the\ngenerated `host:` in `.cotal/auth/server.conf` and that nats is listening on that address\nbefore registering the mesh on the manager host. Detached child logs stay under the **project**\n`.cotal/` that `up` ran in (see [When something looks absent](#when-something-looks-absent));\nthey are not `~/.cotal` unless that directory is the mesh root.\n\nA user-auth mesh can expose only its credential exchange through an operator-owned HTTPS reverse\nproxy while leaving the existing local exchange untouched:\n\n```bash\ncotal up --user-auth --idp https://idp.example/api/auth \\\n --exchange-public-port 7443 \\\n --exchange-public-url https://auth.example\n```\n\nThe public listener itself still binds `127.0.0.1:7443`; configure the proxy to terminate TLS and\nforward to it. It serves only `/health`, `/jwks`, `/exchange`, and `/.well-known/cotal-mesh` with\nthe documented methods. It needs no local file capability: the signed IdP JWT or managed-agent\nactor token is the proof, while the original loopback listener remains capability-gated. Add\n`--exchange-trusted-proxy` only when that listener is reachable exclusively through your trusted\nproxy; it keys failure throttling by the last `X-Forwarded-For` hop instead of the socket address.\nThe well-known bundle includes IdP pins and a deny-all sentinel credential, so fetch it only from\nthe configured HTTPS origin. To change these listener flags, stop and restart the mesh; a refresh\nof an already-running service does not replace its bind or proxy policy. See\n[Identity & auth](identity-and-auth.md#per-user-authentication) for the trust boundary.\n\n### Remote supervised seats by enrollment\n\nA remote seat does not need to run `cotal login` when the mesh owner pre-mints a single-use\nenrollment for it. Mount the enrollment URL as a private file, place the seat persona on the remote\nmachine, and launch the foreground seat:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./worker.md --space main\n```\n\nThe URL is redeemed once with an unauthenticated GET. Redirects, off-machine plain HTTP, retries,\nand login fallback are refused. If the seat has no mesh record yet, the enrollment response's stock\nuser-bundle fields register it before the launch. The returned actor token then uses the same remote\nauth-service exchange as a login-provisioned agent. The enrollment URL and file path do not enter the\npreflight or harness environment. A failed or reused enrollment leaves no actor material on disk; ask the owner\nfor a fresh enrollment. When the foreground seat exits, this machine's credential files are removed\nand the mesh-side grant stays until the mesh operator revokes it; the launch line says so. The exact\nserver contract is in\n[Enrollment redeem](identity-and-auth.md#enrollment-redeem).\n\n`cotal status` prints the detailed setup, process, registry, and live mesh status. Its Machine\nsection names the running CLI's source checkout, installed package root, or npx package root beside\nthe version. A stale Claude skills row names the installed and CLI versions it compared. `cotal\nsetup` (after the first run) prints the compact card.\n\nBefore reporting ready, the manager resolves every installed connector's declared harness\nbinaries against its own environment. A missing binary does not stop unrelated manager work: boot\ncontinues, but prints a named `connector <name> unavailable` line and records that reason in the\nmanager's `status` response. Available connector rows record the absolute paths boot resolved.\nSpawn keeps the same pre-mint check as a backstop for connectors registered after boot.\n\nOn an authenticated manager start, unfinished static lifecycle rows reconcile while the control\nendpoint is already serving. The manager `status` response reports\nthe `staticReconciliation` state, the last sweep counts, and each failed alias with its durable\nphase and literal disposition. `cotal status --components` reports the state and per-alias failure\ndetails. A failed exact terminal is retried in the same process after 1, 5,\nand 30 seconds. Each attempt re-reads the durable slot and re-enters the same deterministic terminal\noperation; the delays only schedule work and never release the lifecycle fence.\n\nOn shutdown, the manager fences new reconciliation work and waits for an exact terminal that already\nstarted. The current serial sweep stops before its next alias, and startup cannot publish the manager\nservice after `stop()` completes.\n\nThe four-attempt budget is per manager process. An exhausted row stays held and reports\n`retry-exhausted` with the remedy to restart the manager. The next process derives a fresh budget\nfrom the still-authoritative durable row. A `recovered` row remains visible until the next static\nreconciliation sweep, then clears. This component reports reconciliation outcomes. It does not say\nwhether footprint cleanup completed independently of the terminal result; that separate durable\nprojection remains tracked by #1274.\n\n`cotal service install` is the supported way to run the manager as a user service\n([CLI reference](cli.md#service)): a systemd user unit on Linux, a launchd agent on macOS, one\nper mesh, surviving logout and reboot when lingering is enabled with `--linger`. It installs only\nthe manager; the units below remain the process models for every other component, and they are\nstill **examples of process models** for those: copy them only after you decide which processes\nthe unit should own.\n\n### Supervising the detached stack\n\n`cotal up --detach` is a launcher: it starts the broker, delivery daemon, and manager, reports what\nstarted, then exits. Do not wrap it in a systemd service with `Type=oneshot` and\n`RemainAfterExit=yes` and treat `systemctl is-active` as stack health. That unit becomes `active\n(exited)` when the launcher exits successfully and stays active even if every detached process dies.\nWhen `up --detach` can identify that exact unit shape, it prints a warning but keeps the requested\nstartup behavior.\n\nFor a single-host stack, keep `cotal up` itself in the foreground so systemd tracks a long-running\nprocess and restarts the stack if that process fails:\n\n```ini\n[Service]\nType=simple\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal up --space main --host 0.0.0.0\nRestart=on-failure\nRestartSec=5s\n```\n\nAn active unit then proves the foreground launcher and broker are still running, but it still does\nnot prove that every child component serves. Pair it with the component check below. Also remember\nthat `cotal up` starts a local manager as well as the broker and delivery daemon; run\n`cotal up --no-manager` (add the flag to the unit's `ExecStart` too) on a host intended to be\nbroker-only, so the unit and the host agree.\n\nThat `Type=simple` shape puts nats in the unit's cgroup with the foreground `up` process. A\n`Restart=always` (or `on-failure`) of **this** unit therefore restarts nats as well, so remote\nmanagers drop for the time it takes the broker to come back. Wrapping `cotal up --detach` in\n`Type=oneshot` with `RemainAfterExit=yes` does not move nats out of that cgroup. Detached\nspawn starts a new process group, not a new systemd cgroup, and the default\n`KillMode=control-group` still signals every process left in the service cgroup on stop or\nrestart, including the nats PID. Escaping that cgroup needs an explicit unit setting such as\n`KillMode=process`, or a separate nats unit; this CLI does not ship that escape. The\n`Type=oneshot` unit below is a `cotal status --components` liveness check, not a\n`--detach` launcher. Neither trade is universal from\n`Type=simple` alone; it follows from which processes the unit actually owns. `cotal service\ninstall` covers only the manager, so for the broker and its siblings pick the example that\nmatches the ownership you want, and treat\n`systemctl is-active` as unit health, not mesh health.\n\nIf the deployment deliberately uses `cotal up --detach` as a boot action, monitor observed state\ninstead of the launcher's exit:\n\n```ini\n[Unit]\nDescription=Check Cotal component liveness\n\n[Service]\nType=oneshot\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal status --components --space main\n```\n\nRun that check from a systemd timer or another monitor and alert on a nonzero exit. The command\ndistinguishes `absent`, `not-serving`, and `refused` components and never treats a sibling's health as\nproof. Its delivery-process check is local to the broker host, so run it there. On a split topology,\nalso probe the broker URL from the manager host and monitor the manager's own service there. A remote\nmanager cannot observe the broker host's delivery PID, and an `active` unit on either host says\nnothing about the other host.\n\nStop one part without tearing down the mesh by naming its registered component: `cotal down\nmanager`, `cotal down delivery`, or `cotal down web`. Component names from installed extensions\njoin the same surface; `cotal down` with no names retains whole-stack behavior and\nleaves managed agents running as unmanaged OS processes. `cotal down --with-agents`\nis the previous reap. If a pinned manager has no spare-capability record, stop its managed agents\nexplicitly before running that whole-stack command. The record is also absent when a manager predates\ncapability reporting, and that older manager may not understand the reap request.\n\n## Remote supervised agents\n\nOn a remote user-auth mesh, foreground `cotal spawn` remains the default participant path. A\nparticipant can run detached agents only after the host advertises and operates the remote manager\nauthority service, and the participant's actor-ledger row includes `supervise`. This is not implied\nby `spawn` or `admin`.\n\nThe participant's loopback/operator exchange obtains one closed `manager-service` view for its\nordinary derived owner, a fixed server-selected manager actor, and one opaque manager instance.\nThe host, not the participant, issues the public-nkey JWT material via the replay-safe,\nlifecycle-bound prepare → activate → renew exchange, plus a one-shot target-pinned retirement\nrequest for a host-managed terminal. It never exports the space signer, a static\nprovisioner credential, or generic storage authority. The manager may provision only descendants\nof that same owner, with host validation at each provision.\n\nThe registry entry decides the broker URL `supervise` dials, so a mesh published over `wss://` is\ndialed as a websocket. The manager-authority registration it runs first also takes its TLS\nrequirement from that entry, so the prepare credential is not exchanged over a plaintext\nconnection the record did not describe. `cotal meshes add` records both.\n\nWhen the authority service, login, or renewal is unavailable, the remote manager degrades\nfail-closed: it refuses new agents, restarts, and credential replacement rather than pretending\nlocal authority exists. Existing agents remain live only while their independent credentials are\nvalid. A hosted composition must revoke the managed grant and finish its resumable release before it\nrequests terminal retirement. Deleting DM or delivery consumers is not retirement and must not reset\na resumable lifecycle's frontier or pending state. The alias remains held until the terminal barrier\nconfirms. Restore service and renew successfully before asking it to recover an agent. See\n[Identity & auth](identity-and-auth.md#remote-manager-authority) and the [CLI\nreference](cli.md#supervise).\n\n## Spawning agents\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn reviewer --detach # supervised: the manager runs it in a PTY\ncotal attach --name reviewer # watch/type into a detached agent (Ctrl-] detaches)\ncotal ps # what the manager is running\ncotal stop --name reviewer # stop one\n```\n\nHow a spawn resolves:\n\n- **Persona.** A bare `cotal spawn` uses `.cotal/agents/default.md`; a positional name\n picks `.cotal/agents/<name>.md`; `--config` takes an explicit ref or path. Set\n `COTAL_DEFAULT_PERSONA=<name-or-path>` to change the fallback. Fields and format:\n [agent files](agent-files.md).\n- **Harness.** Resolution order is an explicit `--agent` or `cotal_spawn` `agent` argument,\n then the persona file's `agent:` pin, then the invoking caller's `COTAL_DEFAULT_AGENT`,\n then the manager's `COTAL_DEFAULT_AGENT`, then the product default (Claude). Compared in\n [Connectors](connectors.md); per-connector guides:\n [Claude](connect-claude.md) · [OpenCode](connect-opencode.md) ·\n [Hermes](connect-hermes.md) · [pi](connect-pi.md).\n- **Model.** `--model` overrides the persona file's `model:` (Claude: `opus` / `sonnet` or\n a full id; OpenCode: `provider/model`). Connectors that expose a catalog report it via\n `cotal models --agent opencode`: model ids plus available variants; pick one with\n `--model provider/model --variant high`.\n- **Tools.** A spawned agent gets only the cotal tools by default; share your own MCP\n servers deliberately with `--share-tools` ([config](config.md)).\n- **Launch options.** `--opt key=value` (repeatable) passes a native harness flag straight\n through; a persona or manifest `launchOptions:` mapping does the same declaratively (a\n `--opt` wins per key). It is a **raw passthrough**, with no allow/deny list: Claude renders\n each as `--key value` (a bare `--key` for an empty value), OpenCode merges them into its\n agent config, and Hermes has no option surface so it fails loud. The trust boundary is the\n `spawn` capability itself, not the flag set, so granting `spawn` is host-launch authority\n ([security](security.md)). A key must be a plain flag name; malformed or prototype-polluting\n keys are refused.\n\nDetach from an attached PTY with **Ctrl-]** (the agent keeps running); rebind it with\n`COTAL_DETACH_KEY=ctrl-<char>` when it clashes with a keybinding inside the agent's TUI.\n\n**Runtimes.** The manager spawns into a **pty** by default. On Linux a detached per-seat\ncustodian owns that PTY, so replacing the manager worker does not close the seat. A custodian\nwhose agent has exited exits a few seconds later on its own, and a manager that boots after a\ncrash reaps the seats its predecessor left running. Other\nplatforms still spawn the PTY in-process; `adopt` throws until their transport lands. Optional runtimes are installed\nthrough the extension surface, for example `cotal ext add @cotal-ai/orca`, then selected with\n`--runtime orca` (similarly `@cotal-ai/tmux`, `@cotal-ai/cmux`, and `@cotal-ai/herdr`). They put teammates in native\nterminal surfaces rather than manager-owned PTYs. Runtime names are open-ended and resolved from\nthe registry; a missing provider or app throws, never silently falls back\n([architecture](architecture.md)).\n\n## Mesh registry\n\n`cotal up` records each running mesh in a machine-local registry\n(`~/.cotal/meshes/space.<key>.json`, named by a case-safe hex encoding of the space: broker URL, the project root holding its creds and\npersonas, and its mode). So a bare `cotal spawn <persona>` from *any* directory joins the\nrunning mesh with the right credentials instead of mistaking the cwd for a space:\n\n- `cotal use <name>` sets the default from every directory, including inside another mesh's\n project. `--space <name>` overrides it for one command.\n- When one broker has records for several spaces, `cotal up --space <name>` refreshes that named\n space.\n- With no live selected default, a project with its own `.cotal/` resolves to that project's\n mesh; otherwise one running mesh is used automatically and several are an error.\n- `cotal meshes` lists them (a `*` marks the default); `cotal down` removes the entry.\n\nThe registry stores a *path*, never a secret; trust material stays in each project's\n`.cotal/auth`. If the mesh is down or won't take your creds, spawn fails with one\nsentence, never a raw NATS trace.\n\n### Meshes you did not start here\n\nA mesh running on another machine has no `cotal up` on this one, so register it by hand:\n\n```bash\ncotal meshes add # guided: asks for the broker, probes it, offers what it finds\ncotal meshes add optiplex --server nats://100.90.12.34:4222 --root ~/meshes/optiplex \\\n --allow-unencrypted-overlay # see below: an overlay address needs this\ncotal meshes rm optiplex\n```\n\nOn a terminal, a bare `cotal meshes add` walks you through it: it probes the broker you name and\nreports whether it is open or requires credentials, offers the spaces the folder already holds\ncredentials for, and shows the record before writing it. Scripts and agents keep the flag form -\nwithout a terminal nothing prompts.\n\n`--root` is the local folder holding that mesh's `.cotal/auth` and `.cotal/agents` (its personas);\nthe mode is inferred from what that folder holds.\n\n**Know what you are copying.** For an authenticated mesh that folder carries the space's account\n**signing seed**, which is the authority to mint any identity in the space. A machine holding it\nis a certificate authority for the mesh rather than a client of it: anyone who reads it can\nimpersonate any agent, read every retained channel and DM, change ACLs, and keep issuing\nthemselves credentials. There is no per-machine revocation; undoing it means rotating the signing\nkey and re-minting every credential in the space. Copy it only to machines you would trust with\nthe whole mesh. `cotal mint` on its own does not substitute here: registering an `auth` mesh needs\nsigning material that composes, which a minted user credential is not. The\nbroker is probed before the record is written, so a bad address or a credential that mesh will not\naccept fails at registration rather than at your first `spawn` (`--force` records it without verifying,\nuseful when the mesh is simply down right now).\n\n#### Which addresses you may register\n\nRegistering a mesh is how this machine starts sending agent credentials to a broker it does not\nrun. NATS announces itself in plaintext before anyone authenticates, so an attacker on the path\ncan pose as the broker and read the credential out of the connect unless the connection\n**requires TLS**, which is recorded on the entry and enforced on every dial through it.\n\nWhat the record will require decides what you may register:\n\n- **Without required TLS**, the address is the gate: **loopback** (`127.0.0.0/8`, `::1`), or\n **your private overlay** (`100.64.0.0/10`, `fd7a:115c:a1e0::/48`) with\n `--allow-unencrypted-overlay`. The tunnel provides the protection, and this command cannot check\n its state. Hostnames are refused because the lookup would choose which machine receives your\n credentials.\n- **With required TLS**, set `--tls` or use a `tls://` URL. The recorded scheme enforces the TLS\n requirement. A **hostname or public address** is accepted because the certificate chain and\n hostname check identify the peer. A registration whose broker cannot complete the handshake\n fails unless you pass `--force`, which records the entry without verification.\n\nOrdinary private ranges like `10.x` and `192.168.x` are refused in **both** modes. A café's wifi\nis private but does not belong to you, and no public CA issues certificates for those ranges. An\naddress spelling changes nothing: `[::ffff:192.168.1.10]`, `3232235786`, `0300.0250.01.012`, and\n`192.168.257` all resolve to private addresses and receive the same refusal as the dotted form.\n`--force` exists for a mesh that is down. It never permits an unsafe credential destination.\n\n#### Registering a hosted user-auth mesh\n\nA user-auth space's IdP pins are established where the mesh runs and are never guessed. Register\none from **supplied** trust: `--user-auth-file bundle.json` (exported on the mesh's machine), or\n`--from https://…/.well-known/cotal-mesh`, which asks before it contacts the address at all,\nfetches the discovery document over HTTPS, shows you the pins, and asks again before adopting\nthem. Redirects are refused because a 302 can walk a pinned fetch down to\nplaintext or onto another host, and the pinned exchange must be an `https://` URL too. The one\nexception is an exchange on **this machine**, where nothing leaves the box: plain `http://` is\naccepted for a loopback *literal* (`127.0.0.1`, `::1`, and any spelling of them), but **not** for\n`localhost`, which a hosts entry or poisoned lookup could point elsewhere. Use the\nliteral. Registration checks that the pinned exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also checks that the broker refuses a\nbare connect; that refusal is the pass. The bundle's sentinel credentials are written to a private (0600) file\nunder the entry's root; the registry itself never carries the secret.\n\n**Without required TLS**, an overlay address is **refused unless you accept the dependency\nexplicitly**, with `--allow-unencrypted-overlay`. The address is not the guarantee: it is protected\nwhile the tunnel is up, and if the tunnel is down that range is ordinary carrier-grade NAT and\nwhoever answers the dial receives your credentials. Only you can know which it is, so the command\nasks you to say so. Your acceptance is recorded on the mesh entry rather than printed and\nforgotten, and the guided form asks the same question instead of taking the flag.\n\n**With required TLS** (`--tls`, or a `tls://` URL) that consent is no longer asked for, and the\nflag is not needed: the handshake is what protects the connection, so the acceptance it stood in\nfor has been replaced by proof rather than promise. `cotal meshes add <space> --server\nnats://100.64.0.1 --tls` registers an overlay address with no prompt, no flag and no recorded\nacceptance. This is the \"the flag disappears once the broker can be served over TLS\" case, and it\nhas now arrived.\n\nThis gate is on **registration**. `cotal join --creds --server <url>` deliberately takes an\nexplicit connection at face value and does not consult the registry, so it is not covered. Join\nthat way only to an address you would have registered.\n\nRecords added this way are removed only by something that names them. A failed liveness probe\ndoes not delete any record: an unreachable broker, local or registered by hand, is shown as\n`offline` in `cotal meshes`. A bare command does not count that offline record as running;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the\nproject they are tearing down, and they leave a hand-registered one alone even when `--root`\npointed at that project. A `cotal up` for that space refuses outright (naming `cotal meshes rm`) unless it is\nthat same endpoint: finding a broker already answering there is a refresh that starts nothing and\nleaves the record's provenance alone, while actually starting the broker for that space, server and\nroot makes this machine the one running it, so the record becomes an ordinary local one that\n`cotal down` clears. `cotal meshes rm` drops it and re-registering with `--force` replaces it. `rm`\nonly forgets a mesh. To stop one running here, use `cotal down`.\n\n## Watching\n\n`cotal console` is the terminal view (TUI on a real terminal, plain line stream when\npiped); `cotal web` is the browser dashboard. Both are read-only observers; the\nwalkthrough is [Watch a mesh](watch-a-mesh.md).\n\n## History\n\nRetained history is operator-owned. `cotal clean history --force` purges a space's\nretained channel history; `--dms` also purges DMs (`cotal history clear` is an alias).\nIt is deliberately **not** an agent tool: agents cannot wipe the record\n([identity & auth](identity-and-auth.md)). For a **stopped** mesh, `cotal clean store\n--force` deletes the on-disk JetStream store outright, and `cotal clean all --force`\nalso resets the space identity ([CLI reference](cli.md#clean)).\n\n## Offline backup\n\nFor a coherent durable cut, preserve the whole stack first, then create the artifact while it stays\ndown:\n\n```bash\ncotal down --preserve-state\ncotal backup create ./space-backup # full by default\n# later: deliberately resume the unchanged source\ncotal up --detach\n# or, from another preserved cut, restore before the normal listener opens\ncotal up --restore ./space-backup --detach\n```\n\nUse `--store-dir` on both preservation and backup for a custom JetStream store. A store cap set\nwith `cotal up --max-file-store <bytes>` travels with the preserved state, and the resume renders it\nagain. nats-server reads the cap once at start and refuses a config reload that changes it, so a new\ncap always needs a restart. `registry` is the\nonly partial selection (`backup create ... --only registry`; `up --restore ... --restore-only\nregistry`). Backup never stops or restarts a mesh implicitly, never opens the original store, and\ndoes not contain credentials or trust secrets. Backup/restore in every auth mode, open included,\nuses isolated, operation-specific maintenance logins; normal agent credentials cannot enter that\nlistener. Full\nrestore requires the same space and exact current local trust continuity, recreates conservative\nconsumer checkpoints bound to their snapshot stream sequence state, and resumes retained agents under\ntheir original principals. The trust commitment includes the cryptographically validated full\noperator/system/data-account root chain as well as static/user authority state. A registry-only\nrestore completes canonical empty infrastructure but leaves retained agents stopped because their\nDM/DLV/TASK/ACL state is outside that selection. Authenticated restore validates the complete space\ntrust bundle before staging or changing the preserved store. Interrupted ordinary resume retries the\nsame durable attempt after its prior listener is stopped. Restore re-entry can recover a surviving normal listener\nonly when its attempt nonce, NATS server name, process owner, endpoint, and target-store identity all\nmatch the fsynced proof. A provably dead uncommitted owner is retired under lock and replaced with a\nfresh attempt-bound listener; an occupied foreign listener or ambiguous owner is never adopted. The\nmanager commit validates while retained cleanup is still suppressed; the CLI durably records its\nattempt-bound 64-hex token in `manager-committed` / `resume-committed` before `finalizeResume` can\nrelease suppression. A retry from either committed state goes straight to exact-token finalization;\nfailure preserves the committed gate and retained cleanup suppression. Missing commit evidence,\ninterrupted finalization, a live recorded endpoint despite missing pidfiles, or ambiguous proof fails closed. See the [CLI\nbackup and restore contract](cli.md#backups) for artifact, checkpoint, fallback,\ndisaster-consent, and degraded-recovery details.\n\n## Personas from the CLI\n\n`cotal personas` manages the local catalog offline: `list` (`--running` overlays live\nmarkers), `show <name>`, `edit <name>` (re-validates on save), `new <name>`, `rm <name>\n--force`. The runtime write is `cotal_persona`; the runtime read is `cotal_personas`\n(list / show), both over the wire with the manager's ownership checks. Fields: [agent files](agent-files.md).\n\n## Gate recovery\n\nA manager that dies mid-registration leaves its issuance gate *frozen* under that registration\nop. The freeze is correct: it stops two incarnations serving at once. The successor now completes\nthat dead op on boot, using the same guard as [`cotal reconcile-gate`](cli.md#reconcile-gate): it\nacts only when the freeze-holder is affirmatively gone under a complete CONNZ sweep (`gone` and\n`sweepComplete=true`). If the dead op's spec write committed, it finishes that same freeze\n(promote and reopen at the committed registration revision). If the spec did not advance, it\nabort-reopens the gate (generation+1, processEpoch unchanged) and continues the normal takeover.\nBoot heal and the following re-registration use separate one-shot executor windows, so a large\npredecessor family cannot spend the takeover's credential lifetime. If that later registration\nstill crosses a connection lifetime, it retries the same frozen operation with fresh authority\nand resumes verified-holder progress instead of freezing a new generation.\nA live holder, an incomplete sweep, or an unreachable delivery daemon still\nrefuses. Silence is never evidence of death, and there is no TTL. If holder verification is\ninterrupted, the frozen operation resumes from its durable, operation-and-gate-revision-bound\nprogress after liveness is checked again. A later freeze cannot reuse that progress: the cursor\nbinds the exact op, gate revision, and holder set. Use `cotal reconcile-gate` when the boot path cannot run\n(daemon down, a non-manager endpoint, or you want to lift the freeze without starting a manager). A spawn that hits the same frozen gate names that verb in the refusal\n(`blockedOp=registration`, the holding `opId`, `remedy=cotal reconcile-gate`) instead of a\nwait-timeout: the facts were always in the manager log; they now reach the spawn caller too.\n\nGive reconciliation a **quiet manager**. Suspend systemd restart policies, watchdogs, health-check\nrestart loops, and any other automation that can start or kill `cotal supervise` while boot healing or\n`cotal reconcile-gate` is running. Leave one recovery attempt in control until it finishes.\nRestarting the manager during the walk interrupts the current authority window. Durable progress makes\nthat interruption resumable, but a quiet manager is still the fastest and safest incident procedure.\n\n### Last-resort JetStream store replacement\n\nStore replacement is not normal gate recovery, is never automatic, and is destructive to mesh history.\nUse it only after the retained store cannot be reconciled and after deciding that losing its durable\ncontents is acceptable.\n\n1. Stop every actor touching the space: supervisor, watchdog, manager, delivery daemon, and broker.\n Confirm that no Cotal or NATS process still has the store open.\n2. Preserve the stopped store before changing anything. Move `.cotal/nats` aside to a dated backup and\n archive both `nats` and `auth`. Do not delete the only copy.\n3. Understand the loss: replacing the store removes JetStream message and control history and durable\n consumer state. Agent session files stored outside JetStream remain, but the mesh history they\n referenced does not.\n4. Start the broker against a new empty store, then start one manager. Wait until it reports\n serving successfully.\n5. Repopulate the mesh only after that manager is healthy. Re-enable supervisors, watchdogs, and other\n restart automation last.\n\nKeep the preserved store until the incident is reviewed and any required forensic or manual recovery is\ncomplete. Restoring it later restores the old durable state, including the fault that led to this last\nresort, so do not swap it back into a live mesh casually.\n\n## When something looks absent\n\nPermission denials are **loud, never silent**: an over-tight ACL rejects the endpoint call and\nalso shows up as a logged denial, instead of returning an empty or incomplete result that looks\nsuccessful. Check\n`.cotal/manager.<key>.log`, `.cotal/delivery.<key>.log` (one pair per space, keyed as\n[Config](config.md#project-files) describes), and `.cotal/nats.log`; `cotal status` shows\nwhat is actually running. Those files live under the **project** `.cotal/`, not `~/.cotal`,\nunless the mesh root is the home directory. `cotal up --detach` redirects delivery and manager\nstdio onto those files, so an operator-created systemd unit around that launcher does not put\nthe child logs in that unit's journal. `journalctl -u <unit>` can be empty while the crash\nreason is already in the project log. The access rules are collected in\n[Channels & permissions](channels-and-permissions.md).\n"
228
228
  },
229
229
  {
230
230
  "slug": "security",