@cotal-ai/connector-core 0.71.0 → 0.72.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import type { DocsBundle } from "./docs.js";
|
|
2
2
|
/** The installed Cotal version, for stamping the orientation card and other surfaces. */
|
|
3
|
-
export declare const DOCS_VERSION = "0.
|
|
3
|
+
export declare const DOCS_VERSION = "0.72.0";
|
|
4
4
|
export declare function loadDocsBundle(): DocsBundle;
|
|
5
5
|
//# sourceMappingURL=docs-bundle.generated.d.ts.map
|
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
/** The installed Cotal version, for stamping the orientation card and other surfaces. */
|
|
2
|
-
export const DOCS_VERSION = "0.
|
|
2
|
+
export const DOCS_VERSION = "0.72.0";
|
|
3
3
|
export function loadDocsBundle() {
|
|
4
4
|
return {
|
|
5
|
-
"version": "0.
|
|
5
|
+
"version": "0.72.0",
|
|
6
6
|
"generatedFrom": "docs/*.md + SPEC.md + spec/cotal-lang.md + spec/cotal.schema.json",
|
|
7
7
|
"pages": [
|
|
8
8
|
{
|
|
@@ -136,7 +136,7 @@ export function loadDocsBundle() {
|
|
|
136
136
|
"title": "The control surface",
|
|
137
137
|
"kind": "Concept (informative)",
|
|
138
138
|
"summary": "Cotal once had a privileged control rail: a fixed set of named service tiers (self / manager / admin / delivery) on their own ctl.",
|
|
139
|
-
"body": "# The control surface\n\n> **Concept** (informative) · **For:** operators and client authors who want to know how the manager and other daemons are driven · **Normative:** [SPEC §13](../SPEC.md#13-endpoint-control-surface-v04)\n\nCotal once had a privileged control rail: a fixed set of named service tiers\n(`self` / `manager` / `admin` / `delivery`) on their own `ctl.*` subjects, with the manager\nas a special case the broker recognised by name. That rail is gone. Everything that serves\nstructured commands now, the manager, the delivery daemon, a wrapped MCP server, a\nthird-party service, is an ordinary **endpoint**: a daemon that registers a service\nidentity, publishes its contracts, and answers `describe`. `manager` is an endpoint name\nlike any other; no subject, envelope, or grant in this surface knows it specially. The\nmanager is a service on the mesh, not an authority over it: it holds only the capability\nrows its callers grant it, and serves over a scoped credential.\n\n## The `ep` rails\n\nOne kind, `ep`, carries every request under a mode token that says where the request\nroutes, never which verb it is (the verb rides the envelope): `one` (queue-group\nanycast, one and only one class member), `all` (scatter, every instance), and `inst` (one instance by its\nstable address). Replies come back on a `reply` rail keyed to the serving instance and its\nepoch. Around these sit the sibling planes the composites use: per-goal events, timers,\nsessions, and the journal that holds durable facts. Every request carries the caller as\nthree forge-locked tokens, `owner`, `actor`, and lifecycle `uid`, plus an unguessable\nnonce, so the broker polices who is calling in the subject grammar itself. See\n[SPEC §13.2](../SPEC.md#132-grammar) for the grammar and [§13.5](../SPEC.md#135-verbs) for\nthe verbs (`call`, `cast`, `watch`, `claim`, `scatter`).\n\n## Lifecycle identity\n\nA principal `owner.actor` is a reusable routing alias: a despawn frees the actor name and a\nlater spawn may legitimately reuse it, so the alias alone is never authority. Two further\ncoordinates make an identity durable: a **lifecycle uid**, an unguessable, never-reused id\nfor one managed lifecycle under a principal, and a **process epoch**, the fenced ownership\nepoch of the process currently animating it, advanced on every restart or takeover. At most\none live epoch owns an identity, and a superseded epoch must stop serving. Durables and\ncredentials key on the lifecycle uid, not the reusable name, which is what lets a\nsupervised restart recover the same lifecycle instead of minting a new one. See\n[SPEC §13.1](../SPEC.md#131-lifecycle-identity) and [identity & auth](identity-and-auth.md).\n\n## Service discovery\n\nNo client has compile-time knowledge of any endpoint's commands. `cotal describe\n<endpoint>` resolves a registered endpoint's command set off the wire: the reserved\n`describe` command answers the registered contract digests, the schemas are fetched from the\nspace's content-addressed contract store, recompiled, and verified against those digests.\nEach command prints with its capability class and targeting shape. `cotal invoke <endpoint>\n<command> --args '<json>'` then calls one command by name, validating the arguments as they\nwill be sent (JSON drops a key whose value is undefined) against the fetched input schema before\npublish. A refusal at that check means nothing was sent, and it is marked `not-executed`. A\nsigned-in user invokes the same surface through their bearer, and the broker enforces each\ncommand's existing capability grant. A manager alias supplied through `--name` resolves through\nits name-keyed `inspect` command, so an authorized targeted call does not need the manager-wide\n`ps` enumeration grant. Every built-in manager command uses this\nsame trust chain, so there is nothing the built-ins can reach that a described contract cannot. The registered\n`auth` endpoint is describable the same way `manager` is: `cotal describe auth` lists\n`retire-lifecycle` and its exact-mode target shape.\nSee [SPEC §13.7](../SPEC.md#137-contracts-and-discovery) and [cli.md](cli.md).\n\nThe manager's `resolve-cwd` command is in the `manager.spawn` capability class. It accepts an\nabsolute path on that manager's host and returns its canonical directory plus the host name. It\nrefuses a relative, missing or non-directory path with `failed-precondition`; it creates nothing.\n`spawn` applies the same check at admission, before any credentials or durables are minted.\n\n### Inspecting a managed name\n\nManager `inspect` keeps its successful response as the live managed-agent row. A live hit does\nnot read durable lifecycle state, so a temporary records-store failure cannot break inspection of\nan agent the manager currently holds.\n\nOn a live miss, a static manager point-reads its durable slot row. A name with no slot, or a slot\nwhose phase is `retired`, remains `not-found`. A nonterminal slot returns\n`failed-precondition` with `error.details[].kind =\nai.cotal.manager.static-slot-observation`. The detail carries the slot's `slotPhase`,\n`owner`, `actor`, `slotLifecycleUid`, `cleanupComplete` when recorded, and `slotRevision`. It\nthen carries the separate lifecycle head's `headState`, `headOp` when present,\n`headLifecycleUid`, and `headRevision`. Head fields are absent when provisioning has not written\nthe lifecycle head yet.\nThe error message carries the same diagnostic summary so string-only operator paths do not hide\nthe structured detail.\n\nA slot row records the manager instance that owns it. In a space with more than one manager, the\nclass queue can hand `inspect` to an instance that does not host the name. When the row names a\ndifferent instance and is not `retired`, the miss returns `failed-precondition` with the same\ndetail plus `ownerInstanceId`, which names the only manager that can act on it. The message names\nboth instances, so a caller that reads only the string can tell it from `not-found`. A sibling's\n`retired` row remains `not-found`. A named `cotal_despawn` resolves its target through this read\nand cannot address an instance, so it asks again until the owning instance answers, up to 16\ntimes.\n\nThe slot is read before the head. These records do not form one atomic snapshot, so the detail\nalso carries `readOrder: [\"slot\", \"head\"]` and `consistency: \"ordered-not-atomic\"`. A head can\nadvance between the reads. The issuance gate is not projected because the retirement operation\nneeded for this diagnosis is already recorded on the head, and reading a third record would add\nanother non-atomic edge without changing the per-name result.\n\nIf either durable read fails or exceeds its bound, the miss returns `unavailable` with\n`ai.cotal.manager.static-slot-read-failed` rather than claiming the name is absent. That detail\nnames the inspected `name`, the failed `record` (`slot`, `head`, or `slot-or-head` when the layer\ncannot distinguish them), and `operation: \"read\"`. User-auth managers do not own `mgrslot` rows,\nso their inspect misses remain live-map reads.\n\nFor Linux custodied seats, retirement requires the runtime's process-exit evidence before\nfreeing the alias or deleting its credentials and delivery state. Socket loss alone is not\nproof of exit. The runtime retains the record captured at launch or adoption so a clean\ncustodian exit can unlink its file without losing the recorded boot and process identities.\nIf the file is missing, reaping uses that retained record and the existing kernel identity\nchecks. An unknown reference without either record refuses cleanup. Reused process ids\nare never signalled on the strength of the old record.\n\n### Listing the durable slots\n\nManager `slots` (`manager.read`, untargeted) lists the durable static slot rows this manager\nowns. Only static managers hold these rows: a user-mode or open manager answers\n`failed-precondition`, and a manager whose durable store is not standing answers `unavailable`.\nEach row carries the same `readOrder` and `consistency` fields `inspect` uses, because the list\nis read the same way: torn across rows as well as within each row's slot/head pair. A `retired`\nrow is never listed. `live` reflects the manager's live roster at render time, not the durable\nrow.\n\n## Spawn is a goal\n\nLong-running commands are **actions** ([SPEC §13.6](../SPEC.md#136-composites)): the caller\nsubmits with a client-generated `goalId` and a request fingerprint, the endpoint records a\ndurable accept or reject decision, progress rides per-goal events, and the work ends in one\nterminal outcome (`succeeded`, `failed`, `cancelled`, `expired`, or `uncertain`). Spawn is\nthe reference case. Rather than block the caller for up to 30 seconds while an agent comes\nup, the manager accepts the goal and returns the allocated identity at once:\n\n```json\n{\n \"name\": \"reviewer_2\",\n \"owner\": \"u_...\", \"actor\": \"reviewer_2\", \"uid\": \"...\",\n \"goalId\": \"...\", \"fingerprint\": \"...\",\n \"readinessDeadlineMs\": 30000,\n \"executor\": { \"lifecycleUid\": \"...\", \"epoch\": 3 }\n}\n```\n\nThe `uid` is the lifecycle the agent runs at. On a participant manager whose host enrolls its\nagents, the host picks that uid, so the manager accepts the goal only after the host has answered.\nA host refusal there refuses the spawn, and no goal is bound.\n\nThe name is the one actually allocated: a persona-derived collision is auto-numbered\n(`reviewer`, then `reviewer_2`), while a hard-pinned `--name` that collides with a live\nagent is refused at accept, before anything is minted. Auto-numbering never hands out a numbered\nname it has already issued in that manager process, even after the agent holding it is gone, so\na collision takes the next number. Only numbering consults that history: a hard-pinned `--name`,\nor a persona whose own name is a numbered string, takes that name whenever it is free, and\nnumbering does not skip a string such a spawn held before. The triple plus `goalId` let the\ncaller follow progress (connector handoff, process launched, presence join) and reconcile\nlater against the exact instance that accepted. Presence within the manager's default\n30-second readiness window, or a connector's declared bounded window, settles the goal\n`succeeded`; an early process exit is `failed`; the window passing with neither is `uncertain`,\na bounded, durable outcome that a later `ps` or status read settles against the live roster.\n`uncertain` is a real terminal outcome, not an absence and not a silent hang. It carries the\ndiagnosis of whoever owned the deadline: for a launch that\nnames the agent and says to inspect it rather than re-issue, since re-issuing after a launch\nthat in fact succeeded mints a duplicate. A follower keeps the acceptance as the data of any\nterminal other than `succeeded`, so `cotal_spawn` returns an uncertain launch as a pending result\ninstead of an error: it names the allocated agent, its id, and its manager, and tells the calling\nagent to watch the roster. A committer that supplies no diagnosis falls back to\n\"the success signal did not arrive within the readiness deadline\". The agent's own eventual\nstate is then observable on its presence record.\n\nThe acceptance carries that exact `readinessDeadlineMs`. A synchronous follower treats its own\nrequest deadline as a floor and waits through the accepted readiness budget plus delivery margin,\nso a connector-specific slow boot cannot be reported as a caller timeout while the manager is\nstill legitimately waiting for its terminal.\n\nA spawn that is **refused** because a lifecycle barrier already holds the actor (a frozen\nissuance gate, a retiring alias, a retired uid) is not a wait-timeout. The manager already\nknows the blocked op (`registration` / `retirement` / `activation` / `takeover`), the `opId`\nholding it, and the remedy when one exists (`retry`, `cotal reconcile-gate`). The detail\ncarries `headState` (`active` / `retiring` / `retired`) only when the refusing site read the\nlifecycle head, and `gateState` (`frozen` / `retired`) only when it read the issuance gate. A\ngate frozen by a takeover or a registration says nothing about the head, so that refusal\ncarries `gateState=frozen` and no `headState`. Those facts ride `error.details[]` as\n`kind = ai.cotal.ep.lifecycle-blocked` and are also appended to the error string, so a\ncaller that only prints `error.message` still sees them. The CLI and the connector tools hand a\nrefusal on in one shape, so `cotal spawn -f` keeps the same code, details, rendered facts and\nacceptance data as `cotal spawn --detach` and `cotal_spawn`. A connector that collapses the\nrefusal to \"startup failed (unknown)\" or a SPEC 13.6 wait-timeout is hiding a knowable\nstate, not reporting a missing one.\n\n## Instance routing\n\nA space can run more than one manager. Each manager persists a stable logical instance id\nacross restarts and advances its process epoch when it comes back, so callers address a\nspecific manager without caring which process currently serves it. A start serves only at the\nepoch its own registration committed, never at the epoch of a later start of the same instance.\nOn a static or open mesh,\nan untargeted spawn rides class anycast (any manager may accept, and the acceptance records which one did).\n`cotal spawn <persona> --detach --on <instance>` and `cotal_spawn(instance: \"<instance>\")`\npin one instance by its exact id. A foreground CLI spawn has no manager to pin and refuses the\nflag. An MCP pin that does not resolve is refused without falling back to class anycast. There are no ordinal\naliases and no short forms: wherever a display names an instance you can address, it prints\nthe whole id, because both surfaces take nothing else.\n\nOn a user-auth mesh, manager commands obtain a short-lived `manager-caller` view from the\nexchange. It authorizes one concrete manager instance using the caller's current actor grant and\nthe host's registered service records. Discovery and invocation both use that instance's `inst`\nroute. This view grants no registry scan, class queue, or additional command capability. An absent,\nambiguous or unauthorized selection refuses before the command is sent.\n\nManaged launches carry `COTAL_MANAGER_INSTANCE` so their tools address the manager that launched\nthem. Existing unbound sessions can use the exchange's unique authorized selection without\nreplacing their actor or conversation. The connector uses a separate control connection; the\nstanding message connection and its credential source are unchanged. Accepted spawn goals are\nfollowed on that renewing connection, using its existing caller-scoped progress grant, so a long\nreadiness budget does not depend on the short-lived control credential. The follower confirms its\nprogress subscription with the broker before submitting on the separate connection. A caller still\nchecks the resolved instance and epoch, and never retries an ambiguous mutation outcome.\n\nThe manager's `goal-result` command accepts `{goalId}` and returns `{goalId, result?}`. It reads\nonly the authenticated caller's owner, actor and lifecycle through the manager's separate trusted\ngoal-writer connection. The caller receives an attributed reply, never a raw JetStream reader\ngrant. Each read is admitted by the connection's broker-enforced command grant. A live user-auth\nconnection remains bounded by its bearer expiry after revocation; a renewed connection is checked\nagainst fresh authority. There is no separate per-read ledger check. An absent `result` means no\nterminal is recorded; it does not prove the goal is running or permit another submission. The\nexisting trusted goal-writer's leader-served EPF read is space-wide at the broker; the handler\nconfines it to this endpoint and caller triple.\n\nThe manager's reserved `cancel` command accepts `{goalId, mode?}` and returns `{goalId, state}`.\nIt is served for a turn the manager relays. The goal is the authenticated caller's own, so a caller\nwithdraws only a turn it submitted. The turn ends `cancelled`, its seat is not shown it again, and a\nlater yield of it is answered with that terminal. A goal that already ended is refused\n`failed-precondition` with its cached outcome attached, and a goal this manager does not relay is\nrefused without being changed. So is a second cancel that arrives while a first is still ending the\nturn; a first that fails leaves the turn pending unless something ended it meanwhile. A workflow\nrun sends it for the turn, ask attempt or escalation of a branch it cancelled.\n\nA followed mutation requires a manager whose attributed describe includes `goal-result`. Update\nthe manager, issuer and client together before using that recovery path. Reloading an issuer alone\ncannot change an already-running participant manager. Recovery re-resolves the accepting instance's\nepoch, preserves the caller lifecycle and validates the result against the accepted goal and any\nacceptance fingerprint. Stopping the caller ends its observation, not the already accepted goal.\n\nA followed call resolves the endpoint within its deadline before the submission starts, so a\nrefused or unanswered describe surfaces as its own error. A describe or command publish that the\nbroker refuses reports `not-executed`.\nCancellation before submission reports `not-executed`. Once submission starts, cancellation or a\nlost reply reports an unknown outcome unless an attributed refusal proves otherwise. A received\nrefusal remains a refusal even when stop races it. Local failures do not invent responder identities.\nThe follower owns its subscription, timers and read cancellation signal. Reconciliation begins\nbefore the wait deadline, and late read completions cannot settle an expired observation. Its read\ncallback receives the accepting caller triple, remaining budget and abort signal; borrowed bearer\ncommands and control connections use that signal. An in-flight dial that finishes after cancellation\ncloses without publishing. A local reply-subscription failure prevents publication and is observed\nby the same request promise, including when the transport is closing or draining. Request\ncancellation does not revoke or resubmit the accepted operation.\n\n\"Only one manager per space\" is not the current invariant. A split topology that keeps the\nbroker host manager-free is still a topology choice: `cotal up` on that host starts a\nmanager you then stop with `cotal down manager` after `✓ manager up` in\n`.cotal/manager.<spaceKey>.log` (detach stdout listing `manager` is pidfile liveness, not a\nteardown boundary), and `cotal supervise\n--server` runs the manager elsewhere ([Run a mesh](run-a-mesh.md)). Extra live managers\nare addressable, not an error.\n\nThe reserved `describe` bootstrap is the one request the resolver may repeat while waiting: it is\nread-only, it is re-published under the same request binding, and every attempt stays inside the\noriginal deadline. This covers the startup window where Core NATS discards the first request before\nthe manager has subscribed. If the connection closes while the resolver waits, the describe fails\ncleanly instead of throwing from the retry timer. The resolved command is never repeated by this\nreadiness behavior.\n\nThe resolve and the invoke are separate trips through the same anycast queue, so in a\nmulti-manager space an unpinned call can land on an instance the caller did not resolve. Every\ncall carries the incarnation it resolved against, and a manager that is not that incarnation\n**refuses before running the command**, so the failure an operator sees says the command did\nnot run, and re-issuing it cannot duplicate the effect. That is the difference that matters for\na mutation: the older behaviour detected the mismatch on the reply, after the manager had\nalready acted, and could only tell you to go and check. `--on` still matters for reaching a\nspecific manager (`ps`, `stop`, `attach`, `spawn --detach`), but it is no longer what stands\nbetween a split and a duplicated spawn. Against a manager older than this fence the refusal is\nstill after the fact, and its message says so. The re-issue is automatic only when the refusal\nstates `not-executed` in its `outcome` field; a refusal that omits the field, or states\n`unknown`, is surfaced to the caller instead of repaired, because neither proves the command did\nnot run. The CLI's manager commands, `cotal invoke`, the `cotal run` verbs and the manager row of\n`cotal status` re-describe and re-issue an unpinned call after each such refusal, up to 16 times,\nso a split reaches the operator only when every attempt split. A hosted run's own manager calls\nuse the same bound. A pinned call is never re-issued. An agent's own manager\ntools, such as `cotal_spawn` and `cotal_despawn`, re-describe and re-issue with the same bound,\nincluding the goal-result read that follows a spawn to its outcome.\n\nAn unpinned targeted call, such as `cotal_despawn` or a hosted run's turn relay, can also reach a\nmanager that does not host its target, because each manager resolves targets against the agents\nit runs. That manager refuses with `expired` and `not-executed` and says it holds no mapping for\nthe target, and the same re-issue repairs it within the same bound. An agent that no manager hosts\nstill ends in that refusal once the re-issues run out. A pinned call gets the refusal of the\ninstance it named.\n\nA manager whose boot inventory marked every declared connector unavailable does not subscribe\n`spawn` or `launch` on the class `one` rail. Those commands stay on scatter and on this\ninstance's `inst` rail, so a sibling that can launch them can take an unpinned spawn, and a\ncaller that pins this instance with `--on` still gets a named harness refusal. `describe`\nstill lists the commands: the instance rail serves them, and `describe` itself stays on the\nclass rail (SPEC 13.7). An unpinned `spawn` can therefore bind-fence: `describe` may land on\nthe skip member while `spawn` lands on a sibling, the command was not run, and the caller\nre-issues or pins `--on`. `status` reports `classSpawn: false` when that skip is in effect.\nA manager that can launch some connectors keeps the class rail. If the queue hands it a\nharness its inventory marked unavailable, the refusal names `--on` because the standing serve\ncredential cannot read sibling inventories. Pin the capable instance (the whole id, as `ps`\nprints it).\n\n`ps` and\n`status` become a **scatter** across every registered instance: the caller freezes the\nexpected set from the service registry, invokes each under a shared deadline, and merges the\nresults with per-instance attribution. A non-answering instance is labelled as registered\nwith no answer within the deadline, never silently omitted. See [SPEC §13.5](../SPEC.md#135-verbs) (scatter) and [cli.md](cli.md).\n\nThe expected set comes from the **registry**, which records registration rather than liveness.\nAn instance that crashes never deregisters, so it stays in the set and the gather has nothing\nleft to wait for but an answer that cannot come. It pays the whole deadline, on every scatter,\nindefinitely. A scatter can therefore be given a per-instance liveness probe: when the broker\nitself reports that an instance holds no subscription on its own instance rail, the gather stops\nwaiting for it. Only that affirmative report counts. A lapsed presence entry, a probe that timed\nout, and a probe that failed are all *absence of evidence*, and treating any of them as death\nwould turn a slow correct answer into a fast wrong one, so they leave the full deadline standing.\nNothing about the outcome changes either way: an instance that did not answer is still\nunreachable, still surfaced, and the scatter is still not complete.\n\nThe probe is supplied by the **caller**, not invented by the scatter. Asking about an instance is\na publish on that instance's rail, and a credential that holds no row for it is refused by the\nbroker asynchronously, while the publish itself returns normally. The probe verb watches for that\nrefusal and raises it as `permission-denied` naming the rail, so it is never mistaken for a quiet\ninstance, and it never burns the probe budget waiting out a refusal. Only the layer that\nminted the credential knows which ids it may ask about, so that layer asks about those and no\nothers. `cotal ps` freezes the class on its first connection, re-mints an instrument pinned only\nto the frozen ids, and scatters on a second; a refusal the broker raises anyway is printed and\nthe instance's row says the probe was refused, which is a fact about the credential, not about\nthe instance.\n\nThis does not help against an instance that is **connected but not answering**. A hung manager\nholds its subscriptions, so it is indistinguishable from a slow one, and it still costs the full\ndeadline. That is the correct result, not a gap in the probe.\n\n### Deregistration\n\nA probe makes a dead registration cheap to skip; it does not remove it. Removal is the\nregistration's own exit, and there are two explicit routes to it\n([SPEC §13.5](../SPEC.md#135-verbs): a deleted `svc` spec *is* the deregistration).\n\nA manager that stops cleanly removes its own registration, so an ordinary shutdown leaves no stale\nrow. The delete is pinned to the registration revision that process wrote. When a successor has\nregistered the same instance since then, the stop logs that and leaves the successor's registration\nalone. It refuses that delete while this instance holds the endpoint governance slot at the live\nissuance-gate generation (a registration still completing its reopen). A leftover slot whose\ngeneration is behind that live generation is not in-flight and does not block the stop. A manager\nthat cannot renew or read its lease keeps serving, stays registered, and retries. If another process\nholds the same instance key, that process has taken the instance over, so this one logs the conflict\nand exits without deregistering, leaving the successor's registration alone.\n\nA restart that died *mid-registration* is a different residue: the issuance gate stays frozen under\nthat op. The successor completes the dead registration on boot when the freeze-holder is\naffirmatively gone under a complete CONNZ sweep (the same composition as\n[`cotal reconcile-gate`](cli.md#reconcile-gate)). A committed spec write is finished under that\nsame freeze; only a definite no-commit abort-reopens and then runs the normal takeover.\nIt does not invent a TTL and it does not start a new freeze over a still-held one.\n\nThat residue has a second half, and it is the endpoint governance slot rather than the gate. Every\nregistration takes the endpoint-wide slot before it publishes its spec and holds it until its own\ngate reopens, which is what serializes registration for the endpoint. An instance that stopped\nbetween those two points leaves the slot held with no registration behind it, so the endpoint\nrefuses new registrations while nothing is actually in flight. The slot is stamped with the\ngeneration of the gate its holder had frozen when it took it, and a slot is promoted only at that\nsame generation. So once the holder's gate has reopened past the stamp, the slot can never be\npromoted by anyone, and the next registration for that endpoint replaces it. That reclaim is part of\nan ordinary start and needs no operator step.\n\nA slot whose holder's gate is still at the stamped generation is a registration that is genuinely in\nflight, and it keeps refusing. The two states read differently only in the holder's gate coordinate,\nso reopening that gate is what separates them: the holder's own restart heals it on boot, and\n[`cotal reconcile-gate`](cli.md#reconcile-gate) is the operator's route when the boot path cannot\nrun. The registration path is the slot's only writer, and neither repair command writes it.\nA registration that cannot read the holder's gate at all refuses, because an unreadable gate does\nnot distinguish the two states either. Each of these refusals carries\n`kind = ai.cotal.ep.foreign-slot-held` in `error.details[]` with the holder's instance id and the\n`condition` that refused: `in-flight` for a holder gate still at the stamp, or `no-seam`,\n`unreadable`, `garbled` or `behind` when the registration could not read that gate or read it below\nthe stamp. A remote manager asks its host to reconcile the holder only on `in-flight`, the one\ncondition a gate repair can clear.\n\nFor the instance that cannot cooperate, an operator names it:\n`cotal deregister-instance --instance <id>` ([cli.md](cli.md#deregister-instance)). It removes the\nrecord only on the same evidence `cotal ps` acts on: the broker reporting nothing subscribed on\nthat instance's own rail. It refuses if the instance answers a describe, refuses if the probe could\nnot run at all, and refuses if the instance is merely quiet, because a hung process still holds its\nsubscriptions and is therefore not affirmed gone. It also refuses while that instance holds the\nendpoint governance slot at the live issuance-gate generation (a registration still completing);\na leftover slot behind that generation is not in-flight and does not block. Nothing sweeps the\nregistry on an age threshold or on silence.\nAn instance that is deregistered while it is merely wedged re-registers over the tombstone on its\nnext start, which is what makes the operator's decision a recoverable one.\n\n## Attach sessions\n\n`cotal attach` no longer returns a `ws://127.0.0.1` URL. It creates a one-use, holder-bound\nsession offer: the manager mints a token bound to the caller, the target lifecycle, its own\ninstance id and epoch, and an expiry, and replies with a session id and expiry only, no URL\nand no secret in the reply. The CLI redeems the offer over the mesh (a second redeem is\nrefused). On a registered open mesh that redeem is a bare connection, the same path other\ncontrol commands already use; on a static-auth mesh it is still a session-caller credential\nminted from the resolved root's seed. On a user-auth mesh the CLI holds no seed: it exchanges its\nlogin and the grant for a `session-caller` view bearer, and the callout mints the same caller rails\nwith the grant's expiry. Terminal bytes then stream on core-NATS session subjects\nscoped to the two parties. Backpressure is a bounded in-flight window with an explicit drop notice, never\nsilent loss; a late attach still repaints the full screen from a replayed terminal\nsnapshot. Close, expiry, target despawn, and a manager restart are distinct, surfaced end\nstates: a restarted manager's successor refuses the old epoch's sessions and the client\nshows \"manager restarted; re-attach\".\n\n## Seat input\n\n`attach` is a stream, so it is the wrong shape for a program that wants to send one line: it\nholds a session open and expects a terminal at the caller's end. The `input` command is the\nother half. One authorized call writes text into a running seat's terminal as if it had been\ntyped there, and answers with the seat and the number of bytes delivered.\n\nIt exists for **harness commands**. A line beginning with `/` (`/compact`, `/clear`, `/model`)\nis neither chat nor an event: the agent's own harness handles it, and the keyboard is the only\nway in. An external control surface that can already read a seat's turns and talk to it still\ncannot drive it without this.\n\nThe op is targeted, rides the `manager.lifecycle` capability, and declares authz modes `owner`\nand `any`, the row shape `attach` and `despawn` already carry, checked by the same authorization.\nEnter is appended unless the caller suppresses it, and nothing is echoed back, since the resulting\nturns already have somewhere to go.\n\n**Who may call it is narrower than either of those**, and the reasoning is worth stating because\nthe natural assumption is wrong. `despawn` and `attach` are granted to anything holding `spawn`;\n`input` is granted only to operator credentials. The tempting argument for treating them alike is\nthat an attach session's `write` already reaches the same terminal, so `input` adds nothing. It\ndoes not reach it: an attach yields a signed session offer, and redeeming one needs a per-session\ncredential minted from the space signing seed, which no agent holds. So `input` would be new\nauthority, and the own-owner rule that bounds `despawn` covers every seat under an owner rather\nthan only the ones a caller launched. Killing a peer is denial; typing into a peer is control of\nit. The write therefore sits with the credential that is already the administrative authority for\nthe domain.\n\nOnly a runtime that owns the child's input stream can serve it. The `pty` runtime does; the\nexternal terminal runtimes attach to a process they do not own, and there the command refuses\nand names the runtime rather than dropping the keystroke. A seat that is not running refuses for\nits own reason, and the two are distinguishable, so a caller can tell \"this will never work\"\nfrom \"not right now\". See [cli.md](cli.md#input).\n\n## Grants\n\nThere is no broad control credential. A caller holds one capability row per command it is\nallowed to send, and minting maps each named capability to the request subjects it needs and no\nothers. The manager serves over a scoped serve credential that can answer and\nreply but cannot, for instance, write another endpoint's records or forge a goal terminal;\nthe goal-fact writer and the session writer are separate, narrowly scoped credentials the\nbroker fences by subject. Authorization is checked at the serving boundary, and for actions\nit linearises at acceptance: a spawn refused there mints no reservation and leaves no\nprocess. See [SPEC §13.9](../SPEC.md#139-authority-boundary) and\n[identity & auth](identity-and-auth.md).\n\nA carried resume transcript never rides the rails. The operator-only `transcript-receive` command\nanswers whether to upload and hands back a one-time claim for `spawn`, and the bytes travel through\nthe target instance's own transfer bucket under two one-shot credentials: a writer the operator\nmints for that one transcript, and a reader the target instance mints for its own bucket, or that\nthe host issues a remote manager through its `transferReader` authority operation.\n\n## See also\n\n- [Architecture](architecture.md), where the manager and the wire fit in the whole system.\n- [CLI](cli.md), for `describe`, `invoke`, `spawn`, `ps`, `status`, `attach`, and `input`.\n- [SPEC §13](../SPEC.md#13-endpoint-control-surface-v04), the normative contract.\n"
|
|
139
|
+
"body": "# The control surface\n\n> **Concept** (informative) · **For:** operators and client authors who want to know how the manager and other daemons are driven · **Normative:** [SPEC §13](../SPEC.md#13-endpoint-control-surface-v04)\n\nCotal once had a privileged control rail: a fixed set of named service tiers\n(`self` / `manager` / `admin` / `delivery`) on their own `ctl.*` subjects, with the manager\nas a special case the broker recognised by name. That rail is gone. Everything that serves\nstructured commands now, the manager, the delivery daemon, a wrapped MCP server, a\nthird-party service, is an ordinary **endpoint**: a daemon that registers a service\nidentity, publishes its contracts, and answers `describe`. `manager` is an endpoint name\nlike any other; no subject, envelope, or grant in this surface knows it specially. The\nmanager is a service on the mesh, not an authority over it: it holds only the capability\nrows its callers grant it, and serves over a scoped credential.\n\n## The `ep` rails\n\nOne kind, `ep`, carries every request under a mode token that says where the request\nroutes, never which verb it is (the verb rides the envelope): `one` (queue-group\nanycast, one and only one class member), `all` (scatter, every instance), and `inst` (one instance by its\nstable address). Replies come back on a `reply` rail keyed to the serving instance and its\nepoch. Around these sit the sibling planes the composites use: per-goal events, timers,\nsessions, and the journal that holds durable facts. Every request carries the caller as\nthree forge-locked tokens, `owner`, `actor`, and lifecycle `uid`, plus an unguessable\nnonce, so the broker polices who is calling in the subject grammar itself. See\n[SPEC §13.2](../SPEC.md#132-grammar) for the grammar and [§13.5](../SPEC.md#135-verbs) for\nthe verbs (`call`, `cast`, `watch`, `claim`, `scatter`).\n\n## Lifecycle identity\n\nA principal `owner.actor` is a reusable routing alias: a despawn frees the actor name and a\nlater spawn may legitimately reuse it, so the alias alone is never authority. Two further\ncoordinates make an identity durable: a **lifecycle uid**, an unguessable, never-reused id\nfor one managed lifecycle under a principal, and a **process epoch**, the fenced ownership\nepoch of the process currently animating it, advanced on every restart or takeover. At most\none live epoch owns an identity, and a superseded epoch must stop serving. Durables and\ncredentials key on the lifecycle uid, not the reusable name, which is what lets a\nsupervised restart recover the same lifecycle instead of minting a new one. See\n[SPEC §13.1](../SPEC.md#131-lifecycle-identity) and [identity & auth](identity-and-auth.md).\n\n## Service discovery\n\nNo client has compile-time knowledge of any endpoint's commands. `cotal describe\n<endpoint>` resolves a registered endpoint's command set off the wire: the reserved\n`describe` command answers the registered contract digests, the schemas are fetched from the\nspace's content-addressed contract store, recompiled, and verified against those digests.\nEach command prints with its capability class and targeting shape. `cotal invoke <endpoint>\n<command> --args '<json>'` then calls one command by name, validating the arguments as they\nwill be sent (JSON drops a key whose value is undefined) against the fetched input schema before\npublish. A refusal at that check means nothing was sent, and it is marked `not-executed`. A\nsigned-in user invokes the same surface through their bearer, and the broker enforces each\ncommand's existing capability grant. A manager alias supplied through `--name` resolves through\nits name-keyed `inspect` command, so an authorized targeted call does not need the manager-wide\n`ps` enumeration grant. Every built-in manager command uses this\nsame trust chain, so there is nothing the built-ins can reach that a described contract cannot. The registered\n`auth` endpoint is describable the same way `manager` is: `cotal describe auth` lists\n`retire-lifecycle` and its exact-mode target shape.\nSee [SPEC §13.7](../SPEC.md#137-contracts-and-discovery) and [cli.md](cli.md).\n\nThe manager's `resolve-cwd` command is in the `manager.spawn` capability class. It accepts an\nabsolute path on that manager's host and returns its canonical directory plus the host name. It\nrefuses a relative, missing or non-directory path with `failed-precondition`; it creates nothing.\n`spawn` applies the same check at admission, before any credentials or durables are minted.\n\n### Inspecting a managed name\n\nManager `inspect` keeps its successful response as the live managed-agent row. A live hit does\nnot read durable lifecycle state, so a temporary records-store failure cannot break inspection of\nan agent the manager currently holds.\n\nOn a live miss, a static manager point-reads its durable slot row. A name with no slot, or a slot\nwhose phase is `retired`, remains `not-found`. A nonterminal slot returns\n`failed-precondition` with `error.details[].kind =\nai.cotal.manager.static-slot-observation`. The detail carries the slot's `slotPhase`,\n`owner`, `actor`, `slotLifecycleUid`, `cleanupComplete` when recorded, and `slotRevision`. It\nthen carries the separate lifecycle head's `headState`, `headOp` when present,\n`headLifecycleUid`, and `headRevision`. Head fields are absent when provisioning has not written\nthe lifecycle head yet.\nThe error message carries the same diagnostic summary so string-only operator paths do not hide\nthe structured detail.\n\nA slot row records the manager instance that owns it. In a space with more than one manager, the\nclass queue can hand `inspect` to an instance that does not host the name. When the row names a\ndifferent instance and is not `retired`, the miss returns `failed-precondition` with the same\ndetail plus `ownerInstanceId`, which names the only manager that can act on it. The message names\nboth instances, so a caller that reads only the string can tell it from `not-found`. A sibling's\n`retired` row remains `not-found`. A named `cotal_despawn` resolves its target through this read\nand cannot address an instance, so it asks again until the owning instance answers, up to 16\ntimes.\n\nThe slot is read before the head. These records do not form one atomic snapshot, so the detail\nalso carries `readOrder: [\"slot\", \"head\"]` and `consistency: \"ordered-not-atomic\"`. A head can\nadvance between the reads. The issuance gate is not projected because the retirement operation\nneeded for this diagnosis is already recorded on the head, and reading a third record would add\nanother non-atomic edge without changing the per-name result.\n\nIf either durable read fails or exceeds its bound, the miss returns `unavailable` with\n`ai.cotal.manager.static-slot-read-failed` rather than claiming the name is absent. That detail\nnames the inspected `name`, the failed `record` (`slot`, `head`, or `slot-or-head` when the layer\ncannot distinguish them), and `operation: \"read\"`. User-auth managers do not own `mgrslot` rows,\nso their inspect misses remain live-map reads.\n\nFor Linux custodied seats, retirement requires the runtime's process-exit evidence before\nfreeing the alias or deleting its credentials and delivery state. Socket loss alone is not\nproof of exit. The runtime retains the record captured at launch or adoption so a clean\ncustodian exit can unlink its file without losing the recorded boot and process identities.\nIf the file is missing, reaping uses that retained record and the existing kernel identity\nchecks. An unknown reference without either record refuses cleanup. Reused process ids\nare never signalled on the strength of the old record.\n\n### Listing the durable slots\n\nManager `slots` (`manager.read`, untargeted) lists the durable static slot rows this manager\nowns. Only static managers hold these rows: a user-mode or open manager answers\n`failed-precondition`, and a manager whose durable store is not standing answers `unavailable`.\nEach row carries the same `readOrder` and `consistency` fields `inspect` uses, because the list\nis read the same way: torn across rows as well as within each row's slot/head pair. A `retired`\nrow is never listed. `live` reflects the manager's live roster at render time, not the durable\nrow. A slot row is never deleted, so a row whose latest operation is a DEL or PURGE marker is\ncorruption: the list answers `unavailable` naming that row, as `inspect` does for its name.\n\n## Spawn is a goal\n\nLong-running commands are **actions** ([SPEC §13.6](../SPEC.md#136-composites)): the caller\nsubmits with a client-generated `goalId` and a request fingerprint, the endpoint records a\ndurable accept or reject decision, progress rides per-goal events, and the work ends in one\nterminal outcome (`succeeded`, `failed`, `cancelled`, `expired`, or `uncertain`). Spawn is\nthe reference case. Rather than block the caller for up to 30 seconds while an agent comes\nup, the manager accepts the goal and returns the allocated identity at once:\n\n```json\n{\n \"name\": \"reviewer_2\",\n \"owner\": \"u_...\", \"actor\": \"reviewer_2\", \"uid\": \"...\",\n \"goalId\": \"...\", \"fingerprint\": \"...\",\n \"readinessDeadlineMs\": 30000,\n \"executor\": { \"lifecycleUid\": \"...\", \"epoch\": 3 }\n}\n```\n\nThe `uid` is the lifecycle the agent runs at. On a participant manager whose host enrolls its\nagents, the host picks that uid, so the manager accepts the goal only after the host has answered.\nA host refusal there refuses the spawn, and no goal is bound.\n\nThe name is the one actually allocated: a persona-derived collision is auto-numbered\n(`reviewer`, then `reviewer_2`), while a hard-pinned `--name` that collides with a live\nagent is refused at accept, before anything is minted. Auto-numbering never hands out a numbered\nname it has already issued in that manager process, even after the agent holding it is gone, so\na collision takes the next number. Only numbering consults that history: a hard-pinned `--name`,\nor a persona whose own name is a numbered string, takes that name whenever it is free, and\nnumbering does not skip a string such a spawn held before. The triple plus `goalId` let the\ncaller follow progress (connector handoff, process launched, presence join) and reconcile\nlater against the exact instance that accepted. Presence within the manager's default\n30-second readiness window, or a connector's declared bounded window, settles the goal\n`succeeded`; an early process exit is `failed`; the window passing with neither is `uncertain`,\na bounded, durable outcome that a later `ps` or status read settles against the live roster.\n`uncertain` is a real terminal outcome, not an absence and not a silent hang. It carries the\ndiagnosis of whoever owned the deadline: for a launch that\nnames the agent and says to inspect it rather than re-issue, since re-issuing after a launch\nthat in fact succeeded mints a duplicate. A follower keeps the acceptance as the data of any\nterminal other than `succeeded`, so `cotal_spawn` returns an uncertain launch as a pending result\ninstead of an error: it names the allocated agent, its id, and its manager, and tells the calling\nagent to watch the roster. A committer that supplies no diagnosis falls back to\n\"the success signal did not arrive within the readiness deadline\". The agent's own eventual\nstate is then observable on its presence record.\n\nThe acceptance carries that exact `readinessDeadlineMs`. A synchronous follower treats its own\nrequest deadline as a floor and waits through the accepted readiness budget plus delivery margin,\nso a connector-specific slow boot cannot be reported as a caller timeout while the manager is\nstill legitimately waiting for its terminal.\n\nA spawn that is **refused** because a lifecycle barrier already holds the actor (a frozen\nissuance gate, a retiring alias, a retired uid) is not a wait-timeout. The manager already\nknows the blocked op (`registration` / `retirement` / `activation` / `takeover`), the `opId`\nholding it, and the remedy when one exists (`retry`, `cotal reconcile-gate`). The detail\ncarries `headState` (`active` / `retiring` / `retired`) only when the refusing site read the\nlifecycle head, and `gateState` (`frozen` / `retired`) only when it read the issuance gate. A\ngate frozen by a takeover or a registration says nothing about the head, so that refusal\ncarries `gateState=frozen` and no `headState`. Those facts ride `error.details[]` as\n`kind = ai.cotal.ep.lifecycle-blocked` and are also appended to the error string, so a\ncaller that only prints `error.message` still sees them. The CLI and the connector tools hand a\nrefusal on in one shape, so `cotal spawn -f` keeps the same code, details, rendered facts and\nacceptance data as `cotal spawn --detach` and `cotal_spawn`. A connector that collapses the\nrefusal to \"startup failed (unknown)\" or a SPEC 13.6 wait-timeout is hiding a knowable\nstate, not reporting a missing one.\n\n## Instance routing\n\nA space can run more than one manager. Each manager persists a stable logical instance id\nacross restarts and advances its process epoch when it comes back, so callers address a\nspecific manager without caring which process currently serves it. A start serves only at the\nepoch its own registration committed, never at the epoch of a later start of the same instance.\nOn a static or open mesh,\nan untargeted spawn rides class anycast (any manager may accept, and the acceptance records which one did).\n`cotal spawn <persona> --detach --on <instance>` and `cotal_spawn(instance: \"<instance>\")`\npin one instance by its exact id. A foreground CLI spawn has no manager to pin and refuses the\nflag. An MCP pin that does not resolve is refused without falling back to class anycast. There are no ordinal\naliases and no short forms: wherever a display names an instance you can address, it prints\nthe whole id, because both surfaces take nothing else.\n\nOn a user-auth mesh, manager commands obtain a short-lived `manager-caller` view from the\nexchange. It authorizes one concrete manager instance using the caller's current actor grant and\nthe host's registered service records. Discovery and invocation both use that instance's `inst`\nroute. This view grants no registry scan, class queue, or additional command capability. An absent,\nambiguous or unauthorized selection refuses before the command is sent.\n\nManaged launches carry `COTAL_MANAGER_INSTANCE` so their tools address the manager that launched\nthem. Existing unbound sessions can use the exchange's unique authorized selection without\nreplacing their actor or conversation. The connector uses a separate control connection; the\nstanding message connection and its credential source are unchanged. Accepted spawn goals are\nfollowed on that renewing connection, using its existing caller-scoped progress grant, so a long\nreadiness budget does not depend on the short-lived control credential. The follower confirms its\nprogress subscription with the broker before submitting on the separate connection. A caller still\nchecks the resolved instance and epoch, and never retries an ambiguous mutation outcome.\n\nThe manager's `goal-result` command accepts `{goalId}` and returns `{goalId, result?}`. It reads\nonly the authenticated caller's owner, actor and lifecycle through the manager's separate trusted\ngoal-writer connection. The caller receives an attributed reply, never a raw JetStream reader\ngrant. Each read is admitted by the connection's broker-enforced command grant. A live user-auth\nconnection remains bounded by its bearer expiry after revocation; a renewed connection is checked\nagainst fresh authority. There is no separate per-read ledger check. An absent `result` means no\nterminal is recorded; it does not prove the goal is running or permit another submission. The\nexisting trusted goal-writer's leader-served EPF read is space-wide at the broker; the handler\nconfines it to this endpoint and caller triple.\n\nThe manager's reserved `cancel` command accepts `{goalId, mode?}` and returns `{goalId, state}`.\nIt is served for a turn the manager relays. The goal is the authenticated caller's own, so a caller\nwithdraws only a turn it submitted. The turn ends `cancelled`, its seat is not shown it again, and a\nlater yield of it is answered with that terminal. A goal that already ended is refused\n`failed-precondition` with its cached outcome attached, and a goal this manager does not relay is\nrefused without being changed. So is a second cancel that arrives while a first is still ending the\nturn; a first that fails leaves the turn pending unless something ended it meanwhile. A workflow\nrun sends it for the turn, ask attempt or escalation of a branch it cancelled.\n\nA followed mutation requires a manager whose attributed describe includes `goal-result`. Update\nthe manager, issuer and client together before using that recovery path. Reloading an issuer alone\ncannot change an already-running participant manager. Recovery re-resolves the accepting instance's\nepoch, preserves the caller lifecycle and validates the result against the accepted goal and any\nacceptance fingerprint. Stopping the caller ends its observation, not the already accepted goal.\n\nA followed call resolves the endpoint within its deadline before the submission starts, so a\nrefused or unanswered describe surfaces as its own error. A describe or command publish that the\nbroker refuses reports `not-executed`.\nCancellation before submission reports `not-executed`. Once submission starts, cancellation or a\nlost reply reports an unknown outcome unless an attributed refusal proves otherwise. A received\nrefusal remains a refusal even when stop races it. Local failures do not invent responder identities.\nThe follower owns its subscription, timers and read cancellation signal. Reconciliation begins\nbefore the wait deadline, and late read completions cannot settle an expired observation. Its read\ncallback receives the accepting caller triple, remaining budget and abort signal; borrowed bearer\ncommands and control connections use that signal. An in-flight dial that finishes after cancellation\ncloses without publishing. A local reply-subscription failure prevents publication and is observed\nby the same request promise, including when the transport is closing or draining. Request\ncancellation does not revoke or resubmit the accepted operation.\n\n\"Only one manager per space\" is not the current invariant. A split topology that keeps the\nbroker host manager-free is still a topology choice: `cotal up` on that host starts a\nmanager you then stop with `cotal down manager` after `✓ manager up` in\n`.cotal/manager.<spaceKey>.log` (detach stdout listing `manager` is pidfile liveness, not a\nteardown boundary), and `cotal supervise\n--server` runs the manager elsewhere ([Run a mesh](run-a-mesh.md)). Extra live managers\nare addressable, not an error.\n\nThe reserved `describe` bootstrap is the one request the resolver may repeat while waiting: it is\nread-only, it is re-published under the same request binding, and every attempt stays inside the\noriginal deadline. This covers the startup window where Core NATS discards the first request before\nthe manager has subscribed. If the connection closes while the resolver waits, the describe fails\ncleanly instead of throwing from the retry timer. The resolved command is never repeated by this\nreadiness behavior.\n\nThe resolve and the invoke are separate trips through the same anycast queue, so in a\nmulti-manager space an unpinned call can land on an instance the caller did not resolve. Every\ncall carries the incarnation it resolved against, and a manager that is not that incarnation\n**refuses before running the command**, so the failure an operator sees says the command did\nnot run, and re-issuing it cannot duplicate the effect. That is the difference that matters for\na mutation: the older behaviour detected the mismatch on the reply, after the manager had\nalready acted, and could only tell you to go and check. `--on` still matters for reaching a\nspecific manager (`ps`, `stop`, `attach`, `spawn --detach`), but it is no longer what stands\nbetween a split and a duplicated spawn. Against a manager older than this fence the refusal is\nstill after the fact, and its message says so. The re-issue is automatic only when the refusal\nstates `not-executed` in its `outcome` field; a refusal that omits the field, or states\n`unknown`, is surfaced to the caller instead of repaired, because neither proves the command did\nnot run. The CLI's manager commands, `cotal invoke`, the `cotal run` verbs and the manager row of\n`cotal status` re-describe and re-issue an unpinned call after each such refusal, up to 16 times,\nso a split reaches the operator only when every attempt split. A hosted run's own manager calls\nuse the same bound. A pinned call is never re-issued. An agent's own manager\ntools, such as `cotal_spawn` and `cotal_despawn`, re-describe and re-issue with the same bound,\nincluding the goal-result read that follows a spawn to its outcome.\n\nAn unpinned targeted call, such as `cotal_despawn` or a hosted run's turn relay, can also reach a\nmanager that does not host its target, because each manager resolves targets against the agents\nit runs. That manager refuses with `expired` and `not-executed` and says it holds no mapping for\nthe target, and the same re-issue repairs it within the same bound. An agent that no manager hosts\nstill ends in that refusal once the re-issues run out. A pinned call gets the refusal of the\ninstance it named.\n\nA manager whose boot inventory marked every declared connector unavailable does not subscribe\n`spawn` or `launch` on the class `one` rail. Those commands stay on scatter and on this\ninstance's `inst` rail, so a sibling that can launch them can take an unpinned spawn, and a\ncaller that pins this instance with `--on` still gets a named harness refusal. `describe`\nstill lists the commands: the instance rail serves them, and `describe` itself stays on the\nclass rail (SPEC 13.7). An unpinned `spawn` can therefore bind-fence: `describe` may land on\nthe skip member while `spawn` lands on a sibling, the command was not run, and the caller\nre-issues or pins `--on`. `status` reports `classSpawn: false` when that skip is in effect.\nA manager that can launch some connectors keeps the class rail. If the queue hands it a\nharness its inventory marked unavailable, the refusal names `--on` because the standing serve\ncredential cannot read sibling inventories. Pin the capable instance (the whole id, as `ps`\nprints it).\n\n`ps` and\n`status` become a **scatter** across every registered instance: the caller freezes the\nexpected set from the service registry, invokes each under a shared deadline, and merges the\nresults with per-instance attribution. A non-answering instance is labelled as registered\nwith no answer within the deadline, never silently omitted. See [SPEC §13.5](../SPEC.md#135-verbs) (scatter) and [cli.md](cli.md).\n\nThe expected set comes from the **registry**, which records registration rather than liveness.\nAn instance that crashes never deregisters, so it stays in the set and the gather has nothing\nleft to wait for but an answer that cannot come. It pays the whole deadline, on every scatter,\nindefinitely. A scatter can therefore be given a per-instance liveness probe: when the broker\nitself reports that an instance holds no subscription on its own instance rail, the gather stops\nwaiting for it. Only that affirmative report counts. A lapsed presence entry, a probe that timed\nout, and a probe that failed are all *absence of evidence*, and treating any of them as death\nwould turn a slow correct answer into a fast wrong one, so they leave the full deadline standing.\nNothing about the outcome changes either way: an instance that did not answer is still\nunreachable, still surfaced, and the scatter is still not complete.\n\nThe probe is supplied by the **caller**, not invented by the scatter. Asking about an instance is\na publish on that instance's rail, and a credential that holds no row for it is refused by the\nbroker asynchronously, while the publish itself returns normally. The probe verb watches for that\nrefusal and raises it as `permission-denied` naming the rail, so it is never mistaken for a quiet\ninstance, and it never burns the probe budget waiting out a refusal. Only the layer that\nminted the credential knows which ids it may ask about, so that layer asks about those and no\nothers. `cotal ps` freezes the class on its first connection, re-mints an instrument pinned only\nto the frozen ids, and scatters on a second; a refusal the broker raises anyway is printed and\nthe instance's row says the probe was refused, which is a fact about the credential, not about\nthe instance.\n\nThis does not help against an instance that is **connected but not answering**. A hung manager\nholds its subscriptions, so it is indistinguishable from a slow one, and it still costs the full\ndeadline. That is the correct result, not a gap in the probe.\n\n### Deregistration\n\nA probe makes a dead registration cheap to skip; it does not remove it. Removal is the\nregistration's own exit, and there are two explicit routes to it\n([SPEC §13.5](../SPEC.md#135-verbs): a deleted `svc` spec *is* the deregistration).\n\nA manager that stops cleanly removes its own registration, so an ordinary shutdown leaves no stale\nrow. The delete is pinned to the registration revision that process wrote. When a successor has\nregistered the same instance since then, the stop logs that and leaves the successor's registration\nalone. It refuses that delete while this instance holds the endpoint governance slot at the live\nissuance-gate generation (a registration still completing its reopen). A leftover slot whose\ngeneration is behind that live generation is not in-flight and does not block the stop. A manager\nthat cannot renew or read its lease keeps serving, stays registered, and retries. If another process\nholds the same instance key, that process has taken the instance over, so this one logs the conflict\nand exits without deregistering, leaving the successor's registration alone.\n\nA restart that died *mid-registration* is a different residue: the issuance gate stays frozen under\nthat op. The successor completes the dead registration on boot when the freeze-holder is\naffirmatively gone under a complete CONNZ sweep (the same composition as\n[`cotal reconcile-gate`](cli.md#reconcile-gate)). A committed spec write is finished under that\nsame freeze; only a definite no-commit abort-reopens and then runs the normal takeover.\nIt does not invent a TTL and it does not start a new freeze over a still-held one.\n\nThat residue has a second half, and it is the endpoint governance slot rather than the gate. Every\nregistration takes the endpoint-wide slot before it publishes its spec and holds it until its own\ngate reopens, which is what serializes registration for the endpoint. An instance that stopped\nbetween those two points leaves the slot held with no registration behind it, so the endpoint\nrefuses new registrations while nothing is actually in flight. The slot is stamped with the\ngeneration of the gate its holder had frozen when it took it, and a slot is promoted only at that\nsame generation. So once the holder's gate has reopened past the stamp, the slot can never be\npromoted by anyone, and the next registration for that endpoint replaces it. That reclaim is part of\nan ordinary start and needs no operator step.\n\nA slot whose holder's gate is still at the stamped generation is a registration that is genuinely in\nflight, and it keeps refusing. The two states read differently only in the holder's gate coordinate,\nso reopening that gate is what separates them: the holder's own restart heals it on boot, and\n[`cotal reconcile-gate`](cli.md#reconcile-gate) is the operator's route when the boot path cannot\nrun. The registration path is the slot's only writer, and neither repair command writes it.\nA registration that cannot read the holder's gate at all refuses, because an unreadable gate does\nnot distinguish the two states either. Each of these refusals carries\n`kind = ai.cotal.ep.foreign-slot-held` in `error.details[]` with the holder's instance id and the\n`condition` that refused: `in-flight` for a holder gate still at the stamp, or `no-seam`,\n`unreadable`, `garbled` or `behind` when the registration could not read that gate or read it below\nthe stamp. A remote manager asks its host to reconcile the holder only on `in-flight`, the one\ncondition a gate repair can clear.\n\nFor the instance that cannot cooperate, an operator names it:\n`cotal deregister-instance --instance <id>` ([cli.md](cli.md#deregister-instance)). It removes the\nrecord only on the same evidence `cotal ps` acts on: the broker reporting nothing subscribed on\nthat instance's own rail. It refuses if the instance answers a describe, refuses if the probe could\nnot run at all, and refuses if the instance is merely quiet, because a hung process still holds its\nsubscriptions and is therefore not affirmed gone. It also refuses while that instance holds the\nendpoint governance slot at the live issuance-gate generation (a registration still completing);\na leftover slot behind that generation is not in-flight and does not block. Nothing sweeps the\nregistry on an age threshold or on silence.\nAn instance that is deregistered while it is merely wedged re-registers over the tombstone on its\nnext start, which is what makes the operator's decision a recoverable one.\n\n## Attach sessions\n\n`cotal attach` no longer returns a `ws://127.0.0.1` URL. It creates a one-use, holder-bound\nsession offer: the manager mints a token bound to the caller, the target lifecycle, its own\ninstance id and epoch, and an expiry, and replies with a session id and expiry only, no URL\nand no secret in the reply. The CLI redeems the offer over the mesh (a second redeem is\nrefused). On a registered open mesh that redeem is a bare connection, the same path other\ncontrol commands already use; on a static-auth mesh it is still a session-caller credential\nminted from the resolved root's seed. On a user-auth mesh the CLI holds no seed: it exchanges its\nlogin and the grant for a `session-caller` view bearer, and the callout mints the same caller rails\nwith the grant's expiry. Terminal bytes then stream on core-NATS session subjects\nscoped to the two parties. Backpressure is a bounded in-flight window with an explicit drop notice, never\nsilent loss; a late attach still repaints the full screen from a replayed terminal\nsnapshot. Close, expiry, target despawn, and a manager restart are distinct, surfaced end\nstates: a restarted manager's successor refuses the old epoch's sessions and the client\nshows \"manager restarted; re-attach\".\n\n## Seat input\n\n`attach` is a stream, so it is the wrong shape for a program that wants to send one line: it\nholds a session open and expects a terminal at the caller's end. The `input` command is the\nother half. One authorized call writes text into a running seat's terminal as if it had been\ntyped there, and answers with the seat and the number of bytes delivered.\n\nIt exists for **harness commands**. A line beginning with `/` (`/compact`, `/clear`, `/model`)\nis neither chat nor an event: the agent's own harness handles it, and the keyboard is the only\nway in. An external control surface that can already read a seat's turns and talk to it still\ncannot drive it without this.\n\nThe op is targeted, rides the `manager.lifecycle` capability, and declares authz modes `owner`\nand `any`, the row shape `attach` and `despawn` already carry, checked by the same authorization.\nEnter is appended unless the caller suppresses it, and nothing is echoed back, since the resulting\nturns already have somewhere to go.\n\n**Who may call it is narrower than either of those**, and the reasoning is worth stating because\nthe natural assumption is wrong. `despawn` and `attach` are granted to anything holding `spawn`;\n`input` is granted only to operator credentials. The tempting argument for treating them alike is\nthat an attach session's `write` already reaches the same terminal, so `input` adds nothing. It\ndoes not reach it: an attach yields a signed session offer, and redeeming one needs a per-session\ncredential minted from the space signing seed, which no agent holds. So `input` would be new\nauthority, and the own-owner rule that bounds `despawn` covers every seat under an owner rather\nthan only the ones a caller launched. Killing a peer is denial; typing into a peer is control of\nit. The write therefore sits with the credential that is already the administrative authority for\nthe domain.\n\nOnly a runtime that owns the child's input stream can serve it. The `pty` runtime does; the\nexternal terminal runtimes attach to a process they do not own, and there the command refuses\nand names the runtime rather than dropping the keystroke. A seat that is not running refuses for\nits own reason, and the two are distinguishable, so a caller can tell \"this will never work\"\nfrom \"not right now\". See [cli.md](cli.md#input).\n\n## Grants\n\nThere is no broad control credential. A caller holds one capability row per command it is\nallowed to send, and minting maps each named capability to the request subjects it needs and no\nothers. The manager serves over a scoped serve credential that can answer and\nreply but cannot, for instance, write another endpoint's records or forge a goal terminal;\nthe goal-fact writer and the session writer are separate, narrowly scoped credentials the\nbroker fences by subject. Authorization is checked at the serving boundary, and for actions\nit linearises at acceptance: a spawn refused there mints no reservation and leaves no\nprocess. See [SPEC §13.9](../SPEC.md#139-authority-boundary) and\n[identity & auth](identity-and-auth.md).\n\nA carried resume transcript never rides the rails. The operator-only `transcript-receive` command\nanswers whether to upload and hands back a one-time claim for `spawn`, and the bytes travel through\nthe target instance's own transfer bucket under two one-shot credentials: a writer the operator\nmints for that one transcript, and a reader the target instance mints for its own bucket, or that\nthe host issues a remote manager through its `transferReader` authority operation.\n\n## See also\n\n- [Architecture](architecture.md), where the manager and the wire fit in the whole system.\n- [CLI](cli.md), for `describe`, `invoke`, `spawn`, `ps`, `status`, `attach`, and `input`.\n- [SPEC §13](../SPEC.md#13-endpoint-control-surface-v04), the normative contract.\n"
|
|
140
140
|
},
|
|
141
141
|
{
|
|
142
142
|
"slug": "define-a-team",
|
|
@@ -164,7 +164,7 @@ export function loadDocsBundle() {
|
|
|
164
164
|
"title": "Embedding Cotal",
|
|
165
165
|
"kind": "Guide (informative)",
|
|
166
166
|
"summary": "The cotal binary in this repo is one composition root: an operator CLI.",
|
|
167
|
-
"body": "# Embedding Cotal\n\n> **Guide** (informative) · **For:** implementers building a service on top of Cotal · **Prereqs:** [Architecture](architecture.md), [Identity and auth](identity-and-auth.md), [Delivery daemon](delivery-daemon.md)\n\nThe `cotal` binary in this repo is one composition root: an operator CLI. A separate service\n(for example a hosted, multi-tenant Cotal) does not fork this repo. It writes its **own**\ncomposition root that depends on the published `@cotal-ai/*` packages and imports the surfaces it\nwants. `bin/cotal.ts` uses the same composition pattern. This page is the contract for that: what is a real library\nexport you can build against, how to boot the server-side daemons from those exports, and where the\ncurrent export surface stops short of a fully hosted composition.\n\nThis is the \"guarded substrate\" boundary in practice. Nothing here reveals or assumes a specific\nhost; it documents the public seams any embedder composes.\n\n## What you embed\n\nThe supported reference shape here is **one broker operator serving one space** (one tenant: a\ndedicated data account, under an operator that also holds the system account and a quarantined\nauth-callout account) plus three standalone processes. The trust layer itself composes many spaces\nunder one broker operator today (`createBrokerAuth` + `createSpaceAccountAuth` + N-space\n`serverConfig`); what does not exist yet is the per-space **lifecycle** on a shared broker (see\n[Known gaps](#hosted-composition-gaps)). The three processes:\n\n| daemon | package | what it is |\n|---|---|---|\n| auth-service | `@cotal-ai/auth` | the NATS auth callout, the IdP token exchange, and JWKS. Plane 1 to Plane 2. |\n| delivery | `@cotal-ai/delivery` | the Plane-3 durable backstop: fan-out writer plus trusted reader, per space. |\n| supervise | `@cotal-ai/manager` | the per-machine agent lifecycle (spawn/despawn/attach), per space. |\n\n`mint`, `deliver`, and `auth-service` expose their behavior as direct library primitives, and the\nsupported one-space bootstrap below re-composes from exported low-level primitives. `supervise` and\nthe full `up` orchestration are **not** public runners: `up` also does broker bring-up, restore,\nprocess and registry management, and lifecycle work, and `supervise`'s orchestration is private (see\n[Supervisor signing authority](#supervisor-signing-authority)).\n\n## The export surface\n\nEverything below is a real export of a published package, reachable from the package root (each\npackage publishes only `.` via `dist/index.{js,d.ts}` and ships `files: [\"dist\"]`). Type-only names\nare marked; import them with `import type`.\n\n**Daemon runners and lifecycle**\n\n| symbol | package | purpose |\n|---|---|---|\n| `runAuthService(args, store?)` | `@cotal-ai/auth` | boot the auth-service daemon; `store` injects the secret material. |\n| `runDelivery(args, store?)` | `@cotal-ai/delivery` | boot the delivery daemon; `store` injects the scoped `delivery` cred. |\n| `startAuthService(inputs)` | `@cotal-ai/auth` | start one account-scoped auth-service context and return an `AuthServiceHandle` with the loopback `url`, the per-start `cap`, `readiness`, `drain`, and idempotent `close`. `close` rejects when the context did not release its plane claim: the claim row was no longer its own, or the release write failed. The row then stays held, and the next start reclaims it through the liveness oracle. With the optional `publicFace` input it also serves the public exchange face and carries `publicUrl`. With the optional `platformControl` input the handle also has `platformControlAuthority`, the in-process platform control door, `platformControlReadiness`, its read-only readiness read, `observeManagerGate`, the manager gate read a host composing the delegated user intent decisions passes them, and `activateManagedLifecycle`, the activation that host runs at a delegated launch's pinned lifecycle UID. With `platformControl.host` it also has `registerHostIncarnation`, which registers the host process's own endpoint instance and returns the incarnation a delegated execution pins, `observeHostGate`, the read of that endpoint's issuance gate, and `awaitHostFence`, which resolves once a later registration or barrier fences an incarnation. `runAuthService` remains the CLI entry. |\n| `PlatformControlAuthorityRequest`, `PlatformControlInnerRequest`, `PlatformControlAuthorityResult`, `PlatformControlAssignment` *(types)* | `@cotal-ai/core` | the closed envelope, its inner request union, its result and the backend's assignment row for `platformControlAuthority`. `platformControlOwner` in `@cotal-ai/auth` derives the `p_` owner the door issues under. |\n| `startDeliveryService(inputs)` | `@cotal-ai/delivery` | start one account-scoped delivery instance and return a `HostedServiceHandle` with `readiness`, `drain`, and idempotent `close`. The process runner remains the CLI entry. |\n| `deliveryCredsKey(space, composition)`, `membershipRwCredsKey(space, composition)` | `@cotal-ai/workspace` | build the secret-store keys the delivery cred and the membership feed's rw cred are read/re-signed under. Keys are **per-space**: `space.<hex>/<kind>`. A hosted composition passes `{ injected: true }`. |\n| `retireManagerInstanceIdentity(root, space, expected)` | `@cotal-ai/workspace` | remove a persisted manager identity only if its complete instance id and serve identity still match `expected`. Returns `removed` or `absent`; refuses malformed, nonregular, and changed records. `absent` is not proof of ownership or successful teardown. The caller must separately prove stop and retirement ownership before using it. |\n| `DELIVERY_CREDS_KIND`, `MEMBERSHIP_RW_CREDS_KIND` | `@cotal-ai/workspace` | the operator-facing KIND names (`delivery.creds`, `membership-rw.creds`) those keys are built from, and what renewal results report. A kind is **not** a key: putting a cred under the bare kind writes the pre-0.4 flat location, which nothing reads. |\n| `Manager`, `ManagerOptions` *(type)* | `@cotal-ai/manager` | construct and run a supervisor in-process; `ManagerOptions.secretStore` injects the one store it reads/writes every secret through. `ManagerOptions.remoteAuthority` is the hosted manager-service authority bundle, including host-owned release, retained-validation, goal-index, and serve-time admin-authorization callbacks. |\n| `ManagerOptions.pooled` | `@cotal-ai/manager` | require signerless remote authority and an explicit non-custodial runtime before local execution starts. A pooled composition must supply the assigned account key and all-duty renewal callback; the CLI's default remains unchanged. A signed-in human's manager gets that material from `managerServiceAuthority`. A platform-run control manager gets it from `AuthServiceHandle.platformControlAuthority` with the shipped `remoteManagerClient` builders, as the [platform control authority](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/platform-pooled-control-authority.md) design describes. |\n| `createRuntime`, `Runtime` *(type)* | `@cotal-ai/manager` | resolve the spawn backend (pty built in). |\n| `liveKvEntries(kv, filterOrOptions?, options?)`, `LiveKvEntriesOptions` *(type)* | `@cotal-ai/core` | read live KV entries in one finite scan. Pass `{ signal }` as the second argument or after a key filter to cancel. An interrupted scan throws `IncompleteKvScan`; cancellation throws the signal reason, including during an empty-bucket bind. The scan deletes only its owned consumers, including each one nats.js rebuilt from it, after the broker has answered every create those rebuilds sent. A not-found delete counts as gone, a refused one leaves the consumers to broker inactivity expiry (which also covers a crash), and any other delete failure is thrown, as is a not-found delete of a consumer whose create got no reply from the broker before a timeout or a closed connection. A cleanup failure is thrown only when the scan would otherwise return; the scan's own error, its cancellation reason and `IncompleteKvScan` take precedence. |\n\nThe remote manager authority parser accepts `renewStandingBundle` and `renewRunDriver` only with\nan assigned account nkey, the current manager process epoch, and a host-authenticated registration\nproof. A run renewal also names its active holder, takeover, epoch, fencing token, and the two\nexisting nkeys. The host must fresh-check those coordinates against its registration gate and run\njournal before issuing server-selected profiles. A host without that renewal authorization refuses\nthe request. Until the host issuer wires the operations and validates them on real connections,\nthe presence of these types is not an operational pooled renewal guarantee.\n\nWith `renewStandingBundle` configured, the manager renews all five standing credentials together.\nIt checks every returned credential for the held nkey and assigned account, test-connects each one,\nand adopts them only if the serve epoch has not moved. A refused or failed candidate leaves the\ncurrent credentials in place and records the refusal as cleanup debt until a later renewal succeeds.\nOn shutdown, after the standing context is drained, an expired maintenance executor is renewed\nthrough the existing scoped host operation so deregistration can finish without restarting duties.\n\n**Provisioning and minting** (all `@cotal-ai/core`)\n\n| symbol | purpose |\n|---|---|\n| `createBrokerAuth(label)` | mint BROKER trust: the operator and system account one nats-server trusts. One per broker, shared by every space on it. |\n| `createSpaceAccountAuth(broker, space)` | mint one space's own data account, signed by that broker's operator: the add-a-tenant primitive. |\n| `createSpaceAuth(space)` | the one-space convenience: broker trust + one account in a single composed bundle. |\n| `setupSpaceStreams({ servers, space, creds })` | create the space's JetStream streams. |\n| `ensureDefaultDeliveryClass({ servers, space, creds?, deliveryClass })` | write the space's default delivery class at creation so it is wire-discoverable (SPEC section 4). |\n| `serverConfig(broker, spaces, { storeDir, maxFileStore?, extraAccounts?, port?, host? })` | render the broker config: one operator, N space accounts. `storeDir` is required, `maxFileStore` caps JetStream file storage in bytes (omitted, nats-server's dynamic default applies), and `extraAccounts` preloads the auth-callout account. |\n| `mintCreds(auth, identity, profile, opts?)` | mint a scoped cred for any `Profile`. |\n| `mintMembershipObserverCreds`, `mintConnectionEvictorCreds` | mint the membership/eviction scoped creds. |\n| `provisionAgent`, `provisionAgentDurables` | create a principal's bind-only durables. |\n| `newIdentity`, `stripSpaceAuth` | a fresh nkey identity; a stripped signer bundle (data signing seed only). |\n| `Profile`, `CredentialKind`, `MintOpts`, `SpaceAuth` *(types)*, `CREDENTIAL_LIFETIMES` | the profile matrix and cred lifetime policy. |\n\n**Auth building blocks** (all `@cotal-ai/auth`)\n\n| symbol | purpose |\n|---|---|\n| `createCalloutAuth`, `startAuthCallout` | the NATS auth-callout responder. |\n| `createUserTokenIssuer`, `pinnedJwksResolver` | mint and verify the Cotal user bearer. |\n| `createIdpBridge` | exchange a verified IdP JWT for a Cotal bearer (see [the callout contract](identity-and-auth.md#the-idp-callout-contract)). |\n| `deriveOwnerToken`, `validateUserToken` | owner derivation; strict bearer validation. |\n| `cotalAuthProvider` | the self-registering `auth-provider` extension. |\n| `ensureCalloutAuth`/`loadCalloutAuth`, `ensureIssuer`/`loadIssuer`, `ensureOwnerSecret`/`loadOwnerSecret` | read/write the auth secret kinds through a `SecretStore`. |\n| `PLANE_CLAIM_REFUSED`, `planeClaimRefusal`, `PlaneClaimRefusal` *(type)* | every plane-claim refusal carries a `PLANE_CLAIM_REFUSED` detail, and `planeClaimRefusal(err)` reads its `reason`: `corrupt`, `live-peer`, `unknown`, `concurrent`, `fenced`, `released` or `lost`. Only `unknown`, an inconclusive liveness observation, is coded `unavailable`. A host can retry contention and stop on a corrupt row without matching message text. |\n\n**Seams and the wire** (all `@cotal-ai/core` unless noted)\n\n| symbol | purpose |\n|---|---|\n| `SecretStore` *(type)* | the durable hosted-secret seam (get/put/delete); `get()` returns raw seeds/keys into process memory, so it is a blob seam, not HSM/KMS signing. |\n| `FsSecretStore`, `workspaceSecretStore(root)` | the filesystem default. **These live in `@cotal-ai/workspace`, not core.** |\n| `AuthProvider` *(type)*, `Connector` *(type)*, `Runtime` *(type)*, `Command` *(type)* | the extension contracts; implementations self-register on import. |\n| `registry` | the shared registry a composition root pulls surfaces into. |\n| `CotalEndpoint`, subjects, message types | the wire client and shapes. |\n| `ParsedArgs` *(type)* | the shape the daemon runners take (see below). |\n\nFor a Linux Unix-socket adapter, `peerCredentials(socket)` from `@cotal-ai/seat` returns\nkernel-observed peer `pid`, `uid` and `gid`. Compare these against the host's authorization\npolicy; request-supplied identity and process liveness do not replace that policy or a\nlifecycle fence. The helper starts no custodian and refuses unsupported platforms or a\nmissing native helper.\n\nThe runners take a CLI-shaped `ParsedArgs`, not a typed options object, so a host fabricates one:\n\n```ts\nconst args: ParsedArgs = { values: { space, server, port: \"0\" }, positionals: [], raw: [] };\n```\n\nFor an embedded delivery instance, use `startDeliveryService` instead. Its `HostedContextInputs`\ninclude the account public key and lifecycle UID, space, broker URL, injected store, stable\n`storeIdentity`, and an explicit `stateDir`. The store must declare that same injected identity.\nThe initial delivery credential must belong to the assigned account. The function returns only\nafter the delivery responder is bound. `close()` withdraws serving and releases only the lease\nowned by that instance. It closes both membership connections even when a disconnected drain\nfails, so they cannot reconnect after closure. Credential-expiry health state clears after\nsuccessful broker-verified adoption through the existing `reloadCreds` rail. A failed start\nrefuses locally without exiting the host process or stopping another account's delivery service.\nIf a health fault occurs during an asynchronous store read, startup rejects when the read returns\nand closes any resources created by that late completion.\n\n`startAuthService` takes the same `HostedContextInputs`. The store must declare the assigned\ninjected identity, and its data account must be the assigned account. The IdP pin and ledger live\nunder the explicit `stateDir`, and the auth plane's instance identity in its `.cotal/space.<hex>/`.\nThe context never resolves a workspace root from the working directory and has no local manager, so\nonly remote manager gates can be selected. It returns after\nthe authority plane, the callout subscription and the loopback listener are bound. A start that\nfails closes the connections it opened and releases its plane claim, so a retry on the same space\ncan claim it. A fenced plane or\na lost broker connection makes that context `unavailable` and closes it without exiting the process.\nThe host writes no discovery file for it. The handle carries what `auth-service.json` holds for a\nCLI start: the loopback `url`, `publicUrl` when a public face runs, and the per-start `cap`. The cap\nalone authorizes the loopback host actions, lifecycle retirement and managed-agent enrollment\nverification, so keep it in the authority process. `publicFace` takes the CLI's public face inputs\n(`port`, `url`, `trustedProxy`, `advertisedServer`, `agentProvisioningUrl`) under the same rules,\nand a face without a port refuses to start.\n\nTwo optional inputs serve a platform composition. `platformControl: { observeAssignment }` adds\n`platformControlAuthority` to the handle. It is a typed in-process method, served on no listener,\nthat issues the manager-service request family for the one control manager the backend assigned to\nthis account, under a derived `p_` owner. It reads the assignment fresh on every call and refuses\nan IdP token, another account, a stale revision, another instance or lifecycle, and an instance\nanother owner registered. It refuses `prepare` and `activate` while the assignment's named\npredecessor is still registered or frozen. The same input adds `platformControlReadiness(instanceId)`,\nthe route for a host that needs to know whether its assigned control manager is serving. It returns\nthat instance's attributed `status` reply and refuses any instance the current assignment does not\nname, or one whose gate another owner holds. It reads over the context's own connection, whose\ngrant is the assigned instance's `describe` and `status` and its own reply rail. That connection\nrenews in process like the context's other connections and never leaves it, so the host lends no\nhuman or operator credential to a worker, mints no control instrument per read, and does not read\nliveness off the manager process. Without the input both members are `undefined`.\n`standingRenewableTtlSeconds` is forwarded unchanged to the authority plane, which bounds it to 5\nto 86400 seconds. It is a trusted-host input for the renewal rehearsal, and no request or CLI flag\nsets it. SPEC §13.1 and §13.6 define the view.\n\nThe auth plane can renew a registered manager's\nfive standing credentials from the current service registration. Run-driver renewal still\nrefuses without an authoritative activated-run reader, so these handles do not yet make a\ncomplete pooled auth and delivery host. A fresh auth plane can initialize without a\ndelivery-admin responder. Reclaiming a held claim from a dead predecessor needs the delivery\ninstance first: its admin rail must complete the broker connection-liveness sweep before the\nauth plane takes the claim. An absent or inconclusive oracle refuses the reclaim.\n\n### Remote manager client composition\n\n`@cotal-ai/manager` exports the `remoteManagerClient` namespace, containing the stock remote\nrequest builders and response validators, and `registerRemoteManagerAuthority` for registration\nwith a host-issued prepare credential. It returns the process epoch and registration revision its\nown registration committed. When a later start of the same instance registers before this start\nauthorizes its serve grant, this start is refused with `expired`.\n`RemoteManagerIdentityState` describes the five private\nmanager identities stored under an explicit account-local root. Use these public exports when\ncomposing `ManagerOptions.remoteAuthority`; do not copy CLI validators or import private modules.\n`managerClusterArtifacts()` returns the canonical document, manifest and their digests used by\nregistration. Pass its `[document, manifest]` pair as `contractArtifacts` to both\n`remoteManagerRegistrationProof(owner, state, contractArtifacts)` from `@cotal-ai/core` and\n`remoteManagerClient.remoteManagerAuthorityRequest(state, actor, \"activate\", { registrationProof, contractArtifacts })`.\nThe request builder takes each operation's coordinates as named fields of its last argument.\nThe host still validates the artifact closure and current registration before activation.\nThe namespace includes closed standing/run renewal, admission, maintenance, enrollment and\nretirement helpers. `remoteRunHosting` builds the four `runHosting` callbacks from the registration\nand the transport you supply for the host's run admission, run attempt and authority requests, so\nyour manager sends the run requests the stock manager sends. The namespace provides no signer or\nnew grant. The host still owns authenticated issuance, current registration and activated-run\nobservations, and any guarded foreign-holder repair.\nAn embedding must preserve those checks and supply a supported runtime; the client exports alone\ndo not provide a pooled runtime, an authority service or a complete hosted context.\n\n### Long-lived endpoints take a bearer function\n\n`EndpointOptions.bearer` accepts either a string or a function, and the difference is not stylistic.\nA string is minted once, so when it expires (which it will: callout bearers live minutes) the\nendpoint has nothing to renew with. It will not present the dead token to the broker, since that is\na guaranteed denial that still costs a full auth-callout round trip. It refuses to reconnect, emits\n`warning` saying which case it is in, and retries on a widening backoff until the process\nre-authenticates and rebuilds it. Retry notices use `warning` rather than `error` because Node\nrethrows an unhandled `error` event and would kill a host the endpoint is still trying to recover.\n\nPass a **function** for anything that outlives one bearer. That is a renewal source: it is called\nahead of each expiry and again whenever a reconnect finds the cached bearer dead, and it requires\nexplicit `card.owner` and `card.actor`. The first-party surfaces already do this\n(`UserViewAuth.source`, the connector's `agentBearerCommand`). A string bearer is for a short\none-shot connection.\n\nLong-lived hosts must also subscribe to the endpoint's `warning` event. It carries conditions the\nendpoint is surviving, including failed credential renewal, reconnect retries, and a durable leave\nit keeps retrying after the broker refuses a durable channel's live subscription. A host may choose\nto ignore warnings for a one-shot endpoint whose awaited operation owns the verdict, but that choice\nshould be explicit. An unhandled warning is nonfatal and silent.\n\nThe `error` event carries a fault the endpoint cannot return from a call, such as a refused\nsubscription or a refused publish that no request was waiting on. Attach a listener before `start()`,\nsince Node throws on an unhandled `error`. A denial the broker returns to a request, such as an\nobserver's read of the DM stream, reaches only that call, which decides what it means, and is not\nemitted again as an `error`.\n\n## Booting the daemons\n\n### auth-service\n\n`runAuthService(args, store?)` reads its provisioned long-lived secret kinds (service keys, callout\naccount, issuer keys, owner secret) through the injected `SecretStore`; a host provisions those into\nthe store first. It is a **signer and identity authority**, not a scoped daemon: at runtime it holds\nthe data-account and callout-account signing seeds, the issuer's private JWKs, and the\nowner-derivation secret in process memory (`SecretStore.get` exports raw values). The IdP pin and the\nactor ledger are **not** store-injected: `runAuthService` resolves them under\n`userAuthStateDir(findCotalRoot(), space)`, a path relative to the process working directory, so a\nhost provisions those into that exact directory (neither `store` nor `COTAL_HOME` selects it). It\nalso writes an ephemeral `auth-service.json` discovery file there that carries the live exchange\ncapability. That file appears only after every plane is bound, so waiting on it is the readiness\nsignal: a host that also passes the daemon's pid to the provider's `ready()` gets a process-bound\nwait. The wait extends past the base timeout while that pid is alive, up to a fixed bound, and it\nends at once when the pid exits.\n\n```ts\nimport { runAuthService } from \"@cotal-ai/auth\";\n// store implements SecretStore over your secret backend; get() returns raw seeds into memory.\n// Provision the auth secret kinds into the store, AND the IdP pin + actor ledger under\n// userAuthStateDir(findCotalRoot(), space), before this call.\nawait runAuthService(\n { values: { space, server: brokerUrl, port: \"8081\" }, positionals: [], raw: [] },\n store,\n);\n```\n\n### delivery\n\n`runDelivery(args, store?)` runs from a **pre-minted scoped `delivery` cred** and never loads the\nsigner. Provide the cred either through the injected store (under\n`deliveryCredsKey(space, { injected: true })`) or with a\n`--creds` file; the two are mutually exclusive. The daemon re-fetches the cred from the store at 75%\nof its JWT lifetime and fails loud rather than riding to expiry, so **something must re-sign a fresh\ncred into that same store**. When that read finds the previous generation still there, the daemon\nreports the missed remint and retries in 60 seconds. The current cred stays live until its expiry,\nand the store is read once per retry rather than once per second.\n\n```ts\nimport { runDelivery } from \"@cotal-ai/delivery\";\nawait runDelivery({ values: { space, server: brokerUrl }, positionals: [], raw: [] }, store);\n```\n\nThat renewal is a **signer** operation, not the delivery daemon's:\n`remintDaemonCreds(root, space, store?, { preflight? })` (`@cotal-ai/workspace`) reads the `SpaceAuth`\nsigner **through the same resolved `store`** (`getSpaceAuth(store ?? workspaceSecretStore(root), space)`,\nkeys `auth/broker.json` + `auth/account.<key>.json`; the pre-split `auth/auth.json` monolith is\nmigration input and the container signer mount only) and re-signs the daemon creds (`delivery.creds` and the membership feed's\n`membership-rw.creds`) back into that store. The injected `store` is both the signer source and the\ncredential destination, never a split. `space` is **required** and validated against the store's signer, so a\nstore swapped to a different space cannot re-sign over the wrong broker's creds. `preflight` is a\ncaller-supplied proof that the broker accepts the credential. The reference `Manager` passes a\n`probeConnect` over its `servers`. It gates **every** candidate before overwriting the last-good,\nwhether the signer is a full bundle or a stripped projection: a bundle's JWT chain proves only that\nit is self-consistent and\nnamed the space, NOT that its account is the broker's *current* account for that space (two\n`createSpaceAuth(space)` calls yield same-named, different-account chains), so a same-label alternate\nsigner would otherwise mint a broker-dead cred and clobber the good one. The offline local repair (`doctor auth --fix`) has no preflight. It permits the overwrite only\nunder **authority continuity**: the candidate must be signed by the same account signing key (`iss`) as the current\n(already broker-accepted) cred. A same-label alternate account breaks continuity and is refused, full or\nstripped; a legitimate local re-sign is continuous and proceeds without a network. The reference\n`Manager` runs it on a schedule against its **own**\n`secretStore` (see below), so passing the manager and the delivery daemon the *same* store closes the\nrenewal loop end-to-end on an injected backend: the manager reads the signer from the store, re-signs\ninto it, and the daemon adopts each generation on a preflight-proven 75% timer. The stock\ncross-host composition cannot satisfy that by writing one filesystem and fingerprinting another:\n`Manager.start()` and every later remint challenge the daemon's `reloadStoreIdentity` and a\ndivergent pair is refused naming both stores. The identity is the store the daemon actually\nreloads: an injected coordinate, the workstation root only when `--creds` is\n`<root>/.cotal/<spaceSegment(space)>/delivery.creds` (matching the canonical arm), the\nfile's own directory for any other `--creds` path, or the workstation root. A filesystem\nstore is also named by a random id it records in `store.id` inside its own directory, so two\nhosts that use the same root path are two stores. No key and no `--creds` file may be that\nfile under any name, and a `store.id` that is a symbolic link or holds anything but a lowercase UUID is refused. Uninjected\n`--creds` that names one real workstation while process cwd resolves another is refused\nat start, naming both, because membership-rw still uses `findCotalRoot`. A `--creds`\npath that is not under any `.cotal` tree is not that case and is not refused here. It never\nwalks ancestors with `findCotalRoot`. No bound daemon is not a named\nstore, so start proceeds; a later daemon on a foreign store is refused on the next remint.\nThe first-party filesystem adapter declares its workspace-root identity on the store itself. Other\ninjected adapters declare their stable coordinate on `SecretStore.identity`, or name it in\n`COTAL_SECRET_STORE` on both processes. It never throws: it\nreturns per-file results (`skipped: \"no-auth\"` when the store holds no signer records),\nso the caller must check them or the cred still rides to expiry. A composition whose signer lives in\nKMS/Vault simply injects that store; no bespoke renewal is needed. A `--creds` file path must be\nreplaced atomically before the 75% read. The signer can now be injected behind the store seam, which\nresolves custody. The remaining hosted gap is signer **isolation**. The seed is still decrypted\nin-process at the manager's uid, so it needs an OS sandbox or remote signer.\n\n### Supervisor signing authority\n\n`@cotal-ai/manager` exports the `Manager` class; there is **no** `runSupervise(opts)` runner. The\nprivate CLI `runManager` also does broker-reachability checks, space/default resolution,\nroster/launch parsing and materialization, installed-extension resolution, signal handling, staged\npre-spawn, and the forever wait. A host composes that lifecycle itself around `Manager`:\n\n```ts\nimport { Manager } from \"@cotal-ai/manager\";\nconst mgr = new Manager({ space, servers: brokerUrl, workspaceRoot });\nawait mgr.start(); // then wire your own SIGINT/SIGTERM -> mgr.stop()\n```\n\n`stop()` runs once. A later call joins the stop in progress and settles with it, and a call that\nasks for a different `withAgents` than the running stop is refused.\n\nUnlike delivery, the manager is **not** a pre-minted-scoped-cred daemon (auth-service is also a\nsigner: it holds fewer artifacts than the full trust bundle, but its data-account signing seed still\ngrants complete data-account mint authority on compromise, so this is not least-privilege). On `start()`\nthe manager reads its space's full trust chain **through its `secretStore`** (`getSpaceAuth(this.secrets,\nthis.space)`, composed from `auth/broker.json` + `auth/account.<key>.json`; a container may instead\nmount a stripped signer bundle at the legacy `auth/auth.json` key) and **self-mints** its supervisor cred and renewals from the\ndata-account signing seed. In static mode it also mints every per-agent cred from that seed; in user\nmode agents instead receive callout-minted bearers, but the manager still holds the signing seed for\nits own creds and renewal. So a hosted supervisor is a **trusted per-tenant account-signer process**,\nnot a least-privilege connect client. It additionally requires a `~/.cotal/meshes/space.<key>.json`\nregistry record and the workspace user-auth marker to start in user mode. `ManagerOptions.secretStore`\ninjects the one `SecretStore` the manager uses for **the signer itself (the split trust\nrecords)**, daemon-credential renewal (`remintDaemonCreds`), and per-agent secrets,\ndefaulting to the workspace filesystem store; pass the delivery daemon the *same* store for end-to-end\nhosted renewal. The store declares the same identity on both processes, or both set\n`COTAL_SECRET_STORE` to the same coordinate. The manager\nremints no daemon credential when the daemon names a different store, including a daemon that binds\nafter start; it keeps running and serving its own agents, so one space can carry a manager on more\nthan one workspace root. That manager also stays off the space's renewal lease, so a manager or a\n`cotal doctor auth --fix` on the daemon's store can still take it. Pointing several managers at one coordinate is safe: the store identity\nalone cannot pick an owner (it carries no holder and no tiebreak, so every manager sharing the store\nmatches), so the manager that also holds the space's renewal lease is the one that remints and the\nrest skip it. Without that lease two owners would remint on independent timers with no ordering\nbetween them, and one write would land between the other's re-sign and its fingerprint-only\n`reloadCreds`. `cotal doctor auth --fix` takes the same lease before it re-signs, so a live manager\nand a local repair never race each other either. The signer IS now injectable: a hosted composition injects a KMS/Vault store and no\nsigning seed lands on the hosted disk. What remains is signer **isolation**. The seed is decrypted\nin-process at the manager's uid. That issue needs an OS sandbox or remote signer; it is no longer a\ncustody problem. The other knobs are `workspaceRoot` and the process-global `COTAL_HOME`.\n\n> Scope note: the **static-auth** operator paths (`cotal spawn`/`join`/`status`/`web`, via\n> `mesh-target` → `connect`/`preflight`) still read the signer from the local split records (sync\n> `loadSpaceAuth`). That is the single-machine composition, where the signer is on local disk by the\n> static-auth model; multi-tenant hosting runs **user mode**, which never mints from on-disk trust. The\n> store-injectable signer path is the hosted-server set: the manager, `remintDaemonCreds`, and delivery.\n\nThe typed remote-manager authority contract includes a one-shot terminal phase. A host implements\n`remoteAuthority.prepareAgentRetirement` to revoke the managed grant and finish its resumable\nrelease while preserving the UID, then `remoteAuthority.mintRetirementRequester` returns the\nhost-signed JWT for a fresh participant-owned nkey. The credential is pinned to the authenticated\nowner, server-derived manager serve principal, current instance epoch, and exact target lifecycle.\nThe manager then uses the existing auth `retireLifecycle` rail with the operation id derived by\n`managedRetirementOpId(target.lifecycleUid)`. This derivation is the reference remote-Manager\ncomposition's closed contract, not a rule for every retirement entry point; interactive retirement\nkeeps its existing operation identity and remains compatible. The `retireLifecycle` rail independently\nrecomputes the managed id from its broker-pinned target before any gate, head, intent, or barrier\naccess, so mint-time validation is not the terminal boundary. A failure keeps\nthe alias held. This does not expose the auth barrier or give the participant signer authority.\n\nIf the participant disappears after prepare, the host finishes the retirement itself on the auth\nservice's loopback face: `POST /managed-lifecycle/retire` (exported as `MANAGED_RETIRE_PATH` from\n`@cotal-ai/auth`) with the `Bearer <cap>` from `auth-service.json` and only\n`{ owner, actor, lifecycleUid }`. It has the interactive door's guards (POST only, no `Origin`, JSON,\ncapability, closed body) and is never served on the public face. The managed grant must already be\nrevoked at that uid, or it answers 409. It runs the same `managedRetirementOpId(uid)` operation as\nthe rail, and the rail and the door share one in-process flight, so a late participant request and\nthe host call converge on one barrier.\n\n| lifecycle head | answer |\n| --- | --- |\n| absent, or `retired` at another uid | `200 { retired: false, lifecycleUid, notStarted: true }` |\n| `active`/`retiring` at another uid | `409` |\n| `retired` at this uid | `200 { retired: true, lifecycleUid, alreadyRetired: true }` |\n| `active`/`retiring` at this uid | the barrier runs, then `200 { retired: true, lifecycleUid }` |\n\nDeprovisioning durables stays with `deprovisionAgent` and a `deprovisioner` credential. Pass `memberChannels` to both to also purge the retired lifecycle's durable membership rows on those concrete channels. The manager fills that list from the launch's concrete read channels and the delivery daemon's read-only `lifecycleMemberships` admin verb. When that verb cannot answer, rows on other channels stay retained, the teardown logs the inventory as incomplete, and retirement remains held pending retry rather than releasing the alias. Each teardown examines two exact consumers, one ACL key and the named member keys. KV deletion uses a native revision condition; a lost condition with a live replacement refuses rather than claiming absence. Consumer INFO checks before and after DELETE distinguish verified prior absence from disappearance. `acknowledged` counts native DELETE success replies. `disappeared` counts observed live-to-absent consumers and live KV rows whose conditional purge lost to a competing deletion. These KV rows were present at the first read, so they never count as prior `absent` or as this caller's `deleted`. The ACL and membership subtotals preserve that distinction. Neither establishes which concurrent caller uniquely removed a consumer, so `consumers.deleted` and the total `deleted` are `null` when a consumer disappears without a winner token. A repeat after verified absence reports zero, not an invented deletion. `refused` counts slots whose cleanup or state remains uncertain. A partial failure raises `DeprovisionError` carrying these bounded observations. The Manager's static sweep preserves unknown uniqueness rather than adding acknowledged requests as physical removals; its slot totals remain separate.\n\nA host that resumes retained managed actors also implements\n`remoteAuthority.validateRetainedAgent`. The participant sends back the actor token and sentinel it\nalready holds, plus the `nextRegistrationProof` returned by the activation response. That proof is\nhost-issued after registration and binds the manager owner, actor, lifecycle, identity nkeys, current\nregistration revision, and serving epoch. The host checks it against the current open manager gate,\nvalidates the retained secrets against its current managed row, and returns only the non-secret\nauthority shape. The manager binds every result coordinate and the returned authority back to its\ninventory before use. Do not copy the provider's `issuer.json` or `callout.json` into the participant\nstore. Both contain private signing or exchange authority.\n\nThe same composition supplies `remoteAuthority.agentBearerExchangeUrl`, the pinned public auth-service\nbase used by retained children. Remote adoption launches `agent-bearer --exchange-url <base>`; it must\nnot select the local `--dir` arm, which depends on a host-only auth-service process record.\n\nA host that lets a remote participant spawn FRESH managed agents implements\n`remoteAuthority.enrollManagedAgent`. The participant generates the standing actor token, writes it\nat mode 0600, and passes only its SHA-256 digest with the requested actor, label, role,\ncapabilities, and channel lists, so the plaintext secret never leaves the participant machine. There\nis deliberately no `lifecycleUid` input: the host selects the UID, because only the host sees the\nretirement tombstones that make a UID permanently unusable, and a participant-chosen UID could aim a\nfresh grant at a dead incarnation. The host authors the ledger grant, pre-creates the lifecycle-keyed\ndurables, clamps the requested lists to what the spawning owner already holds, and returns the owner,\nactor, chosen `lifecycleUid`, the space sentinel credentials, the effective lists, and\n`agentBearerExchangeUrl`. The manager binds every returned coordinate, re-keys the secret family onto\nthe returned UID, and launches `agent-bearer --exchange-url <base>`. When the hook is absent a\nsignerless manager refuses the user-mode spawn rather than authoring a local grant the host knows\nnothing about.\n\nBoth managed-agent operations ride the one verified `POST /manager-service-authority` transport as\n`kind: \"manager-managed-agent-enrollment\"` and `kind: \"manager-managed-agent-prepare-retirement\"`.\nStock `cotal auth-service` answers both itself, because it owns the actor ledger and the space's\nprovisioning authority. An enrollment writes the managed grant at a fresh UID with the supervising\nactor as its parent, provisions that UID's durables, and returns the daemon's public exchange URL as\n`agentBearerExchangeUrl`; a daemon started without `--exchange-public-port` refuses enrollment. A\nretry with the same token digest answers the same UID while the supervising actor's current grant\ncovers it, and a fresh enrollment's refusal otherwise. While the agent's grant stands, an enrollment\nwith another digest is refused with `conflict` until that lifecycle's retirement is prepared. A\nprepare-retirement releases the target UID's broker footprint and then revokes its grant, so the\nmanager's terminal rail finds the grant gone. A platform that keeps these writers in its own storage\nintercepts both kinds instead. It terminates its own public route, authenticates the human there,\nand asks the auth service for the decision at\n`POST /manager-service-authority/verify-enrollment` (exported as `VERIFY_ENROLLMENT_PATH` from\n`@cotal-ai/auth`) with the `Bearer <cap>` from `auth-service.json` and only `{ owner, request }`. That\ndoor has the managed retirement door's guards, derives the caller's scope from the local ledger rather\nthan the body, checks the manager gate and registration proof in-process, and answers\n`{ authorized: true, owner, actor, instanceId, serveEpoch }`. `authorizeRemoteManagedAgentEnrollment`\nand `authorizeRemoteManagedAgentPrepareRetirement` are exported too, for a host that composes the\ndecision without the HTTP hop. Both require ledger scope `supervise`; `spawn` and `admin` do not\nimply it.\n\nA host that runs managed agents on its own hosted runtime adds two more kinds on the same transport.\n`kind: \"manager-managed-agent-runtime-create\"` asks the host to create the runtime for one agent it\nalready enrolled, and `kind: \"manager-managed-agent-runtime-status\"` reads that runtime's state. Both\ncarry the manager envelope plus `target: { owner, actor, lifecycleUid }`, the coordinate the\nenrollment returned. Both schemas are closed. An unknown top-level or target field, including\n`providerRef`, `handle`, or `name`, is refused as `bad-request`, because the host alone issues and\nholds provider references. There is no stop, adopt, or probe kind: stop goes through\nprepare-retirement. Stock dispatch refuses both kinds with `unimplemented`, and the verify-enrollment\ndoor decides them. `authorizeRemoteManagedAgentRuntimeCreate` and\n`authorizeRemoteManagedAgentRuntimeStatus` apply the enrollment door's checks: host space, a target\nowner equal to the authenticated owner, the manager actor's own ledger row with `supervise`, the open\ngate, the current serve epoch, and the registration proof. Each returns only\n`{ owner, instanceId, actor, target }`. The door touches no provider and writes nothing. The host\nmatches the decision to its own intent record and performs the create afterwards. The host answers\nwith `state` (`reserved`, `creating`, `bound`, `create-unknown`, `closing`, or `closed`), `readiness`\n(`ready`, `bound-not-ready`, or `none`), and an optional `retirementPhase`. A manager builds requests\nwith `remoteManagerClient.remoteManagedAgentRuntimeRequest` and binds the answer with\n`remoteManagedAgentRuntimeState`.\n\nAn enrollment result may also carry `runtimeIntent: { state: \"reserved\" }` when the host reserved a\nhosted runtime for the agent. It is display-only. Older hosts omit it, the manager binds both shapes\nto the same material, and nothing reads it as authority.\n\nNo stock door lets a platform control holder launch or retire an agent for a signed-in user. The\nmanaged-agent kinds above act only under the authenticated owner and refuse a caller that is not\nthat user, so a platform could only run a user's agent by holding the user's login or by enrolling\nthe agent under its own owner. Both are refused. The\n[delegated user launch intent](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/delegated-user-launch-intent.md)\ndesign and SPEC §13.16 define the smallest addition. The user admits one launch or one retirement\non the host's authenticated route. The holder consumes that intent once, from its current\nregistration, epoch and lifecycle. The host then enrolls the agent under the user's `u_` owner with\nthe user's own actor as its ledger parent, so the envelope walk, membership and channel lists match\nwhat the user's own manager would produce. Retirement keeps the prepare, provider closure and\nterminal barrier order, and the host finishes it when the holder is gone. A launch the host had to\nundo keeps its agent name held until the host process that ran it confirms it has stopped, and\nwhile the name is held the host also refuses it to the user's own manager. `@cotal-ai/auth` ships the\ntwo decisions, `authorizeDelegatedUserIntentAdmission` and `authorizeDelegatedUserIntentExecution`,\nfor a host that owns an intent store and those writers to compose on its own routes. Both read the\nholder's gate through the handle's `observeManagerGate`, present with `platformControl`, which reads\nover the context's own connection, so the host opens no second data-account connection. It is an\nobservation for the decision: the consuming CAS and the writers still apply their own checks. The\nconsuming CAS pins the incarnation of the host process that runs the flight as the executor, never\nthe auth plane's, because the two can restart independently. `platformControl.host` names that\nprocess's reverse-DNS endpoint, the closure digest of its §13.7 cluster and the contract artifacts\nregistration reads, and the auth plane self-authorizes that one name. Registration reads the closure\nmanifest `{ v: 1, root, members }` at `clusterDigest`, then the cluster document at the manifest's\n`root`, and verifies each against its digest. `artifacts` therefore carries both, and `clusterDigest`\nis the digest of the manifest. `members` stays empty: SPEC §13.7 lists every reachable artifact\nthere, but this implementation registers single-document clusters only and refuses a manifest that\nlists members. `singleDocumentClosure(document)` returns that manifest and its closure digest. The\ninstance id is a lifecycle token, `[a-z0-9]{26,32}`. The minimal construction below has one command\nover the void schema. A host copies it, replaces `document` with its real cluster, and passes `host`\nas `platformControl: { observeAssignment, host }`.\n\n```ts\nimport { mintLifecycleUid, singleDocumentClosure, VOID_SCHEMA_DIGEST } from \"@cotal-ai/core\";\n\nconst document = {\n urn: \"com.example.host\",\n revision: 1,\n attributes: [],\n events: [],\n commands: [{\n name: \"ping\", class: \"ephemeral\", targeted: false, capability: \"host.ping\",\n inputDigest: VOID_SCHEMA_DIGEST, outputDigest: VOID_SCHEMA_DIGEST,\n }],\n};\nconst { manifest, closureDigest } = singleDocumentClosure(document);\nconst host = { endpoint: \"com.example.host\", clusterDigest: closureDigest, artifacts: [document, manifest] };\nconst instanceId = mintLifecycleUid(); // first start only; later starts reuse the persisted id\n```\n\n`registerHostIncarnation(instanceId)` publishes the artifacts, registers that instance through the\nceremony the plane runs for itself, and returns `{ instanceId, processEpoch }` with the epoch that\nregistration committed. The host calls it at every start with its persisted instance id, before it\nadmits or recovers any flight, so a restart fences its predecessor. The first registration of an\ninstance commits epoch 0, which is open and serving like any later epoch, and each later start of\nthat instance commits the previous epoch plus one. A consumer compares epochs for equality and never\nreads 0 as absent or not ready. A start\nwhose confirming read of the gate finds that a later start of the same instance registered is\nrefused with `conflict`. The returned epoch is a committed coordinate and stays current only until\nthe next start registers, which can happen before the call returns. `observeHostGate(instanceId)`\nis a point-in-time read of an executor's gate, and answers null for an absent gate; a sweeper\ndecides on it. `awaitHostFence(instanceId, processEpoch)` resolves with the gate once it is no\nlonger open at that epoch, and with null once it is absent. It takes any non-negative safe integer\nepoch, 0 included, and refuses any other value with `bad-request`. The host arms it with its returned\nincarnation before it admits or recovers any flight, and stops serving when it resolves. It polls\nthe gate, because no runtime credential may watch the auth bucket. The\nhost's launch writer first activates the agent's lifecycle at the pinned UID through the handle's\n`activateManagedLifecycle`, before any ledger row or durable, and its compensation runs the same\ncall before the terminal barrier, so a launch whose agent never exchanged its bearer still reaches\nthe terminal barrier at that UID. The launched agent exchanges its bearer on the context's public\nface, so the host starts it with `publicFace`, and its retirement writer ends at\n`POST /managed-lifecycle/retire` with the handle's `cap`. Stock dispatch refuses both kinds as\n`unimplemented`. The holder's composition passes\n`remoteAuthority.executeDelegatedUserIntent`, which posts the execution request and binds the answer\nwith `parseRemoteDelegatedUserIntentExecutionResult`. It then starts the agent with\n`startAgent({ ..., delegatedIntent: { intentId, owner, parent } })` and retires it with\n`retireDelegatedAgent(name, intentId)`, which stops the agent only after the host confirms\n`retired: true` for its exact target. A delegated agent's stop, exit or failed launch keeps its name\nheld until that retirement confirms.\n\nRemote user-mode managers must also supply `remoteAuthority.authorizeAdmin`. The manager builds each\nrequest only from the caller tuple parsed from the broker-authenticated endpoint subject, then relays\nthat tuple over the current registered manager lifecycle. HTTPS does not separately authenticate the\nrelayed caller. The host authenticates the manager operator, binds the request to the current open\nmanager gate, registration proof, serving epoch, and identity nkeys, then reads the caller's unified\nauthoritative row fresh. It returns only the manager owner and `authorized: boolean`, with every request\ncoordinate echoed. Missing, revoked, narrowed, foreign-owner, and stale-lifecycle callers all return\n`false`; malformed coordinates or corrupt and unavailable authority state fail the operation. The\nparticipant never reads or mirrors the host ledger, and the remote branch has no local fallback. The\nsame callback gates all `manager.admin` handlers, any-mode cross-owner control, and `ps` or `inspect`\ncross-owner visibility. Launch keeps its owner-equality policy.\n\nThe remote authority's instance executor remains the scoped maintenance credential for clean service\nderegistration and exact instance registration operations. It carries no records-stream consumer\nlifecycle authority. The manager's boot `goalidx` sweep uses the authenticated host operation, which\nreturns parsed `goalidx.manager.<owner>.>` entries for that owner only. The host keeps the sealed\nconsumer connection and its create/delete rights. The five-minute executor already renews through\n`remoteAuthority.renewExecutor`. The signerless supervisor, serve, goal-writer, session-ledger and\nper-run driver credentials do not yet have a complete remote renewal and adoption path. The\n[hosted runtime contract](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/hosted-runtime-contracts.md) records the bounded additions and\ntheir ownership; it is not a shipped pooled service.\n\n**Signer isolation needs an OS sandbox.** The default pty runtime\nruns agent children under the *same* OS uid and the *same* `workspaceRoot`, so mode-0600 on\nthe trust records does not stop a hostile same-uid agent from reading their absolute paths. The reference\n[deploy](deploy.md) tree does not solve this: it mounts the signer into the agent's own container, so\nits phase-1 boundary isolates agents from each other, not the signer from the agent. A hosted\ncomposition must run the manager/minter that holds the signer in a different uid, container, or mount\nnamespace from the agent children, which mount no signer at all; that split is future\nhosted-composition work, so until it (or a remote/injected minter) exists, do not run untrusted\nagents under this manager.\n\n### Delegated seats outside the manager's filesystem\n\nThe [portable lifecycle bootstrap](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/portable-lifecycle-bootstrap.md)\ndesign and SPEC §13.17 define how a managed agent that `enrollManagedAgent` already enrolled starts\nin a child that cannot see the manager's filesystem. The ordinary `spawn` path writes the token and\nsentinel under the manager's workspace root and hands the runtime a launch whose bearer command and\nmaterial file are paths on that filesystem, so such a child needs this path instead.\n\nThe delegation boundary is one optional runtime method. A runtime that implements\n`Runtime.spawnDelegated(launch, handoff)` receives two values and no paths: a `DelegatedSeatLaunch`\n(connector name, persona text, model and the other launch choices) and a `ManagedLifecycleHandoff`\n(space, owner, actor, the host-chosen `lifecycleUid`, broker and IdP pins, the pinned exchange base,\nthe sentinel, the channel lists, and the raw actor token). The manager enrolls once, as it does\ntoday, and builds the handoff from what it holds. It never sends the token to the host, never copies\na file from its workspace or secret store, and never builds a local launch for that seat. No signer,\nissuer or callout record, loopback capability, provisioner or manager credential, control token, or\nmanager path crosses. A spawn choice that only the manager's host can honour is refused before\nenrollment: `--resume`, a manifest agent's `continuity: exact`, `--cwd`, and any shared MCP server,\nwhether from `--share-tools` or the config default (`--share-tools none` passes).\n\nThe runtime creates one provider resource under `managedRuntimeKey(target)`, writes the handoff into\nit as one 0600 file and the persona beside it, and runs the stock bootstrap there:\n`cotal spawn --config <persona-file> --space <space> --name <actor> --expect-owner <owner> --expect-lifecycle-uid <uid>`\nwith `COTAL_MANAGED_HANDOFF_FILE` naming the file. `delegatedSeatCommand` builds that argv. The\n`cotal` entry reads the file into memory, deletes it and drops the variable before it parses flags,\nprints help or loads extensions, so every outcome, a refusal of its own flags included, leaves no\nfile. It refuses before any broker connection or exchange request when the\nspace, owner, actor, or lifecycle UID differ from the expected values, and then runs the\nenrollment-redeem consumer: it registers the mesh in its own home, writes the token to its own 0600\nfile, and exchanges it through `agent-bearer --exchange-url` unchanged. It never enrolls, redeems, or\nmints a token or UID.\n\nReadiness is still mesh presence. A create whose answer is lost leaves the handle running, so the\nlaunch settles uncertain and stays held; the manager never retries it. A provider read that finds no\nresource under the key is not an exit, because the create may still land. Every close by\n`managedRuntimeKey` is fenced: it completes only once the create was answered or the provider refuses\nany later create under the key. A provider that names its own resources may run the create as a\ndurable operation keyed by `managedRuntimeKey` and close through the identifier its authenticated\ncreate response returned, kept where the host can read it without the manager. Only that response\nbinds an identifier to the key; one derived from the key or found by name or listing is never closed\nor adopted, and while the response is unknown the launch stays held. Every stop, the reap of a child\nwhose parent exited and `Manager.stop({ withAgents: true })` included, runs `prepareAgentRetirement`\nfor the UID-exact target, then `stop()` on the handle `spawnDelegated` returned, then the terminal\nbarrier. `preparePreservation` refuses a cut that holds a delegated seat, and a `Manager.stop()` after\na refused cut retires the seat through the same steps. After the manager is gone\nthe host runs the same steps, makes the same fenced close by `managedRuntimeKey`, and finishes at\n`MANAGED_RETIRE_PATH`. Supply `spawnDelegated` only from a runtime whose host can make that fenced\nclose without the manager. The enrollment redeem\n(`COTAL_ENROLLMENT_FILE`) stays for lifecycles whose token the host generated itself; a host cannot\nmint one for a manager-enrolled lifecycle because it holds only the digest.\n\n## Provisioning a space (one-space reference shape)\n\n```ts\nimport { createSpaceAuth, setupSpaceStreams, ensureDefaultDeliveryClass, mintCreds, newIdentity } from \"@cotal-ai/core\";\nconst auth = await createSpaceAuth(space); // trust bundle (in-memory seeds)\nconst provisionerCreds = await mintCreds(auth, newIdentity(), \"provisioner\");\nawait setupSpaceStreams({ servers: brokerUrl, space, creds: provisionerCreds });\n// SPEC section 4: write the default delivery class at space creation so it is wire-discoverable,\n// never inferred from the resolution fallback. A daemon-backed space is \"durable\".\nawait ensureDefaultDeliveryClass({ servers: brokerUrl, space, creds: provisionerCreds, deliveryClass: \"durable\" });\nconst deliveryCreds = await mintCreds(auth, newIdentity(), \"delivery\");\n// put deliveryCreds into your SecretStore under deliveryCredsKey(space, { injected: true })\n// (@cotal-ai/workspace) before booting delivery — the key is per-space, not the bare kind.\n```\n\nRendering the broker config for a user-auth space is `serverConfig(broker, spaces, { storeDir,\nmaxFileStore?, extraAccounts })`, where `extraAccounts` must include the callout account from\n`createCalloutAuth` so the auth-service has a broker account to answer on. That account never shares\nthe data account. `maxFileStore` is an optional positive integer byte cap; any other value throws.\n\nBroker trust and space accounts are separate authorities: `createBrokerAuth` mints the one\noperator + system account a broker trusts, `createSpaceAccountAuth(broker, space)` signs each\ntenant's data account under it, and `serverConfig(broker, spaces, opts)` renders them all into one\nconfig. A host composition can therefore provision several spaces on one broker today. `cotal up`\nrenders that config from every tenant the root's auth directory holds, so booting one space keeps\nthe broker trusting its siblings, and it refuses to render at all while any account record is\nunreadable. The rest of the CLI lifecycle is still broker-wide: `down`, `clean` and `backup` refuse\non a multi-space root rather than scoping to one tenant, and the per-space lifecycle is the\nremaining multi-space operator layer. See\n[Known gaps](#hosted-composition-gaps).\n\n## Hazardous provisioning primitives\n\n`mintCreds`, the full `Profile`/`CredentialKind` matrix, `createSpaceAuth`, and `stripSpaceAuth` are\nlow-level operator primitives. Handle them as account-authority material:\n\n- A holder of a `SpaceAuth` (or a `stripSpaceAuth` bundle, which **keeps** the data signing seed) is\n a fully-trusted tenant-account authority: it can mint `admin`, `provisioner`, and destructive\n profiles, not merely `supervisor`, and mint a DM-reading identity. `createSpaceAuth`'s full result\n holds operator, system, and account seeds in memory.\n- Choose `profile` and `MintOpts` from **server-side constants**, never from tenant input. `MintOpts`\n can widen the bounded TTL defaults; cap it at your boundary. `CREDENTIAL_LIFETIMES` is a policy\n record, not an authorization boundary.\n- Never log signer material or export it into env. Do not co-locate signer access with an untrusted\n connector/runtime process at the same OS uid (file permissions do not contain a same-uid reader;\n see the manager's isolation note). Segregate per tenant; rotate on compromise\n (`rotateDataAccountSigningKey`).\n\n## Hosted composition gaps\n\nThe primitives above are present as exports, but three capabilities are **not** cleanly composable\nfrom the public contract today. Each is tied to work in flight; a host either waits for the seam or\nscopes the capability out. None is a wire concern.\n\n1. **Delivery immediate live eviction and a fully-hosted membership feed.** The renewable\n `membership-rw.creds` is now a `SecretStore` kind. `startMembership` reads it through the\n injected store, and the manager re-signs it there. The graph-feed writer therefore renews on a hosted\n backend (its data connection adopts each generation on a preflight-proven 75% timer). What still\n reads from a fixed on-disk path are the *static* `membership-observer.creds` and\n `connection-evictor.creds` ($SYS creds, minted at the `up` that provisions the account and renewed by `up --rotate-sys`) and `membership.json`\n (`{accountId}`, non-secret config); those, plus the private provisioning wrapper, keep immediate\n live eviction and a fully-hosted feed a partial gap. Missing files degrade membership to\n traffic-only and make live eviction refuse (loudly). The supported delivery contract here is the\n Plane-3 durable backstop.\n2. **Supervisor signer isolation.** `ManagerOptions.secretStore` now injects the one `SecretStore` the\n manager reads/writes every secret through, including the composed `SpaceAuth`\n signer (the split trust records), its daemon-cred renewal, and its per-agent kinds. What remains is process\n isolation: the manager still decrypts the signer in-process at its uid, so untrusted agent children\n must run under a different uid/container/mount namespace or behind a future remote signer.\n3. **Per-space lifecycle on a shared broker.** The trust layer is multi-space\n (`createBrokerAuth` + `createSpaceAccountAuth` + N-space `serverConfig`, persisted as\n `broker.json` + `account.<key>.json`) and `cotal up` renders the whole tenant list, but there is\n no per-space provisioning verb and no per-space teardown/backup/restore: the CLI's broker-wide\n lifecycle verbs refuse on a multi-space root, naming the tenants.\n This is the remaining multi-space operator layer.\n4. **A non-Better-Auth production IdP.** The exchange core (`createIdpBridge`) is EdDSA-generic, but\n the stock provider and login client are Better-Auth-endpoint-shaped, `cotalAuthProvider`\n self-registers on import (colliding with a host-owned provider under `resolveAuthProvider`), and\n the login flow speaks Better Auth's device-code endpoints. A different IdP is a host-built auth\n composition on the low-level primitives, not a configuration change (see\n [the IdP callout contract](identity-and-auth.md#the-idp-callout-contract)).\n\n## Hosted durability\n\nSpace-durable **coordination** state (chat/DM/task history, live presence, membership runtime, the\ndurable ACL registry, leases) lives in **JetStream**, written by the delivery daemon and the\nendpoints. It is broker-resident and needs no host-side durable path.\n\nWhat is **not** in JetStream, and is hosting-critical, is trust and authorization state a host must\nplace and keep:\n\n| state | class | where today | hosted injection |\n|---|---|---|---|\n| full `SpaceAuth` trust chain (`auth/broker.json` + `auth/account.<key>.json`, composed; a stripped signer bundle may instead be mounted at the legacy `auth/auth.json` key) | signing authority | `SecretStore` | `SecretStore` (manager + renewal) |\n| auth kinds: callout account/creds/xkey, issuer private keys, owner-derivation secret, data-signer projection | signing/identity authority | four `SecretStore` kinds | `SecretStore` (auth-service) |\n| `delivery.creds` | standing scoped cred | `SecretStore` or `--creds` | `SecretStore` (delivery) |\n| actor ledger, IdP pin | authorization + trust config | ambient `userAuthStateDir(findCotalRoot(), space)` | none (root-relative; not `store`/`COTAL_HOME`) |\n| `membership-rw.creds` | standing scoped cred | `SecretStore` | `SecretStore` (delivery + manager renewal) |\n| membership-observer / connection-evictor creds + `membership.json` | scoped $SYS creds / config | workspace filesystem | none (see gap 1) |\n| manager agent creds, actor tokens, sentinel creds | lifecycle authority | `SecretStore` | `SecretStore` (manager `secretStore`) |\n| `~/.cotal/meshes/space.<key>.json` record (holds IdP trust pins/root pointers) | non-secret, integrity-critical | machine home | process-global `COTAL_HOME` only |\n| auth-health, renewal records | non-secret diagnostics | workspace filesystem | `workspaceRoot` |\n\nThe `SpaceAuth` trust chain and the auth-service store kinds are **separate** identities/projections,\nnever parts of one document. `auth-service.json` (the live exchange capability) is ephemeral runtime\nstate, not durable, but is sensitive while the daemon runs. `@cotal-ai/workspace` is machine-local\noperator tooling by design; personas, PID files, and the `current-mesh` pointer are truly local and\nmust **not** sit on a hosted durable path. Everything classed above as an authority is what a hosted\ncomposition must provision and persist: signer-bearing server secrets now have `SecretStore` seams;\nthe remaining non-injectable rows are the explicit ambient `workspaceRoot`/cwd paths above.\n\n## See also\n\n- [Substrate stability](stability.md): what v0.3 and the 0.x packages guarantee, and the projected v0.4 break.\n- [Identity and auth](identity-and-auth.md): the profile matrix, the signer, and the IdP callout contract.\n- [Delivery daemon](delivery-daemon.md): the Plane-3 durable backstop.\n- [Deploy](deploy.md): the reference container against an external broker.\n"
|
|
167
|
+
"body": "# Embedding Cotal\n\n> **Guide** (informative) · **For:** implementers building a service on top of Cotal · **Prereqs:** [Architecture](architecture.md), [Identity and auth](identity-and-auth.md), [Delivery daemon](delivery-daemon.md)\n\nThe `cotal` binary in this repo is one composition root: an operator CLI. A separate service\n(for example a hosted, multi-tenant Cotal) does not fork this repo. It writes its **own**\ncomposition root that depends on the published `@cotal-ai/*` packages and imports the surfaces it\nwants. `bin/cotal.ts` uses the same composition pattern. This page is the contract for that: what is a real library\nexport you can build against, how to boot the server-side daemons from those exports, and where the\ncurrent export surface stops short of a fully hosted composition.\n\nThis is the \"guarded substrate\" boundary in practice. Nothing here reveals or assumes a specific\nhost; it documents the public seams any embedder composes.\n\n## What you embed\n\nThe supported reference shape here is **one broker operator serving one space** (one tenant: a\ndedicated data account, under an operator that also holds the system account and a quarantined\nauth-callout account) plus three standalone processes. The trust layer itself composes many spaces\nunder one broker operator today (`createBrokerAuth` + `createSpaceAccountAuth` + N-space\n`serverConfig`); what does not exist yet is the per-space **lifecycle** on a shared broker (see\n[Known gaps](#hosted-composition-gaps)). The three processes:\n\n| daemon | package | what it is |\n|---|---|---|\n| auth-service | `@cotal-ai/auth` | the NATS auth callout, the IdP token exchange, and JWKS. Plane 1 to Plane 2. |\n| delivery | `@cotal-ai/delivery` | the Plane-3 durable backstop: fan-out writer plus trusted reader, per space. |\n| supervise | `@cotal-ai/manager` | the per-machine agent lifecycle (spawn/despawn/attach), per space. |\n\n`mint`, `deliver`, and `auth-service` expose their behavior as direct library primitives, and the\nsupported one-space bootstrap below re-composes from exported low-level primitives. `supervise` and\nthe full `up` orchestration are **not** public runners: `up` also does broker bring-up, restore,\nprocess and registry management, and lifecycle work, and `supervise`'s orchestration is private (see\n[Supervisor signing authority](#supervisor-signing-authority)).\n\n## The export surface\n\nEverything below is a real export of a published package, reachable from the package root (each\npackage publishes only `.` via `dist/index.{js,d.ts}` and ships `files: [\"dist\"]`). Type-only names\nare marked; import them with `import type`.\n\n**Daemon runners and lifecycle**\n\n| symbol | package | purpose |\n|---|---|---|\n| `runAuthService(args, store?)` | `@cotal-ai/auth` | boot the auth-service daemon; `store` injects the secret material. |\n| `runDelivery(args, store?)` | `@cotal-ai/delivery` | boot the delivery daemon; `store` injects the scoped `delivery` cred. |\n| `startAuthService(inputs)` | `@cotal-ai/auth` | start one account-scoped auth-service context and return an `AuthServiceHandle` with the loopback `url`, the per-start `cap`, `readiness`, `drain`, and idempotent `close`. `close` rejects when the context did not release its plane claim: the claim row was no longer its own, or the release write failed. The row then stays held, and the next start reclaims it through the liveness oracle. With the optional `publicFace` input it also serves the public exchange face and carries `publicUrl`. With the optional `platformControl` input the handle also has `platformControlAuthority`, the in-process platform control door, `platformControlReadiness`, its read-only readiness read, `observeManagerGate`, the manager gate read a host composing the delegated user intent decisions passes them, and `activateManagedLifecycle`, the activation that host runs at a delegated launch's pinned lifecycle UID. With `platformControl.host` it also has `registerHostIncarnation`, which registers the host process's own endpoint instance and returns the incarnation a delegated execution pins, `observeHostGate`, the read of that endpoint's issuance gate, and `awaitHostFence`, which resolves once a later registration or barrier fences an incarnation. `runAuthService` remains the CLI entry. |\n| `PlatformControlAuthorityRequest`, `PlatformControlInnerRequest`, `PlatformControlAuthorityResult`, `PlatformControlAssignment` *(types)* | `@cotal-ai/core` | the closed envelope, its inner request union, its result and the backend's assignment row for `platformControlAuthority`. `platformControlOwner` in `@cotal-ai/auth` derives the `p_` owner the door issues under. |\n| `startDeliveryService(inputs)` | `@cotal-ai/delivery` | start one account-scoped delivery instance and return a `HostedServiceHandle` with `readiness`, `drain`, and idempotent `close`. The process runner remains the CLI entry. |\n| `deliveryCredsKey(space, composition)`, `membershipRwCredsKey(space, composition)` | `@cotal-ai/workspace` | build the secret-store keys the delivery cred and the membership feed's rw cred are read/re-signed under. Keys are **per-space**: `space.<hex>/<kind>`. A hosted composition passes `{ injected: true }`. |\n| `retireManagerInstanceIdentity(root, space, expected)` | `@cotal-ai/workspace` | remove a persisted manager identity only if its complete instance id and serve identity still match `expected`. Returns `removed` or `absent`; refuses malformed, nonregular, and changed records. `absent` is not proof of ownership or successful teardown. The caller must separately prove stop and retirement ownership before using it. |\n| `DELIVERY_CREDS_KIND`, `MEMBERSHIP_RW_CREDS_KIND` | `@cotal-ai/workspace` | the operator-facing KIND names (`delivery.creds`, `membership-rw.creds`) those keys are built from, and what renewal results report. A kind is **not** a key: putting a cred under the bare kind writes the pre-0.4 flat location, which nothing reads. |\n| `Manager`, `ManagerOptions` *(type)* | `@cotal-ai/manager` | construct and run a supervisor in-process; `ManagerOptions.secretStore` injects the one store it reads/writes every secret through. `ManagerOptions.remoteAuthority` is the hosted manager-service authority bundle, including host-owned release, retained-validation, goal-index, and serve-time admin-authorization callbacks. |\n| `ManagerOptions.pooled` | `@cotal-ai/manager` | require signerless remote authority and an explicit non-custodial runtime before local execution starts. A pooled composition must supply the assigned account key and all-duty renewal callback; the CLI's default remains unchanged. A signed-in human's manager gets that material from `managerServiceAuthority`. A platform-run control manager gets it from `AuthServiceHandle.platformControlAuthority` with the shipped `remoteManagerClient` builders, as the [platform control authority](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/platform-pooled-control-authority.md) design describes. |\n| `createRuntime`, `Runtime` *(type)* | `@cotal-ai/manager` | resolve the spawn backend (pty built in). |\n| `liveKvEntries(kv, filterOrOptions?, options?)`, `LiveKvEntriesOptions` *(type)* | `@cotal-ai/core` | read live KV entries in one finite scan. Pass `{ signal }` as the second argument or after a key filter to cancel. An interrupted scan throws `IncompleteKvScan`; cancellation throws the signal reason, including during an empty-bucket bind. The scan deletes only its owned consumers, including each one nats.js rebuilt from it, after the broker has answered every create those rebuilds sent. A not-found delete counts as gone, a refused one leaves the consumers to broker inactivity expiry (which also covers a crash), and any other delete failure is thrown, as is a not-found delete of a consumer whose create got no reply from the broker before a timeout or a closed connection. A cleanup failure is thrown only when the scan would otherwise return; the scan's own error, its cancellation reason and `IncompleteKvScan` take precedence. |\n\nThe remote manager authority parser accepts `renewStandingBundle` and `renewRunDriver` only with\nan assigned account nkey, the current manager process epoch, and a host-authenticated registration\nproof. A run renewal also names its active holder, takeover, epoch, fencing token, and the two\nexisting nkeys. The host must fresh-check those coordinates against its registration gate and run\njournal before issuing server-selected profiles. A host without that renewal authorization refuses\nthe request. Until the host issuer wires the operations and validates them on real connections,\nthe presence of these types is not an operational pooled renewal guarantee.\n\nWith `renewStandingBundle` configured, the manager renews all five standing credentials together.\nIt checks every returned credential for the held nkey and assigned account, test-connects each one,\nand adopts them only if the serve epoch has not moved. A refused or failed candidate leaves the\ncurrent credentials in place and records the refusal as cleanup debt until a later renewal succeeds.\nOn shutdown, after the standing context is drained, an expired maintenance executor is renewed\nthrough the existing scoped host operation so deregistration can finish without restarting duties.\n\n**Provisioning and minting** (all `@cotal-ai/core`)\n\n| symbol | purpose |\n|---|---|\n| `createBrokerAuth(label)` | mint BROKER trust: the operator and system account one nats-server trusts. One per broker, shared by every space on it. |\n| `createSpaceAccountAuth(broker, space)` | mint one space's own data account, signed by that broker's operator: the add-a-tenant primitive. |\n| `createSpaceAuth(space)` | the one-space convenience: broker trust + one account in a single composed bundle. |\n| `setupSpaceStreams({ servers, space, creds })` | create the space's JetStream streams. |\n| `ensureDefaultDeliveryClass({ servers, space, creds?, deliveryClass })` | write the space's default delivery class at creation so it is wire-discoverable (SPEC section 4). |\n| `serverConfig(broker, spaces, { storeDir, maxFileStore?, extraAccounts?, port?, host? })` | render the broker config: one operator, N space accounts. `storeDir` is required, `maxFileStore` caps JetStream file storage in bytes (omitted, nats-server's dynamic default applies), and `extraAccounts` preloads the auth-callout account. |\n| `mintCreds(auth, identity, profile, opts?)` | mint a scoped cred for any `Profile`. |\n| `mintMembershipObserverCreds`, `mintConnectionEvictorCreds` | mint the membership/eviction scoped creds. |\n| `provisionAgent`, `provisionAgentDurables` | create a principal's bind-only durables. |\n| `newIdentity`, `stripSpaceAuth` | a fresh nkey identity; a stripped signer bundle (data signing seed only). |\n| `Profile`, `CredentialKind`, `MintOpts`, `SpaceAuth` *(types)*, `CREDENTIAL_LIFETIMES` | the profile matrix and cred lifetime policy. |\n\n**Auth building blocks** (all `@cotal-ai/auth`)\n\n| symbol | purpose |\n|---|---|\n| `createCalloutAuth`, `startAuthCallout` | the NATS auth-callout responder. |\n| `createUserTokenIssuer`, `pinnedJwksResolver` | mint and verify the Cotal user bearer. |\n| `createIdpBridge` | exchange a verified IdP JWT for a Cotal bearer (see [the callout contract](identity-and-auth.md#the-idp-callout-contract)). |\n| `deriveOwnerToken`, `validateUserToken` | owner derivation; strict bearer validation. |\n| `cotalAuthProvider` | the self-registering `auth-provider` extension. |\n| `ensureCalloutAuth`/`loadCalloutAuth`, `ensureIssuer`/`loadIssuer`, `ensureOwnerSecret`/`loadOwnerSecret` | read/write the auth secret kinds through a `SecretStore`. |\n| `PLANE_CLAIM_REFUSED`, `planeClaimRefusal`, `PlaneClaimRefusal` *(type)* | every plane-claim refusal carries a `PLANE_CLAIM_REFUSED` detail, and `planeClaimRefusal(err)` reads its `reason`: `corrupt`, `live-peer`, `unknown`, `concurrent`, `fenced`, `released` or `lost`. Only `unknown`, an inconclusive liveness observation, is coded `unavailable`. A host can retry contention and stop on a corrupt row without matching message text. |\n\n**Seams and the wire** (all `@cotal-ai/core` unless noted)\n\n| symbol | purpose |\n|---|---|\n| `SecretStore` *(type)* | the durable hosted-secret seam (get/put/delete); `get()` returns raw seeds/keys into process memory, so it is a blob seam, not HSM/KMS signing. |\n| `FsSecretStore`, `workspaceSecretStore(root)` | the filesystem default. **These live in `@cotal-ai/workspace`, not core.** |\n| `AuthProvider` *(type)*, `Connector` *(type)*, `Runtime` *(type)*, `Command` *(type)* | the extension contracts; implementations self-register on import. |\n| `registry` | the shared registry a composition root pulls surfaces into. |\n| `CotalEndpoint`, subjects, message types | the wire client and shapes. |\n| `ParsedArgs` *(type)* | the shape the daemon runners take (see below). |\n\nFor a Linux Unix-socket adapter, `peerCredentials(socket)` from `@cotal-ai/seat` returns\nkernel-observed peer `pid`, `uid` and `gid`. Compare these against the host's authorization\npolicy; request-supplied identity and process liveness do not replace that policy or a\nlifecycle fence. The helper starts no custodian and refuses unsupported platforms or a\nmissing native helper.\n\nThe runners take a CLI-shaped `ParsedArgs`, not a typed options object, so a host fabricates one:\n\n```ts\nconst args: ParsedArgs = { values: { space, server, port: \"0\" }, positionals: [], raw: [] };\n```\n\nFor an embedded delivery instance, use `startDeliveryService` instead. Its `HostedContextInputs`\ninclude the account public key and lifecycle UID, space, broker URL, injected store, stable\n`storeIdentity`, and an explicit `stateDir`. The store must declare that same injected identity.\nThe initial delivery credential must belong to the assigned account. The function returns only\nafter the delivery responder is bound. `close()` withdraws serving and releases only the lease\nowned by that instance. It closes both membership connections even when a disconnected drain\nfails, so they cannot reconnect after closure. Credential-expiry health state clears after\nsuccessful broker-verified adoption through the existing `reloadCreds` rail. A failed start\nrefuses locally without exiting the host process or stopping another account's delivery service.\nIf a health fault occurs during an asynchronous store read, startup rejects when the read returns\nand closes any resources created by that late completion.\n\n`startAuthService` takes the same `HostedContextInputs`. The store must declare the assigned\ninjected identity, and its data account must be the assigned account. The IdP pin and ledger live\nunder the explicit `stateDir`, and the auth plane's instance identity in its `.cotal/space.<hex>/`.\nThe context never resolves a workspace root from the working directory and has no local manager, so\nonly remote manager gates can be selected. It returns after\nthe authority plane, the callout subscription and the loopback listener are bound. A start that\nfails closes the connections it opened and releases its plane claim, so a retry on the same space\ncan claim it. A fenced plane or\na lost broker connection makes that context `unavailable` and closes it without exiting the process.\nThe host writes no discovery file for it. The handle carries what `auth-service.json` holds for a\nCLI start: the loopback `url`, `publicUrl` when a public face runs, and the per-start `cap`. The cap\nalone authorizes the loopback host actions, lifecycle retirement and managed-agent enrollment\nverification, so keep it in the authority process. `publicFace` takes the CLI's public face inputs\n(`port`, `url`, `trustedProxy`, `advertisedServer`, `agentProvisioningUrl`) under the same rules,\nand a face without a port refuses to start.\n\nTwo optional inputs serve a platform composition. `platformControl: { observeAssignment }` adds\n`platformControlAuthority` to the handle. It is a typed in-process method, served on no listener,\nthat issues the manager-service request family for the one control manager the backend assigned to\nthis account, under a derived `p_` owner. It reads the assignment fresh on every call and refuses\nan IdP token, another account, a stale revision, another instance or lifecycle, and an instance\nanother owner registered. It refuses `prepare` and `activate` while the assignment's named\npredecessor is still registered or frozen. The same input adds `platformControlReadiness(instanceId)`,\nthe route for a host that needs to know whether its assigned control manager is serving. It returns\nthat instance's attributed `status` reply and refuses any instance the current assignment does not\nname, or one whose gate another owner holds. It reads over the context's own connection, whose\ngrant is the assigned instance's `describe` and `status` and its own reply rail. That connection\nrenews in process like the context's other connections and never leaves it, so the host lends no\nhuman or operator credential to a worker, mints no control instrument per read, and does not read\nliveness off the manager process. Without the input both members are `undefined`.\n`standingRenewableTtlSeconds` is forwarded unchanged to the authority plane, which bounds it to 5\nto 86400 seconds. It is a trusted-host input for the renewal rehearsal, and no request or CLI flag\nsets it. SPEC §13.1 and §13.6 define the view.\n\nThe auth plane can renew a registered manager's\nfive standing credentials from the current service registration. Run-driver renewal still\nrefuses without an authoritative activated-run reader, so these handles do not yet make a\ncomplete pooled auth and delivery host. A fresh auth plane can initialize without a\ndelivery-admin responder. Reclaiming a held claim from a dead predecessor needs the delivery\ninstance first: its admin rail must complete the broker connection-liveness sweep before the\nauth plane takes the claim. An absent or inconclusive oracle refuses the reclaim.\n\n### Remote manager client composition\n\n`@cotal-ai/manager` exports the `remoteManagerClient` namespace, containing the stock remote\nrequest builders and response validators, and `registerRemoteManagerAuthority` for registration\nwith a host-issued prepare credential. It returns the process epoch and registration revision its\nown registration committed. When a later start of the same instance registers before this start\nauthorizes its serve grant, this start is refused with `expired`.\n`RemoteManagerIdentityState` describes the five private\nmanager identities stored under an explicit account-local root. Use these public exports when\ncomposing `ManagerOptions.remoteAuthority`; do not copy CLI validators or import private modules.\n`managerClusterArtifacts()` returns the canonical document, manifest and their digests used by\nregistration. Pass its `[document, manifest]` pair as `contractArtifacts` to both\n`remoteManagerRegistrationProof(owner, state, contractArtifacts)` from `@cotal-ai/core` and\n`remoteManagerClient.remoteManagerAuthorityRequest(state, actor, \"activate\", { registrationProof, contractArtifacts })`.\nThe request builder takes each operation's coordinates as named fields of its last argument.\nThe host still validates the artifact closure and current registration before activation.\nThe namespace includes closed standing/run renewal, admission, maintenance, enrollment and\nretirement helpers. `remoteRunHosting` builds the four `runHosting` callbacks from the registration\nand the transport you supply for the host's run admission, run attempt and authority requests, so\nyour manager sends the run requests the stock manager sends. The namespace provides no signer or\nnew grant. The host still owns authenticated issuance, current registration and activated-run\nobservations, and any guarded foreign-holder repair.\nAn embedding must preserve those checks and supply a supported runtime; the client exports alone\ndo not provide a pooled runtime, an authority service or a complete hosted context.\n\n### Long-lived endpoints take a bearer function\n\n`EndpointOptions.bearer` accepts either a string or a function, and the difference is not stylistic.\nA string is minted once, so when it expires (which it will: callout bearers live minutes) the\nendpoint has nothing to renew with. It will not present the dead token to the broker, since that is\na guaranteed denial that still costs a full auth-callout round trip. It refuses to reconnect, emits\n`warning` saying which case it is in, and retries on a widening backoff until the process\nre-authenticates and rebuilds it. Retry notices use `warning` rather than `error` because Node\nrethrows an unhandled `error` event and would kill a host the endpoint is still trying to recover.\n\nPass a **function** for anything that outlives one bearer. That is a renewal source: it is called\nahead of each expiry and again whenever a reconnect finds the cached bearer dead, and it requires\nexplicit `card.owner` and `card.actor`. The first-party surfaces already do this\n(`UserViewAuth.source`, the connector's `agentBearerCommand`). A string bearer is for a short\none-shot connection.\n\nLong-lived hosts must also subscribe to the endpoint's `warning` event. It carries conditions the\nendpoint is surviving, including failed credential renewal, reconnect retries, and a durable leave\nit keeps retrying after the broker refuses a durable channel's live subscription. A host may choose\nto ignore warnings for a one-shot endpoint whose awaited operation owns the verdict, but that choice\nshould be explicit. An unhandled warning is nonfatal and silent.\n\nThe `error` event carries a fault the endpoint cannot return from a call, such as a refused\nsubscription or a refused publish that no request was waiting on. Attach a listener before `start()`,\nsince Node throws on an unhandled `error`. A denial the broker returns to a request, such as an\nobserver's read of the DM stream, reaches only that call, which decides what it means, and is not\nemitted again as an `error`.\n\n## Booting the daemons\n\n### auth-service\n\n`runAuthService(args, store?)` reads its provisioned long-lived secret kinds (service keys, callout\naccount, issuer keys, owner secret) through the injected `SecretStore`; a host provisions those into\nthe store first. It is a **signer and identity authority**, not a scoped daemon: at runtime it holds\nthe data-account and callout-account signing seeds, the issuer's private JWKs, and the\nowner-derivation secret in process memory (`SecretStore.get` exports raw values). The IdP pin and the\nactor ledger are **not** store-injected: `runAuthService` resolves them under\n`userAuthStateDir(findCotalRoot(), space)`, a path relative to the process working directory, so a\nhost provisions those into that exact directory (neither `store` nor `COTAL_HOME` selects it). It\nalso writes an ephemeral `auth-service.json` discovery file there that carries the live exchange\ncapability. That file appears only after every plane is bound, so waiting on it is the readiness\nsignal: a host that also passes the daemon's pid to the provider's `ready()` gets a process-bound\nwait. The wait extends past the base timeout while that pid is alive, up to a fixed bound, and it\nends at once when the pid exits.\n\n```ts\nimport { runAuthService } from \"@cotal-ai/auth\";\n// store implements SecretStore over your secret backend; get() returns raw seeds into memory.\n// Provision the auth secret kinds into the store, AND the IdP pin + actor ledger under\n// userAuthStateDir(findCotalRoot(), space), before this call.\nawait runAuthService(\n { values: { space, server: brokerUrl, port: \"8081\" }, positionals: [], raw: [] },\n store,\n);\n```\n\n### delivery\n\n`runDelivery(args, store?)` runs from a **pre-minted scoped `delivery` cred** and never loads the\nsigner. Provide the cred either through the injected store (under\n`deliveryCredsKey(space, { injected: true })`) or with a\n`--creds` file; the two are mutually exclusive. The daemon re-fetches the cred from the store at 75%\nof its JWT lifetime and fails loud rather than riding to expiry, so **something must re-sign a fresh\ncred into that same store**. When that read finds the previous generation still there, the daemon\nreports the missed remint and retries in 60 seconds. The current cred stays live until its expiry,\nand the store is read once per retry rather than once per second.\n\n```ts\nimport { runDelivery } from \"@cotal-ai/delivery\";\nawait runDelivery({ values: { space, server: brokerUrl }, positionals: [], raw: [] }, store);\n```\n\nThat renewal is a **signer** operation, not the delivery daemon's:\n`remintDaemonCreds(root, space, store?, { preflight? })` (`@cotal-ai/workspace`) reads the `SpaceAuth`\nsigner **through the same resolved `store`** (`getSpaceAuth(store ?? workspaceSecretStore(root), space)`,\nkeys `auth/broker.json` + `auth/account.<key>.json`; the pre-split `auth/auth.json` monolith is\nmigration input and the container signer mount only) and re-signs the daemon creds (`delivery.creds` and the membership feed's\n`membership-rw.creds`) back into that store. The injected `store` is both the signer source and the\ncredential destination, never a split. `space` is **required** and validated against the store's signer, so a\nstore swapped to a different space cannot re-sign over the wrong broker's creds. `preflight` is a\ncaller-supplied proof that the broker accepts the credential. The reference `Manager` passes a\n`probeConnect` over its `servers`. It gates **every** candidate before overwriting the last-good,\nwhether the signer is a full bundle or a stripped projection: a bundle's JWT chain proves only that\nit is self-consistent and\nnamed the space, NOT that its account is the broker's *current* account for that space (two\n`createSpaceAuth(space)` calls yield same-named, different-account chains), so a same-label alternate\nsigner would otherwise mint a broker-dead cred and clobber the good one. The offline local repair (`doctor auth --fix`) has no preflight. It permits the overwrite only\nunder **authority continuity**: the candidate must be signed by the same account signing key (`iss`) as the current\n(already broker-accepted) cred. A same-label alternate account breaks continuity and is refused, full or\nstripped; a legitimate local re-sign is continuous and proceeds without a network. The reference\n`Manager` runs it on a schedule against its **own**\n`secretStore` (see below), so passing the manager and the delivery daemon the *same* store closes the\nrenewal loop end-to-end on an injected backend: the manager reads the signer from the store, re-signs\ninto it, and the daemon adopts each generation on a preflight-proven 75% timer. The stock\ncross-host composition cannot satisfy that by writing one filesystem and fingerprinting another:\n`Manager.start()` and every later remint challenge the daemon's `reloadStoreIdentity` and a\ndivergent pair is refused naming both stores. The identity is the store the daemon actually\nreloads: an injected coordinate, the workstation root only when `--creds` is\n`<root>/.cotal/<spaceSegment(space)>/delivery.creds` (matching the canonical arm), the\nfile's own directory for any other `--creds` path, or the workstation root. A filesystem\nstore is also named by a random id it records in `store.id` inside its own directory, so two\nhosts that use the same root path are two stores. No key and no `--creds` file may be that\nfile under any name, and a `store.id` that is a symbolic link or holds anything but a lowercase UUID is refused. Uninjected\n`--creds` that names one real workstation while process cwd resolves another is refused\nat start, naming both, because membership-rw still uses `findCotalRoot`. A `--creds`\npath that is not under any `.cotal` tree is not that case and is not refused here. It never\nwalks ancestors with `findCotalRoot`. No bound daemon is not a named\nstore, so start proceeds; a later daemon on a foreign store is refused on the next remint.\nThe first-party filesystem adapter declares its workspace-root identity on the store itself. Other\ninjected adapters declare their stable coordinate on `SecretStore.identity`, or name it in\n`COTAL_SECRET_STORE` on both processes. It never throws: it\nreturns per-file results (`skipped: \"no-auth\"` when the store holds no signer records),\nso the caller must check them or the cred still rides to expiry. A composition whose signer lives in\nKMS/Vault simply injects that store; no bespoke renewal is needed. A `--creds` file path must be\nreplaced atomically before the 75% read. The signer can now be injected behind the store seam, which\nresolves custody. The remaining hosted gap is signer **isolation**. The seed is still decrypted\nin-process at the manager's uid, so it needs an OS sandbox or remote signer.\n\n### Supervisor signing authority\n\n`@cotal-ai/manager` exports the `Manager` class; there is **no** `runSupervise(opts)` runner. The\nprivate CLI `runManager` also does broker-reachability checks, space/default resolution,\nroster/launch parsing and materialization, installed-extension resolution, signal handling, staged\npre-spawn, and the forever wait. A host composes that lifecycle itself around `Manager`:\n\n```ts\nimport { Manager } from \"@cotal-ai/manager\";\nconst mgr = new Manager({ space, servers: brokerUrl, workspaceRoot });\nawait mgr.start(); // then wire your own SIGINT/SIGTERM -> mgr.stop()\n```\n\n`stop()` runs once. A later call joins the stop in progress and settles with it, and a call that\nasks for a different `withAgents` than the running stop is refused.\n\nUnlike delivery, the manager is **not** a pre-minted-scoped-cred daemon (auth-service is also a\nsigner: it holds fewer artifacts than the full trust bundle, but its data-account signing seed still\ngrants complete data-account mint authority on compromise, so this is not least-privilege). On `start()`\nthe manager reads its space's full trust chain **through its `secretStore`** (`getSpaceAuth(this.secrets,\nthis.space)`, composed from `auth/broker.json` + `auth/account.<key>.json`; a container may instead\nmount a stripped signer bundle at the legacy `auth/auth.json` key) and **self-mints** its supervisor cred and renewals from the\ndata-account signing seed. In static mode it also mints every per-agent cred from that seed; in user\nmode agents instead receive callout-minted bearers, but the manager still holds the signing seed for\nits own creds and renewal. So a hosted supervisor is a **trusted per-tenant account-signer process**,\nnot a least-privilege connect client. It additionally requires a `~/.cotal/meshes/space.<key>.json`\nregistry record and the workspace user-auth marker to start in user mode. `ManagerOptions.secretStore`\ninjects the one `SecretStore` the manager uses for **the signer itself (the split trust\nrecords)**, daemon-credential renewal (`remintDaemonCreds`), and per-agent secrets,\ndefaulting to the workspace filesystem store; pass the delivery daemon the *same* store for end-to-end\nhosted renewal. The store declares the same identity on both processes, or both set\n`COTAL_SECRET_STORE` to the same coordinate. The manager\nremints no daemon credential when the daemon names a different store, including a daemon that binds\nafter start; it keeps running and serving its own agents, so one space can carry a manager on more\nthan one workspace root. That manager also stays off the space's renewal lease, so a manager or a\n`cotal doctor auth --fix` on the daemon's store can still take it. Pointing several managers at one coordinate is safe: the store identity\nalone cannot pick an owner (it carries no holder and no tiebreak, so every manager sharing the store\nmatches), so the manager that also holds the space's renewal lease is the one that remints and the\nrest skip it. Without that lease two owners would remint on independent timers with no ordering\nbetween them, and one write would land between the other's re-sign and its fingerprint-only\n`reloadCreds`. `cotal doctor auth --fix` takes the same lease before it re-signs, so a live manager\nand a local repair never race each other either. The signer IS now injectable: a hosted composition injects a KMS/Vault store and no\nsigning seed lands on the hosted disk. What remains is signer **isolation**. The seed is decrypted\nin-process at the manager's uid. That issue needs an OS sandbox or remote signer; it is no longer a\ncustody problem. The other knobs are `workspaceRoot` and the process-global `COTAL_HOME`.\n\nA store, runtime or extension the manager calls may reject with any value, including `null`. The\nmanager logs or refuses with that value's `message`, or with the value itself as text when it has\nnone, and keeps serving. A value it cannot read reports as `an unreadable rejection`.\n\n> Scope note: the **static-auth** operator paths (`cotal spawn`/`join`/`status`/`web`, via\n> `mesh-target` → `connect`/`preflight`) still read the signer from the local split records (sync\n> `loadSpaceAuth`). That is the single-machine composition, where the signer is on local disk by the\n> static-auth model; multi-tenant hosting runs **user mode**, which never mints from on-disk trust. The\n> store-injectable signer path is the hosted-server set: the manager, `remintDaemonCreds`, and delivery.\n\nThe typed remote-manager authority contract includes a one-shot terminal phase. A host implements\n`remoteAuthority.prepareAgentRetirement` to revoke the managed grant and finish its resumable\nrelease while preserving the UID, then `remoteAuthority.mintRetirementRequester` returns the\nhost-signed JWT for a fresh participant-owned nkey. The credential is pinned to the authenticated\nowner, server-derived manager serve principal, current instance epoch, and exact target lifecycle.\nThe manager then uses the existing auth `retireLifecycle` rail with the operation id derived by\n`managedRetirementOpId(target.lifecycleUid)`. This derivation is the reference remote-Manager\ncomposition's closed contract, not a rule for every retirement entry point; interactive retirement\nkeeps its existing operation identity and remains compatible. The `retireLifecycle` rail independently\nrecomputes the managed id from its broker-pinned target before any gate, head, intent, or barrier\naccess, so mint-time validation is not the terminal boundary. A failure keeps\nthe alias held. This does not expose the auth barrier or give the participant signer authority.\n\nIf the participant disappears after prepare, the host finishes the retirement itself on the auth\nservice's loopback face: `POST /managed-lifecycle/retire` (exported as `MANAGED_RETIRE_PATH` from\n`@cotal-ai/auth`) with the `Bearer <cap>` from `auth-service.json` and only\n`{ owner, actor, lifecycleUid }`. It has the interactive door's guards (POST only, no `Origin`, JSON,\ncapability, closed body) and is never served on the public face. The managed grant must already be\nrevoked at that uid, or it answers 409. It runs the same `managedRetirementOpId(uid)` operation as\nthe rail, and the rail and the door share one in-process flight, so a late participant request and\nthe host call converge on one barrier.\n\n| lifecycle head | answer |\n| --- | --- |\n| absent, or `retired` at another uid | `200 { retired: false, lifecycleUid, notStarted: true }` |\n| `active`/`retiring` at another uid | `409` |\n| `retired` at this uid | `200 { retired: true, lifecycleUid, alreadyRetired: true }` |\n| `active`/`retiring` at this uid | the barrier runs, then `200 { retired: true, lifecycleUid }` |\n\nDeprovisioning durables stays with `deprovisionAgent` and a `deprovisioner` credential. Pass `memberChannels` to both to also purge the retired lifecycle's durable membership rows on those concrete channels. The manager fills that list from the launch's concrete read channels and the delivery daemon's read-only `lifecycleMemberships` admin verb. When that verb cannot answer, rows on other channels stay retained, the teardown logs the inventory as incomplete, and retirement remains held pending retry rather than releasing the alias. Each teardown examines two exact consumers, one ACL key and the named member keys. KV deletion uses a native revision condition; a lost condition with a live replacement refuses rather than claiming absence. Consumer INFO checks before and after DELETE distinguish verified prior absence from disappearance. `acknowledged` counts native DELETE success replies. `disappeared` counts observed live-to-absent consumers and live KV rows whose conditional purge lost to a competing deletion. These KV rows were present at the first read, so they never count as prior `absent` or as this caller's `deleted`. The ACL and membership subtotals preserve that distinction. Neither establishes which concurrent caller uniquely removed a consumer, so `consumers.deleted` and the total `deleted` are `null` when a consumer disappears without a winner token. A repeat after verified absence reports zero, not an invented deletion. `refused` counts slots whose cleanup or state remains uncertain. A partial failure raises `DeprovisionError` carrying these bounded observations. The Manager's static sweep preserves unknown uniqueness rather than adding acknowledged requests as physical removals; its slot totals remain separate.\n\nA host that resumes retained managed actors also implements\n`remoteAuthority.validateRetainedAgent`. The participant sends back the actor token and sentinel it\nalready holds, plus the `nextRegistrationProof` returned by the activation response. That proof is\nhost-issued after registration and binds the manager owner, actor, lifecycle, identity nkeys, current\nregistration revision, and serving epoch. The host checks it against the current open manager gate,\nvalidates the retained secrets against its current managed row, and returns only the non-secret\nauthority shape. The manager binds every result coordinate and the returned authority back to its\ninventory before use. Do not copy the provider's `issuer.json` or `callout.json` into the participant\nstore. Both contain private signing or exchange authority.\n\nThe same composition supplies `remoteAuthority.agentBearerExchangeUrl`, the pinned public auth-service\nbase used by retained children. Remote adoption launches `agent-bearer --exchange-url <base>`; it must\nnot select the local `--dir` arm, which depends on a host-only auth-service process record.\n\nA host that lets a remote participant spawn FRESH managed agents implements\n`remoteAuthority.enrollManagedAgent`. The participant generates the standing actor token, writes it\nat mode 0600, and passes only its SHA-256 digest with the requested actor, label, role,\ncapabilities, and channel lists, so the plaintext secret never leaves the participant machine. There\nis deliberately no `lifecycleUid` input: the host selects the UID, because only the host sees the\nretirement tombstones that make a UID permanently unusable, and a participant-chosen UID could aim a\nfresh grant at a dead incarnation. The host authors the ledger grant, pre-creates the lifecycle-keyed\ndurables, clamps the requested lists to what the spawning owner already holds, and returns the owner,\nactor, chosen `lifecycleUid`, the space sentinel credentials, the effective lists, and\n`agentBearerExchangeUrl`. The manager binds every returned coordinate, re-keys the secret family onto\nthe returned UID, and launches `agent-bearer --exchange-url <base>`. When the hook is absent a\nsignerless manager refuses the user-mode spawn rather than authoring a local grant the host knows\nnothing about.\n\nBoth managed-agent operations ride the one verified `POST /manager-service-authority` transport as\n`kind: \"manager-managed-agent-enrollment\"` and `kind: \"manager-managed-agent-prepare-retirement\"`.\nStock `cotal auth-service` answers both itself, because it owns the actor ledger and the space's\nprovisioning authority. An enrollment writes the managed grant at a fresh UID with the supervising\nactor as its parent, provisions that UID's durables, and returns the daemon's public exchange URL as\n`agentBearerExchangeUrl`; a daemon started without `--exchange-public-port` refuses enrollment. A\nretry with the same token digest answers the same UID while the supervising actor's current grant\ncovers it, and a fresh enrollment's refusal otherwise. While the agent's grant stands, an enrollment\nwith another digest is refused with `conflict` until that lifecycle's retirement is prepared. A\nprepare-retirement releases the target UID's broker footprint and then revokes its grant, so the\nmanager's terminal rail finds the grant gone. A platform that keeps these writers in its own storage\nintercepts both kinds instead. It terminates its own public route, authenticates the human there,\nand asks the auth service for the decision at\n`POST /manager-service-authority/verify-enrollment` (exported as `VERIFY_ENROLLMENT_PATH` from\n`@cotal-ai/auth`) with the `Bearer <cap>` from `auth-service.json` and only `{ owner, request }`. That\ndoor has the managed retirement door's guards, derives the caller's scope from the local ledger rather\nthan the body, checks the manager gate and registration proof in-process, and answers\n`{ authorized: true, owner, actor, instanceId, serveEpoch }`. `authorizeRemoteManagedAgentEnrollment`\nand `authorizeRemoteManagedAgentPrepareRetirement` are exported too, for a host that composes the\ndecision without the HTTP hop. Both require ledger scope `supervise`; `spawn` and `admin` do not\nimply it.\n\nA host that runs managed agents on its own hosted runtime adds two more kinds on the same transport.\n`kind: \"manager-managed-agent-runtime-create\"` asks the host to create the runtime for one agent it\nalready enrolled, and `kind: \"manager-managed-agent-runtime-status\"` reads that runtime's state. Both\ncarry the manager envelope plus `target: { owner, actor, lifecycleUid }`, the coordinate the\nenrollment returned. Both schemas are closed. An unknown top-level or target field, including\n`providerRef`, `handle`, or `name`, is refused as `bad-request`, because the host alone issues and\nholds provider references. There is no stop, adopt, or probe kind: stop goes through\nprepare-retirement. Stock dispatch refuses both kinds with `unimplemented`, and the verify-enrollment\ndoor decides them. `authorizeRemoteManagedAgentRuntimeCreate` and\n`authorizeRemoteManagedAgentRuntimeStatus` apply the enrollment door's checks: host space, a target\nowner equal to the authenticated owner, the manager actor's own ledger row with `supervise`, the open\ngate, the current serve epoch, and the registration proof. Each returns only\n`{ owner, instanceId, actor, target }`. The door touches no provider and writes nothing. The host\nmatches the decision to its own intent record and performs the create afterwards. The host answers\nwith `state` (`reserved`, `creating`, `bound`, `create-unknown`, `closing`, or `closed`), `readiness`\n(`ready`, `bound-not-ready`, or `none`), and an optional `retirementPhase`. A manager builds requests\nwith `remoteManagerClient.remoteManagedAgentRuntimeRequest` and binds the answer with\n`remoteManagedAgentRuntimeState`.\n\nAn enrollment result may also carry `runtimeIntent: { state: \"reserved\" }` when the host reserved a\nhosted runtime for the agent. It is display-only. Older hosts omit it, the manager binds both shapes\nto the same material, and nothing reads it as authority.\n\nNo stock door lets a platform control holder launch or retire an agent for a signed-in user. The\nmanaged-agent kinds above act only under the authenticated owner and refuse a caller that is not\nthat user, so a platform could only run a user's agent by holding the user's login or by enrolling\nthe agent under its own owner. Both are refused. The\n[delegated user launch intent](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/delegated-user-launch-intent.md)\ndesign and SPEC §13.16 define the smallest addition. The user admits one launch or one retirement\non the host's authenticated route. The holder consumes that intent once, from its current\nregistration, epoch and lifecycle. The host then enrolls the agent under the user's `u_` owner with\nthe user's own actor as its ledger parent, so the envelope walk, membership and channel lists match\nwhat the user's own manager would produce. Retirement keeps the prepare, provider closure and\nterminal barrier order, and the host finishes it when the holder is gone. A launch the host had to\nundo keeps its agent name held until the host process that ran it confirms it has stopped, and\nwhile the name is held the host also refuses it to the user's own manager. `@cotal-ai/auth` ships the\ntwo decisions, `authorizeDelegatedUserIntentAdmission` and `authorizeDelegatedUserIntentExecution`,\nfor a host that owns an intent store and those writers to compose on its own routes. Both read the\nholder's gate through the handle's `observeManagerGate`, present with `platformControl`, which reads\nover the context's own connection, so the host opens no second data-account connection. It is an\nobservation for the decision: the consuming CAS and the writers still apply their own checks. The\nconsuming CAS pins the incarnation of the host process that runs the flight as the executor, never\nthe auth plane's, because the two can restart independently. `platformControl.host` names that\nprocess's reverse-DNS endpoint, the closure digest of its §13.7 cluster and the contract artifacts\nregistration reads, and the auth plane self-authorizes that one name. Registration reads the closure\nmanifest `{ v: 1, root, members }` at `clusterDigest`, then the cluster document at the manifest's\n`root`, and verifies each against its digest. `artifacts` therefore carries both, and `clusterDigest`\nis the digest of the manifest. `members` stays empty: SPEC §13.7 lists every reachable artifact\nthere, but this implementation registers single-document clusters only and refuses a manifest that\nlists members. `singleDocumentClosure(document)` returns that manifest and its closure digest. The\ninstance id is a lifecycle token, `[a-z0-9]{26,32}`. The minimal construction below has one command\nover the void schema. A host copies it, replaces `document` with its real cluster, and passes `host`\nas `platformControl: { observeAssignment, host }`.\n\n```ts\nimport { mintLifecycleUid, singleDocumentClosure, VOID_SCHEMA_DIGEST } from \"@cotal-ai/core\";\n\nconst document = {\n urn: \"com.example.host\",\n revision: 1,\n attributes: [],\n events: [],\n commands: [{\n name: \"ping\", class: \"ephemeral\", targeted: false, capability: \"host.ping\",\n inputDigest: VOID_SCHEMA_DIGEST, outputDigest: VOID_SCHEMA_DIGEST,\n }],\n};\nconst { manifest, closureDigest } = singleDocumentClosure(document);\nconst host = { endpoint: \"com.example.host\", clusterDigest: closureDigest, artifacts: [document, manifest] };\nconst instanceId = mintLifecycleUid(); // first start only; later starts reuse the persisted id\n```\n\n`registerHostIncarnation(instanceId)` publishes the artifacts, registers that instance through the\nceremony the plane runs for itself, and returns `{ instanceId, processEpoch }` with the epoch that\nregistration committed. The host calls it at every start with its persisted instance id, before it\nadmits or recovers any flight, so a restart fences its predecessor. The first registration of an\ninstance commits epoch 0, which is open and serving like any later epoch, and each later start of\nthat instance commits the previous epoch plus one. A consumer compares epochs for equality and never\nreads 0 as absent or not ready. A start\nwhose confirming read of the gate finds that a later start of the same instance registered is\nrefused with `conflict`. The returned epoch is a committed coordinate and stays current only until\nthe next start registers, which can happen before the call returns. `observeHostGate(instanceId)`\nis a point-in-time read of an executor's gate, and answers null for an absent gate; a sweeper\ndecides on it. `awaitHostFence(instanceId, processEpoch)` resolves with the gate once it is no\nlonger open at that epoch, and with null once it is absent. It takes any non-negative safe integer\nepoch, 0 included, and refuses any other value with `bad-request`. The host arms it with its returned\nincarnation before it admits or recovers any flight, and stops serving when it resolves. It polls\nthe gate, because no runtime credential may watch the auth bucket. The\nhost's launch writer first activates the agent's lifecycle at the pinned UID through the handle's\n`activateManagedLifecycle`, before any ledger row or durable, and its compensation runs the same\ncall before the terminal barrier, so a launch whose agent never exchanged its bearer still reaches\nthe terminal barrier at that UID. The launched agent exchanges its bearer on the context's public\nface, so the host starts it with `publicFace`, and its retirement writer ends at\n`POST /managed-lifecycle/retire` with the handle's `cap`. Stock dispatch refuses both kinds as\n`unimplemented`. The holder's composition passes\n`remoteAuthority.executeDelegatedUserIntent`, which posts the execution request and binds the answer\nwith `parseRemoteDelegatedUserIntentExecutionResult`. It then starts the agent with\n`startAgent({ ..., delegatedIntent: { intentId, owner, parent } })` and retires it with\n`retireDelegatedAgent(name, intentId)`, which stops the agent only after the host confirms\n`retired: true` for its exact target. A delegated agent's stop, exit or failed launch keeps its name\nheld until that retirement confirms.\n\nRemote user-mode managers must also supply `remoteAuthority.authorizeAdmin`. The manager builds each\nrequest only from the caller tuple parsed from the broker-authenticated endpoint subject, then relays\nthat tuple over the current registered manager lifecycle. HTTPS does not separately authenticate the\nrelayed caller. The host authenticates the manager operator, binds the request to the current open\nmanager gate, registration proof, serving epoch, and identity nkeys, then reads the caller's unified\nauthoritative row fresh. It returns only the manager owner and `authorized: boolean`, with every request\ncoordinate echoed. Missing, revoked, narrowed, foreign-owner, and stale-lifecycle callers all return\n`false`; malformed coordinates or corrupt and unavailable authority state fail the operation. The\nparticipant never reads or mirrors the host ledger, and the remote branch has no local fallback. The\nsame callback gates all `manager.admin` handlers, any-mode cross-owner control, and `ps` or `inspect`\ncross-owner visibility. Launch keeps its owner-equality policy.\n\nThe remote authority's instance executor remains the scoped maintenance credential for clean service\nderegistration and exact instance registration operations. It carries no records-stream consumer\nlifecycle authority. The manager's boot `goalidx` sweep uses the authenticated host operation, which\nreturns parsed `goalidx.manager.<owner>.>` entries for that owner only. The host keeps the sealed\nconsumer connection and its create/delete rights. The five-minute executor already renews through\n`remoteAuthority.renewExecutor`. The signerless supervisor, serve, goal-writer, session-ledger and\nper-run driver credentials do not yet have a complete remote renewal and adoption path. The\n[hosted runtime contract](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/hosted-runtime-contracts.md) records the bounded additions and\ntheir ownership; it is not a shipped pooled service.\n\n**Signer isolation needs an OS sandbox.** The default pty runtime\nruns agent children under the *same* OS uid and the *same* `workspaceRoot`, so mode-0600 on\nthe trust records does not stop a hostile same-uid agent from reading their absolute paths. The reference\n[deploy](deploy.md) tree does not solve this: it mounts the signer into the agent's own container, so\nits phase-1 boundary isolates agents from each other, not the signer from the agent. A hosted\ncomposition must run the manager/minter that holds the signer in a different uid, container, or mount\nnamespace from the agent children, which mount no signer at all; that split is future\nhosted-composition work, so until it (or a remote/injected minter) exists, do not run untrusted\nagents under this manager.\n\n### Delegated seats outside the manager's filesystem\n\nThe [portable lifecycle bootstrap](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/portable-lifecycle-bootstrap.md)\ndesign and SPEC §13.17 define how a managed agent that `enrollManagedAgent` already enrolled starts\nin a child that cannot see the manager's filesystem. The ordinary `spawn` path writes the token and\nsentinel under the manager's workspace root and hands the runtime a launch whose bearer command and\nmaterial file are paths on that filesystem, so such a child needs this path instead.\n\nThe delegation boundary is one optional runtime method. A runtime that implements\n`Runtime.spawnDelegated(launch, handoff)` receives two values and no paths: a `DelegatedSeatLaunch`\n(connector name, persona text, model and the other launch choices) and a `ManagedLifecycleHandoff`\n(space, owner, actor, the host-chosen `lifecycleUid`, broker and IdP pins, the pinned exchange base,\nthe sentinel, the channel lists, and the raw actor token). The manager enrolls once, as it does\ntoday, and builds the handoff from what it holds. It never sends the token to the host, never copies\na file from its workspace or secret store, and never builds a local launch for that seat. No signer,\nissuer or callout record, loopback capability, provisioner or manager credential, control token, or\nmanager path crosses. A spawn choice that only the manager's host can honour is refused before\nenrollment: `--resume`, a manifest agent's `continuity: exact`, `--cwd`, and any shared MCP server,\nwhether from `--share-tools` or the config default (`--share-tools none` passes).\n\nThe runtime creates one provider resource under `managedRuntimeKey(target)`, writes the handoff into\nit as one 0600 file and the persona beside it, and runs the stock bootstrap there:\n`cotal spawn --config <persona-file> --space <space> --name <actor> --expect-owner <owner> --expect-lifecycle-uid <uid>`\nwith `COTAL_MANAGED_HANDOFF_FILE` naming the file. `delegatedSeatCommand` builds that argv. The\n`cotal` entry reads the file into memory, deletes it and drops the variable before it parses flags,\nprints help or loads extensions, so every outcome, a refusal of its own flags included, leaves no\nfile. It refuses before any broker connection or exchange request when the\nspace, owner, actor, or lifecycle UID differ from the expected values, and then runs the\nenrollment-redeem consumer: it registers the mesh in its own home, writes the token to its own 0600\nfile, and exchanges it through `agent-bearer --exchange-url` unchanged. It never enrolls, redeems, or\nmints a token or UID.\n\nReadiness is still mesh presence. A create whose answer is lost leaves the handle running, so the\nlaunch settles uncertain and stays held; the manager never retries it. A provider read that finds no\nresource under the key is not an exit, because the create may still land. Every close by\n`managedRuntimeKey` is fenced: it completes only once the create was answered or the provider refuses\nany later create under the key. A provider that names its own resources may run the create as a\ndurable operation keyed by `managedRuntimeKey` and close through the identifier its authenticated\ncreate response returned, kept where the host can read it without the manager. Only that response\nbinds an identifier to the key; one derived from the key or found by name or listing is never closed\nor adopted, and while the response is unknown the launch stays held. Every stop, the reap of a child\nwhose parent exited and `Manager.stop({ withAgents: true })` included, runs `prepareAgentRetirement`\nfor the UID-exact target, then `stop()` on the handle `spawnDelegated` returned, then the terminal\nbarrier. `preparePreservation` refuses a cut that holds a delegated seat, and a `Manager.stop()` after\na refused cut retires the seat through the same steps. After the manager is gone\nthe host runs the same steps, makes the same fenced close by `managedRuntimeKey`, and finishes at\n`MANAGED_RETIRE_PATH`. Supply `spawnDelegated` only from a runtime whose host can make that fenced\nclose without the manager. The enrollment redeem\n(`COTAL_ENROLLMENT_FILE`) stays for lifecycles whose token the host generated itself; a host cannot\nmint one for a manager-enrolled lifecycle because it holds only the digest.\n\n## Provisioning a space (one-space reference shape)\n\n```ts\nimport { createSpaceAuth, setupSpaceStreams, ensureDefaultDeliveryClass, mintCreds, newIdentity } from \"@cotal-ai/core\";\nconst auth = await createSpaceAuth(space); // trust bundle (in-memory seeds)\nconst provisionerCreds = await mintCreds(auth, newIdentity(), \"provisioner\");\nawait setupSpaceStreams({ servers: brokerUrl, space, creds: provisionerCreds });\n// SPEC section 4: write the default delivery class at space creation so it is wire-discoverable,\n// never inferred from the resolution fallback. A daemon-backed space is \"durable\".\nawait ensureDefaultDeliveryClass({ servers: brokerUrl, space, creds: provisionerCreds, deliveryClass: \"durable\" });\nconst deliveryCreds = await mintCreds(auth, newIdentity(), \"delivery\");\n// put deliveryCreds into your SecretStore under deliveryCredsKey(space, { injected: true })\n// (@cotal-ai/workspace) before booting delivery — the key is per-space, not the bare kind.\n```\n\nRendering the broker config for a user-auth space is `serverConfig(broker, spaces, { storeDir,\nmaxFileStore?, extraAccounts })`, where `extraAccounts` must include the callout account from\n`createCalloutAuth` so the auth-service has a broker account to answer on. That account never shares\nthe data account. `maxFileStore` is an optional positive integer byte cap; any other value throws.\n\nBroker trust and space accounts are separate authorities: `createBrokerAuth` mints the one\noperator + system account a broker trusts, `createSpaceAccountAuth(broker, space)` signs each\ntenant's data account under it, and `serverConfig(broker, spaces, opts)` renders them all into one\nconfig. A host composition can therefore provision several spaces on one broker today. `cotal up`\nrenders that config from every tenant the root's auth directory holds, so booting one space keeps\nthe broker trusting its siblings, and it refuses to render at all while any account record is\nunreadable. The rest of the CLI lifecycle is still broker-wide: `down`, `clean` and `backup` refuse\non a multi-space root rather than scoping to one tenant, and the per-space lifecycle is the\nremaining multi-space operator layer. See\n[Known gaps](#hosted-composition-gaps).\n\n## Hazardous provisioning primitives\n\n`mintCreds`, the full `Profile`/`CredentialKind` matrix, `createSpaceAuth`, and `stripSpaceAuth` are\nlow-level operator primitives. Handle them as account-authority material:\n\n- A holder of a `SpaceAuth` (or a `stripSpaceAuth` bundle, which **keeps** the data signing seed) is\n a fully-trusted tenant-account authority: it can mint `admin`, `provisioner`, and destructive\n profiles, not merely `supervisor`, and mint a DM-reading identity. `createSpaceAuth`'s full result\n holds operator, system, and account seeds in memory.\n- Choose `profile` and `MintOpts` from **server-side constants**, never from tenant input. `MintOpts`\n can widen the bounded TTL defaults; cap it at your boundary. `CREDENTIAL_LIFETIMES` is a policy\n record, not an authorization boundary.\n- Never log signer material or export it into env. Do not co-locate signer access with an untrusted\n connector/runtime process at the same OS uid (file permissions do not contain a same-uid reader;\n see the manager's isolation note). Segregate per tenant; rotate on compromise\n (`rotateDataAccountSigningKey`).\n\n## Hosted composition gaps\n\nThe primitives above are present as exports, but three capabilities are **not** cleanly composable\nfrom the public contract today. Each is tied to work in flight; a host either waits for the seam or\nscopes the capability out. None is a wire concern.\n\n1. **Delivery immediate live eviction and a fully-hosted membership feed.** The renewable\n `membership-rw.creds` is now a `SecretStore` kind. `startMembership` reads it through the\n injected store, and the manager re-signs it there. The graph-feed writer therefore renews on a hosted\n backend (its data connection adopts each generation on a preflight-proven 75% timer). What still\n reads from a fixed on-disk path are the *static* `membership-observer.creds` and\n `connection-evictor.creds` ($SYS creds, minted at the `up` that provisions the account and renewed by `up --rotate-sys`) and `membership.json`\n (`{accountId}`, non-secret config); those, plus the private provisioning wrapper, keep immediate\n live eviction and a fully-hosted feed a partial gap. Missing files degrade membership to\n traffic-only and make live eviction refuse (loudly). The supported delivery contract here is the\n Plane-3 durable backstop.\n2. **Supervisor signer isolation.** `ManagerOptions.secretStore` now injects the one `SecretStore` the\n manager reads/writes every secret through, including the composed `SpaceAuth`\n signer (the split trust records), its daemon-cred renewal, and its per-agent kinds. What remains is process\n isolation: the manager still decrypts the signer in-process at its uid, so untrusted agent children\n must run under a different uid/container/mount namespace or behind a future remote signer.\n3. **Per-space lifecycle on a shared broker.** The trust layer is multi-space\n (`createBrokerAuth` + `createSpaceAccountAuth` + N-space `serverConfig`, persisted as\n `broker.json` + `account.<key>.json`) and `cotal up` renders the whole tenant list, but there is\n no per-space provisioning verb and no per-space teardown/backup/restore: the CLI's broker-wide\n lifecycle verbs refuse on a multi-space root, naming the tenants.\n This is the remaining multi-space operator layer.\n4. **A non-Better-Auth production IdP.** The exchange core (`createIdpBridge`) is EdDSA-generic, but\n the stock provider and login client are Better-Auth-endpoint-shaped, `cotalAuthProvider`\n self-registers on import (colliding with a host-owned provider under `resolveAuthProvider`), and\n the login flow speaks Better Auth's device-code endpoints. A different IdP is a host-built auth\n composition on the low-level primitives, not a configuration change (see\n [the IdP callout contract](identity-and-auth.md#the-idp-callout-contract)).\n\n## Hosted durability\n\nSpace-durable **coordination** state (chat/DM/task history, live presence, membership runtime, the\ndurable ACL registry, leases) lives in **JetStream**, written by the delivery daemon and the\nendpoints. It is broker-resident and needs no host-side durable path.\n\nWhat is **not** in JetStream, and is hosting-critical, is trust and authorization state a host must\nplace and keep:\n\n| state | class | where today | hosted injection |\n|---|---|---|---|\n| full `SpaceAuth` trust chain (`auth/broker.json` + `auth/account.<key>.json`, composed; a stripped signer bundle may instead be mounted at the legacy `auth/auth.json` key) | signing authority | `SecretStore` | `SecretStore` (manager + renewal) |\n| auth kinds: callout account/creds/xkey, issuer private keys, owner-derivation secret, data-signer projection | signing/identity authority | four `SecretStore` kinds | `SecretStore` (auth-service) |\n| `delivery.creds` | standing scoped cred | `SecretStore` or `--creds` | `SecretStore` (delivery) |\n| actor ledger, IdP pin | authorization + trust config | ambient `userAuthStateDir(findCotalRoot(), space)` | none (root-relative; not `store`/`COTAL_HOME`) |\n| `membership-rw.creds` | standing scoped cred | `SecretStore` | `SecretStore` (delivery + manager renewal) |\n| membership-observer / connection-evictor creds + `membership.json` | scoped $SYS creds / config | workspace filesystem | none (see gap 1) |\n| manager agent creds, actor tokens, sentinel creds | lifecycle authority | `SecretStore` | `SecretStore` (manager `secretStore`) |\n| `~/.cotal/meshes/space.<key>.json` record (holds IdP trust pins/root pointers) | non-secret, integrity-critical | machine home | process-global `COTAL_HOME` only |\n| auth-health, renewal records | non-secret diagnostics | workspace filesystem | `workspaceRoot` |\n\nThe `SpaceAuth` trust chain and the auth-service store kinds are **separate** identities/projections,\nnever parts of one document. `auth-service.json` (the live exchange capability) is ephemeral runtime\nstate, not durable, but is sensitive while the daemon runs. `@cotal-ai/workspace` is machine-local\noperator tooling by design; personas, PID files, and the `current-mesh` pointer are truly local and\nmust **not** sit on a hosted durable path. Everything classed above as an authority is what a hosted\ncomposition must provision and persist: signer-bearing server secrets now have `SecretStore` seams;\nthe remaining non-injectable rows are the explicit ambient `workspaceRoot`/cwd paths above.\n\n## See also\n\n- [Substrate stability](stability.md): what v0.3 and the 0.x packages guarantee, and the projected v0.4 break.\n- [Identity and auth](identity-and-auth.md): the profile matrix, the signer, and the IdP callout contract.\n- [Delivery daemon](delivery-daemon.md): the Plane-3 durable backstop.\n- [Deploy](deploy.md): the reference container against an external broker.\n"
|
|
168
168
|
},
|
|
169
169
|
{
|
|
170
170
|
"slug": "examples",
|
|
@@ -227,7 +227,7 @@ export function loadDocsBundle() {
|
|
|
227
227
|
"title": "Run a mesh",
|
|
228
228
|
"kind": "Guide (informative)",
|
|
229
229
|
"summary": "Day-to-day operation of a local mesh: what cotal up actually runs, how spawning resolves personas, harnesses, and models, how to reach a mesh from any directory, and the operator-only maintenance v…",
|
|
230
|
-
"body": "# Run a mesh\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nDay-to-day operation of a local mesh: what `cotal up` actually runs, how spawning\nresolves personas, harnesses, and models, how to reach a mesh from any directory, and the\noperator-only maintenance verbs. Every command's full flag set is in the\n[CLI reference](cli.md).\n\n## The stack\n\n`cotal up` brings up the whole local stack and bare `cotal down` stops it. Managed\nagents stay running as unmanaged OS processes; pass `--with-agents` to take them\nwith the stack. Seats of the built-in pty runtime run inside the manager process, so\nthey stop with the manager either way. Ctrl-C on a foreground `up` stops the manager through the\nsame stop as bare down and prints the same report; when that stop is refused, for example because\nthe manager cannot prove it can spare, Ctrl-C leaves the stack running, and you end it with\n`cotal down --with-agents`. A current manager records what its stop does with its seats before\nbare down signals it. A pre-pin legacy manager instead receives a reduced-guarantee\nwarning and is signalled according to the documented upgrade contract. Its running binary\nmay still carry the older destructive SIGTERM handler, so the CLI does not claim its\npre-signal agent inventory was spared; those agents may have been reaped.\n\n- **Broker**: a local `nats-server` (logs to `.cotal/nats.log`).\n- **Delivery daemon**: the durable backstop, auth mode only\n ([what it does](delivery-daemon.md)).\n- **Manager**: a detached supervisor answering the control plane, so\n `cotal spawn --detach` and the `cotal_spawn` tool work right after `up`.\n\nCotal creates the presence bucket in memory storage. Its records are liveness that every endpoint\nrewrites each heartbeat, so nothing is lost when a broker restart empties it, and nats-server's file\nstore write latch cannot reach it. A broker stop removes the memory stream itself, so every `cotal up`,\nincluding the resume after `cotal down --preserve-state`, creates it again before any daemon starts.\nJetStream fixes a stream's storage class when it is created, so a presence bucket created file-backed\nby an older cotal stays file-backed until that stream is recreated.\n\nA file-backed presence bucket can remain open and watchable while refusing every write. A bound\nendpoint reports this as `presence-write-stuck` after one full presence TTL of consecutive failures.\nThe roster is last-known while that condition is active. Restarting the broker clears nats-server's\nin-memory store latch and preserves the JetStream root. Current credentials split the required stream\nauthority: the `cotal up` provisioner can create the presence stream but cannot delete it, while the\nteardown credential can delete it but cannot recreate it. Cotal therefore reports the condition but\ndoes not attempt an unsafe partial delete-and-recreate. Stop and restart the broker to recover.\nA broker below nats-server 2.14.5 carries the latch (nats-server fixed it in 2.14.5). When `cotal up` starts or finds such a broker and the space's presence bucket is file-backed, it says so. A memory-backed bucket gets no warning. A broker below the SPEC §13.12 floor of 2.12 is refused at connect with the floor sentence.\n\nThree modes:\n\n- **Default (static auth).** JWT-authed, on by default: sender authenticity and per-agent\n ACLs, enforced by the broker ([how](identity-and-auth.md)).\n- **`--user-auth --idp <url>`.** Per-user auth: people `cotal login` once, the operator\n grants their agents on the actor ledger, and every connect is authorized live against\n that grant. Starts the space's auth service alongside the broker\n ([how](identity-and-auth.md)).\n- **`--open`.** An unauthenticated, live-only dev mesh (no auth, no delivery daemon). For\n quick local experiments.\n\nThe broker and local services bind **loopback** by default. `--host 0.0.0.0` widens the broker\nbind independently of the auth mode, so \"network-reachable\" never silently means\n\"unauthenticated\". With no explicit `--server`, `cotal up` auto-selects a free local port when\nthe default address is already held by another project; an explicit `--server` fails loud on\ncollision.\n\n`--host` is a boot flag, not a live rebind. A fresh `cotal up` writes the generated\n`.cotal/auth/server.conf` (project-local, not `~/.cotal`) with that bind and starts nats against\nit. If anything is already answering at the mesh URL, `up` refreshes the recorded mesh and\nleaves the running nats listener alone, so passing `--host 0.0.0.0` on a live or orphaned\nbroker does not change who can connect. To change the bind: `cotal down`, then `cotal up --host\n<addr>` against a stopped broker so the generated file is rewritten. Do not edit `server.conf`\nby hand; the next real boot overwrites it.\n\nOn a stopped shared broker, `up` renders every persisted space account and every enabled\nspace's auth-callout account into the resolver preload, regardless of which space starts\nthe broker. A missing callout account for an enabled space stops the boot rather than\nstarting with a reduced resolver. An already-running broker is refreshed without rewriting\nits config.\n\nA broker-only host is a first-class `up` mode. `cotal up --no-manager` boots the broker and, in\nauth mode, the delivery daemon, and no local manager, so the broker host never has a manager to\nstop and never leaves a manager slot stale. A refresh under the flag of a mesh whose manager is\nlive refuses rather than keeping or stopping it: `cotal down manager` first. Without the flag,\nauth-mode `up` still starts nats, the delivery daemon, and a\nlocal manager. A space may run more than one manager, addressed by instance id\n([control surface](control-surface.md#instance-routing)); putting no manager on the broker host\nis a topology choice, not a singleton invariant. A manager whose boot inventory has no\navailable connector does not take unpinned `spawn`/`launch` on the class rail, so a sibling\nthat can launch the harness can. `describe` still rides the class rail, so an unpinned spawn\ncan bind-fence when that skip member answered describe; re-issue, or pin `--on`. Pin one\ninstance with `--on` when a partial inventory still answers with a harness refusal. The\nsupported split is:\n\n```bash\n# broker host (project root that owns the generated conf, pidfiles, and logs)\ncotal up --detach --host 0.0.0.0 --space main --no-manager\n# no local manager starts: the summary lists nats-server + delivery daemon, and there is no\n# `.cotal/manager.<spaceKey>.log` to wait for on this host\n\n# manager host (registered remote mesh, same space)\ncotal meshes add --server nats://broker.example:4222 --root ~/meshes/main\ncotal supervise --space main --server nats://broker.example:4222\n```\n\nWait for `✓ manager up` in `.cotal/manager.<spaceKey>.log` on the manager host before spawning\nagents. On a broker host started without `--no-manager`, `cotal up --detach` prints `✓ running in\nthe background:` with `manager` listed once the manager pidfile is live; stop that local manager\nonly after the `✓ manager up` line. A host started WITH `--no-manager` never runs one, so neither\nthe wait nor the stop applies there. That detach stdout is not a safe teardown boundary: it is\npidfile liveness, not `✓ manager up`. `✓ manager up` is supervise's post-start line after\n`await mgr.start()`. `cotal down manager` after only the detach line can still default-terminate\nthe child during registration after it has taken the governance slot. Stopping before that\npost-start log line can leave the endpoint governance slot held until the holder's gate\nreopens past the stamp (the successor's boot heal, or\n[`cotal reconcile-gate`](cli.md#reconcile-gate) when that boot cannot run). See\n[Gate recovery](#gate-recovery).\n\nStandalone `cotal deliver --creds` is not a repair for that split. Production renewal needs\nthe manager and the daemon to address one credential store. The manager renews its own service\ncredential inside that credential's own window and re-dials its service connection with the\nrenewed credential; if the connection closes and cannot be restored within about forty seconds\nit releases its lease and exits so a restart can serve, while a broker that is briefly gone is\nwaited out. Separate host filesystems still\nleave manager root A writing and the daemon reloading root B; that composition is refused\nwhile the daemon stays up. Before every remint the manager challenges the delivery daemon's\nstore identity, and the answer must come from the process holding the delivery lease: the\nreply names the answering endpoint and the manager reads the lease row itself under its own\ncredential, so a non-holder answering on the queue-grouped admin rail is refused instead of\ncounting as the daemon's store. A rail that reports no responder is also settled from the\nlease row, so a live holder on record makes that outcome a refusal rather than an absent\ndaemon. Keep delivery on the broker host under `up`, and share one store\nonly when you are composing a hosted pair ([embedding](embedding.md#supervisor-signing-authority)).\nOn the `--no-manager` split above, the manager host's manager stays off the daemon-credential\nrenewal lease once its store check finds the daemon on another store. A filesystem store is named\nby its root and by a random id in `.cotal/store.id`, which the copied `.cotal/auth` does not carry,\nso this holds when both hosts use the same root path. `cotal doctor auth --fix` on\nthe broker host then renews the daemon credentials once they pass their renewal point.\n\n### Split host bind\n\nA remote manager cannot reach a loopback broker. After changing `--host`, confirm the\ngenerated `host:` in `.cotal/auth/server.conf` and that nats is listening on that address\nbefore registering the mesh on the manager host. Detached child logs stay under the **project**\n`.cotal/` that `up` ran in (see [When something looks absent](#when-something-looks-absent));\nthey are not `~/.cotal` unless that directory is the mesh root.\n\nA user-auth mesh can expose only its credential exchange through an operator-owned HTTPS reverse\nproxy while leaving the existing local exchange untouched:\n\n```bash\ncotal up --user-auth --idp https://idp.example/api/auth \\\n --exchange-public-port 7443 \\\n --exchange-public-url https://auth.example\n```\n\nThe public listener itself still binds `127.0.0.1:7443`; configure the proxy to terminate TLS and\nforward to it. It serves only `/health`, `/jwks`, `/exchange`, and `/.well-known/cotal-mesh` with\nthe documented methods. It needs no local file capability: the signed IdP JWT or managed-agent\nactor token is the proof, while the original loopback listener remains capability-gated. Add\n`--exchange-trusted-proxy` only when that listener is reachable exclusively through your trusted\nproxy; it keys failure throttling by the last `X-Forwarded-For` hop instead of the socket address.\nThe well-known bundle includes IdP pins and a deny-all sentinel credential, so fetch it only from\nthe configured HTTPS origin. To change these listener flags, stop and restart the mesh; a refresh\nof an already-running service does not replace its bind or proxy policy. See\n[Identity & auth](identity-and-auth.md#per-user-authentication) for the trust boundary.\n\n### Remote supervised seats by enrollment\n\nA remote seat does not need to run `cotal login` when the mesh owner pre-mints a single-use\nenrollment for it. Mount the enrollment URL as a private file, place the seat persona on the remote\nmachine, and launch the foreground seat:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./worker.md --space main\n```\n\nThe URL is redeemed once with an unauthenticated GET. Redirects, off-machine plain HTTP, retries,\nand login fallback are refused. If the seat has no mesh record yet, the enrollment response's stock\nuser-bundle fields register it before the launch. The returned actor token then uses the same remote\nauth-service exchange as a login-provisioned agent. The enrollment URL and file path do not enter the\npreflight or harness environment. A failed or reused enrollment leaves no actor material on disk; ask the owner\nfor a fresh enrollment. When the foreground seat exits, this machine's credential files are removed\nand the mesh-side grant stays until the mesh operator revokes it; the launch line says so. The exact\nserver contract is in\n[Enrollment redeem](identity-and-auth.md#enrollment-redeem).\n\n`cotal status` prints the detailed setup, process, registry, and live mesh status. Its Machine\nsection names the running CLI's source checkout, installed package root, or npx package root beside\nthe version. It has one row per installed connector, which reports whether the executables that\nconnector declares in `requires` are on PATH. Status, setup and the manager's preflight resolve them\nthe same way: an entry written as a path is checked as given, and a directory never counts as the\nexecutable. A connector whose setup provider reports health adds its\nown rows above those. The Claude Code connector reports its plugin and its skills plugin, and a stale\nskills row names the installed and CLI versions it compared. `cotal\nsetup` (after the first run) prints the compact card.\n\nBefore reporting ready, the manager resolves every installed connector's declared harness\nbinaries against its own environment. A missing binary does not stop unrelated manager work: boot\ncontinues, but prints a named `connector <name> unavailable` line and records that reason in the\nmanager's `status` response. Available connector rows record the absolute paths boot resolved.\nA spawned seat and a seat resumed after `cotal down --preserve-state` both launch from those paths,\nand both are refused with the recorded reason when their connector's row is unavailable.\n`cotal models` takes the same rule and reports that reason in place of the catalog, so it agrees\nwith a launch about a harness installed or removed after boot. The manager looks again only when it\nrestarts. A connector registered after boot has no row, so spawn, resume and `cotal models` check\nits binaries on PATH when they run.\n\nOn an authenticated manager start, unfinished static lifecycle rows reconcile while the control\nendpoint is already serving. The manager `status` response reports\nthe `staticReconciliation` state, the last sweep counts, and each failed alias with its durable\nphase and literal disposition. `cotal status --components` reports the state and per-alias failure\ndetails. A failed exact terminal is retried in the same process after 1, 5,\nand 30 seconds. Each attempt re-reads the durable slot and re-enters the same deterministic terminal\noperation; the delays only schedule work and never release the lifecycle fence. The terminal's\ncleanup removes the lifecycle's credential file and its broker durables and read-ACL row as separate\nsteps. A file that cannot be removed does not leave the broker footprint behind, and its failure\nkeeps the alias held for the next attempt.\n\nOn shutdown, the manager fences new reconciliation work and waits for an exact terminal that already\nstarted. The current serial sweep stops before its next alias, and startup cannot publish the manager\nservice after `stop()` completes.\n\nThe four-attempt budget is per manager process. An exhausted row stays held and reports\n`retry-exhausted` with the remedy to restart the manager. The next process derives a fresh budget\nfrom the still-authoritative durable row. A `recovered` row remains visible until the next static\nreconciliation sweep, then clears. This component reports reconciliation outcomes. It does not say\nwhether footprint cleanup completed independently of the terminal result; that separate durable\nprojection remains tracked by #1274.\n\n`cotal service install` is the supported way to run the manager as a user service\n([CLI reference](cli.md#service)): a systemd user unit on Linux, a launchd agent on macOS, one\nper mesh, surviving logout and reboot. On Linux that needs user lingering: install refuses while\nit is off and prints the root command that enables it. It installs only\nthe manager; the units below remain the process models for every other component, and they are\nstill **examples of process models** for those: copy them only after you decide which processes\nthe unit should own.\n\n### Supervising the detached stack\n\n`cotal up --detach` is a launcher: it starts the broker, delivery daemon, and manager, reports what\nstarted, then exits. Do not wrap it in a systemd service with `Type=oneshot` and\n`RemainAfterExit=yes` and treat `systemctl is-active` as stack health. That unit becomes `active\n(exited)` when the launcher exits successfully and stays active even if every detached process dies.\nWhen `up --detach` can identify that exact unit shape, it prints a warning but keeps the requested\nstartup behavior.\n\nFor a single-host stack, keep `cotal up` itself in the foreground so systemd tracks a long-running\nprocess and restarts the stack if that process fails:\n\n```ini\n[Service]\nType=simple\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal up --space main --host 0.0.0.0\nRestart=on-failure\nRestartSec=5s\n```\n\nAn active unit then proves the foreground launcher and broker are still running, but it still does\nnot prove that every child component serves. Pair it with the component check below. Also remember\nthat `cotal up` starts a local manager as well as the broker and delivery daemon; run\n`cotal up --no-manager` (add the flag to the unit's `ExecStart` too) on a host intended to be\nbroker-only, so the unit and the host agree.\n\nSeats spawned by the built-in `pty` runtime run with `oom_score_adj` 500, so under memory\npressure the kernel prefers a seat over the broker, manager and delivery daemon, which are left as\nthey were started; the extension runtimes do not own the seat's process and get no preference.\n\nThat `Type=simple` shape puts nats in the unit's cgroup with the foreground `up` process. A\n`Restart=always` (or `on-failure`) of **this** unit therefore restarts nats as well, so remote\nmanagers drop for the time it takes the broker to come back. Wrapping `cotal up --detach` in\n`Type=oneshot` with `RemainAfterExit=yes` does not move nats out of that cgroup. Detached\nspawn starts a new process group, not a new systemd cgroup, and the default\n`KillMode=control-group` still signals every process left in the service cgroup on stop or\nrestart, including the nats PID. Escaping that cgroup needs an explicit unit setting such as\n`KillMode=process`, or a separate nats unit; this CLI does not ship that escape. The\n`Type=oneshot` unit below is a `cotal status --components` liveness check, not a\n`--detach` launcher. Neither trade is universal from\n`Type=simple` alone; it follows from which processes the unit actually owns. `cotal service\ninstall` covers only the manager, so for the broker and its siblings pick the example that\nmatches the ownership you want, and treat\n`systemctl is-active` as unit health, not mesh health.\n\nA broker that crashes under that foreground `up` keeps its mesh record and exits non-zero, so the\nunit's restart takes the repair path against the recorded store rather than starting a second one.\n\nIf the deployment deliberately uses `cotal up --detach` as a boot action, monitor observed state\ninstead of the launcher's exit:\n\n```ini\n[Unit]\nDescription=Check Cotal component liveness\n\n[Service]\nType=oneshot\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal status --components --space main\n```\n\nRun that check from a systemd timer or another monitor and alert on a nonzero exit. The command\ndistinguishes `absent`, `not-serving`, and `refused` components and never treats a sibling's health as\nproof. Its delivery-process check is local to the broker host, so run it there. On a split topology,\nalso probe the broker URL from the manager host and monitor the manager's own service there. A remote\nmanager cannot observe the broker host's delivery PID, and an `active` unit on either host says\nnothing about the other host.\n\nStop one part without tearing down the mesh by naming its registered component: `cotal down\nmanager`, `cotal down delivery`, or `cotal down web`. Component names from installed extensions\njoin the same surface; `cotal down` with no names retains whole-stack behavior and\nleaves managed agents running as unmanaged OS processes, except pty seats, which stop with the\nmanager. `cotal down --with-agents` is the previous reap. If a pinned manager has no\nspare-capability record, stop its managed agents explicitly before running that whole-stack\ncommand. A current manager always publishes the record, so it is absent only for an older manager,\nwhich may not understand the reap request.\n\n## Remote supervised agents\n\nOn a remote user-auth mesh, foreground `cotal spawn` remains the default participant path. A\nparticipant can run detached agents only after the host advertises and operates the remote manager\nauthority service, and the participant's actor-ledger row includes `supervise`. This is not implied\nby `spawn` or `admin`.\n\nThe participant's loopback/operator exchange obtains one closed `manager-service` view for its\nordinary derived owner, a fixed server-selected manager actor, and one opaque manager instance.\nThe host, not the participant, issues the public-nkey JWT material via the replay-safe,\nlifecycle-bound prepare → activate → renew exchange, plus a one-shot target-pinned retirement\nrequest for a host-managed terminal. It never exports the space signer, a static\nprovisioner credential, or generic storage authority. Remote registration publishes its service\nstatus at the registered revision and current process epoch, so manager-caller selection can find it.\n\nStock participant supervision asks its host to enroll a detached agent and to prepare its terminal\nretirement, over the same manager-authority transport. The stock auth service answers both when it\nruns with a public exchange face: it grants the agent under the participant's owner at a lifecycle\nUID it picks, bounded by the participant actor's own grant, provisions that UID's durables, and on\nretirement releases them and revokes the grant before the manager's terminal rail. It refuses a\nsecond enrollment of a name whose grant still stands until that agent's retirement is prepared. A\nhost platform that keeps these writers in its own storage intercepts both requests on its own route\ninstead. Copying host secrets or actor-ledger files to a participant is not supported. Foreground\nspawning and operator-local hosted managers use their existing paths.\n\nThe remote manager that `cotal supervise` starts can host workflow runs through its host: the host\nadmits each run and signs only the run's own driver, mediator and operator credentials. A logged-in\nuser's `cotal run start` against it is admitted: the auth callout issues the user's manager\nconnection, and the host binds each run to the owner who registered the manager. The run spawns\nagents that user owns, enrolled by the host like any detached spawn, with the reach the user's own\nrow grants when the spawn runs. A spawn may be placed on that manager and on no other instance. The\nhost's own manager refuses user-auth runs by name.\n[User-auth run start](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/user-auth-run-start.md)\nrecords the path.\n\nThe registry entry decides the broker URL `supervise` dials, so a mesh published over `wss://` is\ndialed as a websocket. The manager-authority registration it runs first also takes its TLS\nrequirement from that entry, so the prepare credential is not exchanged over a plaintext\nconnection the record did not describe. `cotal meshes add` records both.\n\nWhen the authority service, login, or renewal is unavailable, the remote manager degrades\nfail-closed: it refuses new agents, restarts, and credential replacement rather than pretending\nlocal authority exists. Existing agents remain live only while their independent credentials are\nvalid. A hosted composition must revoke the managed grant and finish its resumable release before it\nrequests terminal retirement. Deleting DM or delivery consumers is not retirement and must not reset\na resumable lifecycle's frontier or pending state. The alias remains held until the terminal barrier\nconfirms. Restore service and renew successfully before asking it to recover an agent. See\n[Identity & auth](identity-and-auth.md#remote-manager-authority) and the [CLI\nreference](cli.md#supervise).\n\n## Spawning agents\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn reviewer --detach # supervised: the manager runs it in a PTY\ncotal attach --name reviewer # watch/type into a detached agent (Ctrl-] detaches)\ncotal ps # what the manager is running\ncotal stop --name reviewer # stop one\n```\n\nHow a spawn resolves:\n\n- **Persona.** A bare `cotal spawn` uses `.cotal/agents/default.md`; a positional name\n picks `.cotal/agents/<name>.md`; `--config` takes an explicit ref or path. Set\n `COTAL_DEFAULT_PERSONA=<name-or-path>` to change the fallback. Fields and format:\n [agent files](agent-files.md).\n- **Harness.** Resolution order is an explicit `--agent` or `cotal_spawn` `agent` argument,\n then the persona file's `agent:` pin, then the invoking caller's `COTAL_DEFAULT_AGENT`,\n then the manager's `COTAL_DEFAULT_AGENT`, then the product default (Claude). Compared in\n [Connectors](connectors.md); per-connector guides:\n [Claude](connect-claude.md) · [OpenCode](connect-opencode.md) ·\n [Hermes](connect-hermes.md) · [pi](connect-pi.md).\n- **Model.** `--model` overrides the persona file's `model:` (Claude: `opus` / `sonnet` or\n a full id; OpenCode: `provider/model`). Connectors that expose a catalog report it via\n `cotal models --agent opencode`: model ids plus available variants; pick one with\n `--model provider/model --variant high`.\n- **Tools.** A spawned Claude Code agent gets the cotal tools plus the MCP servers the cotal\n config shares, which first-run `cotal setup` fills with your own; narrow them per spawn with\n `--share-tools` ([config](config.md)).\n- **Launch options.** `--opt key=value` (repeatable) passes a native harness flag straight\n through; a persona or manifest `launchOptions:` mapping does the same declaratively (a\n `--opt` wins per key). It is a **raw passthrough**, with no allow/deny list: Claude renders\n each as `--key value` (a bare `--key` for an empty value), OpenCode merges them into its\n agent config, and Hermes has no option surface so it fails loud. The trust boundary is the\n `spawn` capability itself, not the flag set, so granting `spawn` is host-launch authority\n ([security](security.md)). A key must be a plain flag name; malformed or prototype-polluting\n keys are refused.\n\nDetach from an attached PTY with **Ctrl-]** (the agent keeps running); rebind it with\n`COTAL_DETACH_KEY=ctrl-<char>` when it clashes with a keybinding inside the agent's TUI.\n\n**Runtimes.** The manager spawns into a **pty** by default. It spawns the PTY in-process on\nevery platform, so replacing the manager worker closes its seats and the pty runtime gives no hot\nupdate. Any manager stop, bare `cotal down` included, stops and deprovisions those seats. A stopping\nmanager refuses new spawns and first waits for the ones it already accepted, so their seats stop too. On Linux\nit can still adopt and reap seats that an earlier manager launched under a detached per-seat\ncustodian, so those seats drain under the new manager; it starts no new custodian. A custodian whose agent has exited exits a few seconds later on its own. `cotal seats`\nlists the custodians left on the machine, and `cotal seats --drain` retires the ones whose agent\nhas exited while keeping every seat whose agent still runs ([cli.md](cli.md#seats)). When a pty\nagent exits on its own, in-process or under a custodian, the manager logs a `seat reaped:` line\nwith the exit code and, for a signalled child, the signal number. The line ends with the last line\nthe child printed that starts with a connector's `[cotal-<name>]` or `[cotal-<name>/<part>]`\nprefix, cut to 240 characters, when it printed one. A custodian keeps the same record beside the\nseat's custody record, so a later reap of that seat, including one by a\nsuccessor manager, reports how the child ended. When the custodian cannot write that record, it\nsays why in the seat's `custodian.log`, and a later reap of a child that ended on its own reports\nthe record as missing or unreadable. Optional runtimes are installed\nthrough the extension surface, for example `cotal ext add @cotal-ai/orca`, then selected with\n`--runtime orca` (similarly `@cotal-ai/tmux`, `@cotal-ai/cmux`, and `@cotal-ai/herdr`). They put teammates in native\nterminal surfaces rather than manager-owned PTYs. Runtime names are open-ended and resolved from\nthe registry; a missing provider or app throws, never silently falls back\n([architecture](architecture.md)).\n\n## Mesh registry\n\n`cotal up` records each running mesh in a machine-local registry\n(`~/.cotal/meshes/space.<key>.json`, named by a case-safe hex encoding of the space: broker URL, the project root holding its creds and\npersonas, and its mode). So a bare `cotal spawn <persona>` from *any* directory joins the\nrunning mesh with the right credentials instead of mistaking the cwd for a space:\n\n- `cotal use <name>` sets the default from every directory, including inside another mesh's\n project. `--space <name>` overrides it for one command.\n- When one broker has records for several spaces, `cotal up --space <name>` refreshes that named\n space.\n- A refresh rewrites only what that command decided: the server, root and mode, the user-auth\n endpoints, and an explicit `--host` or `--max-sessions`. Every other field, such as the TLS\n requirement, is kept as the record stands when the refresh writes it, so a change another\n command made during the refresh survives. If the record was removed during the refresh, `up`\n fails instead of writing it back.\n- With no live selected default, a project with its own `.cotal/` resolves to that project's\n mesh; otherwise one running mesh is used automatically and several are an error.\n- `cotal meshes` lists them (a `*` marks the default); `cotal down` removes the entry.\n\nThe registry stores a *path*, never a secret; trust material stays in each project's\n`.cotal/auth`. If the mesh is down or won't take your creds, spawn fails with one\nsentence, never a raw NATS trace.\n\n### Meshes you did not start here\n\nA mesh running on another machine has no `cotal up` on this one, so register it by hand:\n\n```bash\ncotal meshes add # guided: asks for the broker, probes it, offers what it finds\ncotal meshes add optiplex --server nats://100.90.12.34:4222 --root ~/meshes/optiplex \\\n --allow-unencrypted-overlay # see below: an overlay address needs this\ncotal meshes rm optiplex\n```\n\nOn a terminal, a bare `cotal meshes add` walks you through it: it probes the broker you name and\nreports whether it is open or requires credentials, offers the spaces the folder already holds\ncredentials for, and shows the record before writing it. Scripts and agents keep the flag form -\nwithout a terminal nothing prompts.\n\n`--root` is the local folder holding that mesh's `.cotal/auth` and `.cotal/agents` (its personas);\nthe mode is inferred from what that folder holds.\n\nThe instance identities of the manager and the user-auth service are not part of that folder.\nEach root keeps its own in `.cotal/space.<hex>/`, so `cotal supervise` or `cotal up --user-auth` in\nthe root you copied the folder to starts an instance of its own. A root last run by an older Cotal\nstill holds them in `.cotal/auth`, as `manager-instance.<hex>.json`, `manager-siblings.<hex>.json`\nand `space.<hex>/.cotal/auth/auth-instance.<hex>.json`. Delete those files from a copy of such a\nfolder before the first `cotal supervise` or `cotal up` there.\n\n**Know what you are copying.** For an authenticated mesh that folder carries the space's account\n**signing seed**, which is the authority to mint any identity in the space. A machine holding it\nis a certificate authority for the mesh rather than a client of it: anyone who reads it can\nimpersonate any agent, read every retained channel and DM, change ACLs, and keep issuing\nthemselves credentials. There is no per-machine revocation; undoing it means rotating the signing\nkey and re-minting every credential in the space. Copy it only to machines you would trust with\nthe whole mesh. `cotal mint` on its own does not substitute here: registering an `auth` mesh needs\nsigning material that composes, which a minted user credential is not. The\nbroker is probed before the record is written, so a bad address or a credential that mesh will not\naccept fails at registration rather than at your first `spawn` (`--force` records it without verifying,\nuseful when the mesh is simply down right now).\n\n#### Which addresses you may register\n\nRegistering a mesh is how this machine starts sending agent credentials to a broker it does not\nrun. NATS announces itself in plaintext before anyone authenticates, so an attacker on the path\ncan pose as the broker and read the credential out of the connect unless the connection\n**requires TLS**, which is recorded on the entry and enforced on every dial through it.\n\nWhat the record will require decides what you may register:\n\n- **Without required TLS**, the address is the gate: **loopback** (`127.0.0.0/8`, `::1`), or\n **your private overlay** (`100.64.0.0/10`, `fd7a:115c:a1e0::/48`) with\n `--allow-unencrypted-overlay`. The tunnel provides the protection, and this command cannot check\n its state. Hostnames are refused because the lookup would choose which machine receives your\n credentials.\n- **With required TLS**, set `--tls` or use a `tls://` URL. The recorded scheme enforces the TLS\n requirement. A **hostname or public address** is accepted because the certificate chain and\n hostname check identify the peer. A registration whose broker cannot complete the handshake\n fails unless you pass `--force`, which records the entry without verification.\n\nOrdinary private ranges like `10.x` and `192.168.x` are refused in **both** modes. A café's wifi\nis private but does not belong to you, and no public CA issues certificates for those ranges. An\naddress spelling changes nothing: `[::ffff:192.168.1.10]`, `3232235786`, `0300.0250.01.012`, and\n`192.168.257` all resolve to private addresses and receive the same refusal as the dotted form.\n`--force` exists for a mesh that is down. It never permits an unsafe credential destination.\n\n#### Registering a hosted user-auth mesh\n\nA user-auth space's IdP pins are established where the mesh runs and are never guessed. Register\none from **supplied** trust: `--user-auth-file bundle.json` (exported on the mesh's machine), or\n`--from https://auth.example`, which asks before it contacts the address at all, fetches the\ndiscovery document at `/.well-known/cotal-mesh` under that address over HTTPS, shows you the pins,\nand asks again before adopting them. A URL that already ends in `/.well-known/cotal-mesh` is\nfetched as given. Redirects are refused because a 302 can walk a pinned fetch down to\nplaintext or onto another host, and the pinned exchange must be an `https://` URL too. The one\nexception is an exchange on **this machine**, where nothing leaves the box: plain `http://` is\naccepted for a loopback *literal* (`127.0.0.1`, `::1`, and any spelling of them), but **not** for\n`localhost`, which a hosts entry or poisoned lookup could point elsewhere. Use the\nliteral. Registration checks that the pinned exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also checks that the broker refuses a\nbare connect; that refusal is the pass. The bundle's sentinel credentials are written to a private (0600) file\nunder the entry's root; the registry itself never carries the secret.\n\n**Without required TLS**, an overlay address is **refused unless you accept the dependency\nexplicitly**, with `--allow-unencrypted-overlay`. The address is not the guarantee: it is protected\nwhile the tunnel is up, and if the tunnel is down that range is ordinary carrier-grade NAT and\nwhoever answers the dial receives your credentials. Only you can know which it is, so the command\nasks you to say so. Your acceptance is recorded on the mesh entry rather than printed and\nforgotten, and the guided form asks the same question instead of taking the flag.\n\n**With required TLS** (`--tls`, or a `tls://` URL) that consent is no longer asked for, and the\nflag is not needed: the handshake is what protects the connection, so the acceptance it stood in\nfor has been replaced by proof rather than promise. `cotal meshes add <space> --server\nnats://100.64.0.1 --tls` registers an overlay address with no prompt, no flag and no recorded\nacceptance. This is the \"the flag disappears once the broker can be served over TLS\" case, and it\nhas now arrived.\n\nThis gate is on **registration**. `cotal join --creds --server <url>` deliberately takes an\nexplicit connection at face value and does not consult the registry, so it is not covered. Join\nthat way only to an address you would have registered.\n\nThe connection is still probed first, with the same second try at the longer budget the registry\npreflight uses, so a slow link reads as a connect that did not finish within that budget and a\nrefused port reads as a broker that is not running.\n\nRecords added this way are removed only by something that names them. A failed liveness probe\ndoes not delete any record: an unreachable broker, local or registered by hand, is shown as\n`offline` in `cotal meshes`. A bare command does not count that offline record as running;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the\nproject they are tearing down, and they leave a hand-registered one alone even when `--root`\npointed at that project. A `cotal up` for that space refuses outright unless it is that same\nendpoint: finding a broker already answering there is a refresh that starts nothing and leaves the\nrecord's provenance alone, while actually starting the broker for that space, server and root\nmakes this machine the one running it, so the record becomes an ordinary local one that\n`cotal down` clears. The refusal names `cotal supervise --space <s> --server <url>` (plus `cotal\ndeliver`) when the registered broker is on another host, and `cotal meshes rm` when it is local.\n`cotal meshes rm` drops it and re-registering with `--force` replaces it. `rm` only forgets a\nmesh. To stop one running here, use `cotal down`.\n\n## Watching\n\n`cotal console` is the terminal view (TUI on a real terminal, plain line stream when\npiped); `cotal web` is the browser dashboard. Both are read-only observers; the\nwalkthrough is [Watch a mesh](watch-a-mesh.md).\n\n## History\n\nRetained history is operator-owned. `cotal clean history --force` purges a space's\nretained channel history; `--dms` also purges DMs (`cotal history clear` is an alias).\nIt is deliberately **not** an agent tool: agents cannot wipe the record\n([identity & auth](identity-and-auth.md)). For a **stopped** mesh, `cotal clean store\n--force` deletes the on-disk JetStream store outright, and `cotal clean all --force`\nalso resets the space identity ([CLI reference](cli.md#clean)).\n\n## Offline backup\n\nFor a coherent durable cut, preserve the whole stack first, then create the artifact while it stays\ndown:\n\n```bash\ncotal down --preserve-state\ncotal backup create ./space-backup # full by default\n# later: deliberately resume the unchanged source\ncotal up --detach\n# or, from another preserved cut, restore before the normal listener opens\ncotal up --restore ./space-backup --detach\n```\n\nA refused cut leaves the mesh running and unfenced: fix what the refusal names and run\n`cotal down --preserve-state` again.\n\nUse `--store-dir` on both preservation and backup for a custom JetStream store. A store cap set\nwith `cotal up --max-file-store <bytes>` travels with the preserved state, and the resume renders it\nagain. nats-server reads the cap once at start and refuses a config reload that changes it, so a new\ncap always needs a restart. The cut records the chat stream's frontier per retained seat, so a\nresumed seat catches up from there instead of replaying its channels. `registry` is the\nonly partial selection (`backup create ... --only registry`; `up --restore ... --restore-only\nregistry`). Backup never stops or restarts a mesh implicitly, never opens the original store, and\ndoes not contain credentials or trust secrets. Backup/restore in every auth mode, open included,\nuses isolated, operation-specific maintenance logins; normal agent credentials cannot enter that\nlistener. Full\nrestore requires the same space and exact current local trust continuity, recreates conservative\nconsumer checkpoints bound to their snapshot stream sequence state, and resumes retained agents under\ntheir original principals. The trust commitment includes the cryptographically validated full\noperator/system/data-account root chain as well as static/user authority state. A registry-only\nrestore completes canonical empty infrastructure but leaves retained agents stopped because their\nDM/DLV/TASK/ACL state is outside that selection. Authenticated restore validates the complete space\ntrust bundle before staging or changing the preserved store. Interrupted ordinary resume retries the\nsame durable attempt after its prior listener is stopped. Restore re-entry can recover a surviving normal listener\nonly when its attempt nonce, NATS server name, process owner, endpoint, and target-store identity all\nmatch the fsynced proof. A provably dead uncommitted owner is retired under lock and replaced with a\nfresh attempt-bound listener; an occupied foreign listener or ambiguous owner is never adopted. The\nmanager commit validates while retained cleanup is still suppressed; the CLI durably records its\nattempt-bound 64-hex token in `manager-committed` / `resume-committed` before `finalizeResume` can\nrelease suppression. A retry from either committed state goes straight to exact-token finalization;\nfailure preserves the committed gate and retained cleanup suppression. Missing commit evidence,\ninterrupted finalization, a live recorded endpoint despite missing pidfiles, or ambiguous proof fails closed. See the [CLI\nbackup and restore contract](cli.md#backups) for artifact, checkpoint, fallback,\ndisaster-consent, and degraded-recovery details.\n\n## Personas from the CLI\n\n`cotal personas` manages the local catalog offline: `list` (`--running` overlays live\nmarkers), `show <name>`, `edit <name>` (re-validates on save), `new <name>`, `rm <name>\n--force`. The runtime write is `cotal_persona`; the runtime read is `cotal_personas`\n(list / show), both over the wire with the manager's ownership checks. Fields: [agent files](agent-files.md).\n\n## Gate recovery\n\nA manager that dies mid-registration leaves its issuance gate *frozen* under that registration\nop. The freeze is correct: it stops two incarnations serving at once. The successor now completes\nthat dead op on boot, using the same guard as [`cotal reconcile-gate`](cli.md#reconcile-gate): it\nacts only when the freeze-holder is affirmatively gone under a complete CONNZ sweep (`gone` and\n`sweepComplete=true`). If the dead op's spec write committed, it finishes that same freeze\n(promote and reopen at the committed registration revision). If the spec did not advance, it\nabort-reopens the gate (generation+1, processEpoch unchanged) and continues the normal takeover.\nBoot heal and the following re-registration use separate one-shot executor windows, so a large\npredecessor family cannot spend the takeover's credential lifetime. If that later registration\nstill crosses a connection lifetime, it retries the same frozen operation with fresh authority\nand resumes verified-holder progress instead of freezing a new generation.\nA live holder, an incomplete sweep, or an unreachable delivery daemon still\nrefuses. Silence is never evidence of death, and there is no TTL. If holder verification is\ninterrupted, the frozen operation resumes from its durable, operation-and-gate-revision-bound\nprogress after liveness is checked again. A later freeze cannot reuse that progress: the cursor\nbinds the exact op, gate revision, and holder set. Use `cotal reconcile-gate` when the boot path cannot run\n(daemon down, a non-manager endpoint, or you want to lift the freeze without starting a manager). A spawn that hits the same frozen gate names that verb in the refusal\n(`blockedOp=registration`, the holding `opId`, `remedy=cotal reconcile-gate`) instead of a\nwait-timeout: the facts were always in the manager log; they now reach the spawn caller too.\n\nGive reconciliation a **quiet manager**. Suspend systemd restart policies, watchdogs, health-check\nrestart loops, and any other automation that can start or kill `cotal supervise` while boot healing or\n`cotal reconcile-gate` is running. Leave one recovery attempt in control until it finishes.\nRestarting the manager during the walk interrupts the current authority window. Durable progress makes\nthat interruption resumable, but a quiet manager is still the fastest and safest incident procedure.\n\n### Last-resort JetStream store replacement\n\nStore replacement is not normal gate recovery, is never automatic, and is destructive to mesh history.\nUse it only after the retained store cannot be reconciled and after deciding that losing its durable\ncontents is acceptable.\n\n1. Stop every actor touching the space: supervisor, watchdog, manager, delivery daemon, and broker.\n Confirm that no Cotal or NATS process still has the store open.\n2. Preserve the stopped store before changing anything. Move `.cotal/nats` aside to a dated backup and\n archive both `nats` and `auth`. Do not delete the only copy.\n3. Understand the loss: replacing the store removes JetStream message and control history and durable\n consumer state. Agent session files stored outside JetStream remain, but the mesh history they\n referenced does not.\n4. Start the broker against a new empty store, then start one manager. Wait until it reports\n serving successfully.\n5. Repopulate the mesh only after that manager is healthy. Re-enable supervisors, watchdogs, and other\n restart automation last.\n\nKeep the preserved store until the incident is reviewed and any required forensic or manual recovery is\ncomplete. Restoring it later restores the old durable state, including the fault that led to this last\nresort, so do not swap it back into a live mesh casually.\n\n## When something looks absent\n\nPermission denials are **loud, never silent**: an over-tight ACL rejects the endpoint call it\nrefuses instead of returning an empty or incomplete result that looks successful, and a denial no\ncall is waiting on, such as a refused subscription, shows up as a logged denial on the endpoint. Check\n`.cotal/manager.<key>.log`, `.cotal/delivery.<key>.log` (one pair per space, keyed as\n[Config](config.md#project-files) describes), and `.cotal/nats.log`; `cotal status` shows\nwhat is actually running. Those files live under the **project** `.cotal/`, not `~/.cotal`,\nunless the mesh root is the home directory. `cotal up --detach` redirects delivery and manager\nstdio onto those files, so an operator-created systemd unit around that launcher does not put\nthe child logs in that unit's journal. `journalctl -u <unit>` can be empty while the crash\nreason is already in the project log. Manager log lines start with the UTC time they were\nwritten. The access rules are collected in\n[Channels & permissions](channels-and-permissions.md).\n"
|
|
230
|
+
"body": "# Run a mesh\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nDay-to-day operation of a local mesh: what `cotal up` actually runs, how spawning\nresolves personas, harnesses, and models, how to reach a mesh from any directory, and the\noperator-only maintenance verbs. Every command's full flag set is in the\n[CLI reference](cli.md).\n\n## The stack\n\n`cotal up` brings up the whole local stack and bare `cotal down` stops it. Managed\nagents stay running as unmanaged OS processes; pass `--with-agents` to take them\nwith the stack. Seats of the built-in pty runtime run inside the manager process, so\nthey stop with the manager either way. Ctrl-C on a foreground `up` stops the manager through the\nsame stop as bare down and prints the same report; when that stop is refused, for example because\nthe manager cannot prove it can spare, Ctrl-C leaves the stack running, and you end it with\n`cotal down --with-agents`. A current manager records what its stop does with its seats before\nbare down signals it. A pre-pin legacy manager instead receives a reduced-guarantee\nwarning and is signalled according to the documented upgrade contract. Its running binary\nmay still carry the older destructive SIGTERM handler, so the CLI does not claim its\npre-signal agent inventory was spared; those agents may have been reaped.\n\n- **Broker**: a local `nats-server` (logs to `.cotal/nats.log`).\n- **Delivery daemon**: the durable backstop, auth mode only\n ([what it does](delivery-daemon.md)).\n- **Manager**: a detached supervisor answering the control plane, so\n `cotal spawn --detach` and the `cotal_spawn` tool work right after `up`.\n\nCotal creates the presence bucket in memory storage. Its records are liveness that every endpoint\nrewrites each heartbeat, so nothing is lost when a broker restart empties it, and nats-server's file\nstore write latch cannot reach it. A broker stop removes the memory stream itself, so every `cotal up`,\nincluding the resume after `cotal down --preserve-state`, creates it again before any daemon starts.\nJetStream fixes a stream's storage class when it is created, so a presence bucket created file-backed\nby an older cotal stays file-backed until that stream is recreated.\n\nA file-backed presence bucket can remain open and watchable while refusing every write. A bound\nendpoint reports this as `presence-write-stuck` after one full presence TTL of consecutive failures.\nThe roster is last-known while that condition is active. Restarting the broker clears nats-server's\nin-memory store latch and preserves the JetStream root. Current credentials split the required stream\nauthority: the `cotal up` provisioner can create the presence stream but cannot delete it, while the\nteardown credential can delete it but cannot recreate it. Cotal therefore reports the condition but\ndoes not attempt an unsafe partial delete-and-recreate. Stop and restart the broker to recover.\nA broker below nats-server 2.14.5 carries the latch (nats-server fixed it in 2.14.5). When `cotal up` starts or finds such a broker and the space's presence bucket is file-backed, it says so. A memory-backed bucket gets no warning. A broker below the SPEC §13.12 floor of 2.12 is refused at connect with the floor sentence.\n\nThree modes:\n\n- **Default (static auth).** JWT-authed, on by default: sender authenticity and per-agent\n ACLs, enforced by the broker ([how](identity-and-auth.md)).\n- **`--user-auth --idp <url>`.** Per-user auth: people `cotal login` once, the operator\n grants their agents on the actor ledger, and every connect is authorized live against\n that grant. Starts the space's auth service alongside the broker\n ([how](identity-and-auth.md)).\n- **`--open`.** An unauthenticated, live-only dev mesh (no auth, no delivery daemon). For\n quick local experiments.\n\nThe broker and local services bind **loopback** by default. `--host 0.0.0.0` widens the broker\nbind independently of the auth mode, so \"network-reachable\" never silently means\n\"unauthenticated\". With no explicit `--server`, `cotal up` auto-selects a free local port when\nthe default address is already held by another project; an explicit `--server` fails loud on\ncollision.\n\n`--host` is a boot flag, not a live rebind. A fresh `cotal up` writes the generated\n`.cotal/auth/server.conf` (project-local, not `~/.cotal`) with that bind and starts nats against\nit. If anything is already answering at the mesh URL, `up` refreshes the recorded mesh and\nleaves the running nats listener alone, so passing `--host 0.0.0.0` on a live or orphaned\nbroker does not change who can connect. To change the bind: `cotal down`, then `cotal up --host\n<addr>` against a stopped broker so the generated file is rewritten. Do not edit `server.conf`\nby hand; the next real boot overwrites it.\n\nOn a stopped shared broker, `up` renders every persisted space account and every enabled\nspace's auth-callout account into the resolver preload, regardless of which space starts\nthe broker. A missing callout account for an enabled space stops the boot rather than\nstarting with a reduced resolver. An already-running broker is refreshed without rewriting\nits config.\n\nA broker-only host is a first-class `up` mode. `cotal up --no-manager` boots the broker and, in\nauth mode, the delivery daemon, and no local manager, so the broker host never has a manager to\nstop and never leaves a manager slot stale. A refresh under the flag of a mesh whose manager is\nlive refuses rather than keeping or stopping it: `cotal down manager` first. Without the flag,\nauth-mode `up` still starts nats, the delivery daemon, and a\nlocal manager. A space may run more than one manager, addressed by instance id\n([control surface](control-surface.md#instance-routing)); putting no manager on the broker host\nis a topology choice, not a singleton invariant. A manager whose boot inventory has no\navailable connector does not take unpinned `spawn`/`launch` on the class rail, so a sibling\nthat can launch the harness can. `describe` still rides the class rail, so an unpinned spawn\ncan bind-fence when that skip member answered describe; re-issue, or pin `--on`. Pin one\ninstance with `--on` when a partial inventory still answers with a harness refusal. The\nsupported split is:\n\n```bash\n# broker host (project root that owns the generated conf, pidfiles, and logs)\ncotal up --detach --host 0.0.0.0 --space main --no-manager\n# no local manager starts: the summary lists nats-server + delivery daemon, and there is no\n# `.cotal/manager.<spaceKey>.log` to wait for on this host\n\n# manager host (registered remote mesh, same space)\ncotal meshes add --server nats://broker.example:4222 --root ~/meshes/main\ncotal supervise --space main --server nats://broker.example:4222\n```\n\nWait for `✓ manager up` in `.cotal/manager.<spaceKey>.log` on the manager host before spawning\nagents. On a broker host started without `--no-manager`, `cotal up --detach` prints `✓ running in\nthe background:` with `manager` listed once the manager pidfile is live; stop that local manager\nonly after the `✓ manager up` line. A host started WITH `--no-manager` never runs one, so neither\nthe wait nor the stop applies there. That detach stdout is not a safe teardown boundary: it is\npidfile liveness, not `✓ manager up`. `✓ manager up` is supervise's post-start line after\n`await mgr.start()`. `cotal down manager` after only the detach line can still default-terminate\nthe child during registration after it has taken the governance slot. Stopping before that\npost-start log line can leave the endpoint governance slot held until the holder's gate\nreopens past the stamp (the successor's boot heal, or\n[`cotal reconcile-gate`](cli.md#reconcile-gate) when that boot cannot run). See\n[Gate recovery](#gate-recovery).\n\nStandalone `cotal deliver --creds` is not a repair for that split. Production renewal needs\nthe manager and the daemon to address one credential store. The manager renews its own service\ncredential inside that credential's own window and re-dials its service connection with the\nrenewed credential; if the connection closes and cannot be restored within about forty seconds\nit releases its lease and exits so a restart can serve, while a broker that is briefly gone is\nwaited out. Separate host filesystems still\nleave manager root A writing and the daemon reloading root B; that composition is refused\nwhile the daemon stays up. Before every remint the manager challenges the delivery daemon's\nstore identity, and the answer must come from the process holding the delivery lease: the\nreply names the answering endpoint and the manager reads the lease row itself under its own\ncredential, so a non-holder answering on the queue-grouped admin rail is refused instead of\ncounting as the daemon's store. A rail that reports no responder is also settled from the\nlease row, so a live holder on record makes that outcome a refusal rather than an absent\ndaemon. Keep delivery on the broker host under `up`, and share one store\nonly when you are composing a hosted pair ([embedding](embedding.md#supervisor-signing-authority)).\nOn the `--no-manager` split above, the manager host's manager stays off the daemon-credential\nrenewal lease once its store check finds the daemon on another store. A filesystem store is named\nby its root and by a random id in `.cotal/store.id`, which the copied `.cotal/auth` does not carry,\nso this holds when both hosts use the same root path. `cotal doctor auth --fix` on\nthe broker host then renews the daemon credentials once they pass their renewal point.\n\n### Split host bind\n\nA remote manager cannot reach a loopback broker. After changing `--host`, confirm the\ngenerated `host:` in `.cotal/auth/server.conf` and that nats is listening on that address\nbefore registering the mesh on the manager host. Detached child logs stay under the **project**\n`.cotal/` that `up` ran in (see [When something looks absent](#when-something-looks-absent));\nthey are not `~/.cotal` unless that directory is the mesh root.\n\nA user-auth mesh can expose only its credential exchange through an operator-owned HTTPS reverse\nproxy while leaving the existing local exchange untouched:\n\n```bash\ncotal up --user-auth --idp https://idp.example/api/auth \\\n --exchange-public-port 7443 \\\n --exchange-public-url https://auth.example\n```\n\nThe public listener itself still binds `127.0.0.1:7443`; configure the proxy to terminate TLS and\nforward to it. It serves only `/health`, `/jwks`, `/exchange`, and `/.well-known/cotal-mesh` with\nthe documented methods. It needs no local file capability: the signed IdP JWT or managed-agent\nactor token is the proof, while the original loopback listener remains capability-gated. Add\n`--exchange-trusted-proxy` only when that listener is reachable exclusively through your trusted\nproxy; it keys failure throttling by the last `X-Forwarded-For` hop instead of the socket address.\nThe well-known bundle includes IdP pins and a deny-all sentinel credential, so fetch it only from\nthe configured HTTPS origin. To change these listener flags, stop and restart the mesh; a refresh\nof an already-running service does not replace its bind or proxy policy. See\n[Identity & auth](identity-and-auth.md#per-user-authentication) for the trust boundary.\n\n### Remote supervised seats by enrollment\n\nA remote seat does not need to run `cotal login` when the mesh owner pre-mints a single-use\nenrollment for it. Mount the enrollment URL as a private file, place the seat persona on the remote\nmachine, and launch the foreground seat:\n\n```bash\nCOTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \\\n cotal spawn --config ./worker.md --space main\n```\n\nThe URL is redeemed once with an unauthenticated GET. Redirects, off-machine plain HTTP, retries,\nand login fallback are refused. If the seat has no mesh record yet, the enrollment response's stock\nuser-bundle fields register it before the launch. The returned actor token then uses the same remote\nauth-service exchange as a login-provisioned agent. The enrollment URL and file path do not enter the\npreflight or harness environment. A failed or reused enrollment leaves no actor material on disk; ask the owner\nfor a fresh enrollment. When the foreground seat exits, this machine's credential files are removed\nand the mesh-side grant stays until the mesh operator revokes it; the launch line says so. The exact\nserver contract is in\n[Enrollment redeem](identity-and-auth.md#enrollment-redeem).\n\n`cotal status` prints the detailed setup, process, registry, and live mesh status. Its Machine\nsection names the running CLI's source checkout, installed package root, or npx package root beside\nthe version. It has one row per installed connector, which reports whether the executables that\nconnector declares in `requires` are on PATH. Status, setup and the manager's preflight resolve them\nthe same way: an entry written as a path is checked as given, and a directory never counts as the\nexecutable. A connector whose setup provider reports health adds its\nown rows above those. The Claude Code connector reports its plugin and its skills plugin, and a stale\nskills row names the installed and CLI versions it compared. `cotal\nsetup` (after the first run) prints the compact card.\n\nBefore reporting ready, the manager resolves every installed connector's declared harness\nbinaries against its own environment. A missing binary does not stop unrelated manager work: boot\ncontinues, but prints a named `connector <name> unavailable` line and records that reason in the\nmanager's `status` response. Available connector rows record the absolute paths boot resolved.\nA spawned seat and a seat resumed after `cotal down --preserve-state` both launch from those paths,\nand both are refused with the recorded reason when their connector's row is unavailable.\n`cotal models` takes the same rule and reports that reason in place of the catalog, so it agrees\nwith a launch about a harness installed or removed after boot. The manager looks again only when it\nrestarts. A connector registered after boot has no row, so spawn, resume and `cotal models` check\nits binaries on PATH when they run.\n\nOn an authenticated manager start, unfinished static lifecycle rows reconcile while the control\nendpoint is already serving. The manager `status` response reports\nthe `staticReconciliation` state, the last sweep counts, and each failed alias with its durable\nphase and literal disposition. `cotal status --components` reports the state and per-alias failure\ndetails. A slot row carrying a DEL or PURGE marker stops the sweep before it plans any alias, and\nthe manager log names the row. A failed exact terminal is retried in the same process after 1, 5,\nand 30 seconds. Each attempt re-reads the durable slot and re-enters the same deterministic terminal\noperation; the delays only schedule work and never release the lifecycle fence. The terminal's\ncleanup removes the lifecycle's credential file and its broker durables and read-ACL row as separate\nsteps. A file that cannot be removed does not leave the broker footprint behind, and its failure\nkeeps the alias held for the next attempt.\n\nOn shutdown, the manager fences new reconciliation work and waits for an exact terminal that already\nstarted. The current serial sweep stops before its next alias, and startup cannot publish the manager\nservice after `stop()` completes.\n\nThe four-attempt budget is per manager process. An exhausted row stays held and reports\n`retry-exhausted` with the remedy to restart the manager. The next process derives a fresh budget\nfrom the still-authoritative durable row. A `recovered` row remains visible until the next static\nreconciliation sweep, then clears. This component reports reconciliation outcomes. It does not say\nwhether footprint cleanup completed independently of the terminal result; that separate durable\nprojection remains tracked by #1274.\n\n`cotal service install` is the supported way to run the manager as a user service\n([CLI reference](cli.md#service)): a systemd user unit on Linux, a launchd agent on macOS, one\nper mesh, surviving logout and reboot. On Linux that needs user lingering: install refuses while\nit is off and prints the root command that enables it. It installs only\nthe manager; the units below remain the process models for every other component, and they are\nstill **examples of process models** for those: copy them only after you decide which processes\nthe unit should own.\n\n### Supervising the detached stack\n\n`cotal up --detach` is a launcher: it starts the broker, delivery daemon, and manager, reports what\nstarted, then exits. Do not wrap it in a systemd service with `Type=oneshot` and\n`RemainAfterExit=yes` and treat `systemctl is-active` as stack health. That unit becomes `active\n(exited)` when the launcher exits successfully and stays active even if every detached process dies.\nWhen `up --detach` can identify that exact unit shape, it prints a warning but keeps the requested\nstartup behavior.\n\nFor a single-host stack, keep `cotal up` itself in the foreground so systemd tracks a long-running\nprocess and restarts the stack if that process fails:\n\n```ini\n[Service]\nType=simple\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal up --space main --host 0.0.0.0\nRestart=on-failure\nRestartSec=5s\n```\n\nAn active unit then proves the foreground launcher and broker are still running, but it still does\nnot prove that every child component serves. Pair it with the component check below. Also remember\nthat `cotal up` starts a local manager as well as the broker and delivery daemon; run\n`cotal up --no-manager` (add the flag to the unit's `ExecStart` too) on a host intended to be\nbroker-only, so the unit and the host agree.\n\nSeats spawned by the built-in `pty` runtime run with `oom_score_adj` 500, so under memory\npressure the kernel prefers a seat over the broker, manager and delivery daemon, which are left as\nthey were started; the extension runtimes do not own the seat's process and get no preference.\n\nThat `Type=simple` shape puts nats in the unit's cgroup with the foreground `up` process. A\n`Restart=always` (or `on-failure`) of **this** unit therefore restarts nats as well, so remote\nmanagers drop for the time it takes the broker to come back. Wrapping `cotal up --detach` in\n`Type=oneshot` with `RemainAfterExit=yes` does not move nats out of that cgroup. Detached\nspawn starts a new process group, not a new systemd cgroup, and the default\n`KillMode=control-group` still signals every process left in the service cgroup on stop or\nrestart, including the nats PID. Escaping that cgroup needs an explicit unit setting such as\n`KillMode=process`, or a separate nats unit; this CLI does not ship that escape. The\n`Type=oneshot` unit below is a `cotal status --components` liveness check, not a\n`--detach` launcher. Neither trade is universal from\n`Type=simple` alone; it follows from which processes the unit actually owns. `cotal service\ninstall` covers only the manager, so for the broker and its siblings pick the example that\nmatches the ownership you want, and treat\n`systemctl is-active` as unit health, not mesh health.\n\nA broker that crashes under that foreground `up` keeps its mesh record and exits non-zero, so the\nunit's restart takes the repair path against the recorded store rather than starting a second one.\n\nIf the deployment deliberately uses `cotal up --detach` as a boot action, monitor observed state\ninstead of the launcher's exit:\n\n```ini\n[Unit]\nDescription=Check Cotal component liveness\n\n[Service]\nType=oneshot\nWorkingDirectory=/srv/cotal-mesh\nExecStart=/usr/bin/cotal status --components --space main\n```\n\nRun that check from a systemd timer or another monitor and alert on a nonzero exit. The command\ndistinguishes `absent`, `not-serving`, and `refused` components and never treats a sibling's health as\nproof. Its delivery-process check is local to the broker host, so run it there. On a split topology,\nalso probe the broker URL from the manager host and monitor the manager's own service there. A remote\nmanager cannot observe the broker host's delivery PID, and an `active` unit on either host says\nnothing about the other host.\n\nStop one part without tearing down the mesh by naming its registered component: `cotal down\nmanager`, `cotal down delivery`, or `cotal down web`. Component names from installed extensions\njoin the same surface; `cotal down` with no names retains whole-stack behavior and\nleaves managed agents running as unmanaged OS processes, except pty seats, which stop with the\nmanager. `cotal down --with-agents` is the previous reap. If a pinned manager has no\nspare-capability record, stop its managed agents explicitly before running that whole-stack\ncommand. A current manager always publishes the record, so it is absent only for an older manager,\nwhich may not understand the reap request.\n\n## Remote supervised agents\n\nOn a remote user-auth mesh, foreground `cotal spawn` remains the default participant path. A\nparticipant can run detached agents only after the host advertises and operates the remote manager\nauthority service, and the participant's actor-ledger row includes `supervise`. This is not implied\nby `spawn` or `admin`.\n\nThe participant's loopback/operator exchange obtains one closed `manager-service` view for its\nordinary derived owner, a fixed server-selected manager actor, and one opaque manager instance.\nThe host, not the participant, issues the public-nkey JWT material via the replay-safe,\nlifecycle-bound prepare → activate → renew exchange, plus a one-shot target-pinned retirement\nrequest for a host-managed terminal. It never exports the space signer, a static\nprovisioner credential, or generic storage authority. Remote registration publishes its service\nstatus at the registered revision and current process epoch, so manager-caller selection can find it.\n\nStock participant supervision asks its host to enroll a detached agent and to prepare its terminal\nretirement, over the same manager-authority transport. The stock auth service answers both when it\nruns with a public exchange face: it grants the agent under the participant's owner at a lifecycle\nUID it picks, bounded by the participant actor's own grant, provisions that UID's durables, and on\nretirement releases them and revokes the grant before the manager's terminal rail. It refuses a\nsecond enrollment of a name whose grant still stands until that agent's retirement is prepared. A\nhost platform that keeps these writers in its own storage intercepts both requests on its own route\ninstead. Copying host secrets or actor-ledger files to a participant is not supported. Foreground\nspawning and operator-local hosted managers use their existing paths.\n\nThe remote manager that `cotal supervise` starts can host workflow runs through its host: the host\nadmits each run and signs only the run's own driver, mediator and operator credentials. A logged-in\nuser's `cotal run start` against it is admitted: the auth callout issues the user's manager\nconnection, and the host binds each run to the owner who registered the manager. The run spawns\nagents that user owns, enrolled by the host like any detached spawn, with the reach the user's own\nrow grants when the spawn runs. A spawn may be placed on that manager and on no other instance. The\nhost's own manager refuses user-auth runs by name.\n[User-auth run start](https://github.com/Cotal-AI/Cotal/blob/main/docs/design/user-auth-run-start.md)\nrecords the path.\n\nThe registry entry decides the broker URL `supervise` dials, so a mesh published over `wss://` is\ndialed as a websocket. The manager-authority registration it runs first also takes its TLS\nrequirement from that entry, so the prepare credential is not exchanged over a plaintext\nconnection the record did not describe. `cotal meshes add` records both.\n\nWhen the authority service, login, or renewal is unavailable, the remote manager degrades\nfail-closed: it refuses new agents, restarts, and credential replacement rather than pretending\nlocal authority exists. Existing agents remain live only while their independent credentials are\nvalid. A hosted composition must revoke the managed grant and finish its resumable release before it\nrequests terminal retirement. Deleting DM or delivery consumers is not retirement and must not reset\na resumable lifecycle's frontier or pending state. The alias remains held until the terminal barrier\nconfirms. Restore service and renew successfully before asking it to recover an agent. See\n[Identity & auth](identity-and-auth.md#remote-manager-authority) and the [CLI\nreference](cli.md#supervise).\n\n## Spawning agents\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn reviewer --detach # supervised: the manager runs it in a PTY\ncotal attach --name reviewer # watch/type into a detached agent (Ctrl-] detaches)\ncotal ps # what the manager is running\ncotal stop --name reviewer # stop one\n```\n\nHow a spawn resolves:\n\n- **Persona.** A bare `cotal spawn` uses `.cotal/agents/default.md`; a positional name\n picks `.cotal/agents/<name>.md`; `--config` takes an explicit ref or path. Set\n `COTAL_DEFAULT_PERSONA=<name-or-path>` to change the fallback. Fields and format:\n [agent files](agent-files.md).\n- **Harness.** Resolution order is an explicit `--agent` or `cotal_spawn` `agent` argument,\n then the persona file's `agent:` pin, then the invoking caller's `COTAL_DEFAULT_AGENT`,\n then the manager's `COTAL_DEFAULT_AGENT`, then the product default (Claude). Compared in\n [Connectors](connectors.md); per-connector guides:\n [Claude](connect-claude.md) · [OpenCode](connect-opencode.md) ·\n [Hermes](connect-hermes.md) · [pi](connect-pi.md).\n- **Model.** `--model` overrides the persona file's `model:` (Claude: `opus` / `sonnet` or\n a full id; OpenCode: `provider/model`). Connectors that expose a catalog report it via\n `cotal models --agent opencode`: model ids plus available variants; pick one with\n `--model provider/model --variant high`.\n- **Tools.** A spawned Claude Code agent gets the cotal tools plus the MCP servers the cotal\n config shares, which first-run `cotal setup` fills with your own; narrow them per spawn with\n `--share-tools` ([config](config.md)).\n- **Launch options.** `--opt key=value` (repeatable) passes a native harness flag straight\n through; a persona or manifest `launchOptions:` mapping does the same declaratively (a\n `--opt` wins per key). It is a **raw passthrough**, with no allow/deny list: Claude renders\n each as `--key value` (a bare `--key` for an empty value), OpenCode merges them into its\n agent config, and Hermes has no option surface so it fails loud. The trust boundary is the\n `spawn` capability itself, not the flag set, so granting `spawn` is host-launch authority\n ([security](security.md)). A key must be a plain flag name; malformed or prototype-polluting\n keys are refused.\n\nDetach from an attached PTY with **Ctrl-]** (the agent keeps running); rebind it with\n`COTAL_DETACH_KEY=ctrl-<char>` when it clashes with a keybinding inside the agent's TUI.\n\n**Runtimes.** The manager spawns into a **pty** by default. It spawns the PTY in-process on\nevery platform, so replacing the manager worker closes its seats and the pty runtime gives no hot\nupdate. Any manager stop, bare `cotal down` included, stops and deprovisions those seats. A stopping\nmanager refuses new spawns and first waits for the ones it already accepted, so their seats stop too. On Linux\nit can still adopt and reap seats that an earlier manager launched under a detached per-seat\ncustodian, so those seats drain under the new manager; it starts no new custodian. A custodian whose agent has exited exits a few seconds later on its own. `cotal seats`\nlists the custodians left on the machine, and `cotal seats --drain` retires the ones whose agent\nhas exited while keeping every seat whose agent still runs ([cli.md](cli.md#seats)). When a pty\nagent exits on its own, in-process or under a custodian, the manager logs a `seat reaped:` line\nwith the exit code and, for a signalled child, the signal number. The line ends with the last line\nthe child printed that starts with a connector's `[cotal-<name>]` or `[cotal-<name>/<part>]`\nprefix, cut to 240 characters, when it printed one. A custodian keeps the same record beside the\nseat's custody record, so a later reap of that seat, including one by a\nsuccessor manager, reports how the child ended. When the custodian cannot write that record, it\nsays why in the seat's `custodian.log`, and a later reap of a child that ended on its own reports\nthe record as missing or unreadable. Optional runtimes are installed\nthrough the extension surface, for example `cotal ext add @cotal-ai/orca`, then selected with\n`--runtime orca` (similarly `@cotal-ai/tmux`, `@cotal-ai/cmux`, and `@cotal-ai/herdr`). They put teammates in native\nterminal surfaces rather than manager-owned PTYs. Runtime names are open-ended and resolved from\nthe registry; a missing provider or app throws, never silently falls back\n([architecture](architecture.md)).\n\n## Mesh registry\n\n`cotal up` records each running mesh in a machine-local registry\n(`~/.cotal/meshes/space.<key>.json`, named by a case-safe hex encoding of the space: broker URL, the project root holding its creds and\npersonas, and its mode). So a bare `cotal spawn <persona>` from *any* directory joins the\nrunning mesh with the right credentials instead of mistaking the cwd for a space:\n\n- `cotal use <name>` sets the default from every directory, including inside another mesh's\n project. `--space <name>` overrides it for one command.\n- When one broker has records for several spaces, `cotal up --space <name>` refreshes that named\n space.\n- A refresh rewrites only what that command decided: the server, root and mode, the user-auth\n endpoints, and an explicit `--host` or `--max-sessions`. Every other field, such as the TLS\n requirement, is kept as the record stands when the refresh writes it, so a change another\n command made during the refresh survives. If the record was removed during the refresh, `up`\n fails instead of writing it back.\n- With no live selected default, a project with its own `.cotal/` resolves to that project's\n mesh; otherwise one running mesh is used automatically and several are an error.\n- `cotal meshes` lists them (a `*` marks the default); `cotal down` removes the entry.\n\nThe registry stores a *path*, never a secret; trust material stays in each project's\n`.cotal/auth`. If the mesh is down or won't take your creds, spawn fails with one\nsentence, never a raw NATS trace.\n\n### Meshes you did not start here\n\nA mesh running on another machine has no `cotal up` on this one, so register it by hand:\n\n```bash\ncotal meshes add # guided: asks for the broker, probes it, offers what it finds\ncotal meshes add optiplex --server nats://100.90.12.34:4222 --root ~/meshes/optiplex \\\n --allow-unencrypted-overlay # see below: an overlay address needs this\ncotal meshes rm optiplex\n```\n\nOn a terminal, a bare `cotal meshes add` walks you through it: it probes the broker you name and\nreports whether it is open or requires credentials, offers the spaces the folder already holds\ncredentials for, and shows the record before writing it. Scripts and agents keep the flag form -\nwithout a terminal nothing prompts.\n\n`--root` is the local folder holding that mesh's `.cotal/auth` and `.cotal/agents` (its personas);\nthe mode is inferred from what that folder holds.\n\nThe instance identities of the manager and the user-auth service are not part of that folder.\nEach root keeps its own in `.cotal/space.<hex>/`, so `cotal supervise` or `cotal up --user-auth` in\nthe root you copied the folder to starts an instance of its own. A root last run by an older Cotal\nstill holds them in `.cotal/auth`, as `manager-instance.<hex>.json`, `manager-siblings.<hex>.json`\nand `space.<hex>/.cotal/auth/auth-instance.<hex>.json`. Delete those files from a copy of such a\nfolder before the first `cotal supervise` or `cotal up` there.\n\n**Know what you are copying.** For an authenticated mesh that folder carries the space's account\n**signing seed**, which is the authority to mint any identity in the space. A machine holding it\nis a certificate authority for the mesh rather than a client of it: anyone who reads it can\nimpersonate any agent, read every retained channel and DM, change ACLs, and keep issuing\nthemselves credentials. There is no per-machine revocation; undoing it means rotating the signing\nkey and re-minting every credential in the space. Copy it only to machines you would trust with\nthe whole mesh. `cotal mint` on its own does not substitute here: registering an `auth` mesh needs\nsigning material that composes, which a minted user credential is not. The\nbroker is probed before the record is written, so a bad address or a credential that mesh will not\naccept fails at registration rather than at your first `spawn` (`--force` records it without verifying,\nuseful when the mesh is simply down right now).\n\n#### Which addresses you may register\n\nRegistering a mesh is how this machine starts sending agent credentials to a broker it does not\nrun. NATS announces itself in plaintext before anyone authenticates, so an attacker on the path\ncan pose as the broker and read the credential out of the connect unless the connection\n**requires TLS**, which is recorded on the entry and enforced on every dial through it.\n\nWhat the record will require decides what you may register:\n\n- **Without required TLS**, the address is the gate: **loopback** (`127.0.0.0/8`, `::1`), or\n **your private overlay** (`100.64.0.0/10`, `fd7a:115c:a1e0::/48`) with\n `--allow-unencrypted-overlay`. The tunnel provides the protection, and this command cannot check\n its state. Hostnames are refused because the lookup would choose which machine receives your\n credentials.\n- **With required TLS**, set `--tls` or use a `tls://` URL. The recorded scheme enforces the TLS\n requirement. A **hostname or public address** is accepted because the certificate chain and\n hostname check identify the peer. A registration whose broker cannot complete the handshake\n fails unless you pass `--force`, which records the entry without verification.\n\nOrdinary private ranges like `10.x` and `192.168.x` are refused in **both** modes. A café's wifi\nis private but does not belong to you, and no public CA issues certificates for those ranges. An\naddress spelling changes nothing: `[::ffff:192.168.1.10]`, `3232235786`, `0300.0250.01.012`, and\n`192.168.257` all resolve to private addresses and receive the same refusal as the dotted form.\n`--force` exists for a mesh that is down. It never permits an unsafe credential destination.\n\n#### Registering a hosted user-auth mesh\n\nA user-auth space's IdP pins are established where the mesh runs and are never guessed. Register\none from **supplied** trust: `--user-auth-file bundle.json` (exported on the mesh's machine), or\n`--from https://auth.example`, which asks before it contacts the address at all, fetches the\ndiscovery document at `/.well-known/cotal-mesh` under that address over HTTPS, shows you the pins,\nand asks again before adopting them. A URL that already ends in `/.well-known/cotal-mesh` is\nfetched as given. Redirects are refused because a 302 can walk a pinned fetch down to\nplaintext or onto another host, and the pinned exchange must be an `https://` URL too. The one\nexception is an exchange on **this machine**, where nothing leaves the box: plain `http://` is\naccepted for a loopback *literal* (`127.0.0.1`, `::1`, and any spelling of them), but **not** for\n`localhost`, which a hosts entry or poisoned lookup could point elsewhere. Use the\nliteral. Registration checks that the pinned exchange\nanswers `/health` and `/jwks` as the pinned issuer. It also checks that the broker refuses a\nbare connect; that refusal is the pass. The bundle's sentinel credentials are written to a private (0600) file\nunder the entry's root; the registry itself never carries the secret.\n\n**Without required TLS**, an overlay address is **refused unless you accept the dependency\nexplicitly**, with `--allow-unencrypted-overlay`. The address is not the guarantee: it is protected\nwhile the tunnel is up, and if the tunnel is down that range is ordinary carrier-grade NAT and\nwhoever answers the dial receives your credentials. Only you can know which it is, so the command\nasks you to say so. Your acceptance is recorded on the mesh entry rather than printed and\nforgotten, and the guided form asks the same question instead of taking the flag.\n\n**With required TLS** (`--tls`, or a `tls://` URL) that consent is no longer asked for, and the\nflag is not needed: the handshake is what protects the connection, so the acceptance it stood in\nfor has been replaced by proof rather than promise. `cotal meshes add <space> --server\nnats://100.64.0.1 --tls` registers an overlay address with no prompt, no flag and no recorded\nacceptance. This is the \"the flag disappears once the broker can be served over TLS\" case, and it\nhas now arrived.\n\nThis gate is on **registration**. `cotal join --creds --server <url>` deliberately takes an\nexplicit connection at face value and does not consult the registry, so it is not covered. Join\nthat way only to an address you would have registered.\n\nThe connection is still probed first, with the same second try at the longer budget the registry\npreflight uses, so a slow link reads as a connect that did not finish within that budget and a\nrefused port reads as a broker that is not running.\n\nRecords added this way are removed only by something that names them. A failed liveness probe\ndoes not delete any record: an unreachable broker, local or registered by hand, is shown as\n`offline` in `cotal meshes`. A bare command does not count that offline record as running;\nname it with `--space` to restart it. `cotal down` / `cotal clean all` still drop an `up` record for the\nproject they are tearing down, and they leave a hand-registered one alone even when `--root`\npointed at that project. A `cotal up` for that space refuses outright unless it is that same\nendpoint: finding a broker already answering there is a refresh that starts nothing and leaves the\nrecord's provenance alone, while actually starting the broker for that space, server and root\nmakes this machine the one running it, so the record becomes an ordinary local one that\n`cotal down` clears. The refusal names `cotal supervise --space <s> --server <url>` (plus `cotal\ndeliver`) when the registered broker is on another host, and `cotal meshes rm` when it is local.\n`cotal meshes rm` drops it and re-registering with `--force` replaces it. `rm` only forgets a\nmesh. To stop one running here, use `cotal down`.\n\n## Watching\n\n`cotal console` is the terminal view (TUI on a real terminal, plain line stream when\npiped); `cotal web` is the browser dashboard. Both are read-only observers; the\nwalkthrough is [Watch a mesh](watch-a-mesh.md).\n\n## History\n\nRetained history is operator-owned. `cotal clean history --force` purges a space's\nretained channel history; `--dms` also purges DMs (`cotal history clear` is an alias).\nIt is deliberately **not** an agent tool: agents cannot wipe the record\n([identity & auth](identity-and-auth.md)). For a **stopped** mesh, `cotal clean store\n--force` deletes the on-disk JetStream store outright, and `cotal clean all --force`\nalso resets the space identity ([CLI reference](cli.md#clean)).\n\n## Offline backup\n\nFor a coherent durable cut, preserve the whole stack first, then create the artifact while it stays\ndown:\n\n```bash\ncotal down --preserve-state\ncotal backup create ./space-backup # full by default\n# later: deliberately resume the unchanged source\ncotal up --detach\n# or, from another preserved cut, restore before the normal listener opens\ncotal up --restore ./space-backup --detach\n```\n\nA refused cut leaves the mesh running and unfenced: fix what the refusal names and run\n`cotal down --preserve-state` again.\n\nUse `--store-dir` on both preservation and backup for a custom JetStream store. A store cap set\nwith `cotal up --max-file-store <bytes>` travels with the preserved state, and the resume renders it\nagain. nats-server reads the cap once at start and refuses a config reload that changes it, so a new\ncap always needs a restart. The cut records the chat stream's frontier per retained seat, so a\nresumed seat catches up from there instead of replaying its channels. `registry` is the\nonly partial selection (`backup create ... --only registry`; `up --restore ... --restore-only\nregistry`). Backup never stops or restarts a mesh implicitly, never opens the original store, and\ndoes not contain credentials or trust secrets. Backup/restore in every auth mode, open included,\nuses isolated, operation-specific maintenance logins; normal agent credentials cannot enter that\nlistener. Full\nrestore requires the same space and exact current local trust continuity, recreates conservative\nconsumer checkpoints bound to their snapshot stream sequence state, and resumes retained agents under\ntheir original principals. The trust commitment includes the cryptographically validated full\noperator/system/data-account root chain as well as static/user authority state. A registry-only\nrestore completes canonical empty infrastructure but leaves retained agents stopped because their\nDM/DLV/TASK/ACL state is outside that selection. Authenticated restore validates the complete space\ntrust bundle before staging or changing the preserved store. Interrupted ordinary resume retries the\nsame durable attempt after its prior listener is stopped. Restore re-entry can recover a surviving normal listener\nonly when its attempt nonce, NATS server name, process owner, endpoint, and target-store identity all\nmatch the fsynced proof. A provably dead uncommitted owner is retired under lock and replaced with a\nfresh attempt-bound listener; an occupied foreign listener or ambiguous owner is never adopted. The\nmanager commit validates while retained cleanup is still suppressed; the CLI durably records its\nattempt-bound 64-hex token in `manager-committed` / `resume-committed` before `finalizeResume` can\nrelease suppression. A retry from either committed state goes straight to exact-token finalization;\nfailure preserves the committed gate and retained cleanup suppression. Missing commit evidence,\ninterrupted finalization, a live recorded endpoint despite missing pidfiles, or ambiguous proof fails closed. See the [CLI\nbackup and restore contract](cli.md#backups) for artifact, checkpoint, fallback,\ndisaster-consent, and degraded-recovery details.\n\n## Personas from the CLI\n\n`cotal personas` manages the local catalog offline: `list` (`--running` overlays live\nmarkers), `show <name>`, `edit <name>` (re-validates on save), `new <name>`, `rm <name>\n--force`. The runtime write is `cotal_persona`; the runtime read is `cotal_personas`\n(list / show), both over the wire with the manager's ownership checks. Fields: [agent files](agent-files.md).\n\n## Gate recovery\n\nA manager that dies mid-registration leaves its issuance gate *frozen* under that registration\nop. The freeze is correct: it stops two incarnations serving at once. The successor now completes\nthat dead op on boot, using the same guard as [`cotal reconcile-gate`](cli.md#reconcile-gate): it\nacts only when the freeze-holder is affirmatively gone under a complete CONNZ sweep (`gone` and\n`sweepComplete=true`). If the dead op's spec write committed, it finishes that same freeze\n(promote and reopen at the committed registration revision). If the spec did not advance, it\nabort-reopens the gate (generation+1, processEpoch unchanged) and continues the normal takeover.\nBoot heal and the following re-registration use separate one-shot executor windows, so a large\npredecessor family cannot spend the takeover's credential lifetime. If that later registration\nstill crosses a connection lifetime, it retries the same frozen operation with fresh authority\nand resumes verified-holder progress instead of freezing a new generation.\nA live holder, an incomplete sweep, or an unreachable delivery daemon still\nrefuses. Silence is never evidence of death, and there is no TTL. If holder verification is\ninterrupted, the frozen operation resumes from its durable, operation-and-gate-revision-bound\nprogress after liveness is checked again. A later freeze cannot reuse that progress: the cursor\nbinds the exact op, gate revision, and holder set. Use `cotal reconcile-gate` when the boot path cannot run\n(daemon down, a non-manager endpoint, or you want to lift the freeze without starting a manager). A spawn that hits the same frozen gate names that verb in the refusal\n(`blockedOp=registration`, the holding `opId`, `remedy=cotal reconcile-gate`) instead of a\nwait-timeout: the facts were always in the manager log; they now reach the spawn caller too.\n\nGive reconciliation a **quiet manager**. Suspend systemd restart policies, watchdogs, health-check\nrestart loops, and any other automation that can start or kill `cotal supervise` while boot healing or\n`cotal reconcile-gate` is running. Leave one recovery attempt in control until it finishes.\nRestarting the manager during the walk interrupts the current authority window. Durable progress makes\nthat interruption resumable, but a quiet manager is still the fastest and safest incident procedure.\n\n### Last-resort JetStream store replacement\n\nStore replacement is not normal gate recovery, is never automatic, and is destructive to mesh history.\nUse it only after the retained store cannot be reconciled and after deciding that losing its durable\ncontents is acceptable.\n\n1. Stop every actor touching the space: supervisor, watchdog, manager, delivery daemon, and broker.\n Confirm that no Cotal or NATS process still has the store open.\n2. Preserve the stopped store before changing anything. Move `.cotal/nats` aside to a dated backup and\n archive both `nats` and `auth`. Do not delete the only copy.\n3. Understand the loss: replacing the store removes JetStream message and control history and durable\n consumer state. Agent session files stored outside JetStream remain, but the mesh history they\n referenced does not.\n4. Start the broker against a new empty store, then start one manager. Wait until it reports\n serving successfully.\n5. Repopulate the mesh only after that manager is healthy. Re-enable supervisors, watchdogs, and other\n restart automation last.\n\nKeep the preserved store until the incident is reviewed and any required forensic or manual recovery is\ncomplete. Restoring it later restores the old durable state, including the fault that led to this last\nresort, so do not swap it back into a live mesh casually.\n\n## When something looks absent\n\nPermission denials are **loud, never silent**: an over-tight ACL rejects the endpoint call it\nrefuses instead of returning an empty or incomplete result that looks successful, and a denial no\ncall is waiting on, such as a refused subscription, shows up as a logged denial on the endpoint. Check\n`.cotal/manager.<key>.log`, `.cotal/delivery.<key>.log` (one pair per space, keyed as\n[Config](config.md#project-files) describes), and `.cotal/nats.log`; `cotal status` shows\nwhat is actually running. Those files live under the **project** `.cotal/`, not `~/.cotal`,\nunless the mesh root is the home directory. `cotal up --detach` redirects delivery and manager\nstdio onto those files, so an operator-created systemd unit around that launcher does not put\nthe child logs in that unit's journal. `journalctl -u <unit>` can be empty while the crash\nreason is already in the project log. Manager log lines start with the UTC time they were\nwritten. The access rules are collected in\n[Channels & permissions](channels-and-permissions.md).\n"
|
|
231
231
|
},
|
|
232
232
|
{
|
|
233
233
|
"slug": "security",
|
|
@@ -269,7 +269,7 @@ export function loadDocsBundle() {
|
|
|
269
269
|
"title": "Upgrading a running deployment",
|
|
270
270
|
"kind": "Guide (informative)",
|
|
271
271
|
"summary": "Substrate stability tells you what the version numbers promise.",
|
|
272
|
-
"body": "# Upgrading a running deployment\n\n> **Guide** (informative) · **For:** operators upgrading a mesh that already exists · **See also:** [Substrate stability](stability.md), [Run a mesh](run-a-mesh.md), [Identity and auth](identity-and-auth.md)\n\n[Substrate stability](stability.md) tells you what the version numbers promise. This page is the\nother half: what to actually do when the deployment already exists, has credentials in it, and\ncannot simply be recreated. Every release that breaks a running deployment gets a section here,\nnaming what migrates on its own, what does not, and the order to move the pieces in.\n\n## The pre-1.0 upgrade contract\n\nThe packages are pre-1.0, so a minor bump may break an API or an on-disk expectation. Four\ncommitments make that survivable for someone with a fleet:\n\n- **Pin an exact version.** `0.N.P`, never `^0.N.P`. A range can pull a breaking minor in during an\n unrelated reinstall.\n- **Every break that touches a running deployment gets a section on this page**, written in terms of\n what an operator does, not in terms of which module changed.\n- **Read the section before you start, not halfway through.** A section names the work up front\n precisely so the operation does not change shape once it is underway.\n- **A break that cannot be made automatic says so.** Where credentials or state must be recreated by\n hand, the section says which ones and when, rather than leaving you to discover it at the moment\n the first one stops working.\n- **A change to the shape of a credential, or to who may renew one, is breaking whatever the commit\n marker says.** This rule is stated because the marker is a judgement made while writing the code\n and the consequence is felt by someone running it a day later. A fleet that keeps authenticating\n looks compatible and is not, if nothing in it can renew. Any automated check of this rule would\n read commit markers, so a break recorded as a feature is the one case it could not see, which is\n why the rule is written for people first. **The marker held for this release: the 0.49.0 change\n that caused all of this, `36d177951 feat(core)!`, did carry its `!`.** The rule exists for the\n next one that does not.\n\nWhat this page does not promise is a rolling upgrade. Nothing in the current line dual-serves two\nauthority versions, so where broker and manager run separately there is a window in which the mesh\nis down. The sections below give that window's shape so it can be scheduled rather than endured.\n\n## Hermes model from the environment in 0.68.0\n\nA connector now launches on the model and variant its launcher resolved (the `--model` or\n`--variant` flag, else the agent file's `model:` or `variant:`) and no longer reads them again from\nthe agent file. The Hermes connector also no longer takes a model from `HERMES_MODEL` in the\nenvironment of the process that spawns the seat, including when `spawn.env` lists it.\n\n### What stops working\n\nA Hermes spawn whose only model was `HERMES_MODEL` in the spawning environment is refused at launch,\nand the refusal names both ways to set a model. Spawns that set `--model` or `model:` are unchanged,\non every connector.\n\nCode that calls a connector's `buildLaunch` directly with only `configPath` now gets no model or\nvariant from that file. Pass them as `model` and `variant`.\n\n### Before the upgrade\n\nMove each Hermes seat's model from `HERMES_MODEL` onto its spawn with `--model`, or into its\npersona's `model:`.\n\n## Run answers on a participant manager in 0.68.0\n\nA participant manager now asks its issuing host for an answering credential by naming the run and\nstep it answers. The host reads the pause's token off that run's journal and no longer accepts a\ntoken from the manager. Runs on a mesh with no participant manager are unaffected.\n\n### What stops working\n\nWhile a participant manager and its issuing host run different sides of this release, the host\nrefuses every `cotal run answer` and every amendment that manager serves, because each side refuses\nthe other's request shape. Starting, resuming and reading runs is unchanged. A pause stays waiting\nthrough the window, or follows its timeout if it has one.\n\n### Before the upgrade\n\nUpgrade the auth service and every participant manager registered with it in the same window, then\nanswer the pauses that waited.\n\n## Headless OpenCode handshake in 0.69.0\n\nWith `COTAL_SERVE_HEADLESS=1`, the OpenCode launcher's `[cotal-serve]` line on stdout now carries\nonly `port` and `session`. The server password no longer appears in it, and the 1.x TUI no longer\nreceives the password on its command line.\n\n### What stops working\n\nA headless host that read `password` from that line has no password, and the server refuses its\nrequests. Seats with a TUI, and headless seats that no host drives, are unaffected.\n\n### Before the upgrade\n\nHave each headless host mint a password and pass it to the launcher as `OPENCODE_SERVER_PASSWORD`,\nthen use it for basic auth as before. Without that variable the launcher mints its own.\n\n## Filesystem store identity in 0.69.0\n\nThe delivery daemon's answer to the manager's store check now names a filesystem store by its root\nand by a random `id` that the store records once in `store.id` inside its own directory:\n`.cotal/store.id` for a workspace root, or the directory of the file for `cotal deliver --creds\n<file>`. A manager no longer counts the daemon's store as its own because the two roots have the\nsame path. On a split whose broker host and manager host use one root path, the manager host now\nstays off the daemon-credential renewal lease, so `cotal doctor auth --fix` on the broker host can\nrenew the daemon credentials.\n\n### What stops working\n\nA manager and a delivery daemon on different sides of this release refuse each other's answer to\nthe store check. The manager then remints no daemon credential, and a manager that is booting does\nnot start. This is read from the code and was not measured across two releases. A\n`cotal deliver --creds <file>` whose directory is a read-only mount and holds no `store.id` stops at\nstart. So does a `--creds` file that is its directory's `store.id` under any name, and a `store.id`\nthat is a symbolic link or holds anything but a lowercase UUID.\n\n### Before the upgrade\n\nUpgrade the broker host and every manager host of a space in the same window. For a `--creds` file\non a read-only mount, add a regular `store.id` file beside it that holds a new lowercase UUID and no newline,\nas `node -e 'process.stdout.write(crypto.randomUUID())' > store.id` writes. Move a `--creds` file\nnamed or linked as `store.id` to a file of its own.\n\n## Detached spawns with `--share-tools` in 0.69.0\n\nThe manager's `spawn` operation now takes `shareTools` as a list of MCP server names. The CLI parses\n`--share-tools` into that list before it sends the request, and the manager cluster document moves\nto revision 22. A cut taken with `cotal down --preserve-state` before the upgrade still resumes: the\nmanager reads its `cotal-manager-resume/v1` inventory and writes new cuts as\n`cotal-manager-resume/v2`.\n\n### What stops working\n\nA CLI and a manager on different sides of this release refuse a detached spawn that passes\n`--share-tools`, because the CLI checks each request against the contract the manager serves. This\nis read from the code and was not measured across two releases. A detached spawn without the flag,\na foreground spawn and a roster entry are unaffected. A manager older than this release cannot\nresume a cut that this release took.\n\n### Before the upgrade\n\nUpgrade the CLI on every host that runs `cotal spawn --detach` in the same window as the managers\nit reaches.\n\n## Shared MCP server checks in 0.69.0\n\nThe cotal config reader now checks each server under `connectors.<name>.mcpServers` when it reads\nthe file, and refuses one that cannot launch as written, naming the file and the field. The rules\nare in [the config file](config.md#the-config-file).\n\n### What stops working\n\nA config file that holds such a server refuses every Claude spawn that reads it, including one with\n`--share-tools none`. Before, a field of the wrong type failed each Claude spawn that shared the\nserver with a `TypeError` that named neither the file nor the server, a spawn that did not share it\nlaunched, and a server with no `command` or `url` was passed to `claude`, which never started it.\nRead from the code and not measured: spawns on other connectors, a manager resume and the step of\n`cotal setup` that records the shared list read the same files, so each stops at the same refusal.\n\n### Before the upgrade\n\nCheck `connectors.<name>.mcpServers` in the operator-level config file and in each space's\n`.cotal/config.json`. Give each server a string `command`, or a `type` of `http`, `sse` or `ws` with\na string `url`. Write `args` as a list of strings and `env` and `headers` as objects of strings, or\nremove the server.\n\n## Remote manager family eviction in 0.69.0\n\nA remote manager registered through its host now asks the host to evict up to 256 holders of its\ncredential family in one maintenance request, and the host reads the family once for the whole set.\nBefore, a restart sent one request per holder and the host read the whole family for each one.\nMeshes with no remote manager are unaffected.\n\n### What stops working\n\nWhile a remote manager and its issuing host run different sides of this release, each side refuses\nthe other's eviction request shape. A restart whose credential family already has holders then fails\nat its eviction step and leaves the manager's registration gate frozen. A first start, a clean stop\nand the host's reconciliation of a foreign slot holder are unchanged.\n\n### Before the upgrade\n\nUpgrade the auth service and every remote manager registered with it in the same window. A manager\nthat restarted inside the window resumes its frozen registration on its next start once both sides\nrun this release.\n\n## AG-UI emitter holder hooks in 0.69.0\n\n`AguiEmitterHolder` from `@cotal-ai/connector-core` now takes its hooks as one named object after\nthe emitter factory: `new AguiEmitterHolder(startEmitter, { onError, onRunClosed, waitLive, runMeta })`.\nOnly `onError` is required. Nothing about a running mesh changes, and every shipped connector passes\nits hooks by name. Only a connector of your own that builds a holder is affected.\n\n### What stops working\n\nA holder built with positional hooks, such as `new AguiEmitterHolder(start, onError, onRunClosed)`,\nno longer compiles, because the constructor takes two arguments. Plain JavaScript that keeps the\npositional form still runs, but the holder calls none of its hooks, so a failure never reaches\n`onError`.\n\n### Before the upgrade\n\nPass each hook by name, for example `new AguiEmitterHolder(start, { onError, onRunClosed })`, and\ndrop any `undefined` that filled an earlier slot to reach a later hook.\n\n## Worker run failure type in 0.70.0\n\n`WorkerRunFailed`, the failed result of `runInWorker` in `@cotal-ai/lang`, is now a union on\n`class`: `released`, `held`, `effect`, `too-large`, `rejected` or `error`. A running mesh needs\nnothing, because the runtime host and the engine thread ship in the same install. A run on the\ncompiled engine whose program throws an object with `code: \"L5012\"` or `code: \"L5025\"` used to end\nreleased and now ends failed, as it does on the walker.\n\n### What stops working\n\nTypeScript code that reads `code`, `reason`, `step`, `pending`, `kind`, `detail` or `tooLarge` on a\n`WorkerRunFailed` it has not narrowed fails with TS2339. A released, held, too-large or rejected\nresult no longer carries `code`, so JavaScript that branched on `L5012`, `L5025`, `L5006` or\n`L5010` stops matching with no error. `tooLarge` is gone.\n\n### Before the upgrade\n\nBranch on `class` where such code read `code`: `released` for L5012, `held` for L5025, `too-large`\nfor L5006 and `rejected` for L5010. An `effect` or `error` result keeps its `code`. Once narrowed to\n`too-large`, a result carries the `stepKey`, `bytes` and `bound` that `tooLarge` held.\n\n## Remote manager request builder in 0.70.0\n\n`remoteManagerClient.remoteManagerAuthorityRequest` from `@cotal-ai/manager` now takes an\noperation's coordinates as one object, and `remoteManagerRegistrationProof` from `@cotal-ai/core`\ncomputes the proof from the manager's identity state instead of a request. Nothing about a running\nmesh changes: the proof digest and the request on the wire are the same, so a manager and a host on\ndifferent sides of this release still accept each other. Only code that builds remote manager\nrequests itself is affected, in TypeScript and in plain JavaScript.\n\n### What stops working\n\nA call that passes the registration proof, contract artifacts, session, retirement or transfer\nreader as positional arguments after the operation no longer compiles. A call that passes a request\nto `remoteManagerRegistrationProof` no longer compiles either, because the second argument now names\nthe lifecycle `lifecycleUid`, as the identity state does.\n\nPlain JavaScript runs both old calls without an error. The builder drops the positional coordinates,\nso the host refuses the request with `requires a sha256 registrationProof`. A proof computed from a\nrequest leaves out the lifecycle, so the host refuses a request that carries it as a proof mismatch.\n\n### Before the upgrade\n\nName the coordinates, for example\n`remoteManagerAuthorityRequest(state, \"cli\", \"retire\", { registrationProof, retirement })`.\nCompute the proof as `remoteManagerRegistrationProof(owner, state)`, adding the contract artifacts\nas a third argument for activation only. A host that recomputes the proof from a received request\npasses `{ space, instanceId, lifecycleUid: managerLifecycleUid, identities }` from that request.\n\n## Bearer validator lifetime cap in 0.70.0\n\n`validateUserToken` from `@cotal-ai/auth` no longer takes `maxTtlSec`. It caps a bearer's lifetime\nat the cap of the bearer's view, the same cap the issuer applies when it mints: 900 seconds, or 300\nfor a `transfer-writer` bearer. The auth callout never passed the option, so a running mesh behaves\nas before. Only code of your own that calls the validator with `maxTtlSec` is affected.\n\n### What stops working\n\nA call that passes `maxTtlSec` in an object literal no longer compiles. Plain JavaScript that keeps\nit still runs, and the value is ignored. A `NaN` value, such as `Number()` of an unset environment\nvariable, used to turn the lifetime check off and accept a bearer of any lifetime. That bearer is\nnow refused at its view's cap.\n\n### Before the upgrade\n\nRemove `maxTtlSec` from each call. A test that needs a bearer to expire sooner mints one with a\nshorter lifetime.\n\n## Persisted identity records in 0.70.0\n\nThe manager instance identity, the manager sibling identities, the auth plane instance identity and\na participant manager's remote authority state now share one reader and one first mint in\n`@cotal-ai/workspace`, exported as `claimIdentityRecord` with the nkey check `identityOf`. Each\nrecord is read as a regular file, must hold non-empty nkeys and is created exclusively, so\nconcurrent first starts of a participant manager on one root now settle on one identity where each\nused to keep its own. `saveManagerInstanceIdentity` and `saveAuthInstanceIdentity` are gone. A\nrunning mesh whose records are plain files needs nothing.\n\n### What stops working\n\nA manager instance, auth instance or remote authority record that is a symlink, a directory or any\nother non-regular entry is refused where it used to be followed. The manager, the auth plane and a\nparticipant manager fail to start on it, and `cotal reconcile-gate` and `cotal deregister-instance`\nrefuse it. Retirement already refused it. A remote authority record with an empty nkey id or seed\nis refused too. A first mint that loses its race and cannot read the winner now refuses with\n`identity-record-create-lost` in place of `manager-instance-identity-create-lost` or\n`auth-instance-identity-create-lost`. Code that imports either `save` function no longer compiles.\n\n### Before the upgrade\n\nReplace a symlinked identity record with a copy of the file it points to. Code that wrote a record\nwith a `save` function plants it with `createManagerInstanceIdentity` or\n`createAuthInstanceIdentity`, which create the record when it is absent and otherwise return the\nstored one unchanged. Nothing replaces an overwrite of a stored identity.\n\n## Manager instance in user credentials in 0.70.0\n\n`AuthProvider.userCredentials` from `@cotal-ai/core` no longer returns `managerInstanceId`. A\n`manager-caller` credential's manager instance is the signed `act.managerInstanceId` claim in its\nbearer, which the broker verifies and the CLI already used. The reference provider in\n`@cotal-ai/auth` stops copying the exchange response's field into its result, where nothing\ncompared it with the bearer. The exchange still answers with the field, so a running mesh behaves as\nbefore.\n\n### What stops working\n\nCode of your own that reads `managerInstanceId` from a `userCredentials` result no longer compiles,\nand plain JavaScript reads `undefined` there.\n\n### Before the upgrade\n\nRead the instance from the bearer's `act.managerInstanceId` claim.\n\n## Auth plane identity location in 0.70.0\n\nThe user-auth service keeps its instance identity in the root's `.cotal/space.<hex>/auth-instance.json`,\nbeside the manager's. It used to sit inside `.cotal/auth`, at\n`space.<hex>/.cotal/auth/auth-instance.<hex>.json`, so a copy of that folder carried it. The first\nstart of an upgraded root moves the record and keeps the instance. A hosted context started through\n`startAuthService` has its record moved the same way inside its `stateDir`.\n\n### What stops working\n\nCode that calls `openAuthAuthorityPlane` without the new `identityRoot` option no longer compiles. A\nstart that finds a record both in `.cotal/space.<hex>/` and at its older place refuses and names the\ntwo files. A start also refuses when the older place of the auth or manager identity holds a symlink,\na directory or anything else that is not a regular file. The manager used to skip a dangling symlink\nthere and mint a new identity.\n\n### Before the upgrade\n\nPass `identityRoot` to `openAuthAuthorityPlane`. When `dir` is a workspace root's user-auth state\ndir, `<root>/.cotal/auth/space.<hex>`, pass that root. A plane with no workspace root, as\n`startAuthService` runs, passes `dir` itself. Either keeps the identity the plane already has: on the\nfirst start it moves from `<dir>/.cotal/auth/` to `<identityRoot>/.cotal/space.<hex>/`. Never pass a\ndirectory inside `.cotal/auth`: the record would land in the folder an operator copies and travel\nwith it again.\n\nA copy of `.cotal/auth` taken from a root last run by an older Cotal carries that root's record.\nDelete `.cotal/auth/space.<hex>/.cotal/auth/auth-instance.<hex>.json` from the root you copied it to\nbefore the first `cotal up --user-auth` there.\n\n## Per-seat `COTAL_` names in `spawn.env` in 0.71.0\n\n`spawn.env` in the cotal config no longer forwards a `COTAL_` name the launcher sets for each seat,\nsuch as `COTAL_ROLE`, `COTAL_MODEL` or `COTAL_SUBSCRIBE`. Before, a seat launched with no value of\nits own took the spawning process's value and ran under that role, model or read set. The\nmachine-wide knobs a seat already receives, such as `COTAL_HOME`, may still be listed.\n\n### What stops working\n\nEvery spawn and resume under a config whose `spawn.env` lists such a name is refused before\nlaunch, and the refusal names the entry. Code that calls `launchEnv` from `@cotal-ai/connector-core`\nwith such a name in `envAllow` gets the same error.\n\n### Before the upgrade\n\nRemove those names from `spawn.env`. Give each seat its role, model and channels with `--role`,\n`--model` and `--subscribe`, or in its persona's `role:`, `model:` and `subscribe:`.\n\n## Role addresses in 0.71.0\n\nA role must be one `[A-Za-z0-9_-]` token. Before 0.71.0 any other spelling was rewritten into one:\n` probe ` reached the `probe` queue and `pro.be` reached `pro_be`, while the message kept the\nspelling sent. An anycast to `*` was accepted and stored where no holder reads it.\n\n### What stops working\n\nAn agent whose role is outside the token set no longer starts, however it is launched:\n`cotal join --role`, `cotal spawn --role`, an agent file's `role:`, `COTAL_ROLE` and an embedded\nendpoint's `card.role` are all refused before the agent joins.\n\nA send to such a role, or to `*`, through `cotal send ask`, `/anycast` or `cotal_anycast` is refused,\nand nothing is stored.\n\n`routeToken` is no longer exported from `@cotal-ai/core`. A role routes as spelled, so code that\nused it to name a role's queue uses the role itself, and `assertValidRole` checks one.\n\n### What migrates on its own\n\nEvery task queue. A `svc_<role>` durable was always named from the rewritten token, so its pending\nrequests and its holders carry over.\n\n### Before the upgrade\n\nRename each role outside the token set to the token it already routed to: remove the surrounding\nspaces and replace every other character outside the set with `_`. Rename it where the holder is\nlaunched and in every script or prompt that sends to it.\n\n## Carrying a resumed Claude session to another host in 0.67.0\n\n`cotal spawn --resume <id> --detach --on <instance>` now carries a Claude session held on the\noperator's host to the target manager instance. Both sides need this release: an older manager does\nnot serve `transcript-receive`, and the CLI then stops with that manager's refusal instead of\nlaunching. The manager cluster document moves to revision 21, and the `ps` row's `resume` object\ngains `host` and `transferredAt`.\n\nA manager host that runs carried seats needs `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_AUTH_TOKEN` or a\ncloud provider selection in its environment, because each carried seat runs in its own Claude home\nwith no stored login. On an authenticated mesh the CLI mints the transfer writer from the space's\nsigning seed, so the carrying host needs that seed, as for any other operator command. On a user-auth\nmesh it exchanges the operator's login for a `transfer-writer` view instead, so the operator's grant\nneeds scope `admin`, and the auth service must run this release. A remote manager receives a carry once\nits host serves the manager-service `transferReader` operation. A seat launched without carrying,\nincluding any `--resume` whose id this host does not hold, is unchanged.\n\n## Lifecycle head type in 0.67.0\n\n`LifecycleMapping`, the type `parseLifecycleHead` returns, is now a union on `state`. Nothing about\na running mesh changes: heads that parsed before parse the same way, and the refusals are\nunchanged. Only TypeScript code that compiles against `@cotal-ai/core` is affected.\n\n### What stops working\n\nAn `interface` that extends `LifecycleMapping` fails with TS2312, because an interface cannot extend\na union. Code that builds a head in memory no longer compiles when the head is `retiring` without\nits `op`, or `active` or `retired` with one. The parser already refused those heads.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type ActiveMapping = LifecycleMapping & { state: \"active\" }`. A reader that has checked\n`state === \"retiring\"` reads `op` without a guard.\n\n## Issuance gate types in 0.67.0\n\n`EpGateRow` and `EndpointGateRow`, which `parseIssuanceGate` and `parseEndpointGate` return, and\n`EpGateState`, which an `EpIssuanceGate` or `EpIssuanceBarrier` returns from `observe`, are now\nunions on `state`. Nothing about a running mesh changes: gates that parsed before parse the same\nway, and the refusals are unchanged. Only TypeScript code that compiles against `@cotal-ai/core`\nis affected.\n\n### What stops working\n\nAn `interface` that extends one of these types fails with TS2312, because an interface cannot\nextend a union. Code that builds a gate in memory, such as a custom barrier's `observe`, no longer\ncompiles when the gate is `frozen` or `retired` without its `op`, or `open` with one. The gate\nparsers already refused those rows.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type CustomGateRow = EpGateRow & { custom: string }`.\n\n## Lifecycle-blocked refusals in 0.66.0\n\nA refusal that carries `ai.cotal.ep.lifecycle-blocked` now reports only the lifecycle state it\nread. Nothing about a running mesh changes. A client that branches on the detail must read the new\nfield.\n\n### What stops working\n\nA refusal raised at the issuance gate used to carry `headState` without reading the head:\n`retiring` for a frozen gate and `retired` for a retired one. It now carries `gateState`\n(`frozen` or `retired`) and no `headState`. A client that treats `headState: \"retired\"` as a\nburned uid, or `headState: \"retiring\"` as a retirement in flight, no longer matches those\nrefusals, and the `[lifecycle ...]` suffix on the error string changes the same way. A custom\nissuance barrier whose `observe` returns a frozen gate without a valid `op` (a string `opId` and\none of the four op kinds) is now refused as `internal` by `registerServiceInstance`.\n\n### Before the upgrade\n\nUpdate such a client to read `gateState` for a gate refusal and `blockedOp` for the operation that\nholds the gate. `headState` is present only when the refusal read the head, for example an\nactivation refused because the head is still retiring.\n\n## Workflow programs that bind `once` in 0.65.0\n\n`once` is now a scope of the workflow language, so it is a reserved name. A program that declares\nits own `once` binding (`const once = ...`, a parameter or a function named `once`) is refused at\nvalidation with L2002. Nothing else about a running mesh changes.\n\n### What stops working\n\nA run whose recorded program binds `once` cannot be resumed after the upgrade, because a resume\nvalidates the recorded program again. A new `cotal run start` of such a program is refused before\nanything is recorded.\n\n### Before the upgrade\n\nList the runs with `cotal run ps` and check each program that is still running or held for a\nbinding named `once`. Let those runs finish on the old version before you upgrade the manager, and\nrename the binding in the program before you start it again.\n\n## From 0.58.0 to 0.59.0\n\nEvery connector now publishes a failed run's `RUN_ERROR` on `events.<owner>.<actor>` with the fixed\nmessage `run failed` and no `code` or `rawEvent`. The error text and error kind a harness reports\ncan echo a prompt, a peer message or tool output, and that channel has a different read ACL. A\nreader that showed the message or branched on `code` gets neither after the upgrade. Where a\nconnector reports the error kind as the agent's presence condition, that is unchanged.\n\n### Settle pending event frames before the upgrade\n\nEach session's events are frozen in its event write-ahead log before they are published. A session\nrestarted on 0.59.0 whose log still holds an unacknowledged frame with an older `RUN_ERROR` does not\nrepublish it: its event emitter halts with `egress-run-error` and publishes nothing further for that\nsession. The broker may or may not already hold that frame, so the halt cannot settle it.\n\n1. Stop the seats cleanly on 0.58.0, with the broker still up.\n2. List the logs that still hold a pending frame. The logs live under the events state root\n (`COTAL_WORKSPACE_ROOT`). Empty output means there is nothing to settle.\n\n ```sh\n find \"$COTAL_WORKSPACE_ROOT/.cotal/events\" -name wal.json \\\n -exec jq -r 'select(.pending != null) | input_filename' {} +\n ```\n\n3. For each session listed, start it again on 0.58.0 while the broker is reachable, let it recover,\n stop it, and run step 2 again. Recovery publishes the frame as 0.58.0 would have, error text\n included, so it only finishes what 0.58.0 had already started.\n\n If that start halts with `cas-loss` instead, the agent's subject is no longer at the sequence this\n log expects, and no restart settles that log, on 0.58.0 or later. A lost acknowledgement is one\n cause: the broker stored the frame, so it and its error text are already on the channel, and every\n retry halts the same way because the stream checks the frozen expectation before it deduplicates.\n The halt message names the other causes, such as a second emitter for the same agent under a\n different state root, a restored stream or frontier record, or a purged channel. With those the\n pending frame may never have reached the broker, so a `cas-loss` does not tell you whether it\n landed. Find and stop any second writer and rule out a restored state first. Clearing the halt\n then means purging the agent's event channel and removing the agent's directory under the events\n state root whole (see [Event plane](connect-claude.md#event-plane)). That abandons the pending\n frame whether or not the broker has it, and the purge also drops the earlier frames of every\n session of that agent.\n4. Upgrade once step 2 prints nothing.\n\nIf a session halts with `egress-run-error` after the upgrade, go back to step 3 for that session on\n0.58.0. Do not edit or delete `wal.json` on its own to get past either halt: clearing the pending\nframe abandons that epoch, an event the broker never received is lost, and removing part of the\ndirectory leaves a state the next start refuses.\n\n## Explicit actor grants in 0.59.0\n\n`cotal actor grant` no longer fills an omitted ACL flag with its wide default. A grant names\n`--scope`, `--allow-subscribe` and `--allow-publish`, or passes `--full` to give the ones it leaves\noff their wide defaults (`spawn,role:default`, `>` read, `>` post). Any other grant is refused. The\nbreak is in the CLI on the machine that holds the actor ledger, the one that ran\n`cotal up --user-auth --idp <url>`. No stored row, credential or wire message changes.\n\n### What keeps working\n\nExisting actor ledger rows keep the authority they were granted, and their users and agents connect\nas before. `actor revoke`, `actor list` and a `grant` that names all three ACL flags behave as they\ndid on 0.58.0. Nothing on disk is converted.\n\n### What stops working\n\nA grant that leaves off any of the three flags without `--full` exits 1 with\n`refusing to grant \"<actor>\" with --scope, --allow-subscribe, --allow-publish left off`, naming the\nflags it is missing, and then prints both accepted forms. It writes no row and does not retire the\nactor's current lifecycle. An existing row stays as it was, and an actor granted for the first time\nstays out until the grant is run again. This includes the bare grant printed on 0.58.0 by\n`cotal login`, `cotal status`, `actor list` and the not-granted refusal. Look for it in provisioning\nscripts, onboarding runbooks and anything that pastes those hints.\n\n### Upgrade order\n\nChange the scripts before the ledger machine is upgraded, and make each grant name all three flags.\n0.58.0 and 0.59.0 both accept that form. To keep a wide row, write its defaults out:\n\n```sh\ncotal actor grant <actor> --sub <IdP subject> \\\n --scope spawn,role:default --allow-subscribe '>' --allow-publish '>'\n```\n\nSwitch to `--full` only once the ledger machine runs 0.59.0. 0.58.0 refuses it with\n`Unknown option '--full'` before it reads the ledger. Brokers, managers and participant machines\nneed nothing for this break, so their order is the one the section above gives.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused grant changes nothing. The\nexposure is a grant script that runs against 0.59.0 before it was changed: it fails and grants\nnothing.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. On the ledger machine, save the output\nof `cotal actor list` to compare rows after the changed scripts run, and list the scripts that call\n`cotal actor grant`.\n\n### The upgrade end to end\n\n```sh\n# on the ledger machine, still on 0.58.0\ncotal actor list > actors-before.txt\ngrep -rn 'actor grant' <your provisioning scripts>\n# make every grant name --scope, --allow-subscribe and --allow-publish, run them, then upgrade\nnpm i -g cotal-ai@0.59.0\ncotal actor list | diff actors-before.txt -\n```\n\nBoth refusals quoted here were run on 0.58.0 and on the 0.59.0 code. That brokers, managers and stored\nrows need nothing is read from the change, which touches only the CLI and its hints, and was not run\non a live split deployment.\n\n## Repeated flags refused in 0.59.0\n\nA `cotal` flag given more than once is now a usage error unless the command declares it repeatable.\nOn 0.58.0 the last value won with no message, so `cotal down web --space a --space b` acted on `b`\nwhile a wrapper that checked the first `--space` verified `a`. The break is in the command-line\nparser on the machine that runs the command, including commands added with `cotal ext add`. No\nstored state, credential or wire message changes.\n\n### What keeps working\n\nA command line that gives each flag once parses as it did on 0.58.0, in any order and in the\n`--flag=value` form. Flags whose help says repeatable, such as `--opt` and `down --session-store`,\nstill collect every value. A flag-shaped word after `--` is still a positional. The daemons, units\nand agents that `cotal` starts for itself are given each flag once, so a fleet driven only by `cotal`\ncommands typed by hand needs no action.\n\n### What stops working\n\nA command line that repeats any other flag exits 1 before the command runs. It prints\n`Option '--space' cannot be repeated`, or `Option '-f, --file' cannot be repeated` for a flag with a\nshort form, followed by the command's help. `-f` and `--file` count as the same flag. Look for it in\nscripts, aliases and wrappers that append a flag to override one set earlier, such as a fixed\n`--space` followed by `\"$@\"`.\n\n### Upgrade order\n\nChange those scripts first so each flag is given once. 0.58.0 and 0.59.0 both accept that form.\nBrokers, managers and participant machines need nothing for this break, and each machine's CLI\napplies it when that machine is upgraded, so their order is the one the sections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused command does nothing. The\nexposure is a script that still repeats a flag when it runs on 0.59.0: it exits 1 instead of acting on\nthe last value.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the scripts, aliases and wrappers\nthat call `cotal` so each one can be checked.\n\n### The upgrade end to end\n\n```sh\n# still on 0.58.0\ngrep -rn 'cotal ' <your scripts and wrappers>\n# give each non-repeatable flag once, then upgrade\nnpm i -g cotal-ai@0.59.0\n# run each changed script; a repeat left behind exits 1 with the usage error and does nothing\n```\n\nThe refusal and its messages were run against the 0.59.0 parser and `cotal topology view`. That the\nargument lists `cotal` builds for its own processes give each flag once is read from the code, and\nwas not run on a live split deployment.\n\n## Detached spawns from a seat's shell in 0.62.0\n\nOn a static or open mesh, `cotal spawn --detach` run inside a managed seat's shell now launches as\nthat seat. On 0.61.0 it minted a one-shot operator instrument, so the manager recorded that\ninstrument as the spawner and the seat's own `cotal_despawn` of the child was refused with\n`not authorized: <seat> was not spawned by <caller> (admin tier required)`. The break is in the CLI\non the machine where the seats run. No stored state, credential or wire message changes.\n\n### What keeps working\n\n`cotal spawn --detach` from an operator terminal or from a script outside any seat launches as\nbefore, and so does any call with `--creds`, one aimed at a space other than the seat's own, or a raw\nopen target named with `--server` and an unregistered `--space`. A user-auth mesh is unchanged. A seat with\n`capabilities: [spawn]` still spawns from its shell, and can now stop that child with\n`cotal_despawn`. `--on <instance>` from a seat's shell still lands on that manager instance, now as\nthe seat.\n\n### What stops working\n\n- On a static mesh, a seat without `capabilities: [spawn]` can no longer spawn from its shell. Its\n own credential holds no spawn subject, so the broker refuses the request and the command exits 1.\n- A child launched from a seat's shell is now that seat's child, so the manager stops it when the\n seat exits, as it does for a `cotal_spawn` child. A child that has to outlive the seat that\n started it now goes with the seat.\n- A seat launched without `COTAL_SPACE` is placed by its static credential. Every connector sets\n that variable, so this only reaches a hand-built launch: from such a seat's shell, a spawn aimed at\n a static space that holds no credential for the seat is refused instead of running as the operator.\n\n### Upgrade order\n\nOnly the CLI that seats run from their shell changes, which is the one installed on the host where\nthe seats run. Brokers and managers need nothing for this break, so their order is the one the\nsections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it. A child already running when you upgrade\nkeeps the spawner the manager recorded at its launch.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the agent files whose seats run\n`cotal spawn --detach` from their shell, note which of them lack `capabilities: [spawn]`, and note\nwhich of their children must outlive the seat.\n\n### The upgrade end to end\n\n```sh\n# still on 0.61.0: find the seats that spawn from their shell\ngrep -rln 'cotal spawn' .cotal/agents\n# add `capabilities: [spawn]` to each of those agent files that lacks it, and launch any child\n# that must outlive its seat from an operator terminal instead\nnpm i -g cotal-ai@0.62.0\n```\n\nThe attribution, the despawn, the refusal of a seat without `spawn`, the stop on seat exit and a\nseat's `--on` spawn were run on a local static mesh, and the attribution and the despawn on a local\nopen mesh.\n\n## From 0.53.0 to 0.54.0\n\nManager calls now borrow an instance-bound `manager-caller` credential. Followed mutations require\n`manager.goal-result` on the selected manager, so a compatible issuer, manager and client must be\nloaded together. An older manager is refused before a followed mutation; upgrading an installed\nbinary alone does not replace code in a running manager, connector or embedded client.\n\n### Preserve state before changing processes\n\nSnapshot the broker's durable storage using its supported backup procedure, the host authority and\nactor ledgers, and each participant's manager identity, runtime custody records, credentials and\nsaved sessions. Include the embedding application's database and configuration under its supported\nbackup procedure. Record the loaded package versions and the CLI path used by bearer helpers.\nKeep these copies private. Do not change the IdP issuer, regenerate manager identities, rotate agent\ncredentials or recreate tenant storage to make the upgrade pass.\n\nNo ledger, goal-history or session conversion is required for this change. Existing ordinary\nmessaging credentials retain their normal expiry rules. New manager-caller credentials are obtained\non demand from the current grant; old manager-call credentials do not gain the new view automatically.\nExisting accepted goals remain durable and must not be submitted again merely because observation\nwas interrupted. Fresh remote registration publishes its service status at the current revision and\nepoch; do not seed that status manually.\n\n### Upgrade the split deployment\n\n1. Stage one pinned 0.54.0 package set for the host and participants, including the embedding SDKs.\n Pause new manager mutations and let accepted work settle where possible before reloading processes.\n2. Upgrade the host issuer and embedding first. Keep the broker, its account identities and durable\n storage in place. Then load the matching manager release on each participating machine.\n3. Preserve active seats through the runtime's supported update path. A Linux custodial runtime may\n release and re-adopt seats within its 600-second unattended window; verify the actual runtime,\n custody records and process identities before relying on it. A legacy PTY runtime without release\n support cannot preserve active seats through a generic manager restart. Drain it at an approved\n idle window instead of signalling the manager or replacing conversations.\n4. Reload the clients and connectors through their session-preserving host controls. Refresh any\n bearer helper captured from an older immutable CLI path. A transport-only reconnect does not reload\n JavaScript. Verify authenticated instance selection, a read-only manager command and canonical\n result recovery before allowing new followed mutations.\n\nTreat the interval from issuer reload through compatible manager/client reload as a manager-control\noutage. Mixed versions can refuse discovery or commands; there is no promised rolling transition.\nOrdinary agent sessions survive only where their runtime and credentials permit it. If verification\nfails, keep mutations paused and repair forward from the preserved state rather than resetting it.\nThis release does not add host-backed enrollment or terminal release for stock participant detached\nagents; see [Remote supervised agents](run-a-mesh.md#remote-supervised-agents).\n\n## From 0.48.2 to 0.49.0\n\n0.49.0 changes how a credential's authority is recorded. A credential is no longer only a signed\nfile: it is an *issuance*, with a generation the issuer chose and durable evidence of the ceiling it\nwas granted under. The important consequence for a running deployment is not at connect time. It is\nat renewal time.\n\n### What keeps working without any action\n\n- **Existing agent credentials keep authenticating.** A credential minted under 0.48.2 is not\n revoked and is not rejected at connect. Nothing needs to be re-issued to bring the fleet back up\n after the upgrade.\n- **The channel registry survives.** Channels, their replay settings, descriptions, and usage text\n are ordinary durable state and are not rewritten by the upgrade.\n- **`cotal deliver` is still a standalone command.** Running the delivery daemon as its own process\n remains supported; it is not restricted to being a child of `cotal up`.\n- **`cotal join` keeps its flags.** In particular `--lifecycle-uid` is not new in 0.49.0. It has\n been required alongside `--creds` since well before this release, and the pairing rule did not\n change here. A scripted external join that worked under 0.48.2 works unchanged.\n\n### What does not migrate\n\n**A credential minted before 0.49.0 cannot be renewed.** Managed agent credentials carry a\n24-hour lifetime and the manager re-signs one once it passes **75%** of its life, ticking every\nquarter of the TTL so a tick always lands inside that window. When the manager reaches a credential\nthat carries no issuance, it refuses to renew it and logs the agent by name:\n\n```\n! managed cred renewal <agent>: renewManagedStaticCred: <agent> carries no issuance;\n a static credential minted before SPEC 13.15 is not renewed under an unbound generation\n - respawn the agent\n - the agent dies loud at this cred's expiry unless it is reminted\n```\n\nSo the fleet comes up fine, runs normally, and then each agent stops at its own credential's\nexpiry, within roughly a day of the upgrade, one at a time rather than together. The refusal is\ndeliberate: the renewal would otherwise have to invent a generation nobody issued, which is the\nstate the release exists to remove.\n\n**Respawn the managed agents as the last step of the upgrade.** For this particular upgrade the\nrespawn is not optional: stopping a 0.48.2 manager ends its agent processes whichever CLI you use,\nfor the reason given under the outage window below. The respawn is how they come back, and it is\nalso what mints each credential as an issuance so it renews from then on. One planned pass over the\nfleet is the whole job. Skipping it leaves agents stopped and, for any credential that survived\ninto 0.49.0 unminted, brings the renewal cliff above a day later, one agent at a time.\n\n### Credentials you minted yourself\n\n**A credential you minted with `cotal mint` is a different case, and it very likely needs\nnothing.** The distinction that matters is not the word \"static\", which covers both. It is **what\nminted the credential and who owns its renewal**. A credential the **manager** minted for an agent\nit spawned carries a lifetime and is renewed by the manager, so it is the subject of everything\nabove. A credential **you** minted with `cotal mint` and handed to an external peer is issued with\n**no expiry at all**, and no manager renews it: it is not in the sweep, so there is no renewal to\nfail. It keeps working after the upgrade, and re-minting it would mean coordinating with a third\nparty for no gain.\n\nThe manager says which one it is holding. Where a credential has no expiry to reach, the sweep\nnames it and moves on rather than refusing:\n\n```\n! managed cred renewal <agent>: credential is unbounded - not renewed\n (a pre-TTL credential stays as minted until respawn)\n```\n\nRe-mint an external peer's credential only if you want it to carry a lifetime, and at a time you\nchoose.\n\n### How to read the boot log\n\nA 0.49.0 manager starting over an existing space may print lines like:\n\n```\n verified evicted: <holder-key> (3/12)\n already verified (durable): <holder-key>\n✓ boot self-heal: manager/<id> registration gate reopened at generation <n>\n```\n\nThese are **not** a credential migration, and reading them as one is the most likely way to\nconclude the fleet is fine when it is not. They come from the manager repairing **one** endpoint\nregistration gate that a previous restart left frozen, and they enumerate that single gate's\ncredential-family holders as it verifies each one evicted. `already verified (durable)` on a later\nstart is the repair cursor resuming, not a credential that became durable. The repair is real and\nuseful (it is what previously needed `cotal reconcile-gate` by hand), but it says nothing about\nwhether your agent credentials carry issuances. The renewal refusal above is the signal that does.\n\n### Which side to upgrade first in a split topology\n\nMove the manager first.\n\nThe stores 0.49.0 introduces are created by the **manager** at its own boot, not by the broker.\nThey are create-or-verify and idempotent, so a 0.49.0 manager brings the space's authority stores\nup to the new shape itself, and it does so against whichever broker is answering.\n\nBeing honest about the evidence behind each direction, because they are not equally established:\n\n- **Broker-first was measured on a live 30-agent deployment** (issue #1578). Upgrading the broker\n first locks the old manager out immediately: `cotal up` re-renders the broker's generated config\n from the trust record, and after the restart the still-0.48.2 manager is refused on every\n connection with an `authentication error` naming the Nkey, continuously. That text comes from the\n broker process, not from a Cotal command, so match on its shape rather than on an exact string.\n `cotal ps` reports zero agents while\n the agent processes are still alive, because the manager has lost its view of them, not because\n they died. Upgrading the manager clears it immediately.\n- **Manager-first is reasoned from where the new stores are provisioned**, not from a measured\n fleet upgrade. It is the recommended order because the manager is the component that creates what\n 0.49.0 adds, but it has not been run end to end on a production split topology at the time of\n writing. Treat it as the better-supported order rather than a guaranteed one, and keep the\n rollback below ready either way.\n\nWhichever order you pick, **this is not a rolling upgrade**. Between the two steps the mesh is down\nand the manager cannot see its agents. Go straight through rather than pausing between them, and\nschedule it as an outage window.\n\n### What the window looks like\n\n- **The managed agent processes do not survive step 1, in either order.** This is the one place\n where the obvious reordering does not rescue you, so it is worth understanding rather than\n working around. Sparing agents on a bare manager stop is a **handshake**: a 0.49.0 manager\n publishes a capability file proving it can release its agents, and a 0.49.0 `cotal down` refuses\n the stop unless it finds one. **A 0.48.2 manager never publishes that file**, because the\n mechanism ships in the release you are installing. So the old CLI against the old manager sends a\n plain stop and takes every seat with it, and the new CLI against the old manager either refuses\n (leaving `--with-agents`, which reaps deliberately) or falls to the legacy path, warns that it\n cannot verify the manager can spare its agents, and signals it anyway.\n- **You can confirm which side you are on in one command, without stopping anything.** The flag that\n marks the newer behaviour is absent from the older CLI, and its summary line makes the difference\n plain:\n\n ```\n $ cotal down --help # on 0.48.2\n cotal down - stop the whole local stack, or name only the components to stop\n\n $ cotal down --help # on 0.49.0\n cotal down - stop the whole local stack (managed agents stay running unless --with-agents), ...\n ```\n\n If your `cotal down --help` does not mention `--with-agents`, stopping the manager stops the\n agents with it.\n- **Therefore the respawn in step 5 is mandatory recovery for this upgrade, not an optional pass.**\n It is also the step that re-mints credentials as issuances, so it is the same action either way.\n Plan the window to include it rather than treating it as cleanup.\n- The **manager's view** of them is lost while the two sides disagree, so `cotal ps` reports zero\n and control commands do not reach seats.\n- **Messages are not delivered** while the mesh is down.\n- The window is as long as it takes to restart the second component, plus the manager's own start.\n It is minutes, not hours, provided you do not stop between the steps.\n- **Nothing self-heals if you stop halfway.** The refusal is continuous until both sides match.\n\n### Snapshot this before you start\n\nTake these while the deployment is still on 0.48.2. The two `cotal` reads are live reads and must\nhappen before anything stops.\n\n- **A filesystem or volume snapshot of both containers**, if your platform offers one. This is the\n only rollback that covers every case, and it is what the reporting deployment used.\n- **`cotal backup create <dir>`**, for the durable space state, **but read the next paragraph before\n you rely on it**: on a split broker and manager topology it is very likely unavailable to you, and\n the volume snapshot above is your actual rollback.\n- **The trust records and credential directory** under `.cotal/auth` on the manager host, including\n the per-space material directory. These are what a re-mint would otherwise have to replace.\n- **A copy of the channel registry**, so you can verify it came back rather than assuming it did:\n `cotal channels list` before and after.\n- **The output of `cotal ps`**, so you know how many seats you expect to see afterwards and can tell\n a lost view from a lost agent.\n\n#### `cotal backup` on a split topology\n\n**`cotal backup create` cannot read a running stack.** It requires a completed cut, and only\n`cotal down --preserve-state` publishes one:\n\n```\n$ cotal backup create ./backup.0482\n✗ backup requires a completed cut; run `cotal down --preserve-state` first\n```\n\n**And `cotal down --preserve-state` requires a manager alive on the host you run it from.** It uses\nthat manager to attest that every retained child stopped, and the check is deliberately fail-closed:\na manager that is dead or merely uncertain refuses rather than preserving an unproven cut. The check\nreads a local pidfile, so a **remote** manager does not satisfy it. On a split topology the broker\nhost has no local manager, which means the documented durable-backup path is not available there.\n\n**Measured rather than assumed, at 0.48.2**: the backup refusal above is executed output. The\npreservation requirement is read from `down.ts` at the same tag, where the preserve path asks a\nmanager to prepare an inventory and then requires that manager to be locally alive before it\ncommits. The part not executed end to end is a genuine two-host split, which needs two real hosts.\n\n**What to do instead.** Use the filesystem or volume snapshot of both containers. That is the\nrollback the reporting deployment actually used, it covers the broker's durable state and the\nmanager's credential material together, and it does not depend on either component being able to\nattest for the other. If you want `cotal backup` as well, take it from a host that does have a live\nlocal manager, and understand it is a second copy rather than the primary rollback.\n\n**This looks like a product limitation rather than a documentation gap**, and it is written here as\none so an operator is not left thinking they mis-typed a command. The upgrade path for the exact\ntopology this page is addressed to cannot use the documented backup command.\n\n### The upgrade end to end\n\n```bash\n# 0. on 0.48.2, STILL RUNNING: record what you expect to see afterwards.\n# These two are live reads, so they must happen before anything stops.\ncotal channels list > channels.before\ncotal ps > ps.before\n\n# 1. manager host. READ THE NOTE BELOW THE BLOCK FIRST: this step ends the\n# managed agent processes whichever order you choose, and the respawn in\n# step 5 is how they come back. It is recovery, not tidying.\n#\n# STOP THE MANAGER WITH THE 0.48.2 CLI, BEFORE INSTALLING 0.49.0. The\n# order matters and it is not recoverable once you install: a 0.49.0\n# `down manager` REFUSES to stop a 0.48.2 manager whose pid record carries\n# a start token, which is every manager on a platform that can read one\n# (Linux can):\n# refusing bare manager stop: ... does not prove this manager can detach\n# its agents; use --with-agents or stop the agents explicitly\n# The refusal names two remedies and NEITHER clears it for this case. The\n# check reads a capability file that only a 0.49.0 manager writes; it never\n# counts agents, so stopping them first changes nothing. And `--with-agents`\n# is whole-stack only, so `down manager --with-agents` is refused by its own\n# flag rule. See #1592.\ncotal down manager # the 0.48.2 CLI, still installed.\n # 0.48.2 has no --with-agents; this\n # is the whole route. On a host that\n # runs the whole stack, the 0.49.0\n # `cotal down --with-agents` after\n # installing is the alternative.\nnpm install -g cotal-ai@0.49.0 # ONLY after the stop above\n# `supervise` RUNS IN THE FOREGROUND and holds the terminal until you stop\n# it. There is no --detach on this command. Start it under whatever keeps\n# your manager alive normally (systemd unit, container entrypoint, or a\n# second terminal), and run the remaining steps from another shell.\ncotal supervise --space <space> --server nats://<broker>:4222\n\n# 2. broker host: stop the stack.\n# NOT `--preserve-state` on a split topology: it needs a manager alive on\n# THIS host to attest its children stopped, and yours is on the other one.\n# Your rollback is the volume snapshot from \"Snapshot this before you\n# start\", not `cotal backup`.\n# See \"cotal backup on a split topology\" above.\ncotal down\n\n# 3. broker host: install 0.49.0 and start it again\nnpm install -g cotal-ai@0.49.0\n# Record the manager log's size BEFORE starting, so step 3a can tell THIS\n# boot's output from every earlier one. It must be captured here, ahead of\n# the start: taken afterwards it sits past the new line and the wait hangs.\n# `<spaceKey>` is NOT the space name. It is lowercase hex of the name's\n# UTF-8 bytes, so space `prod` is `manager.70726f64.log`. Do not guess it:\n# `cotal up` prints the real path on its launch line. Substituting the\n# plain name points at a file that does not exist, and the wait below then\n# burns its full timeout before telling you.\nLOG=.cotal/manager.<spaceKey>.log\nOFF=$( [ -f \"$LOG\" ] && wc -c < \"$LOG\" || echo 0 )\ncotal up --detach --host 0.0.0.0 --space <space> --no-manager\n\n# 3a. SPLIT TOPOLOGY ONLY: `--no-manager` above boots the broker (and the\n# delivery daemon) with NO local manager on the broker host, so there is\n# no wait-and-stop step on a current cotal-ai. The rest of this step is\n# the OLDER-host recipe, kept because the flag is refused there and that\n# refusal is your signal you are on it: without the flag the `up` also\n# starts a local manager, and you must wait for the log to show it is up,\n# then stop it, or you finish the upgrade with two managers and the one\n# you did not intend is the one nobody is watching.\n# A bare `grep -q` does NOT wait: it reads once and exits 1 immediately\n# if the line has not been written yet. Bound the wait instead, so a\n# manager that never comes up fails loudly rather than reading as ready.\n# The log is opened APPEND-ONLY, so on any host that has run a manager\n# before, this file ALREADY carries a `manager up` line from an earlier\n# boot. Grepping the whole file therefore matches instantly and waits for\n# nothing. Read only what THIS boot appended, using the $OFF captured in\n# step 3 above (before the start, which is the only point it is correct):\ntimeout 60 bash -c \\\n \"until tail -c +$((OFF+1)) \\\"$LOG\\\" | grep -q '. manager up'; do sleep 1; done\"\n# exit 0 = THIS boot logged it; exit 124 = it never did, so STOP and look.\n# This manager is 0.49.0 and publishes its own spare-capability file, so\n# the bare stop below is NOT the refusal case from step 1.\ncotal down manager # broker + delivery remain\n# On a current cotal-ai the two commands above are unnecessary (nothing\n# to wait for, nothing to stop) and `cotal down manager` simply reports\n# no manager to stop.\n\n# 4. verify the mesh is whole again before touching the fleet.\n# Do NOT compare `cotal ps` against ps.before yet: step 1 ended the agent\n# processes, so at this point it is EXPECTED to be empty, and an empty\n# `ps` is also the signature of the broker/manager mismatch described\n# above. The two are indistinguishable here, so compare what the mesh\n# itself should have carried across instead:\ncotal channels list # compare against channels.before: this SHOULD match now\ncotal ps # expect it to be EMPTY here; ps.before is the target for\n # step 5, not for this step\n\n# 5. the step that is easy to skip: respawn the managed agents so their\n# credentials are re-minted as issuances and can renew. Persona is a\n# POSITIONAL argument here, unlike `cotal stop`, which requires --name.\n# One call per agent:\ncotal spawn <persona> --detach --name <n> --space <space>\n# then the comparison step 4 could not make:\ncotal ps # NOW compare against ps.before: seat count should match\n```\n\nThe mesh is down from step 2 until step 3 finishes. That is the window. On a split topology there is\nno cut and no backup inside it, so the window is the stop, the install and the restart, nothing more.\n\n## Adding a section for a future release\n\n**Every changeset marked breaking adds a section to this page.** A release that changes what an\noperator must do, in what order, or what stops working, is not finished until the section exists.\n`scripts/upgrade-section-gate.mjs` grades a commit range for this: run it as\n`pnpm upgrade-section-gate --base <ref>` and it reds when the range carries a breaking change and\nadds no new release section. CI runs its self-test and, as a step of the `attribution` job, grades\neach pull request's own range as `HEAD^1..HEAD` over the merge snapshot it checked out. That job is\nthe only context in the branch protection rule set, so a red gate FAILS A REQUIRED CHECK AND BLOCKS\nTHE MERGE. The section is not optional and a reviewer cannot wave it through without an\nadministrator overriding branch protection. Be precise about what the check proves either\nway, because one trusted past its evidence is worse than none. It proves a section for a release\n**was written here**. It cannot prove the section is **correct**, or that it describes the break\nthat actually landed, and it cannot see a breaking change that carries no marker at all. Reviewing\nthe words remains a person's job.\n\n**Mark the break, or the gate cannot see it.** Any one of these is enough, and they are the only\nthings it reads:\n\n- a `!` before the colon in the commit subject, as in `feat(core)!: bind hosted runs to the caller`\n- a `BREAKING CHANGE:` footer in the commit body\n- a changeset in `.changeset/` declaring a `major` bump for any package\n\nThe marker must survive the squash. A `!` that lives only in a commit you squash away is not in the\nrange the gate grades, so put it in the subject that lands on `main`.\n\n**The heading is a `##` and names the release**, like `## From 0.48.2 to 0.49.0`. Both matter, and\nneither is a style preference. Coverage is claimed by a heading, so a heading that names\nno release claims every release and distinguishes none: `## Notes` with a sentence under it would\notherwise satisfy the rule. Naming the release also makes the section the one an operator upgrading\nthat release will search for. Use `###` freely for detail inside a section. Subsections belong to\ntheir release rather than counting as separate coverage.\n\nName the release that first carries the change: the next version Changesets publishes, which\n`pnpm changeset status --verbose` lists. `bin/package.json` on `main` still reads the release already\npublished. If a release is cut while the change is open, the change ships in the release after it,\nso move the heading before merging. The gate accepts any version in a heading, so before merging a\nrelease pull request, check every heading added since the previous tag against the version it\npublishes.\n\nA section is written for the operator, not for the reviewer. It answers, in this order:\n\n1. What keeps working with no action at all.\n2. What does **not** migrate, and when that becomes visible. Name the log line if there is one.\n3. The order to move components in for a split topology, and why that order.\n4. What the outage window looks like, including what survives it.\n5. What to snapshot before starting.\n6. The commands, end to end.\n\n**Where an answer was not measured, say so in the document rather than guessing.** An operator who\nknows which half of a recommendation is reasoned and which is measured can plan around it; one who\nfinds out afterwards cannot.\n"
|
|
272
|
+
"body": "# Upgrading a running deployment\n\n> **Guide** (informative) · **For:** operators upgrading a mesh that already exists · **See also:** [Substrate stability](stability.md), [Run a mesh](run-a-mesh.md), [Identity and auth](identity-and-auth.md)\n\n[Substrate stability](stability.md) tells you what the version numbers promise. This page is the\nother half: what to actually do when the deployment already exists, has credentials in it, and\ncannot simply be recreated. Every release that breaks a running deployment gets a section here,\nnaming what migrates on its own, what does not, and the order to move the pieces in.\n\n## The pre-1.0 upgrade contract\n\nThe packages are pre-1.0, so a minor bump may break an API or an on-disk expectation. Four\ncommitments make that survivable for someone with a fleet:\n\n- **Pin an exact version.** `0.N.P`, never `^0.N.P`. A range can pull a breaking minor in during an\n unrelated reinstall.\n- **Every break that touches a running deployment gets a section on this page**, written in terms of\n what an operator does, not in terms of which module changed.\n- **Read the section before you start, not halfway through.** A section names the work up front\n precisely so the operation does not change shape once it is underway.\n- **A break that cannot be made automatic says so.** Where credentials or state must be recreated by\n hand, the section says which ones and when, rather than leaving you to discover it at the moment\n the first one stops working.\n- **A change to the shape of a credential, or to who may renew one, is breaking whatever the commit\n marker says.** This rule is stated because the marker is a judgement made while writing the code\n and the consequence is felt by someone running it a day later. A fleet that keeps authenticating\n looks compatible and is not, if nothing in it can renew. Any automated check of this rule would\n read commit markers, so a break recorded as a feature is the one case it could not see, which is\n why the rule is written for people first. **The marker held for this release: the 0.49.0 change\n that caused all of this, `36d177951 feat(core)!`, did carry its `!`.** The rule exists for the\n next one that does not.\n\nWhat this page does not promise is a rolling upgrade. Nothing in the current line dual-serves two\nauthority versions, so where broker and manager run separately there is a window in which the mesh\nis down. The sections below give that window's shape so it can be scheduled rather than endured.\n\n## Auth context closure in 0.71.0 (unreleased)\n\nExisting deployments need no credential migration or restart for these additive APIs. Embedded\nhosts can now inspect `handle.connections()` and await `handle.closed` after `close()` or `drain()`\nto prove every owned transport ended, including the callout, replaced readiness readers and\nshort-lived clients. The inventory is a detached snapshot.\n\nA transport close failure now rejects with its connection label. The terminal signal stays pending\nwhile any connection remains live. Repair the failure and retry `close()` before awaiting\n`handle.closed`. Closing one hosted context does not close another account's context.\n\nRead a space's claim with `readPlaneClaim(kv, space)` on that account's leader-only auth bucket.\nAn unclaimed space returns `undefined`; held and released rows retain their claim identity.\nDeleted, malformed and foreign-space rows refuse. `PlaneClaimRow` and `PLANE_CLAIM_KEY` are exported.\n\nUse `observeAccountLivenessWithCreds({ servers, observerCreds, accountId, options })` with the\naccount-scoped membership-observer credential to list that account's connections. It never widens\ncredentials or evicts connections. Zero rows prove absence only with a complete sweep and the\nsingle-server proof. An embedded endpoint's trusted composition can retain transport custody\nthrough `EndpointOptions.onConnection`.\n\n## Hermes model from the environment in 0.68.0\n\nA connector now launches on the model and variant its launcher resolved (the `--model` or\n`--variant` flag, else the agent file's `model:` or `variant:`) and no longer reads them again from\nthe agent file. The Hermes connector also no longer takes a model from `HERMES_MODEL` in the\nenvironment of the process that spawns the seat, including when `spawn.env` lists it.\n\n### What stops working\n\nA Hermes spawn whose only model was `HERMES_MODEL` in the spawning environment is refused at launch,\nand the refusal names both ways to set a model. Spawns that set `--model` or `model:` are unchanged,\non every connector.\n\nCode that calls a connector's `buildLaunch` directly with only `configPath` now gets no model or\nvariant from that file. Pass them as `model` and `variant`.\n\n### Before the upgrade\n\nMove each Hermes seat's model from `HERMES_MODEL` onto its spawn with `--model`, or into its\npersona's `model:`.\n\n## Run answers on a participant manager in 0.68.0\n\nA participant manager now asks its issuing host for an answering credential by naming the run and\nstep it answers. The host reads the pause's token off that run's journal and no longer accepts a\ntoken from the manager. Runs on a mesh with no participant manager are unaffected.\n\n### What stops working\n\nWhile a participant manager and its issuing host run different sides of this release, the host\nrefuses every `cotal run answer` and every amendment that manager serves, because each side refuses\nthe other's request shape. Starting, resuming and reading runs is unchanged. A pause stays waiting\nthrough the window, or follows its timeout if it has one.\n\n### Before the upgrade\n\nUpgrade the auth service and every participant manager registered with it in the same window, then\nanswer the pauses that waited.\n\n## Headless OpenCode handshake in 0.69.0\n\nWith `COTAL_SERVE_HEADLESS=1`, the OpenCode launcher's `[cotal-serve]` line on stdout now carries\nonly `port` and `session`. The server password no longer appears in it, and the 1.x TUI no longer\nreceives the password on its command line.\n\n### What stops working\n\nA headless host that read `password` from that line has no password, and the server refuses its\nrequests. Seats with a TUI, and headless seats that no host drives, are unaffected.\n\n### Before the upgrade\n\nHave each headless host mint a password and pass it to the launcher as `OPENCODE_SERVER_PASSWORD`,\nthen use it for basic auth as before. Without that variable the launcher mints its own.\n\n## Filesystem store identity in 0.69.0\n\nThe delivery daemon's answer to the manager's store check now names a filesystem store by its root\nand by a random `id` that the store records once in `store.id` inside its own directory:\n`.cotal/store.id` for a workspace root, or the directory of the file for `cotal deliver --creds\n<file>`. A manager no longer counts the daemon's store as its own because the two roots have the\nsame path. On a split whose broker host and manager host use one root path, the manager host now\nstays off the daemon-credential renewal lease, so `cotal doctor auth --fix` on the broker host can\nrenew the daemon credentials.\n\n### What stops working\n\nA manager and a delivery daemon on different sides of this release refuse each other's answer to\nthe store check. The manager then remints no daemon credential, and a manager that is booting does\nnot start. This is read from the code and was not measured across two releases. A\n`cotal deliver --creds <file>` whose directory is a read-only mount and holds no `store.id` stops at\nstart. So does a `--creds` file that is its directory's `store.id` under any name, and a `store.id`\nthat is a symbolic link or holds anything but a lowercase UUID.\n\n### Before the upgrade\n\nUpgrade the broker host and every manager host of a space in the same window. For a `--creds` file\non a read-only mount, add a regular `store.id` file beside it that holds a new lowercase UUID and no newline,\nas `node -e 'process.stdout.write(crypto.randomUUID())' > store.id` writes. Move a `--creds` file\nnamed or linked as `store.id` to a file of its own.\n\n## Detached spawns with `--share-tools` in 0.69.0\n\nThe manager's `spawn` operation now takes `shareTools` as a list of MCP server names. The CLI parses\n`--share-tools` into that list before it sends the request, and the manager cluster document moves\nto revision 22. A cut taken with `cotal down --preserve-state` before the upgrade still resumes: the\nmanager reads its `cotal-manager-resume/v1` inventory and writes new cuts as\n`cotal-manager-resume/v2`.\n\n### What stops working\n\nA CLI and a manager on different sides of this release refuse a detached spawn that passes\n`--share-tools`, because the CLI checks each request against the contract the manager serves. This\nis read from the code and was not measured across two releases. A detached spawn without the flag,\na foreground spawn and a roster entry are unaffected. A manager older than this release cannot\nresume a cut that this release took.\n\n### Before the upgrade\n\nUpgrade the CLI on every host that runs `cotal spawn --detach` in the same window as the managers\nit reaches.\n\n## Shared MCP server checks in 0.69.0\n\nThe cotal config reader now checks each server under `connectors.<name>.mcpServers` when it reads\nthe file, and refuses one that cannot launch as written, naming the file and the field. The rules\nare in [the config file](config.md#the-config-file).\n\n### What stops working\n\nA config file that holds such a server refuses every Claude spawn that reads it, including one with\n`--share-tools none`. Before, a field of the wrong type failed each Claude spawn that shared the\nserver with a `TypeError` that named neither the file nor the server, a spawn that did not share it\nlaunched, and a server with no `command` or `url` was passed to `claude`, which never started it.\nRead from the code and not measured: spawns on other connectors, a manager resume and the step of\n`cotal setup` that records the shared list read the same files, so each stops at the same refusal.\n\n### Before the upgrade\n\nCheck `connectors.<name>.mcpServers` in the operator-level config file and in each space's\n`.cotal/config.json`. Give each server a string `command`, or a `type` of `http`, `sse` or `ws` with\na string `url`. Write `args` as a list of strings and `env` and `headers` as objects of strings, or\nremove the server.\n\n## Remote manager family eviction in 0.69.0\n\nA remote manager registered through its host now asks the host to evict up to 256 holders of its\ncredential family in one maintenance request, and the host reads the family once for the whole set.\nBefore, a restart sent one request per holder and the host read the whole family for each one.\nMeshes with no remote manager are unaffected.\n\n### What stops working\n\nWhile a remote manager and its issuing host run different sides of this release, each side refuses\nthe other's eviction request shape. A restart whose credential family already has holders then fails\nat its eviction step and leaves the manager's registration gate frozen. A first start, a clean stop\nand the host's reconciliation of a foreign slot holder are unchanged.\n\n### Before the upgrade\n\nUpgrade the auth service and every remote manager registered with it in the same window. A manager\nthat restarted inside the window resumes its frozen registration on its next start once both sides\nrun this release.\n\n## AG-UI emitter holder hooks in 0.69.0\n\n`AguiEmitterHolder` from `@cotal-ai/connector-core` now takes its hooks as one named object after\nthe emitter factory: `new AguiEmitterHolder(startEmitter, { onError, onRunClosed, waitLive, runMeta })`.\nOnly `onError` is required. Nothing about a running mesh changes, and every shipped connector passes\nits hooks by name. Only a connector of your own that builds a holder is affected.\n\n### What stops working\n\nA holder built with positional hooks, such as `new AguiEmitterHolder(start, onError, onRunClosed)`,\nno longer compiles, because the constructor takes two arguments. Plain JavaScript that keeps the\npositional form still runs, but the holder calls none of its hooks, so a failure never reaches\n`onError`.\n\n### Before the upgrade\n\nPass each hook by name, for example `new AguiEmitterHolder(start, { onError, onRunClosed })`, and\ndrop any `undefined` that filled an earlier slot to reach a later hook.\n\n## Worker run failure type in 0.70.0\n\n`WorkerRunFailed`, the failed result of `runInWorker` in `@cotal-ai/lang`, is now a union on\n`class`: `released`, `held`, `effect`, `too-large`, `rejected` or `error`. A running mesh needs\nnothing, because the runtime host and the engine thread ship in the same install. A run on the\ncompiled engine whose program throws an object with `code: \"L5012\"` or `code: \"L5025\"` used to end\nreleased and now ends failed, as it does on the walker.\n\n### What stops working\n\nTypeScript code that reads `code`, `reason`, `step`, `pending`, `kind`, `detail` or `tooLarge` on a\n`WorkerRunFailed` it has not narrowed fails with TS2339. A released, held, too-large or rejected\nresult no longer carries `code`, so JavaScript that branched on `L5012`, `L5025`, `L5006` or\n`L5010` stops matching with no error. `tooLarge` is gone.\n\n### Before the upgrade\n\nBranch on `class` where such code read `code`: `released` for L5012, `held` for L5025, `too-large`\nfor L5006 and `rejected` for L5010. An `effect` or `error` result keeps its `code`. Once narrowed to\n`too-large`, a result carries the `stepKey`, `bytes` and `bound` that `tooLarge` held.\n\n## Remote manager request builder in 0.70.0\n\n`remoteManagerClient.remoteManagerAuthorityRequest` from `@cotal-ai/manager` now takes an\noperation's coordinates as one object, and `remoteManagerRegistrationProof` from `@cotal-ai/core`\ncomputes the proof from the manager's identity state instead of a request. Nothing about a running\nmesh changes: the proof digest and the request on the wire are the same, so a manager and a host on\ndifferent sides of this release still accept each other. Only code that builds remote manager\nrequests itself is affected, in TypeScript and in plain JavaScript.\n\n### What stops working\n\nA call that passes the registration proof, contract artifacts, session, retirement or transfer\nreader as positional arguments after the operation no longer compiles. A call that passes a request\nto `remoteManagerRegistrationProof` no longer compiles either, because the second argument now names\nthe lifecycle `lifecycleUid`, as the identity state does.\n\nPlain JavaScript runs both old calls without an error. The builder drops the positional coordinates,\nso the host refuses the request with `requires a sha256 registrationProof`. A proof computed from a\nrequest leaves out the lifecycle, so the host refuses a request that carries it as a proof mismatch.\n\n### Before the upgrade\n\nName the coordinates, for example\n`remoteManagerAuthorityRequest(state, \"cli\", \"retire\", { registrationProof, retirement })`.\nCompute the proof as `remoteManagerRegistrationProof(owner, state)`, adding the contract artifacts\nas a third argument for activation only. A host that recomputes the proof from a received request\npasses `{ space, instanceId, lifecycleUid: managerLifecycleUid, identities }` from that request.\n\n## Bearer validator lifetime cap in 0.70.0\n\n`validateUserToken` from `@cotal-ai/auth` no longer takes `maxTtlSec`. It caps a bearer's lifetime\nat the cap of the bearer's view, the same cap the issuer applies when it mints: 900 seconds, or 300\nfor a `transfer-writer` bearer. The auth callout never passed the option, so a running mesh behaves\nas before. Only code of your own that calls the validator with `maxTtlSec` is affected.\n\n### What stops working\n\nA call that passes `maxTtlSec` in an object literal no longer compiles. Plain JavaScript that keeps\nit still runs, and the value is ignored. A `NaN` value, such as `Number()` of an unset environment\nvariable, used to turn the lifetime check off and accept a bearer of any lifetime. That bearer is\nnow refused at its view's cap.\n\n### Before the upgrade\n\nRemove `maxTtlSec` from each call. A test that needs a bearer to expire sooner mints one with a\nshorter lifetime.\n\n## Persisted identity records in 0.70.0\n\nThe manager instance identity, the manager sibling identities, the auth plane instance identity and\na participant manager's remote authority state now share one reader and one first mint in\n`@cotal-ai/workspace`, exported as `claimIdentityRecord` with the nkey check `identityOf`. Each\nrecord is read as a regular file, must hold non-empty nkeys and is created exclusively, so\nconcurrent first starts of a participant manager on one root now settle on one identity where each\nused to keep its own. `saveManagerInstanceIdentity` and `saveAuthInstanceIdentity` are gone. A\nrunning mesh whose records are plain files needs nothing.\n\n### What stops working\n\nA manager instance, auth instance or remote authority record that is a symlink, a directory or any\nother non-regular entry is refused where it used to be followed. The manager, the auth plane and a\nparticipant manager fail to start on it, and `cotal reconcile-gate` and `cotal deregister-instance`\nrefuse it. Retirement already refused it. A remote authority record with an empty nkey id or seed\nis refused too. A first mint that loses its race and cannot read the winner now refuses with\n`identity-record-create-lost` in place of `manager-instance-identity-create-lost` or\n`auth-instance-identity-create-lost`. Code that imports either `save` function no longer compiles.\n\n### Before the upgrade\n\nReplace a symlinked identity record with a copy of the file it points to. Code that wrote a record\nwith a `save` function plants it with `createManagerInstanceIdentity` or\n`createAuthInstanceIdentity`, which create the record when it is absent and otherwise return the\nstored one unchanged. Nothing replaces an overwrite of a stored identity.\n\n## Manager instance in user credentials in 0.70.0\n\n`AuthProvider.userCredentials` from `@cotal-ai/core` no longer returns `managerInstanceId`. A\n`manager-caller` credential's manager instance is the signed `act.managerInstanceId` claim in its\nbearer, which the broker verifies and the CLI already used. The reference provider in\n`@cotal-ai/auth` stops copying the exchange response's field into its result, where nothing\ncompared it with the bearer. The exchange still answers with the field, so a running mesh behaves as\nbefore.\n\n### What stops working\n\nCode of your own that reads `managerInstanceId` from a `userCredentials` result no longer compiles,\nand plain JavaScript reads `undefined` there.\n\n### Before the upgrade\n\nRead the instance from the bearer's `act.managerInstanceId` claim.\n\n## Auth plane identity location in 0.70.0\n\nThe user-auth service keeps its instance identity in the root's `.cotal/space.<hex>/auth-instance.json`,\nbeside the manager's. It used to sit inside `.cotal/auth`, at\n`space.<hex>/.cotal/auth/auth-instance.<hex>.json`, so a copy of that folder carried it. The first\nstart of an upgraded root moves the record and keeps the instance. A hosted context started through\n`startAuthService` has its record moved the same way inside its `stateDir`.\n\n### What stops working\n\nCode that calls `openAuthAuthorityPlane` without the new `identityRoot` option no longer compiles. A\nstart that finds a record both in `.cotal/space.<hex>/` and at its older place refuses and names the\ntwo files. A start also refuses when the older place of the auth or manager identity holds a symlink,\na directory or anything else that is not a regular file. The manager used to skip a dangling symlink\nthere and mint a new identity.\n\n### Before the upgrade\n\nPass `identityRoot` to `openAuthAuthorityPlane`. When `dir` is a workspace root's user-auth state\ndir, `<root>/.cotal/auth/space.<hex>`, pass that root. A plane with no workspace root, as\n`startAuthService` runs, passes `dir` itself. Either keeps the identity the plane already has: on the\nfirst start it moves from `<dir>/.cotal/auth/` to `<identityRoot>/.cotal/space.<hex>/`. Never pass a\ndirectory inside `.cotal/auth`: the record would land in the folder an operator copies and travel\nwith it again.\n\nA copy of `.cotal/auth` taken from a root last run by an older Cotal carries that root's record.\nDelete `.cotal/auth/space.<hex>/.cotal/auth/auth-instance.<hex>.json` from the root you copied it to\nbefore the first `cotal up --user-auth` there.\n\n## Per-seat `COTAL_` names in `spawn.env` in 0.71.0\n\n`spawn.env` in the cotal config no longer forwards a `COTAL_` name the launcher sets for each seat,\nsuch as `COTAL_ROLE`, `COTAL_MODEL` or `COTAL_SUBSCRIBE`. Before, a seat launched with no value of\nits own took the spawning process's value and ran under that role, model or read set. The\nmachine-wide knobs a seat already receives, such as `COTAL_HOME`, may still be listed.\n\n### What stops working\n\nEvery spawn and resume under a config whose `spawn.env` lists such a name is refused before\nlaunch, and the refusal names the entry. Code that calls `launchEnv` from `@cotal-ai/connector-core`\nwith such a name in `envAllow` gets the same error.\n\n### Before the upgrade\n\nRemove those names from `spawn.env`. Give each seat its role, model and channels with `--role`,\n`--model` and `--subscribe`, or in its persona's `role:`, `model:` and `subscribe:`.\n\n## Role addresses in 0.71.0\n\nA role must be one `[A-Za-z0-9_-]` token. Before 0.71.0 any other spelling was rewritten into one:\n` probe ` reached the `probe` queue and `pro.be` reached `pro_be`, while the message kept the\nspelling sent. An anycast to `*` was accepted and stored where no holder reads it.\n\n### What stops working\n\nAn agent whose role is outside the token set no longer starts, however it is launched:\n`cotal join --role`, `cotal spawn --role`, an agent file's `role:`, `COTAL_ROLE` and an embedded\nendpoint's `card.role` are all refused before the agent joins.\n\nA send to such a role, or to `*`, through `cotal send ask`, `/anycast` or `cotal_anycast` is refused,\nand nothing is stored.\n\n`routeToken` is no longer exported from `@cotal-ai/core`. A role routes as spelled, so code that\nused it to name a role's queue uses the role itself, and `assertValidRole` checks one.\n\n### What migrates on its own\n\nEvery task queue. A `svc_<role>` durable was always named from the rewritten token, so its pending\nrequests and its holders carry over.\n\n### Before the upgrade\n\nRename each role outside the token set to the token it already routed to: remove the surrounding\nspaces and replace every other character outside the set with `_`. Rename it where the holder is\nlaunched and in every script or prompt that sends to it.\n\n## Carrying a resumed Claude session to another host in 0.67.0\n\n`cotal spawn --resume <id> --detach --on <instance>` now carries a Claude session held on the\noperator's host to the target manager instance. Both sides need this release: an older manager does\nnot serve `transcript-receive`, and the CLI then stops with that manager's refusal instead of\nlaunching. The manager cluster document moves to revision 21, and the `ps` row's `resume` object\ngains `host` and `transferredAt`.\n\nA manager host that runs carried seats needs `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_AUTH_TOKEN` or a\ncloud provider selection in its environment, because each carried seat runs in its own Claude home\nwith no stored login. On an authenticated mesh the CLI mints the transfer writer from the space's\nsigning seed, so the carrying host needs that seed, as for any other operator command. On a user-auth\nmesh it exchanges the operator's login for a `transfer-writer` view instead, so the operator's grant\nneeds scope `admin`, and the auth service must run this release. A remote manager receives a carry once\nits host serves the manager-service `transferReader` operation. A seat launched without carrying,\nincluding any `--resume` whose id this host does not hold, is unchanged.\n\n## Lifecycle head type in 0.67.0\n\n`LifecycleMapping`, the type `parseLifecycleHead` returns, is now a union on `state`. Nothing about\na running mesh changes: heads that parsed before parse the same way, and the refusals are\nunchanged. Only TypeScript code that compiles against `@cotal-ai/core` is affected.\n\n### What stops working\n\nAn `interface` that extends `LifecycleMapping` fails with TS2312, because an interface cannot extend\na union. Code that builds a head in memory no longer compiles when the head is `retiring` without\nits `op`, or `active` or `retired` with one. The parser already refused those heads.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type ActiveMapping = LifecycleMapping & { state: \"active\" }`. A reader that has checked\n`state === \"retiring\"` reads `op` without a guard.\n\n## Issuance gate types in 0.67.0\n\n`EpGateRow` and `EndpointGateRow`, which `parseIssuanceGate` and `parseEndpointGate` return, and\n`EpGateState`, which an `EpIssuanceGate` or `EpIssuanceBarrier` returns from `observe`, are now\nunions on `state`. Nothing about a running mesh changes: gates that parsed before parse the same\nway, and the refusals are unchanged. Only TypeScript code that compiles against `@cotal-ai/core`\nis affected.\n\n### What stops working\n\nAn `interface` that extends one of these types fails with TS2312, because an interface cannot\nextend a union. Code that builds a gate in memory, such as a custom barrier's `observe`, no longer\ncompiles when the gate is `frozen` or `retired` without its `op`, or `open` with one. The gate\nparsers already refused those rows.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type CustomGateRow = EpGateRow & { custom: string }`.\n\n## Lifecycle-blocked refusals in 0.66.0\n\nA refusal that carries `ai.cotal.ep.lifecycle-blocked` now reports only the lifecycle state it\nread. Nothing about a running mesh changes. A client that branches on the detail must read the new\nfield.\n\n### What stops working\n\nA refusal raised at the issuance gate used to carry `headState` without reading the head:\n`retiring` for a frozen gate and `retired` for a retired one. It now carries `gateState`\n(`frozen` or `retired`) and no `headState`. A client that treats `headState: \"retired\"` as a\nburned uid, or `headState: \"retiring\"` as a retirement in flight, no longer matches those\nrefusals, and the `[lifecycle ...]` suffix on the error string changes the same way. A custom\nissuance barrier whose `observe` returns a frozen gate without a valid `op` (a string `opId` and\none of the four op kinds) is now refused as `internal` by `registerServiceInstance`.\n\n### Before the upgrade\n\nUpdate such a client to read `gateState` for a gate refusal and `blockedOp` for the operation that\nholds the gate. `headState` is present only when the refusal read the head, for example an\nactivation refused because the head is still retiring.\n\n## Workflow programs that bind `once` in 0.65.0\n\n`once` is now a scope of the workflow language, so it is a reserved name. A program that declares\nits own `once` binding (`const once = ...`, a parameter or a function named `once`) is refused at\nvalidation with L2002. Nothing else about a running mesh changes.\n\n### What stops working\n\nA run whose recorded program binds `once` cannot be resumed after the upgrade, because a resume\nvalidates the recorded program again. A new `cotal run start` of such a program is refused before\nanything is recorded.\n\n### Before the upgrade\n\nList the runs with `cotal run ps` and check each program that is still running or held for a\nbinding named `once`. Let those runs finish on the old version before you upgrade the manager, and\nrename the binding in the program before you start it again.\n\n## From 0.58.0 to 0.59.0\n\nEvery connector now publishes a failed run's `RUN_ERROR` on `events.<owner>.<actor>` with the fixed\nmessage `run failed` and no `code` or `rawEvent`. The error text and error kind a harness reports\ncan echo a prompt, a peer message or tool output, and that channel has a different read ACL. A\nreader that showed the message or branched on `code` gets neither after the upgrade. Where a\nconnector reports the error kind as the agent's presence condition, that is unchanged.\n\n### Settle pending event frames before the upgrade\n\nEach session's events are frozen in its event write-ahead log before they are published. A session\nrestarted on 0.59.0 whose log still holds an unacknowledged frame with an older `RUN_ERROR` does not\nrepublish it: its event emitter halts with `egress-run-error` and publishes nothing further for that\nsession. The broker may or may not already hold that frame, so the halt cannot settle it.\n\n1. Stop the seats cleanly on 0.58.0, with the broker still up.\n2. List the logs that still hold a pending frame. The logs live under the events state root\n (`COTAL_WORKSPACE_ROOT`). Empty output means there is nothing to settle.\n\n ```sh\n find \"$COTAL_WORKSPACE_ROOT/.cotal/events\" -name wal.json \\\n -exec jq -r 'select(.pending != null) | input_filename' {} +\n ```\n\n3. For each session listed, start it again on 0.58.0 while the broker is reachable, let it recover,\n stop it, and run step 2 again. Recovery publishes the frame as 0.58.0 would have, error text\n included, so it only finishes what 0.58.0 had already started.\n\n If that start halts with `cas-loss` instead, the agent's subject is no longer at the sequence this\n log expects, and no restart settles that log, on 0.58.0 or later. A lost acknowledgement is one\n cause: the broker stored the frame, so it and its error text are already on the channel, and every\n retry halts the same way because the stream checks the frozen expectation before it deduplicates.\n The halt message names the other causes, such as a second emitter for the same agent under a\n different state root, a restored stream or frontier record, or a purged channel. With those the\n pending frame may never have reached the broker, so a `cas-loss` does not tell you whether it\n landed. Find and stop any second writer and rule out a restored state first. Clearing the halt\n then means purging the agent's event channel and removing the agent's directory under the events\n state root whole (see [Event plane](connect-claude.md#event-plane)). That abandons the pending\n frame whether or not the broker has it, and the purge also drops the earlier frames of every\n session of that agent.\n4. Upgrade once step 2 prints nothing.\n\nIf a session halts with `egress-run-error` after the upgrade, go back to step 3 for that session on\n0.58.0. Do not edit or delete `wal.json` on its own to get past either halt: clearing the pending\nframe abandons that epoch, an event the broker never received is lost, and removing part of the\ndirectory leaves a state the next start refuses.\n\n## Explicit actor grants in 0.59.0\n\n`cotal actor grant` no longer fills an omitted ACL flag with its wide default. A grant names\n`--scope`, `--allow-subscribe` and `--allow-publish`, or passes `--full` to give the ones it leaves\noff their wide defaults (`spawn,role:default`, `>` read, `>` post). Any other grant is refused. The\nbreak is in the CLI on the machine that holds the actor ledger, the one that ran\n`cotal up --user-auth --idp <url>`. No stored row, credential or wire message changes.\n\n### What keeps working\n\nExisting actor ledger rows keep the authority they were granted, and their users and agents connect\nas before. `actor revoke`, `actor list` and a `grant` that names all three ACL flags behave as they\ndid on 0.58.0. Nothing on disk is converted.\n\n### What stops working\n\nA grant that leaves off any of the three flags without `--full` exits 1 with\n`refusing to grant \"<actor>\" with --scope, --allow-subscribe, --allow-publish left off`, naming the\nflags it is missing, and then prints both accepted forms. It writes no row and does not retire the\nactor's current lifecycle. An existing row stays as it was, and an actor granted for the first time\nstays out until the grant is run again. This includes the bare grant printed on 0.58.0 by\n`cotal login`, `cotal status`, `actor list` and the not-granted refusal. Look for it in provisioning\nscripts, onboarding runbooks and anything that pastes those hints.\n\n### Upgrade order\n\nChange the scripts before the ledger machine is upgraded, and make each grant name all three flags.\n0.58.0 and 0.59.0 both accept that form. To keep a wide row, write its defaults out:\n\n```sh\ncotal actor grant <actor> --sub <IdP subject> \\\n --scope spawn,role:default --allow-subscribe '>' --allow-publish '>'\n```\n\nSwitch to `--full` only once the ledger machine runs 0.59.0. 0.58.0 refuses it with\n`Unknown option '--full'` before it reads the ledger. Brokers, managers and participant machines\nneed nothing for this break, so their order is the one the section above gives.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused grant changes nothing. The\nexposure is a grant script that runs against 0.59.0 before it was changed: it fails and grants\nnothing.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. On the ledger machine, save the output\nof `cotal actor list` to compare rows after the changed scripts run, and list the scripts that call\n`cotal actor grant`.\n\n### The upgrade end to end\n\n```sh\n# on the ledger machine, still on 0.58.0\ncotal actor list > actors-before.txt\ngrep -rn 'actor grant' <your provisioning scripts>\n# make every grant name --scope, --allow-subscribe and --allow-publish, run them, then upgrade\nnpm i -g cotal-ai@0.59.0\ncotal actor list | diff actors-before.txt -\n```\n\nBoth refusals quoted here were run on 0.58.0 and on the 0.59.0 code. That brokers, managers and stored\nrows need nothing is read from the change, which touches only the CLI and its hints, and was not run\non a live split deployment.\n\n## Repeated flags refused in 0.59.0\n\nA `cotal` flag given more than once is now a usage error unless the command declares it repeatable.\nOn 0.58.0 the last value won with no message, so `cotal down web --space a --space b` acted on `b`\nwhile a wrapper that checked the first `--space` verified `a`. The break is in the command-line\nparser on the machine that runs the command, including commands added with `cotal ext add`. No\nstored state, credential or wire message changes.\n\n### What keeps working\n\nA command line that gives each flag once parses as it did on 0.58.0, in any order and in the\n`--flag=value` form. Flags whose help says repeatable, such as `--opt` and `down --session-store`,\nstill collect every value. A flag-shaped word after `--` is still a positional. The daemons, units\nand agents that `cotal` starts for itself are given each flag once, so a fleet driven only by `cotal`\ncommands typed by hand needs no action.\n\n### What stops working\n\nA command line that repeats any other flag exits 1 before the command runs. It prints\n`Option '--space' cannot be repeated`, or `Option '-f, --file' cannot be repeated` for a flag with a\nshort form, followed by the command's help. `-f` and `--file` count as the same flag. Look for it in\nscripts, aliases and wrappers that append a flag to override one set earlier, such as a fixed\n`--space` followed by `\"$@\"`.\n\n### Upgrade order\n\nChange those scripts first so each flag is given once. 0.58.0 and 0.59.0 both accept that form.\nBrokers, managers and participant machines need nothing for this break, and each machine's CLI\napplies it when that machine is upgraded, so their order is the one the sections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused command does nothing. The\nexposure is a script that still repeats a flag when it runs on 0.59.0: it exits 1 instead of acting on\nthe last value.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the scripts, aliases and wrappers\nthat call `cotal` so each one can be checked.\n\n### The upgrade end to end\n\n```sh\n# still on 0.58.0\ngrep -rn 'cotal ' <your scripts and wrappers>\n# give each non-repeatable flag once, then upgrade\nnpm i -g cotal-ai@0.59.0\n# run each changed script; a repeat left behind exits 1 with the usage error and does nothing\n```\n\nThe refusal and its messages were run against the 0.59.0 parser and `cotal topology view`. That the\nargument lists `cotal` builds for its own processes give each flag once is read from the code, and\nwas not run on a live split deployment.\n\n## Detached spawns from a seat's shell in 0.62.0\n\nOn a static or open mesh, `cotal spawn --detach` run inside a managed seat's shell now launches as\nthat seat. On 0.61.0 it minted a one-shot operator instrument, so the manager recorded that\ninstrument as the spawner and the seat's own `cotal_despawn` of the child was refused with\n`not authorized: <seat> was not spawned by <caller> (admin tier required)`. The break is in the CLI\non the machine where the seats run. No stored state, credential or wire message changes.\n\n### What keeps working\n\n`cotal spawn --detach` from an operator terminal or from a script outside any seat launches as\nbefore, and so does any call with `--creds`, one aimed at a space other than the seat's own, or a raw\nopen target named with `--server` and an unregistered `--space`. A user-auth mesh is unchanged. A seat with\n`capabilities: [spawn]` still spawns from its shell, and can now stop that child with\n`cotal_despawn`. `--on <instance>` from a seat's shell still lands on that manager instance, now as\nthe seat.\n\n### What stops working\n\n- On a static mesh, a seat without `capabilities: [spawn]` can no longer spawn from its shell. Its\n own credential holds no spawn subject, so the broker refuses the request and the command exits 1.\n- A child launched from a seat's shell is now that seat's child, so the manager stops it when the\n seat exits, as it does for a `cotal_spawn` child. A child that has to outlive the seat that\n started it now goes with the seat.\n- A seat launched without `COTAL_SPACE` is placed by its static credential. Every connector sets\n that variable, so this only reaches a hand-built launch: from such a seat's shell, a spawn aimed at\n a static space that holds no credential for the seat is refused instead of running as the operator.\n\n### Upgrade order\n\nOnly the CLI that seats run from their shell changes, which is the one installed on the host where\nthe seats run. Brokers and managers need nothing for this break, so their order is the one the\nsections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it. A child already running when you upgrade\nkeeps the spawner the manager recorded at its launch.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the agent files whose seats run\n`cotal spawn --detach` from their shell, note which of them lack `capabilities: [spawn]`, and note\nwhich of their children must outlive the seat.\n\n### The upgrade end to end\n\n```sh\n# still on 0.61.0: find the seats that spawn from their shell\ngrep -rln 'cotal spawn' .cotal/agents\n# add `capabilities: [spawn]` to each of those agent files that lacks it, and launch any child\n# that must outlive its seat from an operator terminal instead\nnpm i -g cotal-ai@0.62.0\n```\n\nThe attribution, the despawn, the refusal of a seat without `spawn`, the stop on seat exit and a\nseat's `--on` spawn were run on a local static mesh, and the attribution and the despawn on a local\nopen mesh.\n\n## From 0.53.0 to 0.54.0\n\nManager calls now borrow an instance-bound `manager-caller` credential. Followed mutations require\n`manager.goal-result` on the selected manager, so a compatible issuer, manager and client must be\nloaded together. An older manager is refused before a followed mutation; upgrading an installed\nbinary alone does not replace code in a running manager, connector or embedded client.\n\n### Preserve state before changing processes\n\nSnapshot the broker's durable storage using its supported backup procedure, the host authority and\nactor ledgers, and each participant's manager identity, runtime custody records, credentials and\nsaved sessions. Include the embedding application's database and configuration under its supported\nbackup procedure. Record the loaded package versions and the CLI path used by bearer helpers.\nKeep these copies private. Do not change the IdP issuer, regenerate manager identities, rotate agent\ncredentials or recreate tenant storage to make the upgrade pass.\n\nNo ledger, goal-history or session conversion is required for this change. Existing ordinary\nmessaging credentials retain their normal expiry rules. New manager-caller credentials are obtained\non demand from the current grant; old manager-call credentials do not gain the new view automatically.\nExisting accepted goals remain durable and must not be submitted again merely because observation\nwas interrupted. Fresh remote registration publishes its service status at the current revision and\nepoch; do not seed that status manually.\n\n### Upgrade the split deployment\n\n1. Stage one pinned 0.54.0 package set for the host and participants, including the embedding SDKs.\n Pause new manager mutations and let accepted work settle where possible before reloading processes.\n2. Upgrade the host issuer and embedding first. Keep the broker, its account identities and durable\n storage in place. Then load the matching manager release on each participating machine.\n3. Preserve active seats through the runtime's supported update path. A Linux custodial runtime may\n release and re-adopt seats within its 600-second unattended window; verify the actual runtime,\n custody records and process identities before relying on it. A legacy PTY runtime without release\n support cannot preserve active seats through a generic manager restart. Drain it at an approved\n idle window instead of signalling the manager or replacing conversations.\n4. Reload the clients and connectors through their session-preserving host controls. Refresh any\n bearer helper captured from an older immutable CLI path. A transport-only reconnect does not reload\n JavaScript. Verify authenticated instance selection, a read-only manager command and canonical\n result recovery before allowing new followed mutations.\n\nTreat the interval from issuer reload through compatible manager/client reload as a manager-control\noutage. Mixed versions can refuse discovery or commands; there is no promised rolling transition.\nOrdinary agent sessions survive only where their runtime and credentials permit it. If verification\nfails, keep mutations paused and repair forward from the preserved state rather than resetting it.\nThis release does not add host-backed enrollment or terminal release for stock participant detached\nagents; see [Remote supervised agents](run-a-mesh.md#remote-supervised-agents).\n\n## From 0.48.2 to 0.49.0\n\n0.49.0 changes how a credential's authority is recorded. A credential is no longer only a signed\nfile: it is an *issuance*, with a generation the issuer chose and durable evidence of the ceiling it\nwas granted under. The important consequence for a running deployment is not at connect time. It is\nat renewal time.\n\n### What keeps working without any action\n\n- **Existing agent credentials keep authenticating.** A credential minted under 0.48.2 is not\n revoked and is not rejected at connect. Nothing needs to be re-issued to bring the fleet back up\n after the upgrade.\n- **The channel registry survives.** Channels, their replay settings, descriptions, and usage text\n are ordinary durable state and are not rewritten by the upgrade.\n- **`cotal deliver` is still a standalone command.** Running the delivery daemon as its own process\n remains supported; it is not restricted to being a child of `cotal up`.\n- **`cotal join` keeps its flags.** In particular `--lifecycle-uid` is not new in 0.49.0. It has\n been required alongside `--creds` since well before this release, and the pairing rule did not\n change here. A scripted external join that worked under 0.48.2 works unchanged.\n\n### What does not migrate\n\n**A credential minted before 0.49.0 cannot be renewed.** Managed agent credentials carry a\n24-hour lifetime and the manager re-signs one once it passes **75%** of its life, ticking every\nquarter of the TTL so a tick always lands inside that window. When the manager reaches a credential\nthat carries no issuance, it refuses to renew it and logs the agent by name:\n\n```\n! managed cred renewal <agent>: renewManagedStaticCred: <agent> carries no issuance;\n a static credential minted before SPEC 13.15 is not renewed under an unbound generation\n - respawn the agent\n - the agent dies loud at this cred's expiry unless it is reminted\n```\n\nSo the fleet comes up fine, runs normally, and then each agent stops at its own credential's\nexpiry, within roughly a day of the upgrade, one at a time rather than together. The refusal is\ndeliberate: the renewal would otherwise have to invent a generation nobody issued, which is the\nstate the release exists to remove.\n\n**Respawn the managed agents as the last step of the upgrade.** For this particular upgrade the\nrespawn is not optional: stopping a 0.48.2 manager ends its agent processes whichever CLI you use,\nfor the reason given under the outage window below. The respawn is how they come back, and it is\nalso what mints each credential as an issuance so it renews from then on. One planned pass over the\nfleet is the whole job. Skipping it leaves agents stopped and, for any credential that survived\ninto 0.49.0 unminted, brings the renewal cliff above a day later, one agent at a time.\n\n### Credentials you minted yourself\n\n**A credential you minted with `cotal mint` is a different case, and it very likely needs\nnothing.** The distinction that matters is not the word \"static\", which covers both. It is **what\nminted the credential and who owns its renewal**. A credential the **manager** minted for an agent\nit spawned carries a lifetime and is renewed by the manager, so it is the subject of everything\nabove. A credential **you** minted with `cotal mint` and handed to an external peer is issued with\n**no expiry at all**, and no manager renews it: it is not in the sweep, so there is no renewal to\nfail. It keeps working after the upgrade, and re-minting it would mean coordinating with a third\nparty for no gain.\n\nThe manager says which one it is holding. Where a credential has no expiry to reach, the sweep\nnames it and moves on rather than refusing:\n\n```\n! managed cred renewal <agent>: credential is unbounded - not renewed\n (a pre-TTL credential stays as minted until respawn)\n```\n\nRe-mint an external peer's credential only if you want it to carry a lifetime, and at a time you\nchoose.\n\n### How to read the boot log\n\nA 0.49.0 manager starting over an existing space may print lines like:\n\n```\n verified evicted: <holder-key> (3/12)\n already verified (durable): <holder-key>\n✓ boot self-heal: manager/<id> registration gate reopened at generation <n>\n```\n\nThese are **not** a credential migration, and reading them as one is the most likely way to\nconclude the fleet is fine when it is not. They come from the manager repairing **one** endpoint\nregistration gate that a previous restart left frozen, and they enumerate that single gate's\ncredential-family holders as it verifies each one evicted. `already verified (durable)` on a later\nstart is the repair cursor resuming, not a credential that became durable. The repair is real and\nuseful (it is what previously needed `cotal reconcile-gate` by hand), but it says nothing about\nwhether your agent credentials carry issuances. The renewal refusal above is the signal that does.\n\n### Which side to upgrade first in a split topology\n\nMove the manager first.\n\nThe stores 0.49.0 introduces are created by the **manager** at its own boot, not by the broker.\nThey are create-or-verify and idempotent, so a 0.49.0 manager brings the space's authority stores\nup to the new shape itself, and it does so against whichever broker is answering.\n\nBeing honest about the evidence behind each direction, because they are not equally established:\n\n- **Broker-first was measured on a live 30-agent deployment** (issue #1578). Upgrading the broker\n first locks the old manager out immediately: `cotal up` re-renders the broker's generated config\n from the trust record, and after the restart the still-0.48.2 manager is refused on every\n connection with an `authentication error` naming the Nkey, continuously. That text comes from the\n broker process, not from a Cotal command, so match on its shape rather than on an exact string.\n `cotal ps` reports zero agents while\n the agent processes are still alive, because the manager has lost its view of them, not because\n they died. Upgrading the manager clears it immediately.\n- **Manager-first is reasoned from where the new stores are provisioned**, not from a measured\n fleet upgrade. It is the recommended order because the manager is the component that creates what\n 0.49.0 adds, but it has not been run end to end on a production split topology at the time of\n writing. Treat it as the better-supported order rather than a guaranteed one, and keep the\n rollback below ready either way.\n\nWhichever order you pick, **this is not a rolling upgrade**. Between the two steps the mesh is down\nand the manager cannot see its agents. Go straight through rather than pausing between them, and\nschedule it as an outage window.\n\n### What the window looks like\n\n- **The managed agent processes do not survive step 1, in either order.** This is the one place\n where the obvious reordering does not rescue you, so it is worth understanding rather than\n working around. Sparing agents on a bare manager stop is a **handshake**: a 0.49.0 manager\n publishes a capability file proving it can release its agents, and a 0.49.0 `cotal down` refuses\n the stop unless it finds one. **A 0.48.2 manager never publishes that file**, because the\n mechanism ships in the release you are installing. So the old CLI against the old manager sends a\n plain stop and takes every seat with it, and the new CLI against the old manager either refuses\n (leaving `--with-agents`, which reaps deliberately) or falls to the legacy path, warns that it\n cannot verify the manager can spare its agents, and signals it anyway.\n- **You can confirm which side you are on in one command, without stopping anything.** The flag that\n marks the newer behaviour is absent from the older CLI, and its summary line makes the difference\n plain:\n\n ```\n $ cotal down --help # on 0.48.2\n cotal down - stop the whole local stack, or name only the components to stop\n\n $ cotal down --help # on 0.49.0\n cotal down - stop the whole local stack (managed agents stay running unless --with-agents), ...\n ```\n\n If your `cotal down --help` does not mention `--with-agents`, stopping the manager stops the\n agents with it.\n- **Therefore the respawn in step 5 is mandatory recovery for this upgrade, not an optional pass.**\n It is also the step that re-mints credentials as issuances, so it is the same action either way.\n Plan the window to include it rather than treating it as cleanup.\n- The **manager's view** of them is lost while the two sides disagree, so `cotal ps` reports zero\n and control commands do not reach seats.\n- **Messages are not delivered** while the mesh is down.\n- The window is as long as it takes to restart the second component, plus the manager's own start.\n It is minutes, not hours, provided you do not stop between the steps.\n- **Nothing self-heals if you stop halfway.** The refusal is continuous until both sides match.\n\n### Snapshot this before you start\n\nTake these while the deployment is still on 0.48.2. The two `cotal` reads are live reads and must\nhappen before anything stops.\n\n- **A filesystem or volume snapshot of both containers**, if your platform offers one. This is the\n only rollback that covers every case, and it is what the reporting deployment used.\n- **`cotal backup create <dir>`**, for the durable space state, **but read the next paragraph before\n you rely on it**: on a split broker and manager topology it is very likely unavailable to you, and\n the volume snapshot above is your actual rollback.\n- **The trust records and credential directory** under `.cotal/auth` on the manager host, including\n the per-space material directory. These are what a re-mint would otherwise have to replace.\n- **A copy of the channel registry**, so you can verify it came back rather than assuming it did:\n `cotal channels list` before and after.\n- **The output of `cotal ps`**, so you know how many seats you expect to see afterwards and can tell\n a lost view from a lost agent.\n\n#### `cotal backup` on a split topology\n\n**`cotal backup create` cannot read a running stack.** It requires a completed cut, and only\n`cotal down --preserve-state` publishes one:\n\n```\n$ cotal backup create ./backup.0482\n✗ backup requires a completed cut; run `cotal down --preserve-state` first\n```\n\n**And `cotal down --preserve-state` requires a manager alive on the host you run it from.** It uses\nthat manager to attest that every retained child stopped, and the check is deliberately fail-closed:\na manager that is dead or merely uncertain refuses rather than preserving an unproven cut. The check\nreads a local pidfile, so a **remote** manager does not satisfy it. On a split topology the broker\nhost has no local manager, which means the documented durable-backup path is not available there.\n\n**Measured rather than assumed, at 0.48.2**: the backup refusal above is executed output. The\npreservation requirement is read from `down.ts` at the same tag, where the preserve path asks a\nmanager to prepare an inventory and then requires that manager to be locally alive before it\ncommits. The part not executed end to end is a genuine two-host split, which needs two real hosts.\n\n**What to do instead.** Use the filesystem or volume snapshot of both containers. That is the\nrollback the reporting deployment actually used, it covers the broker's durable state and the\nmanager's credential material together, and it does not depend on either component being able to\nattest for the other. If you want `cotal backup` as well, take it from a host that does have a live\nlocal manager, and understand it is a second copy rather than the primary rollback.\n\n**This looks like a product limitation rather than a documentation gap**, and it is written here as\none so an operator is not left thinking they mis-typed a command. The upgrade path for the exact\ntopology this page is addressed to cannot use the documented backup command.\n\n### The upgrade end to end\n\n```bash\n# 0. on 0.48.2, STILL RUNNING: record what you expect to see afterwards.\n# These two are live reads, so they must happen before anything stops.\ncotal channels list > channels.before\ncotal ps > ps.before\n\n# 1. manager host. READ THE NOTE BELOW THE BLOCK FIRST: this step ends the\n# managed agent processes whichever order you choose, and the respawn in\n# step 5 is how they come back. It is recovery, not tidying.\n#\n# STOP THE MANAGER WITH THE 0.48.2 CLI, BEFORE INSTALLING 0.49.0. The\n# order matters and it is not recoverable once you install: a 0.49.0\n# `down manager` REFUSES to stop a 0.48.2 manager whose pid record carries\n# a start token, which is every manager on a platform that can read one\n# (Linux can):\n# refusing bare manager stop: ... does not prove this manager can detach\n# its agents; use --with-agents or stop the agents explicitly\n# The refusal names two remedies and NEITHER clears it for this case. The\n# check reads a capability file that only a 0.49.0 manager writes; it never\n# counts agents, so stopping them first changes nothing. And `--with-agents`\n# is whole-stack only, so `down manager --with-agents` is refused by its own\n# flag rule. See #1592.\ncotal down manager # the 0.48.2 CLI, still installed.\n # 0.48.2 has no --with-agents; this\n # is the whole route. On a host that\n # runs the whole stack, the 0.49.0\n # `cotal down --with-agents` after\n # installing is the alternative.\nnpm install -g cotal-ai@0.49.0 # ONLY after the stop above\n# `supervise` RUNS IN THE FOREGROUND and holds the terminal until you stop\n# it. There is no --detach on this command. Start it under whatever keeps\n# your manager alive normally (systemd unit, container entrypoint, or a\n# second terminal), and run the remaining steps from another shell.\ncotal supervise --space <space> --server nats://<broker>:4222\n\n# 2. broker host: stop the stack.\n# NOT `--preserve-state` on a split topology: it needs a manager alive on\n# THIS host to attest its children stopped, and yours is on the other one.\n# Your rollback is the volume snapshot from \"Snapshot this before you\n# start\", not `cotal backup`.\n# See \"cotal backup on a split topology\" above.\ncotal down\n\n# 3. broker host: install 0.49.0 and start it again\nnpm install -g cotal-ai@0.49.0\n# Record the manager log's size BEFORE starting, so step 3a can tell THIS\n# boot's output from every earlier one. It must be captured here, ahead of\n# the start: taken afterwards it sits past the new line and the wait hangs.\n# `<spaceKey>` is NOT the space name. It is lowercase hex of the name's\n# UTF-8 bytes, so space `prod` is `manager.70726f64.log`. Do not guess it:\n# `cotal up` prints the real path on its launch line. Substituting the\n# plain name points at a file that does not exist, and the wait below then\n# burns its full timeout before telling you.\nLOG=.cotal/manager.<spaceKey>.log\nOFF=$( [ -f \"$LOG\" ] && wc -c < \"$LOG\" || echo 0 )\ncotal up --detach --host 0.0.0.0 --space <space> --no-manager\n\n# 3a. SPLIT TOPOLOGY ONLY: `--no-manager` above boots the broker (and the\n# delivery daemon) with NO local manager on the broker host, so there is\n# no wait-and-stop step on a current cotal-ai. The rest of this step is\n# the OLDER-host recipe, kept because the flag is refused there and that\n# refusal is your signal you are on it: without the flag the `up` also\n# starts a local manager, and you must wait for the log to show it is up,\n# then stop it, or you finish the upgrade with two managers and the one\n# you did not intend is the one nobody is watching.\n# A bare `grep -q` does NOT wait: it reads once and exits 1 immediately\n# if the line has not been written yet. Bound the wait instead, so a\n# manager that never comes up fails loudly rather than reading as ready.\n# The log is opened APPEND-ONLY, so on any host that has run a manager\n# before, this file ALREADY carries a `manager up` line from an earlier\n# boot. Grepping the whole file therefore matches instantly and waits for\n# nothing. Read only what THIS boot appended, using the $OFF captured in\n# step 3 above (before the start, which is the only point it is correct):\ntimeout 60 bash -c \\\n \"until tail -c +$((OFF+1)) \\\"$LOG\\\" | grep -q '. manager up'; do sleep 1; done\"\n# exit 0 = THIS boot logged it; exit 124 = it never did, so STOP and look.\n# This manager is 0.49.0 and publishes its own spare-capability file, so\n# the bare stop below is NOT the refusal case from step 1.\ncotal down manager # broker + delivery remain\n# On a current cotal-ai the two commands above are unnecessary (nothing\n# to wait for, nothing to stop) and `cotal down manager` simply reports\n# no manager to stop.\n\n# 4. verify the mesh is whole again before touching the fleet.\n# Do NOT compare `cotal ps` against ps.before yet: step 1 ended the agent\n# processes, so at this point it is EXPECTED to be empty, and an empty\n# `ps` is also the signature of the broker/manager mismatch described\n# above. The two are indistinguishable here, so compare what the mesh\n# itself should have carried across instead:\ncotal channels list # compare against channels.before: this SHOULD match now\ncotal ps # expect it to be EMPTY here; ps.before is the target for\n # step 5, not for this step\n\n# 5. the step that is easy to skip: respawn the managed agents so their\n# credentials are re-minted as issuances and can renew. Persona is a\n# POSITIONAL argument here, unlike `cotal stop`, which requires --name.\n# One call per agent:\ncotal spawn <persona> --detach --name <n> --space <space>\n# then the comparison step 4 could not make:\ncotal ps # NOW compare against ps.before: seat count should match\n```\n\nThe mesh is down from step 2 until step 3 finishes. That is the window. On a split topology there is\nno cut and no backup inside it, so the window is the stop, the install and the restart, nothing more.\n\n## Adding a section for a future release\n\n**Every changeset marked breaking adds a section to this page.** A release that changes what an\noperator must do, in what order, or what stops working, is not finished until the section exists.\n`scripts/upgrade-section-gate.mjs` grades a commit range for this: run it as\n`pnpm upgrade-section-gate --base <ref>` and it reds when the range carries a breaking change and\nadds no new release section. CI runs its self-test and, as a step of the `attribution` job, grades\neach pull request's own range as `HEAD^1..HEAD` over the merge snapshot it checked out. That job is\nthe only context in the branch protection rule set, so a red gate FAILS A REQUIRED CHECK AND BLOCKS\nTHE MERGE. The section is not optional and a reviewer cannot wave it through without an\nadministrator overriding branch protection. Be precise about what the check proves either\nway, because one trusted past its evidence is worse than none. It proves a section for a release\n**was written here**. It cannot prove the section is **correct**, or that it describes the break\nthat actually landed, and it cannot see a breaking change that carries no marker at all. Reviewing\nthe words remains a person's job.\n\n**Mark the break, or the gate cannot see it.** Any one of these is enough, and they are the only\nthings it reads:\n\n- a `!` before the colon in the commit subject, as in `feat(core)!: bind hosted runs to the caller`\n- a `BREAKING CHANGE:` footer in the commit body\n- a changeset in `.changeset/` declaring a `major` bump for any package\n\nThe marker must survive the squash. A `!` that lives only in a commit you squash away is not in the\nrange the gate grades, so put it in the subject that lands on `main`.\n\n**The heading is a `##` and names the release**, like `## From 0.48.2 to 0.49.0`. Both matter, and\nneither is a style preference. Coverage is claimed by a heading, so a heading that names\nno release claims every release and distinguishes none: `## Notes` with a sentence under it would\notherwise satisfy the rule. Naming the release also makes the section the one an operator upgrading\nthat release will search for. Use `###` freely for detail inside a section. Subsections belong to\ntheir release rather than counting as separate coverage.\n\nName the release that first carries the change: the next version Changesets publishes, which\n`pnpm changeset status --verbose` lists. `bin/package.json` on `main` still reads the release already\npublished. If a release is cut while the change is open, the change ships in the release after it,\nso move the heading before merging. The gate accepts any version in a heading, so before merging a\nrelease pull request, check every heading added since the previous tag against the version it\npublishes.\n\nA section is written for the operator, not for the reviewer. It answers, in this order:\n\n1. What keeps working with no action at all.\n2. What does **not** migrate, and when that becomes visible. Name the log line if there is one.\n3. The order to move components in for a split topology, and why that order.\n4. What the outage window looks like, including what survives it.\n5. What to snapshot before starting.\n6. The commands, end to end.\n\n**Where an answer was not measured, say so in the document rather than guessing.** An operator who\nknows which half of a recommendation is reasoned and which is measured can plan around it; one who\nfinds out afterwards cannot.\n"
|
|
273
273
|
},
|
|
274
274
|
{
|
|
275
275
|
"slug": "watch-a-mesh",
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"docs-bundle.generated.js","sourceRoot":"","sources":["../src/docs-bundle.generated.ts"],"names":[],"mappings":"AAKA,yFAAyF;AACzF,MAAM,CAAC,MAAM,YAAY,GAAG,QAAQ,CAAC;AAErC,MAAM,UAAU,cAAc;IAC5B,OAAO;QACP,SAAS,EAAE,QAAQ;QACnB,eAAe,EAAE,mEAAmE;QACpF,OAAO,EAAE;YACP;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,+FAA+F;gBAC1G,MAAM,EAAE,81IAA81I;aACv2I;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,gHAAgH;gBAC3H,MAAM,EAAE,g2XAAg2X;aACz2X;YACD;gBACE,MAAM,EAAE,cAAc;gBACtB,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,upnBAAupnB;aAChqnB;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,kBAAkB;gBAC3B,MAAM,EAAE,mEAAmE;gBAC3E,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,sgqCAAsgqC;aAC/gqC;YACD;gBACE,MAAM,EAAE,0BAA0B;gBAClC,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,mCAAmC;gBAC3C,SAAS,EAAE,gGAAgG;gBAC3G,MAAM,EAAE,6qMAA6qM;aACtrM;YACD;gBACE,MAAM,EAAE,mBAAmB;gBAC3B,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,oDAAoD;gBAC/D,MAAM,EAAE,4+1CAA4+1C;aACr/1C;YACD;gBACE,MAAM,EAAE,aAAa;gBACrB,OAAO,EAAE,aAAa;gBACtB,MAAM,EAAE,yFAAyF;gBACjG,SAAS,EAAE,gJAAgJ;gBAC3J,MAAM,EAAE,gjWAAgjW;aACzjW;YACD;gBACE,MAAM,EAAE,uBAAuB;gBAC/B,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,sFAAsF;gBAC9F,SAAS,EAAE,6GAA6G;gBACxH,MAAM,EAAE,0hTAA0hT;aACniT;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,sMAAsM;gBACjN,MAAM,EAAE,4+SAA4+S;aACr/S;YACD;gBACE,MAAM,EAAE,KAAK;gBACb,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,wGAAwG;gBAChH,SAAS,EAAE,iKAAiK;gBAC5K,MAAM,EAAE,ws9KAAws9K;aACjt9K;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,uHAAuH;gBAC/H,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,05/BAA05/B;aACn6/B;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,gBAAgB;gBACzB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,+EAA+E;gBAC1F,MAAM,EAAE,6pzCAA6pzC;aACtqzC;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,iisBAAiisB;aAC1isB;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,wBAAwB;gBACjC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,kJAAkJ;gBAC7J,MAAM,EAAE,4ifAA4if;aACrjf;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,6CAA6C;gBACxD,MAAM,EAAE,4z2BAA4z2B;aACr02B;YACD;gBACE,MAAM,EAAE,kBAAkB;gBAC1B,OAAO,EAAE,yBAAyB;gBAClC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wJAAwJ;gBACnK,MAAM,EAAE,+3XAA+3X;aACx4X;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,oBAAoB;gBAC7B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,8DAA8D;gBACzE,MAAM,EAAE,g3QAAg3Q;aACz3Q;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,4HAA4H;gBACvI,MAAM,EAAE,6wMAA6wM;aACtxM;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,oIAAoI;gBAC/I,MAAM,EAAE,
|
|
1
|
+
{"version":3,"file":"docs-bundle.generated.js","sourceRoot":"","sources":["../src/docs-bundle.generated.ts"],"names":[],"mappings":"AAKA,yFAAyF;AACzF,MAAM,CAAC,MAAM,YAAY,GAAG,QAAQ,CAAC;AAErC,MAAM,UAAU,cAAc;IAC5B,OAAO;QACP,SAAS,EAAE,QAAQ;QACnB,eAAe,EAAE,mEAAmE;QACpF,OAAO,EAAE;YACP;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,+FAA+F;gBAC1G,MAAM,EAAE,81IAA81I;aACv2I;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,gHAAgH;gBAC3H,MAAM,EAAE,g2XAAg2X;aACz2X;YACD;gBACE,MAAM,EAAE,cAAc;gBACtB,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,upnBAAupnB;aAChqnB;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,kBAAkB;gBAC3B,MAAM,EAAE,mEAAmE;gBAC3E,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,sgqCAAsgqC;aAC/gqC;YACD;gBACE,MAAM,EAAE,0BAA0B;gBAClC,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,mCAAmC;gBAC3C,SAAS,EAAE,gGAAgG;gBAC3G,MAAM,EAAE,6qMAA6qM;aACtrM;YACD;gBACE,MAAM,EAAE,mBAAmB;gBAC3B,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,oDAAoD;gBAC/D,MAAM,EAAE,4+1CAA4+1C;aACr/1C;YACD;gBACE,MAAM,EAAE,aAAa;gBACrB,OAAO,EAAE,aAAa;gBACtB,MAAM,EAAE,yFAAyF;gBACjG,SAAS,EAAE,gJAAgJ;gBAC3J,MAAM,EAAE,gjWAAgjW;aACzjW;YACD;gBACE,MAAM,EAAE,uBAAuB;gBAC/B,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,sFAAsF;gBAC9F,SAAS,EAAE,6GAA6G;gBACxH,MAAM,EAAE,0hTAA0hT;aACniT;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,sMAAsM;gBACjN,MAAM,EAAE,4+SAA4+S;aACr/S;YACD;gBACE,MAAM,EAAE,KAAK;gBACb,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,wGAAwG;gBAChH,SAAS,EAAE,iKAAiK;gBAC5K,MAAM,EAAE,ws9KAAws9K;aACjt9K;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,uHAAuH;gBAC/H,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,05/BAA05/B;aACn6/B;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,gBAAgB;gBACzB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,+EAA+E;gBAC1F,MAAM,EAAE,6pzCAA6pzC;aACtqzC;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,iisBAAiisB;aAC1isB;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,wBAAwB;gBACjC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,kJAAkJ;gBAC7J,MAAM,EAAE,4ifAA4if;aACrjf;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,6CAA6C;gBACxD,MAAM,EAAE,4z2BAA4z2B;aACr02B;YACD;gBACE,MAAM,EAAE,kBAAkB;gBAC1B,OAAO,EAAE,yBAAyB;gBAClC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wJAAwJ;gBACnK,MAAM,EAAE,+3XAA+3X;aACx4X;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,oBAAoB;gBAC7B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,8DAA8D;gBACzE,MAAM,EAAE,g3QAAg3Q;aACz3Q;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,4HAA4H;gBACvI,MAAM,EAAE,6wMAA6wM;aACtxM;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,oIAAoI;gBAC/I,MAAM,EAAE,2qkCAA2qkC;aACprkC;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,uMAAuM;gBAClN,MAAM,EAAE,+4SAA+4S;aACx5S;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,+BAA+B;gBACxC,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,6HAA6H;gBACxI,MAAM,EAAE,irfAAirf;aAC1rf;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,6GAA6G;gBACxH,MAAM,EAAE,q6MAAq6M;aAC96M;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,iBAAiB;gBAC1B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,yEAAyE;gBACpF,MAAM,EAAE,gthEAAgthE;aACzthE;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,6DAA6D;gBACxE,MAAM,EAAE,+6GAA+6G;aACx7G;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,yBAAyB;gBACjC,SAAS,EAAE,wEAAwE;gBACnF,MAAM,EAAE,q5PAAq5P;aAC95P;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,yBAAyB;gBACjC,SAAS,EAAE,+CAA+C;gBAC1D,MAAM,EAAE,i0NAAi0N;aAC10N;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,8BAA8B;gBACvC,MAAM,EAAE,8CAA8C;gBACtD,SAAS,EAAE,qIAAqI;gBAChJ,MAAM,EAAE,o6SAAo6S;aAC76S;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,sDAAsD;gBAC9D,SAAS,EAAE,uJAAuJ;gBAClK,MAAM,EAAE,srZAAsrZ;aAC/rZ;YACD;gBACE,MAAM,EAAE,uBAAuB;gBAC/B,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,0IAA0I;gBACrJ,MAAM,EAAE,0yhBAA0yhB;aACnzhB;YACD;gBACE,MAAM,EAAE,SAAS;gBACjB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,0CAA0C;gBAClD,SAAS,EAAE,gIAAgI;gBAC3I,MAAM,EAAE,i0aAAi0a;aAC10a;YACD;gBACE,MAAM,EAAE,SAAS;gBACjB,OAAO,EAAE,SAAS;gBAClB,MAAM,EAAE,yBAAyB;gBACjC,SAAS,EAAE,mLAAmL;gBAC9L,MAAM,EAAE,y7KAAy7K;aACl8K;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,07+CAA07+C;aACn8+C;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,gBAAgB;gBACzB,MAAM,EAAE,oCAAoC;gBAC5C,SAAS,EAAE,kGAAkG;gBAC7G,MAAM,EAAE,6xbAA6xb;aACtyb;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,oCAAoC;gBAC7C,MAAM,EAAE,0CAA0C;gBAClD,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,4xqBAA4xqB;aACryqB;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,QAAQ;gBACjB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,2DAA2D;gBACtE,MAAM,EAAE,swKAAswK;aAC/wK;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,kGAAkG;gBAC7G,MAAM,EAAE,8uPAA8uP;aACvvP;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,kGAAkG;gBAC7G,MAAM,EAAE,09OAA09O;aACn+O;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,gCAAgC;gBACzC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,iEAAiE;gBAC5E,MAAM,EAAE,s//DAAs//D;aAC///D;YACD;gBACE,MAAM,EAAE,cAAc;gBACtB,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,uHAAuH;gBAClI,MAAM,EAAE,s+qBAAs+qB;aAC/+qB;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,kHAAkH;gBAC7H,MAAM,EAAE,084CAA084C;aACn94C;SACF;QACD,MAAM,EAAE;YACN,OAAO,EAAE,0BAA0B;YACnC,MAAM,EAAE,g7kgBAAg7kgB;SACz7kgB;QACD,MAAM,EAAE;YACN,OAAO,EAAE,mCAAmC;YAC5C,MAAM,EAAE,i9mGAAi9mG;SAC19mG;QACD,QAAQ,EAAE;YACR,OAAO,EAAE,oCAAoC;YAC7C,MAAM,EAAE,m4pBAAm4pB;SAC54pB;KACF,CAAC;AACF,CAAC"}
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@cotal-ai/connector-core",
|
|
3
3
|
"description": "Shared MCP-bridge runtime for Cotal connectors: the mesh agent, cotal_* tools, and hook relay.",
|
|
4
|
-
"version": "0.
|
|
4
|
+
"version": "0.72.0",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"repository": {
|
|
7
7
|
"type": "git",
|
|
@@ -35,7 +35,7 @@
|
|
|
35
35
|
"devDependencies": {
|
|
36
36
|
"@cotal-ai/smoke-kit": "0.0.0",
|
|
37
37
|
"@ag-ui/core": "0.0.57",
|
|
38
|
-
"@cotal-ai/core": "0.
|
|
38
|
+
"@cotal-ai/core": "0.72.0",
|
|
39
39
|
"@nats-io/transport-node": "^3.4.0"
|
|
40
40
|
},
|
|
41
41
|
"files": [
|