@cotal-ai/connector-core 0.72.0 → 0.72.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import type { DocsBundle } from "./docs.js";
|
|
2
2
|
/** The installed Cotal version, for stamping the orientation card and other surfaces. */
|
|
3
|
-
export declare const DOCS_VERSION = "0.72.
|
|
3
|
+
export declare const DOCS_VERSION = "0.72.1";
|
|
4
4
|
export declare function loadDocsBundle(): DocsBundle;
|
|
5
5
|
//# sourceMappingURL=docs-bundle.generated.d.ts.map
|
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
/** The installed Cotal version, for stamping the orientation card and other surfaces. */
|
|
2
|
-
export const DOCS_VERSION = "0.72.
|
|
2
|
+
export const DOCS_VERSION = "0.72.1";
|
|
3
3
|
export function loadDocsBundle() {
|
|
4
4
|
return {
|
|
5
|
-
"version": "0.72.
|
|
5
|
+
"version": "0.72.1",
|
|
6
6
|
"generatedFrom": "docs/*.md + SPEC.md + spec/cotal-lang.md + spec/cotal.schema.json",
|
|
7
7
|
"pages": [
|
|
8
8
|
{
|
|
@@ -87,7 +87,7 @@ export function loadDocsBundle() {
|
|
|
87
87
|
"title": "Connect Claude",
|
|
88
88
|
"kind": "Guide (informative)",
|
|
89
89
|
"summary": "The Claude Code connector turns a real claude session into a Cotal mesh peer.",
|
|
90
|
-
"body": "# Connect Claude\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nThe Claude Code connector turns a real `claude` session into a Cotal mesh peer. A bundled\nplugin inside the session joins NATS, maps lifecycle hooks to presence, and exposes the\nmesh tools. Nothing wraps Claude; it is an ordinary session that happens to be on the\nmesh.\n\nThe shared mesh runtime (agent, `cotal_*` tools, hook relay) lives in\n[`@cotal-ai/connector-core`](../extensions/connector-core); this connector is the thin\nClaude-specific adapter over it. Its lifecycle hook imports the relay from the\n`@cotal-ai/connector-core/relay` subpath, so each hook process loads the relay and its environment\nreaders and none of the NATS client, zod or yaml. Siblings: [OpenCode](connect-opencode.md) (beta),\n[Hermes](connect-hermes.md) (alpha), [pi](connect-pi.md) (alpha); the\n[Connectors](connectors.md) matrix compares them feature-by-feature.\n\n## Set up\n\n```bash\ncotal setup # one-time: installs the plugin, seeds one agent; launches nothing\ncotal up # brings up the mesh + delivery daemon + a detached manager\n```\n\n`cotal setup` installs the cotal plugin (so the repo's Claude sessions get the `cotal_*`\ntools), shares your own MCP servers with spawned sessions on its first run (see\n[Sharing your MCP servers](#sharing-your-mcp-servers)), and seeds one `default` persona; `cotal up` brings up the local stack so\n`cotal spawn --detach` / `cotal_spawn` work right away. Re-running either is idempotent.\nThe install mechanics and the invariants behind them are in\n[setup internals](setup-internals.md).\n\n`cotal setup` also installs Cotal's authored Agent Skills (`SKILL.md`, the agentskills.io format) for\ncoordinating agent teams (today `team-topology`), from one canonical source, on two channels:\n\n- **Claude Code** gets a second, skills-only plugin, `cotal-skills`, from the same `cotal-mesh`\n marketplace, at **user scope** (machine-wide). The Claude connector declares and implements this\n setup provider, including the marketplace assets and native plugin commands; the base CLI only passes\n the vendor-neutral Agent Skills directory. The plugin carries no code and no core dependency,\n and uninstalls on its own with `claude plugin uninstall cotal-skills --scope user`. Its plugin version\n is stamped from the running CLI release, so an upgrade + `cotal setup --skills` runs `claude plugin update` and\n the deployed install actually gets the new skill. `cotal setup` installs it on first run and on repeat\n runs, so upgraders are not left behind. The same provider reports the plugin and skills plugin rows\n in `cotal status`, which point a stale or missing skills plugin at `cotal setup --skills`.\n- **Every other harness** (Codex, Cursor, OpenCode, Gemini CLI, Windsurf/Devin) reads the cross-vendor\n `~/.agents/skills/` directory convention, which has no remote index, so `cotal setup` **reconciles** it\n (and `cotal setup --skills` does only that):\n it installs/updates each Cotal skill, backs up a copy you have edited to `SKILL.md.bak` before\n replacing it, and removes a Cotal skill that is no longer shipped. Only skills Cotal owns are touched;\n your own or third-party skills there are left alone. `cotal status` reports whether the drop is current,\n stale, missing, or has a retired skill to reconcile, and names `cotal setup --skills` as the remedy. This is the working cross-vendor path.\n\nCotal also generates an [Agent Skills discovery index](https://cotal.ai/.well-known/agent-skills/index.json)\non cotal.ai, but that RFC is still a draft with no harness consuming it yet, so it is a forward bet,\nnot a channel to rely on today.\n\n## Spawn a session\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn dave --detach # supervised: the manager runs it in a PTY\n```\n\nA spawn resolves a persona from `.cotal/agents/<name>.md` ([agent files](agent-files.md));\n`--model`, `--variant`, `--cwd`, `--prompt`, ACL overrides, and `--share-tools` apply to\nboth forms ([run a mesh](run-a-mesh.md) has the full resolution rules). The session joins\nwith identity from its environment and auto-registers presence by the time it is\ninteractive.\n\nInside the session, the agent orients with one read-only tool, `cotal_orientation`: its\nidentity, the channels it reads and may post to, its capabilities, the tools available,\nwho's present, and unread counts. The full tool surface is the\n[MCP tool catalog](mcp-tools.md). In auth mode the team-supervision tools\n(`cotal_spawn` / `cotal_persona` / `cotal_personas`) are injected **only** for personas declaring\n`capabilities: [spawn]` (the same grant that opens the privileged control subject), so an\nagent's toolset matches its declared capabilities. `cotal_run` is gated separately by\n`run`; use `capabilities: [spawn, run]` for both. Fresh setup defaults include both.\nSee [workflow tool setup](workflows.md#from-an-agent-session) for a first run and missing-tool checks.\nClearing retained history is\noperator-only ([run a mesh](run-a-mesh.md)), never an agent tool.\n\n## How it binds\n\nClaude Code exposes four integration surfaces, and three of them collapse into a single\ndual-purpose MCP server:\n\n| Surface | Mechanism |\n|---|---|\n| Outbound, ambient | `http` lifecycle hooks → POST to the connector (presence, activity) |\n| Outbound, deliberate | MCP tools `cotal_send` / `cotal_dm` / `cotal_anycast` (+ `cotal_feedback`) |\n| Inbound, pull | MCP tool `cotal_inbox` (same server) |\n| Inbound, push | Channel nudge + hook drain (below) |\n\nThe manager launches the *real* `claude` (no wrapper):\n\n```\nclaude --strict-mcp-config --mcp-config '{\"mcpServers\":{\"cotal\":{…}}}' \\\n --dangerously-load-development-channels server:cotal\n# env: COTAL_SPACE, COTAL_NAME, COTAL_ROLE, COTAL_CHANNEL=1, plus claude's documented auth vars\n```\n\n- **Model auth.** Locally, `claude` still reads macOS Keychain / `~/.claude`. In a container or\n CI there is no Keychain, so the connector forwards the documented credential set:\n `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`), `ANTHROPIC_API_KEY` /\n `ANTHROPIC_AUTH_TOKEN`, and the cloud-provider flags plus their credential vars. Host-session\n markers (`CLAUDE_CODE_CHILD_SESSION`, `CLAUDECODE`) stay out so a nested seat still saves a\n transcript. See [Deploy](deploy.md).\n- **Persona privacy.** The persona body is written to a private file and Claude receives only\n `--append-system-prompt-file <path>`. The body never appears in the spawned process argv. The\n carrier is a 0600 file inside a 0700 directory on POSIX, with equivalent owner-only ACL hardening\n on Windows. That is OS-user isolation: any process running as your user can read it while it\n exists. The manager or the foreground `cotal spawn` removes it, and the shared-server MCP config\n file, once it has proved the `claude` process gone. If the launcher is killed first, a watcher\n started beside `claude` removes them when `claude` exits.\n- **MCP servers.** `--strict-mcp-config` ignores every ambient MCP source, so a spawned agent\n loads the cotal server plus the servers the cotal config shares. First-run `cotal setup`\n fills that list with your own user-scope servers, so a spawned session has the tools you know\n (see below).\n- **Installed plugin.** The plugin is installed once (`claude plugin install\n cotal@cotal-mesh --scope local`) because its hooks bind only to an *installed* plugin.\n The repo's `.claude-plugin/marketplace.json` lists the committed plugin tree under\n `claude-plugin/`, which each release regenerates with the built bundles, the skills and the\n release version, so an install from the repo or from a pinned commit runs without a build\n ([Release](release.md)). `cotal setup` (npx, no clone) materializes the same marketplace under\n `~/.cotal/claude-plugin/` from the installed CLI (each plugin dir is rebuilt from scratch and\n atomically replaced, never merged, so no stale file rides in). The\n `cotal-skills` plugin installs from that same marketplace at user scope (`claude plugin install\n cotal-skills@cotal-mesh --scope user`); its manifest and install behavior ship inside the Claude connector, and\n its version tracks the CLI release so updates land.\n- **Identity-gated.** Connector code requires `COTAL_NAME`, `COTAL_LINK` or `COTAL_AGENT_FILE`.\n A plain `claude` with none of them never joins, so your own sessions in a repo do not appear\n as stray peers. Its MCP server still answers `initialize` and lists one static tool,\n `cotal_how_to_join`, which explains how to launch a session on a mesh. It builds no mesh\n agent, opens no broker connection and binds no control socket.\n- **Hands-free.** The dev-channels flag prints a one-time confirm prompt. The runtime waits for the\n dialog title in normalized terminal text and presses Enter once when it appears, so startup speed\n does not affect a supervised launch. The PTY runtime reads the child's output, and the tmux, cmux,\n Orca and Herdr runtimes read the pane's screen. If the declared prompt never appears within 15\n seconds, the seat ends with a bounded error naming the unmatched prompt instead of hanging\n silently or answering another dialog. The PTY runtime writes that error to the seat's output, and\n the other runtimes write it to the manager's log.\n- **Trusted directory.** Claude opens a directory it has not trusted on its workspace-trust dialog,\n and the dialog's default answer exits. No one is at a supervised seat to answer it, so a launch\n whose directory the manager host's own Claude does not trust is refused before it starts, naming\n the directory and the dialog. Trust is read as Claude reads it: trust given to a parent directory\n counts up to the root of the directory's own Git repository, and a linked worktree shares the trust\n of its repository's main checkout. Open `claude` in that directory on the manager host once and\n trust it, then spawn again. A foreground `cotal spawn` shows the dialog in your own terminal instead.\n\nInbound mesh messages arrive in context as\n`<channel source=\"cotal\" from=\"bob\" kind=\"dm\" …>…</channel>`: each meta key a tag\nattribute the agent can read for routing.\n\n## How messages reach the session\n\nDurable deliveries land in the connector's inbox from JetStream consumers\n([SPEC §8](../SPEC.md#8-nats--jetstream-binding)); live channel traffic can instead arrive\nthrough an at-most-once core subscription. A durable message sent while the agent is busy\nor offline waits on the stream. Two things move a message from inbox to model; one\ndelivers, the other only wakes:\n\n- **Hook drain (delivery).** `SessionStart` / `UserPromptSubmit` hooks read automatic inbox items and\n inject them as `additionalContext`. This is the single authoritative path: deterministic and works\n on any Claude Code build. Quiet ambient is excluded and stays buffered for `cotal_inbox`.\n A message is **acked only once the hook reply carrying it has cleared both legs of its journey**:\n the connector's control socket to the hook process (which gives up after 2s), and the hook\n process's own stdout to Claude Code (which it force-exits 1s after starting to write). The relay\n sends a receipt back down the control socket from that stdout write's callback, and only on a\n clean write (a runtime whose pipe has gone away fails it), and the connector treats that receipt,\n not its own socket write, as delivery. So a large injection killed mid-flush, or one written to a\n broken pipe, leaves the message un-acked and JetStream redelivers it. What this does *not* prove is\n that Claude Code read or applied the reply: a payload small enough to fit the pipe buffer is\n reported written the moment the kernel takes it. That residual is why the path errs toward\n at-least-once rather than treating a confirmed write as a confirmed read. Acking when\n the reply was merely *formatted* meant a lost reply was a lost message: it was already marked\n handled, so its own redelivery was silently acked on arrival.\n A hook whose handler throws still returns an empty reply so the session is never blocked, and\n that reply carries nothing, so it commits nothing: the batch it had started to surface stays\n un-acked and goes out on a later frame. The seat also drops any `turn-pending` row that breaks\n the manager contract, such as one with no integer deadline, and says so once in its log. A reply\n with no `turns` array changes nothing: the seat keeps the turns it already holds.\n This errs toward **at-least-once**: if a reply lands but its confirmation does not, the batch is\n surfaced again and flagged as a possible repeat. A duplicate injection is noise; a buried DM stops\n the peer answering at all.\n- **Channel nudge (wake).** An arriving message fires a `notifications/claude/channel`\n event that wakes an *idle* session into a turn, so the drain runs *now* instead of at\n the next prompt. The nudge never acks anything. A nudge that the host rejects is retried with a\n bounded backoff while anything is still pending. For an idle session it is the only wake source,\n so dropping it means silence until someone types. When the channel becomes active, the connector\n first re-fires a focus mention remembered during startup, otherwise one buffered wake. A rejected\n push keeps its bounded retry, and JetStream redelivery remains the durable backstop for unacked\n inbox items. Neither a redelivery nor that retry repeats a nudge already pushed for that message,\n whether the message had its own nudge or was counted in a batch one, so a session held in a long\n tool call gets one nudge per message. Once a hook frame carries the message, or the push that\n announced it fails, its next redelivery nudges again, so a reply that never reached Claude Code\n still recovers. If the channel cannot run at all, delivery still waits for the next hook. Live-only\n traffic has no durable retry.\n\n**Two priority tiers.** A *directed* message (DM, anycast, or a channel message that\n`@mentions` us) always nudges. *Ambient* channel chatter does not nudge mid-turn; it\naccumulates, and the `Stop` → idle transition fires one batch nudge so the backlog drains\ntogether.\n\n**Constraints (accepted).** Channels are a Claude Code research preview (≥ v2.1.80;\npermission relay ≥ v2.1.81): Anthropic auth only, admin-enabled on Team/Enterprise, and a\ncustom channel needs the `--dangerously-load-development-channels` launch flag. The hook\ndrain does not depend on any of that; the channel only adds \"wake me when idle.\"\n\nThe same channel also relays **tool-permission requests** onto the mesh, so a peer (a\nhuman at the CLI, a policy node) can approve or deny an agent's pending tool call through\nCotal rather than a per-terminal prompt.\n\n### Attention\n\nAn agent picks how aggressively peer traffic reaches it with\n`cotal_status({ attention })` (three modes, orthogonal to presence):\n\n| arrival | open (default) | dnd | focus |\n|---|---|---|---|\n| directed (dm / anycast) | wake + inject | wake + inject | wake + inject |\n| channel `@mention` | wake + inject | wake + inject | ack-drop; wake to *pull*; not injected |\n| ambient channel chatter | wake when idle; hold while working | never wakes; injects next turn | ack-drop; recall via `cotal_inbox` |\n\nPer-channel overrides refine this: **quiet** (delivered, never wakes; `@mention` still\nwakes) and **muted** (dropped on receive, mentions included; DMs/anycast unaffected), set\nwith `cotal_channel_mode` or as agent-file defaults (`quiet:` / `muted:`,\n[agent files](agent-files.md)). A per-channel override is the final word for that channel.\nQuiet ambient is pull-only: it never hitchhikes on a human prompt, DM, mention, or other\nconnector-driven turn. `cotal_inbox` explicitly surfaces and clears it. A quiet-channel\n`@mention` remains automatic and injects normally.\n\nA pull is bounded too, and clears only what it hands over. One `cotal_inbox` call carries at most a\nreceivable window (direct messages and role requests first, then channel traffic, replayed history\nlast); whatever does not fit stays buffered, is named in the reply, and comes back on the next call.\nA message too large for one whole response is delivered in parts: once no smaller mail is waiting,\neach call carries its next part, and it is cleared only after its last part goes out, because clearing\nwhat was not handed over is the loss this bound exists to stop.\nThat matters most on the path where it is easiest to lose mail: reconnecting brings a channel-history\nreplay with it, so the largest payload and the least expendable message arrive in the same read.\n\nThe local inbox is bounded. On pathological overflow it evicts pull-only items first, then other\nchannel traffic, and a direct message or role request only when the whole buffer is directed mail.\nAn evicted channel item is acknowledged. An evicted direct message or role request never is, and\nthe broker redelivers it after the ack wait until a redelivery finds room. A direct message stays\npending on the session's DM durable, where `cotal deliver pending <name>` counts it. A role request\nstays on its role's shared queue, which that command does not read. A full inbox therefore delays\ndirected mail until the session drains it.\nIf the bounded live/durable classification guard also fills, the connector fails closed:\notherwise-normal ambient becomes pull-only until restart. Muted hard-drop and normal focus recall\nstill take precedence. Focus also keeps a bounded exclusion list so mode toggles cannot recall\nquiet/muted traffic; if that safety bound fills, recall skips the affected channel and reports it\nas incomplete rather than risk resurfacing excluded content. Recall cannot tell one message with an\nempty id from an identical one with another disposition, so in focus such a message is held in the\nlocal inbox as pull-only instead of being dropped, and a mention of it still wakes the agent. When\nthe session settles an id-less copy, it reads the chat stream's last sequence. Identical copies\narrive in stream order, and those reads can answer out of order, so the first read that can see the\ncopies binds them in arrival order, latest first, each to the latest unbound stream copy at or below\nthe lowest sequence read for it or any later identical copy.\nWhile the connection stays up, every stream copy at or below that sequence reached the session\nfirst, so a later identical copy sent during a reconnect gap is above it and stays unbound. A copy\nthat arrives while a read runs may not be in that read, so it binds nothing there and recall reads\nthe channel again, up to three times; if it still could be a stream copy in the last read, recall\nleaves that stream copy in the stream and reports the channel as incomplete. A settled copy a\ncomplete read cannot bind is behind the focus start or out of retention, and is forgotten. Recall\nhands back into the inbox only the stream copies nothing is bound to, such as one sent during a\nreconnect gap or one the inbox evicted on overflow, and `cotal_inbox` hands each over once. Overflow\nfrees only the evicted copy, held or quiet, and a copy no read has bound yet keeps its place in\narrival order, so an identical muted copy stays out of recall. When\nthe inbox is full, recall leaves them in the stream for a later call and reports the channel as\nincomplete. A history read that fails, or a channel with replay off, settles nothing and is reported\nas incomplete, and recall calls run one at a time. If the sequence read for a settled copy fails,\nor answers only after the connection dropped, recall skips that channel for the rest of the focus\nperiod and reports it as incomplete. Recall cannot tell a late copy of a message it handed back from\na new identical message, so every copy takes its own disposition: a new identical quiet mention is\nstill delivered automatically, and a late copy can surface a second time.\nA recalled message that already went out in part is read to its last part, even if an exclusion\nlands after its first part. One session reads its inbox one call at a time: a `cotal_inbox` call\nthat overlaps another waits for it to finish, so neither decides from a view the other has already\nmoved past.\nIf the separate hard-drop disposition guard fills, channel traffic is dropped for the rest of the\nsession rather than risk a late copy bypassing an earlier muted/focus decision; DMs and anycast are\nunaffected.\n\nAttention is **advisory UX, not a boundary**: any peer can wake a dnd/focus agent by\nnaming it, and `muted` means \"I opted out of receiving\", not \"the channel is blocked\";\nthe broker still authorizes and delivers. Focus's real effect is shrinking the\nuntrusted-ambient injection surface (only subject-authenticated dm/anycast auto-inject).\nIt resets to **open** on `SessionStart`, so a restarted agent never stays silently deaf.\nYour attention is mirrored into presence so peers can see it.\n\nWhatever does reach a turn is framed so a peer cannot write the frame. A line that begins at column\nzero is written by the connector; one message is one line plus indented continuations, with the\nsender inside a single bracket pair. A message body, a sender name and role, and a service or\nchannel label are all peer-controlled, so each passes through the same neutralization the\n`cotal_inbox` reply uses: no line break a splitter may honour and no bracket survives into a\nrendered attribution. This matters more for an injected block than for a reply, because the agent\ndid not ask for it and so never had the chance to distrust it.\n\n## Presence mapping\n\nThe connector wires a small subset of Claude Code hooks to presence states; presence is\ncoarse, and \"what it is doing\" rides on activity updates. Presence is **advisory**: a presence\npublish that fails (the endpoint mid-reconnect, say) is swallowed and never prevents the same hook\nfrom delivering messages or flushing held ones.\nA `SessionStart` during an open turn, including compaction, preserves the current `working` or\n`waiting` status until `Stop`, `StopFailure`, or `SessionEnd` closes the turn.\n\n| Hook | → state |\n|---|---|\n| `SessionStart` | `idle` only when no turn is open (join; surfaces the inbox; captures the live model into `meta.model` when no pin) |\n| `UserPromptSubmit` | `working` (turn starts; surfaces the inbox) |\n| `PreToolUse` | no change; records *what* is about to run, so a permission wait can name it |\n| `Notification` (`permission_prompt` / `agent_needs_input`) | `waiting` with condition `approval` / `input` (activity leads with the pending tool, e.g. `Bash: git push …`) |\n| `Stop` / `StopFailure` | `idle` (turn done / died on an API error; flushes anything held while busy). `StopFailure` also relays Claude Code's native error value as `condition.source` and maps it to the closed condition vocabulary. On the [event plane](#event-plane) it closes the run with `RUN_ERROR`. |\n| `SessionEnd` | `offline` (graceful leave) |\n\nThe connector also leaves gracefully when its stdin closes. An MCP client closes it to end the\nsession, and a killed `claude` closes it with no `SessionEnd`, so a dead session drops off the\nroster instead of staying on it as a live peer.\n\n`StopFailure` maps `rate_limit` and `overloaded` directly; auth and credential failures to\n`auth`; account and billing failures to `billing`; `invalid_request` to `request`;\n`model_not_found` to `model`; `server_error` to `server`; `max_output_tokens` to `context`; and\n`unknown` to `failed`. The native value remains in `condition.source`.\n\nHooks are relayed over the connector's **authenticated** local control endpoint (per-user\nsocket + per-launch token, constant-time checked), so a local process that finds the path\nstill can't drive presence or stop the agent. The full Claude Code hook-event list lives\nwith the adapter:\n[`extensions/connector-claude-code`](../extensions/connector-claude-code/README.md).\n\n## Event plane\n\nA spawned session publishes a **structured** account of what it\ndid: run boundaries per turn, assistant text, reasoning, and each tool call with its start\nand its end. Not prose about the work, the work itself, in a vocabulary a program can\nread. The launcher sets `COTAL_EVENTS` by default; pass `--no-events` to opt out on an unrestricted\nspace. A user-auth registration with `policy: { events: \"required\" }` carries `eventsRequired` in the\nprivate launch material, so the connector arms even without `COTAL_EVENTS`; `--no-events` is refused.\nA hand-driven user-mode session may carry the same decision as `COTAL_EVENTS_REQUIRED=1`. Its own\npublish grant must cover `events.<owner>.<actor>` or the connector refuses before joining. An unmanaged\nsession with no launch material and no required-policy fallback keeps the generic default behavior.\n\nIf the event plane stops for good, the space's policy decides what happens to the seat, on every\nconnector. On a space that requires events the seat stops and leaves the mesh. On any other space\nit keeps running without events, and the connector log records `AG-UI emitter stopped` with the\nreason. For Claude Code the connector is the MCP server: it leaves the mesh and exits with code 1,\nand its stderr carries that line.\n\nA new session includes its first run even when Claude writes a positional startup prompt before the\nconnector receives `SessionStart`. That from-zero read is keyed only to Claude's explicit\n`source: \"startup\"`; resumed, forked, cleared, and compacted sessions adopt at the transcript boundary\ncaptured at that adopt, before the mesh link connects, so nothing Claude appends while the connector\nis still starting up lands behind the cursor and is silently dropped. Crash recovery follows the\ncursor already stored in the event write-ahead log, regardless of the new process's startup label.\n\nClaude starts each hook in its own process, so a prompt or stop relay can reach Cotal before the\n`SessionStart` relay. The connector holds those event flushes and the terminal until `SessionStart`\nsupplies the source, then enqueues adopt, flush, and close in that order.\n\n`SessionStart` can also run before the connector process has bound its local control socket. The hook\nthe `SessionStart` relay retries only transient pre-connect listener errors, with capped backoff\ninside its existing two-second budget. Later hooks and permanent local faults still fail open\nimmediately. Once a socket has connected, a broken exchange is not retried: the connector may\nalready have handled the frame, so replaying it could apply one lifecycle event twice.\nThat retained `SessionStart` can itself arrive before Claude creates the transcript path. A genuinely\nnew startup waits up to five seconds for that file with capped backoff, and the same deadline bounds\none stalled file read; expiry fails loud instead of silently losing the first run. A forked session\ngets the same wait, because Claude copies the parent transcript into the fork's own file after the\nhook, and then adopts at the end of that copy. Resumed, cleared and compacted starts and recovered\ncursors still require their existing source at once.\n\nTool arguments (`TOOL_CALL_ARGS`) and tool results (`TOOL_CALL_RESULT`) are not republished\nonto this channel. The durable emitter drops those events before they are written to the\nwrite-ahead log, because this channel's read ACL is not the ACL the tool ran under. Content is\nmandatory on both kinds, so the event is suppressed rather than emptied or replaced with a\nplaceholder. Tool start and end still go out. A restart that finds a pending pre-fix frame\nstill carrying those kinds HALTS rather than republishing it.\n\nThe channel is **`events.<owner>.<actor>`**, named after the session's principal. What the actor\nhalf is depends on the mesh, and the difference matters when you go looking for it: on a static mesh\nit is a key the manager allocated, never the display name, so two live agents sharing a display name\ndo not share a stream; on a user-auth mesh it is the agent's own name, because that is what the\nledger row is keyed on. Spelled out again with both halves below. The launch grants publish rights\non that channel alone. A spawn\nthat asks for a *different* agent's event channel is refused at the door rather than granted, since\nthat channel is that session's event stream. The same rule runs on restart: a manager\nresume document that names another agent's event channel is refused rather than adopted, because the\nmanaged row is re-armed from that document and the credential is re-minted from the row.\n\nThe rule reads a **concrete** channel, two principal tokens and nothing else. A pattern such as\n`events.<owner>.>` is not an event channel to it and passes untouched, governed by ordinary ACL\nauthority: on a user mesh the delegation envelope, on a static mesh the spawning credential itself.\nThat is deliberate, because the pattern is the form an operator writes on purpose for an observer,\nand it is worth knowing rather than assuming the fence is total.\n\nTo let something else read a plane, grant it out of band. The refusal prints the command for the\nmesh it is running on, spelled out in full, and only that one.\n\nOn a **user-auth** mesh:\n\n```bash\ncotal actor grant <reader> --owner <owner> --scope '' --allow-subscribe 'events.<owner>.<actor>' --allow-publish ''\n```\n\nEvery field, deliberately. `actor grant` is an upsert of the whole row, so it refuses a grant that\nleaves off any of the three ACL flags. Only `--full` turns an omitted flag into the wide default\n(`>` read, `>` post, `spawn,role:default` scope), which is the opposite of what a scoped watcher is for.\n\nOn a **static** mesh there is no actor ledger for `actor grant` to write to, and the refusal says\nso; mint the reader instead:\n\n```bash\ncotal mint watcher --profile agent --allow-subscribe 'events.<owner>.<actor>' --provision\n```\n\nThe **agent** profile, not the observer one. `mint` reads `--allow-subscribe` only for that\nprofile, and refuses it anywhere else: `--profile observer --allow-subscribe <channel>` exits\nnon-zero and writes no creds file, because the observer profile carries a fixed read set over the\nwhole chat plane, which is the opposite of what a scoped watcher is for. The agent profile also prints the lifecycle uid the\nreader needs, since an authed consuming endpoint refuses to start without one.\n\nOn an **open** mesh there is nothing to grant: the mesh has no credentials and no ACLs, so any peer\nthat lists the channel reads it, and the refusal says so instead of naming a command. The\nown-channel rule still applies there, because a spawn is not the place to hand out a read on\nanother agent's tool inputs and outputs.\n\nTwo things a reader has to do that are not obvious, both on `CotalEndpoint`. It must pass the event\nchannel in `channels`: an endpoint reads the channels it lists, so one constructed without\nthe event channel joins nothing and the frames never arrive. And it reads history with `readHistory(channel)`, the delivery daemon's mediated read, not\n`channelHistory(channel)`: a scoped credential is denied the ad-hoc consumer the direct read\ncreates, by design. `cotal console` and the web console already do both.\n\nThe `<owner>.<actor>` pair is the session's principal. On a user-auth mesh the actor half is the\nagent's own name, so the channel is `events.<your-owner>.<agent-name>`. On a\nstatic mesh the owner half is the literal `local` and the actor is a key the manager allocated, so\nthe channel is `events.local.<key>`; the spawn reply carries that key as `id`. Note\nthat `cotal console` and the web console keep event channels out of their channel lists on purpose,\nsince a plane is a machine feed rather than a conversation; they draw the frames when you open the\nchannel by name.\n\nThe rule governs the manager's doors, which are the ones a caller other than you can reach. A\nforeground `cotal spawn` on your own machine mints from your own signing material, so it can still\ngrant any channel you name: that is the out-of-band grant, not a way around the rule.\n\n**Failed turns publish run errors.** Claude Code decides for itself\nwhether a turn finished or died and fires one of two hooks accordingly, so the connector relays that\ndecision rather than making one of its own: a turn that ended on an API error ends its run with\n`RUN_ERROR` carrying the fixed message `run failed` and no code. Neither the detail Claude Code\nreported nor its error kind is published there: both are upstream values that can echo your prompt or\ntool output, and the events channel has a different read ACL. The error kind still reaches presence\nas the agent's condition (`rate_limit`, `auth`, `billing` and the rest). A turn that ended normally still\nends with a run-finished event carrying no outcome, which says the turn ended and does not claim it\nsucceeded.\n\nEvents are written to a per-session write-ahead log before they are published, so a hook that fires\nafter a restart resumes at the cursor it left rather than replaying or skipping, and a run that was\nopen when the session stopped is closed rather than left dangling.\n\nOne channel carries **every session of one agent**, because it is named after the principal and not\nafter the session. Alongside the per-session logs the connector keeps one small record per principal,\nholding the last sequence the broker assigned on that channel, so a new session continues the stream\nits predecessor left instead of starting again from nothing. Both live under the events state root\n(`COTAL_WORKSPACE_ROOT`), and neither is something you edit by hand.\n\nA **missing** record is not a fault: the connector rebuilds it from the session logs beside it,\nwhich is how an agent that was already running before this record existed keeps its stream. That\nrebuild stops if any one of those session logs is damaged. Unreadable, not valid JSON, and written\nfor a different principal all count, and so does a session directory or a log that is a link rather\nthan the real file the connector wrote, or a log that has more than one name. A tip taken from the\nrest would be too low, and it would stop publication later with nothing left to point at the cause.\nThe connector names the file instead, and the only way past it is the directory removal described\nbelow, under the same condition. A record that **disagrees with the broker** is a fault, and the\nconnector stops publishing and says why rather than guessing. A record that **moved while a session\nwas writing to it** is refused the same way: it means something else wrote the principal's record,\nand the connector reports which value it held and which the file holds rather than writing over the\nlater one. There is no command to clear it. The state is the principal's directory under the events\nroot, and clearing it by hand means removing that directory whole: the sequence, the cursor and the\nper-session logs only mean anything together, so removing part of it leaves a state the next start\nrefuses. Removing it is only half a remedy, and the half that comes first is the channel. The\ndirectory is where the agent's memory of the tip lives, not the tip itself, so on a channel that\nstill holds frames the next session opens expecting an empty one and stops on the same\ndisagreement, with the logs a tip could have been rebuilt from now gone. Purge the channel first,\nthen remove the directory.\n\nReading it: `cotal console` and the web console draw event frames directly. A frame carries no text\npart by design, so a surface that renders a message as flat text shows a marker instead of prose.\n\n**On a per-user-auth mesh, the default event plane needs the spawner's grant to cover the channel.** The event\nchannel is added to the child's publish set, and delegation only narrows: an agent may hand down\na subset of what it holds and no more. So a peer-initiated spawn is refused unless the\nspawning identity's own grant already covers the child's event channel. The refusal prints the\nexact `cotal actor grant` command that widens it. An operator launch, whose chain reaches an\nadmin-scoped or roster row, is unaffected. Passing `events: false` is the explicit opt-out.\n\nArming the event plane through a typed spawn request (`manager.spawn` with `events`, including\nthe CLI's `cotal spawn --detach --events`) additionally requires the caller's admin tier on a\nuser mesh. A non-admin caller that asks for the plane is refused before anything is provisioned,\nand one that stays silent gets a spawn without it, with the reply saying so.\n\n## Resume a session\n\n`--resume <session-id>` pulls an existing Claude session, its context and transcript,\ninto the mesh. It **forks**: Claude mints a *new* session id from that transcript\n(`--resume <id> --fork-session`), so the meshed agent gets its own session and the\noriginal is untouched.\n\n- `cotal spawn --resume <id>` (foreground) is the primary surface: the transcript is on\n *your* machine, and errors are Claude's own stderr, inline.\n- `--detach --resume <id> --on <instance>` carries a session held on *your* machine to that\n manager instance, which may run on another host. The CLI finds the transcript under your\n Claude config (`~/.claude`, or `$CLAUDE_CONFIG_DIR`), sends it through a JetStream Object\n Store bucket only that instance reads, under a writer credential pinned to that one transcript,\n and prints `carried session <id> to <instance>:\n sha256:<hex>, <sent> of <size> bytes sent in <chunks> chunks`. A re-run of the same bytes\n sends nothing, and an interrupted carry continues where it stopped. The seat forks it in a\n private Claude home under the manager's `.cotal/seat-homes/`, which no other seat's Claude\n lists or finds, and which is removed when the seat stops. When Claude starts the fork, the seat\n records the SHA-256 of the transcript it read from its own project; the manager stops a seat\n whose record names other bytes than the carried ones, or that records none within the join\n timeout after it joins, an uncertain launch included, and otherwise shows that record as the\n seat's provenance. `cotal attach` to such a seat names its source after the seat\n name, as `(resumed from <host>:<id>)`. A remote manager receives a carry when its host issues it a\n transfer reader. On a user-auth mesh the CLI exchanges the operator's login for a one-object\n `transfer-writer` view, which needs scope `admin`.\n- A session name in place of an id is refused, listing each session on this host that carries\n that name with its id, SHA-256 and modification time. An id this host does not hold resolves\n against the **manager host's** `~/.claude`, as before.\n- A seat-private home holds no login. The manager host needs `CLAUDE_CODE_OAUTH_TOKEN` (from\n `claude setup-token`), `ANTHROPIC_AUTH_TOKEN`, or a cloud provider selection in its\n environment; `ANTHROPIC_API_KEY` alone is refused. The launch directory must already be\n trusted by the manager host's own Claude, and Claude must be 2.1.234 or later.\n- The manager waits for a real outcome: `✓ started` means the agent *joined the mesh*,\n `✗ exited on launch` carries Claude's last output, and an uncertain launch (~30 s) is\n reported without tearing the agent down.\n- Resume is an **operator surface only**, deliberately not exposed on MCP `cotal_spawn`\n (a mesh peer naming host-local transcripts would widen `spawn` into transcript\n disclosure). Only the Claude connector supports it today; OpenCode and Hermes fail loud.\n- Needs a `claude` new enough for `--resume … --fork-session` (verified on 2.1.197).\n\n## Sharing your MCP servers\n\nA spawned session keeps your own MCP servers by default. On its first run, `cotal setup` copies\nthe user-scope servers from your Claude Code config (`~/.claude.json`, or the one under\n`$CLAUDE_CONFIG_DIR`) into the cotal config file (`~/.config/cotal/config.json`) under\n`connectors.claude.mcpServers`, and names them in its output. With none to copy it writes an\nempty list. Each entry is the familiar `.mcp.json` shape ([full format](config.md)). A cotal\nconfig that already declares that list keeps it, and a later `cotal setup` never changes it.\n\nThe cotal config holds secrets only as `${VAR}` references. Setup cannot tell literal text from\na secret, so it leaves out a server with an `env` or `headers` value that is anything but `${VAR}`\nreferences (a `Bearer ${TOKEN}` header among them) and names it in its output. To share one,\nadd it to the cotal config with each secret written as a `${VAR}` reference, and export that\nvariable where you spawn. Setup also leaves out and names an entry no session can start, such as\none with a missing or empty `command` or `url`, or one whose `command` is not a string.\n\nAt launch the connector forwards *only* the named vars the chosen servers declare and\npasses the merged config as an owner-only temp file; `--strict-mcp-config` stays on, so\nonly cotal + the shared servers load.\n\nFor a lighter seat, share fewer. Remove an entry from the cotal config to drop it from every\nspawn, or scope one spawn with `--share-tools tavily,figma` (or `--share-tools none` for cotal\nalone). An empty list (`\"mcpServers\": {}`) in `~/.config/cotal/config.json` keeps every spawn\nisolated, and setup leaves it as it is.\n\nTwo caveats: sharing a server grants its credential to the agent (the var lives in the\nClaude process's environment, so share only when you're fine with that teammate holding\nthe key), and memory adds up, because a heavy server boots once per spawn, multiplied\nacross a team, and can starve a small machine.\n\n## Feedback\n\n`cotal_feedback` works out of the box: without a key it posts to the public intake at\n`https://cotal.ai/v1/feedback` (needs a contact email: `COTAL_FEEDBACK_EMAIL`, then\n`git config user.email`, else the agent asks). Set `COTAL_FEEDBACK_KEY=fbk_<key>` in a\nbeta tester's environment to route to the keyed intake (`Authorization: Bearer`, identity\nderived from the key); `COTAL_FEEDBACK_URL` overrides either endpoint. The CLI can send\ntoo: `cotal feedback \"<summary>\" [--type bug]`. Each submission carries\n`origin: human | agent`, whether the tester asked, or the agent auto-reported a major\nissue.\n"
|
|
90
|
+
"body": "# Connect Claude\n\n> **Guide** (informative) · **For:** operators · **Prereqs:** [Quickstart](getting-started.md)\n\nThe Claude Code connector turns a real `claude` session into a Cotal mesh peer. A bundled\nplugin inside the session joins NATS, maps lifecycle hooks to presence, and exposes the\nmesh tools. Nothing wraps Claude; it is an ordinary session that happens to be on the\nmesh.\n\nThe shared mesh runtime (agent, `cotal_*` tools, hook relay) lives in\n[`@cotal-ai/connector-core`](../extensions/connector-core); this connector is the thin\nClaude-specific adapter over it. Its lifecycle hook imports the relay from the\n`@cotal-ai/connector-core/relay` subpath, so each hook process loads the relay and its environment\nreaders and none of the NATS client, zod or yaml. Siblings: [OpenCode](connect-opencode.md) (beta),\n[Hermes](connect-hermes.md) (alpha), [pi](connect-pi.md) (alpha); the\n[Connectors](connectors.md) matrix compares them feature-by-feature.\n\n## Set up\n\n```bash\ncotal setup # one-time: installs the plugin, seeds one agent; launches nothing\ncotal up # brings up the mesh + delivery daemon + a detached manager\n```\n\n`cotal setup` installs the cotal plugin (so the repo's Claude sessions get the `cotal_*`\ntools), shares your own MCP servers with spawned sessions on its first run (see\n[Sharing your MCP servers](#sharing-your-mcp-servers)), and seeds one `default` persona; `cotal up` brings up the local stack so\n`cotal spawn --detach` / `cotal_spawn` work right away. Re-running either is idempotent.\nThe install mechanics and the invariants behind them are in\n[setup internals](setup-internals.md).\n\n`cotal setup` also installs Cotal's authored Agent Skills (`SKILL.md`, the agentskills.io format) for\ncoordinating agent teams (today `team-topology`), from one canonical source, on two channels:\n\n- **Claude Code** gets a second, skills-only plugin, `cotal-skills`, from the same `cotal-mesh`\n marketplace, at **user scope** (machine-wide). The Claude connector declares and implements this\n setup provider, including the marketplace assets and native plugin commands; the base CLI only passes\n the vendor-neutral Agent Skills directory. The plugin carries no code and no core dependency,\n and uninstalls on its own with `claude plugin uninstall cotal-skills --scope user`. Its plugin version\n is stamped from the running CLI release, so an upgrade + `cotal setup --skills` runs `claude plugin update` and\n the deployed install actually gets the new skill. `cotal setup` installs it on first run and on repeat\n runs, so upgraders are not left behind. The same provider reports the plugin and skills plugin rows\n in `cotal status`, which point a stale or missing skills plugin at `cotal setup --skills`.\n- **Every other harness** (Codex, Cursor, OpenCode, Gemini CLI, Windsurf/Devin) reads the cross-vendor\n `~/.agents/skills/` directory convention, which has no remote index, so `cotal setup` **reconciles** it\n (and `cotal setup --skills` does only that):\n it installs/updates each Cotal skill, backs up a copy you have edited to `SKILL.md.bak` before\n replacing it, and removes a Cotal skill that is no longer shipped. Only skills Cotal owns are touched;\n your own or third-party skills there are left alone. `cotal status` reports whether the drop is current,\n stale, missing, or has a retired skill to reconcile, and names `cotal setup --skills` as the remedy. This is the working cross-vendor path.\n\nCotal also generates an [Agent Skills discovery index](https://cotal.ai/.well-known/agent-skills/index.json)\non cotal.ai, but that RFC is still a draft with no harness consuming it yet, so it is a forward bet,\nnot a channel to rely on today.\n\n## Spawn a session\n\n```bash\ncotal spawn # foreground: your default agent, in this terminal\ncotal spawn dave --detach # supervised: the manager runs it in a PTY\n```\n\nA spawn resolves a persona from `.cotal/agents/<name>.md` ([agent files](agent-files.md));\n`--model`, `--variant`, `--cwd`, `--prompt`, ACL overrides, and `--share-tools` apply to\nboth forms ([run a mesh](run-a-mesh.md) has the full resolution rules). The session joins\nwith identity from its environment and auto-registers presence by the time it is\ninteractive.\n\nInside the session, the agent orients with one read-only tool, `cotal_orientation`: its\nidentity, the channels it reads and may post to, its capabilities, the tools available,\nwho's present, and unread counts. The full tool surface is the\n[MCP tool catalog](mcp-tools.md). In auth mode the team-supervision tools\n(`cotal_spawn` / `cotal_persona` / `cotal_personas`) are injected **only** for personas declaring\n`capabilities: [spawn]` (the same grant that opens the privileged control subject), so an\nagent's toolset matches its declared capabilities. `cotal_run` is gated separately by\n`run`; use `capabilities: [spawn, run]` for both. Fresh setup defaults include both.\nSee [workflow tool setup](workflows.md#from-an-agent-session) for a first run and missing-tool checks.\nClearing retained history is\noperator-only ([run a mesh](run-a-mesh.md)), never an agent tool.\n\n## How it binds\n\nClaude Code exposes four integration surfaces, and three of them collapse into a single\ndual-purpose MCP server:\n\n| Surface | Mechanism |\n|---|---|\n| Outbound, ambient | `http` lifecycle hooks → POST to the connector (presence, activity) |\n| Outbound, deliberate | MCP tools `cotal_send` / `cotal_dm` / `cotal_anycast` (+ `cotal_feedback`) |\n| Inbound, pull | MCP tool `cotal_inbox` (same server) |\n| Inbound, push | Channel nudge + hook drain (below) |\n\nThe manager launches the *real* `claude` (no wrapper):\n\n```\nclaude --strict-mcp-config --mcp-config '{\"mcpServers\":{\"cotal\":{…}}}' \\\n --dangerously-load-development-channels server:cotal\n# env: COTAL_SPACE, COTAL_NAME, COTAL_ROLE, COTAL_CHANNEL=1, plus claude's documented auth vars\n```\n\n- **Model auth.** Locally, `claude` still reads macOS Keychain / `~/.claude`. In a container or\n CI there is no Keychain, so the connector forwards the documented credential set:\n `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`), `ANTHROPIC_API_KEY` /\n `ANTHROPIC_AUTH_TOKEN`, and the cloud-provider flags plus their credential vars. Host-session\n markers (`CLAUDE_CODE_CHILD_SESSION`, `CLAUDECODE`) stay out so a nested seat still saves a\n transcript. See [Deploy](deploy.md).\n- **Persona privacy.** The persona body is written to a private file and Claude receives only\n `--append-system-prompt-file <path>`. The body never appears in the spawned process argv. The\n carrier is a 0600 file inside a 0700 directory on POSIX, with equivalent owner-only ACL hardening\n on Windows. That is OS-user isolation: any process running as your user can read it while it\n exists. The manager or the foreground `cotal spawn` removes it, and the shared-server MCP config\n file, once it has proved the `claude` process gone. If the launcher is killed first, a watcher\n started beside `claude` removes them when `claude` exits.\n- **MCP servers.** `--strict-mcp-config` ignores every ambient MCP source, so a spawned agent\n loads the cotal server plus the servers the cotal config shares. First-run `cotal setup`\n fills that list with your own user-scope servers, so a spawned session has the tools you know\n (see below).\n- **Installed plugin.** The plugin is installed once (`claude plugin install\n cotal@cotal-mesh --scope local`) because its hooks bind only to an *installed* plugin.\n The repo's `.claude-plugin/marketplace.json` lists the committed plugin tree under\n `claude-plugin/`, which each release regenerates with the built bundles, the skills and the\n release version, so an install from the repo or from a pinned commit runs without a build\n ([Release](release.md)). `cotal setup` (npx, no clone) materializes the same marketplace under\n `~/.cotal/claude-plugin/` from the installed CLI (each plugin dir is rebuilt from scratch and\n atomically replaced, never merged, so no stale file rides in). The\n `cotal-skills` plugin installs from that same marketplace at user scope (`claude plugin install\n cotal-skills@cotal-mesh --scope user`); its manifest and install behavior ship inside the Claude connector, and\n its version tracks the CLI release so updates land.\n- **Identity-gated.** Connector code requires `COTAL_NAME`, `COTAL_LINK` or `COTAL_AGENT_FILE`.\n A plain `claude` with none of them never joins, so your own sessions in a repo do not appear\n as stray peers. Its MCP server still answers `initialize` and lists one static tool,\n `cotal_how_to_join`, which explains how to launch a session on a mesh. It builds no mesh\n agent, opens no broker connection and binds no control socket.\n- **Hands-free.** The dev-channels flag prints a one-time confirm prompt. The runtime waits for the\n dialog title in normalized terminal text and presses Enter once when it appears, so startup speed\n does not affect a supervised launch. The PTY runtime reads the child's output, and the tmux, cmux,\n Orca and Herdr runtimes read the pane's screen. If the declared prompt never appears within 15\n seconds, the seat ends with a bounded error naming the unmatched prompt instead of hanging\n silently or answering another dialog. The PTY runtime writes that error to the seat's output, and\n the other runtimes write it to the manager's log.\n- **Trusted directory.** Claude opens a directory it has not trusted on its workspace-trust dialog,\n and the dialog's default answer exits. No one is at a supervised seat to answer it, so a launch\n whose directory the manager host's own Claude does not trust is refused before it starts, naming\n the directory and the dialog. Trust is read as Claude reads it: trust given to a parent directory\n counts up to the root of the directory's own Git repository, and a linked worktree shares the trust\n of its repository's main checkout. Open `claude` in that directory on the manager host once and\n trust it, then spawn again. A foreground `cotal spawn` shows the dialog in your own terminal instead.\n\nInbound mesh messages arrive in context as\n`<channel source=\"cotal\" from=\"bob\" kind=\"dm\" …>…</channel>`: each meta key a tag\nattribute the agent can read for routing.\n\n## How messages reach the session\n\nDurable deliveries land in the connector's inbox from JetStream consumers\n([SPEC §8](../SPEC.md#8-nats--jetstream-binding)); live channel traffic can instead arrive\nthrough an at-most-once core subscription. A durable message sent while the agent is busy\nor offline waits on the stream. Two things move a message from inbox to model; one\ndelivers, the other only wakes:\n\n- **Hook drain (delivery).** `SessionStart` / `UserPromptSubmit` hooks read automatic inbox items and\n inject them as `additionalContext`. This is the single authoritative path: deterministic and works\n on any Claude Code build. Quiet ambient is excluded and stays buffered for `cotal_inbox`.\n A message is **acked only once the hook reply carrying it has cleared both legs of its journey**:\n the connector's control socket to the hook process (which gives up after 2s), and the hook\n process's own stdout to Claude Code (which it force-exits 1s after starting to write). The relay\n sends a receipt back down the control socket from that stdout write's callback, and only on a\n clean write (a runtime whose pipe has gone away fails it), and the connector treats that receipt,\n not its own socket write, as delivery. So a large injection killed mid-flush, or one written to a\n broken pipe, leaves the message un-acked and JetStream redelivers it. What this does *not* prove is\n that Claude Code read or applied the reply: a payload small enough to fit the pipe buffer is\n reported written the moment the kernel takes it. That residual is why the path errs toward\n at-least-once rather than treating a confirmed write as a confirmed read. Acking when\n the reply was merely *formatted* meant a lost reply was a lost message: it was already marked\n handled, so its own redelivery was silently acked on arrival.\n A hook whose handler throws still returns an empty reply so the session is never blocked, and\n that reply carries nothing, so it commits nothing: the batch it had started to surface stays\n un-acked and goes out on a later frame. The seat also drops any `turn-pending` row that breaks\n the manager contract, such as one with no integer deadline, and says so once in its log. A reply\n with no `turns` array changes nothing: the seat keeps the turns it already holds.\n This errs toward **at-least-once**: if a reply lands but its confirmation does not, the batch is\n surfaced again and flagged as a possible repeat. A duplicate injection is noise; a buried DM stops\n the peer answering at all.\n- **Channel nudge (wake).** An arriving message fires a `notifications/claude/channel`\n event that wakes an *idle* session into a turn, so the drain runs *now* instead of at\n the next prompt. The nudge never acks anything. A nudge that the host rejects is retried with a\n bounded backoff while anything is still pending. For an idle session it is the only wake source,\n so dropping it means silence until someone types. When the channel becomes active, the connector\n first re-fires a focus mention remembered during startup, otherwise one buffered wake. A rejected\n push keeps its bounded retry, and JetStream redelivery remains the durable backstop for unacked\n inbox items. Neither a redelivery nor that retry repeats a nudge already pushed for that message,\n whether the message had its own nudge or was counted in a batch one, so a session held in a long\n tool call gets one nudge per message. Once a hook frame carries the message, or the push that\n announced it fails, its next redelivery nudges again, so a reply that never reached Claude Code\n still recovers. If the channel cannot run at all, delivery still waits for the next hook. Live-only\n traffic has no durable retry.\n\n**Two priority tiers.** A *directed* message (DM, anycast, or a channel message that\n`@mentions` us) always nudges. *Ambient* channel chatter does not nudge mid-turn; it\naccumulates, and the `Stop` → idle transition fires one batch nudge so the backlog drains\ntogether.\n\n**Constraints (accepted).** Channels are a Claude Code research preview (≥ v2.1.80;\npermission relay ≥ v2.1.81): Anthropic auth only, admin-enabled on Team/Enterprise, and a\ncustom channel needs the `--dangerously-load-development-channels` launch flag. The hook\ndrain does not depend on any of that; the channel only adds \"wake me when idle.\"\n\nThe same channel also relays **tool-permission requests** onto the mesh, so a peer (a\nhuman at the CLI, a policy node) can approve or deny an agent's pending tool call through\nCotal rather than a per-terminal prompt.\n\n### Attention\n\nAn agent picks how aggressively peer traffic reaches it with\n`cotal_status({ attention })` (three modes, orthogonal to presence):\n\n| arrival | open (default) | dnd | focus |\n|---|---|---|---|\n| directed (dm / anycast) | wake + inject | wake + inject | wake + inject |\n| channel `@mention` | wake + inject | wake + inject | ack-drop; wake to *pull*; not injected |\n| ambient channel chatter | wake when idle; hold while working | never wakes; injects next turn | ack-drop; recall via `cotal_inbox` |\n\nPer-channel overrides refine this: **quiet** (delivered, never wakes; `@mention` still\nwakes) and **muted** (dropped on receive, mentions included; DMs/anycast unaffected), set\nwith `cotal_channel_mode` or as agent-file defaults (`quiet:` / `muted:`,\n[agent files](agent-files.md)). A per-channel override is the final word for that channel.\nQuiet ambient is pull-only: it never hitchhikes on a human prompt, DM, mention, or other\nconnector-driven turn. `cotal_inbox` explicitly surfaces and clears it. A quiet-channel\n`@mention` remains automatic and injects normally.\n\nA pull is bounded too, and clears only what it hands over. One `cotal_inbox` call carries at most a\nreceivable window (direct messages and role requests first, then channel traffic, replayed history\nlast); whatever does not fit stays buffered, is named in the reply, and comes back on the next call.\nA message too large for one whole response is delivered in parts: once no smaller mail is waiting,\neach call carries its next part, and it is cleared only after its last part goes out, because clearing\nwhat was not handed over is the loss this bound exists to stop.\nThat matters most on the path where it is easiest to lose mail: reconnecting brings a channel-history\nreplay with it, so the largest payload and the least expendable message arrive in the same read.\n\nThe local inbox is bounded. On pathological overflow it evicts pull-only items first, then other\nchannel traffic, and a direct message or role request only when the whole buffer is directed mail.\nAn evicted channel item is acknowledged. An evicted direct message or role request never is, and\nthe broker redelivers it after the ack wait until a redelivery finds room. A direct message stays\npending on the session's DM durable, where `cotal deliver pending <name>` counts it. A role request\nstays on its role's shared queue, which that command does not read. A full inbox therefore delays\ndirected mail until the session drains it.\nIf the bounded live/durable classification guard also fills, the connector fails closed:\notherwise-normal ambient becomes pull-only until restart. Muted hard-drop and normal focus recall\nstill take precedence. Focus also keeps a bounded exclusion list so mode toggles cannot recall\nquiet/muted traffic; if that safety bound fills, recall skips the affected channel and reports it\nas incomplete rather than risk resurfacing excluded content. Recall cannot tell one message with an\nempty id from an identical one with another disposition, so in focus such a message is held in the\nlocal inbox as pull-only instead of being dropped, and a mention of it still wakes the agent. When\nthe session settles an id-less copy, it reads the chat stream's last sequence. Identical copies\narrive in stream order, and those reads can answer out of order, so the first read that can see the\ncopies binds them in arrival order, latest first, each to the latest unbound stream copy at or below\nthe lowest sequence read for it or any later identical copy.\nWhile the connection stays up, every stream copy at or below that sequence reached the session\nfirst, so a later identical copy sent during a reconnect gap is above it and stays unbound. A copy\nthat arrives while a read runs may not be in that read, so it binds nothing there and recall reads\nthe channel again, up to three times; if it still could be a stream copy in the last read, recall\nleaves that stream copy in the stream and reports the channel as incomplete. A settled copy a\ncomplete read cannot bind is behind the focus start or out of retention, and is forgotten. Recall\nhands back into the inbox only the stream copies nothing is bound to, such as one sent during a\nreconnect gap or one the inbox evicted on overflow, and `cotal_inbox` hands each over once. Overflow\nfrees only the evicted copy, held or quiet, and a copy no read has bound yet keeps its place in\narrival order, so an identical muted copy stays out of recall. When\nthe inbox is full, recall leaves them in the stream for a later call and reports the channel as\nincomplete. A history read that fails, or a channel with replay off, settles nothing and is reported\nas incomplete, and recall calls run one at a time. If the sequence read for a settled copy fails,\nor answers only after the connection dropped, recall skips that channel for the rest of the focus\nperiod and reports it as incomplete. Recall cannot tell a late copy of a message it handed back from\na new identical message, so every copy takes its own disposition: a new identical quiet mention is\nstill delivered automatically, and a late copy can surface a second time.\nA recalled message that already went out in part is read to its last part, even if an exclusion\nlands after its first part. One session reads its inbox one call at a time: a `cotal_inbox` call\nthat overlaps another waits for it to finish, so neither decides from a view the other has already\nmoved past.\nIf the separate hard-drop disposition guard fills, channel traffic is dropped for the rest of the\nsession rather than risk a late copy bypassing an earlier muted/focus decision; DMs and anycast are\nunaffected.\n\nAttention is **advisory UX, not a boundary**: any peer can wake a dnd/focus agent by\nnaming it, and `muted` means \"I opted out of receiving\", not \"the channel is blocked\";\nthe broker still authorizes and delivers. Focus's real effect is shrinking the\nuntrusted-ambient injection surface (only subject-authenticated dm/anycast auto-inject).\nIt resets to **open** on `SessionStart`, so a restarted agent never stays silently deaf.\nYour attention is mirrored into presence so peers can see it.\n\nWhatever does reach a turn is framed so a peer cannot write the frame. A line that begins at column\nzero is written by the connector; one message is one line plus indented continuations, with the\nsender inside a single bracket pair. A message body, a sender name and role, and a service or\nchannel label are all peer-controlled, so each passes through the same neutralization the\n`cotal_inbox` reply uses: no line break a splitter may honour and no bracket survives into a\nrendered attribution. This matters more for an injected block than for a reply, because the agent\ndid not ask for it and so never had the chance to distrust it.\n\n## Presence mapping\n\nThe connector wires a small subset of Claude Code hooks to presence states; presence is\ncoarse, and \"what it is doing\" rides on activity updates. Presence is **advisory**: a presence\npublish that fails (the endpoint mid-reconnect, say) is swallowed and never prevents the same hook\nfrom delivering messages or flushing held ones.\nA `SessionStart` during an open turn, including compaction, preserves the current `working` or\n`waiting` status until `Stop`, `StopFailure`, or `SessionEnd` closes the turn.\n\n| Hook | → state |\n|---|---|\n| `SessionStart` | `idle` only when no turn is open (join; surfaces the inbox; captures the live model into `meta.model` when no pin) |\n| `UserPromptSubmit` | `working` (turn starts; surfaces the inbox) |\n| `PreToolUse` | no change; records *what* is about to run, so a permission wait can name it |\n| `Notification` (`permission_prompt` / `agent_needs_input`) | `waiting` with condition `approval` / `input` (activity leads with the pending tool, e.g. `Bash: git push …`) |\n| `Stop` / `StopFailure` | `idle` (turn done / died on an API error; flushes anything held while busy). `StopFailure` also relays Claude Code's native error value as `condition.source` and maps it to the closed condition vocabulary. On the [event plane](#event-plane) it closes the run with `RUN_ERROR`. |\n| `SessionEnd` | `offline` (graceful leave) |\n\nThe connector also leaves gracefully when its stdin closes. An MCP client closes it to end the\nsession, and a killed `claude` closes it with no `SessionEnd`, so a dead session drops off the\nroster instead of staying on it as a live peer.\n\n`StopFailure` maps `rate_limit` and `overloaded` directly; auth and credential failures to\n`auth`; account and billing failures to `billing`; `invalid_request` to `request`;\n`model_not_found` to `model`; `server_error` to `server`; `max_output_tokens` to `context`; and\n`unknown` to `failed`. The native value remains in `condition.source`.\n\nHooks are relayed over the connector's **authenticated** local control endpoint (per-user\nsocket + per-launch token, constant-time checked), so a local process that finds the path\nstill can't drive presence or stop the agent. The full Claude Code hook-event list lives\nwith the adapter:\n[`extensions/connector-claude-code`](../extensions/connector-claude-code/README.md).\n\n## Event plane\n\nA spawned session publishes a **structured** account of what it\ndid: run boundaries per turn, assistant text, reasoning, and each tool call with its start\nand its end. Not prose about the work, the work itself, in a vocabulary a program can\nread. The launcher sets `COTAL_EVENTS` by default; pass `--no-events` to opt out on an unrestricted\nspace. A user-auth registration with `policy: { events: \"required\" }` carries `eventsRequired` in the\nprivate launch material, so the connector arms even without `COTAL_EVENTS`; `--no-events` is refused.\nA hand-driven user-mode session may carry the same decision as `COTAL_EVENTS_REQUIRED=1`. Its own\npublish grant must cover `events.<owner>.<actor>` or the connector refuses before joining. An unmanaged\nsession with no launch material and no required-policy fallback keeps the generic default behavior.\n\nIf the event plane stops for good, the space's policy decides what happens to the seat, on every\nconnector. On a space that requires events the seat stops and leaves the mesh. On any other space\nit keeps running without events, and the connector log records `AG-UI emitter stopped` with the\nreason. For Claude Code the connector is the MCP server: it leaves the mesh and exits with code 1,\nand its stderr carries that line.\n\nA new session includes its first run even when Claude writes a positional startup prompt before the\nconnector receives `SessionStart`. That from-zero read is keyed only to Claude's explicit\n`source: \"startup\"`; resumed, forked, cleared, and compacted sessions adopt at the transcript boundary\ncaptured at that adopt, before the mesh link connects, so nothing Claude appends while the connector\nis still starting up lands behind the cursor and is silently dropped. Crash recovery follows the\ncursor already stored in the event write-ahead log, regardless of the new process's startup label.\n\nClaude starts each hook in its own process, so a prompt or stop relay can reach Cotal before the\n`SessionStart` relay. The connector holds those event flushes and the terminal until `SessionStart`\nsupplies the source, then enqueues adopt, flush, and close in that order.\n\n`SessionStart` can also run before the connector process has bound its local control socket. The hook\nthe `SessionStart` relay retries only transient pre-connect listener errors, with capped backoff\ninside its existing two-second budget. Later hooks and permanent local faults still fail open\nimmediately. Once a socket has connected, a broken exchange is not retried: the connector may\nalready have handled the frame, so replaying it could apply one lifecycle event twice.\nThat retained `SessionStart` can itself arrive before Claude creates the transcript path. A genuinely\nnew startup waits up to five seconds for that file with capped backoff, and the same deadline bounds\none stalled file read; expiry fails loud instead of silently losing the first run. A forked session\ngets the same wait, because Claude copies the parent transcript into the fork's own file after the\nhook, and then adopts at the end of that copy. Resumed, cleared and compacted starts and recovered\ncursors still require their existing source at once.\n\nTool arguments (`TOOL_CALL_ARGS`) and tool results (`TOOL_CALL_RESULT`) are not republished\nonto this channel. The durable emitter drops those events before they are written to the\nwrite-ahead log, because this channel's read ACL is not the ACL the tool ran under. Content is\nmandatory on both kinds, so the event is suppressed rather than emptied or replaced with a\nplaceholder. Tool start and end still go out. A restart that finds a pending pre-fix frame\nstill carrying those kinds HALTS rather than republishing it.\n\nThe channel is **`events.<owner>.<actor>`**, named after the session's principal. What the actor\nhalf is depends on the mesh, and the difference matters when you go looking for it: on a static mesh\nit is a key the manager allocated, never the display name, so two live agents sharing a display name\ndo not share a stream; on a user-auth mesh it is the agent's own name, because that is what the\nledger row is keyed on. Spelled out again with both halves below. The launch grants publish rights\non that channel alone. A spawn\nthat asks for a *different* agent's event channel is refused at the door rather than granted, since\nthat channel is that session's event stream. The same rule runs on restart: a manager\nresume document that names another agent's event channel is refused rather than adopted, because the\nmanaged row is re-armed from that document and the credential is re-minted from the row.\n\nThe rule reads a **concrete** channel, two principal tokens and nothing else. A pattern such as\n`events.<owner>.>` is not an event channel to it and passes untouched, governed by ordinary ACL\nauthority: on a user mesh the delegation envelope, on a static mesh the spawning credential itself.\nThat is deliberate, because the pattern is the form an operator writes on purpose for an observer,\nand it is worth knowing rather than assuming the fence is total.\n\nTo let something else read a plane, grant it out of band. The refusal prints the command for the\nmesh it is running on, spelled out in full, and only that one.\n\nOn a **user-auth** mesh:\n\n```bash\ncotal actor grant <reader> --owner <owner> --scope '' --allow-subscribe 'events.<owner>.<actor>' --allow-publish ''\n```\n\nEvery field, deliberately. `actor grant` is an upsert of the whole row, so it refuses a grant that\nleaves off any of the three ACL flags. Only `--full` turns an omitted flag into the wide default\n(`>` read, `>` post, `spawn,role:default` scope), which is the opposite of what a scoped watcher is for.\n\nOn a **static** mesh there is no actor ledger for `actor grant` to write to, and the refusal says\nso; mint the reader instead:\n\n```bash\ncotal mint watcher --profile agent --allow-subscribe 'events.<owner>.<actor>' --provision\n```\n\nThe **agent** profile, not the observer one. `mint` reads `--allow-subscribe` only for that\nprofile, and refuses it anywhere else: `--profile observer --allow-subscribe <channel>` exits\nnon-zero and writes no creds file, because the observer profile carries a fixed read set over the\nwhole chat plane, which is the opposite of what a scoped watcher is for. The agent profile also prints the lifecycle uid the\nreader needs, since an authed consuming endpoint refuses to start without one.\n\nOn an **open** mesh there is nothing to grant: the mesh has no credentials and no ACLs, so any peer\nthat lists the channel reads it, and the refusal says so instead of naming a command. The\nown-channel rule still applies there, because a spawn is not the place to hand out a read on\nanother agent's tool inputs and outputs.\n\nTwo things a reader has to do that are not obvious, both on `CotalEndpoint`. It must pass the event\nchannel in `channels`: an endpoint reads the channels it lists, so one constructed without\nthe event channel joins nothing and the frames never arrive. And it reads history with `readHistory(channel)`, the delivery daemon's mediated read, not\n`channelHistory(channel)`: a scoped credential is denied the ad-hoc consumer the direct read\ncreates, by design. `cotal console` and the web console already do both.\n\nThe `<owner>.<actor>` pair is the session's principal. On a user-auth mesh the actor half is the\nagent's own name, so the channel is `events.<your-owner>.<agent-name>`. On a\nstatic mesh the owner half is the literal `local` and the actor is a key the manager allocated, so\nthe channel is `events.local.<key>`; the spawn reply carries that key as `id`. Note\nthat `cotal console` and the web console keep event channels out of their channel lists on purpose,\nsince a plane is a machine feed rather than a conversation; they draw the frames when you open the\nchannel by name.\n\nThe rule governs the manager's doors, which are the ones a caller other than you can reach. A\nforeground `cotal spawn` on your own machine mints from your own signing material, so it can still\ngrant any channel you name: that is the out-of-band grant, not a way around the rule.\n\n**Failed turns publish run errors.** Claude Code decides for itself\nwhether a turn finished or died and fires one of two hooks accordingly, so the connector relays that\ndecision rather than making one of its own: a turn that ended on an API error ends its run with\n`RUN_ERROR` carrying the fixed message `run failed` and no code. Neither the detail Claude Code\nreported nor its error kind is published there: both are upstream values that can echo your prompt or\ntool output, and the events channel has a different read ACL. The error kind still reaches presence\nas the agent's condition (`rate_limit`, `auth`, `billing` and the rest). A turn that ended normally still\nends with a run-finished event carrying no outcome, which says the turn ended and does not claim it\nsucceeded.\n\nEvents are written to a per-session write-ahead log before they are published, so a hook that fires\nafter a restart resumes at the cursor it left rather than replaying or skipping, and a run that was\nopen when the session stopped is closed rather than left dangling.\n\nOne channel carries **every session of one agent**, because it is named after the principal and not\nafter the session. Alongside the per-session logs the connector keeps one small record per principal,\nholding the last sequence the broker assigned on that channel, so a new session continues the stream\nits predecessor left instead of starting again from nothing. Both live under the events state root\n(`COTAL_WORKSPACE_ROOT`), and neither is something you edit by hand.\n\nA **missing** record is not a fault: the connector rebuilds it from the session logs beside it,\nwhich is how an agent that was already running before this record existed keeps its stream. That\nrebuild stops if any one of those session logs is damaged. Unreadable, not valid JSON, and written\nfor a different principal all count, and so does a session directory or a log that is a link rather\nthan the real file the connector wrote, or a log that has more than one name. A tip taken from the\nrest would be too low, and it would stop publication later with nothing left to point at the cause.\nThe connector names the file instead, and the only way past it is the directory removal described\nbelow, under the same condition. A record that **disagrees with the broker** is a fault, and the\nconnector stops publishing and says why rather than guessing. A record that **moved while a session\nwas writing to it** is refused the same way: it means something else wrote the principal's record,\nand the connector reports which value it held and which the file holds rather than writing over the\nlater one. There is no command to clear it. The state is the principal's directory under the events\nroot, and clearing it by hand means removing that directory whole: the sequence, the cursor and the\nper-session logs only mean anything together, so removing part of it leaves a state the next start\nrefuses. Removing it is only half a remedy, and the half that comes first is the channel. The\ndirectory is where the agent's memory of the tip lives, not the tip itself, so on a channel that\nstill holds frames the next session opens expecting an empty one and stops on the same\ndisagreement, with the logs a tip could have been rebuilt from now gone. Purge the channel first,\nthen remove the directory.\n\nReading it: `cotal console` and the web console draw event frames directly. A frame carries no text\npart by design, so a surface that renders a message as flat text shows a marker instead of prose.\n\n**On a per-user-auth mesh, the default event plane needs the spawner's grant to cover the channel.** The event\nchannel is added to the child's publish set, and delegation only narrows: an agent may hand down\na subset of what it holds and no more. So a peer-initiated spawn is refused unless the\nspawning identity's own grant already covers the child's event channel. The refusal prints the\nexact `cotal actor grant` command that widens it. An operator launch, whose chain reaches an\nadmin-scoped or roster row, is unaffected. Passing `events: false` is the explicit opt-out.\n\nArming the event plane through a typed spawn request (`manager.spawn` with `events`, including\nthe CLI's `cotal spawn --detach --events`) is allowed on a user mesh when the child runs under\nthe caller's owner. This also satisfies a registration policy that requires the plane. A spawn\nunder another owner still needs the caller's admin tier. Without it, an explicit request or a\nrequired plane is refused before provisioning. A silent non-owner caller in a space without that\npolicy has the plane disarmed with a notice in a successful reply. Later provisioning can still\nrefuse the spawn. The child's own-channel rule and the ledger's delegation envelope still apply.\n\n## Resume a session\n\n`--resume <session-id>` pulls an existing Claude session, its context and transcript,\ninto the mesh. It **forks**: Claude mints a *new* session id from that transcript\n(`--resume <id> --fork-session`), so the meshed agent gets its own session and the\noriginal is untouched.\n\n- `cotal spawn --resume <id>` (foreground) is the primary surface: the transcript is on\n *your* machine, and errors are Claude's own stderr, inline.\n- `--detach --resume <id> --on <instance>` carries a session held on *your* machine to that\n manager instance, which may run on another host. The CLI finds the transcript under your\n Claude config (`~/.claude`, or `$CLAUDE_CONFIG_DIR`), sends it through a JetStream Object\n Store bucket only that instance reads, under a writer credential pinned to that one transcript,\n and prints `carried session <id> to <instance>:\n sha256:<hex>, <sent> of <size> bytes sent in <chunks> chunks`. A re-run of the same bytes\n sends nothing, and an interrupted carry continues where it stopped. The seat forks it in a\n private Claude home under the manager's `.cotal/seat-homes/`, which no other seat's Claude\n lists or finds, and which is removed when the seat stops. When Claude starts the fork, the seat\n records the SHA-256 of the transcript it read from its own project; the manager stops a seat\n whose record names other bytes than the carried ones, or that records none within the join\n timeout after it joins, an uncertain launch included, and otherwise shows that record as the\n seat's provenance. `cotal attach` to such a seat names its source after the seat\n name, as `(resumed from <host>:<id>)`. A remote manager receives a carry when its host issues it a\n transfer reader. On a user-auth mesh the CLI exchanges the operator's login for a one-object\n `transfer-writer` view, which needs scope `admin`.\n- A session name in place of an id is refused, listing each session on this host that carries\n that name with its id, SHA-256 and modification time. An id this host does not hold resolves\n against the **manager host's** `~/.claude`, as before.\n- A seat-private home holds no login. The manager host needs `CLAUDE_CODE_OAUTH_TOKEN` (from\n `claude setup-token`), `ANTHROPIC_AUTH_TOKEN`, or a cloud provider selection in its\n environment; `ANTHROPIC_API_KEY` alone is refused. The launch directory must already be\n trusted by the manager host's own Claude, and Claude must be 2.1.234 or later.\n- The manager waits for a real outcome: `✓ started` means the agent *joined the mesh*,\n `✗ exited on launch` carries Claude's last output, and an uncertain launch (~30 s) is\n reported without tearing the agent down.\n- Resume is an **operator surface only**, deliberately not exposed on MCP `cotal_spawn`\n (a mesh peer naming host-local transcripts would widen `spawn` into transcript\n disclosure). Only the Claude connector supports it today; OpenCode and Hermes fail loud.\n- Needs a `claude` new enough for `--resume … --fork-session` (verified on 2.1.197).\n\n## Sharing your MCP servers\n\nA spawned session keeps your own MCP servers by default. On its first run, `cotal setup` copies\nthe user-scope servers from your Claude Code config (`~/.claude.json`, or the one under\n`$CLAUDE_CONFIG_DIR`) into the cotal config file (`~/.config/cotal/config.json`) under\n`connectors.claude.mcpServers`, and names them in its output. With none to copy it writes an\nempty list. Each entry is the familiar `.mcp.json` shape ([full format](config.md)). A cotal\nconfig that already declares that list keeps it, and a later `cotal setup` never changes it.\n\nThe cotal config holds secrets only as `${VAR}` references. Setup cannot tell literal text from\na secret, so it leaves out a server with an `env` or `headers` value that is anything but `${VAR}`\nreferences (a `Bearer ${TOKEN}` header among them) and names it in its output. To share one,\nadd it to the cotal config with each secret written as a `${VAR}` reference, and export that\nvariable where you spawn. Setup also leaves out and names an entry no session can start, such as\none with a missing or empty `command` or `url`, or one whose `command` is not a string.\n\nAt launch the connector forwards *only* the named vars the chosen servers declare and\npasses the merged config as an owner-only temp file; `--strict-mcp-config` stays on, so\nonly cotal + the shared servers load.\n\nFor a lighter seat, share fewer. Remove an entry from the cotal config to drop it from every\nspawn, or scope one spawn with `--share-tools tavily,figma` (or `--share-tools none` for cotal\nalone). An empty list (`\"mcpServers\": {}`) in `~/.config/cotal/config.json` keeps every spawn\nisolated, and setup leaves it as it is.\n\nTwo caveats: sharing a server grants its credential to the agent (the var lives in the\nClaude process's environment, so share only when you're fine with that teammate holding\nthe key), and memory adds up, because a heavy server boots once per spawn, multiplied\nacross a team, and can starve a small machine.\n\n## Feedback\n\n`cotal_feedback` works out of the box: without a key it posts to the public intake at\n`https://cotal.ai/v1/feedback` (needs a contact email: `COTAL_FEEDBACK_EMAIL`, then\n`git config user.email`, else the agent asks). Set `COTAL_FEEDBACK_KEY=fbk_<key>` in a\nbeta tester's environment to route to the keyed intake (`Authorization: Bearer`, identity\nderived from the key); `COTAL_FEEDBACK_URL` overrides either endpoint. The CLI can send\ntoo: `cotal feedback \"<summary>\" [--type bug]`. Each submission carries\n`origin: human | agent`, whether the tester asked, or the agent auto-reported a major\nissue.\n"
|
|
91
91
|
},
|
|
92
92
|
{
|
|
93
93
|
"slug": "connect-codex",
|
|
@@ -269,7 +269,7 @@ export function loadDocsBundle() {
|
|
|
269
269
|
"title": "Upgrading a running deployment",
|
|
270
270
|
"kind": "Guide (informative)",
|
|
271
271
|
"summary": "Substrate stability tells you what the version numbers promise.",
|
|
272
|
-
"body": "# Upgrading a running deployment\n\n> **Guide** (informative) · **For:** operators upgrading a mesh that already exists · **See also:** [Substrate stability](stability.md), [Run a mesh](run-a-mesh.md), [Identity and auth](identity-and-auth.md)\n\n[Substrate stability](stability.md) tells you what the version numbers promise. This page is the\nother half: what to actually do when the deployment already exists, has credentials in it, and\ncannot simply be recreated. Every release that breaks a running deployment gets a section here,\nnaming what migrates on its own, what does not, and the order to move the pieces in.\n\n## The pre-1.0 upgrade contract\n\nThe packages are pre-1.0, so a minor bump may break an API or an on-disk expectation. Four\ncommitments make that survivable for someone with a fleet:\n\n- **Pin an exact version.** `0.N.P`, never `^0.N.P`. A range can pull a breaking minor in during an\n unrelated reinstall.\n- **Every break that touches a running deployment gets a section on this page**, written in terms of\n what an operator does, not in terms of which module changed.\n- **Read the section before you start, not halfway through.** A section names the work up front\n precisely so the operation does not change shape once it is underway.\n- **A break that cannot be made automatic says so.** Where credentials or state must be recreated by\n hand, the section says which ones and when, rather than leaving you to discover it at the moment\n the first one stops working.\n- **A change to the shape of a credential, or to who may renew one, is breaking whatever the commit\n marker says.** This rule is stated because the marker is a judgement made while writing the code\n and the consequence is felt by someone running it a day later. A fleet that keeps authenticating\n looks compatible and is not, if nothing in it can renew. Any automated check of this rule would\n read commit markers, so a break recorded as a feature is the one case it could not see, which is\n why the rule is written for people first. **The marker held for this release: the 0.49.0 change\n that caused all of this, `36d177951 feat(core)!`, did carry its `!`.** The rule exists for the\n next one that does not.\n\nWhat this page does not promise is a rolling upgrade. Nothing in the current line dual-serves two\nauthority versions, so where broker and manager run separately there is a window in which the mesh\nis down. The sections below give that window's shape so it can be scheduled rather than endured.\n\n## Auth context closure in 0.71.0 (unreleased)\n\nExisting deployments need no credential migration or restart for these additive APIs. Embedded\nhosts can now inspect `handle.connections()` and await `handle.closed` after `close()` or `drain()`\nto prove every owned transport ended, including the callout, replaced readiness readers and\nshort-lived clients. The inventory is a detached snapshot.\n\nA transport close failure now rejects with its connection label. The terminal signal stays pending\nwhile any connection remains live. Repair the failure and retry `close()` before awaiting\n`handle.closed`. Closing one hosted context does not close another account's context.\n\nRead a space's claim with `readPlaneClaim(kv, space)` on that account's leader-only auth bucket.\nAn unclaimed space returns `undefined`; held and released rows retain their claim identity.\nDeleted, malformed and foreign-space rows refuse. `PlaneClaimRow` and `PLANE_CLAIM_KEY` are exported.\n\nUse `observeAccountLivenessWithCreds({ servers, observerCreds, accountId, options })` with the\naccount-scoped membership-observer credential to list that account's connections. It never widens\ncredentials or evicts connections. Zero rows prove absence only with a complete sweep and the\nsingle-server proof. An embedded endpoint's trusted composition can retain transport custody\nthrough `EndpointOptions.onConnection`.\n\n## Hermes model from the environment in 0.68.0\n\nA connector now launches on the model and variant its launcher resolved (the `--model` or\n`--variant` flag, else the agent file's `model:` or `variant:`) and no longer reads them again from\nthe agent file. The Hermes connector also no longer takes a model from `HERMES_MODEL` in the\nenvironment of the process that spawns the seat, including when `spawn.env` lists it.\n\n### What stops working\n\nA Hermes spawn whose only model was `HERMES_MODEL` in the spawning environment is refused at launch,\nand the refusal names both ways to set a model. Spawns that set `--model` or `model:` are unchanged,\non every connector.\n\nCode that calls a connector's `buildLaunch` directly with only `configPath` now gets no model or\nvariant from that file. Pass them as `model` and `variant`.\n\n### Before the upgrade\n\nMove each Hermes seat's model from `HERMES_MODEL` onto its spawn with `--model`, or into its\npersona's `model:`.\n\n## Run answers on a participant manager in 0.68.0\n\nA participant manager now asks its issuing host for an answering credential by naming the run and\nstep it answers. The host reads the pause's token off that run's journal and no longer accepts a\ntoken from the manager. Runs on a mesh with no participant manager are unaffected.\n\n### What stops working\n\nWhile a participant manager and its issuing host run different sides of this release, the host\nrefuses every `cotal run answer` and every amendment that manager serves, because each side refuses\nthe other's request shape. Starting, resuming and reading runs is unchanged. A pause stays waiting\nthrough the window, or follows its timeout if it has one.\n\n### Before the upgrade\n\nUpgrade the auth service and every participant manager registered with it in the same window, then\nanswer the pauses that waited.\n\n## Headless OpenCode handshake in 0.69.0\n\nWith `COTAL_SERVE_HEADLESS=1`, the OpenCode launcher's `[cotal-serve]` line on stdout now carries\nonly `port` and `session`. The server password no longer appears in it, and the 1.x TUI no longer\nreceives the password on its command line.\n\n### What stops working\n\nA headless host that read `password` from that line has no password, and the server refuses its\nrequests. Seats with a TUI, and headless seats that no host drives, are unaffected.\n\n### Before the upgrade\n\nHave each headless host mint a password and pass it to the launcher as `OPENCODE_SERVER_PASSWORD`,\nthen use it for basic auth as before. Without that variable the launcher mints its own.\n\n## Filesystem store identity in 0.69.0\n\nThe delivery daemon's answer to the manager's store check now names a filesystem store by its root\nand by a random `id` that the store records once in `store.id` inside its own directory:\n`.cotal/store.id` for a workspace root, or the directory of the file for `cotal deliver --creds\n<file>`. A manager no longer counts the daemon's store as its own because the two roots have the\nsame path. On a split whose broker host and manager host use one root path, the manager host now\nstays off the daemon-credential renewal lease, so `cotal doctor auth --fix` on the broker host can\nrenew the daemon credentials.\n\n### What stops working\n\nA manager and a delivery daemon on different sides of this release refuse each other's answer to\nthe store check. The manager then remints no daemon credential, and a manager that is booting does\nnot start. This is read from the code and was not measured across two releases. A\n`cotal deliver --creds <file>` whose directory is a read-only mount and holds no `store.id` stops at\nstart. So does a `--creds` file that is its directory's `store.id` under any name, and a `store.id`\nthat is a symbolic link or holds anything but a lowercase UUID.\n\n### Before the upgrade\n\nUpgrade the broker host and every manager host of a space in the same window. For a `--creds` file\non a read-only mount, add a regular `store.id` file beside it that holds a new lowercase UUID and no newline,\nas `node -e 'process.stdout.write(crypto.randomUUID())' > store.id` writes. Move a `--creds` file\nnamed or linked as `store.id` to a file of its own.\n\n## Detached spawns with `--share-tools` in 0.69.0\n\nThe manager's `spawn` operation now takes `shareTools` as a list of MCP server names. The CLI parses\n`--share-tools` into that list before it sends the request, and the manager cluster document moves\nto revision 22. A cut taken with `cotal down --preserve-state` before the upgrade still resumes: the\nmanager reads its `cotal-manager-resume/v1` inventory and writes new cuts as\n`cotal-manager-resume/v2`.\n\n### What stops working\n\nA CLI and a manager on different sides of this release refuse a detached spawn that passes\n`--share-tools`, because the CLI checks each request against the contract the manager serves. This\nis read from the code and was not measured across two releases. A detached spawn without the flag,\na foreground spawn and a roster entry are unaffected. A manager older than this release cannot\nresume a cut that this release took.\n\n### Before the upgrade\n\nUpgrade the CLI on every host that runs `cotal spawn --detach` in the same window as the managers\nit reaches.\n\n## Shared MCP server checks in 0.69.0\n\nThe cotal config reader now checks each server under `connectors.<name>.mcpServers` when it reads\nthe file, and refuses one that cannot launch as written, naming the file and the field. The rules\nare in [the config file](config.md#the-config-file).\n\n### What stops working\n\nA config file that holds such a server refuses every Claude spawn that reads it, including one with\n`--share-tools none`. Before, a field of the wrong type failed each Claude spawn that shared the\nserver with a `TypeError` that named neither the file nor the server, a spawn that did not share it\nlaunched, and a server with no `command` or `url` was passed to `claude`, which never started it.\nRead from the code and not measured: spawns on other connectors, a manager resume and the step of\n`cotal setup` that records the shared list read the same files, so each stops at the same refusal.\n\n### Before the upgrade\n\nCheck `connectors.<name>.mcpServers` in the operator-level config file and in each space's\n`.cotal/config.json`. Give each server a string `command`, or a `type` of `http`, `sse` or `ws` with\na string `url`. Write `args` as a list of strings and `env` and `headers` as objects of strings, or\nremove the server.\n\n## Remote manager family eviction in 0.69.0\n\nA remote manager registered through its host now asks the host to evict up to 256 holders of its\ncredential family in one maintenance request, and the host reads the family once for the whole set.\nBefore, a restart sent one request per holder and the host read the whole family for each one.\nMeshes with no remote manager are unaffected.\n\n### What stops working\n\nWhile a remote manager and its issuing host run different sides of this release, each side refuses\nthe other's eviction request shape. A restart whose credential family already has holders then fails\nat its eviction step and leaves the manager's registration gate frozen. A first start, a clean stop\nand the host's reconciliation of a foreign slot holder are unchanged.\n\n### Before the upgrade\n\nUpgrade the auth service and every remote manager registered with it in the same window. A manager\nthat restarted inside the window resumes its frozen registration on its next start once both sides\nrun this release.\n\n## AG-UI emitter holder hooks in 0.69.0\n\n`AguiEmitterHolder` from `@cotal-ai/connector-core` now takes its hooks as one named object after\nthe emitter factory: `new AguiEmitterHolder(startEmitter, { onError, onRunClosed, waitLive, runMeta })`.\nOnly `onError` is required. Nothing about a running mesh changes, and every shipped connector passes\nits hooks by name. Only a connector of your own that builds a holder is affected.\n\n### What stops working\n\nA holder built with positional hooks, such as `new AguiEmitterHolder(start, onError, onRunClosed)`,\nno longer compiles, because the constructor takes two arguments. Plain JavaScript that keeps the\npositional form still runs, but the holder calls none of its hooks, so a failure never reaches\n`onError`.\n\n### Before the upgrade\n\nPass each hook by name, for example `new AguiEmitterHolder(start, { onError, onRunClosed })`, and\ndrop any `undefined` that filled an earlier slot to reach a later hook.\n\n## Worker run failure type in 0.70.0\n\n`WorkerRunFailed`, the failed result of `runInWorker` in `@cotal-ai/lang`, is now a union on\n`class`: `released`, `held`, `effect`, `too-large`, `rejected` or `error`. A running mesh needs\nnothing, because the runtime host and the engine thread ship in the same install. A run on the\ncompiled engine whose program throws an object with `code: \"L5012\"` or `code: \"L5025\"` used to end\nreleased and now ends failed, as it does on the walker.\n\n### What stops working\n\nTypeScript code that reads `code`, `reason`, `step`, `pending`, `kind`, `detail` or `tooLarge` on a\n`WorkerRunFailed` it has not narrowed fails with TS2339. A released, held, too-large or rejected\nresult no longer carries `code`, so JavaScript that branched on `L5012`, `L5025`, `L5006` or\n`L5010` stops matching with no error. `tooLarge` is gone.\n\n### Before the upgrade\n\nBranch on `class` where such code read `code`: `released` for L5012, `held` for L5025, `too-large`\nfor L5006 and `rejected` for L5010. An `effect` or `error` result keeps its `code`. Once narrowed to\n`too-large`, a result carries the `stepKey`, `bytes` and `bound` that `tooLarge` held.\n\n## Remote manager request builder in 0.70.0\n\n`remoteManagerClient.remoteManagerAuthorityRequest` from `@cotal-ai/manager` now takes an\noperation's coordinates as one object, and `remoteManagerRegistrationProof` from `@cotal-ai/core`\ncomputes the proof from the manager's identity state instead of a request. Nothing about a running\nmesh changes: the proof digest and the request on the wire are the same, so a manager and a host on\ndifferent sides of this release still accept each other. Only code that builds remote manager\nrequests itself is affected, in TypeScript and in plain JavaScript.\n\n### What stops working\n\nA call that passes the registration proof, contract artifacts, session, retirement or transfer\nreader as positional arguments after the operation no longer compiles. A call that passes a request\nto `remoteManagerRegistrationProof` no longer compiles either, because the second argument now names\nthe lifecycle `lifecycleUid`, as the identity state does.\n\nPlain JavaScript runs both old calls without an error. The builder drops the positional coordinates,\nso the host refuses the request with `requires a sha256 registrationProof`. A proof computed from a\nrequest leaves out the lifecycle, so the host refuses a request that carries it as a proof mismatch.\n\n### Before the upgrade\n\nName the coordinates, for example\n`remoteManagerAuthorityRequest(state, \"cli\", \"retire\", { registrationProof, retirement })`.\nCompute the proof as `remoteManagerRegistrationProof(owner, state)`, adding the contract artifacts\nas a third argument for activation only. A host that recomputes the proof from a received request\npasses `{ space, instanceId, lifecycleUid: managerLifecycleUid, identities }` from that request.\n\n## Bearer validator lifetime cap in 0.70.0\n\n`validateUserToken` from `@cotal-ai/auth` no longer takes `maxTtlSec`. It caps a bearer's lifetime\nat the cap of the bearer's view, the same cap the issuer applies when it mints: 900 seconds, or 300\nfor a `transfer-writer` bearer. The auth callout never passed the option, so a running mesh behaves\nas before. Only code of your own that calls the validator with `maxTtlSec` is affected.\n\n### What stops working\n\nA call that passes `maxTtlSec` in an object literal no longer compiles. Plain JavaScript that keeps\nit still runs, and the value is ignored. A `NaN` value, such as `Number()` of an unset environment\nvariable, used to turn the lifetime check off and accept a bearer of any lifetime. That bearer is\nnow refused at its view's cap.\n\n### Before the upgrade\n\nRemove `maxTtlSec` from each call. A test that needs a bearer to expire sooner mints one with a\nshorter lifetime.\n\n## Persisted identity records in 0.70.0\n\nThe manager instance identity, the manager sibling identities, the auth plane instance identity and\na participant manager's remote authority state now share one reader and one first mint in\n`@cotal-ai/workspace`, exported as `claimIdentityRecord` with the nkey check `identityOf`. Each\nrecord is read as a regular file, must hold non-empty nkeys and is created exclusively, so\nconcurrent first starts of a participant manager on one root now settle on one identity where each\nused to keep its own. `saveManagerInstanceIdentity` and `saveAuthInstanceIdentity` are gone. A\nrunning mesh whose records are plain files needs nothing.\n\n### What stops working\n\nA manager instance, auth instance or remote authority record that is a symlink, a directory or any\nother non-regular entry is refused where it used to be followed. The manager, the auth plane and a\nparticipant manager fail to start on it, and `cotal reconcile-gate` and `cotal deregister-instance`\nrefuse it. Retirement already refused it. A remote authority record with an empty nkey id or seed\nis refused too. A first mint that loses its race and cannot read the winner now refuses with\n`identity-record-create-lost` in place of `manager-instance-identity-create-lost` or\n`auth-instance-identity-create-lost`. Code that imports either `save` function no longer compiles.\n\n### Before the upgrade\n\nReplace a symlinked identity record with a copy of the file it points to. Code that wrote a record\nwith a `save` function plants it with `createManagerInstanceIdentity` or\n`createAuthInstanceIdentity`, which create the record when it is absent and otherwise return the\nstored one unchanged. Nothing replaces an overwrite of a stored identity.\n\n## Manager instance in user credentials in 0.70.0\n\n`AuthProvider.userCredentials` from `@cotal-ai/core` no longer returns `managerInstanceId`. A\n`manager-caller` credential's manager instance is the signed `act.managerInstanceId` claim in its\nbearer, which the broker verifies and the CLI already used. The reference provider in\n`@cotal-ai/auth` stops copying the exchange response's field into its result, where nothing\ncompared it with the bearer. The exchange still answers with the field, so a running mesh behaves as\nbefore.\n\n### What stops working\n\nCode of your own that reads `managerInstanceId` from a `userCredentials` result no longer compiles,\nand plain JavaScript reads `undefined` there.\n\n### Before the upgrade\n\nRead the instance from the bearer's `act.managerInstanceId` claim.\n\n## Auth plane identity location in 0.70.0\n\nThe user-auth service keeps its instance identity in the root's `.cotal/space.<hex>/auth-instance.json`,\nbeside the manager's. It used to sit inside `.cotal/auth`, at\n`space.<hex>/.cotal/auth/auth-instance.<hex>.json`, so a copy of that folder carried it. The first\nstart of an upgraded root moves the record and keeps the instance. A hosted context started through\n`startAuthService` has its record moved the same way inside its `stateDir`.\n\n### What stops working\n\nCode that calls `openAuthAuthorityPlane` without the new `identityRoot` option no longer compiles. A\nstart that finds a record both in `.cotal/space.<hex>/` and at its older place refuses and names the\ntwo files. A start also refuses when the older place of the auth or manager identity holds a symlink,\na directory or anything else that is not a regular file. The manager used to skip a dangling symlink\nthere and mint a new identity.\n\n### Before the upgrade\n\nPass `identityRoot` to `openAuthAuthorityPlane`. When `dir` is a workspace root's user-auth state\ndir, `<root>/.cotal/auth/space.<hex>`, pass that root. A plane with no workspace root, as\n`startAuthService` runs, passes `dir` itself. Either keeps the identity the plane already has: on the\nfirst start it moves from `<dir>/.cotal/auth/` to `<identityRoot>/.cotal/space.<hex>/`. Never pass a\ndirectory inside `.cotal/auth`: the record would land in the folder an operator copies and travel\nwith it again.\n\nA copy of `.cotal/auth` taken from a root last run by an older Cotal carries that root's record.\nDelete `.cotal/auth/space.<hex>/.cotal/auth/auth-instance.<hex>.json` from the root you copied it to\nbefore the first `cotal up --user-auth` there.\n\n## Per-seat `COTAL_` names in `spawn.env` in 0.71.0\n\n`spawn.env` in the cotal config no longer forwards a `COTAL_` name the launcher sets for each seat,\nsuch as `COTAL_ROLE`, `COTAL_MODEL` or `COTAL_SUBSCRIBE`. Before, a seat launched with no value of\nits own took the spawning process's value and ran under that role, model or read set. The\nmachine-wide knobs a seat already receives, such as `COTAL_HOME`, may still be listed.\n\n### What stops working\n\nEvery spawn and resume under a config whose `spawn.env` lists such a name is refused before\nlaunch, and the refusal names the entry. Code that calls `launchEnv` from `@cotal-ai/connector-core`\nwith such a name in `envAllow` gets the same error.\n\n### Before the upgrade\n\nRemove those names from `spawn.env`. Give each seat its role, model and channels with `--role`,\n`--model` and `--subscribe`, or in its persona's `role:`, `model:` and `subscribe:`.\n\n## Role addresses in 0.71.0\n\nA role must be one `[A-Za-z0-9_-]` token. Before 0.71.0 any other spelling was rewritten into one:\n` probe ` reached the `probe` queue and `pro.be` reached `pro_be`, while the message kept the\nspelling sent. An anycast to `*` was accepted and stored where no holder reads it.\n\n### What stops working\n\nAn agent whose role is outside the token set no longer starts, however it is launched:\n`cotal join --role`, `cotal spawn --role`, an agent file's `role:`, `COTAL_ROLE` and an embedded\nendpoint's `card.role` are all refused before the agent joins.\n\nA send to such a role, or to `*`, through `cotal send ask`, `/anycast` or `cotal_anycast` is refused,\nand nothing is stored.\n\n`routeToken` is no longer exported from `@cotal-ai/core`. A role routes as spelled, so code that\nused it to name a role's queue uses the role itself, and `assertValidRole` checks one.\n\n### What migrates on its own\n\nEvery task queue. A `svc_<role>` durable was always named from the rewritten token, so its pending\nrequests and its holders carry over.\n\n### Before the upgrade\n\nRename each role outside the token set to the token it already routed to: remove the surrounding\nspaces and replace every other character outside the set with `_`. Rename it where the holder is\nlaunched and in every script or prompt that sends to it.\n\n## Carrying a resumed Claude session to another host in 0.67.0\n\n`cotal spawn --resume <id> --detach --on <instance>` now carries a Claude session held on the\noperator's host to the target manager instance. Both sides need this release: an older manager does\nnot serve `transcript-receive`, and the CLI then stops with that manager's refusal instead of\nlaunching. The manager cluster document moves to revision 21, and the `ps` row's `resume` object\ngains `host` and `transferredAt`.\n\nA manager host that runs carried seats needs `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_AUTH_TOKEN` or a\ncloud provider selection in its environment, because each carried seat runs in its own Claude home\nwith no stored login. On an authenticated mesh the CLI mints the transfer writer from the space's\nsigning seed, so the carrying host needs that seed, as for any other operator command. On a user-auth\nmesh it exchanges the operator's login for a `transfer-writer` view instead, so the operator's grant\nneeds scope `admin`, and the auth service must run this release. A remote manager receives a carry once\nits host serves the manager-service `transferReader` operation. A seat launched without carrying,\nincluding any `--resume` whose id this host does not hold, is unchanged.\n\n## Lifecycle head type in 0.67.0\n\n`LifecycleMapping`, the type `parseLifecycleHead` returns, is now a union on `state`. Nothing about\na running mesh changes: heads that parsed before parse the same way, and the refusals are\nunchanged. Only TypeScript code that compiles against `@cotal-ai/core` is affected.\n\n### What stops working\n\nAn `interface` that extends `LifecycleMapping` fails with TS2312, because an interface cannot extend\na union. Code that builds a head in memory no longer compiles when the head is `retiring` without\nits `op`, or `active` or `retired` with one. The parser already refused those heads.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type ActiveMapping = LifecycleMapping & { state: \"active\" }`. A reader that has checked\n`state === \"retiring\"` reads `op` without a guard.\n\n## Issuance gate types in 0.67.0\n\n`EpGateRow` and `EndpointGateRow`, which `parseIssuanceGate` and `parseEndpointGate` return, and\n`EpGateState`, which an `EpIssuanceGate` or `EpIssuanceBarrier` returns from `observe`, are now\nunions on `state`. Nothing about a running mesh changes: gates that parsed before parse the same\nway, and the refusals are unchanged. Only TypeScript code that compiles against `@cotal-ai/core`\nis affected.\n\n### What stops working\n\nAn `interface` that extends one of these types fails with TS2312, because an interface cannot\nextend a union. Code that builds a gate in memory, such as a custom barrier's `observe`, no longer\ncompiles when the gate is `frozen` or `retired` without its `op`, or `open` with one. The gate\nparsers already refused those rows.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type CustomGateRow = EpGateRow & { custom: string }`.\n\n## Lifecycle-blocked refusals in 0.66.0\n\nA refusal that carries `ai.cotal.ep.lifecycle-blocked` now reports only the lifecycle state it\nread. Nothing about a running mesh changes. A client that branches on the detail must read the new\nfield.\n\n### What stops working\n\nA refusal raised at the issuance gate used to carry `headState` without reading the head:\n`retiring` for a frozen gate and `retired` for a retired one. It now carries `gateState`\n(`frozen` or `retired`) and no `headState`. A client that treats `headState: \"retired\"` as a\nburned uid, or `headState: \"retiring\"` as a retirement in flight, no longer matches those\nrefusals, and the `[lifecycle ...]` suffix on the error string changes the same way. A custom\nissuance barrier whose `observe` returns a frozen gate without a valid `op` (a string `opId` and\none of the four op kinds) is now refused as `internal` by `registerServiceInstance`.\n\n### Before the upgrade\n\nUpdate such a client to read `gateState` for a gate refusal and `blockedOp` for the operation that\nholds the gate. `headState` is present only when the refusal read the head, for example an\nactivation refused because the head is still retiring.\n\n## Workflow programs that bind `once` in 0.65.0\n\n`once` is now a scope of the workflow language, so it is a reserved name. A program that declares\nits own `once` binding (`const once = ...`, a parameter or a function named `once`) is refused at\nvalidation with L2002. Nothing else about a running mesh changes.\n\n### What stops working\n\nA run whose recorded program binds `once` cannot be resumed after the upgrade, because a resume\nvalidates the recorded program again. A new `cotal run start` of such a program is refused before\nanything is recorded.\n\n### Before the upgrade\n\nList the runs with `cotal run ps` and check each program that is still running or held for a\nbinding named `once`. Let those runs finish on the old version before you upgrade the manager, and\nrename the binding in the program before you start it again.\n\n## From 0.58.0 to 0.59.0\n\nEvery connector now publishes a failed run's `RUN_ERROR` on `events.<owner>.<actor>` with the fixed\nmessage `run failed` and no `code` or `rawEvent`. The error text and error kind a harness reports\ncan echo a prompt, a peer message or tool output, and that channel has a different read ACL. A\nreader that showed the message or branched on `code` gets neither after the upgrade. Where a\nconnector reports the error kind as the agent's presence condition, that is unchanged.\n\n### Settle pending event frames before the upgrade\n\nEach session's events are frozen in its event write-ahead log before they are published. A session\nrestarted on 0.59.0 whose log still holds an unacknowledged frame with an older `RUN_ERROR` does not\nrepublish it: its event emitter halts with `egress-run-error` and publishes nothing further for that\nsession. The broker may or may not already hold that frame, so the halt cannot settle it.\n\n1. Stop the seats cleanly on 0.58.0, with the broker still up.\n2. List the logs that still hold a pending frame. The logs live under the events state root\n (`COTAL_WORKSPACE_ROOT`). Empty output means there is nothing to settle.\n\n ```sh\n find \"$COTAL_WORKSPACE_ROOT/.cotal/events\" -name wal.json \\\n -exec jq -r 'select(.pending != null) | input_filename' {} +\n ```\n\n3. For each session listed, start it again on 0.58.0 while the broker is reachable, let it recover,\n stop it, and run step 2 again. Recovery publishes the frame as 0.58.0 would have, error text\n included, so it only finishes what 0.58.0 had already started.\n\n If that start halts with `cas-loss` instead, the agent's subject is no longer at the sequence this\n log expects, and no restart settles that log, on 0.58.0 or later. A lost acknowledgement is one\n cause: the broker stored the frame, so it and its error text are already on the channel, and every\n retry halts the same way because the stream checks the frozen expectation before it deduplicates.\n The halt message names the other causes, such as a second emitter for the same agent under a\n different state root, a restored stream or frontier record, or a purged channel. With those the\n pending frame may never have reached the broker, so a `cas-loss` does not tell you whether it\n landed. Find and stop any second writer and rule out a restored state first. Clearing the halt\n then means purging the agent's event channel and removing the agent's directory under the events\n state root whole (see [Event plane](connect-claude.md#event-plane)). That abandons the pending\n frame whether or not the broker has it, and the purge also drops the earlier frames of every\n session of that agent.\n4. Upgrade once step 2 prints nothing.\n\nIf a session halts with `egress-run-error` after the upgrade, go back to step 3 for that session on\n0.58.0. Do not edit or delete `wal.json` on its own to get past either halt: clearing the pending\nframe abandons that epoch, an event the broker never received is lost, and removing part of the\ndirectory leaves a state the next start refuses.\n\n## Explicit actor grants in 0.59.0\n\n`cotal actor grant` no longer fills an omitted ACL flag with its wide default. A grant names\n`--scope`, `--allow-subscribe` and `--allow-publish`, or passes `--full` to give the ones it leaves\noff their wide defaults (`spawn,role:default`, `>` read, `>` post). Any other grant is refused. The\nbreak is in the CLI on the machine that holds the actor ledger, the one that ran\n`cotal up --user-auth --idp <url>`. No stored row, credential or wire message changes.\n\n### What keeps working\n\nExisting actor ledger rows keep the authority they were granted, and their users and agents connect\nas before. `actor revoke`, `actor list` and a `grant` that names all three ACL flags behave as they\ndid on 0.58.0. Nothing on disk is converted.\n\n### What stops working\n\nA grant that leaves off any of the three flags without `--full` exits 1 with\n`refusing to grant \"<actor>\" with --scope, --allow-subscribe, --allow-publish left off`, naming the\nflags it is missing, and then prints both accepted forms. It writes no row and does not retire the\nactor's current lifecycle. An existing row stays as it was, and an actor granted for the first time\nstays out until the grant is run again. This includes the bare grant printed on 0.58.0 by\n`cotal login`, `cotal status`, `actor list` and the not-granted refusal. Look for it in provisioning\nscripts, onboarding runbooks and anything that pastes those hints.\n\n### Upgrade order\n\nChange the scripts before the ledger machine is upgraded, and make each grant name all three flags.\n0.58.0 and 0.59.0 both accept that form. To keep a wide row, write its defaults out:\n\n```sh\ncotal actor grant <actor> --sub <IdP subject> \\\n --scope spawn,role:default --allow-subscribe '>' --allow-publish '>'\n```\n\nSwitch to `--full` only once the ledger machine runs 0.59.0. 0.58.0 refuses it with\n`Unknown option '--full'` before it reads the ledger. Brokers, managers and participant machines\nneed nothing for this break, so their order is the one the section above gives.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused grant changes nothing. The\nexposure is a grant script that runs against 0.59.0 before it was changed: it fails and grants\nnothing.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. On the ledger machine, save the output\nof `cotal actor list` to compare rows after the changed scripts run, and list the scripts that call\n`cotal actor grant`.\n\n### The upgrade end to end\n\n```sh\n# on the ledger machine, still on 0.58.0\ncotal actor list > actors-before.txt\ngrep -rn 'actor grant' <your provisioning scripts>\n# make every grant name --scope, --allow-subscribe and --allow-publish, run them, then upgrade\nnpm i -g cotal-ai@0.59.0\ncotal actor list | diff actors-before.txt -\n```\n\nBoth refusals quoted here were run on 0.58.0 and on the 0.59.0 code. That brokers, managers and stored\nrows need nothing is read from the change, which touches only the CLI and its hints, and was not run\non a live split deployment.\n\n## Repeated flags refused in 0.59.0\n\nA `cotal` flag given more than once is now a usage error unless the command declares it repeatable.\nOn 0.58.0 the last value won with no message, so `cotal down web --space a --space b` acted on `b`\nwhile a wrapper that checked the first `--space` verified `a`. The break is in the command-line\nparser on the machine that runs the command, including commands added with `cotal ext add`. No\nstored state, credential or wire message changes.\n\n### What keeps working\n\nA command line that gives each flag once parses as it did on 0.58.0, in any order and in the\n`--flag=value` form. Flags whose help says repeatable, such as `--opt` and `down --session-store`,\nstill collect every value. A flag-shaped word after `--` is still a positional. The daemons, units\nand agents that `cotal` starts for itself are given each flag once, so a fleet driven only by `cotal`\ncommands typed by hand needs no action.\n\n### What stops working\n\nA command line that repeats any other flag exits 1 before the command runs. It prints\n`Option '--space' cannot be repeated`, or `Option '-f, --file' cannot be repeated` for a flag with a\nshort form, followed by the command's help. `-f` and `--file` count as the same flag. Look for it in\nscripts, aliases and wrappers that append a flag to override one set earlier, such as a fixed\n`--space` followed by `\"$@\"`.\n\n### Upgrade order\n\nChange those scripts first so each flag is given once. 0.58.0 and 0.59.0 both accept that form.\nBrokers, managers and participant machines need nothing for this break, and each machine's CLI\napplies it when that machine is upgraded, so their order is the one the sections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused command does nothing. The\nexposure is a script that still repeats a flag when it runs on 0.59.0: it exits 1 instead of acting on\nthe last value.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the scripts, aliases and wrappers\nthat call `cotal` so each one can be checked.\n\n### The upgrade end to end\n\n```sh\n# still on 0.58.0\ngrep -rn 'cotal ' <your scripts and wrappers>\n# give each non-repeatable flag once, then upgrade\nnpm i -g cotal-ai@0.59.0\n# run each changed script; a repeat left behind exits 1 with the usage error and does nothing\n```\n\nThe refusal and its messages were run against the 0.59.0 parser and `cotal topology view`. That the\nargument lists `cotal` builds for its own processes give each flag once is read from the code, and\nwas not run on a live split deployment.\n\n## Detached spawns from a seat's shell in 0.62.0\n\nOn a static or open mesh, `cotal spawn --detach` run inside a managed seat's shell now launches as\nthat seat. On 0.61.0 it minted a one-shot operator instrument, so the manager recorded that\ninstrument as the spawner and the seat's own `cotal_despawn` of the child was refused with\n`not authorized: <seat> was not spawned by <caller> (admin tier required)`. The break is in the CLI\non the machine where the seats run. No stored state, credential or wire message changes.\n\n### What keeps working\n\n`cotal spawn --detach` from an operator terminal or from a script outside any seat launches as\nbefore, and so does any call with `--creds`, one aimed at a space other than the seat's own, or a raw\nopen target named with `--server` and an unregistered `--space`. A user-auth mesh is unchanged. A seat with\n`capabilities: [spawn]` still spawns from its shell, and can now stop that child with\n`cotal_despawn`. `--on <instance>` from a seat's shell still lands on that manager instance, now as\nthe seat.\n\n### What stops working\n\n- On a static mesh, a seat without `capabilities: [spawn]` can no longer spawn from its shell. Its\n own credential holds no spawn subject, so the broker refuses the request and the command exits 1.\n- A child launched from a seat's shell is now that seat's child, so the manager stops it when the\n seat exits, as it does for a `cotal_spawn` child. A child that has to outlive the seat that\n started it now goes with the seat.\n- A seat launched without `COTAL_SPACE` is placed by its static credential. Every connector sets\n that variable, so this only reaches a hand-built launch: from such a seat's shell, a spawn aimed at\n a static space that holds no credential for the seat is refused instead of running as the operator.\n\n### Upgrade order\n\nOnly the CLI that seats run from their shell changes, which is the one installed on the host where\nthe seats run. Brokers and managers need nothing for this break, so their order is the one the\nsections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it. A child already running when you upgrade\nkeeps the spawner the manager recorded at its launch.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the agent files whose seats run\n`cotal spawn --detach` from their shell, note which of them lack `capabilities: [spawn]`, and note\nwhich of their children must outlive the seat.\n\n### The upgrade end to end\n\n```sh\n# still on 0.61.0: find the seats that spawn from their shell\ngrep -rln 'cotal spawn' .cotal/agents\n# add `capabilities: [spawn]` to each of those agent files that lacks it, and launch any child\n# that must outlive its seat from an operator terminal instead\nnpm i -g cotal-ai@0.62.0\n```\n\nThe attribution, the despawn, the refusal of a seat without `spawn`, the stop on seat exit and a\nseat's `--on` spawn were run on a local static mesh, and the attribution and the despawn on a local\nopen mesh.\n\n## From 0.53.0 to 0.54.0\n\nManager calls now borrow an instance-bound `manager-caller` credential. Followed mutations require\n`manager.goal-result` on the selected manager, so a compatible issuer, manager and client must be\nloaded together. An older manager is refused before a followed mutation; upgrading an installed\nbinary alone does not replace code in a running manager, connector or embedded client.\n\n### Preserve state before changing processes\n\nSnapshot the broker's durable storage using its supported backup procedure, the host authority and\nactor ledgers, and each participant's manager identity, runtime custody records, credentials and\nsaved sessions. Include the embedding application's database and configuration under its supported\nbackup procedure. Record the loaded package versions and the CLI path used by bearer helpers.\nKeep these copies private. Do not change the IdP issuer, regenerate manager identities, rotate agent\ncredentials or recreate tenant storage to make the upgrade pass.\n\nNo ledger, goal-history or session conversion is required for this change. Existing ordinary\nmessaging credentials retain their normal expiry rules. New manager-caller credentials are obtained\non demand from the current grant; old manager-call credentials do not gain the new view automatically.\nExisting accepted goals remain durable and must not be submitted again merely because observation\nwas interrupted. Fresh remote registration publishes its service status at the current revision and\nepoch; do not seed that status manually.\n\n### Upgrade the split deployment\n\n1. Stage one pinned 0.54.0 package set for the host and participants, including the embedding SDKs.\n Pause new manager mutations and let accepted work settle where possible before reloading processes.\n2. Upgrade the host issuer and embedding first. Keep the broker, its account identities and durable\n storage in place. Then load the matching manager release on each participating machine.\n3. Preserve active seats through the runtime's supported update path. A Linux custodial runtime may\n release and re-adopt seats within its 600-second unattended window; verify the actual runtime,\n custody records and process identities before relying on it. A legacy PTY runtime without release\n support cannot preserve active seats through a generic manager restart. Drain it at an approved\n idle window instead of signalling the manager or replacing conversations.\n4. Reload the clients and connectors through their session-preserving host controls. Refresh any\n bearer helper captured from an older immutable CLI path. A transport-only reconnect does not reload\n JavaScript. Verify authenticated instance selection, a read-only manager command and canonical\n result recovery before allowing new followed mutations.\n\nTreat the interval from issuer reload through compatible manager/client reload as a manager-control\noutage. Mixed versions can refuse discovery or commands; there is no promised rolling transition.\nOrdinary agent sessions survive only where their runtime and credentials permit it. If verification\nfails, keep mutations paused and repair forward from the preserved state rather than resetting it.\nThis release does not add host-backed enrollment or terminal release for stock participant detached\nagents; see [Remote supervised agents](run-a-mesh.md#remote-supervised-agents).\n\n## From 0.48.2 to 0.49.0\n\n0.49.0 changes how a credential's authority is recorded. A credential is no longer only a signed\nfile: it is an *issuance*, with a generation the issuer chose and durable evidence of the ceiling it\nwas granted under. The important consequence for a running deployment is not at connect time. It is\nat renewal time.\n\n### What keeps working without any action\n\n- **Existing agent credentials keep authenticating.** A credential minted under 0.48.2 is not\n revoked and is not rejected at connect. Nothing needs to be re-issued to bring the fleet back up\n after the upgrade.\n- **The channel registry survives.** Channels, their replay settings, descriptions, and usage text\n are ordinary durable state and are not rewritten by the upgrade.\n- **`cotal deliver` is still a standalone command.** Running the delivery daemon as its own process\n remains supported; it is not restricted to being a child of `cotal up`.\n- **`cotal join` keeps its flags.** In particular `--lifecycle-uid` is not new in 0.49.0. It has\n been required alongside `--creds` since well before this release, and the pairing rule did not\n change here. A scripted external join that worked under 0.48.2 works unchanged.\n\n### What does not migrate\n\n**A credential minted before 0.49.0 cannot be renewed.** Managed agent credentials carry a\n24-hour lifetime and the manager re-signs one once it passes **75%** of its life, ticking every\nquarter of the TTL so a tick always lands inside that window. When the manager reaches a credential\nthat carries no issuance, it refuses to renew it and logs the agent by name:\n\n```\n! managed cred renewal <agent>: renewManagedStaticCred: <agent> carries no issuance;\n a static credential minted before SPEC 13.15 is not renewed under an unbound generation\n - respawn the agent\n - the agent dies loud at this cred's expiry unless it is reminted\n```\n\nSo the fleet comes up fine, runs normally, and then each agent stops at its own credential's\nexpiry, within roughly a day of the upgrade, one at a time rather than together. The refusal is\ndeliberate: the renewal would otherwise have to invent a generation nobody issued, which is the\nstate the release exists to remove.\n\n**Respawn the managed agents as the last step of the upgrade.** For this particular upgrade the\nrespawn is not optional: stopping a 0.48.2 manager ends its agent processes whichever CLI you use,\nfor the reason given under the outage window below. The respawn is how they come back, and it is\nalso what mints each credential as an issuance so it renews from then on. One planned pass over the\nfleet is the whole job. Skipping it leaves agents stopped and, for any credential that survived\ninto 0.49.0 unminted, brings the renewal cliff above a day later, one agent at a time.\n\n### Credentials you minted yourself\n\n**A credential you minted with `cotal mint` is a different case, and it very likely needs\nnothing.** The distinction that matters is not the word \"static\", which covers both. It is **what\nminted the credential and who owns its renewal**. A credential the **manager** minted for an agent\nit spawned carries a lifetime and is renewed by the manager, so it is the subject of everything\nabove. A credential **you** minted with `cotal mint` and handed to an external peer is issued with\n**no expiry at all**, and no manager renews it: it is not in the sweep, so there is no renewal to\nfail. It keeps working after the upgrade, and re-minting it would mean coordinating with a third\nparty for no gain.\n\nThe manager says which one it is holding. Where a credential has no expiry to reach, the sweep\nnames it and moves on rather than refusing:\n\n```\n! managed cred renewal <agent>: credential is unbounded - not renewed\n (a pre-TTL credential stays as minted until respawn)\n```\n\nRe-mint an external peer's credential only if you want it to carry a lifetime, and at a time you\nchoose.\n\n### How to read the boot log\n\nA 0.49.0 manager starting over an existing space may print lines like:\n\n```\n verified evicted: <holder-key> (3/12)\n already verified (durable): <holder-key>\n✓ boot self-heal: manager/<id> registration gate reopened at generation <n>\n```\n\nThese are **not** a credential migration, and reading them as one is the most likely way to\nconclude the fleet is fine when it is not. They come from the manager repairing **one** endpoint\nregistration gate that a previous restart left frozen, and they enumerate that single gate's\ncredential-family holders as it verifies each one evicted. `already verified (durable)` on a later\nstart is the repair cursor resuming, not a credential that became durable. The repair is real and\nuseful (it is what previously needed `cotal reconcile-gate` by hand), but it says nothing about\nwhether your agent credentials carry issuances. The renewal refusal above is the signal that does.\n\n### Which side to upgrade first in a split topology\n\nMove the manager first.\n\nThe stores 0.49.0 introduces are created by the **manager** at its own boot, not by the broker.\nThey are create-or-verify and idempotent, so a 0.49.0 manager brings the space's authority stores\nup to the new shape itself, and it does so against whichever broker is answering.\n\nBeing honest about the evidence behind each direction, because they are not equally established:\n\n- **Broker-first was measured on a live 30-agent deployment** (issue #1578). Upgrading the broker\n first locks the old manager out immediately: `cotal up` re-renders the broker's generated config\n from the trust record, and after the restart the still-0.48.2 manager is refused on every\n connection with an `authentication error` naming the Nkey, continuously. That text comes from the\n broker process, not from a Cotal command, so match on its shape rather than on an exact string.\n `cotal ps` reports zero agents while\n the agent processes are still alive, because the manager has lost its view of them, not because\n they died. Upgrading the manager clears it immediately.\n- **Manager-first is reasoned from where the new stores are provisioned**, not from a measured\n fleet upgrade. It is the recommended order because the manager is the component that creates what\n 0.49.0 adds, but it has not been run end to end on a production split topology at the time of\n writing. Treat it as the better-supported order rather than a guaranteed one, and keep the\n rollback below ready either way.\n\nWhichever order you pick, **this is not a rolling upgrade**. Between the two steps the mesh is down\nand the manager cannot see its agents. Go straight through rather than pausing between them, and\nschedule it as an outage window.\n\n### What the window looks like\n\n- **The managed agent processes do not survive step 1, in either order.** This is the one place\n where the obvious reordering does not rescue you, so it is worth understanding rather than\n working around. Sparing agents on a bare manager stop is a **handshake**: a 0.49.0 manager\n publishes a capability file proving it can release its agents, and a 0.49.0 `cotal down` refuses\n the stop unless it finds one. **A 0.48.2 manager never publishes that file**, because the\n mechanism ships in the release you are installing. So the old CLI against the old manager sends a\n plain stop and takes every seat with it, and the new CLI against the old manager either refuses\n (leaving `--with-agents`, which reaps deliberately) or falls to the legacy path, warns that it\n cannot verify the manager can spare its agents, and signals it anyway.\n- **You can confirm which side you are on in one command, without stopping anything.** The flag that\n marks the newer behaviour is absent from the older CLI, and its summary line makes the difference\n plain:\n\n ```\n $ cotal down --help # on 0.48.2\n cotal down - stop the whole local stack, or name only the components to stop\n\n $ cotal down --help # on 0.49.0\n cotal down - stop the whole local stack (managed agents stay running unless --with-agents), ...\n ```\n\n If your `cotal down --help` does not mention `--with-agents`, stopping the manager stops the\n agents with it.\n- **Therefore the respawn in step 5 is mandatory recovery for this upgrade, not an optional pass.**\n It is also the step that re-mints credentials as issuances, so it is the same action either way.\n Plan the window to include it rather than treating it as cleanup.\n- The **manager's view** of them is lost while the two sides disagree, so `cotal ps` reports zero\n and control commands do not reach seats.\n- **Messages are not delivered** while the mesh is down.\n- The window is as long as it takes to restart the second component, plus the manager's own start.\n It is minutes, not hours, provided you do not stop between the steps.\n- **Nothing self-heals if you stop halfway.** The refusal is continuous until both sides match.\n\n### Snapshot this before you start\n\nTake these while the deployment is still on 0.48.2. The two `cotal` reads are live reads and must\nhappen before anything stops.\n\n- **A filesystem or volume snapshot of both containers**, if your platform offers one. This is the\n only rollback that covers every case, and it is what the reporting deployment used.\n- **`cotal backup create <dir>`**, for the durable space state, **but read the next paragraph before\n you rely on it**: on a split broker and manager topology it is very likely unavailable to you, and\n the volume snapshot above is your actual rollback.\n- **The trust records and credential directory** under `.cotal/auth` on the manager host, including\n the per-space material directory. These are what a re-mint would otherwise have to replace.\n- **A copy of the channel registry**, so you can verify it came back rather than assuming it did:\n `cotal channels list` before and after.\n- **The output of `cotal ps`**, so you know how many seats you expect to see afterwards and can tell\n a lost view from a lost agent.\n\n#### `cotal backup` on a split topology\n\n**`cotal backup create` cannot read a running stack.** It requires a completed cut, and only\n`cotal down --preserve-state` publishes one:\n\n```\n$ cotal backup create ./backup.0482\n✗ backup requires a completed cut; run `cotal down --preserve-state` first\n```\n\n**And `cotal down --preserve-state` requires a manager alive on the host you run it from.** It uses\nthat manager to attest that every retained child stopped, and the check is deliberately fail-closed:\na manager that is dead or merely uncertain refuses rather than preserving an unproven cut. The check\nreads a local pidfile, so a **remote** manager does not satisfy it. On a split topology the broker\nhost has no local manager, which means the documented durable-backup path is not available there.\n\n**Measured rather than assumed, at 0.48.2**: the backup refusal above is executed output. The\npreservation requirement is read from `down.ts` at the same tag, where the preserve path asks a\nmanager to prepare an inventory and then requires that manager to be locally alive before it\ncommits. The part not executed end to end is a genuine two-host split, which needs two real hosts.\n\n**What to do instead.** Use the filesystem or volume snapshot of both containers. That is the\nrollback the reporting deployment actually used, it covers the broker's durable state and the\nmanager's credential material together, and it does not depend on either component being able to\nattest for the other. If you want `cotal backup` as well, take it from a host that does have a live\nlocal manager, and understand it is a second copy rather than the primary rollback.\n\n**This looks like a product limitation rather than a documentation gap**, and it is written here as\none so an operator is not left thinking they mis-typed a command. The upgrade path for the exact\ntopology this page is addressed to cannot use the documented backup command.\n\n### The upgrade end to end\n\n```bash\n# 0. on 0.48.2, STILL RUNNING: record what you expect to see afterwards.\n# These two are live reads, so they must happen before anything stops.\ncotal channels list > channels.before\ncotal ps > ps.before\n\n# 1. manager host. READ THE NOTE BELOW THE BLOCK FIRST: this step ends the\n# managed agent processes whichever order you choose, and the respawn in\n# step 5 is how they come back. It is recovery, not tidying.\n#\n# STOP THE MANAGER WITH THE 0.48.2 CLI, BEFORE INSTALLING 0.49.0. The\n# order matters and it is not recoverable once you install: a 0.49.0\n# `down manager` REFUSES to stop a 0.48.2 manager whose pid record carries\n# a start token, which is every manager on a platform that can read one\n# (Linux can):\n# refusing bare manager stop: ... does not prove this manager can detach\n# its agents; use --with-agents or stop the agents explicitly\n# The refusal names two remedies and NEITHER clears it for this case. The\n# check reads a capability file that only a 0.49.0 manager writes; it never\n# counts agents, so stopping them first changes nothing. And `--with-agents`\n# is whole-stack only, so `down manager --with-agents` is refused by its own\n# flag rule. See #1592.\ncotal down manager # the 0.48.2 CLI, still installed.\n # 0.48.2 has no --with-agents; this\n # is the whole route. On a host that\n # runs the whole stack, the 0.49.0\n # `cotal down --with-agents` after\n # installing is the alternative.\nnpm install -g cotal-ai@0.49.0 # ONLY after the stop above\n# `supervise` RUNS IN THE FOREGROUND and holds the terminal until you stop\n# it. There is no --detach on this command. Start it under whatever keeps\n# your manager alive normally (systemd unit, container entrypoint, or a\n# second terminal), and run the remaining steps from another shell.\ncotal supervise --space <space> --server nats://<broker>:4222\n\n# 2. broker host: stop the stack.\n# NOT `--preserve-state` on a split topology: it needs a manager alive on\n# THIS host to attest its children stopped, and yours is on the other one.\n# Your rollback is the volume snapshot from \"Snapshot this before you\n# start\", not `cotal backup`.\n# See \"cotal backup on a split topology\" above.\ncotal down\n\n# 3. broker host: install 0.49.0 and start it again\nnpm install -g cotal-ai@0.49.0\n# Record the manager log's size BEFORE starting, so step 3a can tell THIS\n# boot's output from every earlier one. It must be captured here, ahead of\n# the start: taken afterwards it sits past the new line and the wait hangs.\n# `<spaceKey>` is NOT the space name. It is lowercase hex of the name's\n# UTF-8 bytes, so space `prod` is `manager.70726f64.log`. Do not guess it:\n# `cotal up` prints the real path on its launch line. Substituting the\n# plain name points at a file that does not exist, and the wait below then\n# burns its full timeout before telling you.\nLOG=.cotal/manager.<spaceKey>.log\nOFF=$( [ -f \"$LOG\" ] && wc -c < \"$LOG\" || echo 0 )\ncotal up --detach --host 0.0.0.0 --space <space> --no-manager\n\n# 3a. SPLIT TOPOLOGY ONLY: `--no-manager` above boots the broker (and the\n# delivery daemon) with NO local manager on the broker host, so there is\n# no wait-and-stop step on a current cotal-ai. The rest of this step is\n# the OLDER-host recipe, kept because the flag is refused there and that\n# refusal is your signal you are on it: without the flag the `up` also\n# starts a local manager, and you must wait for the log to show it is up,\n# then stop it, or you finish the upgrade with two managers and the one\n# you did not intend is the one nobody is watching.\n# A bare `grep -q` does NOT wait: it reads once and exits 1 immediately\n# if the line has not been written yet. Bound the wait instead, so a\n# manager that never comes up fails loudly rather than reading as ready.\n# The log is opened APPEND-ONLY, so on any host that has run a manager\n# before, this file ALREADY carries a `manager up` line from an earlier\n# boot. Grepping the whole file therefore matches instantly and waits for\n# nothing. Read only what THIS boot appended, using the $OFF captured in\n# step 3 above (before the start, which is the only point it is correct):\ntimeout 60 bash -c \\\n \"until tail -c +$((OFF+1)) \\\"$LOG\\\" | grep -q '. manager up'; do sleep 1; done\"\n# exit 0 = THIS boot logged it; exit 124 = it never did, so STOP and look.\n# This manager is 0.49.0 and publishes its own spare-capability file, so\n# the bare stop below is NOT the refusal case from step 1.\ncotal down manager # broker + delivery remain\n# On a current cotal-ai the two commands above are unnecessary (nothing\n# to wait for, nothing to stop) and `cotal down manager` simply reports\n# no manager to stop.\n\n# 4. verify the mesh is whole again before touching the fleet.\n# Do NOT compare `cotal ps` against ps.before yet: step 1 ended the agent\n# processes, so at this point it is EXPECTED to be empty, and an empty\n# `ps` is also the signature of the broker/manager mismatch described\n# above. The two are indistinguishable here, so compare what the mesh\n# itself should have carried across instead:\ncotal channels list # compare against channels.before: this SHOULD match now\ncotal ps # expect it to be EMPTY here; ps.before is the target for\n # step 5, not for this step\n\n# 5. the step that is easy to skip: respawn the managed agents so their\n# credentials are re-minted as issuances and can renew. Persona is a\n# POSITIONAL argument here, unlike `cotal stop`, which requires --name.\n# One call per agent:\ncotal spawn <persona> --detach --name <n> --space <space>\n# then the comparison step 4 could not make:\ncotal ps # NOW compare against ps.before: seat count should match\n```\n\nThe mesh is down from step 2 until step 3 finishes. That is the window. On a split topology there is\nno cut and no backup inside it, so the window is the stop, the install and the restart, nothing more.\n\n## Adding a section for a future release\n\n**Every changeset marked breaking adds a section to this page.** A release that changes what an\noperator must do, in what order, or what stops working, is not finished until the section exists.\n`scripts/upgrade-section-gate.mjs` grades a commit range for this: run it as\n`pnpm upgrade-section-gate --base <ref>` and it reds when the range carries a breaking change and\nadds no new release section. CI runs its self-test and, as a step of the `attribution` job, grades\neach pull request's own range as `HEAD^1..HEAD` over the merge snapshot it checked out. That job is\nthe only context in the branch protection rule set, so a red gate FAILS A REQUIRED CHECK AND BLOCKS\nTHE MERGE. The section is not optional and a reviewer cannot wave it through without an\nadministrator overriding branch protection. Be precise about what the check proves either\nway, because one trusted past its evidence is worse than none. It proves a section for a release\n**was written here**. It cannot prove the section is **correct**, or that it describes the break\nthat actually landed, and it cannot see a breaking change that carries no marker at all. Reviewing\nthe words remains a person's job.\n\n**Mark the break, or the gate cannot see it.** Any one of these is enough, and they are the only\nthings it reads:\n\n- a `!` before the colon in the commit subject, as in `feat(core)!: bind hosted runs to the caller`\n- a `BREAKING CHANGE:` footer in the commit body\n- a changeset in `.changeset/` declaring a `major` bump for any package\n\nThe marker must survive the squash. A `!` that lives only in a commit you squash away is not in the\nrange the gate grades, so put it in the subject that lands on `main`.\n\n**The heading is a `##` and names the release**, like `## From 0.48.2 to 0.49.0`. Both matter, and\nneither is a style preference. Coverage is claimed by a heading, so a heading that names\nno release claims every release and distinguishes none: `## Notes` with a sentence under it would\notherwise satisfy the rule. Naming the release also makes the section the one an operator upgrading\nthat release will search for. Use `###` freely for detail inside a section. Subsections belong to\ntheir release rather than counting as separate coverage.\n\nName the release that first carries the change: the next version Changesets publishes, which\n`pnpm changeset status --verbose` lists. `bin/package.json` on `main` still reads the release already\npublished. If a release is cut while the change is open, the change ships in the release after it,\nso move the heading before merging. The gate accepts any version in a heading, so before merging a\nrelease pull request, check every heading added since the previous tag against the version it\npublishes.\n\nA section is written for the operator, not for the reviewer. It answers, in this order:\n\n1. What keeps working with no action at all.\n2. What does **not** migrate, and when that becomes visible. Name the log line if there is one.\n3. The order to move components in for a split topology, and why that order.\n4. What the outage window looks like, including what survives it.\n5. What to snapshot before starting.\n6. The commands, end to end.\n\n**Where an answer was not measured, say so in the document rather than guessing.** An operator who\nknows which half of a recommendation is reasoned and which is measured can plan around it; one who\nfinds out afterwards cannot.\n"
|
|
272
|
+
"body": "# Upgrading a running deployment\n\n> **Guide** (informative) · **For:** operators upgrading a mesh that already exists · **See also:** [Substrate stability](stability.md), [Run a mesh](run-a-mesh.md), [Identity and auth](identity-and-auth.md)\n\n[Substrate stability](stability.md) tells you what the version numbers promise. This page is the\nother half: what to actually do when the deployment already exists, has credentials in it, and\ncannot simply be recreated. Every release that breaks a running deployment gets a section here,\nnaming what migrates on its own, what does not, and the order to move the pieces in.\n\n## The pre-1.0 upgrade contract\n\nThe packages are pre-1.0, so a minor bump may break an API or an on-disk expectation. Four\ncommitments make that survivable for someone with a fleet:\n\n- **Pin an exact version.** `0.N.P`, never `^0.N.P`. A range can pull a breaking minor in during an\n unrelated reinstall.\n- **Every break that touches a running deployment gets a section on this page**, written in terms of\n what an operator does, not in terms of which module changed.\n- **Read the section before you start, not halfway through.** A section names the work up front\n precisely so the operation does not change shape once it is underway.\n- **A break that cannot be made automatic says so.** Where credentials or state must be recreated by\n hand, the section says which ones and when, rather than leaving you to discover it at the moment\n the first one stops working.\n- **A change to the shape of a credential, or to who may renew one, is breaking whatever the commit\n marker says.** This rule is stated because the marker is a judgement made while writing the code\n and the consequence is felt by someone running it a day later. A fleet that keeps authenticating\n looks compatible and is not, if nothing in it can renew. Any automated check of this rule would\n read commit markers, so a break recorded as a feature is the one case it could not see, which is\n why the rule is written for people first. **The marker held for this release: the 0.49.0 change\n that caused all of this, `36d177951 feat(core)!`, did carry its `!`.** The rule exists for the\n next one that does not.\n\nWhat this page does not promise is a rolling upgrade. Nothing in the current line dual-serves two\nauthority versions, so where broker and manager run separately there is a window in which the mesh\nis down. The sections below give that window's shape so it can be scheduled rather than endured.\n\n## Auth context closure in 0.71.0 (unreleased)\n\nExisting deployments need no credential migration or restart for these additive APIs. Embedded\nhosts can now inspect `handle.connections()` and await `handle.closed` after `close()` or `drain()`\nto prove every owned transport ended, including the callout, replaced readiness readers and\nshort-lived clients. The inventory is a detached snapshot.\n\nA transport close failure now rejects with its connection label. The terminal signal stays pending\nwhile any connection remains live. Repair the failure and retry `close()` before awaiting\n`handle.closed`. Closing one hosted context does not close another account's context.\n\nRead a space's claim with `readPlaneClaim(kv, space)` on that account's leader-only auth bucket.\nAn unclaimed space returns `undefined`; held and released rows retain their claim identity.\nDeleted, malformed and foreign-space rows refuse. `PlaneClaimRow` and `PLANE_CLAIM_KEY` are exported.\n\nUse `observeAccountLivenessWithCreds({ servers, observerCreds, accountId, options })` with the\naccount-scoped membership-observer credential to list that account's connections. It never widens\ncredentials or evicts connections. Zero rows prove absence only with a complete sweep and the\nsingle-server proof. An embedded endpoint's trusted composition can retain transport custody\nthrough `EndpointOptions.onConnection`.\n\n## Unreleased\n\nOn a per-user-auth mesh, a spawn-scoped caller can arm the event plane of a child under its own\nowner without `admin`. This fixes owned spawns in spaces whose registration policy requires the\nplane. Upgrade the manager to pick up the admission change. No credential or state migration is\nneeded. Cross-owner arming still requires `admin`, and the child's own-channel rule and ledger\nenvelope are unchanged. A silent non-owner caller in a space without the policy still has the\nplane disarmed, with a notice if provisioning succeeds.\n\n## Hermes model from the environment in 0.68.0\n\nA connector now launches on the model and variant its launcher resolved (the `--model` or\n`--variant` flag, else the agent file's `model:` or `variant:`) and no longer reads them again from\nthe agent file. The Hermes connector also no longer takes a model from `HERMES_MODEL` in the\nenvironment of the process that spawns the seat, including when `spawn.env` lists it.\n\n### What stops working\n\nA Hermes spawn whose only model was `HERMES_MODEL` in the spawning environment is refused at launch,\nand the refusal names both ways to set a model. Spawns that set `--model` or `model:` are unchanged,\non every connector.\n\nCode that calls a connector's `buildLaunch` directly with only `configPath` now gets no model or\nvariant from that file. Pass them as `model` and `variant`.\n\n### Before the upgrade\n\nMove each Hermes seat's model from `HERMES_MODEL` onto its spawn with `--model`, or into its\npersona's `model:`.\n\n## Run answers on a participant manager in 0.68.0\n\nA participant manager now asks its issuing host for an answering credential by naming the run and\nstep it answers. The host reads the pause's token off that run's journal and no longer accepts a\ntoken from the manager. Runs on a mesh with no participant manager are unaffected.\n\n### What stops working\n\nWhile a participant manager and its issuing host run different sides of this release, the host\nrefuses every `cotal run answer` and every amendment that manager serves, because each side refuses\nthe other's request shape. Starting, resuming and reading runs is unchanged. A pause stays waiting\nthrough the window, or follows its timeout if it has one.\n\n### Before the upgrade\n\nUpgrade the auth service and every participant manager registered with it in the same window, then\nanswer the pauses that waited.\n\n## Headless OpenCode handshake in 0.69.0\n\nWith `COTAL_SERVE_HEADLESS=1`, the OpenCode launcher's `[cotal-serve]` line on stdout now carries\nonly `port` and `session`. The server password no longer appears in it, and the 1.x TUI no longer\nreceives the password on its command line.\n\n### What stops working\n\nA headless host that read `password` from that line has no password, and the server refuses its\nrequests. Seats with a TUI, and headless seats that no host drives, are unaffected.\n\n### Before the upgrade\n\nHave each headless host mint a password and pass it to the launcher as `OPENCODE_SERVER_PASSWORD`,\nthen use it for basic auth as before. Without that variable the launcher mints its own.\n\n## Filesystem store identity in 0.69.0\n\nThe delivery daemon's answer to the manager's store check now names a filesystem store by its root\nand by a random `id` that the store records once in `store.id` inside its own directory:\n`.cotal/store.id` for a workspace root, or the directory of the file for `cotal deliver --creds\n<file>`. A manager no longer counts the daemon's store as its own because the two roots have the\nsame path. On a split whose broker host and manager host use one root path, the manager host now\nstays off the daemon-credential renewal lease, so `cotal doctor auth --fix` on the broker host can\nrenew the daemon credentials.\n\n### What stops working\n\nA manager and a delivery daemon on different sides of this release refuse each other's answer to\nthe store check. The manager then remints no daemon credential, and a manager that is booting does\nnot start. This is read from the code and was not measured across two releases. A\n`cotal deliver --creds <file>` whose directory is a read-only mount and holds no `store.id` stops at\nstart. So does a `--creds` file that is its directory's `store.id` under any name, and a `store.id`\nthat is a symbolic link or holds anything but a lowercase UUID.\n\n### Before the upgrade\n\nUpgrade the broker host and every manager host of a space in the same window. For a `--creds` file\non a read-only mount, add a regular `store.id` file beside it that holds a new lowercase UUID and no newline,\nas `node -e 'process.stdout.write(crypto.randomUUID())' > store.id` writes. Move a `--creds` file\nnamed or linked as `store.id` to a file of its own.\n\n## Detached spawns with `--share-tools` in 0.69.0\n\nThe manager's `spawn` operation now takes `shareTools` as a list of MCP server names. The CLI parses\n`--share-tools` into that list before it sends the request, and the manager cluster document moves\nto revision 22. A cut taken with `cotal down --preserve-state` before the upgrade still resumes: the\nmanager reads its `cotal-manager-resume/v1` inventory and writes new cuts as\n`cotal-manager-resume/v2`.\n\n### What stops working\n\nA CLI and a manager on different sides of this release refuse a detached spawn that passes\n`--share-tools`, because the CLI checks each request against the contract the manager serves. This\nis read from the code and was not measured across two releases. A detached spawn without the flag,\na foreground spawn and a roster entry are unaffected. A manager older than this release cannot\nresume a cut that this release took.\n\n### Before the upgrade\n\nUpgrade the CLI on every host that runs `cotal spawn --detach` in the same window as the managers\nit reaches.\n\n## Shared MCP server checks in 0.69.0\n\nThe cotal config reader now checks each server under `connectors.<name>.mcpServers` when it reads\nthe file, and refuses one that cannot launch as written, naming the file and the field. The rules\nare in [the config file](config.md#the-config-file).\n\n### What stops working\n\nA config file that holds such a server refuses every Claude spawn that reads it, including one with\n`--share-tools none`. Before, a field of the wrong type failed each Claude spawn that shared the\nserver with a `TypeError` that named neither the file nor the server, a spawn that did not share it\nlaunched, and a server with no `command` or `url` was passed to `claude`, which never started it.\nRead from the code and not measured: spawns on other connectors, a manager resume and the step of\n`cotal setup` that records the shared list read the same files, so each stops at the same refusal.\n\n### Before the upgrade\n\nCheck `connectors.<name>.mcpServers` in the operator-level config file and in each space's\n`.cotal/config.json`. Give each server a string `command`, or a `type` of `http`, `sse` or `ws` with\na string `url`. Write `args` as a list of strings and `env` and `headers` as objects of strings, or\nremove the server.\n\n## Remote manager family eviction in 0.69.0\n\nA remote manager registered through its host now asks the host to evict up to 256 holders of its\ncredential family in one maintenance request, and the host reads the family once for the whole set.\nBefore, a restart sent one request per holder and the host read the whole family for each one.\nMeshes with no remote manager are unaffected.\n\n### What stops working\n\nWhile a remote manager and its issuing host run different sides of this release, each side refuses\nthe other's eviction request shape. A restart whose credential family already has holders then fails\nat its eviction step and leaves the manager's registration gate frozen. A first start, a clean stop\nand the host's reconciliation of a foreign slot holder are unchanged.\n\n### Before the upgrade\n\nUpgrade the auth service and every remote manager registered with it in the same window. A manager\nthat restarted inside the window resumes its frozen registration on its next start once both sides\nrun this release.\n\n## AG-UI emitter holder hooks in 0.69.0\n\n`AguiEmitterHolder` from `@cotal-ai/connector-core` now takes its hooks as one named object after\nthe emitter factory: `new AguiEmitterHolder(startEmitter, { onError, onRunClosed, waitLive, runMeta })`.\nOnly `onError` is required. Nothing about a running mesh changes, and every shipped connector passes\nits hooks by name. Only a connector of your own that builds a holder is affected.\n\n### What stops working\n\nA holder built with positional hooks, such as `new AguiEmitterHolder(start, onError, onRunClosed)`,\nno longer compiles, because the constructor takes two arguments. Plain JavaScript that keeps the\npositional form still runs, but the holder calls none of its hooks, so a failure never reaches\n`onError`.\n\n### Before the upgrade\n\nPass each hook by name, for example `new AguiEmitterHolder(start, { onError, onRunClosed })`, and\ndrop any `undefined` that filled an earlier slot to reach a later hook.\n\n## Worker run failure type in 0.70.0\n\n`WorkerRunFailed`, the failed result of `runInWorker` in `@cotal-ai/lang`, is now a union on\n`class`: `released`, `held`, `effect`, `too-large`, `rejected` or `error`. A running mesh needs\nnothing, because the runtime host and the engine thread ship in the same install. A run on the\ncompiled engine whose program throws an object with `code: \"L5012\"` or `code: \"L5025\"` used to end\nreleased and now ends failed, as it does on the walker.\n\n### What stops working\n\nTypeScript code that reads `code`, `reason`, `step`, `pending`, `kind`, `detail` or `tooLarge` on a\n`WorkerRunFailed` it has not narrowed fails with TS2339. A released, held, too-large or rejected\nresult no longer carries `code`, so JavaScript that branched on `L5012`, `L5025`, `L5006` or\n`L5010` stops matching with no error. `tooLarge` is gone.\n\n### Before the upgrade\n\nBranch on `class` where such code read `code`: `released` for L5012, `held` for L5025, `too-large`\nfor L5006 and `rejected` for L5010. An `effect` or `error` result keeps its `code`. Once narrowed to\n`too-large`, a result carries the `stepKey`, `bytes` and `bound` that `tooLarge` held.\n\n## Remote manager request builder in 0.70.0\n\n`remoteManagerClient.remoteManagerAuthorityRequest` from `@cotal-ai/manager` now takes an\noperation's coordinates as one object, and `remoteManagerRegistrationProof` from `@cotal-ai/core`\ncomputes the proof from the manager's identity state instead of a request. Nothing about a running\nmesh changes: the proof digest and the request on the wire are the same, so a manager and a host on\ndifferent sides of this release still accept each other. Only code that builds remote manager\nrequests itself is affected, in TypeScript and in plain JavaScript.\n\n### What stops working\n\nA call that passes the registration proof, contract artifacts, session, retirement or transfer\nreader as positional arguments after the operation no longer compiles. A call that passes a request\nto `remoteManagerRegistrationProof` no longer compiles either, because the second argument now names\nthe lifecycle `lifecycleUid`, as the identity state does.\n\nPlain JavaScript runs both old calls without an error. The builder drops the positional coordinates,\nso the host refuses the request with `requires a sha256 registrationProof`. A proof computed from a\nrequest leaves out the lifecycle, so the host refuses a request that carries it as a proof mismatch.\n\n### Before the upgrade\n\nName the coordinates, for example\n`remoteManagerAuthorityRequest(state, \"cli\", \"retire\", { registrationProof, retirement })`.\nCompute the proof as `remoteManagerRegistrationProof(owner, state)`, adding the contract artifacts\nas a third argument for activation only. A host that recomputes the proof from a received request\npasses `{ space, instanceId, lifecycleUid: managerLifecycleUid, identities }` from that request.\n\n## Bearer validator lifetime cap in 0.70.0\n\n`validateUserToken` from `@cotal-ai/auth` no longer takes `maxTtlSec`. It caps a bearer's lifetime\nat the cap of the bearer's view, the same cap the issuer applies when it mints: 900 seconds, or 300\nfor a `transfer-writer` bearer. The auth callout never passed the option, so a running mesh behaves\nas before. Only code of your own that calls the validator with `maxTtlSec` is affected.\n\n### What stops working\n\nA call that passes `maxTtlSec` in an object literal no longer compiles. Plain JavaScript that keeps\nit still runs, and the value is ignored. A `NaN` value, such as `Number()` of an unset environment\nvariable, used to turn the lifetime check off and accept a bearer of any lifetime. That bearer is\nnow refused at its view's cap.\n\n### Before the upgrade\n\nRemove `maxTtlSec` from each call. A test that needs a bearer to expire sooner mints one with a\nshorter lifetime.\n\n## Persisted identity records in 0.70.0\n\nThe manager instance identity, the manager sibling identities, the auth plane instance identity and\na participant manager's remote authority state now share one reader and one first mint in\n`@cotal-ai/workspace`, exported as `claimIdentityRecord` with the nkey check `identityOf`. Each\nrecord is read as a regular file, must hold non-empty nkeys and is created exclusively, so\nconcurrent first starts of a participant manager on one root now settle on one identity where each\nused to keep its own. `saveManagerInstanceIdentity` and `saveAuthInstanceIdentity` are gone. A\nrunning mesh whose records are plain files needs nothing.\n\n### What stops working\n\nA manager instance, auth instance or remote authority record that is a symlink, a directory or any\nother non-regular entry is refused where it used to be followed. The manager, the auth plane and a\nparticipant manager fail to start on it, and `cotal reconcile-gate` and `cotal deregister-instance`\nrefuse it. Retirement already refused it. A remote authority record with an empty nkey id or seed\nis refused too. A first mint that loses its race and cannot read the winner now refuses with\n`identity-record-create-lost` in place of `manager-instance-identity-create-lost` or\n`auth-instance-identity-create-lost`. Code that imports either `save` function no longer compiles.\n\n### Before the upgrade\n\nReplace a symlinked identity record with a copy of the file it points to. Code that wrote a record\nwith a `save` function plants it with `createManagerInstanceIdentity` or\n`createAuthInstanceIdentity`, which create the record when it is absent and otherwise return the\nstored one unchanged. Nothing replaces an overwrite of a stored identity.\n\n## Manager instance in user credentials in 0.70.0\n\n`AuthProvider.userCredentials` from `@cotal-ai/core` no longer returns `managerInstanceId`. A\n`manager-caller` credential's manager instance is the signed `act.managerInstanceId` claim in its\nbearer, which the broker verifies and the CLI already used. The reference provider in\n`@cotal-ai/auth` stops copying the exchange response's field into its result, where nothing\ncompared it with the bearer. The exchange still answers with the field, so a running mesh behaves as\nbefore.\n\n### What stops working\n\nCode of your own that reads `managerInstanceId` from a `userCredentials` result no longer compiles,\nand plain JavaScript reads `undefined` there.\n\n### Before the upgrade\n\nRead the instance from the bearer's `act.managerInstanceId` claim.\n\n## Auth plane identity location in 0.70.0\n\nThe user-auth service keeps its instance identity in the root's `.cotal/space.<hex>/auth-instance.json`,\nbeside the manager's. It used to sit inside `.cotal/auth`, at\n`space.<hex>/.cotal/auth/auth-instance.<hex>.json`, so a copy of that folder carried it. The first\nstart of an upgraded root moves the record and keeps the instance. A hosted context started through\n`startAuthService` has its record moved the same way inside its `stateDir`.\n\n### What stops working\n\nCode that calls `openAuthAuthorityPlane` without the new `identityRoot` option no longer compiles. A\nstart that finds a record both in `.cotal/space.<hex>/` and at its older place refuses and names the\ntwo files. A start also refuses when the older place of the auth or manager identity holds a symlink,\na directory or anything else that is not a regular file. The manager used to skip a dangling symlink\nthere and mint a new identity.\n\n### Before the upgrade\n\nPass `identityRoot` to `openAuthAuthorityPlane`. When `dir` is a workspace root's user-auth state\ndir, `<root>/.cotal/auth/space.<hex>`, pass that root. A plane with no workspace root, as\n`startAuthService` runs, passes `dir` itself. Either keeps the identity the plane already has: on the\nfirst start it moves from `<dir>/.cotal/auth/` to `<identityRoot>/.cotal/space.<hex>/`. Never pass a\ndirectory inside `.cotal/auth`: the record would land in the folder an operator copies and travel\nwith it again.\n\nA copy of `.cotal/auth` taken from a root last run by an older Cotal carries that root's record.\nDelete `.cotal/auth/space.<hex>/.cotal/auth/auth-instance.<hex>.json` from the root you copied it to\nbefore the first `cotal up --user-auth` there.\n\n## Per-seat `COTAL_` names in `spawn.env` in 0.71.0\n\n`spawn.env` in the cotal config no longer forwards a `COTAL_` name the launcher sets for each seat,\nsuch as `COTAL_ROLE`, `COTAL_MODEL` or `COTAL_SUBSCRIBE`. Before, a seat launched with no value of\nits own took the spawning process's value and ran under that role, model or read set. The\nmachine-wide knobs a seat already receives, such as `COTAL_HOME`, may still be listed.\n\n### What stops working\n\nEvery spawn and resume under a config whose `spawn.env` lists such a name is refused before\nlaunch, and the refusal names the entry. Code that calls `launchEnv` from `@cotal-ai/connector-core`\nwith such a name in `envAllow` gets the same error.\n\n### Before the upgrade\n\nRemove those names from `spawn.env`. Give each seat its role, model and channels with `--role`,\n`--model` and `--subscribe`, or in its persona's `role:`, `model:` and `subscribe:`.\n\n## Role addresses in 0.71.0\n\nA role must be one `[A-Za-z0-9_-]` token. Before 0.71.0 any other spelling was rewritten into one:\n` probe ` reached the `probe` queue and `pro.be` reached `pro_be`, while the message kept the\nspelling sent. An anycast to `*` was accepted and stored where no holder reads it.\n\n### What stops working\n\nAn agent whose role is outside the token set no longer starts, however it is launched:\n`cotal join --role`, `cotal spawn --role`, an agent file's `role:`, `COTAL_ROLE` and an embedded\nendpoint's `card.role` are all refused before the agent joins.\n\nA send to such a role, or to `*`, through `cotal send ask`, `/anycast` or `cotal_anycast` is refused,\nand nothing is stored.\n\n`routeToken` is no longer exported from `@cotal-ai/core`. A role routes as spelled, so code that\nused it to name a role's queue uses the role itself, and `assertValidRole` checks one.\n\n### What migrates on its own\n\nEvery task queue. A `svc_<role>` durable was always named from the rewritten token, so its pending\nrequests and its holders carry over.\n\n### Before the upgrade\n\nRename each role outside the token set to the token it already routed to: remove the surrounding\nspaces and replace every other character outside the set with `_`. Rename it where the holder is\nlaunched and in every script or prompt that sends to it.\n\n## Carrying a resumed Claude session to another host in 0.67.0\n\n`cotal spawn --resume <id> --detach --on <instance>` now carries a Claude session held on the\noperator's host to the target manager instance. Both sides need this release: an older manager does\nnot serve `transcript-receive`, and the CLI then stops with that manager's refusal instead of\nlaunching. The manager cluster document moves to revision 21, and the `ps` row's `resume` object\ngains `host` and `transferredAt`.\n\nA manager host that runs carried seats needs `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_AUTH_TOKEN` or a\ncloud provider selection in its environment, because each carried seat runs in its own Claude home\nwith no stored login. On an authenticated mesh the CLI mints the transfer writer from the space's\nsigning seed, so the carrying host needs that seed, as for any other operator command. On a user-auth\nmesh it exchanges the operator's login for a `transfer-writer` view instead, so the operator's grant\nneeds scope `admin`, and the auth service must run this release. A remote manager receives a carry once\nits host serves the manager-service `transferReader` operation. A seat launched without carrying,\nincluding any `--resume` whose id this host does not hold, is unchanged.\n\n## Lifecycle head type in 0.67.0\n\n`LifecycleMapping`, the type `parseLifecycleHead` returns, is now a union on `state`. Nothing about\na running mesh changes: heads that parsed before parse the same way, and the refusals are\nunchanged. Only TypeScript code that compiles against `@cotal-ai/core` is affected.\n\n### What stops working\n\nAn `interface` that extends `LifecycleMapping` fails with TS2312, because an interface cannot extend\na union. Code that builds a head in memory no longer compiles when the head is `retiring` without\nits `op`, or `active` or `retired` with one. The parser already refused those heads.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type ActiveMapping = LifecycleMapping & { state: \"active\" }`. A reader that has checked\n`state === \"retiring\"` reads `op` without a guard.\n\n## Issuance gate types in 0.67.0\n\n`EpGateRow` and `EndpointGateRow`, which `parseIssuanceGate` and `parseEndpointGate` return, and\n`EpGateState`, which an `EpIssuanceGate` or `EpIssuanceBarrier` returns from `observe`, are now\nunions on `state`. Nothing about a running mesh changes: gates that parsed before parse the same\nway, and the refusals are unchanged. Only TypeScript code that compiles against `@cotal-ai/core`\nis affected.\n\n### What stops working\n\nAn `interface` that extends one of these types fails with TS2312, because an interface cannot\nextend a union. Code that builds a gate in memory, such as a custom barrier's `observe`, no longer\ncompiles when the gate is `frozen` or `retired` without its `op`, or `open` with one. The gate\nparsers already refused those rows.\n\n### Before the upgrade\n\nDeclare such an interface as an intersection instead, for example\n`type CustomGateRow = EpGateRow & { custom: string }`.\n\n## Lifecycle-blocked refusals in 0.66.0\n\nA refusal that carries `ai.cotal.ep.lifecycle-blocked` now reports only the lifecycle state it\nread. Nothing about a running mesh changes. A client that branches on the detail must read the new\nfield.\n\n### What stops working\n\nA refusal raised at the issuance gate used to carry `headState` without reading the head:\n`retiring` for a frozen gate and `retired` for a retired one. It now carries `gateState`\n(`frozen` or `retired`) and no `headState`. A client that treats `headState: \"retired\"` as a\nburned uid, or `headState: \"retiring\"` as a retirement in flight, no longer matches those\nrefusals, and the `[lifecycle ...]` suffix on the error string changes the same way. A custom\nissuance barrier whose `observe` returns a frozen gate without a valid `op` (a string `opId` and\none of the four op kinds) is now refused as `internal` by `registerServiceInstance`.\n\n### Before the upgrade\n\nUpdate such a client to read `gateState` for a gate refusal and `blockedOp` for the operation that\nholds the gate. `headState` is present only when the refusal read the head, for example an\nactivation refused because the head is still retiring.\n\n## Workflow programs that bind `once` in 0.65.0\n\n`once` is now a scope of the workflow language, so it is a reserved name. A program that declares\nits own `once` binding (`const once = ...`, a parameter or a function named `once`) is refused at\nvalidation with L2002. Nothing else about a running mesh changes.\n\n### What stops working\n\nA run whose recorded program binds `once` cannot be resumed after the upgrade, because a resume\nvalidates the recorded program again. A new `cotal run start` of such a program is refused before\nanything is recorded.\n\n### Before the upgrade\n\nList the runs with `cotal run ps` and check each program that is still running or held for a\nbinding named `once`. Let those runs finish on the old version before you upgrade the manager, and\nrename the binding in the program before you start it again.\n\n## From 0.58.0 to 0.59.0\n\nEvery connector now publishes a failed run's `RUN_ERROR` on `events.<owner>.<actor>` with the fixed\nmessage `run failed` and no `code` or `rawEvent`. The error text and error kind a harness reports\ncan echo a prompt, a peer message or tool output, and that channel has a different read ACL. A\nreader that showed the message or branched on `code` gets neither after the upgrade. Where a\nconnector reports the error kind as the agent's presence condition, that is unchanged.\n\n### Settle pending event frames before the upgrade\n\nEach session's events are frozen in its event write-ahead log before they are published. A session\nrestarted on 0.59.0 whose log still holds an unacknowledged frame with an older `RUN_ERROR` does not\nrepublish it: its event emitter halts with `egress-run-error` and publishes nothing further for that\nsession. The broker may or may not already hold that frame, so the halt cannot settle it.\n\n1. Stop the seats cleanly on 0.58.0, with the broker still up.\n2. List the logs that still hold a pending frame. The logs live under the events state root\n (`COTAL_WORKSPACE_ROOT`). Empty output means there is nothing to settle.\n\n ```sh\n find \"$COTAL_WORKSPACE_ROOT/.cotal/events\" -name wal.json \\\n -exec jq -r 'select(.pending != null) | input_filename' {} +\n ```\n\n3. For each session listed, start it again on 0.58.0 while the broker is reachable, let it recover,\n stop it, and run step 2 again. Recovery publishes the frame as 0.58.0 would have, error text\n included, so it only finishes what 0.58.0 had already started.\n\n If that start halts with `cas-loss` instead, the agent's subject is no longer at the sequence this\n log expects, and no restart settles that log, on 0.58.0 or later. A lost acknowledgement is one\n cause: the broker stored the frame, so it and its error text are already on the channel, and every\n retry halts the same way because the stream checks the frozen expectation before it deduplicates.\n The halt message names the other causes, such as a second emitter for the same agent under a\n different state root, a restored stream or frontier record, or a purged channel. With those the\n pending frame may never have reached the broker, so a `cas-loss` does not tell you whether it\n landed. Find and stop any second writer and rule out a restored state first. Clearing the halt\n then means purging the agent's event channel and removing the agent's directory under the events\n state root whole (see [Event plane](connect-claude.md#event-plane)). That abandons the pending\n frame whether or not the broker has it, and the purge also drops the earlier frames of every\n session of that agent.\n4. Upgrade once step 2 prints nothing.\n\nIf a session halts with `egress-run-error` after the upgrade, go back to step 3 for that session on\n0.58.0. Do not edit or delete `wal.json` on its own to get past either halt: clearing the pending\nframe abandons that epoch, an event the broker never received is lost, and removing part of the\ndirectory leaves a state the next start refuses.\n\n## Explicit actor grants in 0.59.0\n\n`cotal actor grant` no longer fills an omitted ACL flag with its wide default. A grant names\n`--scope`, `--allow-subscribe` and `--allow-publish`, or passes `--full` to give the ones it leaves\noff their wide defaults (`spawn,role:default`, `>` read, `>` post). Any other grant is refused. The\nbreak is in the CLI on the machine that holds the actor ledger, the one that ran\n`cotal up --user-auth --idp <url>`. No stored row, credential or wire message changes.\n\n### What keeps working\n\nExisting actor ledger rows keep the authority they were granted, and their users and agents connect\nas before. `actor revoke`, `actor list` and a `grant` that names all three ACL flags behave as they\ndid on 0.58.0. Nothing on disk is converted.\n\n### What stops working\n\nA grant that leaves off any of the three flags without `--full` exits 1 with\n`refusing to grant \"<actor>\" with --scope, --allow-subscribe, --allow-publish left off`, naming the\nflags it is missing, and then prints both accepted forms. It writes no row and does not retire the\nactor's current lifecycle. An existing row stays as it was, and an actor granted for the first time\nstays out until the grant is run again. This includes the bare grant printed on 0.58.0 by\n`cotal login`, `cotal status`, `actor list` and the not-granted refusal. Look for it in provisioning\nscripts, onboarding runbooks and anything that pastes those hints.\n\n### Upgrade order\n\nChange the scripts before the ledger machine is upgraded, and make each grant name all three flags.\n0.58.0 and 0.59.0 both accept that form. To keep a wide row, write its defaults out:\n\n```sh\ncotal actor grant <actor> --sub <IdP subject> \\\n --scope spawn,role:default --allow-subscribe '>' --allow-publish '>'\n```\n\nSwitch to `--full` only once the ledger machine runs 0.59.0. 0.58.0 refuses it with\n`Unknown option '--full'` before it reads the ledger. Brokers, managers and participant machines\nneed nothing for this break, so their order is the one the section above gives.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused grant changes nothing. The\nexposure is a grant script that runs against 0.59.0 before it was changed: it fails and grants\nnothing.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. On the ledger machine, save the output\nof `cotal actor list` to compare rows after the changed scripts run, and list the scripts that call\n`cotal actor grant`.\n\n### The upgrade end to end\n\n```sh\n# on the ledger machine, still on 0.58.0\ncotal actor list > actors-before.txt\ngrep -rn 'actor grant' <your provisioning scripts>\n# make every grant name --scope, --allow-subscribe and --allow-publish, run them, then upgrade\nnpm i -g cotal-ai@0.59.0\ncotal actor list | diff actors-before.txt -\n```\n\nBoth refusals quoted here were run on 0.58.0 and on the 0.59.0 code. That brokers, managers and stored\nrows need nothing is read from the change, which touches only the CLI and its hints, and was not run\non a live split deployment.\n\n## Repeated flags refused in 0.59.0\n\nA `cotal` flag given more than once is now a usage error unless the command declares it repeatable.\nOn 0.58.0 the last value won with no message, so `cotal down web --space a --space b` acted on `b`\nwhile a wrapper that checked the first `--space` verified `a`. The break is in the command-line\nparser on the machine that runs the command, including commands added with `cotal ext add`. No\nstored state, credential or wire message changes.\n\n### What keeps working\n\nA command line that gives each flag once parses as it did on 0.58.0, in any order and in the\n`--flag=value` form. Flags whose help says repeatable, such as `--opt` and `down --session-store`,\nstill collect every value. A flag-shaped word after `--` is still a positional. The daemons, units\nand agents that `cotal` starts for itself are given each flag once, so a fleet driven only by `cotal`\ncommands typed by hand needs no action.\n\n### What stops working\n\nA command line that repeats any other flag exits 1 before the command runs. It prints\n`Option '--space' cannot be repeated`, or `Option '-f, --file' cannot be repeated` for a flag with a\nshort form, followed by the command's help. `-f` and `--file` count as the same flag. Look for it in\nscripts, aliases and wrappers that append a flag to override one set earlier, such as a fixed\n`--space` followed by `\"$@\"`.\n\n### Upgrade order\n\nChange those scripts first so each flag is given once. 0.58.0 and 0.59.0 both accept that form.\nBrokers, managers and participant machines need nothing for this break, and each machine's CLI\napplies it when that machine is upgraded, so their order is the one the sections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it, and a refused command does nothing. The\nexposure is a script that still repeats a flag when it runs on 0.59.0: it exits 1 instead of acting on\nthe last value.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the scripts, aliases and wrappers\nthat call `cotal` so each one can be checked.\n\n### The upgrade end to end\n\n```sh\n# still on 0.58.0\ngrep -rn 'cotal ' <your scripts and wrappers>\n# give each non-repeatable flag once, then upgrade\nnpm i -g cotal-ai@0.59.0\n# run each changed script; a repeat left behind exits 1 with the usage error and does nothing\n```\n\nThe refusal and its messages were run against the 0.59.0 parser and `cotal topology view`. That the\nargument lists `cotal` builds for its own processes give each flag once is read from the code, and\nwas not run on a live split deployment.\n\n## Detached spawns from a seat's shell in 0.62.0\n\nOn a static or open mesh, `cotal spawn --detach` run inside a managed seat's shell now launches as\nthat seat. On 0.61.0 it minted a one-shot operator instrument, so the manager recorded that\ninstrument as the spawner and the seat's own `cotal_despawn` of the child was refused with\n`not authorized: <seat> was not spawned by <caller> (admin tier required)`. The break is in the CLI\non the machine where the seats run. No stored state, credential or wire message changes.\n\n### What keeps working\n\n`cotal spawn --detach` from an operator terminal or from a script outside any seat launches as\nbefore, and so does any call with `--creds`, one aimed at a space other than the seat's own, or a raw\nopen target named with `--server` and an unregistered `--space`. A user-auth mesh is unchanged. A seat with\n`capabilities: [spawn]` still spawns from its shell, and can now stop that child with\n`cotal_despawn`. `--on <instance>` from a seat's shell still lands on that manager instance, now as\nthe seat.\n\n### What stops working\n\n- On a static mesh, a seat without `capabilities: [spawn]` can no longer spawn from its shell. Its\n own credential holds no spawn subject, so the broker refuses the request and the command exits 1.\n- A child launched from a seat's shell is now that seat's child, so the manager stops it when the\n seat exits, as it does for a `cotal_spawn` child. A child that has to outlive the seat that\n started it now goes with the seat.\n- A seat launched without `COTAL_SPACE` is placed by its static credential. Every connector sets\n that variable, so this only reaches a hand-built launch: from such a seat's shell, a spawn aimed at\n a static space that holds no credential for the seat is refused instead of running as the operator.\n\n### Upgrade order\n\nOnly the CLI that seats run from their shell changes, which is the one installed on the host where\nthe seats run. Brokers and managers need nothing for this break, so their order is the one the\nsections above give.\n\n### The window\n\nThis break has no outage. No process restarts for it. A child already running when you upgrade\nkeeps the spawner the manager recorded at its launch.\n\n### Snapshot this first\n\nNothing is rewritten, so this break has no state to back up. List the agent files whose seats run\n`cotal spawn --detach` from their shell, note which of them lack `capabilities: [spawn]`, and note\nwhich of their children must outlive the seat.\n\n### The upgrade end to end\n\n```sh\n# still on 0.61.0: find the seats that spawn from their shell\ngrep -rln 'cotal spawn' .cotal/agents\n# add `capabilities: [spawn]` to each of those agent files that lacks it, and launch any child\n# that must outlive its seat from an operator terminal instead\nnpm i -g cotal-ai@0.62.0\n```\n\nThe attribution, the despawn, the refusal of a seat without `spawn`, the stop on seat exit and a\nseat's `--on` spawn were run on a local static mesh, and the attribution and the despawn on a local\nopen mesh.\n\n## From 0.53.0 to 0.54.0\n\nManager calls now borrow an instance-bound `manager-caller` credential. Followed mutations require\n`manager.goal-result` on the selected manager, so a compatible issuer, manager and client must be\nloaded together. An older manager is refused before a followed mutation; upgrading an installed\nbinary alone does not replace code in a running manager, connector or embedded client.\n\n### Preserve state before changing processes\n\nSnapshot the broker's durable storage using its supported backup procedure, the host authority and\nactor ledgers, and each participant's manager identity, runtime custody records, credentials and\nsaved sessions. Include the embedding application's database and configuration under its supported\nbackup procedure. Record the loaded package versions and the CLI path used by bearer helpers.\nKeep these copies private. Do not change the IdP issuer, regenerate manager identities, rotate agent\ncredentials or recreate tenant storage to make the upgrade pass.\n\nNo ledger, goal-history or session conversion is required for this change. Existing ordinary\nmessaging credentials retain their normal expiry rules. New manager-caller credentials are obtained\non demand from the current grant; old manager-call credentials do not gain the new view automatically.\nExisting accepted goals remain durable and must not be submitted again merely because observation\nwas interrupted. Fresh remote registration publishes its service status at the current revision and\nepoch; do not seed that status manually.\n\n### Upgrade the split deployment\n\n1. Stage one pinned 0.54.0 package set for the host and participants, including the embedding SDKs.\n Pause new manager mutations and let accepted work settle where possible before reloading processes.\n2. Upgrade the host issuer and embedding first. Keep the broker, its account identities and durable\n storage in place. Then load the matching manager release on each participating machine.\n3. Preserve active seats through the runtime's supported update path. A Linux custodial runtime may\n release and re-adopt seats within its 600-second unattended window; verify the actual runtime,\n custody records and process identities before relying on it. A legacy PTY runtime without release\n support cannot preserve active seats through a generic manager restart. Drain it at an approved\n idle window instead of signalling the manager or replacing conversations.\n4. Reload the clients and connectors through their session-preserving host controls. Refresh any\n bearer helper captured from an older immutable CLI path. A transport-only reconnect does not reload\n JavaScript. Verify authenticated instance selection, a read-only manager command and canonical\n result recovery before allowing new followed mutations.\n\nTreat the interval from issuer reload through compatible manager/client reload as a manager-control\noutage. Mixed versions can refuse discovery or commands; there is no promised rolling transition.\nOrdinary agent sessions survive only where their runtime and credentials permit it. If verification\nfails, keep mutations paused and repair forward from the preserved state rather than resetting it.\nThis release does not add host-backed enrollment or terminal release for stock participant detached\nagents; see [Remote supervised agents](run-a-mesh.md#remote-supervised-agents).\n\n## From 0.48.2 to 0.49.0\n\n0.49.0 changes how a credential's authority is recorded. A credential is no longer only a signed\nfile: it is an *issuance*, with a generation the issuer chose and durable evidence of the ceiling it\nwas granted under. The important consequence for a running deployment is not at connect time. It is\nat renewal time.\n\n### What keeps working without any action\n\n- **Existing agent credentials keep authenticating.** A credential minted under 0.48.2 is not\n revoked and is not rejected at connect. Nothing needs to be re-issued to bring the fleet back up\n after the upgrade.\n- **The channel registry survives.** Channels, their replay settings, descriptions, and usage text\n are ordinary durable state and are not rewritten by the upgrade.\n- **`cotal deliver` is still a standalone command.** Running the delivery daemon as its own process\n remains supported; it is not restricted to being a child of `cotal up`.\n- **`cotal join` keeps its flags.** In particular `--lifecycle-uid` is not new in 0.49.0. It has\n been required alongside `--creds` since well before this release, and the pairing rule did not\n change here. A scripted external join that worked under 0.48.2 works unchanged.\n\n### What does not migrate\n\n**A credential minted before 0.49.0 cannot be renewed.** Managed agent credentials carry a\n24-hour lifetime and the manager re-signs one once it passes **75%** of its life, ticking every\nquarter of the TTL so a tick always lands inside that window. When the manager reaches a credential\nthat carries no issuance, it refuses to renew it and logs the agent by name:\n\n```\n! managed cred renewal <agent>: renewManagedStaticCred: <agent> carries no issuance;\n a static credential minted before SPEC 13.15 is not renewed under an unbound generation\n - respawn the agent\n - the agent dies loud at this cred's expiry unless it is reminted\n```\n\nSo the fleet comes up fine, runs normally, and then each agent stops at its own credential's\nexpiry, within roughly a day of the upgrade, one at a time rather than together. The refusal is\ndeliberate: the renewal would otherwise have to invent a generation nobody issued, which is the\nstate the release exists to remove.\n\n**Respawn the managed agents as the last step of the upgrade.** For this particular upgrade the\nrespawn is not optional: stopping a 0.48.2 manager ends its agent processes whichever CLI you use,\nfor the reason given under the outage window below. The respawn is how they come back, and it is\nalso what mints each credential as an issuance so it renews from then on. One planned pass over the\nfleet is the whole job. Skipping it leaves agents stopped and, for any credential that survived\ninto 0.49.0 unminted, brings the renewal cliff above a day later, one agent at a time.\n\n### Credentials you minted yourself\n\n**A credential you minted with `cotal mint` is a different case, and it very likely needs\nnothing.** The distinction that matters is not the word \"static\", which covers both. It is **what\nminted the credential and who owns its renewal**. A credential the **manager** minted for an agent\nit spawned carries a lifetime and is renewed by the manager, so it is the subject of everything\nabove. A credential **you** minted with `cotal mint` and handed to an external peer is issued with\n**no expiry at all**, and no manager renews it: it is not in the sweep, so there is no renewal to\nfail. It keeps working after the upgrade, and re-minting it would mean coordinating with a third\nparty for no gain.\n\nThe manager says which one it is holding. Where a credential has no expiry to reach, the sweep\nnames it and moves on rather than refusing:\n\n```\n! managed cred renewal <agent>: credential is unbounded - not renewed\n (a pre-TTL credential stays as minted until respawn)\n```\n\nRe-mint an external peer's credential only if you want it to carry a lifetime, and at a time you\nchoose.\n\n### How to read the boot log\n\nA 0.49.0 manager starting over an existing space may print lines like:\n\n```\n verified evicted: <holder-key> (3/12)\n already verified (durable): <holder-key>\n✓ boot self-heal: manager/<id> registration gate reopened at generation <n>\n```\n\nThese are **not** a credential migration, and reading them as one is the most likely way to\nconclude the fleet is fine when it is not. They come from the manager repairing **one** endpoint\nregistration gate that a previous restart left frozen, and they enumerate that single gate's\ncredential-family holders as it verifies each one evicted. `already verified (durable)` on a later\nstart is the repair cursor resuming, not a credential that became durable. The repair is real and\nuseful (it is what previously needed `cotal reconcile-gate` by hand), but it says nothing about\nwhether your agent credentials carry issuances. The renewal refusal above is the signal that does.\n\n### Which side to upgrade first in a split topology\n\nMove the manager first.\n\nThe stores 0.49.0 introduces are created by the **manager** at its own boot, not by the broker.\nThey are create-or-verify and idempotent, so a 0.49.0 manager brings the space's authority stores\nup to the new shape itself, and it does so against whichever broker is answering.\n\nBeing honest about the evidence behind each direction, because they are not equally established:\n\n- **Broker-first was measured on a live 30-agent deployment** (issue #1578). Upgrading the broker\n first locks the old manager out immediately: `cotal up` re-renders the broker's generated config\n from the trust record, and after the restart the still-0.48.2 manager is refused on every\n connection with an `authentication error` naming the Nkey, continuously. That text comes from the\n broker process, not from a Cotal command, so match on its shape rather than on an exact string.\n `cotal ps` reports zero agents while\n the agent processes are still alive, because the manager has lost its view of them, not because\n they died. Upgrading the manager clears it immediately.\n- **Manager-first is reasoned from where the new stores are provisioned**, not from a measured\n fleet upgrade. It is the recommended order because the manager is the component that creates what\n 0.49.0 adds, but it has not been run end to end on a production split topology at the time of\n writing. Treat it as the better-supported order rather than a guaranteed one, and keep the\n rollback below ready either way.\n\nWhichever order you pick, **this is not a rolling upgrade**. Between the two steps the mesh is down\nand the manager cannot see its agents. Go straight through rather than pausing between them, and\nschedule it as an outage window.\n\n### What the window looks like\n\n- **The managed agent processes do not survive step 1, in either order.** This is the one place\n where the obvious reordering does not rescue you, so it is worth understanding rather than\n working around. Sparing agents on a bare manager stop is a **handshake**: a 0.49.0 manager\n publishes a capability file proving it can release its agents, and a 0.49.0 `cotal down` refuses\n the stop unless it finds one. **A 0.48.2 manager never publishes that file**, because the\n mechanism ships in the release you are installing. So the old CLI against the old manager sends a\n plain stop and takes every seat with it, and the new CLI against the old manager either refuses\n (leaving `--with-agents`, which reaps deliberately) or falls to the legacy path, warns that it\n cannot verify the manager can spare its agents, and signals it anyway.\n- **You can confirm which side you are on in one command, without stopping anything.** The flag that\n marks the newer behaviour is absent from the older CLI, and its summary line makes the difference\n plain:\n\n ```\n $ cotal down --help # on 0.48.2\n cotal down - stop the whole local stack, or name only the components to stop\n\n $ cotal down --help # on 0.49.0\n cotal down - stop the whole local stack (managed agents stay running unless --with-agents), ...\n ```\n\n If your `cotal down --help` does not mention `--with-agents`, stopping the manager stops the\n agents with it.\n- **Therefore the respawn in step 5 is mandatory recovery for this upgrade, not an optional pass.**\n It is also the step that re-mints credentials as issuances, so it is the same action either way.\n Plan the window to include it rather than treating it as cleanup.\n- The **manager's view** of them is lost while the two sides disagree, so `cotal ps` reports zero\n and control commands do not reach seats.\n- **Messages are not delivered** while the mesh is down.\n- The window is as long as it takes to restart the second component, plus the manager's own start.\n It is minutes, not hours, provided you do not stop between the steps.\n- **Nothing self-heals if you stop halfway.** The refusal is continuous until both sides match.\n\n### Snapshot this before you start\n\nTake these while the deployment is still on 0.48.2. The two `cotal` reads are live reads and must\nhappen before anything stops.\n\n- **A filesystem or volume snapshot of both containers**, if your platform offers one. This is the\n only rollback that covers every case, and it is what the reporting deployment used.\n- **`cotal backup create <dir>`**, for the durable space state, **but read the next paragraph before\n you rely on it**: on a split broker and manager topology it is very likely unavailable to you, and\n the volume snapshot above is your actual rollback.\n- **The trust records and credential directory** under `.cotal/auth` on the manager host, including\n the per-space material directory. These are what a re-mint would otherwise have to replace.\n- **A copy of the channel registry**, so you can verify it came back rather than assuming it did:\n `cotal channels list` before and after.\n- **The output of `cotal ps`**, so you know how many seats you expect to see afterwards and can tell\n a lost view from a lost agent.\n\n#### `cotal backup` on a split topology\n\n**`cotal backup create` cannot read a running stack.** It requires a completed cut, and only\n`cotal down --preserve-state` publishes one:\n\n```\n$ cotal backup create ./backup.0482\n✗ backup requires a completed cut; run `cotal down --preserve-state` first\n```\n\n**And `cotal down --preserve-state` requires a manager alive on the host you run it from.** It uses\nthat manager to attest that every retained child stopped, and the check is deliberately fail-closed:\na manager that is dead or merely uncertain refuses rather than preserving an unproven cut. The check\nreads a local pidfile, so a **remote** manager does not satisfy it. On a split topology the broker\nhost has no local manager, which means the documented durable-backup path is not available there.\n\n**Measured rather than assumed, at 0.48.2**: the backup refusal above is executed output. The\npreservation requirement is read from `down.ts` at the same tag, where the preserve path asks a\nmanager to prepare an inventory and then requires that manager to be locally alive before it\ncommits. The part not executed end to end is a genuine two-host split, which needs two real hosts.\n\n**What to do instead.** Use the filesystem or volume snapshot of both containers. That is the\nrollback the reporting deployment actually used, it covers the broker's durable state and the\nmanager's credential material together, and it does not depend on either component being able to\nattest for the other. If you want `cotal backup` as well, take it from a host that does have a live\nlocal manager, and understand it is a second copy rather than the primary rollback.\n\n**This looks like a product limitation rather than a documentation gap**, and it is written here as\none so an operator is not left thinking they mis-typed a command. The upgrade path for the exact\ntopology this page is addressed to cannot use the documented backup command.\n\n### The upgrade end to end\n\n```bash\n# 0. on 0.48.2, STILL RUNNING: record what you expect to see afterwards.\n# These two are live reads, so they must happen before anything stops.\ncotal channels list > channels.before\ncotal ps > ps.before\n\n# 1. manager host. READ THE NOTE BELOW THE BLOCK FIRST: this step ends the\n# managed agent processes whichever order you choose, and the respawn in\n# step 5 is how they come back. It is recovery, not tidying.\n#\n# STOP THE MANAGER WITH THE 0.48.2 CLI, BEFORE INSTALLING 0.49.0. The\n# order matters and it is not recoverable once you install: a 0.49.0\n# `down manager` REFUSES to stop a 0.48.2 manager whose pid record carries\n# a start token, which is every manager on a platform that can read one\n# (Linux can):\n# refusing bare manager stop: ... does not prove this manager can detach\n# its agents; use --with-agents or stop the agents explicitly\n# The refusal names two remedies and NEITHER clears it for this case. The\n# check reads a capability file that only a 0.49.0 manager writes; it never\n# counts agents, so stopping them first changes nothing. And `--with-agents`\n# is whole-stack only, so `down manager --with-agents` is refused by its own\n# flag rule. See #1592.\ncotal down manager # the 0.48.2 CLI, still installed.\n # 0.48.2 has no --with-agents; this\n # is the whole route. On a host that\n # runs the whole stack, the 0.49.0\n # `cotal down --with-agents` after\n # installing is the alternative.\nnpm install -g cotal-ai@0.49.0 # ONLY after the stop above\n# `supervise` RUNS IN THE FOREGROUND and holds the terminal until you stop\n# it. There is no --detach on this command. Start it under whatever keeps\n# your manager alive normally (systemd unit, container entrypoint, or a\n# second terminal), and run the remaining steps from another shell.\ncotal supervise --space <space> --server nats://<broker>:4222\n\n# 2. broker host: stop the stack.\n# NOT `--preserve-state` on a split topology: it needs a manager alive on\n# THIS host to attest its children stopped, and yours is on the other one.\n# Your rollback is the volume snapshot from \"Snapshot this before you\n# start\", not `cotal backup`.\n# See \"cotal backup on a split topology\" above.\ncotal down\n\n# 3. broker host: install 0.49.0 and start it again\nnpm install -g cotal-ai@0.49.0\n# Record the manager log's size BEFORE starting, so step 3a can tell THIS\n# boot's output from every earlier one. It must be captured here, ahead of\n# the start: taken afterwards it sits past the new line and the wait hangs.\n# `<spaceKey>` is NOT the space name. It is lowercase hex of the name's\n# UTF-8 bytes, so space `prod` is `manager.70726f64.log`. Do not guess it:\n# `cotal up` prints the real path on its launch line. Substituting the\n# plain name points at a file that does not exist, and the wait below then\n# burns its full timeout before telling you.\nLOG=.cotal/manager.<spaceKey>.log\nOFF=$( [ -f \"$LOG\" ] && wc -c < \"$LOG\" || echo 0 )\ncotal up --detach --host 0.0.0.0 --space <space> --no-manager\n\n# 3a. SPLIT TOPOLOGY ONLY: `--no-manager` above boots the broker (and the\n# delivery daemon) with NO local manager on the broker host, so there is\n# no wait-and-stop step on a current cotal-ai. The rest of this step is\n# the OLDER-host recipe, kept because the flag is refused there and that\n# refusal is your signal you are on it: without the flag the `up` also\n# starts a local manager, and you must wait for the log to show it is up,\n# then stop it, or you finish the upgrade with two managers and the one\n# you did not intend is the one nobody is watching.\n# A bare `grep -q` does NOT wait: it reads once and exits 1 immediately\n# if the line has not been written yet. Bound the wait instead, so a\n# manager that never comes up fails loudly rather than reading as ready.\n# The log is opened APPEND-ONLY, so on any host that has run a manager\n# before, this file ALREADY carries a `manager up` line from an earlier\n# boot. Grepping the whole file therefore matches instantly and waits for\n# nothing. Read only what THIS boot appended, using the $OFF captured in\n# step 3 above (before the start, which is the only point it is correct):\ntimeout 60 bash -c \\\n \"until tail -c +$((OFF+1)) \\\"$LOG\\\" | grep -q '. manager up'; do sleep 1; done\"\n# exit 0 = THIS boot logged it; exit 124 = it never did, so STOP and look.\n# This manager is 0.49.0 and publishes its own spare-capability file, so\n# the bare stop below is NOT the refusal case from step 1.\ncotal down manager # broker + delivery remain\n# On a current cotal-ai the two commands above are unnecessary (nothing\n# to wait for, nothing to stop) and `cotal down manager` simply reports\n# no manager to stop.\n\n# 4. verify the mesh is whole again before touching the fleet.\n# Do NOT compare `cotal ps` against ps.before yet: step 1 ended the agent\n# processes, so at this point it is EXPECTED to be empty, and an empty\n# `ps` is also the signature of the broker/manager mismatch described\n# above. The two are indistinguishable here, so compare what the mesh\n# itself should have carried across instead:\ncotal channels list # compare against channels.before: this SHOULD match now\ncotal ps # expect it to be EMPTY here; ps.before is the target for\n # step 5, not for this step\n\n# 5. the step that is easy to skip: respawn the managed agents so their\n# credentials are re-minted as issuances and can renew. Persona is a\n# POSITIONAL argument here, unlike `cotal stop`, which requires --name.\n# One call per agent:\ncotal spawn <persona> --detach --name <n> --space <space>\n# then the comparison step 4 could not make:\ncotal ps # NOW compare against ps.before: seat count should match\n```\n\nThe mesh is down from step 2 until step 3 finishes. That is the window. On a split topology there is\nno cut and no backup inside it, so the window is the stop, the install and the restart, nothing more.\n\n## Adding a section for a future release\n\n**Every changeset marked breaking adds a section to this page.** A release that changes what an\noperator must do, in what order, or what stops working, is not finished until the section exists.\n`scripts/upgrade-section-gate.mjs` grades a commit range for this: run it as\n`pnpm upgrade-section-gate --base <ref>` and it reds when the range carries a breaking change and\nadds no new release section. CI runs its self-test and, as a step of the `attribution` job, grades\neach pull request's own range as `HEAD^1..HEAD` over the merge snapshot it checked out. That job is\nthe only context in the branch protection rule set, so a red gate FAILS A REQUIRED CHECK AND BLOCKS\nTHE MERGE. The section is not optional and a reviewer cannot wave it through without an\nadministrator overriding branch protection. Be precise about what the check proves either\nway, because one trusted past its evidence is worse than none. It proves a section for a release\n**was written here**. It cannot prove the section is **correct**, or that it describes the break\nthat actually landed, and it cannot see a breaking change that carries no marker at all. Reviewing\nthe words remains a person's job.\n\n**Mark the break, or the gate cannot see it.** Any one of these is enough, and they are the only\nthings it reads:\n\n- a `!` before the colon in the commit subject, as in `feat(core)!: bind hosted runs to the caller`\n- a `BREAKING CHANGE:` footer in the commit body\n- a changeset in `.changeset/` declaring a `major` bump for any package\n\nThe marker must survive the squash. A `!` that lives only in a commit you squash away is not in the\nrange the gate grades, so put it in the subject that lands on `main`.\n\n**The heading is a `##` and names the release**, like `## From 0.48.2 to 0.49.0`. Both matter, and\nneither is a style preference. Coverage is claimed by a heading, so a heading that names\nno release claims every release and distinguishes none: `## Notes` with a sentence under it would\notherwise satisfy the rule. Naming the release also makes the section the one an operator upgrading\nthat release will search for. Use `###` freely for detail inside a section. Subsections belong to\ntheir release rather than counting as separate coverage.\n\nName the release that first carries the change: the next version Changesets publishes, which\n`pnpm changeset status --verbose` lists. `bin/package.json` on `main` still reads the release already\npublished. If a release is cut while the change is open, the change ships in the release after it,\nso move the heading before merging. The gate accepts any version in a heading, so before merging a\nrelease pull request, check every heading added since the previous tag against the version it\npublishes.\n\nA section is written for the operator, not for the reviewer. It answers, in this order:\n\n1. What keeps working with no action at all.\n2. What does **not** migrate, and when that becomes visible. Name the log line if there is one.\n3. The order to move components in for a split topology, and why that order.\n4. What the outage window looks like, including what survives it.\n5. What to snapshot before starting.\n6. The commands, end to end.\n\n**Where an answer was not measured, say so in the document rather than guessing.** An operator who\nknows which half of a recommendation is reasoned and which is measured can plan around it; one who\nfinds out afterwards cannot.\n"
|
|
273
273
|
},
|
|
274
274
|
{
|
|
275
275
|
"slug": "watch-a-mesh",
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"docs-bundle.generated.js","sourceRoot":"","sources":["../src/docs-bundle.generated.ts"],"names":[],"mappings":"AAKA,yFAAyF;AACzF,MAAM,CAAC,MAAM,YAAY,GAAG,QAAQ,CAAC;AAErC,MAAM,UAAU,cAAc;IAC5B,OAAO;QACP,SAAS,EAAE,QAAQ;QACnB,eAAe,EAAE,mEAAmE;QACpF,OAAO,EAAE;YACP;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,+FAA+F;gBAC1G,MAAM,EAAE,81IAA81I;aACv2I;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,gHAAgH;gBAC3H,MAAM,EAAE,g2XAAg2X;aACz2X;YACD;gBACE,MAAM,EAAE,cAAc;gBACtB,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,upnBAAupnB;aAChqnB;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,kBAAkB;gBAC3B,MAAM,EAAE,mEAAmE;gBAC3E,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,sgqCAAsgqC;aAC/gqC;YACD;gBACE,MAAM,EAAE,0BAA0B;gBAClC,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,mCAAmC;gBAC3C,SAAS,EAAE,gGAAgG;gBAC3G,MAAM,EAAE,6qMAA6qM;aACtrM;YACD;gBACE,MAAM,EAAE,mBAAmB;gBAC3B,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,oDAAoD;gBAC/D,MAAM,EAAE,4+1CAA4+1C;aACr/1C;YACD;gBACE,MAAM,EAAE,aAAa;gBACrB,OAAO,EAAE,aAAa;gBACtB,MAAM,EAAE,yFAAyF;gBACjG,SAAS,EAAE,gJAAgJ;gBAC3J,MAAM,EAAE,gjWAAgjW;aACzjW;YACD;gBACE,MAAM,EAAE,uBAAuB;gBAC/B,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,sFAAsF;gBAC9F,SAAS,EAAE,6GAA6G;gBACxH,MAAM,EAAE,0hTAA0hT;aACniT;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,sMAAsM;gBACjN,MAAM,EAAE,4+SAA4+S;aACr/S;YACD;gBACE,MAAM,EAAE,KAAK;gBACb,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,wGAAwG;gBAChH,SAAS,EAAE,iKAAiK;gBAC5K,MAAM,EAAE,ws9KAAws9K;aACjt9K;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,uHAAuH;gBAC/H,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,05/BAA05/B;aACn6/B;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,gBAAgB;gBACzB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,+EAA+E;gBAC1F,MAAM,EAAE,
|
|
1
|
+
{"version":3,"file":"docs-bundle.generated.js","sourceRoot":"","sources":["../src/docs-bundle.generated.ts"],"names":[],"mappings":"AAKA,yFAAyF;AACzF,MAAM,CAAC,MAAM,YAAY,GAAG,QAAQ,CAAC;AAErC,MAAM,UAAU,cAAc;IAC5B,OAAO;QACP,SAAS,EAAE,QAAQ;QACnB,eAAe,EAAE,mEAAmE;QACpF,OAAO,EAAE;YACP;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,+FAA+F;gBAC1G,MAAM,EAAE,81IAA81I;aACv2I;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,0BAA0B;gBAClC,SAAS,EAAE,gHAAgH;gBAC3H,MAAM,EAAE,g2XAAg2X;aACz2X;YACD;gBACE,MAAM,EAAE,cAAc;gBACtB,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,upnBAAupnB;aAChqnB;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,kBAAkB;gBAC3B,MAAM,EAAE,mEAAmE;gBAC3E,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,sgqCAAsgqC;aAC/gqC;YACD;gBACE,MAAM,EAAE,0BAA0B;gBAClC,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,mCAAmC;gBAC3C,SAAS,EAAE,gGAAgG;gBAC3G,MAAM,EAAE,6qMAA6qM;aACtrM;YACD;gBACE,MAAM,EAAE,mBAAmB;gBAC3B,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,oDAAoD;gBAC/D,MAAM,EAAE,4+1CAA4+1C;aACr/1C;YACD;gBACE,MAAM,EAAE,aAAa;gBACrB,OAAO,EAAE,aAAa;gBACtB,MAAM,EAAE,yFAAyF;gBACjG,SAAS,EAAE,gJAAgJ;gBAC3J,MAAM,EAAE,gjWAAgjW;aACzjW;YACD;gBACE,MAAM,EAAE,uBAAuB;gBAC/B,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,sFAAsF;gBAC9F,SAAS,EAAE,6GAA6G;gBACxH,MAAM,EAAE,0hTAA0hT;aACniT;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,sMAAsM;gBACjN,MAAM,EAAE,4+SAA4+S;aACr/S;YACD;gBACE,MAAM,EAAE,KAAK;gBACb,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,wGAAwG;gBAChH,SAAS,EAAE,iKAAiK;gBAC5K,MAAM,EAAE,ws9KAAws9K;aACjt9K;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,uHAAuH;gBAC/H,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,05/BAA05/B;aACn6/B;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,gBAAgB;gBACzB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,+EAA+E;gBAC1F,MAAM,EAAE,k9zCAAk9zC;aAC39zC;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,iisBAAiisB;aAC1isB;YACD;gBACE,MAAM,EAAE,gBAAgB;gBACxB,OAAO,EAAE,wBAAwB;gBACjC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,kJAAkJ;gBAC7J,MAAM,EAAE,4ifAA4if;aACrjf;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,6CAA6C;gBACxD,MAAM,EAAE,4z2BAA4z2B;aACr02B;YACD;gBACE,MAAM,EAAE,kBAAkB;gBAC1B,OAAO,EAAE,yBAAyB;gBAClC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wJAAwJ;gBACnK,MAAM,EAAE,+3XAA+3X;aACx4X;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,oBAAoB;gBAC7B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,8DAA8D;gBACzE,MAAM,EAAE,g3QAAg3Q;aACz3Q;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,4HAA4H;gBACvI,MAAM,EAAE,6wMAA6wM;aACtxM;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,oIAAoI;gBAC/I,MAAM,EAAE,2qkCAA2qkC;aACprkC;YACD;gBACE,MAAM,EAAE,eAAe;gBACvB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,uMAAuM;gBAClN,MAAM,EAAE,+4SAA+4S;aACx5S;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,+BAA+B;gBACxC,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,6HAA6H;gBACxI,MAAM,EAAE,irfAAirf;aAC1rf;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,6GAA6G;gBACxH,MAAM,EAAE,q6MAAq6M;aAC96M;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,iBAAiB;gBAC1B,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,yEAAyE;gBACpF,MAAM,EAAE,gthEAAgthE;aACzthE;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,6DAA6D;gBACxE,MAAM,EAAE,+6GAA+6G;aACx7G;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,yBAAyB;gBACjC,SAAS,EAAE,wEAAwE;gBACnF,MAAM,EAAE,q5PAAq5P;aAC95P;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,yBAAyB;gBACjC,SAAS,EAAE,+CAA+C;gBAC1D,MAAM,EAAE,i0NAAi0N;aAC10N;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,8BAA8B;gBACvC,MAAM,EAAE,8CAA8C;gBACtD,SAAS,EAAE,qIAAqI;gBAChJ,MAAM,EAAE,o6SAAo6S;aAC76S;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,UAAU;gBACnB,MAAM,EAAE,sDAAsD;gBAC9D,SAAS,EAAE,uJAAuJ;gBAClK,MAAM,EAAE,srZAAsrZ;aAC/rZ;YACD;gBACE,MAAM,EAAE,uBAAuB;gBAC/B,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,0IAA0I;gBACrJ,MAAM,EAAE,0yhBAA0yhB;aACnzhB;YACD;gBACE,MAAM,EAAE,SAAS;gBACjB,OAAO,EAAE,sBAAsB;gBAC/B,MAAM,EAAE,0CAA0C;gBAClD,SAAS,EAAE,gIAAgI;gBAC3I,MAAM,EAAE,i0aAAi0a;aAC10a;YACD;gBACE,MAAM,EAAE,SAAS;gBACjB,OAAO,EAAE,SAAS;gBAClB,MAAM,EAAE,yBAAyB;gBACjC,SAAS,EAAE,mLAAmL;gBAC9L,MAAM,EAAE,y7KAAy7K;aACl8K;YACD;gBACE,MAAM,EAAE,YAAY;gBACpB,OAAO,EAAE,YAAY;gBACrB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,07+CAA07+C;aACn8+C;YACD;gBACE,MAAM,EAAE,UAAU;gBAClB,OAAO,EAAE,gBAAgB;gBACzB,MAAM,EAAE,oCAAoC;gBAC5C,SAAS,EAAE,kGAAkG;gBAC7G,MAAM,EAAE,6xbAA6xb;aACtyb;YACD;gBACE,MAAM,EAAE,iBAAiB;gBACzB,OAAO,EAAE,oCAAoC;gBAC7C,MAAM,EAAE,0CAA0C;gBAClD,SAAS,EAAE,wMAAwM;gBACnN,MAAM,EAAE,4xqBAA4xqB;aACryqB;YACD;gBACE,MAAM,EAAE,QAAQ;gBAChB,OAAO,EAAE,QAAQ;gBACjB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,2DAA2D;gBACtE,MAAM,EAAE,swKAAswK;aAC/wK;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,qBAAqB;gBAC9B,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,kGAAkG;gBAC7G,MAAM,EAAE,8uPAA8uP;aACvvP;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,uBAAuB;gBAChC,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,kGAAkG;gBAC7G,MAAM,EAAE,09OAA09O;aACn+O;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,gCAAgC;gBACzC,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,iEAAiE;gBAC5E,MAAM,EAAE,oihEAAoihE;aAC7ihE;YACD;gBACE,MAAM,EAAE,cAAc;gBACtB,OAAO,EAAE,cAAc;gBACvB,MAAM,EAAE,qBAAqB;gBAC7B,SAAS,EAAE,uHAAuH;gBAClI,MAAM,EAAE,s+qBAAs+qB;aAC/+qB;YACD;gBACE,MAAM,EAAE,WAAW;gBACnB,OAAO,EAAE,eAAe;gBACxB,MAAM,EAAE,uBAAuB;gBAC/B,SAAS,EAAE,kHAAkH;gBAC7H,MAAM,EAAE,084CAA084C;aACn94C;SACF;QACD,MAAM,EAAE;YACN,OAAO,EAAE,0BAA0B;YACnC,MAAM,EAAE,g7kgBAAg7kgB;SACz7kgB;QACD,MAAM,EAAE;YACN,OAAO,EAAE,mCAAmC;YAC5C,MAAM,EAAE,i9mGAAi9mG;SAC19mG;QACD,QAAQ,EAAE;YACR,OAAO,EAAE,oCAAoC;YAC7C,MAAM,EAAE,m4pBAAm4pB;SAC54pB;KACF,CAAC;AACF,CAAC"}
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"manager-call.d.ts","sourceRoot":"","sources":["../src/manager-call.ts"],"names":[],"mappings":"
|
|
1
|
+
{"version":3,"file":"manager-call.d.ts","sourceRoot":"","sources":["../src/manager-call.ts"],"names":[],"mappings":"AAAA,OAAO,EAWL,KAAK,iBAAiB,EACtB,KAAK,QAAQ,EACb,KAAK,YAAY,EAClB,MAAM,gBAAgB,CAAC;AACxB,OAAO,KAAK,EAAE,WAAW,EAAE,MAAM,aAAa,CAAC;AAE/C,KAAK,aAAa,GAAG,IAAI,CAAC,WAAW,EAAE,OAAO,GAAG,SAAS,GAAG,KAAK,GAAG,cAAc,GAAG,UAAU,GAAG,mBAAmB,CAAC,CAAC;AAExH,kGAAkG;AAClG,wBAAgB,oBAAoB,CAAC,MAAM,EAAE,MAAM,EAAE,MAAM,EAAE,aAAa,GAAG;IAAE,MAAM,EAAE,QAAQ,CAAC;IAAC,UAAU,EAAE,MAAM,CAAA;CAAE,CAuBpH;AAED,mGAAmG;AACnG,wBAAsB,iBAAiB,CACrC,MAAM,EAAE,aAAa,EACrB,MAAM,EAAE,MAAM,EACd,OAAO,EAAE,MAAM,EACf,IAAI,EAAE,MAAM,CAAC,MAAM,EAAE,OAAO,CAAC,GAAG,SAAS,EACzC,IAAI,GAAE;IAAE,MAAM,CAAC,EAAE,YAAY,CAAC;IAAC,UAAU,CAAC,EAAE,MAAM,CAAC;IAAC,MAAM,CAAC,EAAE,OAAO,CAAC;IAAC,MAAM,CAAC,EAAE,WAAW,CAAA;CAAO,GAChG,OAAO,CAAC,iBAAiB,CAAC,CAkE5B"}
|
package/dist/manager-call.js
CHANGED
|
@@ -1,5 +1,4 @@
|
|
|
1
|
-
import {
|
|
2
|
-
import { BASELINE_LIFECYCLE_ENDPOINT, assertLifecycleToken, dialerFor, invokeCommand, isPermissionDenied, issuedUserCaller, replyRefusedBeforeEffect, resolveService, standaloneConnectOpts, } from "@cotal-ai/core";
|
|
1
|
+
import { BASELINE_LIFECYCLE_ENDPOINT, assertLifecycleToken, dialerFor, goalFollowRefusal, invokeCommand, isPermissionDenied, issuedUserCaller, replyRefusedBeforeEffect, resolveService, standaloneConnectOpts, } from "@cotal-ai/core";
|
|
3
2
|
/** Decode only the routing coordinates. The broker verifies the signed bearer before any call. */
|
|
4
3
|
export function managerCallerBinding(bearer, config) {
|
|
5
4
|
let payload;
|
|
@@ -60,30 +59,17 @@ export async function invokeUserManager(config, bearer, command, args, opts = {}
|
|
|
60
59
|
const endpoint = BASELINE_LIFECYCLE_ENDPOINT;
|
|
61
60
|
const resolve = () => resolveService(nc, config.space, endpoint, caller, { instanceId, deadlineMs: opts.deadlineMs ?? 10_000, signal: opts.signal });
|
|
62
61
|
let service = await resolve();
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
v: 1,
|
|
67
|
-
id: randomUUID(),
|
|
68
|
-
ok: false,
|
|
69
|
-
data: undefined,
|
|
70
|
-
error: {
|
|
71
|
-
code: "failed-precondition",
|
|
72
|
-
message: `manager instance ${service.responder.instanceId} does not support "goal-result"; upgrade manager to enable durable goal following (SPEC 13.6)`,
|
|
73
|
-
outcome: "not-executed",
|
|
74
|
-
},
|
|
75
|
-
},
|
|
76
|
-
responder: { endpoint, instanceId: service.responder.instanceId, epoch: service.responder.epoch },
|
|
77
|
-
};
|
|
78
|
-
}
|
|
62
|
+
const unfollowable = opts.follow ? goalFollowRefusal(service) : undefined;
|
|
63
|
+
if (unfollowable)
|
|
64
|
+
return unfollowable;
|
|
79
65
|
const invoke = async () => {
|
|
80
66
|
const result = await invokeCommand(nc, config.space, service, command, args, opts);
|
|
81
67
|
if (result.reply.ok === false && replyRefusedBeforeEffect(result.reply.error)) {
|
|
82
68
|
try {
|
|
83
69
|
service = await resolve();
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
70
|
+
const unfollowable = opts.follow ? goalFollowRefusal(service) : undefined;
|
|
71
|
+
if (unfollowable)
|
|
72
|
+
throw new Error(unfollowable.reply.error.message);
|
|
87
73
|
}
|
|
88
74
|
catch (error) {
|
|
89
75
|
return { ...result, reply: { ...result.reply, error: {
|
package/dist/manager-call.js.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"manager-call.js","sourceRoot":"","sources":["../src/manager-call.ts"],"names":[],"mappings":"AAAA,OAAO,
|
|
1
|
+
{"version":3,"file":"manager-call.js","sourceRoot":"","sources":["../src/manager-call.ts"],"names":[],"mappings":"AAAA,OAAO,EACL,2BAA2B,EAC3B,oBAAoB,EACpB,SAAS,EACT,iBAAiB,EACjB,aAAa,EACb,kBAAkB,EAClB,gBAAgB,EAChB,wBAAwB,EACxB,cAAc,EACd,qBAAqB,GAItB,MAAM,gBAAgB,CAAC;AAKxB,kGAAkG;AAClG,MAAM,UAAU,oBAAoB,CAAC,MAAc,EAAE,MAAqB;IACxE,IAAI,OAAwE,CAAC;IAC7E,IAAI,CAAC;QACH,MAAM,KAAK,GAAG,MAAM,CAAC,KAAK,CAAC,GAAG,CAAC,CAAC;QAChC,IAAI,KAAK,CAAC,MAAM,KAAK,CAAC;YAAE,MAAM,IAAI,KAAK,CAAC,aAAa,CAAC,CAAC;QACvD,OAAO,GAAG,IAAI,CAAC,KAAK,CAAC,MAAM,CAAC,IAAI,CAAC,KAAK,CAAC,CAAC,CAAE,EAAE,WAAW,CAAC,CAAC,QAAQ,CAAC,MAAM,CAAC,CAAC,CAAC;IAC7E,CAAC;IAAC,MAAM,CAAC;QACP,MAAM,IAAI,KAAK,CAAC,qDAAqD,CAAC,CAAC;IACzE,CAAC;IACD,MAAM,IAAI,GAAG,MAAM,CAAC,QAAQ,CAAC;IAC7B,MAAM,GAAG,GAAG,OAAO,EAAE,GAAG,CAAC;IACzB,IAAI,CAAC,IAAI,IAAI,CAAC,GAAG,IAAI,GAAG,CAAC,IAAI,KAAK,gBAAgB,IAAI,OAAO,CAAC,GAAG,KAAK,IAAI,CAAC,KAAK;QAC5E,GAAG,CAAC,KAAK,KAAK,IAAI,CAAC,KAAK,IAAI,GAAG,CAAC,KAAK,KAAK,IAAI,CAAC,KAAK,IAAI,GAAG,CAAC,YAAY,KAAK,MAAM,CAAC,YAAY;QAChG,CAAC,CAAC,OAAO,CAAC,GAAG,KAAK,MAAM,CAAC,KAAK,IAAI,CAAC,KAAK,CAAC,OAAO,CAAC,OAAO,CAAC,GAAG,CAAC,IAAI,OAAO,CAAC,GAAG,CAAC,MAAM,KAAK,CAAC,IAAI,OAAO,CAAC,GAAG,CAAC,CAAC,CAAC,KAAK,MAAM,CAAC,KAAK,CAAC,CAAC;QAChI,MAAM,IAAI,KAAK,CAAC,6FAA6F,CAAC,CAAC;IACjH,IAAI,OAAO,GAAG,CAAC,iBAAiB,KAAK,QAAQ;QAC3C,MAAM,IAAI,KAAK,CAAC,gEAAgE,CAAC,CAAC;IACpF,MAAM,UAAU,GAAG,oBAAoB,CAAC,GAAG,CAAC,iBAAiB,EAAE,mBAAmB,CAAC,CAAC;IACpF,IAAI,MAAM,CAAC,iBAAiB,KAAK,SAAS,IAAI,UAAU,KAAK,MAAM,CAAC,iBAAiB;QACnF,MAAM,IAAI,KAAK,CAAC,gEAAgE,CAAC,CAAC;IACpF,IAAI,OAAO,GAAG,CAAC,YAAY,KAAK,QAAQ;QACtC,MAAM,IAAI,KAAK,CAAC,oDAAoD,CAAC,CAAC;IACxE,OAAO,EAAE,MAAM,EAAE,EAAE,KAAK,EAAE,IAAI,CAAC,KAAK,EAAE,KAAK,EAAE,IAAI,CAAC,KAAK,EAAE,GAAG,EAAE,GAAG,CAAC,YAAY,EAAE,EAAE,UAAU,EAAE,CAAC;AACjG,CAAC;AAED,mGAAmG;AACnG,MAAM,CAAC,KAAK,UAAU,iBAAiB,CACrC,MAAqB,EACrB,MAAc,EACd,OAAe,EACf,IAAyC,EACzC,OAA+F,EAAE;IAEjG,IAAI,CAAC,MAAM,EAAE,cAAc,EAAE,CAAC;IAC9B,MAAM,EAAE,MAAM,EAAE,MAAM,EAAE,UAAU,EAAE,GAAG,oBAAoB,CAAC,MAAM,EAAE,MAAM,CAAC,CAAC;IAC5E,MAAM,WAAW,GAAG,qBAAqB,CAAC,EAAE,MAAM,EAAE,aAAa,EAAE,MAAM,CAAC,QAAS,CAAC,aAAa,EAAE,GAAG,EAAE,MAAM,CAAC,GAAG,EAAE,CAAC,CAAC;IACtH,MAAM,EAAE,GAAG,MAAM,SAAS,CAAC,MAAM,CAAC,OAAO,CAAC,CAAC;QACzC,OAAO,EAAE,MAAM,CAAC,OAAO;QACvB,GAAG,WAAW;QACd,oBAAoB,EAAE,CAAC;QACvB,OAAO,EAAE,IAAI,CAAC,GAAG,CAAC,IAAI,CAAC,UAAU,IAAI,MAAM,EAAE,KAAK,CAAC;KACpD,CAAC,CAAC;IACH,kGAAkG;IAClG,mGAAmG;IACnG,yFAAyF;IACzF,wDAAwD;IACxD,IAAI,MAAM,GAAa,MAAM,CAAC;IAC9B,IAAI,CAAC;QACH,MAAM,GAAG,MAAM,gBAAgB,CAAC,EAAE,EAAE,MAAM,CAAC,KAAK,EAAE,MAAM,CAAC,WAAW,CAAC,IAAI,CAAC,EAAE,MAAM,CAAC,CAAC;IACtF,CAAC;IAAC,OAAO,CAAC,EAAE,CAAC;QACX,IAAI,CAAC,kBAAkB,CAAC,CAAC,CAAC,EAAE,CAAC;YAC3B,MAAM,EAAE,CAAC,KAAK,EAAE,CAAC;YACjB,MAAM,CAAC,CAAC;QACV,CAAC;IACH,CAAC;IACD,MAAM,OAAO,GAAG,GAAG,EAAE,GAAG,KAAK,EAAE,CAAC,KAAK,EAAE,CAAC,CAAC,CAAC,CAAC;IAC3C,IAAI,CAAC,MAAM,EAAE,gBAAgB,CAAC,OAAO,EAAE,OAAO,EAAE,EAAE,IAAI,EAAE,IAAI,EAAE,CAAC,CAAC;IAChE,IAAI,CAAC;QACH,0FAA0F;QAC1F,IAAI,CAAC,MAAM,EAAE,cAAc,EAAE,CAAC;QAC9B,MAAM,QAAQ,GAAG,2BAA2B,CAAC;QAC7C,MAAM,OAAO,GAAG,GAAG,EAAE,CAAC,cAAc,CAAC,EAAE,EAAE,MAAM,CAAC,KAAK,EAAE,QAAQ,EAAE,MAAM,EAAE,EAAE,UAAU,EAAE,UAAU,EAAE,IAAI,CAAC,UAAU,IAAI,MAAM,EAAE,MAAM,EAAE,IAAI,CAAC,MAAM,EAAE,CAAC,CAAC;QACrJ,IAAI,OAAO,GAAG,MAAM,OAAO,EAAE,CAAC;QAC9B,MAAM,YAAY,GAAG,IAAI,CAAC,MAAM,CAAC,CAAC,CAAC,iBAAiB,CAAC,OAAO,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC;QAC1E,IAAI,YAAY;YAAE,OAAO,YAAY,CAAC;QACtC,MAAM,MAAM,GAAG,KAAK,IAAI,EAAE;YACxB,MAAM,MAAM,GAAG,MAAM,aAAa,CAAC,EAAE,EAAE,MAAM,CAAC,KAAK,EAAE,OAAO,EAAE,OAAO,EAAE,IAAI,EAAE,IAAI,CAAC,CAAC;YACnF,IAAI,MAAM,CAAC,KAAK,CAAC,EAAE,KAAK,KAAK,IAAI,wBAAwB,CAAC,MAAM,CAAC,KAAK,CAAC,KAAK,CAAC,EAAE,CAAC;gBAC9E,IAAI,CAAC;oBACH,OAAO,GAAG,MAAM,OAAO,EAAE,CAAC;oBAC1B,MAAM,YAAY,GAAG,IAAI,CAAC,MAAM,CAAC,CAAC,CAAC,iBAAiB,CAAC,OAAO,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC;oBAC1E,IAAI,YAAY;wBAAE,MAAM,IAAI,KAAK,CAAC,YAAY,CAAC,KAAK,CAAC,KAAM,CAAC,OAAO,CAAC,CAAC;gBACvE,CAAC;gBAAC,OAAO,KAAK,EAAE,CAAC;oBACf,OAAO,EAAE,GAAG,MAAM,EAAE,KAAK,EAAE,EAAE,GAAG,MAAM,CAAC,KAAK,EAAE,KAAK,EAAE;gCACnD,GAAG,MAAM,CAAC,KAAK,CAAC,KAAM;gCACtB,OAAO,EAAE,GAAG,OAAO,kCAAkC,UAAU,sDAAsD,KAAK,YAAY,KAAK,CAAC,CAAC,CAAC,KAAK,CAAC,OAAO,CAAC,CAAC,CAAC,MAAM,CAAC,KAAK,CAAC,EAAE;6BAC9K,EAAE,EAAE,CAAC;gBACR,CAAC;gBACD,sFAAsF;gBACtF,gEAAgE;gBAChE,OAAO,aAAa,CAAC,EAAE,EAAE,MAAM,CAAC,KAAK,EAAE,OAAO,EAAE,OAAO,EAAE,IAAI,EAAE,IAAI,CAAC,CAAC;YACvE,CAAC;YACD,OAAO,MAAM,CAAC;QAChB,CAAC,CAAC;QACF,OAAO,MAAM,MAAM,EAAE,CAAC;IACxB,CAAC;YAAS,CAAC;QACT,IAAI,CAAC,MAAM,EAAE,mBAAmB,CAAC,OAAO,EAAE,OAAO,CAAC,CAAC;QACnD,IAAI,KAAgD,CAAC;QACrD,IAAI,CAAC;YACH,MAAM,OAAO,CAAC,IAAI,CAAC;gBACjB,EAAE,CAAC,KAAK,EAAE,CAAC,KAAK,CAAC,GAAG,EAAE,CAAC,EAAE,CAAC,KAAK,EAAE,CAAC;gBAClC,IAAI,OAAO,CAAO,CAAC,OAAO,EAAE,EAAE,GAAG,KAAK,GAAG,UAAU,CAAC,OAAO,EAAE,KAAK,CAAC,CAAC,CAAC,KAAK,CAAC,KAAK,EAAE,CAAC,CAAC,CAAC,CAAC;aACvF,CAAC,CAAC;QACL,CAAC;gBAAS,CAAC;YACT,IAAI,KAAK,KAAK,SAAS;gBAAE,YAAY,CAAC,KAAK,CAAC,CAAC;YAC7C,MAAM,EAAE,CAAC,KAAK,EAAE,CAAC;QACnB,CAAC;IACH,CAAC;AACH,CAAC"}
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@cotal-ai/connector-core",
|
|
3
3
|
"description": "Shared MCP-bridge runtime for Cotal connectors: the mesh agent, cotal_* tools, and hook relay.",
|
|
4
|
-
"version": "0.72.
|
|
4
|
+
"version": "0.72.1",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"repository": {
|
|
7
7
|
"type": "git",
|
|
@@ -35,7 +35,7 @@
|
|
|
35
35
|
"devDependencies": {
|
|
36
36
|
"@cotal-ai/smoke-kit": "0.0.0",
|
|
37
37
|
"@ag-ui/core": "0.0.57",
|
|
38
|
-
"@cotal-ai/core": "0.72.
|
|
38
|
+
"@cotal-ai/core": "0.72.1",
|
|
39
39
|
"@nats-io/transport-node": "^3.4.0"
|
|
40
40
|
},
|
|
41
41
|
"files": [
|