subharness 0.0.5 → 0.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (148) hide show
  1. package/README.md +59 -3
  2. package/dist/adapters/claude-process.js +8 -1
  3. package/dist/adapters/claude-process.js.map +1 -1
  4. package/dist/adapters/claude-tools.d.ts +6 -2
  5. package/dist/adapters/claude-tools.js +14 -12
  6. package/dist/adapters/claude-tools.js.map +1 -1
  7. package/dist/adapters/claude-worker-client.d.ts +49 -0
  8. package/dist/adapters/claude-worker-client.js +359 -0
  9. package/dist/adapters/claude-worker-client.js.map +1 -0
  10. package/dist/adapters/claude-worker-process.d.ts +20 -0
  11. package/dist/adapters/claude-worker-process.js +76 -0
  12. package/dist/adapters/claude-worker-process.js.map +1 -0
  13. package/dist/adapters/claude-worker-protocol.d.ts +38 -0
  14. package/dist/adapters/claude-worker-protocol.js +2 -0
  15. package/dist/adapters/claude-worker-protocol.js.map +1 -0
  16. package/dist/adapters/claude-worker.d.ts +1 -0
  17. package/dist/adapters/claude-worker.js +126 -0
  18. package/dist/adapters/claude-worker.js.map +1 -0
  19. package/dist/adapters/claude.d.ts +16 -6
  20. package/dist/adapters/claude.js +138 -68
  21. package/dist/adapters/claude.js.map +1 -1
  22. package/dist/cli/catalog-worker.js.map +1 -1
  23. package/dist/cli/dashboard-client.js +19 -30
  24. package/dist/cli/dashboard-client.js.map +1 -1
  25. package/dist/cli/main.js +53 -45
  26. package/dist/cli/main.js.map +1 -1
  27. package/dist/errors.d.ts +2 -2
  28. package/dist/errors.js.map +1 -1
  29. package/dist/index.d.ts +6 -0
  30. package/dist/index.js +3 -0
  31. package/dist/index.js.map +1 -1
  32. package/dist/runtime/approval-registry.d.ts +9 -0
  33. package/dist/runtime/approval-registry.js +143 -4
  34. package/dist/runtime/approval-registry.js.map +1 -1
  35. package/dist/runtime/capture.d.ts +22 -0
  36. package/dist/runtime/capture.js +320 -0
  37. package/dist/runtime/capture.js.map +1 -0
  38. package/dist/runtime/catalog.d.ts +39 -0
  39. package/dist/runtime/catalog.js +84 -0
  40. package/dist/runtime/catalog.js.map +1 -0
  41. package/dist/runtime/client.d.ts +2 -1
  42. package/dist/runtime/client.js +52 -65
  43. package/dist/runtime/client.js.map +1 -1
  44. package/dist/runtime/coordinator.d.ts +44 -7
  45. package/dist/runtime/coordinator.js +265 -31
  46. package/dist/runtime/coordinator.js.map +1 -1
  47. package/dist/runtime/daemon.js +12 -6
  48. package/dist/runtime/daemon.js.map +1 -1
  49. package/dist/runtime/dashboard-workspace.js +5 -1
  50. package/dist/runtime/dashboard-workspace.js.map +1 -1
  51. package/dist/runtime/definition.d.ts +6 -1
  52. package/dist/runtime/definition.js +31 -1
  53. package/dist/runtime/definition.js.map +1 -1
  54. package/dist/runtime/native-owner.js +7 -1
  55. package/dist/runtime/native-owner.js.map +1 -1
  56. package/dist/runtime/pagination.d.ts +9 -0
  57. package/dist/runtime/pagination.js +36 -0
  58. package/dist/runtime/pagination.js.map +1 -0
  59. package/dist/runtime/sdk-observation.d.ts +8 -0
  60. package/dist/runtime/sdk-observation.js +75 -0
  61. package/dist/runtime/sdk-observation.js.map +1 -0
  62. package/dist/runtime/sdk-projection.d.ts +8 -0
  63. package/dist/runtime/sdk-projection.js +19 -0
  64. package/dist/runtime/sdk-projection.js.map +1 -0
  65. package/dist/runtime/sdk-service.d.ts +11 -0
  66. package/dist/runtime/sdk-service.js +109 -0
  67. package/dist/runtime/sdk-service.js.map +1 -0
  68. package/dist/runtime/select-native.d.ts +4 -1
  69. package/dist/runtime/select-native.js +14 -17
  70. package/dist/runtime/select-native.js.map +1 -1
  71. package/dist/runtime/service.d.ts +11 -1
  72. package/dist/runtime/service.js +178 -3
  73. package/dist/runtime/service.js.map +1 -1
  74. package/dist/runtime/session-launcher.js +1 -1
  75. package/dist/runtime/session-launcher.js.map +1 -1
  76. package/dist/runtime/snapshots.d.ts +103 -0
  77. package/dist/runtime/snapshots.js +124 -0
  78. package/dist/runtime/snapshots.js.map +1 -0
  79. package/dist/runtime/state.d.ts +26 -0
  80. package/dist/runtime/state.js +86 -17
  81. package/dist/runtime/state.js.map +1 -1
  82. package/dist/runtime/task-data.d.ts +38 -0
  83. package/dist/runtime/task-data.js +2 -0
  84. package/dist/runtime/task-data.js.map +1 -0
  85. package/dist/runtime/transport.d.ts +13 -0
  86. package/dist/runtime/transport.js +166 -0
  87. package/dist/runtime/transport.js.map +1 -0
  88. package/dist/runtime/types.d.ts +28 -9
  89. package/dist/runtime/types.js.map +1 -1
  90. package/dist/runtime/worker-client.js +2 -0
  91. package/dist/runtime/worker-client.js.map +1 -1
  92. package/dist/runtime/worker-server.js +105 -24
  93. package/dist/runtime/worker-server.js.map +1 -1
  94. package/dist/sdk/connected-types.d.ts +42 -0
  95. package/dist/sdk/connected-types.js +2 -0
  96. package/dist/sdk/connected-types.js.map +1 -0
  97. package/dist/sdk/connected.d.ts +3 -0
  98. package/dist/sdk/connected.js +232 -0
  99. package/dist/sdk/connected.js.map +1 -0
  100. package/dist/sdk/execution-definition.d.ts +3 -0
  101. package/dist/sdk/execution-definition.js +59 -0
  102. package/dist/sdk/execution-definition.js.map +1 -0
  103. package/dist/sdk/execution-driver.d.ts +6 -0
  104. package/dist/sdk/execution-driver.js +86 -0
  105. package/dist/sdk/execution-driver.js.map +1 -0
  106. package/dist/sdk/execution-observation.d.ts +3 -0
  107. package/dist/sdk/execution-observation.js +35 -0
  108. package/dist/sdk/execution-observation.js.map +1 -0
  109. package/dist/sdk/execution-options.d.ts +17 -0
  110. package/dist/sdk/execution-options.js +115 -0
  111. package/dist/sdk/execution-options.js.map +1 -0
  112. package/dist/sdk/execution-types.d.ts +72 -0
  113. package/dist/sdk/execution-types.js +2 -0
  114. package/dist/sdk/execution-types.js.map +1 -0
  115. package/dist/sdk/execution.d.ts +3 -0
  116. package/dist/sdk/execution.js +3 -0
  117. package/dist/sdk/execution.js.map +1 -0
  118. package/dist/sdk/hosted-error.d.ts +3 -0
  119. package/dist/sdk/hosted-error.js +14 -0
  120. package/dist/sdk/hosted-error.js.map +1 -0
  121. package/dist/sdk/hosted.d.ts +23 -0
  122. package/dist/sdk/hosted.js +204 -0
  123. package/dist/sdk/hosted.js.map +1 -0
  124. package/dist/sdk/runner.d.ts +5 -0
  125. package/dist/sdk/runner.js +107 -0
  126. package/dist/sdk/runner.js.map +1 -0
  127. package/dist/sdk/tools.js +4 -1
  128. package/dist/sdk/tools.js.map +1 -1
  129. package/package.json +3 -3
  130. package/sdk/access-config.md +2 -0
  131. package/sdk/adapter-contract.md +8 -0
  132. package/sdk/agent.md +1 -1
  133. package/sdk/approvals.md +2 -0
  134. package/sdk/completion-notifications.md +2 -0
  135. package/sdk/distribution.md +4 -2
  136. package/sdk/examples/chat-tool.ts +126 -0
  137. package/sdk/execution.md +743 -0
  138. package/sdk/index.md +45 -21
  139. package/sdk/message-delivery.md +2 -0
  140. package/sdk/permissions.md +2 -0
  141. package/sdk/plugins/sub-agents.md +4 -2
  142. package/sdk/project-team.md +23 -1
  143. package/sdk/sessions.md +2 -0
  144. package/sdk/tools.md +3 -1
  145. package/sdk/v1-runtime.md +7 -5
  146. package/dist/cli/catalog.d.ts +0 -26
  147. package/dist/cli/catalog.js +0 -41
  148. package/dist/cli/catalog.js.map +0 -1
@@ -0,0 +1,743 @@
1
+ # Connected and embedded execution
2
+
3
+ The Node.js SDK runs native agents inside an application or connects to the shared local coordinator used by the CLI. It provides task states, complete responses, and structured approvals. Native harness executables remain required. Applications own their UI transport and user authorization; this is not a browser/edge runtime or an implementation of the A2A wire protocol.
4
+
5
+ Two entry points serve different ownership needs. **`connect()` connects to the same selected local coordinator as the CLI, starting it if absent**: it finds existing conversations without recreating agents, or connects to a newly started empty coordinator. `connect({ start: false })` requires that coordinator to exist already. **`createRunner()` embeds execution in the application process**: it owns isolated sessions and can use application tool closures. Both expose the same session and task handles. A session is a native conversation with a queue; a task is coordinated work including child agents and parent continuations.
6
+
7
+ A connected client disconnects without stopping the shared coordinator or its agents. A coordinator started by `connect()` also survives the calling application exiting. An embedded runner closes by cleaning up its owned execution. Neither mode adds durable restart recovery. These are Node.js APIs used behind a UI; coordinator/model credentials remain on the server, not in browser code.
8
+
9
+ Read the examples in order, or jump to [connection startup](#connect-without-running-the-cli-first), [agent discovery and creation](#discover-check-and-create-connected-sessions), [CLI attachment](#connect-to-cli-work), [paged inspection](#list-and-inspect-execution), [reconnection](#disconnect-and-reconnect), [embedded teams](#define-an-embedded-team), [queueing](#follow-up-or-queue-work), [steering](#correct-or-replace-active-work), [UI controls](#serve-a-ui), [approval buttons](#render-approval-buttons), [grant forms](#select-permission-grants), [chat tools](#consult-from-a-chat-tool), [response replay](#replay-complete-responses), [signatures](#signature-reference), or [CLI/library parity](#one-library-for-sdk-and-cli).
10
+
11
+ Examples awaiting only `task.result` assume no interactive approval is needed, or a separate UI/component handles requests. Pending approvals never auto-resolve. Helpers identified as application-owned are illustrative, not SDK exports. The first version exposes task states, complete responses, and approvals; live tokens and native tool-progress events are outside its scope.
12
+
13
+ ## Connect without running the CLI first
14
+
15
+ The SDK does not require a prior CLI command. By default, `connect()` reuses the selected shared local coordinator or starts it when it is confirmed absent. Starting a coordinator creates no sessions or agents and invokes no model.
16
+
17
+ ```ts
18
+ import { connect } from "subharness";
19
+
20
+ const client = await connect(); // Equivalent to connect({ start: true }).
21
+ try {
22
+ const { items: sessions } = await client.sessions.list();
23
+ console.log(sessions); // [] when no sessions have been created.
24
+ } finally {
25
+ await client.disconnect(); // The shared coordinator keeps running.
26
+ }
27
+ ```
28
+
29
+ Use strict attachment when the application should observe only an existing coordinator:
30
+
31
+ ```ts
32
+ const client = await connect({ start: false });
33
+ try {
34
+ console.log((await client.sessions.list()).items);
35
+ } finally {
36
+ await client.disconnect();
37
+ }
38
+ ```
39
+
40
+ Strict attachment fails when the selected coordinator is absent. A running coordinator with no sessions is still a successful connection. Both forms authenticate and validate the selected instance and protocol support. Authentication failure, protocol incompatibility, or uncertain reachability fails the call; it does not authorize starting a replacement coordinator. Concurrent connection calls that need startup converge on one selected coordinator. The discovery, protocol, and startup rules are specified below under ownership and transport.
41
+
42
+ ## Discover, check, and create connected sessions
43
+
44
+ The agent catalog describes selectable definitions; it is separate from sessions already running. Discovery uses the given existing absolute directory on the coordinator host. Here `cwd` is the application's selected project directory, and `client` is an open connected client.
45
+
46
+ ```ts
47
+ import { codex } from "subharness";
48
+
49
+ const { items: agents } = await client.agents.list({ cwd });
50
+ console.log(agents); // IDs, names, descriptions, and harness/repo/global scope.
51
+
52
+ await client.agents.check({ agent: "repo:researcher", cwd }); // Optional startup check.
53
+ const session = await client.sessions.create({ agent: "repo:researcher", cwd });
54
+ const { task } = await session.prompt("Find the queue's documented guarantees.");
55
+ console.log((await task.result).text);
56
+
57
+ // Alternatively, configure a direct native harness without a definition file.
58
+ const direct = await client.sessions.create({
59
+ agent: codex({ model: "CODEX_MODEL_ID", effort: "high" }),
60
+ cwd,
61
+ });
62
+ ```
63
+
64
+ A string selects a discovered qualified reference or unambiguous name. Reserved strings `"codex"`, `"claude"`, and `"fx"` select direct harnesses with native defaults. OpenCode, Copilot, and Cursor require their existing explicit model configuration through a harness constructor; their bare reserved names reject with `INVALID_CONFIG`. All six implemented harnesses retain their documented access and capability limits. Existing harness constructors accept explicit model/options; their model identifiers remain required. `subagent:` references belong to the private managed-agent context, not unrestricted public creation. Unknown or ambiguous references reject with `UNKNOWN_AGENT` or `AMBIGUOUS_AGENT`.
65
+
66
+ Catalog discovery imports trusted definition modules in isolated coordinator-owned evaluation; it does not start native harnesses or call models. The finite discovered-definition inventory remains unpaginated; there is no artificial catalog-size cap. Discovery is not proof of readiness. `check` opens and closes a disposable native startup session without a prompt, retained conversation, or task; it rejects safely on startup/authentication/configuration/cleanup failure and does not promise a later model call will succeed.
67
+
68
+ `create` resolves the target and captures its loaded definition graph and environment once, then returns an empty session. Editing definition files affects newly created sessions, not an existing captured conversation. Definition functions live on the coordinator host; client closures cannot be transferred. Native harness startup stays lazy until dispatch. Invalid creation returns no session. A separate invalid `prompt` can leave a valid empty session, which the caller may reuse or close. The CLI's atomic first-prompt boundary is preserved separately below.
69
+
70
+ For connected creation/check, omission of `env` snapshots the calling process environment at that operation. An explicit `env` replaces it, with undefined entries omitted. Inherited `AGENT_PARENT_TOKEN` and `AGENT_CLI_PATH` are stripped, and the selected coordinator supplies `AGENT_STATE_DIR`; managed authority is bound separately. This resulting environment governs definition resolution and module evaluation as well as native startup/execution, so environment-dependent imports use the same chosen values. Catalog listing evaluates definitions with its implicit caller snapshot under the same stripping/injection rules. Existing cwd-based project discovery and personal-settings conventions still apply. The environment is privately transferred and captured for the session; follow-ups do not replace it. No caller-supplied environment grants parent authority. Neither the host cwd nor the shared daemon's `process.env` is mutated, and environment/access values never appear in metadata or errors.
71
+
72
+ Connected and CLI-created sessions retain their loaded definition graph in an isolated owner process. Custom-tool callbacks execute in that graph owner's creation directory and captured environment, including callbacks used by managed children. A child's native harness may use a different requested working directory and native environment; that does not change the captured callback's ambient `process.cwd()` or `process.env`. Tools that operate on a child-specific location must use explicit paths or arguments. The graph owner does not inherit caller-supplied managed parent or launcher authority; native delegation instructions and launchers are bound separately to each session. This preserves closure state and avoids re-importing changed source or mutating shared process context.
73
+
74
+ Connected sessions retain the CLI's existing per-project personal-settings and native-worker access resolution. Embedded runners retain their explicit `RunnerOptions.env`/`access` behavior and do not start reading those personal settings. Model credentials remain in the server/coordinator execution path, not UI payloads.
75
+
76
+ ## Connect to CLI work
77
+
78
+ Start a task using the existing CLI. This example uses the repository's configured `repo:architect` agent; use a valid discovered agent reference in another repository.
79
+
80
+ ```sh
81
+ subharness run repo:architect --prompt "Review the SDK contract." --detach --format jsonl
82
+ ```
83
+
84
+ The admission receipt includes `sessionId` and `taskId`. Supply those values to the SDK; the names below represent those returned strings. The connected API retrieves the same records and native conversation, without loading the agent definition again or invoking a model merely to inspect it.
85
+
86
+ ```ts
87
+ import { connect, AgentError } from "subharness";
88
+
89
+ const client = await connect();
90
+ try {
91
+ const { items: sessions } = await client.sessions.list();
92
+ console.log(sessions);
93
+ const { items: tasks } = await client.tasks.list({ sessionId });
94
+ console.log(tasks);
95
+
96
+ const session = await client.sessions.get(sessionId);
97
+ const task = await client.tasks.get(taskId);
98
+ console.log(await session.snapshot());
99
+ console.log(await task.snapshot());
100
+ console.log((await task.result).text);
101
+
102
+ // Continue the same conversation started by the CLI.
103
+ const { task: followup } = await session.prompt("Explain the main tradeoff.");
104
+ console.log((await followup.result).text);
105
+ } finally {
106
+ await client.disconnect();
107
+ }
108
+ ```
109
+
110
+ The client uses the selected shared local coordinator; working directory alone does not select one, and independent embedded runners do not appear in these lists. Use `connect({ start: false })` to attach without starting a coordinator; retrieving the receipt's IDs still requires the original coordinator to retain them. Default `connect()` may start an empty coordinator after the previous one has exited; it cannot restore the receipt's old IDs.
111
+
112
+ Lists expose retained summaries, including terminal tasks and closed sessions. The snippets above show the first page; use `nextCursor` to enumerate more. `tasks.list()` without a session filter covers all retained tasks in the selected coordinator. A supplied session ID must exist. `get` verifies existence asynchronously. Snapshot reads are asynchronous in **both** modes so the same consuming code works locally or through a connection.
113
+
114
+ ## List and inspect execution
115
+
116
+ Use metadata for lists and request expanded task detail only when needed. This example enumerates one session's tasks; `renderRow` and `openDetails` are application-owned UI helpers.
117
+
118
+ ```ts
119
+ let cursor: string | undefined;
120
+ do {
121
+ const page = await client.tasks.list({ sessionId, limit: 100, cursor });
122
+ for (const summary of page.items) {
123
+ renderRow(summary.taskId, summary.title, summary.state, summary.updatedAt);
124
+ }
125
+ cursor = page.nextCursor;
126
+ } while (cursor !== undefined);
127
+
128
+ const task = await client.tasks.get(taskId);
129
+ const details = await task.inspect();
130
+ openDetails({ prompt: details.prompt, response: details.latestResponse });
131
+ ```
132
+
133
+ Session/task pages default to 100 entries, with integer `limit` from 1 through 1,000. Entries follow admission insertion order. The first page captures a membership watermark; later pages exclude newly admitted records. Each page contains current state when read, not a transactionally frozen state snapshot. Opaque cursors bind the collection, coordinator instance, and filter; repeat the same `sessionId` filter when continuing. The limit may change between pages. Invalid or mismatched cursors reject with `INVALID_CURSOR`; cursors remain valid while that coordinator retains its records. No cursor means a fresh traversal.
134
+
135
+ There is no collection-watch API in this version. Poll paged lists to find new work, and attach `task.watch()` for race-free current-state observation of selected tasks. Per-task revisions are not global list cursors. Lists contain no full prompts, responses, credentials, or environment values. `task.inspect()` adds the original admission prompt to the task snapshot; it is separate from every watch update. Steering does not overwrite that original prompt.
136
+
137
+ Task titles use the existing first nonempty prompt line, normalize whitespace/remove terminal controls, and cap at 120 Unicode code points. Task ordinals count admission within a session from one; steering retains the task's ordinal. Timestamps are Unix milliseconds. `updatedAt` starts at admission and records the coordinator's latest task activity update, not an observer read. `dispatchedAt` records first dispatch and is retained across native continuations and private recovery; it is absent for never-dispatched work. `failedAt` records the latest execution or stop-confirmation failure, and `finishedAt` records the terminal transition. Both terminal timestamps are cleared when private recovery reopens the task and are populated again on its next failure/terminal transition. Session summaries include execution cwd and workspace identity. Connected sessions can resolve repository grouping using existing CLI behavior; embedded sessions use their directory without introducing Git discovery.
138
+
139
+ ## Disconnect and reconnect
140
+
141
+ This independent example starts with the CLI receipt's IDs. An existing client may exit while its task continues, including while an approval is pending.
142
+
143
+ ```ts
144
+ const firstClient = await connect({ start: false });
145
+ const firstHandle = await firstClient.tasks.get(taskId);
146
+ console.log(await firstHandle.snapshot());
147
+ await firstClient.disconnect();
148
+ // firstHandle is now invalid. Execution was not cancelled.
149
+
150
+ const nextClient = await connect({ start: false });
151
+ try {
152
+ const task = await nextClient.tasks.get(taskId);
153
+ console.log(await task.snapshot());
154
+ console.log((await task.result).text);
155
+ } finally {
156
+ await nextClient.disconnect();
157
+ }
158
+ ```
159
+
160
+ This example uses strict attachment because it needs retained IDs, not a new empty coordinator. Reconnection works only while the same coordinator retains those IDs. Disconnect releases transport and observers, invalidates that client's handles, and rejects pending connected observations, including `result`, with a safe transport error. This does not fail the server task. A new handle can observe its retained outcome. Coordinator loss/replacement surfaces unavailable records; connection never replays a prompt or substitutes a new session. A new default `connect()` call can start an empty coordinator after the previous instance is confirmed gone, but never restores its sessions, tasks, or queues. Existing handles remain bound to their original instance.
161
+
162
+ ## Define an embedded team
163
+
164
+ These static definition helpers already exist. Model IDs below are placeholders: select real models and supported effort levels through the configured native harness. Cost follows those explicit choices; Subharness does not infer a budget or silently choose another model.
165
+
166
+ ```ts
167
+ import { agent, codex, createRunner, tool, AgentError } from "subharness";
168
+ import { z } from "zod";
169
+
170
+ const researcher = agent({
171
+ name: "researcher",
172
+ description: "Finds source evidence for a bounded question.",
173
+ instructions: "Find relevant sources. Return evidence and uncertainty concisely.",
174
+ harness: codex({ model: "LOW_COST_CODEX_MODEL_ID", effort: "low" }),
175
+ });
176
+
177
+ // Application state captured by a custom-tool closure stays in this process.
178
+ const notes = new Map<string, string>();
179
+ const remember = tool({
180
+ description: "Save a short research note in this application scope.",
181
+ inputSchema: z.object({ key: z.string(), text: z.string() }),
182
+ execute: ({ key, text }) => {
183
+ notes.set(key, text);
184
+ return "Saved";
185
+ },
186
+ });
187
+
188
+ const team = agent({
189
+ name: "lead",
190
+ description: "Reasons about the problem and coordinates its solution.",
191
+ instructions: "Delegate source lookup to research. Evaluate the evidence yourself.",
192
+ harness: codex({ model: "ADVANCED_CODEX_MODEL_ID", effort: "high" }),
193
+ subagents: { research: researcher },
194
+ tools: { remember },
195
+ });
196
+ ```
197
+
198
+ The declaration lets the lead call its research specialist through managed delegation. The runner owns child execution and exposes child state and approvals through the parent. The notes map and tool closure stay in the embedding application; `connect()` does not serialize them into the coordinator. Defining a child does not guarantee the model will delegate every suitable task.
199
+
200
+ ## Run an embedded script
201
+
202
+ This example assumes the operation needs no interactive approval, or another application component handles requests. Awaiting `result` alone never auto-approves a request and can remain pending until one is answered.
203
+
204
+ ```ts
205
+ const runner = createRunner();
206
+ try {
207
+ const session = await runner.createSession(team, { cwd: process.cwd() });
208
+ const { task } = await session.prompt("Explain how this repository queues tasks.");
209
+ const { text } = await task.result;
210
+ console.log(text);
211
+ } finally {
212
+ await runner.close();
213
+ }
214
+ ```
215
+
216
+ Runner options retain their current environment and access semantics. Native harness executables remain required. The library runs in Node.js behind the UI, not in the browser. Closing cancels unfinished owned work and waits for cleanup; it does not undo file or tool effects. Cleanup failures remain explicit, and uncooperative application tool callbacks can delay close.
217
+
218
+ ## Follow up or queue work
219
+
220
+ This and the steering examples use an open `session`, obtained either from `client.sessions.get(id)` or `runner.createSession(...)`. Keep its owner alive for the chat/job scope; the embedded script above closes its own separate scope.
221
+
222
+ ```ts
223
+ const { task: explanation } = await session.prompt("Explain the task queue.");
224
+ console.log((await explanation.result).text);
225
+
226
+ const { task: followup } = await session.prompt("Give an example of that behavior.");
227
+ console.log((await followup.result).text);
228
+
229
+ // Submitting before earlier work finishes queues a separate task by default.
230
+ const { task: first } = await session.prompt("Review the approval contract.");
231
+ const receipt = await session.prompt("Summarize that review in three points.");
232
+ const second = receipt.task;
233
+ if (receipt.paused) console.log("Blocked by:", receipt.blockedByTaskId);
234
+ try {
235
+ await first.result;
236
+ console.log((await second.result).text);
237
+ } catch (error) {
238
+ console.error(error);
239
+ console.log(await session.snapshot());
240
+ }
241
+ ```
242
+
243
+ The native conversation retains history. Default delivery is FIFO `queue`, with one active task and the existing 100-task queue limit. A failed task pauses dispatch; queued results may remain pending. The optional receipt pair `paused: true` and `blockedByTaskId` captures that condition atomically at admission; both fields are absent otherwise, including native steering acceptance. A later snapshot is current state, not a replacement for that admission evidence. Inspect the session before retrying. Under the existing recovery contract, cancelling retained failed work confirms stop and releases its paused queue without erasing its failure. There is no prompt-free SDK resume method. Sending another prompt does not answer an approval.
244
+
245
+ ## Correct or replace active work
246
+
247
+ To correct work, first submit a task and then steer it without awaiting completion:
248
+
249
+ ```ts
250
+ const { task: original } = await session.prompt("Review the public API and propose improvements.");
251
+ const correction = await session.prompt("Keep the existing public API unchanged.", {
252
+ delivery: "steer",
253
+ });
254
+ console.log(correction.requestedDelivery, correction.effectiveDelivery);
255
+ console.log("Same task:", correction.task.id === original.id);
256
+ console.log((await correction.task.result).text);
257
+ ```
258
+
259
+ In a separate case, start work and then explicitly replace the active task:
260
+
261
+ ```ts
262
+ const { task: original } = await session.prompt("Review the implementation and the docs.");
263
+ const replacement = await session.prompt("Stop that approach and review only the docs.", {
264
+ delivery: "interrupt",
265
+ });
266
+ console.log(original.id, replacement.task.id, replacement.effectiveDelivery);
267
+ console.log((await replacement.task.result).text);
268
+ ```
269
+
270
+ These examples assume an initially idle session and that the submitted work remains active when the control is handled. Steering requires an active task; otherwise it rejects with `NO_ACTIVE_TASK`. There is no expected-task guard: control applies to whichever task is active when handled, so completion and queued dispatch can race with the call. Supported native steering can return the same task. Unsupported steering falls back to confirmed interruption and reports `effectiveDelivery: "interrupt"`; unrelated errors do not trigger fallback. Interruption replaces the active task after stopping it and its descendants, preserving unrelated queued work. The interrupted task's result rejects with `INTERRUPTED`. The receipt therefore describes this particular prompt submission, separately from its task.
271
+
272
+ ## Serve a UI
273
+
274
+ This example uses a shared, server-owned connected `client` kept alive across HTTP requests. `User`, `publish`, `authorizeSession`, `authorizeTask`, and `authorizeApproval` belong to the application. Each authorization helper verifies that the specific resource belongs to the authenticated user's chat, including any permitted descendant resources. Local coordinator authentication and opaque IDs do not provide application multi-user authorization.
275
+
276
+ ```ts
277
+ async function submitMessage(user: User, chatId: string, sessionId: string, text: string) {
278
+ await authorizeSession(user, chatId, sessionId);
279
+ const session = await client.sessions.get(sessionId);
280
+ const { task } = await session.prompt(text);
281
+ return { taskId: task.id, sessionId: task.sessionId };
282
+ }
283
+
284
+ async function watchTask(
285
+ user: User, chatId: string, taskId: string,
286
+ signal: AbortSignal, publish: (value: unknown) => Promise<void>,
287
+ ) {
288
+ await authorizeTask(user, chatId, taskId);
289
+ try {
290
+ const task = await client.tasks.get(taskId);
291
+ for await (const snapshot of task.watch({ signal })) {
292
+ await publish(snapshot);
293
+ // Render state, latestResponse, children, and requests authorized for this chat.
294
+ }
295
+ } catch (error) {
296
+ if (error instanceof AgentError && error.code === "OBSERVATION_CLOSED") return;
297
+ throw error; // Route maps transport failures separately from task failures.
298
+ }
299
+ }
300
+
301
+ async function answerApproval(
302
+ user: User, chatId: string, requestId: string, response: ApprovalResponse,
303
+ ) {
304
+ await authorizeApproval(user, chatId, requestId);
305
+ await client.approvals.respond(requestId, response);
306
+ }
307
+
308
+ async function stopTask(user: User, chatId: string, taskId: string) {
309
+ await authorizeTask(user, chatId, taskId);
310
+ return (await client.tasks.get(taskId)).cancel();
311
+ }
312
+ ```
313
+
314
+ A browser disconnect aborts only that watch. Do not disconnect the shared client in an HTTP request's `finally` block. The server disconnects it when its connection scope ends; explicit `task.cancel()` stops a task, and `session.close()` closes server-side execution for that session. An embedded backend instead owns a bounded runner, uses its local lookups and `respond`, and closes that runner when its chat/job scope ends.
315
+
316
+ An initial list followed by each selected task's `watch()` gives a current view and race-free per-task observation: watch begins with the latest snapshot. It does not subscribe to newly created tasks; refresh the list as needed. There is no collection-watch API or global revision cursor. Parent snapshots include descendant approvals even when the parent is `running`. Authorize those relationships and treat displayed native content as untrusted data.
317
+
318
+ ## Render approval buttons
319
+
320
+ For a decision-only request, selecting an action needs only its ID. This illustrative JSX uses application components `Button` and `GrantApprovalForm`, plus an application-owned `submitApproval` function that calls the authenticated approval route above. They are not SDK exports.
321
+
322
+ ```tsx
323
+ function ApprovalButtons({ request }: { request: ApprovalRequest }) {
324
+ if (request.grants !== undefined) {
325
+ return <GrantApprovalForm request={request} />;
326
+ }
327
+ return (
328
+ <section>
329
+ <p>{request.message}</p>
330
+ {request.actions.map(action => (
331
+ <Button key={action.id}
332
+ onClick={() => submitApproval(request.id, { actionId: action.id })}>
333
+ {action.label}
334
+ {action.description && <small>{action.description}</small>}
335
+ </Button>
336
+ ))}
337
+ </section>
338
+ );
339
+ }
340
+ ```
341
+
342
+ The UI might display “Allow once”, “Allow for this session”, and “Deny”, but it renders only the supplied actions. Each action includes a normalized `decision` and, for allow/deny, a `scope` for presentation. Labels and descriptions explain what the choice covers. A generic consumer needs no Codex, Claude Code, or fx branches. The application owns layout, handles submission errors, and refreshes withdrawn requests.
343
+
344
+ An application can use a menu or dialog instead of buttons. Its selected action still comes from the request; it does not construct a decision/scope answer or infer an ID from a label. This server-side example assumes the same explicit authorization as the route above. An embedded owner substitutes `runner.respond` for `client.approvals.respond`.
345
+
346
+ ```ts
347
+ async function answerSelectedAction(request: ApprovalRequest, action: ApprovalAction) {
348
+ if (request.grants !== undefined) {
349
+ throw new Error("Use the grant form for this request.");
350
+ }
351
+ await client.approvals.respond(request.id, { actionId: action.id });
352
+ }
353
+ ```
354
+
355
+ ## Select permission grants
356
+
357
+ For `GrantApprovalForm`, render one checkbox per `request.grants` entry, using its ID, label, and optional description. Start with no checked boxes and require an explicit submit action. The application-owned helpers below prepare the answer without interpreting native paths or permission profiles.
358
+
359
+ ```ts
360
+ function createGrantFormState(request: ApprovalRequest) {
361
+ const selectedGrantIds = new Set<string>(); // Nothing selected for this request.
362
+ return {
363
+ toggle(id: string, checked: boolean) {
364
+ if (!request.grants?.some(grant => grant.id === id)) {
365
+ throw new Error("Unknown grant for this approval request.");
366
+ }
367
+ if (checked) selectedGrantIds.add(id);
368
+ else selectedGrantIds.delete(id);
369
+ },
370
+ async submit(action: ApprovalAction) {
371
+ const response: ApprovalResponse = action.decision === "allow"
372
+ ? { actionId: action.id, grants: [...selectedGrantIds] }
373
+ : { actionId: action.id };
374
+ await submitApproval(request.id, response);
375
+ },
376
+ };
377
+ }
378
+ ```
379
+
380
+ Create one form state for each request ID and bind its handlers to that request's checkboxes and actions. Different requests cannot share selections. Only IDs from that request's grants are selectable; the server validates the selected action and grants again. With no boxes selected, an allow submission explicitly contains `grants: []`, meaning no offered additional grants. Deny/cancel selections never include that array. An action requiring grant selection cannot be submitted without an explicit selection.
381
+
382
+ ## Consult from a chat tool
383
+
384
+ An ordinary callback can be registered with the host chat's tool mechanism. `specialistSession` is an explicitly authorized existing connected session or a session created by this chat's embedded runner, possibly with declared children. `publishTask` is the application's notification helper so its UI can observe and answer approvals while the tool waits.
385
+
386
+ ```ts
387
+ async function consultSpecialist(prompt: string): Promise<string> {
388
+ const { task } = await specialistSession.prompt(prompt);
389
+ await publishTask({ taskId: task.id, sessionId: task.sessionId });
390
+ const { text } = await task.result;
391
+ return text;
392
+ }
393
+ ```
394
+
395
+ The host chat supplies its own input schema and error handling when registering this callback. It receives the final answer or a rejected call; Subharness does not replace the host model loop. Submitting a public prompt from an application callback does not establish managed child lineage. Declared `subagents` use the separate hosted delegation mechanism with its existing ownership rules.
396
+
397
+ The packaged [chat-tool example](examples/chat-tool.ts) exposes an application-owned `onTask(task: TaskHandle)` callback immediately after admission, so a UI can begin watching approvals and offer cancellation before the host tool finishes.
398
+
399
+ For background consultation, the application can return the admitted task ID immediately and await its result separately. The application must then schedule any later host-chat continuation and correlate the result to the original tool call. UI disconnect or a failed notification does not cancel an admitted task.
400
+
401
+ ## Replay complete responses
402
+
403
+ `responses` is useful when every complete reply matters. `CompleteResponse` retains its current `{ responseId, text, state }` fields. `lastSeenResponseId`, `displayResponse`, and `saveCursor` below are application-owned; omit `after` to start from the first retained response.
404
+
405
+ ```ts
406
+ try {
407
+ for await (const response of task.responses({ after: lastSeenResponseId })) {
408
+ await displayResponse(response);
409
+ await saveCursor(response.responseId);
410
+ }
411
+ const final = await task.result;
412
+ console.log("Final answer:", final.text);
413
+ } catch (error) {
414
+ if (error instanceof AgentError) {
415
+ if (error.code === "COORDINATOR_UNAVAILABLE") {
416
+ console.error("Connection lost; the task may still be running.");
417
+ throw error; // Reconnect and obtain a fresh handle before further reads.
418
+ }
419
+ if (error.code === "CANCELLED" || error.code === "INTERRUPTED") {
420
+ console.log("Task stopped:", error.code);
421
+ } else {
422
+ console.error(error.code, error.message);
423
+ }
424
+ } else {
425
+ throw error;
426
+ }
427
+ console.log(await task.snapshot());
428
+ }
429
+ ```
430
+
431
+ The iterator drains retained replies and ends normally after task termination; `task.result` reports execution failure. A waiting reply is not the final answer, and the final result can refer to a response already displayed; deduplicate by `responseId`. Unknown or cross-task cursors reject with `UNKNOWN_RESPONSE`. Passing `signal` aborts only observation with `OBSERVATION_CLOSED`; breaking the loop also releases only the observer. Response readers are independent and do not consume each other's records. Histories remain retained within their execution owner's lifetime; embedded runner scopes should be bounded. Connected transport failure can interrupt this observation without changing the task outcome, and even a snapshot read can then fail.
432
+
433
+ ## Signature reference
434
+
435
+ Agent and harness definitions retain their [definition contracts](agent.md). The signatures below describe execution handles and display data; `TaskSnapshot.requests` contains normalized approval requests. Native [access configuration](access-config.md) remains separate.
436
+
437
+ ```ts
438
+ interface RunnerOptions {
439
+ readonly env?: Readonly<Record<string, string | undefined>>;
440
+ readonly access?: AccessSelection;
441
+ }
442
+ type AccessSelection = Readonly<Partial<Record<HarnessKind, readonly AccessConnection[]>>>;
443
+ type AccessConnection =
444
+ | { readonly type: "subscription"; readonly provider?: string }
445
+ | { readonly type: "api-key"; readonly provider?: string; readonly env?: string; readonly envFile?: string;
446
+ readonly baseUrl?: string; readonly wireApi?: "completions" | "responses";
447
+ readonly apiVersion?: string; readonly wireModel?: string }
448
+ | { readonly type: "vercel-api-key"; readonly env?: string; readonly envFile?: string }
449
+ | { readonly type: "vercel-oidc"; readonly env?: string; readonly envFile?: string; readonly project?: string };
450
+
451
+ interface ConnectOptions {
452
+ readonly start?: boolean; // Default: true. Start the shared coordinator if absent.
453
+ readonly stateDirectory?: string; // Explicit absolute local state directory.
454
+ }
455
+ declare function connect(options?: ConnectOptions): Promise<Client>;
456
+ interface ConnectedSessionOptions {
457
+ readonly agent: string | HarnessConfig;
458
+ readonly cwd: string;
459
+ readonly env?: Readonly<Record<string, string | undefined>>;
460
+ }
461
+ interface AgentSummary {
462
+ readonly id: string;
463
+ readonly name: string;
464
+ readonly description: string;
465
+ readonly scope: "harness" | "repo" | "global";
466
+ }
467
+ interface PageOptions { readonly limit?: number; readonly cursor?: string }
468
+ interface Page<T> { readonly items: readonly T[]; readonly nextCursor?: string }
469
+ interface Client {
470
+ readonly agents: {
471
+ list(options: { readonly cwd: string }): Promise<{ readonly items: readonly AgentSummary[] }>;
472
+ check(options: ConnectedSessionOptions): Promise<void>;
473
+ };
474
+ readonly sessions: {
475
+ create(options: ConnectedSessionOptions): Promise<SessionHandle>;
476
+ list(options?: PageOptions): Promise<Page<SessionSummary>>;
477
+ get(id: string): Promise<SessionHandle>;
478
+ };
479
+ readonly tasks: {
480
+ list(options?: PageOptions & { readonly sessionId?: string }): Promise<Page<TaskSummary>>;
481
+ get(id: string): Promise<TaskHandle>;
482
+ };
483
+ readonly approvals: {
484
+ respond(requestId: string, response: ApprovalResponse): Promise<void>;
485
+ };
486
+ disconnect(): Promise<void>;
487
+ }
488
+
489
+ declare function createRunner(options?: RunnerOptions): Runner;
490
+ type Delivery = "queue" | "steer" | "interrupt";
491
+
492
+ interface Runner {
493
+ createSession(
494
+ definition: AgentDefinition,
495
+ options: { readonly cwd: string },
496
+ ): Promise<SessionHandle>;
497
+ session(id: string): SessionHandle;
498
+ task(id: string): TaskHandle;
499
+ respond(requestId: string, response: ApprovalResponse): Promise<void>;
500
+ close(): Promise<void>;
501
+ }
502
+
503
+ interface SessionHandle {
504
+ readonly id: string;
505
+ prompt(text: string, options?: { readonly delivery?: Delivery }): Promise<PromptReceipt>;
506
+ snapshot(): Promise<SessionSnapshot>;
507
+ close(): Promise<void>;
508
+ }
509
+
510
+ interface PromptReceipt {
511
+ readonly task: TaskHandle;
512
+ readonly requestedDelivery: Delivery;
513
+ readonly effectiveDelivery: Delivery;
514
+ readonly paused?: true;
515
+ readonly blockedByTaskId?: string;
516
+ }
517
+
518
+ interface TaskHandle {
519
+ readonly id: string;
520
+ readonly sessionId: string;
521
+ readonly result: Promise<TaskResult>;
522
+ snapshot(): Promise<TaskSnapshot>;
523
+ inspect(): Promise<TaskDetails>;
524
+ watch(options?: { readonly signal?: AbortSignal }): AsyncIterable<TaskSnapshot>;
525
+ responses(options?: {
526
+ readonly after?: string;
527
+ readonly signal?: AbortSignal;
528
+ }): AsyncIterable<CompleteResponse>;
529
+ cancel(): Promise<TaskSnapshot>;
530
+ }
531
+
532
+ interface TaskSummary {
533
+ readonly taskId: string;
534
+ readonly sessionId: string;
535
+ readonly state: TaskState;
536
+ readonly selectedHarness?: HarnessKind;
537
+ readonly title: string;
538
+ readonly ordinal: number;
539
+ readonly updatedAt: number;
540
+ readonly dispatchedAt?: number;
541
+ readonly failedAt?: number;
542
+ readonly finishedAt?: number;
543
+ }
544
+ interface WorkspaceInfo { readonly directory: string; readonly repository?: string }
545
+ interface SessionSummary {
546
+ readonly sessionId: string;
547
+ readonly state: "open" | "closing" | "closed";
548
+ readonly paused: boolean;
549
+ readonly cwd: string;
550
+ readonly workspace: WorkspaceInfo;
551
+ readonly active?: TaskSummary;
552
+ }
553
+ interface SessionSnapshot extends SessionSummary { readonly queued: readonly TaskSummary[] }
554
+ type TaskState = "queued" | "running" | "waiting" | "awaiting_approval"
555
+ | "completed" | "failed" | "cancelled" | "interrupted";
556
+ type TerminalState = "completed" | "failed" | "cancelled" | "interrupted";
557
+ interface ErrorInfo { readonly code: string; readonly message: string }
558
+ interface CompleteResponse {
559
+ readonly responseId: string;
560
+ readonly text: string;
561
+ readonly state: TaskState;
562
+ }
563
+ interface TaskSnapshot extends TaskSummary {
564
+ readonly revision: number;
565
+ readonly parentTaskId?: string;
566
+ readonly children: readonly TaskSummary[];
567
+ readonly requests: readonly ApprovalRequest[];
568
+ readonly latestResponse?: CompleteResponse;
569
+ readonly error?: ErrorInfo;
570
+ readonly stopUnconfirmed: boolean;
571
+ }
572
+ interface TaskDetails extends TaskSnapshot { readonly prompt: string }
573
+ interface TaskResult {
574
+ readonly taskId: string;
575
+ readonly sessionId: string;
576
+ readonly responseId: string;
577
+ readonly text: string;
578
+ }
579
+ ```
580
+
581
+ `createSession` validates and captures the definition and existing absolute working directory, but starts native execution lazily when work is dispatched. A prompt must contain non-whitespace text and fit the existing 1 MiB limit. `prompt` returns after admission or acceptance of steering, not after native startup or completion. Invalid admission rejects without creating a task. A native startup failure after admission is a failed task.
582
+
583
+ `task.result` resolves only when the parent task reaches `completed`, after children, child-result delivery, and parent continuations finish. Its text comes from the final parent response, not the first waiting response, a child response, or concatenated history. Empty text is valid. Completion reports execution completion, not proof that the requested objective succeeded: an answer explaining a failing test can still be complete.
584
+
585
+ Pending approvals and children keep the promise pending. A child failure that the parent handles does not automatically reject the parent's result. Execution failures reject with a safe `AgentError`; explicit cancellation rejects with `CANCELLED`, and interruption with `INTERRUPTED`. Detailed outcomes remain in snapshots, including `error` and `stopUnconfirmed`. First access to `result` starts observation and memoizes its promise. For a connected task, the coordinator atomically reads current state and registers the observer when it handles that request; the interval does not begin at the earlier JavaScript property-access time. An interval is the current execution through its next terminal state, or the terminal outcome already present at registration. It captures that first terminal outcome even if private recovery occurs before awaiting code resumes; a failure already superseded before registration is not observed retroactively. Embedded observation uses the equivalent atomic in-process boundary. Repeated awaits reuse that promise; a connected transport rejection is not the retained server task outcome. Reconnect and obtain a fresh handle to observe again after connection loss. Settled results and ended iterators never reopen. A previously settled result promise stays settled after client disconnect; first result access on a disconnected handle rejects with `COORDINATOR_UNAVAILABLE`. Snapshot/inspection reads always show current state; a fresh watch or response iterator starts a new observation, while retained response replay remains cursor-based. An unused result promise must not produce an unhandled-rejection side effect; explicitly awaiting it still rejects on failure.
586
+
587
+ `watch` immediately yields a snapshot, then coalesces changes for slow readers and ends after a terminal snapshot. Snapshots expose `state`, `latestResponse`, direct `children`, and pending `requests` throughout the descendant tree. They preserve existing task states and ownership. `responses` yields every retained complete response in order and then waits for more until termination. Both are optional observers; neither drives execution or acknowledges managed child-result delivery.
588
+
589
+ ### Unified approval types
590
+
591
+ Applications receive the actions available for each request and submit the selected action's opaque ID. Normalized decisions and scopes describe actions for presentation; they are not response payloads. Adapters translate the selected action into its exact native response, so applications do not branch on `harness`.
592
+
593
+ ```ts
594
+ type ApprovalScope = "once" | "turn" | "session";
595
+ interface ApprovalResponse {
596
+ readonly actionId: string;
597
+ readonly grants?: readonly string[];
598
+ }
599
+
600
+ type ApprovalAction = {
601
+ readonly id: string;
602
+ readonly label: string;
603
+ readonly description?: string;
604
+ } & (
605
+ | { readonly decision: "allow" | "deny"; readonly scope: ApprovalScope }
606
+ | { readonly decision: "cancel"; readonly scope?: never }
607
+ );
608
+
609
+ interface ApprovalGrant {
610
+ readonly id: string;
611
+ readonly label: string;
612
+ readonly description?: string;
613
+ }
614
+
615
+ interface ApprovalRequest {
616
+ readonly id: string;
617
+ readonly taskId: string;
618
+ readonly sessionId: string;
619
+ readonly harness: HarnessKind;
620
+ readonly message: string;
621
+ readonly actions: readonly ApprovalAction[];
622
+ readonly grants?: readonly ApprovalGrant[];
623
+ readonly context?: Readonly<Record<string, unknown>>;
624
+ }
625
+ ```
626
+
627
+ `id` is the SDK name for the existing request identifier. An action ID is opaque, stable and unique within its request, including across clients observing that request. The application submits it as `actionId`; it is not a native decision string. Distinct native choices retain distinct IDs even when their normalized decision and scope match. Labels and optional descriptions explain the actual operation or permission coverage and reuse duration. A description is required when the label alone cannot communicate a material difference. Optional `context` contains bounded display information about the native operation, not instructions or authority.
628
+
629
+ | Action metadata | Meaning |
630
+ | --- | --- |
631
+ | Allow once | Permit the requested operation once. |
632
+ | Allow for this turn | Permit the offered grants for the current native turn, which can be shorter than the coordinated task. |
633
+ | Allow for this session | Reuse the offered permission within the originating native conversation, not globally, permanently, or across child sessions. |
634
+ | Deny once | Reject this operation using native denial semantics; the agent may continue. |
635
+ | Deny for this turn/session | Reuse the offered denial for the stated native turn/conversation when supported. For fx, an offered `reject_always` is session-scoped denial. It is not silently reduced to denial once. |
636
+ | Cancel | Select a native cancellation option, only when explicitly offered; its effect follows that operation's native semantics. This is distinct from `task.cancel()`. |
637
+
638
+ Only an action ID belonging to the original request is accepted. The coordinator resolves its native mapping from the retained request; caller-supplied labels, decisions, scopes, native payloads, and other extra fields are rejected. An unknown action or invalid grant selection returns `INVALID_APPROVAL_RESPONSE` and leaves a pending request unanswered. A denied operation does not automatically cancel its task. Native policies still determine whether a request can appear; hard sandbox denials cannot be approved.
639
+
640
+ If the request's `grants` is absent, submit only `actionId`; a response containing `grants` is invalid. If the request's `grants` is present, selecting an allow action requires an explicit `grants` array, including `[]` when selecting none. Deny/cancel selections accept no grants. Nothing is preselected or silently granted. Grant IDs are opaque, request-local references to the adapter's immutable permission groups, not caller-defined filesystem paths or profiles. Related restrictions remain grouped so selecting a subset cannot accidentally broaden a permission. Only listed IDs, without duplicates, are valid. Grant order has no semantic meaning: validated selections are canonicalized in the request's grant order before native translation and idempotency comparison.
641
+
642
+ Adapters own exact native mappings and expose the supported choices they can faithfully execute. Unrepresentable native requests fail with `INPUT_REQUIRED` rather than receiving fabricated semantics. Reusable denial and distinct native options are preserved through separate action IDs. The v1 boundary retains original tool arguments and excludes persistent settings changes. Permission approval, argument editing, answering a question, and accepting a plan remain distinct interactions; this API covers permission approval only.
643
+
644
+ `runner.respond` and `client.approvals.respond` acknowledge validated answer submission, not operation completion. They preserve [approval ownership and lifecycle rules](approvals.md): withdrawal races cannot cause late allows; identical answers remain idempotent while retained; conflicting answers fail with `INVALID_APPROVAL_RESPONSE`; withdrawn requests return `REQUEST_EXPIRED`, and unknown request IDs return `UNKNOWN_REQUEST`. No automatic answer or timeout grant is introduced, including when a UI disconnects. Existing schema/size limits and retention of 1,024 settled records remain. SDK normalization does not require changing CLI or hosted-agent tool wire formats; those entry points and SDK action selection resolve the same native requests.
645
+
646
+ The normalization covers the existing adapters without extending their native permission capabilities. Codex maps supported command/file choices and turn/session grants; Claude maps allow-once, conditional session reuse, and denial; fx maps its offered ACP choices including session denial. OpenCode maps its existing `once`/`reject`, Copilot its `approve-once`/`reject`, and Cursor subscription ACP its offered once-only choices. Cursor API-key SDK execution retains its documented lack of an interactive permission callback. Constructor support does not imply every access route can emit every action.
647
+
648
+ ### Ownership and transport boundaries
649
+
650
+ Discovery precedence is explicit `ConnectOptions.stateDirectory`, then `AGENT_STATE_DIR`, then `~/.subharness/state/shared`. Explicit paths must be nonempty absolute local paths and are canonicalized. The CLI uses the same environment/default selection rule; installation location and cwd do not create separate default coordinators. Endpoint metadata is owner-private, with a private authentication token and loopback-only transport. Invalid state-directory or endpoint configuration rejects with `INVALID_CONFIG`.
651
+
652
+ The handshake validates the endpoint's coordinator instance ID, SDK protocol version `1`, and the `sdk-v1` capability; the private CLI bridge additionally requires `cli-v1`. Package versions need not match. A missing/incompatible protocol or required capability rejects with `COORDINATOR_INCOMPATIBLE`; authentication rejection uses `COORDINATOR_AUTH_FAILED`. Missing strict-attachment targets and uncertain reachability use `COORDINATOR_UNAVAILABLE`. A live incompatible or unreachable coordinator is never replaced to make connection succeed.
653
+
654
+ `connect()` defaults to start-or-connect; `start: false` never starts. The entire discovery, health check, handshake, and startup path has a 15-second deadline, including attachment without startup. Final confirmation may add at most one second for unpublished-candidate retirement and 250 milliseconds for a final endpoint check. The final check distinguishes a candidate that published at the deadline from one that retired, without replacing uncertain execution. Attach/health/handshake deadline failures use `COORDINATOR_UNAVAILABLE`; startup failure uses `STARTUP_FAILED`. Neither path can wait indefinitely.
655
+
656
+ Exclusive startup ownership serializes contenders for the selected directory. Under that ownership, absence is established only when endpoint metadata is missing, or when a well-formed stale record names an owner PID proven nonexistent and its endpoint is not responsive. A live owner, unknown owner state, permission-denied PID check, connection refusal, or timeout alone is not proof of absence and yields `COORDINATOR_UNAVAILABLE`. Malformed metadata is `INVALID_CONFIG`. A responsive endpoint with authentication or protocol mismatch returns the corresponding explicit error above; instance mismatch returns `COORDINATOR_UNAVAILABLE`. None is treated as absence.
657
+
658
+ A candidate process may be spawned to acquire startup ownership, but it must prove absence under that ownership before starting/publishing a shared service. The winner publishes one authenticated endpoint atomically. Losers connect to that instance and retire only their own unpublished candidate. Failure to confirm candidate retirement is reported without killing or replacing a process that might have published. Starting a coordinator creates no agents, and the coordinator survives client disconnect and application exit.
659
+
660
+ Embedded `runner.session(id)` and `runner.task(id)` remain synchronous local lookups. Connected `get` methods verify existence asynchronously. All snapshot and inspection methods return promises. Unknown IDs fail rather than constructing records. Connected handles remain bound to the validated coordinator instance and never silently switch after replacement.
661
+
662
+ `client.disconnect()` releases client resources and invalidates its handles without stopping the coordinator or changing tasks or pending approvals. Transport failure rejects affected reads/observations with safe errors such as `COORDINATOR_UNAVAILABLE`; explicit observer signal cancellation uses `OBSERVATION_CLOSED`. Neither implies task cancellation. `task.cancel()` retains confirmed-stop semantics in both modes, including descendant handling and `CANCEL_FAILED` on uncertain stop. `session.close()` explicitly closes execution for that session; `runner.close()` closes all locally owned sessions. Successful closes are idempotent; failed embedded cleanup retains ownership for retry. Application tool callbacks are drained, not force-stopped.
663
+
664
+ There are no automatic mutation retries or public idempotency keys in this version. Losing a creation/prompt/control acknowledgement leaves its outcome uncertain; it is not proof of nonadmission. The application reconnects and inspects authoritative state before deciding what to do, without claiming inspection can always identify the lost operation. Reads and observers can be reissued explicitly against the same instance; automatic transport reconnection is not promised. Approval resubmission retains its separately documented identical-answer idempotency.
665
+
666
+ ## One library for SDK and CLI
667
+
668
+ The CLI consumes the same library connection, discovery, admission, observation, and control core. Its argument parsing, prompt-file/stdin reading, text/JSONL rendering, truncation, and dashboard grouping remain presentation concerns. It does not maintain a parallel direct-HTTP client. A private bridge preserves CLI semantics that are intentionally absent from ordinary application handles.
669
+
670
+ | Private operation | Preserved behavior |
671
+ | --- | --- |
672
+ | Catalog and readiness | Ordinary top-level CLI `list`/`check` use the same library helpers behind `client.agents` without requiring or starting a coordinator. Managed calls carrying `AGENT_PARENT_TOKEN` must first attach to the existing coordinator, validate the active parent, and narrow the child capability before discovery/check. Validation failure never falls back to external-owner authority. Definition loading and disposable readiness startup keep their existing side effects. |
673
+ | Atomic `run` | Resolve/capture the target, environment, and first prompt before publishing one session and its first task together. Invalid input publishes neither; native failure after admission becomes task failure. It is not two independently visible public calls. |
674
+ | Attached run/send and `wait` | Use a one-observation helper returning the existing response, approval, or terminal record. Preserve response-cursor precedence, detach boundaries, and steering acknowledgement. Do not replace first-response waiting with final `task.result`. |
675
+ | Managed delegation | Bind the active parent capability out-of-band. Preserve declared-child discovery, lineage, maximum depth, self/ancestor restrictions, and strict-descendant authority. The public client neither accepts raw parent tokens nor infers them from environment variables. |
676
+ | Managed wait acknowledgement | Only the private active-parent observation path acknowledges delivered terminal child results under existing rules. Public result/watch/replay never perform that coordination effect. |
677
+ | CLI `resume` | Retain the private operation and receipt/observation compatibility with a fresh observation scope. Codex, Claude Code, and fx all reject prompt-free native recovery with `RECOVERY_UNSUPPORTED`; no public resume method or successful recovery guarantee is added. |
678
+ | Status and queue | Preserve the existing status response projection: default response text is sliced to 4,000 characters with its truncation flag, while `--full` returns the complete response. Queue descriptions remain the raw admission prompt's first 120 characters (`prompt.slice(0, 120)`), with existing active/queued records and paused state. Normalized public titles are not a substitute for this CLI output. |
679
+ | Dashboard inspection | Derive summaries and expanded detail from the common records; keep terminal-specific grouping, truncation, and efficient private payload formatting outside the public model. |
680
+
681
+ CLI connection policy stays command-specific: `run` can start the coordinator; follow-up, observation, and control commands attach only. An absent coordinator leaves the dashboard in its existing waiting view without auto-start. These policies use the shared connection implementation, not a second transport.
682
+
683
+ Private recovery does not reset a settled promise or reopen an ended iterator. Its fresh observation scope is registered at the accepted recovery boundary, so it observes the current retained task from that boundary without a later registration race. A fresh public `get` yields a fresh handle; the old handle's snapshot remains a current-state read, but its previously accessed result stays settled. There are no public attempt IDs, automatic prompt replay, or restart recovery. Cancelling retained failed work releases a paused queue only after stop is confirmed.
684
+
685
+ Connected direct harness strings preserve native defaults. Existing public harness constructors retain their required model field. The private CLI bridge keeps its existing direct `HarnessRequest` support for flags such as effort without a model; that CLI compatibility does not require adding another public configuration type. Public connected calls are external-owner operations under local coordinator authentication; applications still enforce their own user/chat authorization.
686
+
687
+ ## SDK and coordinator compatibility
688
+
689
+ CLI syntax and documented output, waiting, and control behavior remain supported through the private bridge.
690
+
691
+ The new shared directory does not scan, migrate, or terminate older installation-specific coordinators. Explicit `stateDirectory` or `AGENT_STATE_DIR` can select an old directory, but connection still requires the new handshake. Incompatible old work remains owned by its old coordinator and can be inspected or finished with the matching old CLI. No existing task IDs are imported into a new instance, and a new empty coordinator is never presented as restored execution.
692
+
693
+ ## Embedded configuration and cleanup
694
+
695
+ Construction validates options synchronously and starts no native processes or listeners. Unknown option fields are invalid. An omitted `env` copies `process.env` at construction. A supplied `env` replaces, rather than merges with, that environment; values must be strings or undefined. Inherited Subharness launcher/parent/state variables are removed from native SDK sessions so they cannot attach to an unrelated CLI parent. The SDK never assigns to `process.env` or changes the host's working directory.
696
+
697
+ Native adapter dependencies are loaded only when selected. Claude's SDK operations run in a private per-session worker with a copied environment so dependency initialization cannot change the application's environment. Tool schemas and callbacks remain in the application process with their original identities. The host continues to own native-process termination and callback draining; a worker exit does not establish that a session has stopped. This internal boundary follows the [native adapter contract](adapter-contract.md).
698
+
699
+ Access is copied and validated at construction. Provider, endpoint, and wire-protocol fields retain the per-harness requirements and restrictions in [access configuration](access-config.md); exposing a field in the union does not enable it for every harness. Omission preserves the native access defaults defined in [access configuration](access-config.md); an explicit empty array disables a harness. Embedded runners do not read CLI personal settings. An explicit selection uses the existing [connection rules](access-config.md), including empty-array disablement, ordered pre-submission fallback, variable/file references, and OIDC validation. The SDK does not implicitly read personal settings or modify Git excludes. Relative credential-file/project paths retain their existing resolution against the main checkout or project root. Agent definitions do not contain credentials.
700
+
701
+ SDK access diagnostics identify the runner's explicit access selection or native discovery. They do not tell applications to configure a CLI personal settings file that the runner does not read.
702
+
703
+ Embedded definitions and their child graphs are captured structurally when creating a session; tool functions and schema objects retain their identity. Later changes to caller-owned maps or harness options do not change admitted configuration. Application objects captured by closures remain live. Embedded admission does not perform Git discovery for dashboard metadata.
704
+
705
+ `cancel` stops the target and its recorded descendants, removes affected queued tasks, waits for native execution and admitted callbacks to stop, and preserves unrelated queued work. Cancellation does not roll back file/tool effects. Failed confirmation rejects with `CANCEL_FAILED`, retains `stopUnconfirmed: true`, and pauses dispatch. Repeated cancellation of confirmed terminal work returns its current snapshot.
706
+
707
+ `session.close` immediately stops new admission to that session, cancels its active and queued tasks and their recorded descendant tasks, then closes its native resources. Independently queued work in another session survives. A closed session remains inspectable. `runner.close` immediately stops all admission, closes all its sessions, retires pending approvals, and settles pending observations. Queued/active tasks become cancelled after confirmed stop; previously terminal responses and outcomes remain intact. Startup racing with close must be awaited and cleaned; it cannot publish an unowned process or submit a model turn after close begins.
708
+
709
+ Concurrent closes share one cleanup attempt. A successful close is idempotent. A failed close rejects with an `AggregateError` of safe `AgentError` entries, retains ownership of uncertain resources, and permits a subsequent close to retry cleanup; admission stays closed. No cleanup failure is silently swallowed. Read-only snapshots and retained response observations remain available after close. Unknown native or callback details are not copied into public failure messages.
710
+
711
+ Host custom tools keep the existing `execute(input)` signature. They run in the application's unchanged working directory and environment; use explicit paths or captured services. They are trusted application code and are not restricted by a native harness sandbox. Callback admission is limited to the active native turn. Cancellation first retires managed waits, then stops native work and drains admitted callbacks. An arbitrary JavaScript callback cannot be force-stopped; uncooperative tools can delay cancellation or close indefinitely. There is no implicit timeout or new callback cancellation-context argument.
712
+
713
+ The SDK exports `AgentError` and `ErrorInfo`. Input/lookup/control errors reject calls; execution failure after admission is task data. Existing error codes remain in use. `RUNNER_CLOSED` and `SESSION_CLOSED` identify closed-owner admission/control attempts; reads and close retries remain available. `cancel` on terminal work is still an inspection even after close. Unknown option/input shapes reject with `INVALID_ARGUMENT`, invalid definitions with `INVALID_DEFINITION`, and invalid access configuration with `INVALID_CONFIG`.
714
+
715
+ ## Contract boundaries
716
+
717
+ Connected and embedded execution share the session, task, and approval contracts above. Native adapters retain their documented capability limits; the SDK does not add a model loop or native permission system.
718
+
719
+ No public native recovery, collection watch, automatic mutation retry, public idempotency keys, arbitrary remote endpoints, closure transfer, or restart persistence is promised. Polling, per-task observation, explicit connection failure, and breaking-release migration are the defined behavior. Original tool input and the exclusion of persistent settings changes remain approval limits; applications own the UI layout.
720
+
721
+ ## Hosted delegation
722
+
723
+ Inline `subagents` declarations are supported. The embedded SDK prepares native custom tools for declared children in the application process, without creating CLI launchers or a coordinator HTTP service. The prepared native definition contains those tools instead of a native-facing `subagents` descriptor; this explicitly avoids CLI shell permission grants. The retained original graph authorizes child selection. Agents without declared children receive no hosted delegation tools. Existing CLI delegation is unchanged.
724
+
725
+ The following reserved tool names are added only to SDK sessions that declare children. A collision with a custom tool name anywhere in an admitted graph rejects before execution. Tool input schemas reject unknown fields. Optional `cwd` defaults to the calling session directory and must be an existing absolute directory.
726
+
727
+ The `agent_run` tool description lists each direct child's map key, display name, and description so the native parent can choose a specialist without prior knowledge of the graph. Its `agent` field enumerates those map keys. Child instructions, tools, and credentials are not included in that catalog.
728
+
729
+ Failed hosted calls expose the safe library error code in the native tool failure, allowing the parent to distinguish invalid lineage, unavailable identifiers, queue limits, and expired approvals. This applies only to SDK-owned delegation errors. Exceptions from application-defined tools retain the generic custom-tool failure behavior and do not expose arbitrary exception details.
730
+
731
+ | Tool | Input | Output |
732
+ | --- | --- | --- |
733
+ | `agent_run` | `{ agent: string, prompt: string, cwd?: string }` | Admission `{ taskId, sessionId, state }`; `agent` is a direct subagent map key. |
734
+ | `agent_send` | `{ sessionId: string, prompt: string, delivery?: "queue" \| "steer" \| "interrupt" }` | `{ taskId, sessionId, requestedDelivery, effectiveDelivery }`. |
735
+ | `agent_wait` | `{ taskId: string, after?: string }` | Private response/approval_required/terminal observation records; native approval schemas are preserved. |
736
+ | `agent_status` | `{ taskId: string }` | Private task snapshot with native approval request schemas. |
737
+ | `agent_queue` | `{ sessionId: string }` | Private session snapshot. |
738
+ | `agent_cancel` | `{ taskId: string }` | The confirmed Private task snapshot with native approval request schemas. |
739
+ | `agent_respond` | `{ requestId: string, content: ApprovalContent }` | `{ requestId, answered: true }` after schema-valid submission. |
740
+
741
+ Private binding identifies the active parent task; the model cannot provide parent identities. Run resolves only direct declarations. Existing identifier-based follow-up/observation behavior remains in effect; these tools are not an application-user authorization boundary. Self/ancestor-session follow-ups remain forbidden. `agent_wait` and `agent_cancel` reject the caller's active task or an ancestor with `INVALID_PARENT`: waiting for or stopping that task would wait for the executing callback itself. `agent_status` remains available for immediate inspection. `agent_respond` can answer only strict-descendant requests, preserving managed-parent authority. Native permissions and tool-deny policies remain authoritative.
742
+
743
+ Child follow-ups belong to the invoking current parent task. Depth is limited to eight and session queue capacity to 100. A terminal child result returned by `agent_wait` through its active parent's managed tool interaction is acknowledged once. Snapshot inspection, including `agent_status`, and UI observation do not acknowledge delivery. Results arriving after a pending parent response continue that same task and native conversation. Parent failure/cancellation retires managed waits and propagates to descendants. Stale tools cannot admit work into a later parent task: their invoking task identity is fixed before asynchronous work and revalidated before mutation. Arbitrary public `runner.createSession` and `session.prompt` calls from an app tool are independent top-level work; they do not imply lineage.