subharness 0.0.3 → 0.0.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -3
- package/dist/adapters/approval-context.d.ts +2 -0
- package/dist/adapters/approval-context.js +9 -0
- package/dist/adapters/approval-context.js.map +1 -0
- package/dist/adapters/approval-lifetime.d.ts +8 -0
- package/dist/adapters/approval-lifetime.js +50 -0
- package/dist/adapters/approval-lifetime.js.map +1 -0
- package/dist/adapters/claude-approvals.d.ts +4 -0
- package/dist/adapters/claude-approvals.js +32 -0
- package/dist/adapters/claude-approvals.js.map +1 -0
- package/dist/adapters/claude-permissions.js +1 -1
- package/dist/adapters/claude-permissions.js.map +1 -1
- package/dist/adapters/claude.js +30 -1
- package/dist/adapters/claude.js.map +1 -1
- package/dist/adapters/codex-approvals.d.ts +6 -0
- package/dist/adapters/codex-approvals.js +111 -0
- package/dist/adapters/codex-approvals.js.map +1 -0
- package/dist/adapters/codex.js +32 -2
- package/dist/adapters/codex.js.map +1 -1
- package/dist/adapters/fx-approvals.d.ts +2 -0
- package/dist/adapters/fx-approvals.js +14 -0
- package/dist/adapters/fx-approvals.js.map +1 -0
- package/dist/adapters/fx.js +49 -18
- package/dist/adapters/fx.js.map +1 -1
- package/dist/adapters/types.d.ts +2 -0
- package/dist/adapters/types.js.map +1 -1
- package/dist/approvals/schema.d.ts +4 -0
- package/dist/approvals/schema.js +70 -0
- package/dist/approvals/schema.js.map +1 -0
- package/dist/approvals/types.d.ts +47 -0
- package/dist/approvals/types.js +2 -0
- package/dist/approvals/types.js.map +1 -0
- package/dist/cli/args.d.ts +2 -1
- package/dist/cli/args.js +20 -2
- package/dist/cli/args.js.map +1 -1
- package/dist/cli/dashboard-client.d.ts +5 -1
- package/dist/cli/dashboard-client.js +161 -14
- package/dist/cli/dashboard-client.js.map +1 -1
- package/dist/cli/dashboard-detail-view.d.ts +16 -0
- package/dist/cli/dashboard-detail-view.js +163 -0
- package/dist/cli/dashboard-detail-view.js.map +1 -0
- package/dist/cli/dashboard-history-view.d.ts +29 -0
- package/dist/cli/dashboard-history-view.js +211 -0
- package/dist/cli/dashboard-history-view.js.map +1 -0
- package/dist/cli/dashboard-input.d.ts +14 -0
- package/dist/cli/dashboard-input.js +118 -0
- package/dist/cli/dashboard-input.js.map +1 -0
- package/dist/cli/dashboard-layout.d.ts +12 -1
- package/dist/cli/dashboard-layout.js +29 -27
- package/dist/cli/dashboard-layout.js.map +1 -1
- package/dist/cli/dashboard-renderer.d.ts +30 -2
- package/dist/cli/dashboard-renderer.js +125 -7
- package/dist/cli/dashboard-renderer.js.map +1 -1
- package/dist/cli/dashboard-style.d.ts +9 -0
- package/dist/cli/dashboard-style.js +36 -0
- package/dist/cli/dashboard-style.js.map +1 -0
- package/dist/cli/dashboard.d.ts +3 -1
- package/dist/cli/dashboard.js +127 -12
- package/dist/cli/dashboard.js.map +1 -1
- package/dist/cli/help.js +7 -0
- package/dist/cli/help.js.map +1 -1
- package/dist/cli/input.d.ts +1 -0
- package/dist/cli/input.js +23 -17
- package/dist/cli/input.js.map +1 -1
- package/dist/cli/main.js +1 -1
- package/dist/cli/main.js.map +1 -1
- package/dist/cli/output.js +17 -1
- package/dist/cli/output.js.map +1 -1
- package/dist/runtime/approval-registry.d.ts +27 -0
- package/dist/runtime/approval-registry.js +126 -0
- package/dist/runtime/approval-registry.js.map +1 -0
- package/dist/runtime/coordinator.d.ts +17 -1
- package/dist/runtime/coordinator.js +124 -6
- package/dist/runtime/coordinator.js.map +1 -1
- package/dist/runtime/dashboard.d.ts +42 -0
- package/dist/runtime/dashboard.js.map +1 -1
- package/dist/runtime/select-native.d.ts +2 -1
- package/dist/runtime/select-native.js +2 -2
- package/dist/runtime/select-native.js.map +1 -1
- package/dist/runtime/service.js +24 -1
- package/dist/runtime/service.js.map +1 -1
- package/dist/runtime/types.d.ts +9 -1
- package/dist/runtime/types.js.map +1 -1
- package/dist/runtime/worker-approvals.d.ts +21 -0
- package/dist/runtime/worker-approvals.js +51 -0
- package/dist/runtime/worker-approvals.js.map +1 -0
- package/dist/runtime/worker-client.js +35 -7
- package/dist/runtime/worker-client.js.map +1 -1
- package/dist/runtime/worker-server.d.ts +2 -0
- package/dist/runtime/worker-server.js +39 -10
- package/dist/runtime/worker-server.js.map +1 -1
- package/package.json +1 -1
- package/sdk/access-config.md +28 -0
- package/sdk/adapter-contract.md +2 -10
- package/sdk/agent-skill.md +6 -3
- package/sdk/approvals.md +118 -0
- package/sdk/cli/dashboard-design.md +38 -4
- package/sdk/cli/dashboard-inspection.md +25 -0
- package/sdk/cli/index.md +11 -5
- package/sdk/cli/output.md +7 -3
- package/sdk/completion-notifications.md +2 -0
- package/sdk/distribution.md +52 -0
- package/sdk/evals.md +9 -3
- package/sdk/fx.md +1 -1
- package/sdk/index.md +1 -0
- package/sdk/message-delivery.md +1 -1
- package/sdk/permissions.md +4 -4
- package/sdk/pr-integration.md +1 -1
- package/sdk/sessions.md +4 -1
- package/sdk/v1-runtime.md +3 -3
package/sdk/distribution.md
CHANGED
|
@@ -40,6 +40,58 @@ Authentication uses npm trusted publishing through GitHub Actions OIDC, without
|
|
|
40
40
|
|
|
41
41
|
To release, choose a new stable tag on `main` that includes this workflow, enter the release notes, leave the GitHub pre-release option unchecked, and publish the GitHub release. The corresponding Actions run reports whether npm publication succeeded. No separate version-bump commit or release-candidate channel is required.
|
|
42
42
|
|
|
43
|
+
### Maintainer commands
|
|
44
|
+
|
|
45
|
+
An agent asked to publish to npm uses this GitHub release path. It needs authenticated `gh` access with permission to create releases in `vercel-labs/subharness`; local npm login and npm two-factor prompts are not part of normal automated publishing. Intended changes must already be merged into `main` with their checks passing. Uncommitted work and commits on other branches are not included.
|
|
46
|
+
|
|
47
|
+
Inspect live state before selecting the version and source commit:
|
|
48
|
+
|
|
49
|
+
```sh
|
|
50
|
+
git fetch origin main --tags
|
|
51
|
+
gh release list --repo vercel-labs/subharness
|
|
52
|
+
gh run list --repo vercel-labs/subharness --workflow release.yml --limit 10
|
|
53
|
+
npm view subharness versions dist-tags --json --registry=https://registry.npmjs.org/
|
|
54
|
+
git log -5 --oneline origin/main
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Use the requested stable version, or settle the version with the maintainer if the request does not specify one. Do not infer it from `package.json`. Verify that it is greater than the latest published version and has no existing npm version, GitHub release, or remote tag. Wait for any in-progress release before creating another. Inspect remote tags with `git ls-remote --tags origin`. Resolve a tag conflict before continuing: `--target` does not move an existing tag.
|
|
58
|
+
|
|
59
|
+
After selecting the version and reviewing release notes, create the release at the exact verified `main` commit. In this example, replace `X.Y.Z` with the selected version and prepare `.context/release-notes.md` with the release notes:
|
|
60
|
+
|
|
61
|
+
```sh
|
|
62
|
+
release_version='X.Y.Z'
|
|
63
|
+
release_commit=$(git rev-parse origin/main)
|
|
64
|
+
gh release create "v$release_version" \
|
|
65
|
+
--repo vercel-labs/subharness \
|
|
66
|
+
--target "$release_commit" \
|
|
67
|
+
--title "v$release_version" \
|
|
68
|
+
--notes-file .context/release-notes.md \
|
|
69
|
+
--latest
|
|
70
|
+
gh run list --repo vercel-labs/subharness --workflow release.yml \
|
|
71
|
+
--event release --commit "$release_commit" \
|
|
72
|
+
--json databaseId,headBranch,headSha,status,conclusion,url
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
Select the run for the new release tag and commit, then run `gh run watch RUN_ID --repo vercel-labs/subharness --exit-status`, replacing `RUN_ID` with its database ID. Inspect the run with `gh run view RUN_ID --repo vercel-labs/subharness --log` if necessary. The job must succeed through **Publish package**, including all behavioral tests and clean-consumer checks. The workflow limits tests to two concurrent files to avoid exhausting subprocess test deadlines on hosted runners; preserve that limit when updating release checks.
|
|
76
|
+
|
|
77
|
+
`gh run watch` does not support fine-grained personal access tokens. If it is unavailable for the current authentication, poll `gh run view RUN_ID --repo vercel-labs/subharness --json status,conclusion,url` until `status` is `completed` and require `conclusion` to be `success`. Authentication must also allow reading Actions runs.
|
|
78
|
+
|
|
79
|
+
Confirm public availability after the job succeeds:
|
|
80
|
+
|
|
81
|
+
```sh
|
|
82
|
+
npm view "subharness@$release_version" version dist.shasum dist.integrity \
|
|
83
|
+
--json --prefer-online --registry=https://registry.npmjs.org/
|
|
84
|
+
npm view subharness dist-tags --json --prefer-online --registry=https://registry.npmjs.org/
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
The requested version must exist, `latest` must point to it, and its `dist.shasum` must match the tarball checksum in the publish log. npm can accept a publish while still processing the package; metadata and tarball downloads can become available at different times. Wait for propagation and verify that the public tarball downloads successfully before reporting completion. Report the GitHub release, successful Actions run, and npm version links.
|
|
88
|
+
|
|
89
|
+
### Failed releases and retries
|
|
90
|
+
|
|
91
|
+
Inspect the failed step and registry state before retrying. A failure before **Publish package** does not publish anything. For a transient failure on an unchanged commit, rerun the failed job only after confirming that publication did not succeed or remain in processing. A rerun uses the original release commit; it does not pick up a later fix on `main`.
|
|
92
|
+
|
|
93
|
+
When a correction requires a new commit, merge and verify it before publishing another release. Never silently delete or move an existing release/tag. Replacing a failed release/tag requires explicit maintainer authorization and confirmation that the npm version was never accepted for publication. Otherwise, use a new stable version. An accepted npm version is immutable, including while it is processing: do not republish it or recreate its tag to attempt another upload. A delayed registry response alone is not evidence that publication failed.
|
|
94
|
+
|
|
43
95
|
## License
|
|
44
96
|
|
|
45
97
|
The current source checkout and packages built from it are distributed under the Apache License, Version 2.0. The already-published `subharness@0.0.1` registry release remains under the MIT license. The root [LICENSE](../LICENSE) contains the complete Apache License text and is included in packages built from this source.
|
package/sdk/evals.md
CHANGED
|
@@ -80,7 +80,7 @@ interface EvalManifest {
|
|
|
80
80
|
effort?: string;
|
|
81
81
|
// Optional permission fields from the matching native harness constructor.
|
|
82
82
|
}>;
|
|
83
|
-
cases?: Array<"auth-smoke" | "concurrent-progress" | "cli-skill" | "cli-background">;
|
|
83
|
+
cases?: Array<"auth-smoke" | "concurrent-progress" | "cli-skill" | "cli-background" | "cli-approval">;
|
|
84
84
|
limits?: {
|
|
85
85
|
startupMs?: number;
|
|
86
86
|
caseMs?: number;
|
|
@@ -94,11 +94,11 @@ There are one to three targets with unique IDs matching `[a-z][a-z0-9-]{0,47}`.
|
|
|
94
94
|
|
|
95
95
|
`access.cwd` is an absolute existing project directory used only to resolve access, never as the model's execution directory. Expected IDs are nonempty strings. Planning may validate directory/file metadata but never reads a credential source, starts a native executable (including `--version`), evaluates repository agent modules, or makes a network request.
|
|
96
96
|
|
|
97
|
-
Cases default to `["auth-smoke"]`. An explicitly supplied list is nonempty, unique, contains only the
|
|
97
|
+
Cases default to `["auth-smoke"]`. An explicitly supplied list is nonempty, unique, contains only the five supported case IDs, and includes `auth-smoke`. Cells are the full target-by-case product, executed in target order and then the fixed order auth-smoke, concurrent-progress, cli-skill, cli-background, cli-approval. Cell IDs are `<target-id>/<case-id>`. Every cell has a fresh caller session and one attempt. Cases after an unsuccessful auth-smoke for that target are `not_run` with reason `auth-prerequisite`; no extra smoke is silently submitted.
|
|
98
98
|
|
|
99
99
|
Limits are positive safe integers. Defaults and accepted ranges in milliseconds are: startupMs 30000 (1000–60000), caseMs 120000 (10000–300000), cleanupMs 10000 (5000–30000), and runMs 900000 (30000–1800000). The run budget must fit at least one startup + case + cleanup allowance. Reserve that full allowance before admitting each cell. The case deadline starts after native startup, separately from the startup deadline. Remaining cells become `not_run/run-budget` when insufficient budget remains.
|
|
100
100
|
|
|
101
|
-
Fixed limits are one caller at a time, one requested caller turn per cell, one deterministic fake child at most, no retries/resume/replay, one invocation of each fixed helper action, helper lifetime at most 30000 ms, fixture output at most 4096 bytes, safe event evidence at most 65536 bytes per cell, and eight evaluated CLI invocations at most. A bounded final result is reserved separately from the event allowance. Overflow stops the affected cell and cannot erase its final classification. Up to
|
|
101
|
+
Fixed limits are one caller at a time, one requested caller turn per cell, one deterministic fake child at most, no retries/resume/replay, one invocation of each fixed helper action, helper lifetime at most 30000 ms, fixture output at most 4096 bytes, safe event evidence at most 65536 bytes per cell, and eight evaluated CLI invocations at most. A bounded final result is reserved separately from the event allowance. Overflow stops the affected cell and cannot erase its final classification. Up to fifteen requested caller turns is not a guarantee of fifteen provider requests or a hard spending cap: native tool loops and internal auxiliary requests are not fully exposed.
|
|
102
102
|
|
|
103
103
|
Planning exits 0. Invalid input, missing build inputs, and a mismatched live plan digest exit 2 without native startup. A run exits 0 only when every required case assertion passes, otherwise 1. Operator cancellation exits 130 after bounded cleanup. Optional unobservable native dimensions are visible coverage gaps rather than part of a concurrent-progress pass denominator.
|
|
104
104
|
|
|
@@ -169,3 +169,9 @@ Native background acknowledgement, notification and idle reactivation are explic
|
|
|
169
169
|
The event writer accepts only known event kinds, bounded identifiers, booleans, numeric timestamps, exit statuses, validated session/task/response IDs, and known fixture values or verifier digests. It records monotonic local ordering and wall time. It excludes tokens and token hashes, full JWT claims, environments, private endpoint secrets, raw CLI response/error text, raw protocol/native settings/stderr, free-form model transcripts, reasoning, and exception messages/stacks before persistence or console output. CLI responses may be passed transiently to the caller but only validated fields enter reports.
|
|
170
170
|
|
|
171
171
|
Each cell's cost is `{ amountUsd: null, source: "unobserved" }` in live mode and `{ amountUsd: null, source: "synthetic-no-billing" }` in fake mode. No actual total is invented from unknown cells. One smoke is a compatibility observation for that exact configuration, not a statistically supported model ranking.
|
|
172
|
+
|
|
173
|
+
## Structured approval case
|
|
174
|
+
|
|
175
|
+
`cli-approval` measures whether the selected caller harness/model follows the forced Subharness skill to inspect and answer a child's native permission requests. It uses the built CLI, a real local coordinator, and one deterministic child with two sequential requests in the same original turn. Request identifiers and offered choice values are generated per fixture; the caller must read the actual schemas. One harmless fixture operation is authorized once and a second is explicitly outside the caller's assignment and must be denied. Session or persistent grants are not authorized for either operation.
|
|
176
|
+
|
|
177
|
+
Mechanical assertions cover both structured answers, correct request/task correlation, schema-valid choices, exactly one authorized effect and zero denied effects, one child session/turn, no retry/resume/replay or policy changes, terminal observation, bounded CLI use, and complete cleanup. The case may use up to eight built CLI invocations. Its single live caller submission uses the manifest's explicit permissions and strict project OIDC route. Child requests and effects are synthetic: passing establishes caller behavior and CLI/coordinator integration, not that a real native child generated those requests. Each `cli-approval` cell records `childExecution: "synthetic"` independently of the caller's `synthetic` flag. Native adapter emission and continuation are verified separately with deterministic external-protocol tests for Codex, Claude and fx. Reports retain only bounded structured evidence, never raw command output or model transcripts.
|
package/sdk/fx.md
CHANGED
|
@@ -59,7 +59,7 @@ Declared custom tools are exposed through an authenticated MCP HTTP endpoint bou
|
|
|
59
59
|
|
|
60
60
|
Native fx 0.0.9 has a verified crash when receiving image-bearing MCP tool results through this ACP/HTTP route. The adapter forwards valid image blocks, but this native release's live image-tool execution is not supported reliably and can fail with `HARNESS_FAILED`. Adding a text label to the image result does not avoid the crash. Text and JSON tools remain supported. Native fx image attachments use a separate path; subharness's CLI currently accepts text prompts only.
|
|
61
61
|
|
|
62
|
-
Native permissions remain authoritative. fx ACP does not expose scoped allow rules equivalent to the Claude adapter's delegation rules. The optional `permissionMode` selects and verifies the native process mode as defined in [Native Permissions](permissions.md). Omission preserves native settings. The library does not infer approval from a permission request.
|
|
62
|
+
Native permissions remain authoritative. fx ACP does not expose scoped allow rules equivalent to the Claude adapter's delegation rules. The optional `permissionMode` selects and verifies the native process mode as defined in [Native Permissions](permissions.md). Omission preserves native settings. The library does not infer approval from a permission request. Supported ACP permission requests expose the native choices through the [structured approval flow](approvals.md); an explicit answer resumes the same operation. Unsupported interactive input and startup-time requests fail with `INPUT_REQUIRED`. Declared subagents use the session launcher and require native permission to execute it and reach the local coordinator. Declaring tools or children does not override a native approval requirement.
|
|
63
63
|
|
|
64
64
|
## Session lifecycle
|
|
65
65
|
|
package/sdk/index.md
CHANGED
|
@@ -9,6 +9,7 @@ The CLI runs generic Codex, Claude Code, and fx agents without definition files.
|
|
|
9
9
|
- [Discovery](config.md): repository/global definitions and personal worktree settings.
|
|
10
10
|
- [Personal access](access-config.md): subscription discovery, explicit API keys, and project OIDC.
|
|
11
11
|
- [Native permissions](permissions.md): session permission options, native limits, and background delegation.
|
|
12
|
+
- [Permission requests](approvals.md): request-specific schemas, structured answers, and approval lifecycle.
|
|
12
13
|
- [Harnesses](harnesses.md): selection, capabilities, and fallback boundaries.
|
|
13
14
|
- [CLI](cli/index.md): commands and response waiting.
|
|
14
15
|
- [Output](cli/output.md): compact text and typed JSONL records.
|
package/sdk/message-delivery.md
CHANGED
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
|
|
13
13
|
Queue dispatch follows task completion, not the return of a CLI command or every native turn. If A returns a response while a reviewer is running, A remains active and B waits. Once descendants and result processing finish, the parent's complete response releases B. The library uses native completion and task relationships, not text classification.
|
|
14
14
|
|
|
15
|
-
A complete response asking a question finishes a task when no descendant work remains. B can then start. The coordinator decides whether its answer should be queued or applied to the current task with another delivery mode. A native approval/input request is different: it has not completed a turn.
|
|
15
|
+
A complete response asking a question finishes a task when no descendant work remains. B can then start. The coordinator decides whether its answer should be queued or applied to the current task with another delivery mode. A native approval/input request is different: it has not completed a turn. Supported native permission requests keep that original turn pending through the [structured approval flow](approvals.md); queued tasks retain their order until it finishes. Unsupported input fails with `INPUT_REQUIRED`. No request is automatically approved.
|
|
16
16
|
|
|
17
17
|
Execution failure pauses pending work. New queued sends remain admissible while dispatch is paused. Their immediate `started` records include `paused: true` and `blockedByTaskId` naming the failed active task, so the caller can inspect it with `subharness status <failed-task-id> --full` and release the queue with `subharness cancel <failed-task-id>`. These fields capture queue state at admission and are omitted when the queue is not paused. A successful explicit native recovery must finish the failed task before pending tasks proceed. Unsupported recovery leaves work paused. Queued command observers may remain waiting until their task starts or is explicitly cancelled. Cancellation must be confirmed before it releases the queue; an unconfirmed stop keeps dispatch paused.
|
|
18
18
|
|
package/sdk/permissions.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Native Permissions
|
|
2
2
|
|
|
3
|
-
Harness constructors accept optional native permission settings directly alongside model options. These settings configure the native session without modifying persistent native settings or creating a Subharness sandbox.
|
|
3
|
+
Harness constructors accept optional native permission settings directly alongside model options. These settings configure the native session without modifying persistent native settings or creating a Subharness sandbox. Later supported native tool approvals are surfaced to the caller through the [structured approval flow](approvals.md); an explicit valid caller decision is required before authorization. The installed harness evaluates the selected policy, and the external execution environment enforces its remaining restrictions.
|
|
4
4
|
|
|
5
5
|
Omitted settings preserve native configuration. Explicit settings remain fixed for session follow-ups. A child uses its own definition and native configuration; it does not inherit the parent's explicit permission options. Invalid definitions fail before execution. Unsupported or observably rejected explicit settings fail without fallback to another harness or permission policy.
|
|
6
6
|
|
|
@@ -33,7 +33,7 @@ claudeCode({
|
|
|
33
33
|
|
|
34
34
|
`ClaudeCodeOptions` and `ClaudeCodeConfig` expose `permissionMode?: "default" | "acceptEdits" | "bypassPermissions" | "plan" | "dontAsk" | "auto"`, `allowedTools?: readonly string[]`, and `disallowedTools?: readonly string[]`. Rule arrays contain nonempty strings and are copied into immutable configuration. Unknown modes, non-string rules, and fields belonging to another harness are invalid definitions. An empty rule list adds no rules.
|
|
35
35
|
|
|
36
|
-
The adapter passes these settings to the native Agent SDK. Explicit allowed rules are combined with the existing rules for declared custom tools and exact child-launcher operations. Disallowed rules remain authoritative under native permission evaluation. Selecting `bypassPermissions` also supplies the native SDK's required explicit bypass enablement.
|
|
36
|
+
The adapter passes these settings to the native Agent SDK. Explicit allowed rules are combined with the existing rules for declared custom tools and exact child-launcher operations. Disallowed rules remain authoritative under native permission evaluation. Selecting `bypassPermissions` also supplies the native SDK's required explicit bypass enablement. Unresolved callbacks use the structured approval flow; the library never chooses an allow decision automatically.
|
|
37
37
|
|
|
38
38
|
Native permission mode and rule evaluation depend on the installed SDK, executable, and managed settings. Observable rejection or a reported different explicit mode is an error. A native initialization event that reports the active mode must agree with the explicit request. Startup success does not claim that every native rule or future command was verified. Diagnostics identify rejected options without exposing rule contents or tool arguments.
|
|
39
39
|
|
|
@@ -54,9 +54,9 @@ fx({
|
|
|
54
54
|
|
|
55
55
|
## Readiness and failures
|
|
56
56
|
|
|
57
|
-
`check` uses the same permission configuration as execution. Its success establishes startup readiness and only the native settings observable during startup. It does not execute a shell or child launcher, prove arbitrary tool access, or claim unobservable native policy values.
|
|
57
|
+
`check` uses the same permission configuration as execution. Its success establishes startup readiness and only the native settings observable during startup. It does not execute a shell or child launcher, prove arbitrary tool access, or claim unobservable native policy values. Startup-time interactive requests still fail with `INPUT_REQUIRED`. During execution, supported permission requests surface through the [structured approval flow](approvals.md), without changing native configuration.
|
|
58
58
|
|
|
59
|
-
|
|
59
|
+
For a surfaced permission request, the caller inspects the action and responds to its schema, then observes the original task. For an unsupported interaction or hard native denial, the caller reports the affected harness and operation and resolves that restriction before starting a new attempt. Retrying unchanged permission failures, broadening policy automatically, or silently selecting another harness is not a recovery strategy.
|
|
60
60
|
|
|
61
61
|
## Background delegation
|
|
62
62
|
|
package/sdk/pr-integration.md
CHANGED
|
@@ -22,7 +22,7 @@ After every mutation, the integrator reads GitHub state again. Merge success req
|
|
|
22
22
|
|
|
23
23
|
The report identifies merged, already merged, blocked, and unattempted entries, their reviewed heads, resulting merge commits when known, and exact blockers. It includes the observed final destination-branch SHA and command outcomes. The integrator writes a nonsecret progress journal only when the assignment names its path. It never exports credentials, dumps the environment, or claims local integration tests ran merely because hosted checks passed.
|
|
24
24
|
|
|
25
|
-
The caller supplies the existing native GitHub authentication and any explicitly authorized native network profile. Model billing continues to use the caller's selected access configuration. Permission failures are reported without bypasses or silent changes of execution route.
|
|
25
|
+
The caller supplies the existing native GitHub authentication and any explicitly authorized native network profile. Model billing continues to use the caller's selected access configuration. Permission failures are reported without bypasses or silent changes of execution route. Supported native permission requests use the [structured approval flow](approvals.md). A permission response does not expand the integrator's assigned merge authority. Unsupported host interactions can still produce `INPUT_REQUIRED`, as specified in the [adapter contract](adapter-contract.md).
|
|
26
26
|
|
|
27
27
|
The allowlist is an instruction-level delegation contract, not a GitHub token scope or a new runtime security boundary. The agent retains native tools; neither the model's instructions nor an API-domain allowlist mechanically limits a credential to those PRs. Native sandbox policy, credential scopes, and GitHub protections remain the enforcement boundaries.
|
|
28
28
|
|
package/sdk/sessions.md
CHANGED
|
@@ -10,6 +10,7 @@ Each `run` creates a new session; `send` creates follow-up tasks or steers the a
|
|
|
10
10
|
| --- | --- |
|
|
11
11
|
| `queued` | Admitted but not yet dispatched |
|
|
12
12
|
| `running` | Native startup or generation is active |
|
|
13
|
+
| `awaiting_approval` | A native operation is waiting for a structured permission answer |
|
|
13
14
|
| `waiting` | Between native turns while delegated work or result processing remains |
|
|
14
15
|
| `completed` | Native response and descendant completion boundary are satisfied |
|
|
15
16
|
| `failed` | Execution or cancellation failed; queue dispatch is paused |
|
|
@@ -24,8 +25,10 @@ Every complete response receives a response identifier and returns control to it
|
|
|
24
25
|
|
|
25
26
|
## Lifetime and recovery
|
|
26
27
|
|
|
27
|
-
The on-demand coordinator retains task records and native session workers beyond individual command exits. Task IDs resolve from other working directories using that coordinator.
|
|
28
|
+
The on-demand coordinator retains task records and native session workers beyond individual command exits. Task IDs resolve from other working directories using that coordinator. Task records remain available for its lifetime; they are not silently evicted. Settled permission-request records have a separate bounded retention policy defined in [Responding to Native Permission Requests](approvals.md).
|
|
28
29
|
|
|
29
30
|
Restarting the coordinator does not restore queues or automatically replay tasks. Native harness history can exist separately, but this version does not reconstruct library sessions from it. Unavailable identifiers return an error. `resume` is limited to supported native recovery of a retained failed task; it does not imply recovery after coordinator loss.
|
|
30
31
|
|
|
31
32
|
Cancellation, interrupt propagation, and queue behavior are defined in [Message Delivery](message-delivery.md). The detailed ownership and bounds are defined in the [runtime contract](v1-runtime.md).
|
|
33
|
+
|
|
34
|
+
Supported native permission requests follow [Responding to Native Permission Requests](approvals.md). Observing an approval returns control without completing or restarting the task.
|
package/sdk/v1-runtime.md
CHANGED
|
@@ -38,9 +38,9 @@ Closing a response-reading command does not cancel work. The coordinator retains
|
|
|
38
38
|
|
|
39
39
|
## Task states and responses
|
|
40
40
|
|
|
41
|
-
The coordinator supplies the terminal dashboard with a read-only snapshot of current session work followed by retained finished tasks, including failed active tasks that pause a queue. Each task appears at most once, and retained history is available to newly opened dashboards for the coordinator’s lifetime. The private snapshot includes an explicit finished-history and grouped-context capability markers so the CLI can reject older incompatible coordinators. It carries only session/task identity, selected harness, bounded prompt title, task state, canonical checkout/repository paths, and timestamps needed by the display; it excludes environment variables, credentials, definitions, transcripts, and native diagnostics. Reading it does not acknowledge responses or change task execution. Workers report the actual selected harness to the coordinator after native startup succeeds. Task timing records the first dispatch, latest failure, and terminal completion or confirmed-stop time without resetting on native continuations. Finished rows retain a fixed elapsed duration; recovery clears terminal timing for the resumed task. Snapshot observation adds no public SDK execution method or persistent history. The
|
|
41
|
+
The coordinator supplies the terminal dashboard with a read-only snapshot of current session work followed by retained finished tasks, including failed active tasks that pause a queue. Each task appears at most once, and retained history is available to newly opened dashboards for the coordinator’s lifetime. The private snapshot includes an explicit finished-history and grouped-context capability markers so the CLI can reject older incompatible coordinators. It carries only session/task identity, selected harness, bounded prompt title, task state, canonical checkout/repository paths, and timestamps needed by the display; it excludes environment variables, credentials, definitions, transcripts, and native diagnostics. Reading it does not acknowledge responses or change task execution. Workers report the actual selected harness to the coordinator after native startup succeeds. Task timing records the first dispatch, latest failure, and terminal completion or confirmed-stop time without resetting on native continuations. Finished rows retain a fixed elapsed duration; recovery clears terminal timing for the resumed task. Snapshot observation adds no public SDK execution method or persistent history. The inspector uses a separate authenticated, read-only session-history projection for retained work requests, stable admission ordinals, queue state and direct child sessions. Full prompt/response content is fetched only for the expanded request through the task-detail projection. It omits already-cached prompt and response text on refresh and never acknowledges delegated outcomes or changes execution. The private transport is defined in [Dashboard inspection](cli/dashboard-inspection.md); presentation is defined in [CLI](cli/index.md#live-dashboard).
|
|
42
42
|
|
|
43
|
-
Task states are `queued`, `running`, `waiting`, `completed`, `failed`, `cancelled`, and `interrupted`. `waiting` means the library task remains active between native turns while delegated work or child-result processing is pending. A native permission request is not a completed response. Terminal outcomes are `completed`, `failed`, `cancelled`, and `interrupted`.
|
|
43
|
+
Task states are `queued`, `running`, `waiting`, `awaiting_approval`, `completed`, `failed`, `cancelled`, and `interrupted`. `waiting` means the library task remains active between native turns while delegated work or child-result processing is pending. A native permission request is not a completed response. Terminal outcomes are `completed`, `failed`, `cancelled`, and `interrupted`.
|
|
44
44
|
|
|
45
45
|
Each complete native response has an opaque response identifier, text, and the task state at that response. `wait` without `--after` returns the first retained response in order, including one that arrived before the command started, or awaits it. Omission does not select the latest response or begin observation at the current time. `wait --after <response-id>` returns the next retained response after that cursor. If no selected response remains and the task is terminal, it returns the terminal outcome. Unknown response identifiers, including identifiers from another task, are errors. Multiple observers can independently read the same response.
|
|
46
46
|
|
|
@@ -60,7 +60,7 @@ New follow-up tasks in a child session belong to the currently invoking parent t
|
|
|
60
60
|
|
|
61
61
|
## Permissions and bounds
|
|
62
62
|
|
|
63
|
-
Adapters preserve native permission restrictions. Declaring subagents authorizes their invocation: Claude Code receives session-only native allow rules for the exact launcher and documented delegation operations, combined with permissions for declared custom tools. These delegation rules do not themselves change native modes, managed policy, or sandbox boundaries. Explicit constructor options separately select session-native permissions under [Native Permissions](permissions.md); persistent user/project settings are unchanged.
|
|
63
|
+
Adapters preserve native permission restrictions. Declaring subagents authorizes their invocation: Claude Code receives session-only native allow rules for the exact launcher and documented delegation operations, combined with permissions for declared custom tools. These delegation rules do not themselves change native modes, managed policy, or sandbox boundaries. Explicit constructor options separately select session-native permissions under [Native Permissions](permissions.md); persistent user/project settings are unchanged. Supported native approval requests remain open and surface through the [structured approval flow](approvals.md). The caller answers the request-specific schema with `respond`, then observes the original task. The permission callback never automatically approves the request. Unsupported input and startup-time interactions still fail with `INPUT_REQUIRED`. Hard native denials remain subject to the selected native policy.
|
|
64
64
|
|
|
65
65
|
Ordinary text/JSON tool results and complete responses are limited to 1 MiB. Rich tool results have the image and aggregate limits defined in [custom tools](tools.md). Oversized results produce explicit errors rather than silent truncation. `status` includes at most 4,000 characters of the latest response and marks truncation; `status --full` retrieves the complete latest response, while `wait` observes the first response or a subsequent response selected by a cursor. Delegation depth is limited to 8, and each session admits at most 100 pending tasks. Exceeding a bound fails admission without dropping existing work. No automatic task-duration deadline is imposed.
|
|
66
66
|
|