subharness 0.0.1 → 0.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. package/LICENSE +201 -21
  2. package/README.md +51 -18
  3. package/dist/adapters/claude-process.d.ts +10 -0
  4. package/dist/adapters/claude-process.js +58 -0
  5. package/dist/adapters/claude-process.js.map +1 -0
  6. package/dist/adapters/claude.js +52 -10
  7. package/dist/adapters/claude.js.map +1 -1
  8. package/dist/adapters/codex-input.d.ts +2 -0
  9. package/dist/adapters/codex-input.js +14 -0
  10. package/dist/adapters/codex-input.js.map +1 -0
  11. package/dist/adapters/codex-permissions.d.ts +13 -0
  12. package/dist/adapters/codex-permissions.js +30 -0
  13. package/dist/adapters/codex-permissions.js.map +1 -0
  14. package/dist/adapters/codex.js +12 -6
  15. package/dist/adapters/codex.js.map +1 -1
  16. package/dist/adapters/fx-auth.js +3 -0
  17. package/dist/adapters/fx-auth.js.map +1 -1
  18. package/dist/adapters/fx-permissions.d.ts +3 -0
  19. package/dist/adapters/fx-permissions.js +128 -0
  20. package/dist/adapters/fx-permissions.js.map +1 -0
  21. package/dist/adapters/fx.js +4 -0
  22. package/dist/adapters/fx.js.map +1 -1
  23. package/dist/cli/args.d.ts +3 -1
  24. package/dist/cli/args.js +18 -12
  25. package/dist/cli/args.js.map +1 -1
  26. package/dist/cli/dashboard-client.d.ts +3 -0
  27. package/dist/cli/dashboard-client.js +78 -0
  28. package/dist/cli/dashboard-client.js.map +1 -0
  29. package/dist/cli/dashboard-layout.d.ts +21 -0
  30. package/dist/cli/dashboard-layout.js +211 -0
  31. package/dist/cli/dashboard-layout.js.map +1 -0
  32. package/dist/cli/dashboard-renderer.d.ts +30 -0
  33. package/dist/cli/dashboard-renderer.js +80 -0
  34. package/dist/cli/dashboard-renderer.js.map +1 -0
  35. package/dist/cli/dashboard.d.ts +31 -0
  36. package/dist/cli/dashboard.js +81 -0
  37. package/dist/cli/dashboard.js.map +1 -0
  38. package/dist/cli/help.d.ts +1 -1
  39. package/dist/cli/help.js +35 -9
  40. package/dist/cli/help.js.map +1 -1
  41. package/dist/cli/main.js +19 -6
  42. package/dist/cli/main.js.map +1 -1
  43. package/dist/cli/output.js +10 -1
  44. package/dist/cli/output.js.map +1 -1
  45. package/dist/config/access.js +4 -4
  46. package/dist/config/access.js.map +1 -1
  47. package/dist/config/loader.js +2 -2
  48. package/dist/config/loader.js.map +1 -1
  49. package/dist/runtime/client.d.ts +5 -1
  50. package/dist/runtime/client.js +91 -40
  51. package/dist/runtime/client.js.map +1 -1
  52. package/dist/runtime/coordinator.d.ts +7 -1
  53. package/dist/runtime/coordinator.js +72 -8
  54. package/dist/runtime/coordinator.js.map +1 -1
  55. package/dist/runtime/daemon.js +100 -18
  56. package/dist/runtime/daemon.js.map +1 -1
  57. package/dist/runtime/dashboard-workspace.d.ts +2 -0
  58. package/dist/runtime/dashboard-workspace.js +38 -0
  59. package/dist/runtime/dashboard-workspace.js.map +1 -0
  60. package/dist/runtime/dashboard.d.ts +28 -0
  61. package/dist/runtime/dashboard.js +19 -0
  62. package/dist/runtime/dashboard.js.map +1 -0
  63. package/dist/runtime/readiness.d.ts +4 -0
  64. package/dist/runtime/readiness.js +17 -0
  65. package/dist/runtime/readiness.js.map +1 -0
  66. package/dist/runtime/select-native.d.ts +2 -1
  67. package/dist/runtime/select-native.js +5 -3
  68. package/dist/runtime/select-native.js.map +1 -1
  69. package/dist/runtime/service.d.ts +1 -1
  70. package/dist/runtime/service.js +10 -3
  71. package/dist/runtime/service.js.map +1 -1
  72. package/dist/runtime/startup-lock.d.ts +5 -0
  73. package/dist/runtime/startup-lock.js +37 -0
  74. package/dist/runtime/startup-lock.js.map +1 -0
  75. package/dist/runtime/startup-protocol.d.ts +17 -0
  76. package/dist/runtime/startup-protocol.js +23 -0
  77. package/dist/runtime/startup-protocol.js.map +1 -0
  78. package/dist/runtime/state.d.ts +1 -2
  79. package/dist/runtime/state.js +8 -43
  80. package/dist/runtime/state.js.map +1 -1
  81. package/dist/runtime/types.d.ts +16 -2
  82. package/dist/runtime/types.js.map +1 -1
  83. package/dist/runtime/worker-client.d.ts +7 -2
  84. package/dist/runtime/worker-client.js +69 -12
  85. package/dist/runtime/worker-client.js.map +1 -1
  86. package/dist/runtime/worker-server.d.ts +7 -0
  87. package/dist/runtime/worker-server.js +101 -0
  88. package/dist/runtime/worker-server.js.map +1 -0
  89. package/dist/runtime/worker.js +2 -66
  90. package/dist/runtime/worker.js.map +1 -1
  91. package/dist/sdk/definitions.js +23 -13
  92. package/dist/sdk/definitions.js.map +1 -1
  93. package/dist/sdk/permission-validation.d.ts +4 -0
  94. package/dist/sdk/permission-validation.js +55 -0
  95. package/dist/sdk/permission-validation.js.map +1 -0
  96. package/dist/sdk/types.d.ts +14 -0
  97. package/dist/sdk/types.js.map +1 -1
  98. package/package.json +17 -17
  99. package/sdk/access-config.md +4 -2
  100. package/sdk/adapter-contract.md +14 -2
  101. package/sdk/agent-skill.md +38 -0
  102. package/sdk/agent.md +6 -2
  103. package/sdk/cli/dashboard-design.md +40 -0
  104. package/sdk/cli/index.md +47 -13
  105. package/sdk/cli/output.md +21 -3
  106. package/sdk/completion-notifications.md +9 -5
  107. package/sdk/config.md +33 -8
  108. package/sdk/distribution.md +21 -7
  109. package/sdk/evals.md +171 -0
  110. package/sdk/fx.md +4 -3
  111. package/sdk/harnesses.md +6 -0
  112. package/sdk/index.md +2 -0
  113. package/sdk/message-delivery.md +6 -2
  114. package/sdk/permissions.md +65 -0
  115. package/sdk/plugins/sub-agents.md +7 -4
  116. package/sdk/pr-integration.md +37 -0
  117. package/sdk/project-team.md +13 -8
  118. package/sdk/sessions.md +1 -1
  119. package/sdk/tools.md +2 -0
  120. package/sdk/v1-runtime.md +22 -7
package/sdk/evals.md ADDED
@@ -0,0 +1,171 @@
1
+ # Harness and Model Evaluations
2
+
3
+ Evaluations are repository-private development tooling for comparing explicit harness/model combinations. They do not extend the public Subharness SDK or CLI. The initial harnesses are Codex App Server, Claude Code Agent SDK, and fx ACP. A headless result does not establish behavior in an interactive terminal, desktop application, or IDE.
4
+
5
+ Ordinary automated tests use deterministic native-protocol fixtures and never request paid generations. Live evaluation requires a separate explicit invocation and a concrete matrix. Development-agent sessions are separate from the matrix they implement. A successful build or fake-harness test is not a live compatibility result.
6
+
7
+ ## Strict project OIDC
8
+
9
+ Every paid participant must use Vercel AI Gateway with a `vercel-oidc` connection for the independently selected project and organization. Subscription, direct API keys, Gateway API keys, alternative providers, model substitutions, and fallback lists are not allowed in an evaluation. Native auxiliary inference is subject to the same routing requirement; an unverifiable route blocks the corresponding claim. The first evaluation scenarios do not use an LLM judge or paid child agents.
10
+
11
+ The operator supplies the expected project and organization IDs independently of the credential. The runner reuses the repository's existing access parser, main-worktree resolution, environment-file precedence, OIDC claim validation, and native adapter bindings. It never implements a second credential-routing system or edits a native login. An explicit environment file cannot be overridden by an ambient variable.
12
+
13
+ All three harness keys must be explicitly present in the selected personal access configuration. Each participating harness has exactly one `vercel-oidc` entry. A nonparticipating harness can have one OIDC entry or an empty disabled list. Omitted keys and non-OIDC alternatives are rejected before native startup, so missing fixture configuration cannot silently discover a subscription. Existing access loading can maintain its documented local Git exclude entry; the eval tooling does not rewrite personal access settings or credential files.
14
+
15
+ Access is resolved before each case, not cached as valid for an entire matrix. Tokens must remain valid for at least the declared startup allowance, case deadline, and cleanup grace. Insufficient lifetime blocks admission. The provider/operator renews tokens outside the runner; there is no automatic login, token issuance, mid-session refresh, or replay after submission. A new attempt requires a new explicit invocation.
16
+
17
+ Local validation proves only token structure, project/organization claims, and time bounds. Native startup verifies the selected route/model to the extent exposed by the adapter. A paid generation separately establishes remote acceptance. Neither local claim decoding nor fake JWT tests prove cryptographic authentication or provider-side cost attribution.
18
+
19
+ Credentials remain in memory or existing private native configuration owned by the adapters. Tokens, raw JWT claims, raw exception messages, environment dumps, native settings, protocol transcripts, and model reasoning are excluded from persisted reports. Isolated directories are not sandboxes; evaluation fixtures contain only benign synthetic data and do not promise to hide the selected credential from every native tool in its process tree.
20
+
21
+ ## Evaluation validity
22
+
23
+ A case distinguishes task correctness, actual CLI use, concurrent command progress, native background execution, completion notification, and autonomous reactivation after idle. These are separate observations. Shell detachment, Subharness `--detach`, managed-child continuation, and native host notifications are not interchangeable. An unobservable native event is blocked by observability, not evidence that a harness lacks the feature.
24
+
25
+ Every selected matrix cell receives a result: `pass`, `fail`, `blocked`, `unsupported`, or `not_run`. Pass requires mechanical evidence, not the model's final text. Unsupported requires affirmative capability evidence for the exact version and transport. Authentication, unavailable models, rate limits, missing tools, permissions, and uncertain cleanup are not model-quality failures. Unknown causes remain unknown. Unstarted cells remain in totals with a reason.
26
+
27
+ Cases are serial and bounded. The runner owns the native session and any fixture helpers, stops admissions on cancellation, and waits for cleanup rather than merely terminating a CLI observer. No automatic retries, detached unowned jobs, or public shutdown commands are invented. Uncertain cleanup prevents the next case. Provider spend is reported only with evidence; unavailable per-case costs remain unknown. Local deadlines and admission limits are not hard dollar ceilings.
28
+
29
+ ## Internal access boundary
30
+
31
+ `evals/access.ts` exports the repository-private function below. It is not a package export and is not a public SDK API.
32
+
33
+ ```ts
34
+ resolveEvalAccess(options: {
35
+ cwd: string;
36
+ harnesses: readonly HarnessKind[];
37
+ expectedProject: { projectId: string; orgId: string };
38
+ minimumValidityMs: number;
39
+ env?: NodeJS.ProcessEnv;
40
+ now?: () => number;
41
+ }): Promise<EvalAccess | EvalAccessError>
42
+ ```
43
+
44
+ `HarnessKind` uses the existing `codex`, `claudeCode`, and `fx` identifiers. The harness list is nonempty and has no duplicates. IDs are nonempty strings. `minimumValidityMs` is a positive finite safe integer. The default environment is `process.env`, and the default clock is `Date.now` in milliseconds. Input errors are returned before native startup; this function never starts a harness or makes a network request.
45
+
46
+ `EvalAccessError` is a typed returned error, not an exception-based domain result. Its `reason` is one of `invalid-input`, `policy`, `credentials-unavailable`, `validation`, `identity`, or `insufficient-lifetime`. Its message is exactly `Evaluation access is unavailable.` for every reason; it never retains raw caught exceptions, credentials, token claims, or native output. A policy error includes absent keys, a disabled participating harness, non-OIDC routes, or fallback lists. Credential reading failures are credentials-unavailable; malformed/expired/not-yet-valid credentials or linkage validation failures are validation. A valid linked token whose project or organization differs from the independent expectation is identity. Insufficient remaining validity uses insufficient-lifetime. Unclassified access-layer failures are validation, not guessed native authentication causes.
47
+
48
+ A successful `EvalAccess` contains `evidence` with only `accessType: "vercel-oidc"`, the expected `projectId` and `orgId`, the participating `harnesses`, the earliest participating `expiresAt` in epoch milliseconds, and `authenticationStage: "local-claims"`. Its `forHarness(kind)` method returns the existing in-memory `ResolvedAccess` for a participating harness, or a safe `EvalAccessError` with reason `invalid-input` for an unselected or invalid kind. Credential bindings are retained in closures, not enumerable object properties; serializing the result or its evidence must not expose a token. This is protection against accidental artifact serialization, not a security boundary against code holding the binding.
49
+
50
+ The access module uses the existing resolvers rather than reimplementing their path or linkage semantics. It additionally compares the validated token identity against the independent expected IDs and enforces the requested remaining lifetime. The converted expiry in milliseconds must be finite; overflow returns `validation` before a binding is admitted. Deterministic tests use synthetic credentials and fixtures, including a linked worktree, and verify that credentials are never included in returned errors or serialized evidence. A test may create and remove linked worktrees of its own disposable Git repository, never worktrees or branches of the developer's real repository.
51
+
52
+ ## Manifest and commands
53
+
54
+ The private entry point is `evals/main.ts`, also available as the repository script `pnpm run evals ...`. Build the checkout before using it. No public `subharness eval` command is introduced.
55
+
56
+ ```sh
57
+ node --import tsx evals/main.ts plan --manifest /absolute/path/matrix.json
58
+ node --import tsx evals/main.ts run --manifest /absolute/path/matrix.json
59
+ node --import tsx evals/main.ts run --manifest /absolute/path/matrix.json --live --approved-plan <sha256>
60
+ ```
61
+
62
+ The plan records hashes of `pnpm-lock.yaml`, `pnpm-workspace.yaml`, and `package.json` as its package inputs, in that order. These files bind the dependency resolution and workspace configuration used by the source checkout. Missing package inputs fail planning before native startup.
63
+
64
+ Each command accepts optional `--out <new-directory>`. `plan` is static. `run` defaults to deterministic fake execution, not paid inference. Live mode additionally requires the current plan digest; that digest binds reviewed inputs but is not itself human spending authorization. Live evaluations require a separately approved concrete matrix and limits. `--approved-plan` without `--live`, live flags on `plan`, duplicate options, and unknown options are invalid. `--help` describes these commands without reading access settings or starting a native executable.
65
+
66
+ The manifest is JSON with this exact shape; unknown fields are rejected at every level. Defaults apply only to omitted optional fields. An explicit `null` is invalid:
67
+
68
+ ```ts
69
+ interface EvalManifest {
70
+ version: 1;
71
+ access: {
72
+ cwd: string;
73
+ expectedProject: { projectId: string; orgId: string };
74
+ };
75
+ targets: Array<{
76
+ id: string;
77
+ harness: "codex" | "claudeCode" | "fx";
78
+ model: string;
79
+ executable: string;
80
+ effort?: string;
81
+ // Optional permission fields from the matching native harness constructor.
82
+ }>;
83
+ cases?: Array<"auth-smoke" | "concurrent-progress" | "cli-skill" | "cli-background">;
84
+ limits?: {
85
+ startupMs?: number;
86
+ caseMs?: number;
87
+ cleanupMs?: number;
88
+ runMs?: number;
89
+ };
90
+ }
91
+ ```
92
+
93
+ There are one to three targets with unique IDs matching `[a-z][a-z0-9-]{0,47}`. Different explicit models of one harness are allowed. Models and optional effort are nonempty strings; effort follows the existing harness constructors' validation. There are no alternative models, implicit model selection, arbitrary prompts, environment maps, or executable shell payloads in a manifest. Codex and Claude use `fast: false`; fx retains its documented native preference. Targets may explicitly include the matching harness's permission fields from [Native Permissions](permissions.md): Codex approvalPolicy/sandboxMode/networkAccessEnabled, Claude permissionMode/allowedTools/disallowedTools, or fx permissionMode. They use the same definition validation, survive normalization, and are bound into the plan digest. Missing options preserve native configuration; the runner never supplies a broader policy implicitly. Invalid and cross-harness fields are rejected before startup. A target's executable is an absolute existing executable regular file, resolved through symlinks and pinned by path and content hash. A Node executable or checked-in native fixture can be used in an offline test manifest; fake execution does not launch the manifest's real executable.
94
+
95
+ `access.cwd` is an absolute existing project directory used only to resolve access, never as the model's execution directory. Expected IDs are nonempty strings. Planning may validate directory/file metadata but never reads a credential source, starts a native executable (including `--version`), evaluates repository agent modules, or makes a network request.
96
+
97
+ Cases default to `["auth-smoke"]`. An explicitly supplied list is nonempty, unique, contains only the four supported case IDs, and includes `auth-smoke`. Cells are the full target-by-case product, executed in target order and then the fixed order auth-smoke, concurrent-progress, cli-skill, cli-background. Cell IDs are `<target-id>/<case-id>`. Every cell has a fresh caller session and one attempt. Cases after an unsuccessful auth-smoke for that target are `not_run` with reason `auth-prerequisite`; no extra smoke is silently submitted.
98
+
99
+ Limits are positive safe integers. Defaults and accepted ranges in milliseconds are: startupMs 30000 (1000–60000), caseMs 120000 (10000–300000), cleanupMs 10000 (5000–30000), and runMs 900000 (30000–1800000). The run budget must fit at least one startup + case + cleanup allowance. Reserve that full allowance before admitting each cell. The case deadline starts after native startup, separately from the startup deadline. Remaining cells become `not_run/run-budget` when insufficient budget remains.
100
+
101
+ Fixed limits are one caller at a time, one requested caller turn per cell, one deterministic fake child at most, no retries/resume/replay, one invocation of each fixed helper action, helper lifetime at most 30000 ms, fixture output at most 4096 bytes, safe event evidence at most 65536 bytes per cell, and eight evaluated CLI invocations at most. A bounded final result is reserved separately from the event allowance. Overflow stops the affected cell and cannot erase its final classification. Up to twelve requested caller turns is not a guarantee of twelve provider requests or a hard spending cap: native tool loops and internal auxiliary requests are not fully exposed.
102
+
103
+ Planning exits 0. Invalid input, missing build inputs, and a mismatched live plan digest exit 2 without native startup. A run exits 0 only when every required case assertion passes, otherwise 1. Operator cancellation exits 130 after bounded cleanup. Optional unobservable native dimensions are visible coverage gaps rather than part of a concurrent-progress pass denominator.
104
+
105
+ ## Driver and lifecycle
106
+
107
+ The private driver starts ownership synchronously and exposes a `ready` promise, a single-use `turn(prompt)` operation, and idempotent `close(reason)` for complete, timeout, cancel, or failure. Closing permanently stops turn admission, including when readiness and closing settle concurrently. Domain failures are typed safe error values; existing adapter exceptions are caught at that boundary, not persisted. Only recognized adapter error codes enter evidence; an unrecognized code becomes a fixed unknown code rather than being copied from the exception. Internal helper types and module layout may be chosen without changing this operator contract.
108
+
109
+ Before each live native startup, the worker calls `resolveEvalAccess` with the manifest access cwd, the one target harness, the independent expected IDs, its controlled environment, and minimum validity `startupMs + caseMs + cleanupMs`. It obtains `forHarness(target.harness)` and passes that resolved binding directly to the existing matching `openCodex`, `openClaude`, or `openFx` adapter. It does not call `selectNative` again or duplicate the adapters' authentication, model checks, or reasoning loop. Even fx's status subprocess starts only after this gate succeeds.
110
+
111
+ Caller executable lookup is pinned to the manifest executable. A private forwarding launcher execs its resolved absolute path so native sibling resources remain discoverable; the fixture does not relocate the installed executable or copy its runtime dependencies. A fake child executable is never put on the caller's PATH. Native permissions remain authoritative. Only explicitly declared target permission options are applied; the runner never broadens or bypasses policy automatically. Live callers retain the operator's native home/configuration context; an empty temporary home is only for fake execution, not a way to hide a login conflict or discard native restrictions. Resolve the selected credential source first, then remove inherited Subharness parent/launcher/state context and competing provider credentials from the caller environment before the matching adapter injects its selected binding. The fixtures require their native command and local coordination facilities; an observed permission denial blocks a measurement instead of being a model failure. An execution directory outside the checkout avoids inheriting its project instructions accidentally; any unobserved native configuration remains identified as unknown.
112
+
113
+ The supervisor owns a worker and its process group from spawn through retirement on supported POSIX environments. An unsupported ownership environment is blocked, not run unsupervised. Group ownership is not a sandbox against arbitrary escaping descendants. Fixed helper programs do not self-detach, execute arbitrary payloads, or create unbounded descendants. Observed escape or uncertain ownership stops further admissions.
114
+
115
+ Timeout or cancellation interrupts active native work, cancels admitted fake tasks, closes sessions and the case's service, and waits for exit. A pending open stays owned; a late native session is closed rather than abandoned by a Promise.race. Grace expiry permits termination only through still-owned child/group handles, never by executable name or a stale PID. Forced local termination does not prove remote inference stopped or private adapter credential files were cleaned up.
116
+
117
+ Cleanup evidence contains `nativeStop` (confirmed, unconfirmed, or not-started), `localExit` (confirmed or unconfirmed), `endpointClosed` (boolean, true when none was opened), and `cleanup` (complete or quarantined). Unconfirmed cleanup overrides the case outcome with `blocked/cleanup-unconfirmed`, quarantines uncertain resources, and leaves remaining cells `not_run/cleanup-unconfirmed`. The separate `stopCause` field preserves a known supervisor trigger: `startup-timeout`, `case-timeout`, `worker-failed`, `cancelled`, or `cleanup-timeout`. It is `null` when no such trigger was observed, including unstarted cells. Cleanup overrides do not erase an already observed trigger or invent a native failure cause. An abnormal worker exit is `worker-failed`, even if a result arrived first; reported cleanup evidence remains available, but the result cannot count as a pass. Confirmed local process exit remains recorded even when native stop or endpoint cleanup is uncertain. Without a worker result, an endpoint is not assumed closed after startup, and `submissionCount` is `null` if execution was authorized but submission cannot be confirmed; zero is reserved for known unsubmitted work. Only verified run-owned disposable paths are removed; artifacts are retained. Hard external termination cannot guarantee cleanup.
118
+
119
+ ## Cases and independent observations
120
+
121
+ ### auth-smoke
122
+
123
+ The prompt requests exactly `eval-ok:<nonce>`. One native turn must settle under the selected binding, and its trimmed text must equal that value. Generation acceptance and answer format are separate assertions, as is confirmed cleanup. The nonce is generated for the case; reports need only the format-match boolean and safe identity, not the raw model response. A settled generation proves more than startup but does not establish provider-side billed cost. A format mismatch after a valid generation is a failed assertion, not an authentication failure.
124
+
125
+ ### concurrent-progress
126
+
127
+ The caller invokes the fixed `hold` and `progress` actions through the supplied absolute helper executable path. A native shell can change PATH, so helper discovery does not depend on retaining the parent's PATH. The supervisor starts and owns the actual bounded helper processes, not merely marker files supplied by the model. It observes this strict sequence: hold started, progress completed while hold is still alive, release, hold completed. Release is caused by the observed progress event, not a guessed sleep. It also verifies exactly-once fixed transformation output and successful helper exits. Each action is admitted once; duplicate or arbitrary actions are rejected and cannot spawn unbounded work.
128
+
129
+ The helper interface accepts only those fixed actions for the current case, not arbitrary paths or shell commands. Helpers read fixed small synthetic input and produce bounded nonce-bearing output. Their event sequence, process liveness, and exits are observed independently of the caller's final text. Model-created files alone do not establish ordering. The prompt explains that hold cannot complete until progress runs and supplies a single-shell background-command recipe that waits for both helpers. The scenario permits native background tools or shell concurrency, but credits only concurrent command progress. Native background acknowledgement, completion notification, and idle reactivation each remain `blocked/observability` with these initial drivers.
130
+
131
+ ### cli-skill
132
+
133
+ The live caller must start the child with `--detach`, retain its returned identifier, perform the supplied independent fixed transformation while the child is still running, and then collect the child result with `wait`. The progress prompt specifies ascending sort order and a comma separator without spaces; surrounding whitespace is ignored by verification. The supervisor holds the deterministic child behind an observed progress barrier; it releases the child only after validating the caller progress artifact. Assertions distinguish detached admission, progress before child completion, identifier correlation, returned-result acknowledgement, and cleanup. Neither elapsed sleeps nor the caller's final text establish those observations.
134
+
135
+ The caller receives the exact current `skills/subharness/SKILL.md` bytes in its prompt. The report records the skill hash and `forced-skill` delivery; this is not a native skill discovery test. The prompt explicitly supplies the quoted absolute path of the private `subharness` launcher, matching managed-session delegation and avoiding native shell changes to PATH. The caller uses that launcher to start one fake Codex child with model `eval-fixture`, explicit fixture cwd and `--detach`, then uses the returned task ID with `wait` without a cursor. The fake child sorts a fixed small string list and returns a nonce not included in the caller prompt. The caller writes that returned nonce to its bounded acknowledgement artifact. Acknowledgement verification reads at most the 4096-byte fixture-output allowance plus one byte to detect overflow; an oversized or nonregular artifact cannot pass.
136
+
137
+ Required observations are actual caller CLI invocations, detached admission, one fake child turn, the matching session/task/response identities, retrieval of the first retained response, exact child output, and exact caller acknowledgement. Task correctness and CLI compliance are separate assertions. Exit 0 from detach alone proves only admission. Runner-issued reads cannot count as caller CLI actions, and fluent summaries cannot replace the evidence.
138
+
139
+ An eval launcher named `subharness` forwards arguments to the unchanged built CLI and observes only bounded safe fields; it never fabricates CLI results. The supervisor starts and tracks the actual CLI subprocesses on the launcher's behalf, so it can retire them before releasing the startup guard, including after service loss. The launcher cannot independently autostart a coordinator. The caller uses its real selected native harness. Only the fake child worker receives the fake Codex executable and synthetic OIDC binding; missing fake executables fail closed, never fall through to installed binaries. The launcher removes inherited provider credentials and managed-parent/launcher context. It uses the case's private coordinator state rather than the developer's coordinator.
140
+
141
+ The eval worker prestarts the existing Coordinator and local service and publishes its normal authenticated endpoint and instance identity. The unchanged built CLI runs with a private, hashed process preloader that rejects attempts to spawn the production daemon; it does not fabricate CLI results or alter parser, client, service, or coordinator code. Endpoint loss therefore fails the case rather than launching a replacement daemon. This guard does not change native harness permissions and is not a public CLI option. Its driver factory admits only the one exact fake Codex reference, model and fixture cwd, at most one child session and one turn, and rejects alternate/extra admissions before startup. This exercises real CLI parser/client/service/coordinator behavior while deliberately excluding production daemon/worker startup coverage. There are no live child models or further delegation in this scenario.
142
+
143
+ ### cli-background
144
+
145
+ `cli-background` complements the existing `cli-skill` detached baseline. It uses the same complete skill, exact private launcher, one deterministic child, independent transformation, nonce acknowledgement, native permissions, and owned cleanup. The caller starts exactly one ordinary attached `run` using its native host background-command controls, performs the independent transformation while that CLI observer remains active, then collects the hosted command's complete output. It must not use `--detach`, a separate Subharness `wait` or `status`, shell `&`, a wrapper process, or an invented background tool. The prompt gives no harness-specific tool name or parameters; it asks the caller to use controls actually exposed in its session. If none are available, the caller reports that limitation without substituting another route.
146
+
147
+ The coordinator holds the child behind a progress barrier. The progress artifact must be absent at observable CLI admission, immediately before the launcher forwards the real `started` record, and must become valid while the one attached CLI observer is still active. Native child startup is not the admission boundary. The launcher preserves streaming admission output so the caller can observe `started` through its host controls before writing progress; no fixed startup sleep or additional Subharness invocation is required. Only then can the child execute and return its nonce. Verification requires one real CLI invocation with successful exit, correlated started and first-response records from that same invocation, one child turn, progress before child completion, exact child output, and the nonce acknowledgement. Prewritten progress, detached invocation, extra CLI observations, nonzero exit, and cleanup uncertainty cannot pass. Cancellation drains the progress observer and pending invocation before fixture cleanup. The case uses the existing supervisor deadlines, access gate, artifact bounds, and fresh-session rules.
148
+
149
+ A passing case establishes concurrent caller progress and result retrieval with one attached Subharness invocation. The current adapter surface does not expose native background acknowledgement or all host tool calls; those observations, notification and idle reactivation remain explicitly blocked by observability. Prompt compliance alone does not prove which native background tool was used. A failed assertion cannot by itself establish that the harness lacks background support. Comparing this case with `cli-skill` measures Subharness invocation count, completion, and duration for the selected native session and model; it is not a statistically conclusive model ranking or a measurement of total host tool calls. Offline protocol fixtures validate the observer/barrier lifecycle but do not provide live native background evidence.
150
+
151
+ ## Fake execution and verification
152
+
153
+ `pnpm run test:evals` runs the focused deterministic evaluation tests. They also participate in ordinary `pnpm test`; neither command starts paid inference.
154
+
155
+ Fake mode never resolves the real manifest access source or launches installed harness binaries. It creates a disposable non-Git access project with all three explicit keys, synthetic linkage and token, and a clean environment, then uses the same access gate against those synthetic expectations. It replaces only external native protocol/SDK boundaries; the ordinary scenario supervisor, helpers, CLI, coordinator, assertions and artifact writer remain real. All fake authentication evidence and results are labeled synthetic. Fakes cannot count as live compatibility or billing evidence.
156
+
157
+ Tests cover all three native boundaries, strict startup gating, fake-child separation, single submission, barrier ordering, real CLI detach/wait and attached run correlation, safe artifacts, startup hangs, cancellation/service loss, and owned process/endpoint cleanup. They need neither installed harnesses nor real credentials, and never fall back to real executables. The scenario fixtures are cooperative functional tests, not a tamper-proof benchmark against arbitrary same-user code.
158
+
159
+ ## Artifacts and interpretation
160
+
161
+ `--out` selects a new directory and never overwrites an existing one. The default is a fresh `.context/evals/<run-id>/` under this checkout. Fixtures are fresh non-Git temporary directories outside both checkouts' Git ancestry. Reports and ledgers are outside those fixture directories; this alone does not enforce a filesystem sandbox against native environments with broader write access.
162
+
163
+ Every command writes `plan.json`. A normalized plan has `version: 1`, `planHash`, `manifest`, `inputs`, and `cells`. `inputs` records the checkout revision, Node/platform identity, hashes of runner/scenario sources, source adapter/runtime dependencies, package manifests and lockfile, skill bytes, built CLI code, and each target's resolved executable path/hash. Hash only declared code/build inputs, never credential files, environment values, tokens, or arbitrary private directories. The live worker rechecks its pinned executable before startup; a changed identity blocks that cell rather than selecting another binary. The deterministic SHA-256 plan digest binds these normalized inputs and excludes the digest itself and output directory. Every cell includes its ID, target ID, case ID, harness, requested model/effort, transport, and instruction delivery (`none` or `forced-skill`). Planning does not request native versions; unknown versions remain null.
164
+
165
+ Runs additionally write `events.jsonl` and `report.json`. A report has `version: 1`, `planHash`, `execution` (fake or live), start/end times, counts for every classification, and every selected cell in order. A cell records its classification, safe reason/code, required and optional assertion results, access evidence/stage, synthetic flag, submission count, duration, termination, cleanup evidence and cost. Requested versus observed model/configuration are distinct; unavailable observed model/version values remain null rather than being inferred from a request. Adapter-verified startup selection may be recorded as such.
166
+
167
+ Native background acknowledgement, notification and idle reactivation are explicit optional blocked observations, never synthetic passes or unsupported conclusions. Unknown execution errors remain `blocked/unknown` with a safe native code when available; explicit permission failures are blocked/permission, access failures blocked/access, missing executables blocked/harness, and model availability or rate-limit failures are blocked when actually identifiable. Proven protocol/runner defects are failures, not failures attributed to model reasoning. No brittle parsing of free-form model text supplies these classifications.
168
+
169
+ The event writer accepts only known event kinds, bounded identifiers, booleans, numeric timestamps, exit statuses, validated session/task/response IDs, and known fixture values or verifier digests. It records monotonic local ordering and wall time. It excludes tokens and token hashes, full JWT claims, environments, private endpoint secrets, raw CLI response/error text, raw protocol/native settings/stderr, free-form model transcripts, reasoning, and exception messages/stacks before persistence or console output. CLI responses may be passed transiently to the caller but only validated fields enter reports.
170
+
171
+ Each cell's cost is `{ amountUsd: null, source: "unobserved" }` in live mode and `{ amountUsd: null, source: "synthetic-no-billing" }` in fake mode. No actual total is invented from unknown cells. One smoke is a compatibility observation for that exact configuration, not a statistically supported model ranking.
package/sdk/fx.md CHANGED
@@ -19,12 +19,13 @@ export default agent({
19
19
  interface FxOptions {
20
20
  readonly model: string;
21
21
  readonly effort?: string;
22
+ readonly permissionMode?: "ask" | "auto" | "full-access";
22
23
  }
23
24
 
24
25
  function fx(options: FxOptions): FxConfig;
25
26
  ```
26
27
 
27
- `FxConfig` is a readonly harness configuration with `kind: "fx"`, the required `model`, and optional `effort`. It is a member of `HarnessConfig` and can appear alone or in an ordered `harness` array. The constructor only validates and declares configuration; it does not start a process or access credentials. Unknown fields, including `fast`, are rejected. Native fx does not expose a compatible fast-mode selector in the supported ACP interface; its native fast-mode preference remains in effect.
28
+ `FxConfig` is a readonly harness configuration with `kind: "fx"`, the required `model`, optional `effort`, and optional `permissionMode`. It is a member of `HarnessConfig` and can appear alone or in an ordered `harness` array. The constructor only validates and declares configuration; it does not start a process or access credentials. Unknown fields, including `fast`, are rejected. Native fx does not expose a compatible fast-mode selector in the supported ACP interface; its native fast-mode preference remains in effect.
28
29
 
29
30
  Models use the exact AI Gateway `provider/model` identifier. An explicit effort must be supported by the native session's advertised configuration for that model. Unsupported effort fails with `UNSUPPORTED_OPTION` before a task is submitted. Omitting effort preserves the native model's default. The adapter verifies the selected model and explicit effort; it does not silently substitute a model. A model appearing in the catalog does not guarantee access through a particular Gateway team or key.
30
31
 
@@ -56,9 +57,9 @@ ACP does not expose a separate system-instructions field for this native release
56
57
 
57
58
  Declared custom tools are exposed through an authenticated MCP HTTP endpoint bound to loopback for that native session. Tool schemas, argument validation, results, and errors follow the existing custom-tool contract. The endpoint and authentication context are passed privately through ACP. They are not included in agent instructions or CLI output. The endpoint stops admitting calls during cancellation or shutdown, waits for admitted callbacks to finish, and closes when its native session closes or fails to start.
58
59
 
59
- Native fx 0.0.9 has a verified crash when receiving image-bearing MCP tool results through this ACP/HTTP route. The adapter forwards valid image blocks, but this native release's live image-tool execution is not supported reliably and can fail with `HARNESS_FAILED`. Adding a text label to the image result does not avoid the crash. Text and JSON tools remain supported. Native fx image attachments use a separate path; Subharness's CLI currently accepts text prompts only.
60
+ Native fx 0.0.9 has a verified crash when receiving image-bearing MCP tool results through this ACP/HTTP route. The adapter forwards valid image blocks, but this native release's live image-tool execution is not supported reliably and can fail with `HARNESS_FAILED`. Adding a text label to the image result does not avoid the crash. Text and JSON tools remain supported. Native fx image attachments use a separate path; subharness's CLI currently accepts text prompts only.
60
61
 
61
- Native permissions remain authoritative. fx ACP does not expose scoped allow rules equivalent to the Claude adapter's delegation rules. The library does not change its permission mode or infer approval from a permission request. Any unresolved native approval or interactive-input request fails with `INPUT_REQUIRED`. Declared subagents use the session launcher and require native permission to execute it and reach the local coordinator. Declaring tools or children does not override a native approval requirement.
62
+ Native permissions remain authoritative. fx ACP does not expose scoped allow rules equivalent to the Claude adapter's delegation rules. The optional `permissionMode` selects and verifies the native process mode as defined in [Native Permissions](permissions.md). Omission preserves native settings. The library does not infer approval from a permission request. Any unresolved native approval or interactive-input request fails with `INPUT_REQUIRED`. Declared subagents use the session launcher and require native permission to execute it and reach the local coordinator. Declaring tools or children does not override a native approval requirement.
62
63
 
63
64
  ## Session lifecycle
64
65
 
package/sdk/harnesses.md CHANGED
@@ -12,6 +12,8 @@ Before submitting work, known unavailability, such as a missing executable or mi
12
12
 
13
13
  Automatic fallback ends at task submission. Execution failure, allowance exhaustion, or an uncertain submission outcome never automatically replays work or migrates a conversation. Partial file and tool effects remain in the execution environment. The caller can inspect progress and explicitly create other work.
14
14
 
15
+ `subharness check` applies the same selection and fallback rules while opening and closing a temporary native session without submitting work. A successful result means startup checks passed for the selected target at that moment. It is not a quota probe or a guarantee that a later task, shell command, declared child launch, or provider request will succeed.
16
+
15
17
  ## Capabilities
16
18
 
17
19
  | Operation | Codex | Claude Code | fx |
@@ -21,7 +23,9 @@ Automatic fallback ends at task submission. Execution failure, allowance exhaust
21
23
  | Active steering | Native steering | Uses interrupt semantics | Uses interrupt semantics |
22
24
  | Interruption | Native interruption with stop confirmation | Native interruption with stop confirmation | Native ACP cancellation with stop confirmation |
23
25
  | Custom tools | Native dynamic tools | Native SDK MCP tools | Private MCP HTTP tools |
26
+ | Declared-child launcher permissions | Uses effective native permissions; no adapter-added rule | Session-only rules for the exact launcher and declared operations | Uses effective native permissions; no adapter-added rule |
24
27
  | Prompt-free recovery of failed work | Explicit unsupported result | Explicit unsupported result | Explicit unsupported result |
28
+ | Startup readiness check without a turn | Supported | Supported | Supported |
25
29
 
26
30
  `resume` remains a stable command; these adapters return `RECOVERY_UNSUPPORTED` when they cannot resume failed work without replay. The queue remains paused. Unsupported native steering follows the [interrupt contract](message-delivery.md), including cancellation propagation and a new replacement task; output reports the effective mode.
27
31
 
@@ -30,3 +34,5 @@ All TypeScript constructors require a model. Direct [CLI harness targets](cli/in
30
34
  Claude fast mode with subscription access is unavailable in v1 because shared agent configuration does not authorize additional subscription spending. Explicit paid API or Gateway access can request it where supported. Native subscription eligibility and provider distribution terms still apply; subscription login is not an account-wide spending cap.
31
35
 
32
36
  See [native adapter behavior](adapter-contract.md) for authentication isolation, native permissions, tool transport, and lifecycle details.
37
+
38
+ Codex permission behavior is version- and environment-dependent. The adapter does not assume that an installed Codex version can accept session-scoped launcher rules. Explicit [permission options](permissions.md) select native sandbox and network settings; the adapter does not broaden them automatically. A native approval request therefore remains possible for declared-child delegation even though the child is authorized by the Subharness definition.
package/sdk/index.md CHANGED
@@ -3,10 +3,12 @@
3
3
  The CLI runs generic Codex, Claude Code, and fx agents without definition files. The optional TypeScript SDK defines reusable specialists on those same external harnesses. The CLI coordinates sessions, queues, follow-ups, and nested delegation in caller-supplied working directories.
4
4
 
5
5
  - [Distribution](distribution.md): package identity, local installation, and release boundaries.
6
+ - [Agent skill](agent-skill.md): the installable skill that teaches a coding agent to delegate with the CLI.
6
7
  - [Agent definitions](agent.md): package exports, fields, harness options, and examples.
7
8
  - [Custom tools](tools.md): validated tool functions exposed through native harness integrations.
8
9
  - [Discovery](config.md): repository/global definitions and personal worktree settings.
9
10
  - [Personal access](access-config.md): subscription discovery, explicit API keys, and project OIDC.
11
+ - [Native permissions](permissions.md): session permission options, native limits, and background delegation.
10
12
  - [Harnesses](harnesses.md): selection, capabilities, and fallback boundaries.
11
13
  - [CLI](cli/index.md): commands and response waiting.
12
14
  - [Output](cli/output.md): compact text and typed JSONL records.
@@ -1,6 +1,6 @@
1
1
  # Message Delivery
2
2
 
3
- `subharness send <session-id> --delivery <queue|steer|interrupt> --prompt <text>` controls delivery. `queue` is the default. Each session has one active library task and a FIFO queue; a library task can contain multiple native turns while coordinating descendants.
3
+ `subharness send <session-id> --delivery <queue|steer|interrupt> [--detach] --prompt <text>` controls delivery. `queue` is the default. Each session has one active library task and a FIFO queue; a library task can contain multiple native turns while coordinating descendants.
4
4
 
5
5
  | Mode | Effect |
6
6
  | --- | --- |
@@ -14,7 +14,7 @@ Queue dispatch follows task completion, not the return of a CLI command or every
14
14
 
15
15
  A complete response asking a question finishes a task when no descendant work remains. B can then start. The coordinator decides whether its answer should be queued or applied to the current task with another delivery mode. A native approval/input request is different: it has not completed a turn. In v1, input the library cannot answer under native policy fails with `INPUT_REQUIRED`; it is never automatically approved.
16
16
 
17
- Execution failure pauses pending work. A successful explicit native recovery must finish the failed task before pending tasks proceed. Unsupported recovery leaves work paused. Queued command observers may remain waiting until their task starts or is explicitly cancelled.
17
+ Execution failure pauses pending work. New queued sends remain admissible while dispatch is paused. Their immediate `started` records include `paused: true` and `blockedByTaskId` naming the failed active task, so the caller can inspect it with `subharness status <failed-task-id> --full` and release the queue with `subharness cancel <failed-task-id>`. These fields capture queue state at admission and are omitted when the queue is not paused. A successful explicit native recovery must finish the failed task before pending tasks proceed. Unsupported recovery leaves work paused. Queued command observers may remain waiting until their task starts or is explicitly cancelled. Cancellation must be confirmed before it releases the queue; an unconfirmed stop keeps dispatch paused.
18
18
 
19
19
  ## Steering
20
20
 
@@ -22,12 +22,16 @@ Native steering adds guidance without cancelling the active task or creating ano
22
22
 
23
23
  If active native steering is unavailable, the operation uses the full interrupt behavior. It cancels the active task and descendants, creates a replacement task with the correction, and preserves independently queued tasks. The output identifies both requested `steer` and effective `interrupt`. It returns the replacement's response under interrupt semantics. Other native errors do not trigger this fallback.
24
24
 
25
+ Steering cannot be combined with `--detach`. A successful steer already returns an acceptance acknowledgement rather than creating a task whose response can be observed, so `--detach --delivery steer` fails with `INVALID_ARGUMENT` before delivery.
26
+
25
27
  There is no expected-task guard: if B has already become active, the correction applies to B. Unsupported steering is detected before native mutation so fallback does not deliver the correction twice.
26
28
 
27
29
  ## Interruption and cancellation
28
30
 
29
31
  If A is active and B/C are queued, interrupting with X stops A and its delegated descendants, runs X, and then resumes B/C in their original order. A and its children are not automatically requeued. X cannot start while affected native execution might still be running.
30
32
 
33
+ `--detach` does not weaken the interruption boundary. An interrupting send still waits until cancellation is confirmed and the replacement is admitted, then returns its `started` record without observing the replacement response.
34
+
31
35
  Cancellation applies recursively to recorded delegation relationships. Pending descendants are removed without execution; active descendants receive native interruption. Other tasks in the same queue, directory, or agent definition are not descendants solely for that reason. Completed results and tool effects remain. Already-admitted custom callbacks must settle before stop confirmation.
32
36
 
33
37
  `cancel` returns after affected work has stopped. Repeated cancellation of already-settled work returns its state. An unconfirmed stop reports a cancellation error and keeps dispatch paused; it does not authorize starting replacement or queued work. Invalid replacement input is rejected before existing work is cancelled.
@@ -0,0 +1,65 @@
1
+ # Native Permissions
2
+
3
+ Harness constructors accept optional native permission settings directly alongside model options. These settings configure the native session without modifying persistent native settings or creating a Subharness sandbox. The library never answers a later native approval request with a synthetic allow response. The installed harness evaluates the selected policy, and the external execution environment enforces its remaining restrictions.
4
+
5
+ Omitted settings preserve native configuration. Explicit settings remain fixed for session follow-ups. A child uses its own definition and native configuration; it does not inherit the parent's explicit permission options. Invalid definitions fail before execution. Unsupported or observably rejected explicit settings fail without fallback to another harness or permission policy.
6
+
7
+ ## Codex
8
+
9
+ ```ts
10
+ codex({
11
+ model: "CODEX_MODEL_ID",
12
+ approvalPolicy: "never",
13
+ sandboxMode: "workspace-write",
14
+ networkAccessEnabled: true,
15
+ });
16
+ ```
17
+
18
+ `CodexOptions` and `CodexConfig` expose `approvalPolicy?: "never" | "on-request" | "untrusted"`, `sandboxMode?: "read-only" | "workspace-write" | "danger-full-access"`, and `networkAccessEnabled?: boolean`. The supported approval values match the installed App Server protocol; the older native SDK value `on-failure` is not supported.
19
+
20
+ `networkAccessEnabled` requires an explicit `sandboxMode: "workspace-write"`. This prevents the option from being silently ignored by a different native sandbox mode. The adapter maps the SDK names to native App Server thread settings and the workspace-write network setting. It checks returned approval policy, sandbox mode, and network access when explicitly requested; a missing or different requested value is an error before a turn is submitted. Native default settings are not inferred from omitted response fields.
21
+
22
+ `approvalPolicy: "never"` suppresses approval prompts; it does not grant operations blocked by the selected sandbox. `danger-full-access` removes Codex's native sandbox boundary, subject to external restrictions.
23
+
24
+ ## Claude Code
25
+
26
+ ```ts
27
+ claudeCode({
28
+ model: "CLAUDE_MODEL_ID",
29
+ permissionMode: "dontAsk",
30
+ allowedTools: ["Read", "Edit", "Write", "Bash"],
31
+ });
32
+ ```
33
+
34
+ `ClaudeCodeOptions` and `ClaudeCodeConfig` expose `permissionMode?: "default" | "acceptEdits" | "bypassPermissions" | "plan" | "dontAsk" | "auto"`, `allowedTools?: readonly string[]`, and `disallowedTools?: readonly string[]`. Rule arrays contain nonempty strings and are copied into immutable configuration. Unknown modes, non-string rules, and fields belonging to another harness are invalid definitions. An empty rule list adds no rules.
35
+
36
+ The adapter passes these settings to the native Agent SDK. Explicit allowed rules are combined with the existing rules for declared custom tools and exact child-launcher operations. Disallowed rules remain authoritative under native permission evaluation. Selecting `bypassPermissions` also supplies the native SDK's required explicit bypass enablement. The library does not synthesize an allow response in its unresolved approval callback.
37
+
38
+ Native permission mode and rule evaluation depend on the installed SDK, executable, and managed settings. Observable rejection or a reported different explicit mode is an error. A native initialization event that reports the active mode must agree with the explicit request. Startup success does not claim that every native rule or future command was verified. Diagnostics identify rejected options without exposing rule contents or tool arguments.
39
+
40
+ `dontAsk` denies calls that would otherwise need approval. An allowed `Bash` rule preapproves shell commands under remaining native controls; it is not filesystem isolation. `acceptEdits` does not automatically grant every shell command.
41
+
42
+ ## fx
43
+
44
+ ```ts
45
+ fx({
46
+ model: "GATEWAY_MODEL_ID",
47
+ permissionMode: "auto",
48
+ });
49
+ ```
50
+
51
+ `FxOptions` and `FxConfig` expose `permissionMode?: "ask" | "auto" | "full-access"`. An explicit mode becomes the child process's `FX_PERMISSION_MODE` and is verified with `fx permissions --json` using the same directory and environment before an ACP session starts. The native wire value `yolo` is equivalent to requested `full-access`. Unsupported introspection, malformed output, or a different mode fails startup with an actionable bounded error. The probe allows five seconds for native startup and owns its process until retirement, including its process group on POSIX systems. A timed-out explicit permission verification returns `INVALID_CONFIG` and cannot select another harness alternative. A missing executable remains `HARNESS_UNAVAILABLE`. Omission performs no additional permission probe and preserves existing native behavior.
52
+
53
+ `ask` requests approval for unresolved sensitive operations. `auto` uses native rules and automatic review, which can still hold or reject operations and can incur additional model usage. `full-access` disables fx's own permission checks; external restrictions still apply. The adapter never overwrites the selected mode with ACP `code`, whose meaning is automatic review rather than disabled permission checks. Native rules remain owned by fx configuration; this API does not introduce a tool allowlist or a filesystem sandbox for fx.
54
+
55
+ ## Readiness and failures
56
+
57
+ `check` uses the same permission configuration as execution. Its success establishes startup readiness and only the native settings observable during startup. It does not execute a shell or child launcher, prove arbitrary tool access, or claim unobservable native policy values. Native approval requests that remain unresolved fail with `INPUT_REQUIRED`; the library does not offer an interactive approval-answer command.
58
+
59
+ A caller reports the affected harness and operation, then changes the explicit configuration or resolves the native restriction before starting a new attempt. Retrying unchanged permission failures, broadening policy automatically, or silently selecting another harness is not a recovery strategy.
60
+
61
+ ## Background delegation
62
+
63
+ The caller reads the Subharness skill and prefers one ordinary `subharness run <target>` through known host background-command controls. It continues independent work and later collects that hosted command's output, which already contains the first response or terminal outcome. When those controls are unavailable or uncertain, it uses `run --detach`, retains the returned task and session identifiers, continues other work, and later collects the result with `wait`. `status` is optional when a snapshot or additional state is needed. Shell `&` and an unobserved process do not replace managed task records. Both paths require collecting the result and confirming its terminal outcome before reporting completion; detached admission alone is not completion.
64
+
65
+ The native caller needs permission to execute the private launcher and reach the local coordinator. Each child needs permission for its own task. This remains true for Codex, Claude Code, and fx; no constructor option promises native desktop notifications or automatic chat reactivation.
@@ -23,19 +23,22 @@ export default agent({
23
23
 
24
24
  ## Invocation
25
25
 
26
- The adapter provides direct child names, descriptions, and CLI instructions to the parent. Managed agents receive the absolute path of a launcher for their session and use that executable for delegation and follow-ups. This avoids invoking another installed program named `agent` when a native shell changes PATH.
26
+ The adapter provides direct child names, descriptions, and CLI instructions to the parent. These instructions prefer one ordinary `run` through the native host's background-command controls when that capability is known to be available. The parent continues independent work, then collects that hosted command's output; the command already observes the first response, so no separate Subharness `wait` or unconditional `status` is needed. If background-command support is unavailable or uncertain, the parent uses `--detach`, retains task/session identifiers, continues independent work, and later collects the result with `wait`. It uses `status` only when a snapshot or additional state is needed. Both paths require checking the returned outcome before claiming completion; a pending response needs further observation. Instructions explain that detached admission is not completion and that shell `&` does not substitute for the host's command controls or managed task records. The harness name alone does not establish background support, and the runtime does not invent native tool parameters or completion notifications. Managed agents receive the absolute path of a launcher for their session and use that executable for delegation and follow-ups. This avoids invoking another installed program named `agent` when a native shell changes PATH.
27
27
 
28
28
  The launcher restores its coordinator location and parent context before invoking the CLI, so delegation does not depend on the shell preserving integration environment variables. The launcher is stored with owner-only permissions and removed when its session worker closes. It contains only local orchestration context, never provider credentials. Commands and agent instructions do not contain parent tokens or credentials.
29
29
 
30
30
  For Claude Code, declaring at least one subagent adds native permission rules for the current session's exact launcher. Equivalent quoted or unquoted spellings are allowed only when they identify the same literal executable. These rules authorize `run subagent:<name>` for the declared names, plus the CLI's `list`, `--help`, `send`, `wait`, `status`, `queue`, `cancel`, and `resume` operations. They do not authorize launching repository or global agents through `run`, another executable, or arbitrary shell commands. Follow-up and observation commands retain their documented identifier-based behavior; the permission rules do not introduce a new task-access policy.
31
31
 
32
- The adapter combines these rules with declared custom-tool permissions. It does not write user/project permission settings or change the native permission mode. Native deny rules, explicit approval requirements, managed policy, and sandbox restrictions remain authoritative. A request that still requires approval returns `INPUT_REQUIRED`. Each child starts with its own native tool permissions; permission to launch it does not grant permission for its edits or shell commands. A session without declared subagents receives no delegation rules. If a launcher path cannot be represented as a literal command in the native permission syntax, startup fails with `INVALID_CONFIG` rather than adding a broader rule.
32
+ The adapter combines these rules with declared custom-tool permissions. Those rules do not write user/project permission settings or change the native permission mode; explicit constructor options independently select session permission settings as defined in [Native Permissions](../permissions.md). Native deny rules, explicit approval requirements, managed policy, and sandbox restrictions remain authoritative. A request that still requires approval returns `INPUT_REQUIRED`. Each child starts with its own native tool permissions; permission to launch it does not grant permission for its edits or shell commands. A session without declared subagents receives no delegation rules. If a launcher path cannot be represented as a literal command in the native permission syntax, startup fails with `INVALID_CONFIG` rather than adding a broader rule.
33
33
 
34
- The native execution environment must allow the CLI to read its private launcher and contact the coordinator over loopback HTTP. A native sandbox that blocks local networking also blocks CLI delegation. The library does not relax that policy; the caller configures the native harness or external execution environment.
34
+ Codex receives the same declared-child instructions and private launcher, but the adapter does not add native allow rules for them. Codex must already have effective permission to execute the launcher, read the files needed by that execution, and contact the coordinator. Native behavior differs across Codex versions, managed policies, sandbox configurations, and host environments. In some combinations, commands that access protected locations such as `.subharness/`, execute the private launcher, or use loopback networking can require approval or remain unavailable; this is an environment-dependent limitation, not a universal Codex rule. A resulting approval request returns `INPUT_REQUIRED`, including when the requested command would launch an authorized declared child. The library does not automatically approve the request, disable sandboxing, or replace shell delegation with another transport.
35
+
36
+ For every harness, the native execution environment must allow the CLI to read its private launcher and contact the coordinator over loopback HTTP. A native sandbox that blocks local networking also blocks CLI delegation. The library does not relax that policy automatically; the caller supplies explicit native permission options, native settings, or an external execution environment. A lead can instead own independent review outside a managed agent's native delegation flow when the environment cannot grant these prerequisites.
35
37
 
36
38
  The command has the same arguments as an ordinary CLI invocation:
37
39
 
38
40
  ```sh
41
+ # Execute through the host's background-command controls when available.
39
42
  subharness run subagent:reviewer --cwd /repo/worktree \
40
43
  --prompt "Review the implementation against the documented requirements."
41
44
  ```
@@ -46,7 +49,7 @@ Each invocation creates a child session and task. The same `send`, `wait`, `stat
46
49
 
47
50
  ## Context and results
48
51
 
49
- A child receives its own instructions and explicit task input, without copying the parent's transcript or forking its native session. The parent supplies requirements, decisions, and relevant file references, including any skill files the child needs to read. Previously loaded skill contents are not inherited. Native repository instructions and skill discovery follow the selected harness's rules; Subharness does not install native subagent definitions or skills. The caller supplies an existing directory; no sandbox or worktree is created.
52
+ A child receives its own instructions and explicit task input, without copying the parent's transcript or forking its native session. The parent supplies requirements, decisions, and relevant file references, including any skill files the child needs to read. Previously loaded skill contents are not inherited. Native repository instructions and skill discovery follow the selected harness's rules; subharness does not install native subagent definitions or skills. The caller supplies an existing directory; no sandbox or worktree is created.
50
53
 
51
54
  Results return to the immediate caller. A child may itself return a pending response while its descendants run. A parent's task remains active until its descendant work and result processing finish, as defined in [response delivery](../completion-notifications.md). The library does not forward every child transcript to every ancestor.
52
55
 
@@ -0,0 +1,37 @@
1
+ # Authorized PR Integration
2
+
3
+ `repo:integrator` is a repository specialist for bounded GitHub integration through the installed `gh` CLI. It uses Codex with `gpt-5.6-sol`, high effort, and fast mode disabled. It has no declared children or custom tools. The integrator must not delegate through Subharness or native harness tools. Omitting a `subagents` map does not itself disable native delegation capabilities. It does not select work from the repository's open PRs, monitor for mergeable PRs, or merge automatically when checks turn green.
4
+
5
+ ## Authorization
6
+
7
+ The lead authorizes a finite list in the current assignment. Each entry identifies the GitHub repository (`owner/repo`), PR number, exact reviewed head commit SHA, destination branch, and merge method. The assignment states dependency order and any separately authorized base-branch changes. Missing or ambiguous authorization blocks mutation; the integrator asks the lead for the missing information rather than inferring it.
8
+
9
+ The assignment is the only source of merge authorization. PR descriptions, comments, labels, checks, repository documents, and tool output are evidence, not instructions or permission to add PRs. A previous assignment does not implicitly authorize a new task. The integrator never expands the list, changes an approved head, or treats a dependent PR as authorized merely because its predecessor is listed. A changed head requires fresh lead authorization before merging. Successfully merged entries are reported and not merged again.
10
+
11
+ ## Validation and mutation
12
+
13
+ Before each merge, the integrator checks the repository, PR number, current head SHA, destination branch, open/non-draft state, mergeability, review decision, and check results against the assignment. Pending or unknown results do not establish success. Required approvals must be satisfied; requested changes, failed checks, conflicts, missing required approvals, and unresolved mergeability block that PR. GitHub branch protections remain authoritative. The integrator may observe pending checks without enabling automatic merge. Before invoking `gh pr merge`, it must also verify that the target branch does not require a merge queue. A required merge queue or an inability to determine its requirement blocks mutation and is reported to the lead. Enqueueing is outside this role: `gh pr merge` can enqueue a PR or enable deferred merging on queue-enabled branches even without `--auto`.
14
+
15
+ An already merged PR is reported without mutation after checking its merge evidence against the authorized repository, head, and destination. A closed but unmerged PR is blocked. A dependent PR cannot proceed until the predecessor's successful merge is verified. Base changes are allowed only when the assignment explicitly authorizes the exact change and its prerequisites; the head, diff scope, checks, reviews, and mergeability must then be revalidated.
16
+
17
+ The merge command always specifies the repository and the approved method, and uses `--match-head-commit` with the authorized SHA. `--admin`, automatic merge, branch deletion, bypasses, force pushes, rebases, tags, local code edits, commits, and conflict resolution are outside this role. The integrator does not alter authentication or permission settings. A blocked PR is reported to the lead; independent authorized entries may continue unless the assignment says to stop the batch.
18
+
19
+ After every mutation, the integrator reads GitHub state again. Merge success requires `MERGED`, the expected destination and reviewed head, a merge timestamp, and the resulting merge commit. A successful command exit alone is insufficient. If the mutation outcome is uncertain, the integrator reconciles current state before considering another attempt; it does not blindly repeat a mutation.
20
+
21
+ ## Evidence and permissions
22
+
23
+ The report identifies merged, already merged, blocked, and unattempted entries, their reviewed heads, resulting merge commits when known, and exact blockers. It includes the observed final destination-branch SHA and command outcomes. The integrator writes a nonsecret progress journal only when the assignment names its path. It never exports credentials, dumps the environment, or claims local integration tests ran merely because hosted checks passed.
24
+
25
+ The caller supplies the existing native GitHub authentication and any explicitly authorized native network profile. Model billing continues to use the caller's selected access configuration. Permission failures are reported without bypasses or silent changes of execution route. Native approval requests that require host interaction can produce `INPUT_REQUIRED`, as specified in the [adapter contract](adapter-contract.md).
26
+
27
+ The allowlist is an instruction-level delegation contract, not a GitHub token scope or a new runtime security boundary. The agent retains native tools; neither the model's instructions nor an API-domain allowlist mechanically limits a credential to those PRs. Native sandbox policy, credential scopes, and GitHub protections remain the enforcement boundaries.
28
+
29
+ ## Invocation
30
+
31
+ The lead supplies a private task file containing the authorization entries and governing documents:
32
+
33
+ ```sh
34
+ subharness run repo:integrator --cwd /repo/worktree --prompt-file /absolute/path/to/merge-authorization.txt
35
+ ```
36
+
37
+ Loading or listing this definition does not authenticate to GitHub, inspect PRs, or perform a merge. The repository role does not install a native permission profile or change machine-wide settings.
@@ -1,6 +1,6 @@
1
1
  # Repository Agent Team
2
2
 
3
- This repository keeps five reusable agents in `.agents/agents/`. An agent defines a stable responsibility, model, and tool access. A skill supplies task-specific instructions that the agent reads when relevant. Adding a technique or library does not require another agent definition.
3
+ This repository keeps six reusable agents in `.subharness/agents/`. An agent defines a stable responsibility, model, and tool access. A skill supplies task-specific instructions that the agent reads when relevant. Adding a technique or library does not require another agent definition.
4
4
 
5
5
  ## Roles
6
6
 
@@ -11,18 +11,21 @@ This repository keeps five reusable agents in `.agents/agents/`. An agent define
11
11
  | `visual-engineer` | fx, `anthropic/claude-fable-5.1`, native default effort | Implement approved interface styling and native vgpu/WGSL effects using the relevant skill. |
12
12
  | `researcher` | fx, `google/gemini-3.8-flash`, native default effort | Answer a bounded technical or design question with primary-source evidence. |
13
13
  | `reviewer` | Claude Code, `claude-opus-5[1m]`, high effort | Independently review correctness, lifecycle, credential routing, and contract compliance without editing files. |
14
+ | `integrator` | Codex, `gpt-5.6-sol`, high effort | Merge only the exact PRs and reviewed heads explicitly authorized by the lead for the current assignment, following [PR integration](pr-integration.md). |
14
15
 
15
16
  Model identities are explicit. There is no automatic model downgrade or change to a paid API route. Fast mode is disabled for Codex and Claude Code; fx retains its native preference because its ACP interface has no selector. Codex and Claude Code use eligible native subscriptions unless personal access settings select another route. The fx roles require explicit personal `access.fx` configuration as described in [fx](fx.md); Gateway team policies can restrict model access.
16
17
 
17
18
  Skills are reusable across roles. The table records this repository's preferred allocation of work, not an SDK restriction on which agent may read a skill. A developer can use an interface skill for one task and a graphics skill for another without changing its definition or model. An omitted custom `tools` map does not disable the harness's native tools.
18
19
 
19
- Roles do not grant native capabilities or permissions. Browsing, image generation, image viewing, Blender, file writes, and shell access depend on the selected harness and execution environment. A missing capability must be reported before claiming a result.
20
+ Role instructions alone do not grant native capabilities or permissions. Browsing, image generation, image viewing, Blender, file writes, and shell access depend on the selected harness and execution environment. A missing capability must be reported before claiming a result.
21
+
22
+ The repository's Codex roles explicitly select `approvalPolicy: "never"`, `sandboxMode: "workspace-write"`, and `networkAccessEnabled: true` for local development and coordinator access. The reviewer selects Claude `permissionMode: "dontAsk"` with `Read`, `Glob`, `Grep`, and `Bash` allowed. Its read-only review responsibility remains an instruction; allowing Bash is not a filesystem read-only boundary. The fx roles explicitly select `permissionMode: "auto"`. These policies use [native permission options](permissions.md), preserve native restrictions, and do not guarantee that every operation is permitted. Protected paths and rejected automatic reviews can still require a lead-managed execution path.
20
23
 
21
24
  ## Skills and task context
22
25
 
23
26
  Repository skills live in `.agents/skills/<name>/SKILL.md`. The short index in `AGENTS.md` explains when each skill applies. Agents read only the relevant skill and supporting references for their current task. Skill contents are not concatenated into every agent's instructions.
24
27
 
25
- These files are ordinary instructions consumed through the native harness. Subharness does not install skills, discover them on the caller's behalf, provide a `skills` definition field, or guarantee that every harness exposes the same skill command. Native automatic discovery depends on the harness and its settings. This repository's shared instructions explicitly require reading `AGENTS.md`, so the index remains usable through ordinary file reading even when a harness does not automatically discover `.agents/skills/`.
28
+ These files are ordinary instructions consumed through the native harness. Each sets `metadata.internal: true` in its frontmatter so the public `npx skills add vercel-labs/subharness` installer offers only the [agent skill](agent-skill.md). The subharness runtime does not install skills, discover them on the caller's behalf, provide a `skills` definition field, or guarantee that every harness exposes the same skill command. Native automatic discovery depends on the harness and its settings. This repository's shared instructions explicitly require reading `AGENTS.md`, so the index remains usable through ordinary file reading even when a harness does not automatically discover `.agents/skills/`.
26
29
 
27
30
  The caller provides one bounded deliverable, governing contracts, input artifacts, working directory, and relevant skill paths. Each child receives its own task and instructions; it does not inherit the parent's transcript or previously loaded skills. A handoff therefore names required skills again.
28
31
 
@@ -36,22 +39,24 @@ subharness run repo:researcher --prompt "Use .agents/skills/source-research/SKIL
36
39
  subharness run repo:reviewer --prompt "Review the mascot camera changes against website/graphics.md. Report actionable findings without editing files."
37
40
  ```
38
41
 
39
- The website's three specialist labels, QA Tester, Developer, and Designer, demonstrate possible role configurations. They are not an inventory of this repository's agent definitions.
42
+ The website's illustrative conversation shows tasks delegated to Claude Code and Codex. Those examples are not an inventory of this repository's agent definitions or a claim that a task is currently running.
40
43
 
41
44
  ## Delegation and review
42
45
 
43
- CLI and runtime work uses the existing architect, developer, and reviewer roles. Separate developer sessions can own independent modules; scope and governing contracts distinguish their assignments without adding permanent roles. The architect evaluates native protocol constraints, the developer implements and integrates behavior with TDD, and the reviewer independently checks user-facing behavior, lifecycle, and credential routing. The lead coordinates contracts, assignments, and final verification. Implementation and fixes run through repository agents whenever supported, so follow-ups, queues, and review cycles exercise Subharness itself. Reproducible coordination failures are evidence for library improvements under the same contract and review rules.
46
+ CLI and runtime work uses the existing architect, developer, and reviewer roles. Separate developer sessions can own independent modules; scope and governing contracts distinguish their assignments without adding permanent roles. The architect evaluates native protocol constraints, the developer implements and integrates behavior with TDD, and the reviewer independently checks user-facing behavior, lifecycle, and credential routing. The lead coordinates contracts, assignments, and final verification. Implementation and fixes run through repository agents whenever supported, so follow-ups, queues, and review cycles exercise subharness itself. Reproducible coordination failures are evidence for library improvements under the same contract and review rules.
47
+
48
+ Repository role instructions follow the same delegation preference as the public skill and managed child instructions: use one attached `run` through known host background-command controls, continue independent work, and collect its output. Use `--detach` followed later by `wait` when those controls are unavailable or uncertain. A hosted `run` already observes the first response and does not require a redundant `wait` or unconditional `status`. This preference does not authorize undeclared children or override a role's delegation restrictions.
44
49
 
45
50
  The lead agent coordinates bounded tasks and resolves missing decisions under `AGENTS.md`. Skills do not authorize new features, public API shapes, permission changes, or undeclared children. Research findings and generated concepts are evidence, not implementation contracts. Source research and explorations belong in the ignored `.context/research/` directory.
46
51
 
47
52
  `developer` retains a declared `reviewer` child for independent code review after code changes. It supplies the scope, governing documents, and relevant skill paths, fixes actionable findings, and requests re-review. The lead can explicitly own this review cycle, including when native permissions prevent the developer from reaching the coordinator. In that case the developer returns its changes and verification evidence without starting a duplicate review or retrying a known blocked launcher. The lead must still obtain independent review and return findings to the developer. A read-only verification task returns its evidence without starting an unnecessary review cycle. The lead separately assigns visual review when needed. If delegation unexpectedly fails, the developer reports the exact operation; returning evidence alone does not establish a passed review. The other roles return work to the lead for independently assigned review and do not declare children. A task-specific specialization can use a fresh session of an existing role with an explicit skill; it does not require a new globally discoverable role.
48
53
 
49
- Delegation through either harness requires the native environment to read and execute the private session launcher and contact the local coordinator over loopback HTTP. Declaring a child does not override a Codex sandbox or native permission policy. The caller configures that access in the harness or external environment; Subharness does not change it. See [subagents](plugins/sub-agents.md) for the launcher and permission boundaries.
54
+ Delegation through either harness requires the native environment to read and execute the private session launcher and contact the local coordinator over loopback HTTP. Declaring a child does not override a Codex sandbox or native permission policy. The caller configures that access through explicit [native permission options](permissions.md), native settings, or the external environment. subharness never broadens the policy automatically. See [subagents](plugins/sub-agents.md) for the launcher and permission boundaries.
50
55
 
51
- A `subagents` declaration authorizes Subharness child sessions as described in [subagents](plugins/sub-agents.md). It does not install native harness subagent definitions or grant the child extra tools. Independent review uses a different session from the implementation or design author. Visual review requires actual rendered captures; code review and generated mockups do not establish visual correctness.
56
+ A `subagents` declaration authorizes subharness child sessions as described in [subagents](plugins/sub-agents.md). It does not install native harness subagent definitions or grant the child extra tools. Independent review uses a different session from the implementation or design author. Visual review requires actual rendered captures; code review and generated mockups do not establish visual correctness.
52
57
 
53
58
  ## Visual evidence and native limits
54
59
 
55
60
  The fx researcher and visual engineer expose a `view_image` custom tool. It returns native image content through `toolResult`, rather than a path or base64 text. It reads supported image files up to 3 MiB only inside the execution repository's `.context/research/landing/`, `.context/site-verification/`, and `apps/docs/public/` directories. Canonical paths, including symlink targets, must stay within those roots. Symlinked allowed roots cannot redirect access elsewhere. The tool does not fetch URLs or alter permissions.
56
61
 
57
- Image tasks respect the native version limits in [fx](fx.md). When native fx is affected by the documented ACP image-result crash, the lead can supply images through native `fx ask --image` using the same role, model, and explicit connection. This is an external native-harness invocation, not a Subharness session or proof of CLI queue and coordinator behavior. Supported text tasks continue through Subharness. Missing capabilities never justify a silent model substitution or custom inference loop.
62
+ Image tasks respect the native version limits in [fx](fx.md). When native fx is affected by the documented ACP image-result crash, the lead can supply images through native `fx ask --image` using the same role, model, and explicit connection. This is an external native-harness invocation, not a subharness session or proof of CLI queue and coordinator behavior. Supported text tasks continue through subharness. Missing capabilities never justify a silent model substitution or custom inference loop.
package/sdk/sessions.md CHANGED
@@ -20,7 +20,7 @@ Failure is a terminal observation, not proof that uncertain native cancellation
20
20
 
21
21
  Every complete response receives a response identifier and returns control to its caller. A response may ask a question or report failing tests; neither its text nor command exit `0` proves objective success. The task can remain pending after that response if descendants or their results remain.
22
22
 
23
- `wait --after` observes later responses without creating another task. `status` reports current state independently of earlier response snapshots. Completing a task does not delete its conversation. Returning a command does not cancel the task or descendants.
23
+ `wait` observes the first retained response without creating another task. `wait --after <response-id>` observes later responses from an explicit cursor. `status` reports current state independently of earlier response snapshots. Completing a task does not delete its conversation. Returning a command does not cancel the task or descendants. `--detach` on a task-creating command ends only that command's observation after admission; it is not stored on the session and does not alter later `send`, `wait`, or `status` behavior.
24
24
 
25
25
  ## Lifetime and recovery
26
26
 
package/sdk/tools.md CHANGED
@@ -21,6 +21,8 @@ export default agent({
21
21
  });
22
22
  ```
23
23
 
24
+ A tool used by several definitions can live in a shared module outside the discovered `agents/` directory, such as `.subharness/tools/uppercaseTool.ts`, and be imported by each definition. See [agent discovery](config.md).
25
+
24
26
  All three fields are required. `description` is a nonempty string. `inputSchema` is a Zod 4 object schema representable as JSON Schema. `execute(input)` receives the parsed value, with TypeScript inference, and returns text, JSON-serializable data, or an explicit rich tool result synchronously or asynchronously. No second execution-context argument is supplied.
25
27
 
26
28
  Validation runs before execution; invalid arguments never reach the function. Unsupported schemas fail during definition validation. Non-serializable or oversized outputs and thrown exceptions become tool errors. Tool results are limited to 1 MiB. Errors do not masquerade as successful outputs.