subharness 0.0.1 → 0.0.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +201 -21
- package/README.md +51 -18
- package/dist/adapters/claude-process.d.ts +10 -0
- package/dist/adapters/claude-process.js +58 -0
- package/dist/adapters/claude-process.js.map +1 -0
- package/dist/adapters/claude.js +52 -10
- package/dist/adapters/claude.js.map +1 -1
- package/dist/adapters/codex-input.d.ts +2 -0
- package/dist/adapters/codex-input.js +14 -0
- package/dist/adapters/codex-input.js.map +1 -0
- package/dist/adapters/codex-permissions.d.ts +13 -0
- package/dist/adapters/codex-permissions.js +30 -0
- package/dist/adapters/codex-permissions.js.map +1 -0
- package/dist/adapters/codex.js +12 -6
- package/dist/adapters/codex.js.map +1 -1
- package/dist/adapters/fx-auth.js +3 -0
- package/dist/adapters/fx-auth.js.map +1 -1
- package/dist/adapters/fx-permissions.d.ts +3 -0
- package/dist/adapters/fx-permissions.js +128 -0
- package/dist/adapters/fx-permissions.js.map +1 -0
- package/dist/adapters/fx.js +4 -0
- package/dist/adapters/fx.js.map +1 -1
- package/dist/cli/args.d.ts +3 -1
- package/dist/cli/args.js +18 -12
- package/dist/cli/args.js.map +1 -1
- package/dist/cli/dashboard-client.d.ts +3 -0
- package/dist/cli/dashboard-client.js +78 -0
- package/dist/cli/dashboard-client.js.map +1 -0
- package/dist/cli/dashboard-layout.d.ts +21 -0
- package/dist/cli/dashboard-layout.js +211 -0
- package/dist/cli/dashboard-layout.js.map +1 -0
- package/dist/cli/dashboard-renderer.d.ts +30 -0
- package/dist/cli/dashboard-renderer.js +80 -0
- package/dist/cli/dashboard-renderer.js.map +1 -0
- package/dist/cli/dashboard.d.ts +31 -0
- package/dist/cli/dashboard.js +81 -0
- package/dist/cli/dashboard.js.map +1 -0
- package/dist/cli/help.d.ts +1 -1
- package/dist/cli/help.js +35 -9
- package/dist/cli/help.js.map +1 -1
- package/dist/cli/main.js +19 -6
- package/dist/cli/main.js.map +1 -1
- package/dist/cli/output.js +10 -1
- package/dist/cli/output.js.map +1 -1
- package/dist/config/access.js +4 -4
- package/dist/config/access.js.map +1 -1
- package/dist/config/loader.js +2 -2
- package/dist/config/loader.js.map +1 -1
- package/dist/runtime/client.d.ts +5 -1
- package/dist/runtime/client.js +91 -40
- package/dist/runtime/client.js.map +1 -1
- package/dist/runtime/coordinator.d.ts +7 -1
- package/dist/runtime/coordinator.js +72 -8
- package/dist/runtime/coordinator.js.map +1 -1
- package/dist/runtime/daemon.js +100 -18
- package/dist/runtime/daemon.js.map +1 -1
- package/dist/runtime/dashboard-workspace.d.ts +2 -0
- package/dist/runtime/dashboard-workspace.js +38 -0
- package/dist/runtime/dashboard-workspace.js.map +1 -0
- package/dist/runtime/dashboard.d.ts +28 -0
- package/dist/runtime/dashboard.js +19 -0
- package/dist/runtime/dashboard.js.map +1 -0
- package/dist/runtime/readiness.d.ts +4 -0
- package/dist/runtime/readiness.js +17 -0
- package/dist/runtime/readiness.js.map +1 -0
- package/dist/runtime/select-native.d.ts +2 -1
- package/dist/runtime/select-native.js +5 -3
- package/dist/runtime/select-native.js.map +1 -1
- package/dist/runtime/service.d.ts +1 -1
- package/dist/runtime/service.js +10 -3
- package/dist/runtime/service.js.map +1 -1
- package/dist/runtime/startup-lock.d.ts +5 -0
- package/dist/runtime/startup-lock.js +37 -0
- package/dist/runtime/startup-lock.js.map +1 -0
- package/dist/runtime/startup-protocol.d.ts +17 -0
- package/dist/runtime/startup-protocol.js +23 -0
- package/dist/runtime/startup-protocol.js.map +1 -0
- package/dist/runtime/state.d.ts +1 -2
- package/dist/runtime/state.js +8 -43
- package/dist/runtime/state.js.map +1 -1
- package/dist/runtime/types.d.ts +16 -2
- package/dist/runtime/types.js.map +1 -1
- package/dist/runtime/worker-client.d.ts +7 -2
- package/dist/runtime/worker-client.js +69 -12
- package/dist/runtime/worker-client.js.map +1 -1
- package/dist/runtime/worker-server.d.ts +7 -0
- package/dist/runtime/worker-server.js +101 -0
- package/dist/runtime/worker-server.js.map +1 -0
- package/dist/runtime/worker.js +2 -66
- package/dist/runtime/worker.js.map +1 -1
- package/dist/sdk/definitions.js +23 -13
- package/dist/sdk/definitions.js.map +1 -1
- package/dist/sdk/permission-validation.d.ts +4 -0
- package/dist/sdk/permission-validation.js +55 -0
- package/dist/sdk/permission-validation.js.map +1 -0
- package/dist/sdk/types.d.ts +14 -0
- package/dist/sdk/types.js.map +1 -1
- package/package.json +17 -17
- package/sdk/access-config.md +4 -2
- package/sdk/adapter-contract.md +14 -2
- package/sdk/agent-skill.md +38 -0
- package/sdk/agent.md +6 -2
- package/sdk/cli/dashboard-design.md +40 -0
- package/sdk/cli/index.md +47 -13
- package/sdk/cli/output.md +21 -3
- package/sdk/completion-notifications.md +9 -5
- package/sdk/config.md +33 -8
- package/sdk/distribution.md +21 -7
- package/sdk/evals.md +171 -0
- package/sdk/fx.md +4 -3
- package/sdk/harnesses.md +6 -0
- package/sdk/index.md +2 -0
- package/sdk/message-delivery.md +6 -2
- package/sdk/permissions.md +65 -0
- package/sdk/plugins/sub-agents.md +7 -4
- package/sdk/pr-integration.md +37 -0
- package/sdk/project-team.md +13 -8
- package/sdk/sessions.md +1 -1
- package/sdk/tools.md +2 -0
- package/sdk/v1-runtime.md +22 -7
package/sdk/adapter-contract.md
CHANGED
|
@@ -4,13 +4,25 @@ Adapters create native Codex, Claude Code, or fx conversations in the caller-sup
|
|
|
4
4
|
|
|
5
5
|
Adapter startup receives the selected harness configuration, loaded agent definition, execution directory, environment, and resolved personal access. Startup verifies compatibility and access before submitting any task. A direct CLI harness target supplies no specialist instructions, custom tools, or declared children. Its internal execution configuration may request a native default model; public SDK constructors still require explicit models. Default resolution happens before the first prompt, within the authorized access route, and the resolved model is retained for follow-ups. An unavailable or unverifiable default fails with guidance to supply `--model`, without probing it through a paid generation. Explicit model and effort validation remains effective. `HARNESS_UNAVAILABLE` and `ACCESS_UNAVAILABLE` permit trying another declared alternative before submission. Invalid configuration and unsupported explicit options do not. No adapter performs automatic fallback after turn submission.
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
The CLI readiness check uses this same startup and pre-submission fallback path in an isolated worker, then closes the native session without calling `turn`. It therefore verifies only capabilities observable during startup. It does not probe quota with generation, exercise native tools, or establish shell and child-launch permissions. Startup may create an empty native conversation and private temporary files; normal close removes library-owned temporary files under the adapter lifecycle rules.
|
|
8
|
+
|
|
9
|
+
A readiness check reports success only after the isolated worker has exited and every adapter-owned native process has confirmed termination. Native close first permits graceful shutdown and then uses bounded termination when required. If native termination cannot be confirmed, or worker cleanup fails, the check emits an error instead of a readiness record. A failed worker close preserves its cleanup diagnostic while forcibly ending the isolated worker; it does not leave the CLI waiting indefinitely for that worker.
|
|
10
|
+
|
|
11
|
+
Agent instructions supplement native instructions. The adapter exposes declared tools without implementing its own model loop. Tool input is validated before the function runs, and returned text or JSON data is encoded as a native text result. Explicit `toolResult` values preserve their text and image blocks through the native protocol, as defined in [custom tools](tools.md). Declared subagents authorize delegation. The Claude adapter translates that authorization into session-only native allow rules for the exact session launcher, declared child invocations, and documented coordination commands. The Codex adapter does not install equivalent launcher rules. Explicit constructor options select the session's native sandbox, network, and approval policy as defined in [Native Permissions](permissions.md); omission preserves native settings. The installed native harness and execution environment continue to enforce their restrictions. Native approval or interactive-input requests that still require a response from this library fail with `INPUT_REQUIRED`; no adapter grants them automatically.
|
|
8
12
|
|
|
9
13
|
Codex uses its native App Server protocol. Claude Code uses its native Agent SDK with the installed Claude executable. A Claude session accepts queued native turns and interruption. Its native steering method reports `UNSUPPORTED_DELIVERY` before any mutation; the coordinator converts that request to the documented interrupt operation and reports the effective delivery mode. Codex steering targets the active native turn. Neither adapter synthesizes native recovery by replaying the original prompt; unavailable recovery returns `RECOVERY_UNSUPPORTED`.
|
|
10
14
|
|
|
11
15
|
Personal subscription selection must verify native subscription access. Ambient API keys, alternate endpoints, provider overrides, and native API-key helpers must not silently change the selected billing method. Explicit API or Gateway selections supply only the selected credential route. Credentials are never included in CLI records or diagnostic output.
|
|
12
16
|
|
|
13
|
-
When Claude Code requires a tool approval, `INPUT_REQUIRED` identifies the native tool if its name is a bounded, valid identifier. Tool arguments, command text, settings, and credentials are not included. Other interactive requests retain a generic actionable message.
|
|
17
|
+
When Claude Code requires a tool approval, `INPUT_REQUIRED` identifies the native tool if its name is a bounded, valid identifier. Tool arguments, command text, settings, and credentials are not included. Other interactive requests retain a generic actionable message.
|
|
18
|
+
|
|
19
|
+
When Codex requests approval through a recognized App Server method, `INPUT_REQUIRED` identifies its fixed operation category and fixed native method name:
|
|
20
|
+
|
|
21
|
+
- `item/commandExecution/requestApproval` is `command execution`.
|
|
22
|
+
- `item/fileChange/requestApproval` is `file change`.
|
|
23
|
+
- `item/permissions/requestApproval` is `additional permissions`.
|
|
24
|
+
|
|
25
|
+
Unknown approval and interactive-input methods retain a generic actionable message. The same classification applies during adapter startup and active turns. Codex request parameters, command text, paths, settings, and credentials are never included. These diagnostics identify what requires native configuration without granting approval, answering the native request, or changing permission rules. The adapter still requests native turn interruption for active-turn input and preserves `INPUT_REQUIRED` and exit code `1` when interruption succeeds.
|
|
14
26
|
|
|
15
27
|
Codex custom tools use the native `subharness` namespace while retaining the declared tool map keys. Explicit API/Gateway model identifiers need not appear in a subscription model catalog. Adapters disable provider model fallback and reject a detectable replacement of the requested model; an upstream rejection after submission is an execution failure. Codex explicitly selects standard service when `fast` is false, rather than inheriting a native accelerated preference.
|
|
16
28
|
|
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
# Agent Skill
|
|
2
|
+
|
|
3
|
+
The subharness agent skill teaches an existing coding agent when and how to delegate work with the subharness CLI. It is an ordinary Markdown skill that the native harness reads; it adds no runtime behavior, tools, or commands to subharness.
|
|
4
|
+
|
|
5
|
+
## Installation
|
|
6
|
+
|
|
7
|
+
The skill is installed with the third-party [skills](https://github.com/vercel-labs/skills) installer:
|
|
8
|
+
|
|
9
|
+
```sh
|
|
10
|
+
npx skills add vercel-labs/subharness
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
The installer reads the source repository, so installation requires access to [vercel-labs/subharness](https://github.com/vercel-labs/subharness) while its GitHub visibility is internal. The installer, not subharness, chooses the destination directories for the supported coding agents. Installing the skill does not install the CLI, a native harness, or any credentials; the agent's environment needs the `subharness` executable on PATH or an `npx subharness` fallback.
|
|
14
|
+
|
|
15
|
+
## Location and discovery
|
|
16
|
+
|
|
17
|
+
The skill's single file is `skills/subharness/SKILL.md` at the repository root. Its frontmatter `name` is `subharness`, and its `description` states that the skill applies when the human asks the agent to delegate, parallelize, or hand off work to Codex, Claude Code, or fx, or to follow up on, check, or cancel delegated work.
|
|
18
|
+
|
|
19
|
+
`npx skills add vercel-labs/subharness` offers only this skill. The repository's development skills under `.agents/skills/` set `metadata.internal: true` in their frontmatter, which the installer hides by default. That field does not change how native harnesses read those skills.
|
|
20
|
+
|
|
21
|
+
The skill is not part of the npm package. [Package Distribution](distribution.md) continues to exclude skills from the package contents.
|
|
22
|
+
|
|
23
|
+
## Content
|
|
24
|
+
|
|
25
|
+
The skill is concise, task-oriented guidance for the main agent. It agrees with the [CLI](cli/index.md), [Output](cli/output.md), [Message delivery](message-delivery.md), and [Sessions](sessions.md) references and introduces no commands, flags, defaults, or behavior they do not document. It covers:
|
|
26
|
+
|
|
27
|
+
- When to delegate: independent, well-scoped work that can run while the conversation continues, and when the human names a harness.
|
|
28
|
+
- Choosing a target: the reserved `claude`, `codex`, and `fx` direct targets, and `subharness list` for discovered specialists.
|
|
29
|
+
- Starting work with `subharness run <target> [--cwd <directory>] <prompt>`, `--prompt`, or `--prompt-file`, including quoting a positional prompt and giving the child explicit task context, because children do not receive the main conversation's transcript or loaded skills.
|
|
30
|
+
- Preferring a single `subharness run <target> <prompt>` through the host's background-task controls when that capability is known to be available. The caller collects that hosted command's output later; a separate `--detach` and `wait` sequence is unnecessary for its first response.
|
|
31
|
+
- Using `subharness run <target> --detach <prompt>` when background-command support is unavailable or uncertain, returning after admission and retaining the task identifier. `wait` returns or awaits the first response and runs only after useful independent work, when the caller is ready to wait. The minimal example contains only `run --detach` and `wait`; `status` is a separate optional nonblocking snapshot, not an unconditional step before or after `wait`.
|
|
32
|
+
- Noting that background execution, process survival, and completion notifications depend on the host; detached admission alone does not provide them.
|
|
33
|
+
- Reading results: the returned session, task, and response identifiers; `subharness wait <task-id>` for the first response; `subharness wait <task-id> --after <response-id>` for the next response; and `subharness status <task-id> --full` for the complete latest response.
|
|
34
|
+
- Following up with `subharness send <session-id>` and its `queue`, `steer`, and `interrupt` delivery modes; inspecting work with `subharness queue <session-id>`; and stopping it with `subharness cancel <task-id>`.
|
|
35
|
+
- Reporting results back to the human in the main conversation, rather than relaying raw output.
|
|
36
|
+
- Limits: subharness does not create worktrees or sandboxes, the caller supplies an existing working directory, native authentication stays with each harness, and records do not survive coordinator loss.
|
|
37
|
+
|
|
38
|
+
`subharness --help` and `subharness run <harness> --help` remain the authoritative usage reference; the skill points the agent to them for options it does not cover.
|
package/sdk/agent.md
CHANGED
|
@@ -31,13 +31,13 @@ Descriptions, instructions, and model identifiers must contain text. Names and m
|
|
|
31
31
|
|
|
32
32
|
`AgentDefinition` exposes readonly configuration fields without execution methods. Functions and schemas make definitions unsuitable as a JSON serialization format. Shared definitions do not contain credentials, provider clients, storage, or lifecycle plugins.
|
|
33
33
|
|
|
34
|
-
Instructions supplement native harness guidance and repository instructions. Authors can compose the instruction string with ordinary TypeScript. Each
|
|
34
|
+
Instructions supplement native harness guidance and repository instructions. Authors can compose the instruction string with ordinary TypeScript. Each `.ts` file under a discovered `agents/` directory default-exports one definition. Shared helpers live outside that directory, as described in [agent discovery](config.md).
|
|
35
35
|
|
|
36
36
|
## Reusing skills
|
|
37
37
|
|
|
38
38
|
A skill is a task-specific instruction file, conventionally named `SKILL.md`, that an agent reads when needed. Keep a role's stable responsibility in `instructions` and reference relevant skill paths in the task or repository guidance. This avoids creating another agent definition for every library or technique.
|
|
39
39
|
|
|
40
|
-
|
|
40
|
+
subharness does not expose a `skills` field, install skills, or normalize native skill discovery. The selected harness owns its discovery rules and file-reading tools. A task can explicitly ask the agent to read an accessible skill file using ordinary file reading. A filesystem path is not a portable argument to a harness's native skill command, which may expect a registered skill name instead. Reading a skill does not grant tools, permissions, or subagent access.
|
|
41
41
|
|
|
42
42
|
Children do not inherit loaded skill contents or the parent's transcript. Include the relevant paths and governing contracts in each child's task. The [repository team](project-team.md) demonstrates this convention with a small set of roles and on-demand repository skills.
|
|
43
43
|
|
|
@@ -50,3 +50,7 @@ Children do not inherit loaded skill contents or the parent's transcript. Includ
|
|
|
50
50
|
The package exports `CodexOptions`, `ClaudeCodeOptions`, and `FxOptions` together with their readonly return types `CodexConfig`, `ClaudeCodeConfig`, and `FxConfig`. Runtime validation also enforces Claude's listed effort values when TypeScript checking is absent; unsupported values fail with `INVALID_DEFINITION` before native execution.
|
|
51
51
|
|
|
52
52
|
The `harness` property always uses that name, even for an array. Alternatives are considered in declaration order; an array does not mean parallel execution. [Personal access settings](access-config.md) determine which connections are eligible without editing a shared agent.
|
|
53
|
+
|
|
54
|
+
## Native permission options
|
|
55
|
+
|
|
56
|
+
Harness constructors also accept the native permission options defined in [Native Permissions](permissions.md). Options are typed and validated independently for each harness. Omitted values preserve native settings; explicit values configure only the new native session and remain fixed for follow-ups. There is no permission profile registry.
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# Dashboard presentation
|
|
2
|
+
|
|
3
|
+
This document defines the grouped dashboard presentation. Command lifecycle and history retention are defined in [Live dashboard](index.md#live-dashboard).
|
|
4
|
+
|
|
5
|
+
The dashboard groups agent rows by repository and worktree. Each group has one compact `repository / worktree` label. A repository with multiple worktrees has a separate group for each checkout rather than a shared enclosing section.
|
|
6
|
+
|
|
7
|
+
The dashboard resolves its launch directory to its owning worktree, including when launched from a subdirectory. If that worktree has displayed runs, its group stays first. Launching from a repository's main checkout pins only that checkout. Other worktrees belonging to the same repository receive no extra priority. No empty group is created for a launch directory without displayed runs.
|
|
8
|
+
|
|
9
|
+
All remaining groups are ordered by their most recent agent update, newest first. A group's recency is the latest update among its displayed runs, including finished runs. Active groups do not receive a separate ordering priority. Equal update times retain a stable order. Polling, elapsed-time redraws, and terminal resizing do not count as agent updates.
|
|
10
|
+
|
|
11
|
+
For example, a dashboard launched inside `subharness / trenton` keeps that group first even when another group has newer activity:
|
|
12
|
+
|
|
13
|
+
```text
|
|
14
|
+
subharness / trenton
|
|
15
|
+
...
|
|
16
|
+
|
|
17
|
+
website / landing
|
|
18
|
+
...
|
|
19
|
+
|
|
20
|
+
subharness / main
|
|
21
|
+
...
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Here, `website / landing` has a more recent update than `subharness / main`. Their repository identities do not change that order. This ordering applies to groups; it does not change ordering of agent rows within a group.
|
|
25
|
+
|
|
26
|
+
## Context and update semantics
|
|
27
|
+
|
|
28
|
+
A session captures its checkout context when admitted. Git subdirectories resolve to the canonical worktree root; linked worktrees share the canonical repository location while separate clones remain separate repositories. Outside Git, the canonical working directory is the group identity and label. Git discovery has a bounded timeout and falls back to directory context on failure. Context is retained for the session and its finished runs; later branch changes or directory deletion do not relabel history. The dashboard resolves its own launch directory once, independently of the coordinator. Context discovery never runs on every refresh.
|
|
29
|
+
|
|
30
|
+
Labels use repository and checkout directory names, not branch names. Ambiguous short labels include a distinguishing parent path. All labels and task titles are sanitized before terminal output.
|
|
31
|
+
|
|
32
|
+
Each task records its latest coordinator-observed update: admission, dispatch or continuation, a selected-harness change, accepted steering, a complete native response, or a settled lifecycle outcome. Token streaming and internal tool activity are not observed by this display. Snapshot reads, response reads, and cancelling an already settled task do not advance recency. A queued follow-up that has no displayed row does not affect group ordering until displayed. Equal group update times preserve the group's first-observed order in the open dashboard.
|
|
33
|
+
|
|
34
|
+
## Row presentation
|
|
35
|
+
|
|
36
|
+
Each group label is followed by two-space-indented rows. Blank lines separate groups. Rows use aligned marker, harness, title, state, and elapsed-duration columns with whitespace rather than dashed leaders. Running uses `●`, queued and waiting use `◌`, failed and interrupted use `!`, and completed and cancelled use a blank marker. State text always accompanies the marker. There are no animated indicators or color requirements.
|
|
37
|
+
|
|
38
|
+
Elapsed durations retain the CLI's `MM:SS` and `H:MM:SS` notation. Terminal rows also show whole-minute completion age (`<1m ago` initially, then `1m ago`, etc.). Completion age is omitted in narrow terminals before reducing the title; titles truncate by grapheme display width, and remaining fixed fields clip only when they cannot fit. Group labels use the same width and terminal-control protections. The interface has no application banner, column headings, footer, selection, scrolling, or additional controls. Viewport clipping applies to the resulting grouped lines. Current rows precede history within each group, but no global active-first ordering overrides checkout pinning or group recency.
|
|
39
|
+
|
|
40
|
+
The private snapshot declares grouped-context support in addition to finished-history support. A dashboard connected to a coordinator without either capability reports `COORDINATOR_OUTDATED` with restart guidance rather than silently omitting grouping or history. The legacy `INVALID_ARGUMENT` response for an unknown coordinator operation also identifies a coordinator without dashboard support and receives the same diagnostic. Other error records and malformed responses remain `COORDINATOR_UNAVAILABLE`. It never restarts the coordinator itself.
|
package/sdk/cli/index.md
CHANGED
|
@@ -4,15 +4,18 @@ The CLI is the primary execution interface for an everyday coding assistant. The
|
|
|
4
4
|
|
|
5
5
|
```sh
|
|
6
6
|
subharness list [--cwd <directory>]
|
|
7
|
-
subharness
|
|
8
|
-
subharness
|
|
9
|
-
subharness
|
|
7
|
+
subharness dashboard
|
|
8
|
+
subharness check <target> [--cwd <directory>] [--format text|jsonl]
|
|
9
|
+
subharness check <harness> [--cwd <directory>] [--model <model>] [--effort <effort>] [--format text|jsonl]
|
|
10
|
+
subharness run <target> [--cwd <directory>] [--detach] <prompt>
|
|
11
|
+
subharness run <target> [--cwd <directory>] [--detach] --prompt <text>
|
|
12
|
+
subharness run <target> [--cwd <directory>] [--detach] --prompt-file <path|->
|
|
10
13
|
subharness run <harness> [--model <model>] [--effort <effort>] <prompt>
|
|
11
14
|
subharness run <harness> --help
|
|
12
|
-
subharness send <session-id> [--delivery queue|steer|interrupt] <prompt>
|
|
13
|
-
subharness send <session-id> [--delivery queue|steer|interrupt] --prompt <text>
|
|
14
|
-
subharness send <session-id> [--delivery queue|steer|interrupt] --prompt-file <path|->
|
|
15
|
-
subharness wait <task-id> --after <response-id>
|
|
15
|
+
subharness send <session-id> [--delivery queue|steer|interrupt] [--detach] <prompt>
|
|
16
|
+
subharness send <session-id> [--delivery queue|steer|interrupt] [--detach] --prompt <text>
|
|
17
|
+
subharness send <session-id> [--delivery queue|steer|interrupt] [--detach] --prompt-file <path|->
|
|
18
|
+
subharness wait <task-id> [--after <response-id>]
|
|
16
19
|
subharness status <task-id> [--full]
|
|
17
20
|
subharness queue <session-id>
|
|
18
21
|
subharness cancel <task-id>
|
|
@@ -21,7 +24,35 @@ subharness resume <session-id>
|
|
|
21
24
|
|
|
22
25
|
Execution commands accept `--format text|jsonl`; text is the default. `--help` and `--version` are plain-text information modes without a `--format` option. `subharness --help` and `<command> --help` show usage without requiring execution arguments or starting execution. Direct harness help includes that harness’s options; other command help shows the general usage. [Output records and exit codes](output.md) define the integration format.
|
|
23
26
|
|
|
24
|
-
`list` includes the built-in harness targets and discovered definitions without starting harnesses or calling models. Inclusion is not a claim that an executable, eligible account, or particular model is available. Definition loading retains its existing validation and trust requirements. `run` starts a new session and its first task, immediately prints their identities, and waits for one complete response or terminal outcome. It returns control even if that task still has delegated work. The result contains a response identifier and task state.
|
|
27
|
+
`list` includes the built-in harness targets and discovered definitions without starting harnesses or calling models. Inclusion is not a claim that an executable, eligible account, or particular model is available. Definition loading retains its existing validation and trust requirements. `check` performs the selected target's native startup checks in an isolated worker, closes the temporary native session, confirms that the worker and adapter-owned native processes have stopped, and returns without submitting a turn or admitting a task. Cleanup or unconfirmed termination is a failed check, not a readiness result. `run` starts a new session and its first task, immediately prints their identities, and waits for one complete response or terminal outcome. It returns control even if that task still has delegated work. The result contains a response identifier and task state.
|
|
28
|
+
|
|
29
|
+
Readiness means only that startup checks for the selected target passed at that moment. It does not guarantee remote quota, task success, shell or child-launch permissions, or continued availability, and it does not check declared descendants. A check may create an empty native conversation and temporary native or launcher files. Loading existing personal access settings can also update Git's local exclude file as described in [personal access configuration](../access-config.md).
|
|
30
|
+
|
|
31
|
+
`run --detach` instead returns immediately after successful admission with only the same `started` record.
|
|
32
|
+
|
|
33
|
+
## Live dashboard
|
|
34
|
+
|
|
35
|
+
`subharness dashboard` opens a read-only live terminal display for the current local coordinator, covering all of its managed directories, with the calling checkout pinned first when it has displayed runs. It includes direct harness sessions, specialists, and delegated agents managed by that coordinator. It does not discover unmanaged native processes or coordinators belonging to other installations or state directories.
|
|
36
|
+
|
|
37
|
+
The list groups runs by repository/worktree, following the context and ordering rules in [Dashboard presentation](dashboard-design.md). Within each group, it shows current work followed by finished runs. Current work has one row per session in session creation order, representing its active task or its first queued task when dispatch has not started. Running and waiting tasks remain visible. A failed active task remains among current work while it pauses the session, including when native stop is unconfirmed. Independently queued follow-ups do not create additional current-work rows.
|
|
38
|
+
|
|
39
|
+
Finished runs have one row per retained completed, failed, cancelled, or interrupted task that is not already shown among current work. They appear newest finished first, with later task admission first when finish times are equal. Follow-up tasks keep their own history rows instead of replacing previous results from the same session. Finished runs appear even when they ended before the dashboard opened. History lasts for the coordinator’s lifetime; restarting the coordinator does not restore earlier runs. Compact repository/worktree labels identify each group; there are no additional controls.
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
subharness / trenton
|
|
43
|
+
● codex Implement the validation running 02:14
|
|
44
|
+
◌ claude Review the current diff waiting 00:38
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
The harness label is `codex`, `claude`, or `fx`. A direct target is known during startup; a specialist displays `pending` until its actual native harness is selected. Selection reports the successfully opened harness, including fallback selection. The title is the first nonempty line of the current task's original prompt, with whitespace normalized and terminal control sequences removed, bounded to 120 Unicode code points before fitting it to the display. An empty sanitized title is `(untitled)`. No model call generates titles. Steering retains the task title; a follow-up uses its own prompt.
|
|
48
|
+
|
|
49
|
+
Time is elapsed wall time since that task first dispatched, including waiting and recovery. A queued task shows `00:00`. Completed, cancelled, and interrupted tasks freeze their elapsed time when that outcome is reached. Failed tasks freeze their elapsed time at failure and resume counting from the original start if explicitly recovered. Tasks cancelled before first dispatch show `00:00`. Repeated observation or cancellation that leaves a settled terminal outcome unchanged does not change its recorded finish time. If stopping previously unconfirmed execution reaches a new outcome, timing records that new outcome. Durations use `MM:SS` below one hour and `H:MM:SS` thereafter. State and elapsed time occupy separate aligned columns. Finished rows include completion age when space permits, as defined in [Dashboard presentation](dashboard-design.md). Narrow terminals truncate titles by display width without splitting grapheme clusters; when the fixed fields alone do not fit, the row is clipped to the available columns.
|
|
50
|
+
|
|
51
|
+
The display contains compact group labels and agent rows: no application heading, border, footer, spinner, key hints, or empty-state message. It uses the alternate screen and restores the normal screen and cursor on exit. When stdin is a terminal, keyboard echo is suppressed and input other than Ctrl+C is ignored; the previous input mode is restored on exit. Non-terminal stdin is not consumed. Rows fit within the terminal viewport without wrapping or scrolling; excess rows remain outside the visible viewport until the terminal is enlarged or changes in row order bring them into view. Current work takes precedence over finished history within each group; group order determines which groups fit in the visible viewport. There is no navigation, selection, or task control. Ctrl+C exits with code `130`; SIGTERM exits with code `143`. Exiting never cancels agents.
|
|
52
|
+
|
|
53
|
+
Snapshots refresh once per second without overlapping requests. Resizing redraws the current snapshot immediately. Unchanged visible rows are not rewritten. Rendering uses native Node streams and terminal escape sequences, with a focused Unicode width utility rather than a UI framework. Slow output does not accumulate unbounded redraws.
|
|
54
|
+
|
|
55
|
+
The command requires terminal stdout and a terminal supporting cursor control; non-terminal output and `TERM=dumb` fail with `INVALID_ARGUMENT` and exit `2` before connecting or emitting terminal escapes. It accepts no positional arguments or options except `--help`, which follows general command help behavior. In particular, it has no `--format` or `--cwd` option. Errors use ordinary text stderr after restoring the terminal. It does not load agent definitions, start a coordinator, start a harness, or submit work. If no coordinator is initially reachable, the list stays blank and checks again once per second so agents started later can appear. After a successful connection, a failed, invalid, or timed-out snapshot exits with an error instead of presenting stale data. The snapshot explicitly identifies support for finished-run history and grouped checkout context. A coordinator without either capability, including one that returns the legacy unknown-operation error for `dashboard`, causes `COORDINATOR_OUTDATED` and exit `1`, with its process ID and guidance to stop that background coordinator after active work ends. Restarting only the dashboard does not reload coordinator code. Coordinator restart discards retained in-memory history; the dashboard never restarts it automatically. Snapshot requests have a five-second timeout.
|
|
25
56
|
|
|
26
57
|
## Direct harnesses and specialists
|
|
27
58
|
|
|
@@ -32,23 +63,24 @@ subharness run claude "Review the current diff."
|
|
|
32
63
|
subharness run codex --cwd ../feature-worktree "Implement the documented validation."
|
|
33
64
|
subharness run fx --model "provider/model" "Compare the proposed implementations."
|
|
34
65
|
subharness run repo:reviewer "Review the current diff."
|
|
66
|
+
subharness check repo:reviewer --format jsonl
|
|
35
67
|
```
|
|
36
68
|
|
|
37
69
|
Replace `provider/model` with an available Gateway model identifier. Nonreserved bare names retain the existing unambiguous specialist lookup. Qualified `repo:`, `global:`, and `subagent:` names retain their existing meanings. A specialist named `claude`, `codex`, or `fx` requires its scope qualifier; it never shadows the built-in target. Unknown targets fail rather than being executed as arbitrary commands. Omitting the target is an argument error; the CLI never chooses the first installed harness.
|
|
38
70
|
|
|
39
71
|
Direct harness sessions use native instructions and project context without adding a specialist role, custom tools, or declared children. They preserve the existing authentication, native permission, queue, cancellation, follow-up, and response contracts. They do not broaden the caller's native permissions or automatically authorize additional delegation.
|
|
40
72
|
|
|
41
|
-
`--model` and `--effort` configure direct harness targets only
|
|
73
|
+
`--model` and `--effort` configure direct harness targets only for `run` and `check`, and can be combined with `--cwd` and, for `run`, any supported prompt source. Empty values are errors. Specialist targets reject these overrides and retain their declared harness configurations. `send` retains the session configuration and does not accept model or effort overrides.
|
|
42
74
|
|
|
43
75
|
Omitting `--model` requests the native default for the authorized access route at session creation. The adapter retains the selected model for that conversation. It does not inherit the calling agent's model, rank models, choose a substitute, or change billing routes. An unavailable or unverifiable native default fails with actionable guidance to supply `--model`; it does not submit a prompt merely to discover a default. Explicit model identifiers retain native validation and substitution checks. TypeScript harness constructors continue to require a model.
|
|
44
76
|
|
|
45
77
|
Effort is harness-specific: Codex and fx accept native effort identifiers, while Claude Code accepts `low`, `medium`, `high`, `xhigh`, or `max`. Explicit effort must be compatible with the selected model where native capabilities expose that validation. An omitted effort retains the native default at session creation. Detected changes to an explicit request are errors. Codex and Claude Code retain standard-speed execution; this interface adds no fast-mode flag. fx retains its documented native preference and Gateway access contract.
|
|
46
78
|
|
|
47
|
-
|
|
79
|
+
Direct-harness help for `run` and `check` describes the relevant options and access requirements without loading definitions, starting a coordinator, or invoking a harness. Native CLI flags are not forwarded. Unsupported flags are errors.
|
|
48
80
|
|
|
49
|
-
`wait` observes the
|
|
81
|
+
`wait` observes the first retained response when `--after` is omitted, returning it immediately when available or waiting for it. If the task ends without a response, it returns the terminal outcome. Omission never means to start observing from the current time or to return the latest response. With `--after`, `wait` observes the next response after the supplied identifier. It returns an already-available response immediately or waits for another response or terminal outcome. It creates no work, consumes no responses, and does not restart the harness. Multiple readers can use independent cursors. Unknown or cross-task response identifiers are errors.
|
|
50
82
|
|
|
51
|
-
`status` returns a nonblocking snapshot with a bounded latest-response preview; `--full` retrieves the complete latest response. `queue` shows active and pending work in order. `cancel` waits for cancellation of the targeted task and its delegated descendants, without removing independently queued tasks. `resume` requests native recovery of a failed task without new input; unsupported recovery reports an error and leaves the queue paused.
|
|
83
|
+
`status` returns a nonblocking snapshot with a bounded latest-response preview; `--full` retrieves the complete latest response. `queue` shows active and pending work in order. When `send` admits a task while execution failure has paused the session queue, the immediate `started` output identifies the failed active task that blocks dispatch and shows the `status --full` and `cancel` commands for it. The new task remains accepted and queued. `cancel` waits for cancellation of the targeted task and its delegated descendants, without removing independently queued tasks. Successfully cancelling the settled failed task releases the queue; failed or unconfirmed cancellation leaves dispatch paused. `resume` requests native recovery of a failed task without new input; unsupported recovery reports an error and leaves the queue paused.
|
|
52
84
|
|
|
53
85
|
## Input and delivery
|
|
54
86
|
|
|
@@ -60,8 +92,10 @@ Effort is harness-specific: Codex and fx accept native effort identifiers, while
|
|
|
60
92
|
|
|
61
93
|
Native steering returns an acceptance acknowledgement. Queued and interrupting sends return their task's first complete response or terminal outcome; they do not wait for the whole queue. There is no expected-task guard. If B starts before a correction is handled, steering targets B.
|
|
62
94
|
|
|
95
|
+
`--detach` is available for `run` and task-creating `send` commands with positional, `--prompt`, and `--prompt-file` input. After successful admission, it ends that command's observation and exits `0` with exactly the existing `started` record; it does not wait for a response or terminal outcome. Admission errors retain their normal records and exit codes. Later startup and execution failures are retrieved with `wait` or `status`, so detached exit `0` means only that admission succeeded. A detached interrupting send still waits for confirmed stop before the replacement is admitted. `--detach --delivery steer` is invalid and exits `2`, because steering uses acknowledgement semantics rather than creating a task to observe. Detachment is not retained in session configuration and does not affect later commands.
|
|
96
|
+
|
|
63
97
|
## Nested execution
|
|
64
98
|
|
|
65
99
|
A managed agent invokes its own declared children with `subharness run subagent:<name>`. Parent context is supplied internally, not through public flags. Names resolve only against that parent's direct declarations. This invocation loads the parent's definition source without discovering unrelated repository or global entries, so an unrelated broken definition does not block a declared child. A generic parent has no declared children and fails this lookup without evaluating a specialist catalog. Missing context or an undeclared child is an error; there is no fallback to another scope. Child execution uses the same commands and lifecycle.
|
|
66
100
|
|
|
67
|
-
The caller supplies an existing working directory. The library does not create worktrees or sandboxes. A local coordinator owns pending execution after an individual response command exits.
|
|
101
|
+
The caller supplies an existing working directory. The library does not create worktrees or sandboxes. A local coordinator owns pending execution after an individual response command exits. `--detach` releases only the CLI observation; it does not override process-lifecycle restrictions imposed by the host. Use host-owned background controls when the coordinator and workers must remain reachable after a shell command exits. Pending execution is not guaranteed to survive coordinator or environment exit, and detached admission does not promise automatic reactivation of an external parent chat.
|
package/sdk/cli/output.md
CHANGED
|
@@ -1,8 +1,20 @@
|
|
|
1
1
|
# CLI Output
|
|
2
2
|
|
|
3
|
+
`dashboard` is a terminal-only live view, not an execution-record command. It accepts no `--format` option and emits no text/JSONL records on success. Its rows, empty state, and terminal lifecycle are defined in the [CLI contract](index.md#live-dashboard).
|
|
4
|
+
|
|
3
5
|
Execution commands accept `--format text|jsonl`; `text` is the default. Both formats report the same operations. JSONL contains complete records separated by newlines, never native token streams or tool transcripts. All records include `version: 1` and a `type` discriminator. Identifiers are opaque strings prefixed with `ses_`, `tsk_`, or `rsp_` and are not paths or process identifiers.
|
|
4
6
|
|
|
5
|
-
Task-creating commands print a `started` record immediately after admission, including when queued.
|
|
7
|
+
Task-creating commands print a `started` record immediately after admission, including when queued. By default, they then return one complete response or terminal outcome for that task. With `--detach`, `run` and task-creating `send` commands end their observation after the `started` record and return no response or terminal record. A successful `steer` prints an `accepted` record. If native steering is unavailable, the operation uses `interrupt` semantics and prints `started` with `requestedDelivery: "steer"` and `delivery: "interrupt"`, then returns the replacement task's response or outcome.
|
|
8
|
+
|
|
9
|
+
When a task is admitted while its session queue is paused by the failed active task, its `started` record includes `paused: true` and `blockedByTaskId` with that failed task's identifier. Both fields describe the admission snapshot and are omitted from other `started` records. Text output says that the task was accepted but cannot dispatch, identifies the failed task, and shows `subharness status <failed-task-id> --full` and `subharness cancel <failed-task-id>` as the inspection and recovery commands. Cancellation must succeed before queue dispatch resumes.
|
|
10
|
+
|
|
11
|
+
A successful readiness check emits one record and admits no task:
|
|
12
|
+
|
|
13
|
+
```jsonl
|
|
14
|
+
{"version":1,"type":"check","target":"repo:developer","ready":true}
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Text output is a short line identifying the ready target. A failed check emits the existing `error` record. Access and capability failures exit `1`; invalid arguments, configuration, or targets exit `2`.
|
|
6
18
|
|
|
7
19
|
Quiet tasks keep their local observation connection alive without printing progress messages or extra JSONL records. Transport keepalives are private to the CLI/coordinator connection. A lost observer connection does not cancel or replay its task; callers can retrieve its state and response with `status`.
|
|
8
20
|
|
|
@@ -11,9 +23,15 @@ Quiet tasks keep their local observation connection alive without printing progr
|
|
|
11
23
|
{"version":1,"type":"response","sessionId":"ses_123","taskId":"tsk_456","responseId":"rsp_789","state":"waiting","text":"Does the limit apply per user?"}
|
|
12
24
|
```
|
|
13
25
|
|
|
26
|
+
A paused admission is represented as:
|
|
27
|
+
|
|
28
|
+
```jsonl
|
|
29
|
+
{"version":1,"type":"started","sessionId":"ses_123","taskId":"tsk_queued","state":"queued","paused":true,"blockedByTaskId":"tsk_failed"}
|
|
30
|
+
```
|
|
31
|
+
|
|
14
32
|
The identifiers above are illustrative. A `response` contains `sessionId`, `taskId`, `responseId`, `state`, and `text`. Its state describes the task at response creation. Current state is available through `status`. A `terminal` record contains `sessionId`, `taskId`, `state`, and an optional `error` object. It is returned when an observation ends without another response.
|
|
15
33
|
|
|
16
|
-
`status` emits a `status` record with `sessionId`, `taskId`, `state`, and optional `response` and `error`. The response preview contains `responseId`, `text`, and `truncated`; at most 4,000 text characters are included. `subharness status <task-id> --full` returns the complete latest response instead.
|
|
34
|
+
`status` emits a `status` record with `sessionId`, `taskId`, `state`, and optional `response` and `error`. The response preview contains `responseId`, `text`, and `truncated`; at most 4,000 text characters are included. `subharness status <task-id> --full` returns the complete latest response instead. `subharness wait <task-id>` returns or awaits the first retained response; adding `--after <response-id>` returns or awaits the next response after that cursor.
|
|
17
35
|
|
|
18
36
|
`list` emits an `agents` record with an `agents` array of `{ id, name, description, scope }`, where scope is `harness`, `repo`, `global`, or `subagent`. The built-in entries have IDs and names `codex`, `claude`, and `fx`, and scope `harness`; they appear in that order before discovered specialists. They describe supported targets, not verified executable, authentication, or model availability. A managed parent's catalog includes its declared `subagent:` entries; generic parents have no declared children. `queue` emits a `queue` record with `sessionId`, `paused`, optional `active`, and a `tasks` array in pending order. Task summaries contain `taskId`, `state`, and a prompt `description` limited to 120 characters.
|
|
19
37
|
|
|
@@ -21,7 +39,7 @@ An `accepted` record contains `sessionId`, `taskId`, `delivery: "steer"`. A succ
|
|
|
21
39
|
|
|
22
40
|
Errors use `{ "version": 1, "type": "error", "error": { "code": "INVALID_ARGUMENT", "message": "..." } }` with task/session identifiers when known. JSONL errors appear on stdout; text errors appear on stderr. Native stderr and credentials are not forwarded. Text responses show identifiers and state above the complete response text, with a `wait --after` hint only while the task is pending. Text layout is intended for reading; integrations should use JSONL.
|
|
23
41
|
|
|
24
|
-
Exit code `0` means the command succeeded or a response was returned; it does not mean the agent fulfilled the objective. Code `1` means execution, access, or capability failure; `2` means invalid input, configuration, or unknown identifiers; and `130` means an observation ended in cancellation or interruption. Successful `cancel` itself exits `0`.
|
|
42
|
+
Exit code `0` means the command succeeded or a response was returned; it does not mean the agent fulfilled the objective. For `check`, it means only that the selected target's startup checks passed at that moment. For `run --detach` and detached task-creating sends, it means admission succeeded, not that native startup or execution succeeded. Admission errors retain their existing error records and exit codes; later failures are available through `wait` or `status`. Code `1` means execution, access, or capability failure; `2` means invalid input, configuration, or unknown identifiers; and `130` means an observation ended in cancellation or interruption. Successful `cancel` itself exits `0`. Combining `--detach` with `--delivery steer` is an `INVALID_ARGUMENT` error and exits `2`.
|
|
25
43
|
|
|
26
44
|
`status` snapshots retain the observed task’s outcome in their exit code: queued, running, waiting, and completed states exit `0`; cancelled or interrupted states exit `130`; failed states use `1` or `2` according to the stored error. A retrieval error uses its own error code.
|
|
27
45
|
|
|
@@ -4,21 +4,23 @@ Returning a response, completing a task, and starting another caller model turn
|
|
|
4
4
|
|
|
5
5
|
The adapter captures complete responses automatically. Agents do not need a reporting tool, special JSON format, or a classifier that recognizes questions. A question is delivered through the same response mechanism as any other text.
|
|
6
6
|
|
|
7
|
-
## Waiting for
|
|
7
|
+
## Waiting for a response
|
|
8
8
|
|
|
9
9
|
```sh
|
|
10
|
-
subharness wait <task-id> --after <response-id>
|
|
10
|
+
subharness wait <task-id> [--after <response-id>]
|
|
11
11
|
```
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
Without `--after`, the command returns the first retained response or waits for it. If the task ends without a response, it returns the terminal outcome. Omitting the cursor never means to start observing from the current time or to return the latest response.
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
With `--after`, the command returns the next retained response after the supplied identifier. If none exists, it waits for another complete response or a terminal outcome. If the task has ended without another response, it returns that outcome rather than waiting indefinitely.
|
|
16
|
+
|
|
17
|
+
An explicit cursor prevents losing a later response that arrives between commands. Reads do not consume responses. Every reader uses its own cursor position, including the implicit position before the first response when `--after` is omitted. Unknown and cross-task identifiers are errors. `status --full` retrieves the complete latest response.
|
|
16
18
|
|
|
17
19
|
## Completion with delegated work
|
|
18
20
|
|
|
19
21
|
A task remains pending while descendants are unfinished or their results still require processing. Ending a native turn does not cancel children or release independently queued tasks. Child results continue the parent in the same native conversation and library task. Normal completion requires the descendants to finish and the parent to produce a response after processing their results.
|
|
20
22
|
|
|
21
|
-
For example, a developer starts a reviewer and asks whether a limit applies per user or organization. The command returns that question with a pending state; the reviewer continues. The caller can steer an answer into the pending developer task and use `wait
|
|
23
|
+
For example, a developer starts a reviewer and asks whether a limit applies per user or organization. The command returns that question with a pending state; the reviewer continues. The caller can steer an answer into the pending developer task and use `subharness wait <task-id> --after <question-response-id>` for another response. If no descendants remain, an ordinary complete response can finish the task and release its queue.
|
|
22
24
|
|
|
23
25
|
A child result returned by a child CLI command to an actively running parent is already delivered through that parent's native tool interaction. It is not duplicated in another continuation. Results arriving after an earlier pending response was returned are retained and delivered through a continuation. Results arriving while a parent is busy are held until they can be delivered without interrupting it.
|
|
24
26
|
|
|
@@ -27,3 +29,5 @@ A child result returned by a child CLI command to an actively running parent is
|
|
|
27
29
|
Managed parents can continue through their harness adapters. An everyday assistant outside the library depends on its host's background-task and notification behavior. Completing a shell command makes its result available; it does not guarantee that every host will start another model turn.
|
|
28
30
|
|
|
29
31
|
A caller can run `run`, task-creating `send`, or `wait` through its own background-task controls and handle the resulting response. The library does not claim to reactivate arbitrary external conversations after their caller has ended its turn. The coordinator still owns pending work after an individual response-reading command exits.
|
|
32
|
+
|
|
33
|
+
For explicit background delegation, `run --detach` returns after admission with a task identifier. The caller can later use `wait <task-id>` to return or await the first retained response, or `status <task-id>` for a nonblocking snapshot. Detached admission ends only the initiating observation; it does not cancel work or configure future observations.
|
package/sdk/config.md
CHANGED
|
@@ -1,19 +1,41 @@
|
|
|
1
1
|
# Agent Discovery
|
|
2
2
|
|
|
3
|
-
Repository agents live in `.
|
|
3
|
+
Repository agents live in `.subharness/agents/` at the Git worktree root. Global agents live in `~/.subharness/agents/`. The `.subharness/` directory holds subharness configuration only. Cross-harness skills conventionally remain in `.agents/skills/`, which subharness does not read.
|
|
4
4
|
|
|
5
5
|
```text
|
|
6
6
|
project/
|
|
7
|
-
.
|
|
7
|
+
.subharness/
|
|
8
8
|
agents/
|
|
9
|
-
developer.
|
|
10
|
-
reviewer.
|
|
11
|
-
|
|
12
|
-
|
|
9
|
+
developer.ts
|
|
10
|
+
reviewer.ts
|
|
11
|
+
tools/
|
|
12
|
+
uppercaseTool.ts
|
|
13
13
|
agents.local.json
|
|
14
14
|
```
|
|
15
15
|
|
|
16
|
-
|
|
16
|
+
Every `.ts` file under `agents/`, including files in nested directories, is a catalog entry that default-exports one `AgentDefinition`. Type declaration files ending in `.d.ts` are ignored. Other extensions are ignored. The file name does not determine the agent's name. There is no registry file or registration command. Symlink directories are not traversed. An invalid entry fails discovery with its source filename; it is not silently skipped.
|
|
17
|
+
|
|
18
|
+
Shared code such as custom tools, instruction fragments, and constants belongs outside `agents/`, for example in `.subharness/tools/`. Definitions import these modules with ordinary relative imports. subharness never loads files outside `agents/` during discovery, and `tools/` is a convention rather than a discovered directory. A helper module placed inside `agents/` is treated as a definition and fails discovery when it does not default-export one.
|
|
19
|
+
|
|
20
|
+
A definition imports shared modules and other definitions, such as a declared subagent, with relative paths:
|
|
21
|
+
|
|
22
|
+
```ts
|
|
23
|
+
// .subharness/agents/developer.ts
|
|
24
|
+
import { agent, codex } from "subharness";
|
|
25
|
+
import reviewer from "./reviewer.js";
|
|
26
|
+
import { uppercaseTool } from "../tools/uppercaseTool.js";
|
|
27
|
+
|
|
28
|
+
export default agent({
|
|
29
|
+
name: "developer",
|
|
30
|
+
description: "Implements features and verifies changes.",
|
|
31
|
+
instructions: "Follow the repository's documented requirements and run its checks.",
|
|
32
|
+
harness: codex({ model: "CODEX_MODEL_ID" }),
|
|
33
|
+
tools: { uppercase: uppercaseTool },
|
|
34
|
+
subagents: { reviewer },
|
|
35
|
+
});
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
The former `.agents/agents/` directory is not read. Definitions moved to `.subharness/agents/` can keep or drop an `.agent` file-name suffix; either way the file is an entry because it is a `.ts` file under `agents/`.
|
|
17
39
|
|
|
18
40
|
Discovery is required for listing and repository/global specialist lookup. Direct harness execution skips it. Declared-child execution loads only the parent's source and direct child map; unrelated catalog entries are not dependencies of a `subagent:` invocation.
|
|
19
41
|
|
|
@@ -21,14 +43,17 @@ The declared name determines identity. Duplicate names in the same scope are err
|
|
|
21
43
|
|
|
22
44
|
```sh
|
|
23
45
|
subharness list --cwd /repo/worktree
|
|
46
|
+
subharness check repo:developer --cwd /repo/worktree
|
|
24
47
|
subharness run repo:developer --cwd /repo/worktree --prompt "Implement the documented feature."
|
|
25
48
|
```
|
|
26
49
|
|
|
27
50
|
An omitted `--cwd` uses the calling command's directory. Inside nested repositories, the nearest Git worktree root is used. Outside Git, the supplied directory is the project root. Bare repositories are not execution directories. Definitions are trusted executable TypeScript: catalog loading evaluates modules.
|
|
28
51
|
|
|
52
|
+
`check` resolves direct harnesses, repository and global specialists, unambiguous bare names, and declared children under the same rules as `run`. Direct harness checks bypass catalog discovery. Specialist checks evaluate the selected definition in an isolated worker, so they retain the same trust requirement and use the selected directory for definition loading and native startup. A check covers only that selected definition, not its declared children.
|
|
53
|
+
|
|
29
54
|
## Personal access and worktrees
|
|
30
55
|
|
|
31
|
-
The optional [access file](access-config.md) is `.
|
|
56
|
+
The optional [access file](access-config.md) is `.subharness/agents.local.json` in the repository's main checkout, shared by its linked worktrees. This means the local checkout associated with those worktrees, not a branch named `main` or another clone of the same remote. Worktree agent definitions and execution directories still come from the selected worktree.
|
|
32
57
|
|
|
33
58
|
| Resource | Source |
|
|
34
59
|
| --- | --- |
|
package/sdk/distribution.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
The npm package name is `subharness`. One package provides the TypeScript SDK and the `subharness` executable and its `agent` compatibility alias. Agent definitions import public helpers from `subharness`. The website workspace remains private and is not part of the npm package.
|
|
4
4
|
|
|
5
|
-
The source repository is [vercel-labs/subharness](https://github.com/vercel-labs/subharness). Its GitHub visibility is internal, so cloning requires repository access. The product name is
|
|
5
|
+
The source repository is [vercel-labs/subharness](https://github.com/vercel-labs/subharness). Its GitHub visibility is internal, so cloning requires repository access. The product name is subharness. The primary CLI command is `subharness`; `agent` remains an identical compatibility alias. Agent discovery and personal configuration live under `.subharness/`.
|
|
6
6
|
|
|
7
7
|
Node.js 22.18 or newer is required. The SDK is ESM and includes TypeScript declarations. The native Codex, Claude Code, and fx executables remain external prerequisites; installing this package does not install harnesses or configure their credentials.
|
|
8
8
|
|
|
@@ -12,20 +12,34 @@ The public npm package is `subharness`, with an initial release version of `0.0.
|
|
|
12
12
|
|
|
13
13
|
For project-local use, install `subharness` with `npm install --save-dev subharness` and invoke its CLI with `npx subharness`. This also makes SDK imports resolve from repository agent definitions. A global CLI installation alone does not make SDK imports resolve from project-local definitions. Global definitions need an SDK dependency reachable from their own directory under normal Node package resolution.
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
Source development uses pnpm 11.20.0, pinned in the root `packageManager` field. `pnpm-workspace.yaml` includes the root SDK and `apps/docs`, with one committed `pnpm-lock.yaml`. Run `pnpm install --frozen-lockfile` from the repository root, then `pnpm run build`. Run the local CLI with `node dist/cli/main.js`; this requires no global link. Direct harness targets do not need a consumer SDK dependency. For use in another project, build a tarball with `pnpm pack` and install that tarball in the consumer so its agent definitions can resolve SDK imports.
|
|
16
16
|
|
|
17
|
-
|
|
17
|
+
Repository scripts use pnpm, including workspace command forwarding and content-generation lifecycle hooks. Dependency build scripts are explicitly allowed only for the native build helpers required by the installed SDK and website dependencies. Package consumers can use any compatible npm package manager; the clean-consumer check deliberately installs with npm to verify that the published artifact does not require pnpm.
|
|
18
|
+
|
|
19
|
+
A tarball built from a source checkout can differ from the registry release with the same version. The current package version is `0.0.3`, producing `subharness-0.0.3.tgz`. Install a local tarball with `npm install --save-dev /absolute/path/to/subharness-0.0.3.tgz`. Run the project's local executable with `npx subharness`.
|
|
18
20
|
|
|
19
21
|
## Package contents and release checks
|
|
20
22
|
|
|
21
|
-
|
|
23
|
+
Packages built from this source contain the built JavaScript, declaration files and source maps under `dist/`, SDK Markdown documentation under `sdk/`, the README, the Apache License 2.0, and package metadata. Source maps embed the original TypeScript so debuggers can display it without a separate source checkout. Website code, development agents, skills, source assets, tests, `.context`, personal settings, and environment files are excluded.
|
|
22
24
|
|
|
23
|
-
Packing builds the SDK from the current source before assembling its files. The CLI's `--version` output matches the package version. A clean consumer must be able to import the SDK and discover a TypeScript agent definition using the installed executable without access to the source checkout or a paid model call. `
|
|
25
|
+
Packing builds the SDK from the current source before assembling its files. The CLI's `--version` output matches the package version. A clean consumer must be able to import the SDK and discover a TypeScript agent definition using the installed executable without access to the source checkout or a paid model call. `pnpm run check:package` packs with pnpm and verifies this clean-consumer behavior and the packaged file boundary without running a coding model. It requires network access to install dependencies from the public npm registry and disables install lifecycle scripts in the temporary consumer.
|
|
24
26
|
|
|
25
|
-
|
|
27
|
+
The committed lockfile pins dependency content with integrity hashes and does not embed company registry URLs. pnpm resolves packages through the configured registry, which defaults to the public npm registry. Installation and package verification do not change the developer's global npm registry configuration. Registry authentication remains with npm.
|
|
26
28
|
|
|
27
29
|
Releases use the public npm registry, public access, and the `latest` distribution tag. A release publishes the verified tarball for its package version. The matching source commit is tagged `v<version>` in the source repository. Publishing the package does not change the repository's internal visibility.
|
|
28
30
|
|
|
31
|
+
## GitHub release publishing
|
|
32
|
+
|
|
33
|
+
Publishing a stable GitHub release runs `.github/workflows/release.yml`. The release tag must have the form `vX.Y.Z`, with three nonnegative integers and no leading zeroes. Drafts and GitHub pre-releases do not publish to npm. Release-candidate tags and other tag formats are rejected. Pushing a tag alone does not publish a package.
|
|
34
|
+
|
|
35
|
+
The workflow checks out the commit associated with the release event. The tag supplies the npm version: for example, `v0.0.3` publishes `subharness@0.0.3`. The workflow sets that version in its temporary checkout before building and testing. It does not commit a version bump or change the tag. The version in a development checkout therefore need not match the latest registry release.
|
|
36
|
+
|
|
37
|
+
One GitHub-hosted job installs the frozen pnpm dependencies, builds the SDK, checks TypeScript, runs the behavioral tests with bounded concurrency, and verifies the package in a clean consumer. Setting `SUBHARNESS_PACK_DESTINATION` to an absolute directory when running `pnpm run check:package` retains the verified tarball there; temporary consumer files are still removed. The workflow publishes that exact tarball to npm with public access and the `latest` tag. Failed checks prevent publication. An existing npm version cannot be overwritten; rerunning an already successful publication fails rather than replacing it. Maintainers publish releases in increasing version order and wait for each release job to finish before publishing another.
|
|
38
|
+
|
|
39
|
+
Authentication uses npm trusted publishing through GitHub Actions OIDC, without an `NPM_TOKEN` secret. In the npm settings for `subharness`, the GitHub Actions trusted publisher must use organization `vercel-labs`, repository `subharness`, and workflow filename `release.yml`, with no environment name. Direct `npm publish` must be enabled in its allowed actions. The workflow requires `contents: read` and `id-token: write` permissions. Provenance is disabled while the source repository is internal, because npm provenance requires a public source repository.
|
|
40
|
+
|
|
41
|
+
To release, choose a new stable tag on `main` that includes this workflow, enter the release notes, leave the GitHub pre-release option unchecked, and publish the GitHub release. The corresponding Actions run reports whether npm publication succeeded. No separate version-bump commit or release-candidate channel is required.
|
|
42
|
+
|
|
29
43
|
## License
|
|
30
44
|
|
|
31
|
-
|
|
45
|
+
The current source checkout and packages built from it are distributed under the Apache License, Version 2.0. The already-published `subharness@0.0.1` registry release remains under the MIT license. The root [LICENSE](../LICENSE) contains the complete Apache License text and is included in packages built from this source.
|