cowork-harness 0.32.0 → 0.33.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 0.32.0
7
- tracks-harness: cowork-harness 0.32.0 (baseline desktop-1.20186.1)
6
+ version: 0.33.0
7
+ tracks-harness: cowork-harness 0.33.0 (baseline desktop-1.20186.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.32.0` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.33.0` (baseline
26
26
  > `desktop-1.20186.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -38,10 +38,22 @@ cowork-harness skill ./my-skill "do X" # run the skill once against the sta
38
38
  Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.32.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.32.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.32.0"`. **Pin `@>=0.32.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published. (≥ 0.32.0 is what gates the commands/assertions this skill teaches: `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes`, batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, `verify-cassettes --allow-domain`/`--allow-email`/`--allow-path`/`--allow-patterns-file` (`path` — local absolute filesystem paths — is the scanner's 4th class, new in 0.21.0; `--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, `/help` in the REPL, `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field, `computer_links_resolve` (new in 0.22.0), `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass; new in 0.27.0), **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly; new in 0.28.0), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, the session-level `agent_env` knob, and resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`; new in 0.30.0), the `lint-skill` (static host-loop footgun + `subagent_type` resolution linter) / `analyze-skill` (advisory `/sessions`-path static scan, `--strict`, `analyze-skill: ignore` marker) / `probe-dispatch` (single-dispatch mechanics probe) commands, `status --latest-for` (resolve a scenario's newest run dir by run time, not directory mtime), the `subagent_dispatch_healthy` composite assertion, and the persisted `result.json` fields `verdict`, `subagents[].referencesRead`/`subagents[].reasoning`, plus the `toolCounts`/`toolErrors`/`toolDurations` shape distinction (all new in 0.31.0), `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict` (new in 0.32.0).)
41
+ - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.33.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.33.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.33.0"`. **Pin `@>=0.33.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
+
44
+ What the ≥ 0.33.0 floor gates, by release:
45
+
46
+ - **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
47
+ - **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
48
+ - **0.22.0:** `computer_links_resolve`.
49
+ - **0.28.0:** `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass), and **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly).
50
+ - **0.30.0:** resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, and the session-level `agent_env` knob.
51
+ - **0.31.0:** the `lint-skill` (static host-loop footgun + `subagent_type` resolution linter) / `analyze-skill` (advisory `/sessions`-path static scan, `--strict`, `analyze-skill: ignore` marker) / `probe-dispatch` (single-dispatch mechanics probe) commands, `status --latest-for` (resolve a scenario's newest run dir by run time, not directory mtime), the `subagent_dispatch_healthy` composite assertion, and the persisted `result.json` fields `verdict`, `subagents[].referencesRead`/`subagents[].reasoning`, plus the `toolCounts`/`toolErrors`/`toolDurations` shape distinction.
52
+ - **0.32.0:** `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict`.
53
+ - **0.33.0:** the `redacted` marker on display-omitted reasoning — `subagents[].reasoning` and the top-level `thinking[]` now carry `{text:"", redacted:true}` when the model returns a signed-but-empty thought (so "reasoned, text omitted" is distinct from "no thought"), plus the fenced `debug.thinking_display` escape hatch.
42
54
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
43
55
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
44
- - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred) or `ANTHROPIC_API_KEY`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
56
+ - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
45
57
  - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip.
46
58
 
47
59
  ## Orient — the three loops
@@ -240,12 +252,12 @@ Full model in `references/scenario-schema.md`.
240
252
  Don't hand-write the YAML from memory — that's how invented keys (`assertions:` vs `assert:`,
241
253
  `json_file`, `answer_policy`) creep in. Start from the bundled generator, which emits the
242
254
  known-good skeleton (right tier, scripted `answers:` + `on_unanswered: fail`, content assertions
243
- separated from live-only ones, one concern per item) and **self-lints its own output**:
255
+ separated from live-only ones, one concern per item) and **self-lints its own output**. The
256
+ generator is the bundled `scripts/scenario.py` — installed as a plugin, point `S` at
257
+ `${CLAUDE_PLUGIN_ROOT}/scripts/scenario.py`; from a repo checkout, use the literal path below:
244
258
 
245
259
  ```bash
246
- S="${CLAUDE_PLUGIN_ROOT}/scripts/scenario.py"
247
- # Working from a repo checkout instead of an installed plugin?
248
- # S=".claude/skills/cowork-harness/scripts/scenario.py"
260
+ S=".claude/skills/cowork-harness/scripts/scenario.py"
249
261
  python3 "$S" scaffold --name report-check --skill ./skills/report-gen \
250
262
  --prompt "Generate the weekly report to outputs/report.md." \
251
263
  --content 'weekly report' --artifact outputs/report.md \
@@ -372,7 +384,7 @@ than a stuck `"running"`.)
372
384
  ### Place assertions in the right CI lane
373
385
 
374
386
  CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
375
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v1` (a packaged GitHub Action with a
387
+ (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@main` (a packaged GitHub Action with a
376
388
  PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
377
389
  the four-stage pipeline.
378
390
 
@@ -1,17 +1,20 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 0.32.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 0.33.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
7
7
 
8
8
  ```yaml
9
- - uses: yaniv-golan/cowork-harness@v1
9
+ - uses: yaniv-golan/cowork-harness@main
10
10
  with:
11
11
  command: replay
12
12
  path: cassettes/
13
13
  ```
14
14
 
15
+ The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
+ release; pin an exact version (e.g. `version: "0.33.0"`) for reproducible CI.
17
+
15
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
16
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
17
20
  independently instead of as one action run per command).
@@ -29,13 +32,13 @@ jobs:
29
32
  - uses: actions/checkout@v4
30
33
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
31
34
  run: |
32
- V=2.1.197 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.205 # match your scenario's pinned baseline's agentVersion
33
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
34
37
  chmod +x "$RUNNER_TEMP/claude-$V"
35
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
36
39
  # before trusting it — see docs/maintenance.md's "Agent-binary provenance" section.
37
40
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
38
- - uses: yaniv-golan/cowork-harness@v1
41
+ - uses: yaniv-golan/cowork-harness@main
39
42
  with:
40
43
  command: run
41
44
  path: scenarios/
@@ -54,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
54
57
  GitHub-hosted runners, no token/Docker/agent:
55
58
 
56
59
  ```yaml
57
- - run: npm i -g "cowork-harness@>=0.32.0"
60
+ - run: npm i -g "cowork-harness@>=0.33.0"
58
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
59
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
60
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -174,7 +177,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
174
177
 
175
178
  ## GitHub Actions sketch
176
179
 
177
- The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v1` does
180
+ The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@main` does
178
181
  in one step (see the top of this doc) — reach for this form when you need independent per-command
179
182
  gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
180
183
  equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
@@ -194,7 +197,7 @@ jobs:
194
197
  with: { node-version: '20' }
195
198
  - uses: actions/setup-python@v5
196
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
197
- - run: npm i -g "cowork-harness@>=0.32.0"
200
+ - run: npm i -g "cowork-harness@>=0.33.0"
198
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
199
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
200
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -223,7 +226,7 @@ jobs:
223
226
  echo "live=true" >> "$GITHUB_OUTPUT"
224
227
  fi
225
228
  - if: steps.guard.outputs.live == 'true'
226
- run: npm i -g "cowork-harness@>=0.32.0"
229
+ run: npm i -g "cowork-harness@>=0.33.0"
227
230
  - if: steps.guard.outputs.live == 'true'
228
231
  run: cowork-harness run scenarios/ --output-format json
229
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 0.32.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 0.33.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -13,10 +13,10 @@ Self-contained reference. Tracks `cowork-harness 0.32.0` (baseline `desktop-1.20
13
13
  | `cowork` | Auto-picks `hostloop` vs `container` the way Cowork itself does for the synced release | "Do what real Cowork does for this release." |
14
14
 
15
15
  - `hostloop` / `cowork` are the production-faithful path; `container` is the practical default.
16
- - Boundary assertions (`egress_*`, `expect_denied`) are enforced at `container`, `microvm`, and
17
- `hostloop` (`container`'s and `hostloop`'s `bash` share the same Docker sandbox + egress proxy —
18
- `hostloop`'s native file tools run with no container at all, gated instead by a path-containment hook).
19
- Only `protocol` is rejected.
16
+ - Boundary assertions (`egress_*`, `expect_denied`) are enforced at `container`, `microvm`, `hostloop`,
17
+ and `cowork` (`cowork` auto-resolves to a sandboxed tier; `container`'s and `hostloop`'s `bash` share
18
+ the same Docker sandbox + egress proxy — `hostloop`'s native file tools run with no container at all,
19
+ gated instead by a path-containment hook). Only `protocol` is rejected.
20
20
  - A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
21
21
  true` — with no container around the native file tools, that combination gives the agent genuine,
22
22
  software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
@@ -176,7 +176,7 @@ and `scaffold` refuse to treat a partial run's half-finished output as a passing
176
176
  - `COWORK_HARNESS_RUNS_DIR` (or `--run-dir <path>`) — override the default run-output root `~/.cowork-harness/runs` (out of any working tree). flag > env > default.
177
177
  - `COWORK_HARNESS_DECIDER_CMD_TIMEOUT_MS` / `COWORK_HARNESS_LLM_TIMEOUT_MS` — decider backstops
178
178
  (default 600 s; **fail loud** on timeout).
179
- - `COWORK_HARNESS_DECIDER_DIR_POLL_MS` / `_TIMEOUT_MS` — the `--decider-dir` rendezvous.
179
+ - `COWORK_HARNESS_DECIDER_DIR_POLL_MS` / `_TIMEOUT_MS` — the `--decider-dir` rendezvous (poll defaults: 300 ms for the run-side rendezvous, 500 ms for `gates --follow`).
180
180
  - `COWORK_HARNESS_DIALOG_TIMEOUT_MS` — dialog auto-cancel (default 6 s).
181
181
  - `COWORK_HARNESS_LLM_MAX_BYTES` — stdout bound on `--decider-llm` (default 8 MiB).
182
182
  - `COWORK_HARNESS_LLM_RETRIES` — bounded retries for a transient non-zero `claude -p` exit on the
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.32.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.33.0`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
@@ -37,7 +37,7 @@ fidelity: container # protocol | container | microvm | hostloop
37
37
  execution: local # OPTIONAL — orthogonal to fidelity (a privilege/sandbox tier, all
38
38
  # local): local (default) | cloud-describe (RESERVED — no runner
39
39
  # exists yet; authoring it is a load-time error, not a silent no-op)
40
- on_unanswered: fail # policy for unscripted gates: fail | prompt | first | llm
40
+ on_unanswered: fail # policy for unscripted gates: fail | prompt | first | llm — run rejects prompt
41
41
  # ("agent" is retired — no longer a valid value)
42
42
 
43
43
  prompt: | # the user turn
@@ -312,7 +312,7 @@ same set live from the schema.
312
312
  | `egress_allowed: <host>` | the host was allowed through |
313
313
  | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
314
314
  | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
315
- | `semantic_matches: {rubric: [...], min_pass?, judge_model?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose (authored-file evidence is unavailable on the **microvm** tier — no pre-run manifest to diff against, so no files are captured; `container`/`hostloop` do capture them). Beyond the microvm case, when the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
315
+ | `semantic_matches: {rubric: [...], min_pass?, judge_model?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose (authored-file evidence is unavailable on the **microvm** tier — no pre-run manifest to diff against, so no files are captured; `container`/`hostloop` do capture them). Beyond the microvm case, when the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
316
316
  | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
317
317
  | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) |
318
318
  | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** |
@@ -329,7 +329,7 @@ paraphrases (that re-records red). Structured JSON → assert it in YAML with **
329
329
  `path` + operator); use the pytest lane (`assert_artifact_json`) only for predicates too complex for a
330
330
  dotted path.
331
331
 
332
- **VerdictSignals in `result.signals`:** `computeVerdict` pushes signals into `result.signals`; most
332
+ **VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; most
333
333
  are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
334
334
  `"success"`) and only three are **warn**-severity (informational, never flip pass/fail). Current signal
335
335
  codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
@@ -353,7 +353,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
353
353
 
354
354
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
355
355
  overall run verdict and exit code — `assert result: success` alone won't catch it; check
356
- `result.signals[].severity` or the run's exit code. Only the three **warn** codes are truly benign.
356
+ `result.verdict.signals[].severity` or the run's exit code. Only the three **warn** codes are truly benign.
357
357
 
358
358
  ## Replay class
359
359
 
@@ -491,7 +491,7 @@ reference omits). Neither list is a strict superset of the other — reach for t
491
491
  the stochastic path flags the run `nonDeterministic`. The LLM decider is one mechanism, two
492
492
  spellings: `on_unanswered: llm` (YAML) and `--decider-llm` (CLI). The bare `--on-unanswered llm`
493
493
  is rejected (use `--decider-llm`). `agent` is **retired** — `on_unanswered: agent` is rejected by
494
- the schema. (`src/types.ts:365` — the `on_unanswered` enum; `src/cli.ts:899` — the CLI-side
494
+ the schema. (`src/types.ts` — the `on_unanswered` enum; `src/cli.ts:899` — the CLI-side
495
495
  `--on-unanswered` value check.)
496
496
 
497
497
  4. **`--on-unanswered first` is non-deterministic too** — it picks option 1 and is flagged
package/CHANGELOG.md CHANGED
@@ -6,6 +6,127 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.33.0] — 2026-07-13
10
+
11
+ ### Fixed
12
+
13
+ - **Empty thinking was surfaced as if the model hadn't reasoned — both `subagents[].reasoning` and the
14
+ top-level `thinking[]`.** The 0.31.0 sub-agent field captured TEXT turns fine but every `thinking` turn
15
+ came through as `{kind:"thinking", text:""}`, and the older top-level `thinking[]` field does the same
16
+ on current-gen models — a consumer couldn't tell "reasoned but text unavailable" from "no thought."
17
+ Root cause (binary-verified against the staged 2.1.205 agent, corpus-corroborated): it is a
18
+ **request-side display mode**, not a persist-time strip — the API's `thinking.display` resolves to
19
+ `"omitted"` (empty thinking text + a signature) for sub-agent turns always, and for the MAIN loop too
20
+ on models whose API default is `"omitted"` (Opus 4.8, Sonnet 5; Sonnet 4.6 defaulted to `"summarized"`).
21
+ Corpus: main-loop thinking text is 1810/1810 on Sonnet-4.6 but **0/2 on Opus-4.8 and 0/15 on Sonnet-5**,
22
+ every empty carrying a signature; sub-agents 0/230. The harness passes no `--thinking-display` —
23
+ faithfully to real Cowork, whose spawn passes none either. Both fields now mark such blocks
24
+ `{text:"", redacted:true}` (`kind:"thinking"` for the sub-agent turn shape) so they read as "reasoned,
25
+ text omitted by request"; TEXT turns are never redacted, and `redacted` is omitted (not `false`) on
26
+ blocks that carry text.
27
+ - **New `debug.thinking_display` escape hatch** (fenced, non-Cowork — like `debug.max_thinking_tokens`).
28
+ Set it to `"summarized"` to emit `--thinking-display summarized`, which flips **both** loops to
29
+ summarized thinking text (the API returns no raw chain-of-thought — `summarized` is the ceiling). It
30
+ diverges from real Cowork (which passes no such flag) and adds token cost, so it is a debug-only opt-in;
31
+ the default stays `"omitted"`, byte-identical argv. `docs/subagents.md`, `docs/session.md`,
32
+ `schema/run-result.json`, and the `RunResult`/`SessionConfig` types state the real capture semantics;
33
+ the released 0.31.0 note is corrected in place.
34
+
35
+ - **Docs pointed the GitHub Action at a ref that doesn't exist.** README, `SKILL.md`, and
36
+ `references/ci-recipe.md` all said `uses: yaniv-golan/cowork-harness@v1`, but no `v1` tag has ever
37
+ been published — a copy-pasted workflow failed with "unable to resolve action" before running
38
+ anything. All six references now bind `@main`; a moving major tag is a 1.0.0 question. A new
39
+ token-free guard (`test/action-docs-sync.test.ts`) locks the ref policy and also validates that
40
+ every documented `with:` key is a real `action.yml` input and every documented `command:` value is
41
+ one the Action describes.
42
+ - **`llms.txt` command list was three commands stale** — `lint-skill`, `analyze-skill`, and
43
+ `probe-dispatch` were missing. Now lists all 29, locked by a new COMMANDS ↔ llms.txt sync test
44
+ (set-equality, so stale names fail too), and the README "Commands at a glance" guard gained the
45
+ reverse direction (a README-only row now fails).
46
+ - **README observability prose still called a `subagents[]` field `model`** — renamed to
47
+ `dispatchModel` (beside `resolvedModel`) in 0.30.0; the prose now names the real fields. Also
48
+ fixed: the effort enum shorthand now lists `xhigh` (`extra` is the accepted alias) in both places,
49
+ the CI stage table matches `ci.yml`'s declaration order, the exit-code summary notes
50
+ `verify-cassettes`' distinct exit-`3` meaning and the reserved exit `4`, a duplicated `--keep`
51
+ example was collapsed, and `on_unanswered: llm` is explicitly marked as a scenario-YAML key (not a
52
+ CLI flag).
53
+ - **`status --latest-for` was undocumented in its own guide** — now covered in `docs/run-status.md`
54
+ (with the real output shape), the `docs/README.md` index row, and the top-level `--help` status
55
+ entry (with its own `0` found / `2` not found exit semantics).
56
+ - **`probe-dispatch --help` didn't start with a `usage:` line** like every other command; it does
57
+ now.
58
+ - **`debugging.md` showed only the scenario run-dir layout** — it now also names `chat`'s
59
+ `runs/chat/<sessionId>/` path. `docs/session.md`'s "all four tiers" now spells out the four
60
+ execution tiers and that `fidelity: cowork` resolves to one of them. `examples/README.md`'s
61
+ answer-policies pointer is now a real link. `effectiveFidelity` (the recorded resolved tier behind
62
+ the `resolved-tier`/`unverifiable-tier` staleness classes) is now mentioned in the README fidelity
63
+ section and `python/README.md`, not only in `docs/cassette.md`.
64
+ - **The docs told a false "default `8899`" story for `COWORK_VM_PROXY_PORT`** — when the var is
65
+ unset, the host binds the egress proxy on an OS-assigned free port and threads that value into the
66
+ guest firewall + `HTTP(S)_PROXY`; `8899` is only the guest-config fallback when a VM is spawned
67
+ without an explicit port, which never happens on a normal run. README and `docs/scenario.md` now
68
+ say so. Also closed: `COWORK_HARNESS_STATUS_CORRUPT_TIMEOUT_MS` (30s corrupt-`status.json`
69
+ backstop on `status --follow`) was the one `COWORK_*` env var documented nowhere — now in
70
+ `docs/run-status.md` + the README knob list; the `COWORK_HARNESS_DECIDER_DIR_POLL_MS` docs name
71
+ the per-subsystem defaults (300 ms rendezvous / 500 ms `gates --follow`); the `semantic_matches`
72
+ judge docs name the pinned default (`claude-opus-4-8`) in README, `docs/scenario.md`, and the
73
+ skill's schema reference; the heartbeat bullet states its 30s default instead of pointing at a
74
+ source file.
75
+ - **`replay --explain` was invisible from the command catalog** — the flagship false-green
76
+ diagnosis flag was documented only in debugging prose; it's now in the `record`/`replay` command
77
+ table row and leads the record/replay "Flags worth knowing" bullet.
78
+ - **`examples/README.md` ships in the npm tarball while the trees it describes don't** — it now
79
+ opens with a "Reading this on npm?" callout (clone for `scenarios/`/`sessions/`/`skills/`/`data/`),
80
+ documents `probes/` (previously invisible: used by `test/live-contract.test.ts` but absent from
81
+ the layout and from schema validation — `examples/probes` is now in `test/examples.test.ts`'s
82
+ scenario sweep), and links the matrices worked example; the README documentation-table row carries
83
+ the same source-checkout caveat.
84
+ - **The companion skill's preflight sent replay-only users through `doctor`**, whose token check
85
+ hard-fails on every tier — a new "Replay-only? Skip `doctor`" bullet carries the same carve-out
86
+ `docs/README.md` already had. The 3,200-character single-bullet 0.32.0 feature parenthetical is
87
+ now a scannable by-release sub-list, with every release bucket tag-verified against git history —
88
+ which caught and fixed a long-standing wrong tag: `semantic_matches` was labeled "new in 0.27.0"
89
+ but shipped in 0.28.0; the five path-gate keys, `agent_env`, and the `hostloop` split/
90
+ `allow_host_writes:` now sit under their real releases (0.30.0 / 0.21.0) too.
91
+ - **`DESIGN.md` still said the staged in-VM agent is 2.1.202** — the `desktop-1.20186.1` baseline
92
+ re-synced it to 2.1.205; DESIGN.md and `docs/protocol.md` now carry dated patch-only notes for
93
+ 1.20186.1 while the 2026-07-11 live pass stays scoped to `desktop-1.20186.0` (not restamped).
94
+
95
+ ### Added
96
+
97
+ - **`test/docs-index-sync.test.ts`** — three token-free guards: every `COWORK_*` env var read in
98
+ `src/` (dot-access **or** helper-read string literal) must be documented in README/`docs/*.md`;
99
+ the judge-model default id in the docs must match `semantic-judge.ts` (fails loud on a const
100
+ rename); `llms.txt` must link every top-level `docs/*.md` guide (both directions — `gotchas.md`,
101
+ `subagents.md`, and `plugin-root.md` were missing and are now linked).
102
+
103
+ - **The Action's `version` default (`latest`) is now documented as intentional** — in `action.yml`'s
104
+ input description, the README inputs sentence, and `references/ci-recipe.md` — with
105
+ pin-an-exact-version guidance for reproducible CI, and why it deliberately differs from the
106
+ companion skill's `@>=0.32.0` floor for ad-hoc CLI installs.
107
+ - **`--help` and unknown-flag structural-guard test coverage now spans all 29 commands** (previously
108
+ 16 and 17 respectively); `lint`/`lint-skill` are asserted against their `scenario.py` passthrough
109
+ usage lines and skip cleanly when `python3` is absent.
110
+
111
+ ### Documentation
112
+
113
+ - **`on_unanswered: prompt` was described as "only valid for `chat`" — wrong on two counts.** `chat`
114
+ never reads a scenario YAML (it runs an inline interactive scenario), and `prompt` is really a
115
+ `skill`-command policy (the adaptive-TTY default, or explicit `skill --on-unanswered prompt`);
116
+ `run` rejects it. The schema `.describe()` (and regenerated `scenario.schema.json`) now say so.
117
+ - **`ANTHROPIC_AUTH_TOKEN` is an accepted auth source but was under-documented.** It resolves
118
+ identically to `ANTHROPIC_API_KEY` (used only when no OAuth token is set) and was already in
119
+ `.env.example`, but the README auth text, the `record`/`doctor` `--help` blurbs, the `doctor`
120
+ no-token detail/remedy, the two `record` credential messages, `docs/cassette.md`, and the companion
121
+ skill's Auth note named only the other two — all now list it as the third alternative.
122
+ - **`llms.txt` exit-code line** now flags that `3`/`1` carry per-command meanings (e.g.
123
+ `verify-cassettes` `3` = could-not-verify, `sync` hard-fail = `1`) and links the authoritative
124
+ SPEC §11 text, instead of implying one global meaning for `3`.
125
+ - **README** gains a platform × tier support matrix (making explicit that Linux live runs are
126
+ `container`-only), doc-index rows for the spawn contract and `docs/decisions/`, the
127
+ `probe-dispatch` fidelity set, and a clearer global-install-ships-`examples/replays/`-only warning;
128
+ `docs/chat.md` notes that scaffolding a `chat` run yields an empty `assert:` block.
129
+
9
130
  ## [0.32.0] — 2026-07-13
10
131
 
11
132
  ### Added
@@ -125,6 +246,11 @@ All notable changes to this project are documented here. The format is based on
125
246
  `thinking[]` field is (~50 entries, ~10KB/entry each), with `reasoningElided` counting the overflow. A
126
247
  missing or malformed child transcript never fails the run — the affected dispatch's `reasoning` is just
127
248
  left absent.
249
+ - **Correction (see Unreleased):** in practice only the TEXT turns carry content — the harness's
250
+ non-interactive spawn forces the API's `thinking.display` to `"omitted"` for sub-agent turns, so
251
+ `thinking` turns come through empty (signature-only). The Unreleased `redacted` marker distinguishes
252
+ "reasoned, text omitted by request" from "no thought"; the "THINKING … turns" wording above
253
+ overstated what this captures by default.
128
254
 
129
255
  - **`status --latest-for <scenario-name-or-slug>`** — resolves and prints the NEWEST run dir for a
130
256
  scenario by actual run time, replacing the fragile `ls -td runs/<scenario>/* | head -1` idiom: bare
package/CONTRIBUTING.md CHANGED
@@ -2,6 +2,8 @@
2
2
 
3
3
  Thanks for helping make Cowork skill-testing reproducible.
4
4
 
5
+ For the full documentation map, see [docs/README.md](./docs/README.md).
6
+
5
7
  ## Development setup
6
8
 
7
9
  ```bash
package/DESIGN.md CHANGED
@@ -38,7 +38,7 @@ flowchart TB
38
38
  A Cowork session is the Desktop app driving an agent that runs **inside an Apple Virtualization.framework microVM**:
39
39
 
40
40
  - VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
41
- - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.202**, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
41
+ - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.205**, per `baselines/desktop-1.20186.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
42
42
  - Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
43
43
  - Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
44
44
 
@@ -165,7 +165,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
165
165
 
166
166
  > **Machine-readable form:** the five shapes below are schema'd as `schema/protocol.v1.json`, with a golden vector pack at `fixtures/protocol/v1/` — see [docs/protocol.md](./docs/protocol.md) for scope, versioning, and how to conformance-test against them.
167
167
 
168
- ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS build 2.1.177+; the staged in-VM agent that L1/L2 run is **2.1.202**, the native host app that `hostloop` runs is **2.1.205**, baseline **`desktop-1.20186.0`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-07-11, superseding the prior `1.19367.0` pin; the intervening `1.19367.0`→`1.20186.0` bump was a minifier re-anchor with a byte/behaviourally-identical value-resolved spawn contract, now confirmed live)
168
+ ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS build 2.1.177+. The staged in-VM agent that L1/L2 ran at the time of this pass was **2.1.202** (the `desktop-1.20186.1` baseline has since re-synced the staged VM ELF to **2.1.205** — see the note below), the native host app that `hostloop` runs is **2.1.205**, baseline **`desktop-1.20186.0`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-07-11, superseding the prior `1.19367.0` pin. The intervening `1.19367.0`→`1.20186.0` bump was a minifier re-anchor with a byte/behaviourally-identical value-resolved spawn contract, now confirmed live.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
169
169
 
170
170
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
171
171