cowork-harness 1.0.6 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.0.6
7
- tracks-harness: cowork-harness 1.0.6 (baseline desktop-1.21459.0)
6
+ version: 1.1.0
7
+ tracks-harness: cowork-harness 1.1.0 (baseline desktop-1.21459.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0.6` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.1.0` (baseline
26
26
  > `desktop-1.21459.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0.6**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.0.6" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.0.6"`. **Pin `@>=1.0.6`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.1.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.1.0"`. **Pin `@>=1.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
- What the ≥ 1.0.6 floor gates, by release:
44
+ What the ≥ 1.1.0 floor gates, by release:
45
45
 
46
46
  - **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
47
47
  - **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
@@ -52,6 +52,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
52
52
  - **0.32.0:** `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict`.
53
53
  - **0.33.0:** the `redacted` marker on display-omitted reasoning — `subagents[].reasoning` and the top-level `thinking[]` now carry `{text:"", redacted:true}` when the model returns a signed-but-empty thought (so "reasoned, text omitted" is distinct from "no thought"), plus the fenced `debug.thinking_display` escape hatch.
54
54
  - **1.0.0:** first stable release — the SPEC §12 compatibility contract takes effect (covered CLI/schema/env/Action surfaces are now stable; breaking changes need a major bump). No new author-facing command; the floor simply tracks the 1.0 release.
55
+ - **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
55
56
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
56
57
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
57
58
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
@@ -277,7 +278,8 @@ quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `
277
278
  (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
278
279
  baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
279
280
  `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
280
- unverifiable), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
281
+ unverifiable), a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) off the
282
+ `container` tier (ERROR on `protocol`/`microvm`/`hostloop`, WARN on `cowork`), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
281
283
  and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
282
284
  (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
283
285
  emits a scenario `lint` would reject.
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.0.6` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.1.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.0.6"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.1.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.0.6"
60
+ - run: npm i -g "cowork-harness@>=1.1.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.0.6"
200
+ - run: npm i -g "cowork-harness@>=1.1.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.0.6"
229
+ run: npm i -g "cowork-harness@>=1.1.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.0.6` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.1.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -171,6 +171,36 @@ writes a `result.json` (marked `partial: true`) with the artifacts the agent pro
171
171
  the work isn't discarded — then exits 2. Inspect it with `cowork-harness inspect <run-dir>`. `verify-run`
172
172
  and `scaffold` refuse to treat a partial run's half-finished output as a passing result.
173
173
 
174
+ ## What a green does NOT prove
175
+
176
+ A passing run is evidence for exactly what it checked, not a blanket certificate. Three gaps come
177
+ up often enough to spell out:
178
+
179
+ - **A green `replay` proves "same as when recorded," not "correct today."** `replay` never touches
180
+ a filesystem or network — it re-evaluates assertions from the frozen cassette. A fixed set of
181
+ keys is live-only and **skipped outright** on replay (absent from `assertions[]`, not vacuously
182
+ passed): `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
183
+ `egress_allowed`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, and `expect_denied`.
184
+ Everything else that *is* evaluated is checked against the **recording**, not fresh behavior — a
185
+ green replay says the skill produced these events when it was recorded, not that it still does
186
+ (`staleness[]` flags skill/baseline drift as a hint; only a live `run` re-confirms current
187
+ behavior). See `docs/cassette.md` § "Still skipped on replay" and `docs/scenario.md` § "Which
188
+ assertions survive replay."
189
+ - **Container-only assertions can't verify off the `container` tier.** `no_scratchpad_leak` and
190
+ `present_files_called` check the `present_files` delivery path, which is served **only** on
191
+ `container` — not `hostloop`/`microvm`. Asserting them off-container hard-fails at runtime (a red
192
+ run, not a false green), so you won't be fooled if you write the assertion. The quieter trap is a
193
+ scenario that runs at `hostloop`/`microvm`/`protocol` and simply omits these assertions: a green
194
+ run there proves **nothing** about scratchpad-leak safety or present_files delivery, because that
195
+ tier never exercises the delivery path the assertions would check. Use `fidelity: container` for
196
+ present_files/scratchpad-delivery coverage.
197
+ - **The harness doesn't observe rendered-artifact interactions, browser downloads, or human
198
+ clicks.** It runs the agent headless — no webview, no browser, no person clicking "Submit." A
199
+ class of Cowork bug (a client-side write-back to a relative URL that resolves-but-fails against
200
+ Cowork's own origin, or a broken blob-download fallback) is invisible to any live run, however
201
+ faithfully sandboxed, because it only manifests in a rendered DOM a human is driving. See
202
+ `docs/fidelity-gaps.md` § "Browser↔webview↔human-interaction boundary."
203
+
174
204
  ## Relevant environment variables
175
205
 
176
206
  - `COWORK_HARNESS_RUNS_DIR` (or `--run-dir <path>`) — override the default run-output root `~/.cowork-harness/runs` (out of any working tree). flag > env > default.
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0.6`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.1.0`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
@@ -17,6 +17,8 @@ Two subcommands:
17
17
  lint flags (see references/scenario-schema.md for the why of each):
18
18
  E egress assertion on `fidelity: protocol` (the harness rejects this run)
19
19
  E `transcript_no_host_path` on hostloop/protocol (fails BY DESIGN at those tiers)
20
+ E `no_scratchpad_leak`/`present_files_called` off (container-only: served only at fidelity:
21
+ protocol/microvm/hostloop container; off-container = cannot-verify)
20
22
  E `requires_capabilities` on `fidelity: protocol` (probe can't run → hard-fails
21
23
  unless allow_missing_capability)
22
24
  E on_unanswered: agent / invalid value (schema rejects `agent`)
@@ -24,6 +26,8 @@ lint flags (see references/scenario-schema.md for the why of each):
24
26
  E `assertions:` instead of `assert:` (block ignored → every check no-ops)
25
27
  W `transcript_no_host_path` on `fidelity: cowork` (tier resolves per baseline gate —
26
28
  incompatible if it lands hostloop)
29
+ W `no_scratchpad_leak`/`present_files_called` on (tier resolves per baseline gate — cannot-verify
30
+ `fidelity: cowork` if it resolves off-container)
27
31
  W no content assertion → no-op on a replay gate (every assertion is fs/egress)
28
32
  W mixed-class assert item → fs/egress half dropped on replay
29
33
  W unknown top-level / assertion key (typo or hallucinated schema)
@@ -127,6 +131,10 @@ LIVE_ONLY_KEYS = {
127
131
  "semantic_matches",
128
132
  }
129
133
  EGRESS_KEYS = {"egress_denied", "egress_allowed"}
134
+ # container-only: served only at fidelity: container (present_files / the scratchpad promotion path
135
+ # it depends on). Off-container these report cannot-verify, not a meaningful pass/fail — same tier-fidelity
136
+ # class as transcript_no_host_path below, just the opposite direction (container-only vs container-hostile).
137
+ CONTAINER_ONLY_KEYS = {"no_scratchpad_leak", "present_files_called"}
130
138
  # verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
131
139
  VERDICT_MODIFIER_KEYS = {
132
140
  "allow_permissive_auto_allow",
@@ -388,6 +396,41 @@ def lint_doc(doc, path, raw_lines):
388
396
  )
389
397
  )
390
398
 
399
+ # E/W: CONTAINER_ONLY_KEYS (no_scratchpad_leak / present_files_called) are served only at
400
+ # fidelity: container — off-container they report cannot-verify, not a meaningful check. Mirrors
401
+ # the transcript_no_host_path tier check above, just the opposite direction: ERROR on the tiers
402
+ # where the runtime deterministically can't serve it, WARN on cowork (baseline-gate-resolution
403
+ # dependent — the linter stays offline, so the message names the dependency instead of resolving it).
404
+ container_only_present = sorted(assert_keys & CONTAINER_ONLY_KEYS)
405
+ if container_only_present:
406
+ if fidelity in ("protocol", "microvm", "hostloop"):
407
+ findings.append(
408
+ Finding(
409
+ "ERROR",
410
+ "container-only-key-off-container",
411
+ f"{container_only_present} on `fidelity: {fidelity}` — container-only (present_files "
412
+ "is served only there); off-container it reports cannot-verify. Use fidelity: "
413
+ "container for present_files/scratchpad delivery.",
414
+ "Use fidelity: container (or drop the assertion for this tier).",
415
+ path,
416
+ )
417
+ )
418
+ elif fidelity == "cowork":
419
+ findings.append(
420
+ Finding(
421
+ "WARN",
422
+ "container-only-key-off-container",
423
+ f"{container_only_present} on `fidelity: cowork` — container-only (present_files "
424
+ "is served only there); off-container it reports cannot-verify. Use fidelity: "
425
+ f"container for present_files/scratchpad delivery. (The tier resolves per the "
426
+ f"baseline's host-loop gate ({HOST_LOOP_GATE_ID}); if it resolves off-container this "
427
+ "assertion cannot verify anything.)",
428
+ "Pin fidelity: container if the assertion is load-bearing; keep cowork only if "
429
+ "you accept the gate-resolution dependency.",
430
+ path,
431
+ )
432
+ )
433
+
391
434
  # E: requires_capabilities on protocol — the capability probe cannot run at protocol tier
392
435
  # (clause b of the requires_capabilities contract), so the run HARD-FAILS unless an assert item
393
436
  # opts out via allow_missing_capability: true. Offline-detectable fails-by-design, same class as
package/CHANGELOG.md CHANGED
@@ -6,6 +6,34 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.1.0] — 2026-07-16
10
+
11
+ Minor: `analyze-skill` gains interactive-artifact write-back detection (static + an optional `--runtime`
12
+ headless-DOM confirmer), `lint` gains a container-only-key tier check, and the `doctor` JSON envelope is
13
+ frozen as a covered SPEC §12 surface. All additive.
14
+
15
+ ### Added
16
+
17
+ - **`analyze-skill` now detects interactive-artifact write-backs lost under Cowork.** Alongside the
18
+ existing `/sessions` path scan, it statically analyzes `.html/.htm/.js/.mjs/.ts/.jsx/.tsx/.py` sources
19
+ under the target for a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back that silently
20
+ fails under Cowork (the artifact is served from Cowork's own origin, so a relative write-back resolves
21
+ non-ok and a page that doesn't check `resp.ok` shows a false "Saved"). Findings: `artifact-write-back-lost`
22
+ (error — gates under `--strict`), `artifact-write-back-suspect` (advisory), and a separate top-level
23
+ `analysisFailures` **could-not-verify** channel (a candidate that couldn't be parsed/analyzed) that
24
+ always exits `3`, `--strict`-independent. A guard that isn't statically provable-truthy is `suspect`,
25
+ never silently clean; the blanket `analyze-skill: ignore` marker does not silence artifact rules.
26
+ Each `SkillFinding` now carries a `severity` (`error|advisory`); `--strict` gates on any error finding.
27
+ - **`analyze-skill --runtime`** — an optional headless-DOM confirmation that drives a materialized `.html`
28
+ artifact in jsdom (stubbed network + synthetic user actions, run twice) to *observe* whether a relative
29
+ write-back fires and is lost. Enrichment only (never changes the exit code); trusted-source scope. `jsdom`
30
+ is an optional, dynamically-imported dependency — absent it reports "run `npm i jsdom` to enable".
31
+ - **`schema/doctor.json`** — the `doctor --output-format json` envelope is now a covered SPEC §12 surface
32
+ (`oneOf` the completed-probe shape and the shared error envelope for every category). `doctor`'s normal
33
+ JSON output is standardized through the shared envelope frame.
34
+ - **`lint` flags container-only assertion keys off-container** — `no_scratchpad_leak`/`present_files_called`
35
+ on `fidelity: protocol|microvm|hostloop` is an ERROR, on `fidelity: cowork` a WARN, clean on `container`.
36
+
9
37
  ## [1.0.6] — 2026-07-15
10
38
 
11
39
  Patch: platform baseline synced to Claude Desktop `1.21459.0`. The spawn contract and rendered system
package/README.md CHANGED
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
91
91
 
92
92
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
93
93
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
94
- > From a global install (`npm i -g "cowork-harness@>=1.0.6"`), point at the package root instead:
94
+ > From a global install (`npm i -g "cowork-harness@>=1.1.0"`), point at the package root instead:
95
95
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
96
96
  > (or copy the cassette into your own project and pass that path).
97
97
 
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
101
101
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
102
102
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
103
103
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
104
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.0.6"`.
104
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.1.0"`.
105
105
 
106
106
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
107
107
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
126
126
  claude plugin install cowork-harness@cowork-harness
127
127
  ```
128
128
 
129
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.0.6"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
129
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.1.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
130
130
 
131
131
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
132
132
 
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
147
147
  To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
148
148
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
149
149
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
150
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.0.6"` — see
150
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.1.0"` — see
151
151
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
152
152
 
153
153
  ### Prerequisites for anything above `protocol` fidelity
@@ -401,7 +401,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
401
401
  | `lint <scenario.yaml \| dir/>…` | Check scenarios for silent false-greens — assertions placed on the wrong CI lane, mixed content/live keys, missing `controlOut`-required keys (files or a directory of `*.yaml`/`*.yml`; bundled `scenario.py`; needs python3 — PyYAML is bundled); `--json` emits findings as machine-readable JSON instead of the text report | before committing a new scenario or after changing assertions |
402
402
  | `lint-skill <SKILL.md \| skill-dir/>…` | Lint a skill body (and any sibling `hooks.json`) for two Cowork host-loop footguns — a `${CLAUDE_PLUGIN_ROOT}` path used in an in-VM bash context, and a hook command that exports an env var or writes into `/tmp` for the in-VM agent — plus static resolution of any pinned `subagent_type` against the enclosing plugin's `agents/*.md` (bundled `scenario.py`; needs python3); the two footguns are WARN-only, and of the three `subagent_type` outcomes, an in-plugin-prefixed agent missing from the enclosing plugin's fully-enumerated `agents/*.md` (`subagent-type-not-found-in-plugin`) is a **provable typo and is WARN too**, while a cross-plugin (`subagent-type-unresolvable`) or unknown-bare (`subagent-type-unknown`) value stays INFO (no built-in agent-type registry to disprove it against) — pass `--strict` (the CI-recommended invocation; plain `lint-skill` is advisory-only) to fail on any WARN, or `--json` for machine-readable findings | authoring or reviewing a skill before a paid Cowork host-loop run exposes the footgun, or before a pinned `subagent_type` typo breaks a `Task` dispatch |
403
403
  | `python3 …/scenario.py resolve-agent-types <plugin-dir> [--json]` | Token-free: print a plugin's valid `<plugin>:<agent>` subagent types, resolved from `.claude-plugin/plugin.json` + `agents/*.md` frontmatter (filename-stem fallback when an agent file has no `name:`); the `…` is `.claude/skills/cowork-harness/scripts/` | "does `founder-skills:deck-review` resolve within this plugin?" without a live dispatch |
404
- | `analyze-skill <SKILL.md \| skill-dir/ \| glob>…` | Token-free ADVISORY scan: flags a `/sessions/...` path handed to a file tool or used as a dispatch/sub-agent output path — that path class is DENIED on host-loop. Accepts multiple positionals (files, dirs, or a simple `*`/`**` glob), walked recursively across a plugin's `SKILL.md`/`agents/`/`references/`/`commands/`. Findings print but exit 0 by default; `--strict` (the CI-recommended invocation) fails on any unsuppressed finding. Three per-file ignore markers (`ignore-next-line`, `ignore-start`/`ignore-end`, file-wide `ignore`) and a `--output-format json` payload are supported — full reference, including the ignore-marker syntax, directory-walk rules, and JSON shape: [docs/subagents.md](./docs/subagents.md#static-path-fidelity-check-analyze-skill) | catching the exact "skill hands `/sessions/...` to a file tool" defect statically, across a whole plugin's contract surface (or several paths/globs in one call), before paying for a live host-loop run to discover it |
404
+ | `analyze-skill <SKILL.md \| skill-dir/ \| glob>…` | Token-free ADVISORY scan: flags a `/sessions/...` path handed to a file tool or used as a dispatch/sub-agent output path — that path class is DENIED on host-loop. Accepts multiple positionals (files, dirs, or a simple `*`/`**` glob), walked recursively across a plugin's `SKILL.md`/`agents/`/`references/`/`commands/`. Findings print but exit 0 by default; `--strict` (the CI-recommended invocation) fails on any unsuppressed finding. Three per-file ignore markers (`ignore-next-line`, `ignore-start`/`ignore-end`, file-wide `ignore`) and a `--output-format json` payload are supported. **It also flags interactive-artifact write-backs lost under Cowork**: it statically analyzes `.html`/`.js`/`.ts`/`.py` sources under the target for a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back that silently fails when the artifact is served from Cowork's own origin (`artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; a candidate that can't be parsed is a could-not-verify exit `3`). `--runtime` adds an optional headless-DOM confirmation (needs `jsdom`) that *observes* the lost write-back. Full reference — ignore-marker syntax, directory-walk rules, JSON shape, and the artifact detector: [docs/subagents.md](./docs/subagents.md#static-path-fidelity-check-analyze-skill) | catching the exact "skill hands `/sessions/...` to a file tool" defect — or an artifact whose Submit silently fails under Cowork — statically, before paying for a live run to discover it |
405
405
  | `probe-dispatch <skill-dir> "<prompt>"` | Cheap single-dispatch mechanics probe: a THIN wrapper over `skill` (fidelity `container`/`microvm`/`hostloop`, default `hostloop`) that scopes a prompt to trigger ONE `Task` dispatch, then prints just that dispatch's `{resolvedAgentType, pathDenials, delivered}` — no new data model, a pure projection of the same `RunResult` `skill` already produces; `--expect-write <suffix>` narrows `delivered`; "one dispatch" is prompt-scoped, not enforced (`--output-format json` for machine consumption) | "did THIS dispatch resolve to the type I expect, avoid a path denial, and actually deliver its write?" without hand-writing a scenario or reading `trace --view dispatches` |
406
406
  | `assertions --list` | List the available scenario assertions (generated from the schema) | "what can I assert?" without grepping the source |
407
407
  | `decide` | Validate a decider against a sample question in ~2 s (no run) | sanity-check a `--decider-*` / `--answer` wiring before a long run |
@@ -665,7 +665,7 @@ jobs:
665
665
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
666
666
  ```
667
667
 
668
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.0.6` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
668
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.1.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
669
669
 
670
670
  The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
671
671
 
package/SPEC.md CHANGED
@@ -674,6 +674,34 @@ already-redacted markers. The TLD list used by the domain scanner was also exten
674
674
  entries (CB-5), adding major European, Asian, and Latin American ccTLDs
675
675
  (`ch|nl|se|no|it|jp|br|nz|in|sg|kr|mx|es|pt|pl|be|at|dk|fi|ie|ru|cn|tw|hu|cz|ro|il|za|ar|cl|pe|tr`).
676
676
 
677
+ ### 11.2 `doctor` — dedicated envelope
678
+
679
+ `doctor [--tier <t>]` is the read-only prerequisite check ("can I run the live tiers — what's
680
+ missing?"). Its `--output-format json` output does **not** reuse the `run/skill/replay` envelope (no
681
+ `RunResult` to judge) — it emits its own, published as **`schema/doctor.json`** (a §12-covered contract
682
+ surface):
683
+
684
+ ```jsonc
685
+ // Completed probe (normal path — routed through the shared jsonPayloadEnvelope, so error is always null):
686
+ { "tool": "cowork-harness", "version": "...", "command": "doctor",
687
+ "ok": true, // false iff any `required` check has status:"fail" for the selected tier
688
+ "error": null,
689
+ "tier": "container", // protocol|container|microvm|hostloop|cowork
690
+ "checks": [ { "id": "node", "title": "Node ≥ 20", "status": "ok", "detail": "node 22.x.x",
691
+ "remedy?": "string", "required": true } ] } // status: ok|fail|warn|skip
692
+
693
+ // Shared error envelope (a thrown failure — bad flag, or the top-level catch):
694
+ { "tool": "cowork-harness", "version": "...", "command": "doctor", "ok": false, "results": [],
695
+ "error": { "category": "usage|unanswered|boundary|runtime|internal", "message": "string", "hint?": "string" } }
696
+ ```
697
+
698
+ `ok = !checks.some(c => c.required && c.status === "fail")`. **Exit codes:** `0` all required checks
699
+ pass · `1` a required check is `status:"fail"` for the selected tier (the completed-probe shape, still
700
+ `error:null`) · `2` usage (bad `--tier`/unexpected args) or an unexpected internal failure caught by the
701
+ top-level catch (`category:"internal"`). The `checks[].id` set is **not** enumerated as a closed
702
+ contract — it grows as tiers/checks are added, the same way `trace` row shapes are excluded from §12's
703
+ covered list.
704
+
677
705
  ## 12. Versioning & the 1.0 compatibility contract
678
706
 
679
707
  From `1.0.0` the project follows [semver](https://semver.org/). The surfaces below are the **covered
@@ -698,6 +726,12 @@ Covered-surface changes follow semver as of `1.0.0` — see [RELEASING.md](./REL
698
726
  `findings` / `staleness` / `unverifiable` / `notes` / `version` / `error` channels, and the exit-code
699
727
  split (`0`/`1`/`2`/`3`, §11) they map to. This is the machine output the CI recipes and the packaged
700
728
  Action steer consumers to parse; renaming or removing a key is breaking, adding one is not.
729
+ - **`doctor` envelope** — `schema/doctor.json` under `--output-format json` (§11.2): the completed-probe
730
+ shape (`tool` / `version` / `command` / `ok` / `error:null` / `tier` / `checks[]`, each check's
731
+ `id` / `title` / `status` / `detail` / `required` / optional `remedy`) and the shared error-envelope
732
+ shape (`results:[]` / `error.category`) it falls back to on a thrown failure; renaming or removing a
733
+ key is breaking, adding one is not. The `checks[].id` set itself is NOT covered — it grows with new
734
+ tiers/checks.
701
735
  - **Cassette format** — the current `cassetteVersion` (**10**, `schema/cassette.v10.json`) and its
702
736
  verdict-modifier assertion keys. The minimum supported read version is **v9**
703
737
  (`MIN_SUPPORTED_CASSETTE_VERSION`): a cassette below the floor is refused at load time with a
package/dist/cli.js CHANGED
@@ -406,7 +406,7 @@ const SUBCOMMAND_USAGE = {
406
406
  rehash: "usage: rehash <dir/> [--dry-run] [--output-format text|json] (migrate cassettes across format bumps using contentSig verification; no re-record needed)",
407
407
  prune: "usage: prune [--keep-last <n>] [--pinned-older-than <N>d|h|m] [--dry-run] [<runs-dir>] (prune accumulated run dirs; default --keep-last 5)",
408
408
  "init-redact": "usage: init-redact [--force] [--output-format json] (copy the packaged reference .cowork-redact.json into the cwd; refuses to overwrite an existing one without --force)",
409
- "analyze-skill": "usage: analyze-skill <SKILL.md | skill-dir/ | glob>… [--strict] [--output-format text|json] (ADVISORY token-free scan: flags a /sessions/... path handed to a file tool or dispatch/sub-agent output — denied on host-loop)\n" +
409
+ "analyze-skill": "usage: analyze-skill <SKILL.md | skill-dir/ | glob>… [--strict] [--runtime] [--output-format text|json] (ADVISORY token-free scan: flags a /sessions/... path handed to a file tool or dispatch/sub-agent output — denied on host-loop; AND interactive-artifact write-backs that are lost under Cowork)\n" +
410
410
  " reuses the harness's own ported /sessions path-gate predicate (isVmSessionsPath) as the deny decision; only the EXTRACTION from SKILL.md text is heuristic.\n" +
411
411
  " a DIRECTORY target scans the UNION of every contract-bearing markdown file present, not just SKILL.md — a plugin's dispatch/output contracts often live in agents/*.md or references/*.md:\n" +
412
412
  " · top-level SKILL.md present → + top-level references/**/*.md\n" +
@@ -414,10 +414,11 @@ const SUBCOMMAND_USAGE = {
414
414
  " · a skill dir INSIDE a plugin (SKILL.md present, walking up finds an enclosing plugin manifest) → + that plugin's root agents/*.md\n" +
415
415
  ' MULTIPLE positionals are accepted (matches lint-skill\'s nargs="+"), each resolved independently (file/dir via the rules above, or a `*`-bearing target via a small hand-rolled glob matcher — `dir/*.md` shallow, `dir/**/*.md` recursive; no dependency, no engines bump).\n' +
416
416
  " every positional's file set is UNIONed and deduped by resolved absolute path across the WHOLE invocation (a file reached both directly and via a dir/glob is analyzed once). ZERO scannable files across ALL positionals is a USAGE ERROR (exit 2), never a silent clean pass; an unresolvable single positional (bad path, unrecognized glob shape) fails the whole invocation the same way.\n" +
417
- " findings are WARNINGS: exit 0 by default even when findings are printed. --strict fails (exit 1) if ANY scanned file (across ALL positionals) has an unsuppressed finding (mirrors lint-skill's --strict). exit 2 = usage error.\n" +
417
+ " findings are WARNINGS: exit 0 by default even when findings are printed. --strict fails (exit 1) if ANY scanned file (across ALL positionals) has an unsuppressed error-severity finding (mirrors lint-skill's --strict). exit 2 = usage error. exit 3 = could-not-verify (an artifact candidate that couldn't be parsed/analyzed — always, --strict-independent; a strict error finding's exit 1 takes precedence).\n" +
418
+ " interactive-artifact write-backs (Item 1): also scans .html/.js/.ts/.jsx/.tsx/.py sources under the target for a relative fetch/XHR/sendBeacon/<form method=post> write-back that is LOST under Cowork (the artifact is served from Cowork's own origin, so a relative write-back resolves non-ok and a page that doesn't check resp.ok shows a false 'Saved'). Rules: artifact-write-back-lost (error, gates under --strict), artifact-write-back-suspect (advisory), plus a could-not-verify channel. These are NOT silenced by the `analyze-skill: ignore` marker. --runtime adds an OPTIONAL headless-DOM confirmation (needs jsdom, dynamically imported; ENRICHMENT only — never changes the exit code; trusted-source scope).\n" +
418
419
  " plain `analyze-skill` (no --strict) is ADVISORY ONLY — CI should invoke `analyze-skill ... --strict` to actually gate on findings, same as `lint-skill --strict`.\n" +
419
420
  " a line containing `analyze-skill: ignore` (bare, or inside an HTML comment) anywhere in one of the scanned files silences EVERY path-fidelity finding for THAT file — even under --strict, even a genuine true positive (an explicit author override); it does not affect other scanned files.\n" +
420
- " --output-format json emits {tool,version,command,ok,files:[{file,findings,suppressed}],scanned,unscanned,strict,error} — NOT a bare array and NOT results[]; `scanned`/`unscanned` name the files analyzed and any contract dirs left out of scope (printed as a banner in text mode too).\n" +
421
+ " --output-format json emits {tool,version,command,ok,files:[{file,findings,suppressed}],scanned,unscanned,analysisFailures,[runtimeConfirmations (only with --runtime)],strict,error} — NOT a bare array and NOT results[]; `scanned`/`unscanned` name the files analyzed and any contract dirs left out of scope (printed as a banner in text mode too); each finding now carries a `severity` (error|advisory).\n" +
421
422
  " a clean/suppressed result is a PRE-FLIGHT signal only — the runtime no_vm_path_file_op / vm_path_denied asserts remain authoritative (see docs/subagents.md).",
422
423
  };
423
424
  // Known subcommands — used by the global value-flag parsers (`--dotenv`, `--run-dir`) to reject a command