cowork-harness 1.0.6 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +8 -6
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +31 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/scripts/scenario.py +43 -0
- package/CHANGELOG.md +28 -0
- package/README.md +6 -6
- package/SPEC.md +34 -0
- package/dist/cli.js +4 -3
- package/dist/run/analyze-artifact-runtime.js +547 -0
- package/dist/run/analyze-artifact.js +921 -0
- package/dist/run/analyze-skill.js +99 -21
- package/dist/run/doctor.js +5 -2
- package/docs/fidelity-gaps.md +42 -0
- package/docs/subagents.md +38 -0
- package/examples/replays/README.md +1 -1
- package/package.json +4 -1
- package/python/test_scenario_lint.py +88 -0
- package/schema/doctor.json +81 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.0
|
|
7
|
-
tracks-harness: cowork-harness 1.0
|
|
6
|
+
version: 1.1.0
|
|
7
|
+
tracks-harness: cowork-harness 1.1.0 (baseline desktop-1.21459.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.1.0` (baseline
|
|
26
26
|
> `desktop-1.21459.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.1.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.1.0"`. **Pin `@>=1.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
|
-
What the ≥ 1.0
|
|
44
|
+
What the ≥ 1.1.0 floor gates, by release:
|
|
45
45
|
|
|
46
46
|
- **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
|
|
47
47
|
- **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
|
|
@@ -52,6 +52,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
52
52
|
- **0.32.0:** `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict`.
|
|
53
53
|
- **0.33.0:** the `redacted` marker on display-omitted reasoning — `subagents[].reasoning` and the top-level `thinking[]` now carry `{text:"", redacted:true}` when the model returns a signed-but-empty thought (so "reasoned, text omitted" is distinct from "no thought"), plus the fenced `debug.thinking_display` escape hatch.
|
|
54
54
|
- **1.0.0:** first stable release — the SPEC §12 compatibility contract takes effect (covered CLI/schema/env/Action surfaces are now stable; breaking changes need a major bump). No new author-facing command; the floor simply tracks the 1.0 release.
|
|
55
|
+
- **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
|
|
55
56
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
56
57
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
57
58
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
@@ -277,7 +278,8 @@ quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `
|
|
|
277
278
|
(ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
|
|
278
279
|
baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
|
|
279
280
|
`allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
|
|
280
|
-
unverifiable), a
|
|
281
|
+
unverifiable), a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) off the
|
|
282
|
+
`container` tier (ERROR on `protocol`/`microvm`/`hostloop`, WARN on `cowork`), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
|
|
281
283
|
and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
|
|
282
284
|
(CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
|
|
283
285
|
emits a scenario `lint` would reject.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.0
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.1.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.0
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.1.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.0
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.1.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.0
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.1.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.0
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.1.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.0
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.1.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -171,6 +171,36 @@ writes a `result.json` (marked `partial: true`) with the artifacts the agent pro
|
|
|
171
171
|
the work isn't discarded — then exits 2. Inspect it with `cowork-harness inspect <run-dir>`. `verify-run`
|
|
172
172
|
and `scaffold` refuse to treat a partial run's half-finished output as a passing result.
|
|
173
173
|
|
|
174
|
+
## What a green does NOT prove
|
|
175
|
+
|
|
176
|
+
A passing run is evidence for exactly what it checked, not a blanket certificate. Three gaps come
|
|
177
|
+
up often enough to spell out:
|
|
178
|
+
|
|
179
|
+
- **A green `replay` proves "same as when recorded," not "correct today."** `replay` never touches
|
|
180
|
+
a filesystem or network — it re-evaluates assertions from the frozen cassette. A fixed set of
|
|
181
|
+
keys is live-only and **skipped outright** on replay (absent from `assertions[]`, not vacuously
|
|
182
|
+
passed): `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
|
|
183
|
+
`egress_allowed`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, and `expect_denied`.
|
|
184
|
+
Everything else that *is* evaluated is checked against the **recording**, not fresh behavior — a
|
|
185
|
+
green replay says the skill produced these events when it was recorded, not that it still does
|
|
186
|
+
(`staleness[]` flags skill/baseline drift as a hint; only a live `run` re-confirms current
|
|
187
|
+
behavior). See `docs/cassette.md` § "Still skipped on replay" and `docs/scenario.md` § "Which
|
|
188
|
+
assertions survive replay."
|
|
189
|
+
- **Container-only assertions can't verify off the `container` tier.** `no_scratchpad_leak` and
|
|
190
|
+
`present_files_called` check the `present_files` delivery path, which is served **only** on
|
|
191
|
+
`container` — not `hostloop`/`microvm`. Asserting them off-container hard-fails at runtime (a red
|
|
192
|
+
run, not a false green), so you won't be fooled if you write the assertion. The quieter trap is a
|
|
193
|
+
scenario that runs at `hostloop`/`microvm`/`protocol` and simply omits these assertions: a green
|
|
194
|
+
run there proves **nothing** about scratchpad-leak safety or present_files delivery, because that
|
|
195
|
+
tier never exercises the delivery path the assertions would check. Use `fidelity: container` for
|
|
196
|
+
present_files/scratchpad-delivery coverage.
|
|
197
|
+
- **The harness doesn't observe rendered-artifact interactions, browser downloads, or human
|
|
198
|
+
clicks.** It runs the agent headless — no webview, no browser, no person clicking "Submit." A
|
|
199
|
+
class of Cowork bug (a client-side write-back to a relative URL that resolves-but-fails against
|
|
200
|
+
Cowork's own origin, or a broken blob-download fallback) is invisible to any live run, however
|
|
201
|
+
faithfully sandboxed, because it only manifests in a rendered DOM a human is driving. See
|
|
202
|
+
`docs/fidelity-gaps.md` § "Browser↔webview↔human-interaction boundary."
|
|
203
|
+
|
|
174
204
|
## Relevant environment variables
|
|
175
205
|
|
|
176
206
|
- `COWORK_HARNESS_RUNS_DIR` (or `--run-dir <path>`) — override the default run-output root `~/.cowork-harness/runs` (out of any working tree). flag > env > default.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.1.0`
|
|
4
4
|
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
@@ -17,6 +17,8 @@ Two subcommands:
|
|
|
17
17
|
lint flags (see references/scenario-schema.md for the why of each):
|
|
18
18
|
E egress assertion on `fidelity: protocol` (the harness rejects this run)
|
|
19
19
|
E `transcript_no_host_path` on hostloop/protocol (fails BY DESIGN at those tiers)
|
|
20
|
+
E `no_scratchpad_leak`/`present_files_called` off (container-only: served only at fidelity:
|
|
21
|
+
protocol/microvm/hostloop container; off-container = cannot-verify)
|
|
20
22
|
E `requires_capabilities` on `fidelity: protocol` (probe can't run → hard-fails
|
|
21
23
|
unless allow_missing_capability)
|
|
22
24
|
E on_unanswered: agent / invalid value (schema rejects `agent`)
|
|
@@ -24,6 +26,8 @@ lint flags (see references/scenario-schema.md for the why of each):
|
|
|
24
26
|
E `assertions:` instead of `assert:` (block ignored → every check no-ops)
|
|
25
27
|
W `transcript_no_host_path` on `fidelity: cowork` (tier resolves per baseline gate —
|
|
26
28
|
incompatible if it lands hostloop)
|
|
29
|
+
W `no_scratchpad_leak`/`present_files_called` on (tier resolves per baseline gate — cannot-verify
|
|
30
|
+
`fidelity: cowork` if it resolves off-container)
|
|
27
31
|
W no content assertion → no-op on a replay gate (every assertion is fs/egress)
|
|
28
32
|
W mixed-class assert item → fs/egress half dropped on replay
|
|
29
33
|
W unknown top-level / assertion key (typo or hallucinated schema)
|
|
@@ -127,6 +131,10 @@ LIVE_ONLY_KEYS = {
|
|
|
127
131
|
"semantic_matches",
|
|
128
132
|
}
|
|
129
133
|
EGRESS_KEYS = {"egress_denied", "egress_allowed"}
|
|
134
|
+
# container-only: served only at fidelity: container (present_files / the scratchpad promotion path
|
|
135
|
+
# it depends on). Off-container these report cannot-verify, not a meaningful pass/fail — same tier-fidelity
|
|
136
|
+
# class as transcript_no_host_path below, just the opposite direction (container-only vs container-hostile).
|
|
137
|
+
CONTAINER_ONLY_KEYS = {"no_scratchpad_leak", "present_files_called"}
|
|
130
138
|
# verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
|
|
131
139
|
VERDICT_MODIFIER_KEYS = {
|
|
132
140
|
"allow_permissive_auto_allow",
|
|
@@ -388,6 +396,41 @@ def lint_doc(doc, path, raw_lines):
|
|
|
388
396
|
)
|
|
389
397
|
)
|
|
390
398
|
|
|
399
|
+
# E/W: CONTAINER_ONLY_KEYS (no_scratchpad_leak / present_files_called) are served only at
|
|
400
|
+
# fidelity: container — off-container they report cannot-verify, not a meaningful check. Mirrors
|
|
401
|
+
# the transcript_no_host_path tier check above, just the opposite direction: ERROR on the tiers
|
|
402
|
+
# where the runtime deterministically can't serve it, WARN on cowork (baseline-gate-resolution
|
|
403
|
+
# dependent — the linter stays offline, so the message names the dependency instead of resolving it).
|
|
404
|
+
container_only_present = sorted(assert_keys & CONTAINER_ONLY_KEYS)
|
|
405
|
+
if container_only_present:
|
|
406
|
+
if fidelity in ("protocol", "microvm", "hostloop"):
|
|
407
|
+
findings.append(
|
|
408
|
+
Finding(
|
|
409
|
+
"ERROR",
|
|
410
|
+
"container-only-key-off-container",
|
|
411
|
+
f"{container_only_present} on `fidelity: {fidelity}` — container-only (present_files "
|
|
412
|
+
"is served only there); off-container it reports cannot-verify. Use fidelity: "
|
|
413
|
+
"container for present_files/scratchpad delivery.",
|
|
414
|
+
"Use fidelity: container (or drop the assertion for this tier).",
|
|
415
|
+
path,
|
|
416
|
+
)
|
|
417
|
+
)
|
|
418
|
+
elif fidelity == "cowork":
|
|
419
|
+
findings.append(
|
|
420
|
+
Finding(
|
|
421
|
+
"WARN",
|
|
422
|
+
"container-only-key-off-container",
|
|
423
|
+
f"{container_only_present} on `fidelity: cowork` — container-only (present_files "
|
|
424
|
+
"is served only there); off-container it reports cannot-verify. Use fidelity: "
|
|
425
|
+
f"container for present_files/scratchpad delivery. (The tier resolves per the "
|
|
426
|
+
f"baseline's host-loop gate ({HOST_LOOP_GATE_ID}); if it resolves off-container this "
|
|
427
|
+
"assertion cannot verify anything.)",
|
|
428
|
+
"Pin fidelity: container if the assertion is load-bearing; keep cowork only if "
|
|
429
|
+
"you accept the gate-resolution dependency.",
|
|
430
|
+
path,
|
|
431
|
+
)
|
|
432
|
+
)
|
|
433
|
+
|
|
391
434
|
# E: requires_capabilities on protocol — the capability probe cannot run at protocol tier
|
|
392
435
|
# (clause b of the requires_capabilities contract), so the run HARD-FAILS unless an assert item
|
|
393
436
|
# opts out via allow_missing_capability: true. Offline-detectable fails-by-design, same class as
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,34 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.1.0] — 2026-07-16
|
|
10
|
+
|
|
11
|
+
Minor: `analyze-skill` gains interactive-artifact write-back detection (static + an optional `--runtime`
|
|
12
|
+
headless-DOM confirmer), `lint` gains a container-only-key tier check, and the `doctor` JSON envelope is
|
|
13
|
+
frozen as a covered SPEC §12 surface. All additive.
|
|
14
|
+
|
|
15
|
+
### Added
|
|
16
|
+
|
|
17
|
+
- **`analyze-skill` now detects interactive-artifact write-backs lost under Cowork.** Alongside the
|
|
18
|
+
existing `/sessions` path scan, it statically analyzes `.html/.htm/.js/.mjs/.ts/.jsx/.tsx/.py` sources
|
|
19
|
+
under the target for a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back that silently
|
|
20
|
+
fails under Cowork (the artifact is served from Cowork's own origin, so a relative write-back resolves
|
|
21
|
+
non-ok and a page that doesn't check `resp.ok` shows a false "Saved"). Findings: `artifact-write-back-lost`
|
|
22
|
+
(error — gates under `--strict`), `artifact-write-back-suspect` (advisory), and a separate top-level
|
|
23
|
+
`analysisFailures` **could-not-verify** channel (a candidate that couldn't be parsed/analyzed) that
|
|
24
|
+
always exits `3`, `--strict`-independent. A guard that isn't statically provable-truthy is `suspect`,
|
|
25
|
+
never silently clean; the blanket `analyze-skill: ignore` marker does not silence artifact rules.
|
|
26
|
+
Each `SkillFinding` now carries a `severity` (`error|advisory`); `--strict` gates on any error finding.
|
|
27
|
+
- **`analyze-skill --runtime`** — an optional headless-DOM confirmation that drives a materialized `.html`
|
|
28
|
+
artifact in jsdom (stubbed network + synthetic user actions, run twice) to *observe* whether a relative
|
|
29
|
+
write-back fires and is lost. Enrichment only (never changes the exit code); trusted-source scope. `jsdom`
|
|
30
|
+
is an optional, dynamically-imported dependency — absent it reports "run `npm i jsdom` to enable".
|
|
31
|
+
- **`schema/doctor.json`** — the `doctor --output-format json` envelope is now a covered SPEC §12 surface
|
|
32
|
+
(`oneOf` the completed-probe shape and the shared error envelope for every category). `doctor`'s normal
|
|
33
|
+
JSON output is standardized through the shared envelope frame.
|
|
34
|
+
- **`lint` flags container-only assertion keys off-container** — `no_scratchpad_leak`/`present_files_called`
|
|
35
|
+
on `fidelity: protocol|microvm|hostloop` is an ERROR, on `fidelity: cowork` a WARN, clean on `container`.
|
|
36
|
+
|
|
9
37
|
## [1.0.6] — 2026-07-15
|
|
10
38
|
|
|
11
39
|
Patch: platform baseline synced to Claude Desktop `1.21459.0`. The spawn contract and rendered system
|
package/README.md
CHANGED
|
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
91
91
|
|
|
92
92
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
93
93
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
94
|
-
> From a global install (`npm i -g "cowork-harness@>=1.0
|
|
94
|
+
> From a global install (`npm i -g "cowork-harness@>=1.1.0"`), point at the package root instead:
|
|
95
95
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
96
96
|
> (or copy the cassette into your own project and pass that path).
|
|
97
97
|
|
|
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
101
101
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
102
102
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
103
103
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
104
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.0
|
|
104
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.1.0"`.
|
|
105
105
|
|
|
106
106
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
107
107
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
126
126
|
claude plugin install cowork-harness@cowork-harness
|
|
127
127
|
```
|
|
128
128
|
|
|
129
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.0
|
|
129
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.1.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
130
130
|
|
|
131
131
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
132
132
|
|
|
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
|
|
|
147
147
|
To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
|
|
148
148
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
149
149
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
150
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.0
|
|
150
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.1.0"` — see
|
|
151
151
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
152
152
|
|
|
153
153
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -401,7 +401,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
401
401
|
| `lint <scenario.yaml \| dir/>…` | Check scenarios for silent false-greens — assertions placed on the wrong CI lane, mixed content/live keys, missing `controlOut`-required keys (files or a directory of `*.yaml`/`*.yml`; bundled `scenario.py`; needs python3 — PyYAML is bundled); `--json` emits findings as machine-readable JSON instead of the text report | before committing a new scenario or after changing assertions |
|
|
402
402
|
| `lint-skill <SKILL.md \| skill-dir/>…` | Lint a skill body (and any sibling `hooks.json`) for two Cowork host-loop footguns — a `${CLAUDE_PLUGIN_ROOT}` path used in an in-VM bash context, and a hook command that exports an env var or writes into `/tmp` for the in-VM agent — plus static resolution of any pinned `subagent_type` against the enclosing plugin's `agents/*.md` (bundled `scenario.py`; needs python3); the two footguns are WARN-only, and of the three `subagent_type` outcomes, an in-plugin-prefixed agent missing from the enclosing plugin's fully-enumerated `agents/*.md` (`subagent-type-not-found-in-plugin`) is a **provable typo and is WARN too**, while a cross-plugin (`subagent-type-unresolvable`) or unknown-bare (`subagent-type-unknown`) value stays INFO (no built-in agent-type registry to disprove it against) — pass `--strict` (the CI-recommended invocation; plain `lint-skill` is advisory-only) to fail on any WARN, or `--json` for machine-readable findings | authoring or reviewing a skill before a paid Cowork host-loop run exposes the footgun, or before a pinned `subagent_type` typo breaks a `Task` dispatch |
|
|
403
403
|
| `python3 …/scenario.py resolve-agent-types <plugin-dir> [--json]` | Token-free: print a plugin's valid `<plugin>:<agent>` subagent types, resolved from `.claude-plugin/plugin.json` + `agents/*.md` frontmatter (filename-stem fallback when an agent file has no `name:`); the `…` is `.claude/skills/cowork-harness/scripts/` | "does `founder-skills:deck-review` resolve within this plugin?" without a live dispatch |
|
|
404
|
-
| `analyze-skill <SKILL.md \| skill-dir/ \| glob>…` | Token-free ADVISORY scan: flags a `/sessions/...` path handed to a file tool or used as a dispatch/sub-agent output path — that path class is DENIED on host-loop. Accepts multiple positionals (files, dirs, or a simple `*`/`**` glob), walked recursively across a plugin's `SKILL.md`/`agents/`/`references/`/`commands/`. Findings print but exit 0 by default; `--strict` (the CI-recommended invocation) fails on any unsuppressed finding. Three per-file ignore markers (`ignore-next-line`, `ignore-start`/`ignore-end`, file-wide `ignore`) and a `--output-format json` payload are supported
|
|
404
|
+
| `analyze-skill <SKILL.md \| skill-dir/ \| glob>…` | Token-free ADVISORY scan: flags a `/sessions/...` path handed to a file tool or used as a dispatch/sub-agent output path — that path class is DENIED on host-loop. Accepts multiple positionals (files, dirs, or a simple `*`/`**` glob), walked recursively across a plugin's `SKILL.md`/`agents/`/`references/`/`commands/`. Findings print but exit 0 by default; `--strict` (the CI-recommended invocation) fails on any unsuppressed finding. Three per-file ignore markers (`ignore-next-line`, `ignore-start`/`ignore-end`, file-wide `ignore`) and a `--output-format json` payload are supported. **It also flags interactive-artifact write-backs lost under Cowork**: it statically analyzes `.html`/`.js`/`.ts`/`.py` sources under the target for a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back that silently fails when the artifact is served from Cowork's own origin (`artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; a candidate that can't be parsed is a could-not-verify exit `3`). `--runtime` adds an optional headless-DOM confirmation (needs `jsdom`) that *observes* the lost write-back. Full reference — ignore-marker syntax, directory-walk rules, JSON shape, and the artifact detector: [docs/subagents.md](./docs/subagents.md#static-path-fidelity-check-analyze-skill) | catching the exact "skill hands `/sessions/...` to a file tool" defect — or an artifact whose Submit silently fails under Cowork — statically, before paying for a live run to discover it |
|
|
405
405
|
| `probe-dispatch <skill-dir> "<prompt>"` | Cheap single-dispatch mechanics probe: a THIN wrapper over `skill` (fidelity `container`/`microvm`/`hostloop`, default `hostloop`) that scopes a prompt to trigger ONE `Task` dispatch, then prints just that dispatch's `{resolvedAgentType, pathDenials, delivered}` — no new data model, a pure projection of the same `RunResult` `skill` already produces; `--expect-write <suffix>` narrows `delivered`; "one dispatch" is prompt-scoped, not enforced (`--output-format json` for machine consumption) | "did THIS dispatch resolve to the type I expect, avoid a path denial, and actually deliver its write?" without hand-writing a scenario or reading `trace --view dispatches` |
|
|
406
406
|
| `assertions --list` | List the available scenario assertions (generated from the schema) | "what can I assert?" without grepping the source |
|
|
407
407
|
| `decide` | Validate a decider against a sample question in ~2 s (no run) | sanity-check a `--decider-*` / `--answer` wiring before a long run |
|
|
@@ -665,7 +665,7 @@ jobs:
|
|
|
665
665
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
666
666
|
```
|
|
667
667
|
|
|
668
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.0
|
|
668
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.1.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
|
|
669
669
|
|
|
670
670
|
The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
671
671
|
|
package/SPEC.md
CHANGED
|
@@ -674,6 +674,34 @@ already-redacted markers. The TLD list used by the domain scanner was also exten
|
|
|
674
674
|
entries (CB-5), adding major European, Asian, and Latin American ccTLDs
|
|
675
675
|
(`ch|nl|se|no|it|jp|br|nz|in|sg|kr|mx|es|pt|pl|be|at|dk|fi|ie|ru|cn|tw|hu|cz|ro|il|za|ar|cl|pe|tr`).
|
|
676
676
|
|
|
677
|
+
### 11.2 `doctor` — dedicated envelope
|
|
678
|
+
|
|
679
|
+
`doctor [--tier <t>]` is the read-only prerequisite check ("can I run the live tiers — what's
|
|
680
|
+
missing?"). Its `--output-format json` output does **not** reuse the `run/skill/replay` envelope (no
|
|
681
|
+
`RunResult` to judge) — it emits its own, published as **`schema/doctor.json`** (a §12-covered contract
|
|
682
|
+
surface):
|
|
683
|
+
|
|
684
|
+
```jsonc
|
|
685
|
+
// Completed probe (normal path — routed through the shared jsonPayloadEnvelope, so error is always null):
|
|
686
|
+
{ "tool": "cowork-harness", "version": "...", "command": "doctor",
|
|
687
|
+
"ok": true, // false iff any `required` check has status:"fail" for the selected tier
|
|
688
|
+
"error": null,
|
|
689
|
+
"tier": "container", // protocol|container|microvm|hostloop|cowork
|
|
690
|
+
"checks": [ { "id": "node", "title": "Node ≥ 20", "status": "ok", "detail": "node 22.x.x",
|
|
691
|
+
"remedy?": "string", "required": true } ] } // status: ok|fail|warn|skip
|
|
692
|
+
|
|
693
|
+
// Shared error envelope (a thrown failure — bad flag, or the top-level catch):
|
|
694
|
+
{ "tool": "cowork-harness", "version": "...", "command": "doctor", "ok": false, "results": [],
|
|
695
|
+
"error": { "category": "usage|unanswered|boundary|runtime|internal", "message": "string", "hint?": "string" } }
|
|
696
|
+
```
|
|
697
|
+
|
|
698
|
+
`ok = !checks.some(c => c.required && c.status === "fail")`. **Exit codes:** `0` all required checks
|
|
699
|
+
pass · `1` a required check is `status:"fail"` for the selected tier (the completed-probe shape, still
|
|
700
|
+
`error:null`) · `2` usage (bad `--tier`/unexpected args) or an unexpected internal failure caught by the
|
|
701
|
+
top-level catch (`category:"internal"`). The `checks[].id` set is **not** enumerated as a closed
|
|
702
|
+
contract — it grows as tiers/checks are added, the same way `trace` row shapes are excluded from §12's
|
|
703
|
+
covered list.
|
|
704
|
+
|
|
677
705
|
## 12. Versioning & the 1.0 compatibility contract
|
|
678
706
|
|
|
679
707
|
From `1.0.0` the project follows [semver](https://semver.org/). The surfaces below are the **covered
|
|
@@ -698,6 +726,12 @@ Covered-surface changes follow semver as of `1.0.0` — see [RELEASING.md](./REL
|
|
|
698
726
|
`findings` / `staleness` / `unverifiable` / `notes` / `version` / `error` channels, and the exit-code
|
|
699
727
|
split (`0`/`1`/`2`/`3`, §11) they map to. This is the machine output the CI recipes and the packaged
|
|
700
728
|
Action steer consumers to parse; renaming or removing a key is breaking, adding one is not.
|
|
729
|
+
- **`doctor` envelope** — `schema/doctor.json` under `--output-format json` (§11.2): the completed-probe
|
|
730
|
+
shape (`tool` / `version` / `command` / `ok` / `error:null` / `tier` / `checks[]`, each check's
|
|
731
|
+
`id` / `title` / `status` / `detail` / `required` / optional `remedy`) and the shared error-envelope
|
|
732
|
+
shape (`results:[]` / `error.category`) it falls back to on a thrown failure; renaming or removing a
|
|
733
|
+
key is breaking, adding one is not. The `checks[].id` set itself is NOT covered — it grows with new
|
|
734
|
+
tiers/checks.
|
|
701
735
|
- **Cassette format** — the current `cassetteVersion` (**10**, `schema/cassette.v10.json`) and its
|
|
702
736
|
verdict-modifier assertion keys. The minimum supported read version is **v9**
|
|
703
737
|
(`MIN_SUPPORTED_CASSETTE_VERSION`): a cassette below the floor is refused at load time with a
|
package/dist/cli.js
CHANGED
|
@@ -406,7 +406,7 @@ const SUBCOMMAND_USAGE = {
|
|
|
406
406
|
rehash: "usage: rehash <dir/> [--dry-run] [--output-format text|json] (migrate cassettes across format bumps using contentSig verification; no re-record needed)",
|
|
407
407
|
prune: "usage: prune [--keep-last <n>] [--pinned-older-than <N>d|h|m] [--dry-run] [<runs-dir>] (prune accumulated run dirs; default --keep-last 5)",
|
|
408
408
|
"init-redact": "usage: init-redact [--force] [--output-format json] (copy the packaged reference .cowork-redact.json into the cwd; refuses to overwrite an existing one without --force)",
|
|
409
|
-
"analyze-skill": "usage: analyze-skill <SKILL.md | skill-dir/ | glob>… [--strict] [--output-format text|json] (ADVISORY token-free scan: flags a /sessions/... path handed to a file tool or dispatch/sub-agent output — denied on host-loop)\n" +
|
|
409
|
+
"analyze-skill": "usage: analyze-skill <SKILL.md | skill-dir/ | glob>… [--strict] [--runtime] [--output-format text|json] (ADVISORY token-free scan: flags a /sessions/... path handed to a file tool or dispatch/sub-agent output — denied on host-loop; AND interactive-artifact write-backs that are lost under Cowork)\n" +
|
|
410
410
|
" reuses the harness's own ported /sessions path-gate predicate (isVmSessionsPath) as the deny decision; only the EXTRACTION from SKILL.md text is heuristic.\n" +
|
|
411
411
|
" a DIRECTORY target scans the UNION of every contract-bearing markdown file present, not just SKILL.md — a plugin's dispatch/output contracts often live in agents/*.md or references/*.md:\n" +
|
|
412
412
|
" · top-level SKILL.md present → + top-level references/**/*.md\n" +
|
|
@@ -414,10 +414,11 @@ const SUBCOMMAND_USAGE = {
|
|
|
414
414
|
" · a skill dir INSIDE a plugin (SKILL.md present, walking up finds an enclosing plugin manifest) → + that plugin's root agents/*.md\n" +
|
|
415
415
|
' MULTIPLE positionals are accepted (matches lint-skill\'s nargs="+"), each resolved independently (file/dir via the rules above, or a `*`-bearing target via a small hand-rolled glob matcher — `dir/*.md` shallow, `dir/**/*.md` recursive; no dependency, no engines bump).\n' +
|
|
416
416
|
" every positional's file set is UNIONed and deduped by resolved absolute path across the WHOLE invocation (a file reached both directly and via a dir/glob is analyzed once). ZERO scannable files across ALL positionals is a USAGE ERROR (exit 2), never a silent clean pass; an unresolvable single positional (bad path, unrecognized glob shape) fails the whole invocation the same way.\n" +
|
|
417
|
-
" findings are WARNINGS: exit 0 by default even when findings are printed. --strict fails (exit 1) if ANY scanned file (across ALL positionals) has an unsuppressed finding (mirrors lint-skill's --strict). exit 2 = usage error.\n" +
|
|
417
|
+
" findings are WARNINGS: exit 0 by default even when findings are printed. --strict fails (exit 1) if ANY scanned file (across ALL positionals) has an unsuppressed error-severity finding (mirrors lint-skill's --strict). exit 2 = usage error. exit 3 = could-not-verify (an artifact candidate that couldn't be parsed/analyzed — always, --strict-independent; a strict error finding's exit 1 takes precedence).\n" +
|
|
418
|
+
" interactive-artifact write-backs (Item 1): also scans .html/.js/.ts/.jsx/.tsx/.py sources under the target for a relative fetch/XHR/sendBeacon/<form method=post> write-back that is LOST under Cowork (the artifact is served from Cowork's own origin, so a relative write-back resolves non-ok and a page that doesn't check resp.ok shows a false 'Saved'). Rules: artifact-write-back-lost (error, gates under --strict), artifact-write-back-suspect (advisory), plus a could-not-verify channel. These are NOT silenced by the `analyze-skill: ignore` marker. --runtime adds an OPTIONAL headless-DOM confirmation (needs jsdom, dynamically imported; ENRICHMENT only — never changes the exit code; trusted-source scope).\n" +
|
|
418
419
|
" plain `analyze-skill` (no --strict) is ADVISORY ONLY — CI should invoke `analyze-skill ... --strict` to actually gate on findings, same as `lint-skill --strict`.\n" +
|
|
419
420
|
" a line containing `analyze-skill: ignore` (bare, or inside an HTML comment) anywhere in one of the scanned files silences EVERY path-fidelity finding for THAT file — even under --strict, even a genuine true positive (an explicit author override); it does not affect other scanned files.\n" +
|
|
420
|
-
" --output-format json emits {tool,version,command,ok,files:[{file,findings,suppressed}],scanned,unscanned,strict,error} — NOT a bare array and NOT results[]; `scanned`/`unscanned` name the files analyzed and any contract dirs left out of scope (printed as a banner in text mode too).\n" +
|
|
421
|
+
" --output-format json emits {tool,version,command,ok,files:[{file,findings,suppressed}],scanned,unscanned,analysisFailures,[runtimeConfirmations (only with --runtime)],strict,error} — NOT a bare array and NOT results[]; `scanned`/`unscanned` name the files analyzed and any contract dirs left out of scope (printed as a banner in text mode too); each finding now carries a `severity` (error|advisory).\n" +
|
|
421
422
|
" a clean/suppressed result is a PRE-FLIGHT signal only — the runtime no_vm_path_file_op / vm_path_denied asserts remain authoritative (see docs/subagents.md).",
|
|
422
423
|
};
|
|
423
424
|
// Known subcommands — used by the global value-flag parsers (`--dotenv`, `--run-dir`) to reject a command
|