cowork-harness 3.0.0 → 3.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.0.0
7
- tracks-harness: cowork-harness 3.0.0 (baseline desktop-1.40609.0)
6
+ version: 3.0.1
7
+ tracks-harness: cowork-harness 3.0.1 (baseline desktop-1.40609.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.1` (baseline
29
29
  > `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.0"`. **Pin `@^3.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.1"`. **Pin `@^3.0.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.0.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.0.1"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^3.0.0"
70
+ - run: npm i -g "cowork-harness@^3.0.1"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^3.0.0"
345
+ - run: npm i -g "cowork-harness@^3.0.1"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^3.0.0"
374
+ run: npm i -g "cowork-harness@^3.0.1"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -20,7 +20,8 @@ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.406
20
20
  - A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
21
21
  true` — with no container around the native file tools, that combination gives the agent genuine,
22
22
  software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
23
- - A `protocol` scenario staging a plugin that declares runnable hooks needs `allow_host_hooks: true`
23
+ - A `protocol` scenario staging a plugin that declares runnable hooks — in `<plugin>/hooks/hooks.json`
24
+ or the plugin manifest's `hooks` key, both live — needs `allow_host_hooks: true`
24
25
  (`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
25
26
  hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
26
27
  Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
@@ -155,6 +156,60 @@ hands `fn` exactly this dict.
155
156
  - `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
156
157
  `["fail", "prompt", "llm", "first"]`.
157
158
 
159
+ ### A script-running skill trips `permissive_auto_allow` — at `protocol` specifically
160
+
161
+ The default `permission_parity: cowork` auto-allows an unscripted, off-registry tool ask — and records
162
+ it, because real Cowork would have BLOCKED for the user. `computeVerdict` then FAILS the run: a green
163
+ carrying one would not be a faithful pass. Only `Read`, `Glob` and `Grep` are default-allow; **`Bash` is
164
+ off-registry**.
165
+
166
+ Two things decide whether you ever see it, and neither is obvious. The harness only decides how to
167
+ ANSWER a permission ask; whether the agent ASKS is the agent binary's own logic, and that varies by
168
+ both the command and the tier. All rows below are measured, same scenario shape, same session:
169
+
170
+ | Tier | Bash command | tool used | ask | verdict |
171
+ |---|---|---|---|---|
172
+ | `protocol` | `echo hello` | `Bash` | none | ✓ green |
173
+ | `protocol` | `python3 -c "print(42)"` | `Bash` | **one** | ✗ `permissive_auto_allow` |
174
+ | `container` | `python3 -c "print(42)"` | `Bash` | none | ✓ green |
175
+ | `hostloop` | `python3 -c "print(42)"` | `mcp__workspace__bash` | none | ✓ green |
176
+
177
+ So the guard is **not** a general hazard for script-running skills — it is a `protocol` one. The host
178
+ CLI that L0 spawns asks for a command like `python3 -c …` and not for `echo`; the staged agent at
179
+ `container` asks for neither, and `hostloop` routes shell through `mcp__workspace__bash` and so is not
180
+ even the same tool. A skill that runs `python3 ${CLAUDE_SKILL_DIR}/scripts/…` therefore goes red at
181
+ `protocol` on an otherwise-correct run, while the same skill is green on the sandboxed tiers — and a
182
+ hello-world probe is green everywhere and teaches the wrong expectation.
183
+
184
+ Three ways through, in preference order: script the gate (`--answer` / `answers:`), which is what the
185
+ warning tells you and keeps the run deterministic; set `permission_parity: strict` to deny instead of
186
+ allow, if refusal is what you want to test; or assert `allow_permissive_auto_allow: true` when the
187
+ permissive behaviour is deliberately what the scenario is about.
188
+
189
+ > **What is NOT established.** Why the host CLI asks for one command and not another was observed, not
190
+ > traced — do not infer a rule for which commands ask from these rows. `microvm` was not measured.
191
+
192
+ ### `baseline agent binary not found` — a Desktop update pruned the pinned ELF
193
+
194
+ A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
195
+ a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
196
+ current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
197
+
198
+ Prefer **repinning `baseline:` to an installed version** for anything
199
+ reproducibility-bound — you keep an exact pin and move it deliberately. `baseline: latest` never rots
200
+ but silently drifts, so two runs weeks apart are not comparable; a pin and `latest` have opposite
201
+ failure modes.
202
+
203
+ To find which versions you actually have, do NOT use `cowork-harness list` — it enumerates the baseline
204
+ definitions shipped with the harness, which are present regardless of what Desktop pruned from this
205
+ machine, so a pruned pin lists as healthy. Test the staged binary: read `agentBinary.stagedPath` from
206
+ the baseline JSON, expand the leading `~`, and check it is a FILE — the pruned case leaves the empty
207
+ directory behind, so a directory test passes on exactly the case that fails.
208
+
209
+ `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` runs the newest sibling instead of the pinned
210
+ binary and downgrades the sha check to advisory — that is the substitution the hard failure exists to
211
+ prevent, so use it to unblock once, never in CI.
212
+
158
213
  ### Determinism contract
159
214
 
160
215
  - `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.1`
4
4
  (baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -100,7 +100,8 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
100
100
  # Read-only folders and folder-less runs need no opt-in.
101
101
 
102
102
  allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
103
- # declares runnable hooks (`<plugin>/hooks/hooks.json`): L0 passes
103
+ # declares runnable hooks — either `<plugin>/hooks/hooks.json` OR
104
+ # the manifest's `hooks` key: L0 passes
104
105
  # --plugin-dir, so the CLI executes those hooks as NATIVE HOST
105
106
  # processes under your account, with no container sandbox. A plugin
106
107
  # that declares no hooks needs no opt-in, and a misplaced root-level
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/AGENTS.md CHANGED
@@ -22,7 +22,7 @@ protocol layer or run-loop bookkeeping in the CLI.
22
22
  Don't add a test that needs a live model or Docker to the default suite; that's the `pytest -m cowork` /
23
23
  `npm run test:live` lane. Python fast lane (from `python/`): `pytest -m 'not cowork'`.
24
24
  - CLI binary `cowork-harness`; env vars `COWORK_HARNESS_*` (+ `COWORK_AGENT_BINARY` / `COWORK_AGENT_IMAGE`) —
25
- see README's [Reproducibility knobs](./README.md#reproducibility-knobs) for the full env-var list.
25
+ see README's [Reproducibility knobs](./docs/cli.md#reproducibility-knobs) for the full env-var list.
26
26
  Node ≥ 22.
27
27
  - `cowork-harness sync` is **local-only** (needs Desktop + `app.asar`; not on CI). The committed
28
28
  `baselines/*.json` are CI's source of truth — never hand-edit release facts into source; they come from
package/CHANGELOG.md CHANGED
@@ -6,6 +6,105 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.0.1] — 2026-08-30
10
+
11
+ ### Changed
12
+
13
+ - **The README now shows the product, not just the argument for it.** Three independent reviews found
14
+ the same gap: for a test harness, the file a user authors is the most persuasive thing the project
15
+ owns, and the router split had left it with zero examples. Adds a worked scenario (prompt, scripted
16
+ answers, assertions) with its verdict output — the scenario is extracted and linted in place, so it is
17
+ a real file rather than plausible-looking YAML — plus a "What that catches" section built from an
18
+ actual run where a named skill was offered, declined, and the green answer looked fine anyway
19
+ (`skillsInvoked: []`, `toolCounts: {}`) — framed as a closed loop, since the author fixed and re-verified
20
+ it with the same instrument roughly ninety minutes later. An outcome example is a claim about a MOMENT;
21
+ the better the tool works the faster its own examples get fixed out from under it, so it needs a tense. Also names two things the docs asserted but never taught:
22
+ diffing the same scenario across two fidelity tiers as a discovery technique, and proving an assertion
23
+ can fail before trusting it. The requirements block moves below the `claude -p` argument — three
24
+ reviewers independently reported reading install prerequisites before any reason to want them.
25
+
26
+ - **The README is now a router, not the whole manual.** It was 975 lines carrying three unrelated
27
+ audiences at once; it is now **282** and branches to per-audience pages: **`docs/cli.md`** (install,
28
+ prerequisites, commands, the two files, run output, env knobs), **`docs/companion-skill.md`** (install
29
+ + orientation; usage stays in `SKILL.md`), and **`docs/ci.md`** (the token-free gate, the packaged
30
+ Action, the live lane — written to stand alone). The harness's own pipeline and contributor suite moved
31
+ to `CONTRIBUTING.md`, which keeps `docs/ci.md` about *consuming* the Action rather than about this repo.
32
+ README keeps what all three audiences share: the fidelity-tier vocabulary, architecture, limitations,
33
+ the docs index, and versioning. Content moved verbatim — this is a relocation, not a rewrite.
34
+
35
+ Nothing is dropped: `Sandboxing` folded into `docs/boundary.md` and `Maintenance` into
36
+ `docs/maintenance.md`; `Discovery` was already covered in `docs/discovery.md`, so it was deleted rather
37
+ than duplicated. Every moved relative link was rewritten for its new depth, and every deep link into a
38
+ moved section was repointed across `docs/`, `examples/README.md`, `llms.txt` and the docs index.
39
+
40
+ ### Fixed
41
+
42
+ - **`bump-version` did not know about the router-split pages.** `docs/cli.md`, `docs/ci.md` and
43
+ `docs/companion-skill.md` carry `cowork-harness@^X.Y.Z` install floors that `check:versions` enforces,
44
+ but the bump tool's target list predated them — so `npm run bump` rewrote every other floor and then
45
+ failed its own post-write lockstep check. Caught at release time by that self-check rather than
46
+ shipping a repo with mismatched floors. All three are now registered, and the test pinning that list
47
+ carries the reason.
48
+
49
+
50
+ - **22 dead anchor links introduced by the README router split, and the guard gap that let them ship.**
51
+ Moving 11 `##` sections out of README left every `](#slug)` link pointing at them dangling — 19 in
52
+ README (including two badges and the two nav lines the page opened with), plus one each in `AGENTS.md`,
53
+ `docs/ci.md` and `docs/companion-skill.md`. GitHub fails these silently: the page simply does not
54
+ scroll. Every file-path link still resolved and every existing guard stayed green, because the anchor
55
+ suite validated links *into* README and *into* sibling docs but never a page's links to its **own**
56
+ headings. `test/repo-docs-anchors.test.ts` now covers that third case across root pages and `docs/`,
57
+ with a canary against a checker that extracts nothing, and was confirmed to fail against the exact
58
+ regression that shipped. The two nav lines are deleted rather than repointed — they were a
59
+ table-of-contents for a 975-line page, and "Pick your path" plus the Documentation table now do that job.
60
+
61
+ Found by two independent reviewers; the second caught the three non-README instances a README-scoped
62
+ fix would have missed.
63
+
64
+ Follow-up from the same review: three cross-page links still read "…\[Prerequisites](./docs/cli.md#…)
65
+ **below**" — resolving correctly while the sentence lied, because Prerequisites had moved to another
66
+ page. Reworded, and guarded: a cross-page link followed closely by "below"/"above" now fails the suite.
67
+ This half is the nastier one — the link works, so every link checker stays green forever.
68
+
69
+ - **The `protocol` host-hook consent gate was defeatable by a spelling choice.** `allow_host_hooks`
70
+ (3.0.0) refuses a spawn until the operator consents to a staged plugin's hooks running as native host
71
+ processes — but detection keyed only on a file named `hooks.json`, and a plugin may equally declare its
72
+ hooks in the manifest's `hooks` key (`.claude-plugin/plugin.json`, or a bare `plugin.json`), which is
73
+ the more common spelling. Such a plugin sailed past the gate: reproduced end-to-end, the hook executed
74
+ under the operator's own account while the run went green, no consent was asked and no disclosure
75
+ printed. The same blind predicate feeds the tier-independent disclosure, so those hooks also ran
76
+ unannounced at `hostloop`. Detection now returns the UNION of both channels, which fixes the gate and
77
+ the disclosure together since both call through it. The `hooks/hooks.json` placement carve-out is
78
+ deliberately NOT extended to the manifest — that carve-out exists because a misplaced `hooks.json` is
79
+ inert, and a manifest-declared hook is live wherever the manifest sits.
80
+
81
+ Reported by a consumer during a 3.0.0 adoption pass, with the defect proven from committed cassettes:
82
+ 10 of theirs carried a `SessionStart:startup` hook_started/hook_response pair from a manifest-only
83
+ declaration, including one recorded at `container`.
84
+
85
+ ### Documentation
86
+
87
+ - **The pruned-agent-binary failure now has a documented remedy, with the tradeoffs named.** A Desktop
88
+ update deletes the prior version's staged ELF while often leaving an empty version directory, so a
89
+ scenario pinning that agent version dies with `baseline agent binary not found`. The code anticipates
90
+ this by name, but the remedy appeared in no skill surface at all — and an installed companion skill
91
+ ships without repo `docs/`, so the agent most likely to hit it had nothing to read. Now in
92
+ `docs/gotchas.md` and the skill's own reference, stating that the three remedies are NOT equivalent:
93
+ repin for a reproducibility-bound suite, `latest` for a one-off (a pin rots silently, `latest` drifts
94
+ silently — opposite failure modes), and `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` only to unblock once,
95
+ never in CI, since it runs the newest sibling and downgrades the sha check to advisory.
96
+
97
+ - **`permissive_auto_allow` and script-running skills is a `protocol`-specific hazard, not a general one.**
98
+ `Bash` is off-registry (only `Read`/`Glob`/`Grep` are default-allow), but the harness only decides how to
99
+ ANSWER a permission ask — whether the agent ASKS varies by command AND tier. Measured: at `protocol`, `echo
100
+ hello` produces no ask (green) while `python3 -c "print(42)"` produces one and fails the guard; the
101
+ identical command at `container` produces no ask (green), and `hostloop` routes shell through
102
+ `mcp__workspace__bash` entirely. So a skill running `python3 ${CLAUDE_SKILL_DIR}/scripts/…` goes red at L0
103
+ and green on the sandboxed tiers, while a hello-world probe is green everywhere and teaches the wrong
104
+ expectation. `references/fidelity-and-answers.md` carries the measured table and the three ways through; why
105
+ the host CLI asks for one command and not another is explicitly left untraced, and `microvm` is named as
106
+ unmeasured.
107
+
9
108
  ## [3.0.0] — 2026-08-29
10
109
 
11
110
  ### Breaking
package/CONTRIBUTING.md CHANGED
@@ -108,3 +108,38 @@ When a Desktop release moves something `sync` doesn't read, it reports an `unkno
108
108
  ## Reporting issues
109
109
 
110
110
  Use the issue templates. For anything security/sandbox related, see [SECURITY.md](./SECURITY.md).
111
+
112
+ ## This repo's own CI pipeline
113
+
114
+ Contributor-facing. For *consuming* the harness in your own CI — the token-free gate and the packaged
115
+ Action — see [docs/ci.md](./docs/ci.md).
116
+
117
+ The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
118
+
119
+ | Stage | Runs | Needs | Gates |
120
+ |---|---|---|---|
121
+ | **build** | format check · version-lockstep guard · typecheck · source guards · build · CLI smoke · token-free `replay` · `verify-cassettes` · `lint` | nothing | every push/PR |
122
+ | **test** | the unit suite (vitest), sharded 4-way | nothing | every push/PR |
123
+ | **floor** | the unit suite once, unsharded, on Node 22 — the version `engines.node` declares — so the floor is exercised rather than asserted (other jobs run Node 24, the Active LTS line) | nothing | every push/PR; gates the merge context |
124
+ | **action-self-test** | packs this commit and runs the packaged `uses: ./` Action across its full case set — pass (committed example cassette), usage-error fail (nonexistent path), assertion fail (checks the reporter renders a ❌ row), `lint`, and `analyze-skill`; `ci.yml` currently carries 11 `command:` invocations, so re-count here rather than trusting this sentence | nothing | every push/PR |
125
+ | **python** | `pytest` helper self-checks (`python/`, run with `-m 'not cowork'` — the token-free subset; the Docker/token `@pytest.mark.cowork` tests are excluded) | nothing (token-free assertions only) | every push/PR |
126
+ | **boundary** | builds the pinned agent image, brings up the default-deny network, runs `boundary-check`, then `npm run test:live` (live contract tests that guard the binary-resolution assumptions, no token needed) | Docker, arm64 runner | proves the sandbox enforces Cowork's limits — **no API key** |
127
+ | **image-recipe** | compiles `docker/Dockerfile.agent` in **both** variants (lean/core and `COWORK_FULL_PARITY=1`) on an arm64 runner, so a recipe change cannot reach the merge gate uncompiled | Docker, arm64 runner | every push/PR; gates the merge context |
128
+ | **scenarios** | the live scenario suite (mixed `protocol` + `container` fidelity across `examples/scenarios/`), plus the `e2e/scenarios/*.yaml` smoke set (this repo's own L0/L1/hostloop self-tests, `microvm` excluded — needs a real VM); uploads transcripts/egress logs as artifacts; relies on a runner-local staged agent binary (no in-workflow download step). | `ANTHROPIC_API_KEY` | fork PRs: the whole job is skipped (`if:` guard); same-repo without a key: warns and exits 0 |
129
+ | **parity-drift** | reminder to re-`sync` when Desktop updates | nothing | **goes red** if the newest committed baseline is &gt; 90 days old, but sits outside `ci-green`'s `needs:` list, so a red run does not block a merge |
130
+
131
+ This ordering means cheap checks fail fast, the **boundary parity gate runs without secrets** (so forks get it too), and expensive live runs only happen when a key is present.
132
+
133
+
134
+ ## The harness's own suite
135
+
136
+ ```bash
137
+ npm run ci # typecheck + build + test (run format:check separately; NOT the same set as CI's `build` job — see CONTRIBUTING.md)
138
+ npm test # vitest: decider, egress allowlist, launch plan, example validation
139
+ cowork-harness boundary-check # self-verify the sandbox (needs Docker; not part of `npm run ci`)
140
+ ```
141
+
142
+ Unit tests cover the scripted-answer logic, the egress allowlist matcher, the session→launch-plan materialization (mounts + discovery settings + env-strip), and a **schema guard** that fails if any shipped baseline/session/scenario stops validating. Add a test alongside any new schema field or `Decider` rule — see [CONTRIBUTING.md](./CONTRIBUTING.md).
143
+
144
+ > Copy your starting scenarios/sessions from **`examples/`**. The **`e2e/`** directory is the harness's *own* fidelity self-tests (smoke scenarios per tier) — not a template to copy.
145
+