cowork-harness 3.0.0 → 3.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +4 -4
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +57 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/AGENTS.md +1 -1
- package/CHANGELOG.md +99 -0
- package/CONTRIBUTING.md +35 -0
- package/README.md +86 -703
- package/dist/run/hook-events.js +38 -5
- package/docs/README.md +9 -6
- package/docs/boundary.md +7 -0
- package/docs/cassette.md +1 -1
- package/docs/ci.md +114 -0
- package/docs/cli.md +532 -0
- package/docs/companion-skill.md +58 -0
- package/docs/debugging.md +3 -3
- package/docs/fidelity-gaps.md +10 -1
- package/docs/gotchas.md +27 -2
- package/docs/maintenance.md +28 -0
- package/docs/session.md +1 -1
- package/docs/stats.md +1 -1
- package/examples/README.md +3 -3
- package/examples/replays/README.md +1 -1
- package/llms.txt +3 -0
- package/package.json +1 -1
- package/scripts/bump-version.ts +12 -1
- package/scripts/check-versions.ts +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.0.
|
|
7
|
-
tracks-harness: cowork-harness 3.0.
|
|
6
|
+
version: 3.0.1
|
|
7
|
+
tracks-harness: cowork-harness 3.0.1 (baseline desktop-1.40609.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.1` (baseline
|
|
29
29
|
> `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.1"`. **Pin `@^3.0.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.0.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.0.
|
|
20
|
+
(e.g. `version: "3.0.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^3.0.
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.0.1"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -342,7 +342,7 @@ jobs:
|
|
|
342
342
|
with: { node-version: '24' }
|
|
343
343
|
- uses: actions/setup-python@v5
|
|
344
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^3.0.
|
|
345
|
+
- run: npm i -g "cowork-harness@^3.0.1"
|
|
346
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +371,7 @@ jobs:
|
|
|
371
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
372
|
fi
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^3.0.
|
|
374
|
+
run: npm i -g "cowork-harness@^3.0.1"
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.0.
|
|
3
|
+
Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.0.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -20,7 +20,8 @@ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.406
|
|
|
20
20
|
- A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
|
|
21
21
|
true` — with no container around the native file tools, that combination gives the agent genuine,
|
|
22
22
|
software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
|
|
23
|
-
- A `protocol` scenario staging a plugin that declares runnable hooks
|
|
23
|
+
- A `protocol` scenario staging a plugin that declares runnable hooks — in `<plugin>/hooks/hooks.json`
|
|
24
|
+
or the plugin manifest's `hooks` key, both live — needs `allow_host_hooks: true`
|
|
24
25
|
(`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
|
|
25
26
|
hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
|
|
26
27
|
Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
|
|
@@ -155,6 +156,60 @@ hands `fn` exactly this dict.
|
|
|
155
156
|
- `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
|
|
156
157
|
`["fail", "prompt", "llm", "first"]`.
|
|
157
158
|
|
|
159
|
+
### A script-running skill trips `permissive_auto_allow` — at `protocol` specifically
|
|
160
|
+
|
|
161
|
+
The default `permission_parity: cowork` auto-allows an unscripted, off-registry tool ask — and records
|
|
162
|
+
it, because real Cowork would have BLOCKED for the user. `computeVerdict` then FAILS the run: a green
|
|
163
|
+
carrying one would not be a faithful pass. Only `Read`, `Glob` and `Grep` are default-allow; **`Bash` is
|
|
164
|
+
off-registry**.
|
|
165
|
+
|
|
166
|
+
Two things decide whether you ever see it, and neither is obvious. The harness only decides how to
|
|
167
|
+
ANSWER a permission ask; whether the agent ASKS is the agent binary's own logic, and that varies by
|
|
168
|
+
both the command and the tier. All rows below are measured, same scenario shape, same session:
|
|
169
|
+
|
|
170
|
+
| Tier | Bash command | tool used | ask | verdict |
|
|
171
|
+
|---|---|---|---|---|
|
|
172
|
+
| `protocol` | `echo hello` | `Bash` | none | ✓ green |
|
|
173
|
+
| `protocol` | `python3 -c "print(42)"` | `Bash` | **one** | ✗ `permissive_auto_allow` |
|
|
174
|
+
| `container` | `python3 -c "print(42)"` | `Bash` | none | ✓ green |
|
|
175
|
+
| `hostloop` | `python3 -c "print(42)"` | `mcp__workspace__bash` | none | ✓ green |
|
|
176
|
+
|
|
177
|
+
So the guard is **not** a general hazard for script-running skills — it is a `protocol` one. The host
|
|
178
|
+
CLI that L0 spawns asks for a command like `python3 -c …` and not for `echo`; the staged agent at
|
|
179
|
+
`container` asks for neither, and `hostloop` routes shell through `mcp__workspace__bash` and so is not
|
|
180
|
+
even the same tool. A skill that runs `python3 ${CLAUDE_SKILL_DIR}/scripts/…` therefore goes red at
|
|
181
|
+
`protocol` on an otherwise-correct run, while the same skill is green on the sandboxed tiers — and a
|
|
182
|
+
hello-world probe is green everywhere and teaches the wrong expectation.
|
|
183
|
+
|
|
184
|
+
Three ways through, in preference order: script the gate (`--answer` / `answers:`), which is what the
|
|
185
|
+
warning tells you and keeps the run deterministic; set `permission_parity: strict` to deny instead of
|
|
186
|
+
allow, if refusal is what you want to test; or assert `allow_permissive_auto_allow: true` when the
|
|
187
|
+
permissive behaviour is deliberately what the scenario is about.
|
|
188
|
+
|
|
189
|
+
> **What is NOT established.** Why the host CLI asks for one command and not another was observed, not
|
|
190
|
+
> traced — do not infer a rule for which commands ask from these rows. `microvm` was not measured.
|
|
191
|
+
|
|
192
|
+
### `baseline agent binary not found` — a Desktop update pruned the pinned ELF
|
|
193
|
+
|
|
194
|
+
A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
|
|
195
|
+
a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
|
|
196
|
+
current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
|
|
197
|
+
|
|
198
|
+
Prefer **repinning `baseline:` to an installed version** for anything
|
|
199
|
+
reproducibility-bound — you keep an exact pin and move it deliberately. `baseline: latest` never rots
|
|
200
|
+
but silently drifts, so two runs weeks apart are not comparable; a pin and `latest` have opposite
|
|
201
|
+
failure modes.
|
|
202
|
+
|
|
203
|
+
To find which versions you actually have, do NOT use `cowork-harness list` — it enumerates the baseline
|
|
204
|
+
definitions shipped with the harness, which are present regardless of what Desktop pruned from this
|
|
205
|
+
machine, so a pruned pin lists as healthy. Test the staged binary: read `agentBinary.stagedPath` from
|
|
206
|
+
the baseline JSON, expand the leading `~`, and check it is a FILE — the pruned case leaves the empty
|
|
207
|
+
directory behind, so a directory test passes on exactly the case that fails.
|
|
208
|
+
|
|
209
|
+
`COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` runs the newest sibling instead of the pinned
|
|
210
|
+
binary and downgrades the sha check to advisory — that is the substitution the hard failure exists to
|
|
211
|
+
prevent, so use it to unblock once, never in CI.
|
|
212
|
+
|
|
158
213
|
### Determinism contract
|
|
159
214
|
|
|
160
215
|
- `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.1`
|
|
4
4
|
(baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -100,7 +100,8 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
|
|
|
100
100
|
# Read-only folders and folder-less runs need no opt-in.
|
|
101
101
|
|
|
102
102
|
allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
|
|
103
|
-
# declares runnable hooks
|
|
103
|
+
# declares runnable hooks — either `<plugin>/hooks/hooks.json` OR
|
|
104
|
+
# the manifest's `hooks` key: L0 passes
|
|
104
105
|
# --plugin-dir, so the CLI executes those hooks as NATIVE HOST
|
|
105
106
|
# processes under your account, with no container sandbox. A plugin
|
|
106
107
|
# that declares no hooks needs no opt-in, and a misplaced root-level
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.0.
|
|
5
|
+
Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/AGENTS.md
CHANGED
|
@@ -22,7 +22,7 @@ protocol layer or run-loop bookkeeping in the CLI.
|
|
|
22
22
|
Don't add a test that needs a live model or Docker to the default suite; that's the `pytest -m cowork` /
|
|
23
23
|
`npm run test:live` lane. Python fast lane (from `python/`): `pytest -m 'not cowork'`.
|
|
24
24
|
- CLI binary `cowork-harness`; env vars `COWORK_HARNESS_*` (+ `COWORK_AGENT_BINARY` / `COWORK_AGENT_IMAGE`) —
|
|
25
|
-
see README's [Reproducibility knobs](./
|
|
25
|
+
see README's [Reproducibility knobs](./docs/cli.md#reproducibility-knobs) for the full env-var list.
|
|
26
26
|
Node ≥ 22.
|
|
27
27
|
- `cowork-harness sync` is **local-only** (needs Desktop + `app.asar`; not on CI). The committed
|
|
28
28
|
`baselines/*.json` are CI's source of truth — never hand-edit release facts into source; they come from
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,105 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.0.1] — 2026-08-30
|
|
10
|
+
|
|
11
|
+
### Changed
|
|
12
|
+
|
|
13
|
+
- **The README now shows the product, not just the argument for it.** Three independent reviews found
|
|
14
|
+
the same gap: for a test harness, the file a user authors is the most persuasive thing the project
|
|
15
|
+
owns, and the router split had left it with zero examples. Adds a worked scenario (prompt, scripted
|
|
16
|
+
answers, assertions) with its verdict output — the scenario is extracted and linted in place, so it is
|
|
17
|
+
a real file rather than plausible-looking YAML — plus a "What that catches" section built from an
|
|
18
|
+
actual run where a named skill was offered, declined, and the green answer looked fine anyway
|
|
19
|
+
(`skillsInvoked: []`, `toolCounts: {}`) — framed as a closed loop, since the author fixed and re-verified
|
|
20
|
+
it with the same instrument roughly ninety minutes later. An outcome example is a claim about a MOMENT;
|
|
21
|
+
the better the tool works the faster its own examples get fixed out from under it, so it needs a tense. Also names two things the docs asserted but never taught:
|
|
22
|
+
diffing the same scenario across two fidelity tiers as a discovery technique, and proving an assertion
|
|
23
|
+
can fail before trusting it. The requirements block moves below the `claude -p` argument — three
|
|
24
|
+
reviewers independently reported reading install prerequisites before any reason to want them.
|
|
25
|
+
|
|
26
|
+
- **The README is now a router, not the whole manual.** It was 975 lines carrying three unrelated
|
|
27
|
+
audiences at once; it is now **282** and branches to per-audience pages: **`docs/cli.md`** (install,
|
|
28
|
+
prerequisites, commands, the two files, run output, env knobs), **`docs/companion-skill.md`** (install
|
|
29
|
+
+ orientation; usage stays in `SKILL.md`), and **`docs/ci.md`** (the token-free gate, the packaged
|
|
30
|
+
Action, the live lane — written to stand alone). The harness's own pipeline and contributor suite moved
|
|
31
|
+
to `CONTRIBUTING.md`, which keeps `docs/ci.md` about *consuming* the Action rather than about this repo.
|
|
32
|
+
README keeps what all three audiences share: the fidelity-tier vocabulary, architecture, limitations,
|
|
33
|
+
the docs index, and versioning. Content moved verbatim — this is a relocation, not a rewrite.
|
|
34
|
+
|
|
35
|
+
Nothing is dropped: `Sandboxing` folded into `docs/boundary.md` and `Maintenance` into
|
|
36
|
+
`docs/maintenance.md`; `Discovery` was already covered in `docs/discovery.md`, so it was deleted rather
|
|
37
|
+
than duplicated. Every moved relative link was rewritten for its new depth, and every deep link into a
|
|
38
|
+
moved section was repointed across `docs/`, `examples/README.md`, `llms.txt` and the docs index.
|
|
39
|
+
|
|
40
|
+
### Fixed
|
|
41
|
+
|
|
42
|
+
- **`bump-version` did not know about the router-split pages.** `docs/cli.md`, `docs/ci.md` and
|
|
43
|
+
`docs/companion-skill.md` carry `cowork-harness@^X.Y.Z` install floors that `check:versions` enforces,
|
|
44
|
+
but the bump tool's target list predated them — so `npm run bump` rewrote every other floor and then
|
|
45
|
+
failed its own post-write lockstep check. Caught at release time by that self-check rather than
|
|
46
|
+
shipping a repo with mismatched floors. All three are now registered, and the test pinning that list
|
|
47
|
+
carries the reason.
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
- **22 dead anchor links introduced by the README router split, and the guard gap that let them ship.**
|
|
51
|
+
Moving 11 `##` sections out of README left every `](#slug)` link pointing at them dangling — 19 in
|
|
52
|
+
README (including two badges and the two nav lines the page opened with), plus one each in `AGENTS.md`,
|
|
53
|
+
`docs/ci.md` and `docs/companion-skill.md`. GitHub fails these silently: the page simply does not
|
|
54
|
+
scroll. Every file-path link still resolved and every existing guard stayed green, because the anchor
|
|
55
|
+
suite validated links *into* README and *into* sibling docs but never a page's links to its **own**
|
|
56
|
+
headings. `test/repo-docs-anchors.test.ts` now covers that third case across root pages and `docs/`,
|
|
57
|
+
with a canary against a checker that extracts nothing, and was confirmed to fail against the exact
|
|
58
|
+
regression that shipped. The two nav lines are deleted rather than repointed — they were a
|
|
59
|
+
table-of-contents for a 975-line page, and "Pick your path" plus the Documentation table now do that job.
|
|
60
|
+
|
|
61
|
+
Found by two independent reviewers; the second caught the three non-README instances a README-scoped
|
|
62
|
+
fix would have missed.
|
|
63
|
+
|
|
64
|
+
Follow-up from the same review: three cross-page links still read "…\[Prerequisites](./docs/cli.md#…)
|
|
65
|
+
**below**" — resolving correctly while the sentence lied, because Prerequisites had moved to another
|
|
66
|
+
page. Reworded, and guarded: a cross-page link followed closely by "below"/"above" now fails the suite.
|
|
67
|
+
This half is the nastier one — the link works, so every link checker stays green forever.
|
|
68
|
+
|
|
69
|
+
- **The `protocol` host-hook consent gate was defeatable by a spelling choice.** `allow_host_hooks`
|
|
70
|
+
(3.0.0) refuses a spawn until the operator consents to a staged plugin's hooks running as native host
|
|
71
|
+
processes — but detection keyed only on a file named `hooks.json`, and a plugin may equally declare its
|
|
72
|
+
hooks in the manifest's `hooks` key (`.claude-plugin/plugin.json`, or a bare `plugin.json`), which is
|
|
73
|
+
the more common spelling. Such a plugin sailed past the gate: reproduced end-to-end, the hook executed
|
|
74
|
+
under the operator's own account while the run went green, no consent was asked and no disclosure
|
|
75
|
+
printed. The same blind predicate feeds the tier-independent disclosure, so those hooks also ran
|
|
76
|
+
unannounced at `hostloop`. Detection now returns the UNION of both channels, which fixes the gate and
|
|
77
|
+
the disclosure together since both call through it. The `hooks/hooks.json` placement carve-out is
|
|
78
|
+
deliberately NOT extended to the manifest — that carve-out exists because a misplaced `hooks.json` is
|
|
79
|
+
inert, and a manifest-declared hook is live wherever the manifest sits.
|
|
80
|
+
|
|
81
|
+
Reported by a consumer during a 3.0.0 adoption pass, with the defect proven from committed cassettes:
|
|
82
|
+
10 of theirs carried a `SessionStart:startup` hook_started/hook_response pair from a manifest-only
|
|
83
|
+
declaration, including one recorded at `container`.
|
|
84
|
+
|
|
85
|
+
### Documentation
|
|
86
|
+
|
|
87
|
+
- **The pruned-agent-binary failure now has a documented remedy, with the tradeoffs named.** A Desktop
|
|
88
|
+
update deletes the prior version's staged ELF while often leaving an empty version directory, so a
|
|
89
|
+
scenario pinning that agent version dies with `baseline agent binary not found`. The code anticipates
|
|
90
|
+
this by name, but the remedy appeared in no skill surface at all — and an installed companion skill
|
|
91
|
+
ships without repo `docs/`, so the agent most likely to hit it had nothing to read. Now in
|
|
92
|
+
`docs/gotchas.md` and the skill's own reference, stating that the three remedies are NOT equivalent:
|
|
93
|
+
repin for a reproducibility-bound suite, `latest` for a one-off (a pin rots silently, `latest` drifts
|
|
94
|
+
silently — opposite failure modes), and `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` only to unblock once,
|
|
95
|
+
never in CI, since it runs the newest sibling and downgrades the sha check to advisory.
|
|
96
|
+
|
|
97
|
+
- **`permissive_auto_allow` and script-running skills is a `protocol`-specific hazard, not a general one.**
|
|
98
|
+
`Bash` is off-registry (only `Read`/`Glob`/`Grep` are default-allow), but the harness only decides how to
|
|
99
|
+
ANSWER a permission ask — whether the agent ASKS varies by command AND tier. Measured: at `protocol`, `echo
|
|
100
|
+
hello` produces no ask (green) while `python3 -c "print(42)"` produces one and fails the guard; the
|
|
101
|
+
identical command at `container` produces no ask (green), and `hostloop` routes shell through
|
|
102
|
+
`mcp__workspace__bash` entirely. So a skill running `python3 ${CLAUDE_SKILL_DIR}/scripts/…` goes red at L0
|
|
103
|
+
and green on the sandboxed tiers, while a hello-world probe is green everywhere and teaches the wrong
|
|
104
|
+
expectation. `references/fidelity-and-answers.md` carries the measured table and the three ways through; why
|
|
105
|
+
the host CLI asks for one command and not another is explicitly left untraced, and `microvm` is named as
|
|
106
|
+
unmeasured.
|
|
107
|
+
|
|
9
108
|
## [3.0.0] — 2026-08-29
|
|
10
109
|
|
|
11
110
|
### Breaking
|
package/CONTRIBUTING.md
CHANGED
|
@@ -108,3 +108,38 @@ When a Desktop release moves something `sync` doesn't read, it reports an `unkno
|
|
|
108
108
|
## Reporting issues
|
|
109
109
|
|
|
110
110
|
Use the issue templates. For anything security/sandbox related, see [SECURITY.md](./SECURITY.md).
|
|
111
|
+
|
|
112
|
+
## This repo's own CI pipeline
|
|
113
|
+
|
|
114
|
+
Contributor-facing. For *consuming* the harness in your own CI — the token-free gate and the packaged
|
|
115
|
+
Action — see [docs/ci.md](./docs/ci.md).
|
|
116
|
+
|
|
117
|
+
The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
118
|
+
|
|
119
|
+
| Stage | Runs | Needs | Gates |
|
|
120
|
+
|---|---|---|---|
|
|
121
|
+
| **build** | format check · version-lockstep guard · typecheck · source guards · build · CLI smoke · token-free `replay` · `verify-cassettes` · `lint` | nothing | every push/PR |
|
|
122
|
+
| **test** | the unit suite (vitest), sharded 4-way | nothing | every push/PR |
|
|
123
|
+
| **floor** | the unit suite once, unsharded, on Node 22 — the version `engines.node` declares — so the floor is exercised rather than asserted (other jobs run Node 24, the Active LTS line) | nothing | every push/PR; gates the merge context |
|
|
124
|
+
| **action-self-test** | packs this commit and runs the packaged `uses: ./` Action across its full case set — pass (committed example cassette), usage-error fail (nonexistent path), assertion fail (checks the reporter renders a ❌ row), `lint`, and `analyze-skill`; `ci.yml` currently carries 11 `command:` invocations, so re-count here rather than trusting this sentence | nothing | every push/PR |
|
|
125
|
+
| **python** | `pytest` helper self-checks (`python/`, run with `-m 'not cowork'` — the token-free subset; the Docker/token `@pytest.mark.cowork` tests are excluded) | nothing (token-free assertions only) | every push/PR |
|
|
126
|
+
| **boundary** | builds the pinned agent image, brings up the default-deny network, runs `boundary-check`, then `npm run test:live` (live contract tests that guard the binary-resolution assumptions, no token needed) | Docker, arm64 runner | proves the sandbox enforces Cowork's limits — **no API key** |
|
|
127
|
+
| **image-recipe** | compiles `docker/Dockerfile.agent` in **both** variants (lean/core and `COWORK_FULL_PARITY=1`) on an arm64 runner, so a recipe change cannot reach the merge gate uncompiled | Docker, arm64 runner | every push/PR; gates the merge context |
|
|
128
|
+
| **scenarios** | the live scenario suite (mixed `protocol` + `container` fidelity across `examples/scenarios/`), plus the `e2e/scenarios/*.yaml` smoke set (this repo's own L0/L1/hostloop self-tests, `microvm` excluded — needs a real VM); uploads transcripts/egress logs as artifacts; relies on a runner-local staged agent binary (no in-workflow download step). | `ANTHROPIC_API_KEY` | fork PRs: the whole job is skipped (`if:` guard); same-repo without a key: warns and exits 0 |
|
|
129
|
+
| **parity-drift** | reminder to re-`sync` when Desktop updates | nothing | **goes red** if the newest committed baseline is > 90 days old, but sits outside `ci-green`'s `needs:` list, so a red run does not block a merge |
|
|
130
|
+
|
|
131
|
+
This ordering means cheap checks fail fast, the **boundary parity gate runs without secrets** (so forks get it too), and expensive live runs only happen when a key is present.
|
|
132
|
+
|
|
133
|
+
|
|
134
|
+
## The harness's own suite
|
|
135
|
+
|
|
136
|
+
```bash
|
|
137
|
+
npm run ci # typecheck + build + test (run format:check separately; NOT the same set as CI's `build` job — see CONTRIBUTING.md)
|
|
138
|
+
npm test # vitest: decider, egress allowlist, launch plan, example validation
|
|
139
|
+
cowork-harness boundary-check # self-verify the sandbox (needs Docker; not part of `npm run ci`)
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
Unit tests cover the scripted-answer logic, the egress allowlist matcher, the session→launch-plan materialization (mounts + discovery settings + env-strip), and a **schema guard** that fails if any shipped baseline/session/scenario stops validating. Add a test alongside any new schema field or `Decider` rule — see [CONTRIBUTING.md](./CONTRIBUTING.md).
|
|
143
|
+
|
|
144
|
+
> Copy your starting scenarios/sessions from **`examples/`**. The **`e2e/`** directory is the harness's *own* fidelity self-tests (smoke scenarios per tier) — not a template to copy.
|
|
145
|
+
|