cowork-harness 3.4.0 → 3.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +16 -4
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +33 -0
- package/DESIGN.md +1 -1
- package/README.md +30 -3
- package/dist/sync/cowork-sync.js +12 -6
- package/docs/ci.md +1 -1
- package/docs/cli.md +3 -3
- package/docs/companion-skill.md +3 -3
- package/docs/fidelity-gaps.md +31 -0
- package/docs/maintenance.md +7 -0
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.4.
|
|
7
|
-
tracks-harness: cowork-harness 3.4.
|
|
6
|
+
version: 3.4.1
|
|
7
|
+
tracks-harness: cowork-harness 3.4.1 (baseline desktop-1.46388.3)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.4.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.4.1` (baseline
|
|
29
29
|
> `desktop-1.46388.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.4.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.4.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.4.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.4.1"`. **Pin `@^3.4.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -158,6 +158,18 @@ Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejec
|
|
|
158
158
|
(it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
|
|
159
159
|
`references/fidelity-and-answers.md`.
|
|
160
160
|
|
|
161
|
+
**Every tier models Cowork's DESKTOP-LOCAL lane** — agent on the user's machine, shell rooted at
|
|
162
|
+
`/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
|
|
163
|
+
**remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
|
|
164
|
+
different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
|
|
165
|
+
reproduces it and none can — that container is not something a local tool can stand up. Which lane a
|
|
166
|
+
real session gets is a Cowork setting ("Only on this computer"), observed **off** on a current install.
|
|
167
|
+
So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
|
|
168
|
+
asserting a **path, mount or delivery mechanism** is a claim about the local lane only. Declare
|
|
169
|
+
`lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
|
|
170
|
+
rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
|
|
171
|
+
above).
|
|
172
|
+
|
|
161
173
|
### Choose an answer path (gates: AskUserQuestion + tool-permission)
|
|
162
174
|
|
|
163
175
|
Default to **deterministic**: scripted `answers:` + `on_unanswered: fail`. Anything that brings a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.4.1` (baseline `desktop-1.46388.3`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.4.
|
|
20
|
+
(e.g. `version: "3.4.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -73,7 +73,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
73
73
|
GitHub-hosted runners, no token/Docker/agent:
|
|
74
74
|
|
|
75
75
|
```yaml
|
|
76
|
-
- run: npm i -g "cowork-harness@^3.4.
|
|
76
|
+
- run: npm i -g "cowork-harness@^3.4.1"
|
|
77
77
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
78
78
|
# no silent false-greens. WITHOUT --strict this
|
|
79
79
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -350,7 +350,7 @@ jobs:
|
|
|
350
350
|
with: { node-version: '24' }
|
|
351
351
|
- uses: actions/setup-python@v5
|
|
352
352
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
353
|
-
- run: npm i -g "cowork-harness@^3.4.
|
|
353
|
+
- run: npm i -g "cowork-harness@^3.4.1"
|
|
354
354
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
355
355
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
356
356
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -379,7 +379,7 @@ jobs:
|
|
|
379
379
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
380
380
|
fi
|
|
381
381
|
- if: steps.guard.outputs.live == 'true'
|
|
382
|
-
run: npm i -g "cowork-harness@^3.4.
|
|
382
|
+
run: npm i -g "cowork-harness@^3.4.1"
|
|
383
383
|
- if: steps.guard.outputs.live == 'true'
|
|
384
384
|
run: cowork-harness run scenarios/ --output-format json
|
|
385
385
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.4.
|
|
3
|
+
Tracks `cowork-harness 3.4.1` (baseline `desktop-1.46388.3`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.4.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.4.1`
|
|
4
4
|
(baseline `desktop-1.46388.3`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.4.
|
|
5
|
+
Tracks `cowork-harness 3.4.1` (baseline `desktop-1.46388.3`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,39 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.4.1] — 2026-09-05
|
|
10
|
+
|
|
11
|
+
### Documentation
|
|
12
|
+
|
|
13
|
+
- **Which Cowork *lane* the harness models is now stated, in the places a consumer reads.** Every
|
|
14
|
+
fidelity tier reproduces the desktop-local lane; Cowork's remote lane runs server-side in a cloud
|
|
15
|
+
container with a different filesystem, shell tool, delivery mechanism and a server-authored prompt.
|
|
16
|
+
Which lane a real session gets is a Cowork setting ("Only on this computer"), and it was observed
|
|
17
|
+
**off** on a current install. README, the companion skill, `docs/fidelity-gaps.md` and
|
|
18
|
+
`docs/maintenance.md` now say so. The distinction that matters: behaviour conclusions (triggering,
|
|
19
|
+
tool sequencing, gate handling) travel between lanes; anything asserting a **path, mount or delivery
|
|
20
|
+
mechanism** is a local-lane claim only. No remote tier is planned — that container is Anthropic's, so
|
|
21
|
+
emulating it would mean authoring an environment rather than reproducing one; `lane: remote` already
|
|
22
|
+
makes the affected assertions refuse to grade instead of passing.
|
|
23
|
+
|
|
24
|
+
### Changed
|
|
25
|
+
|
|
26
|
+
- **The sub-agent override sentinel's note now carries a current probe.** Gate `124685897` reads ON,
|
|
27
|
+
which only enables a server-delivered replacement of the `## Cowork environment` section. Re-probed
|
|
28
|
+
against the 1.46388.3 composition in a real host-loop session: all three composed parts arrived
|
|
29
|
+
byte-identical to the shipped fallback, so the gate is on with no payload. The two non-overridable
|
|
30
|
+
parts are the control that makes it conclusive. The note also states the probe's precondition, which
|
|
31
|
+
it previously lacked.
|
|
32
|
+
|
|
33
|
+
### Verification
|
|
34
|
+
|
|
35
|
+
- **`microvm` exercised after the 3.4.0 tag** — `smoke-l2-microvm` passed 3/3 against the *published*
|
|
36
|
+
3.4.0 artifact (real Apple-VZ VM, separate kernel; guest egress reached allowlisted
|
|
37
|
+
`api.anthropic.com` and denied an off-list telemetry host). That completes all four tiers for the
|
|
38
|
+
1.46388.3 baseline; DESIGN.md's scope note is re-stamped accordingly. The 3.4.0 entry below is left
|
|
39
|
+
as the record of what was verified at tag time. This tier is macOS-arm64 only and CI runners are
|
|
40
|
+
Linux, so it is only ever exercised by hand.
|
|
41
|
+
|
|
9
42
|
## [3.4.0] — 2026-09-05
|
|
10
43
|
|
|
11
44
|
### Upgrade notes
|
package/DESIGN.md
CHANGED
|
@@ -176,7 +176,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
176
176
|
|
|
177
177
|
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
178
178
|
|
|
179
|
-
> **Scope of that claim, stated plainly.** `2026-09-05 / desktop-1.46388.3` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.260** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, and covered **four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 passed, 0 skipped**; `run examples/scenarios/` **7/7 success, 39 assertions, 0 failed** (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe); and the e2e self-tests **
|
|
179
|
+
> **Scope of that claim, stated plainly.** `2026-09-05 / desktop-1.46388.3` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.260** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, and covered **four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 passed, 0 skipped**; `run examples/scenarios/` **7/7 success, 39 assertions, 0 failed** (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe); and the e2e self-tests **9/9 success** (smoke-askuserquestion, smoke-multiselect, smoke-multiselect-deciderdir, smoke-l1-container, smoke-l1-egress, smoke-present-files, smoke-semantic-evidence-files, canary-hostloop, smoke-l2-microvm). **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported none. Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`.** The microvm case ran AFTER the 3.4.0 tag, against the published artifact rather than the checkout, and is recorded here rather than in that release's CHANGELOG entry, which states what was verified at tag time. It is the tier CI can never cover — GitHub's runners are Linux and Apple-VZ is macOS-arm64 only — so it is only ever exercised by hand here: `smoke-l2-microvm` passed 3/3 in a real VM with its own kernel, and its guest egress behaved (allowlisted `api.anthropic.com` reached, an off-list telemetry host denied). **Scope-out, so this is not read as more than it is.** (a) `subagent-manifest-probe` is new in this release and is the first live coverage the sub-agent append has ever had; it proves the composed append is ACTIONABLE — a sub-agent reaches the right files through both tool families, and a bare relative write lands at the host cwd — **not** that the harness's paraphrase matches Desktop's wording, which is the `manifest`/`suffix` fingerprint axes' job. (b) A live pass verifies observed behaviour, not the whole spawn contract by construction. (c) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: `example-pdf-skill` was re-recorded against this baseline and the other two committed cassettes were re-stamped (see the CHANGELOG for why each), so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
180
180
|
|
|
181
181
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
182
182
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^3.4.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^3.4.1"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.4.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.4.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.4.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.4.1"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Not sure a harness is what you need? The next two sections are the argument.
|
|
@@ -215,6 +215,32 @@ L2 microvm parity Optional. Agent inside a real Linux microVM (Lima/Apple-VZ
|
|
|
215
215
|
synced baseline. "Do what real Cowork does."
|
|
216
216
|
```
|
|
217
217
|
|
|
218
|
+
### Which Cowork *lane* this models — read this before trusting an environment assertion
|
|
219
|
+
|
|
220
|
+
Every tier above reproduces Cowork's **desktop-local lane**: the agent runs on your machine, shell
|
|
221
|
+
commands land in a Linux sandbox rooted at `/sessions/<id>`, attached folders appear under
|
|
222
|
+
`/sessions/<id>/mnt/<name>`, and finished files reach the user through `present_files`.
|
|
223
|
+
|
|
224
|
+
Cowork also has a **remote lane**, where the session runs server-side in an ephemeral cloud container
|
|
225
|
+
that reaches your machine over a link. There the filesystem, the shell tool, and file delivery are all
|
|
226
|
+
different — folders arrive under `$HOME/mnt/`, deliverables go to `/mnt/user-data/outputs/` and are
|
|
227
|
+
handed over with `SendUserFile`, and the environment prompt is authored by the server rather than by
|
|
228
|
+
Desktop. **Which lane you get is a Cowork setting** ("Only on this computer", Settings → Cowork), and
|
|
229
|
+
it has been observed **off** — i.e. remote — on a current install.
|
|
230
|
+
|
|
231
|
+
The harness **cannot execute the remote lane**: that container is Anthropic's, not something a local
|
|
232
|
+
tool can stand up. What it does instead is refuse to fake it. Declare `lane: remote` on a scenario and
|
|
233
|
+
the assertions that depend on observing a local filesystem degrade honestly — `file_absent` reports
|
|
234
|
+
evidence-unavailable rather than passing, and delivery is reported as unobservable — so a green never
|
|
235
|
+
means more than it should.
|
|
236
|
+
|
|
237
|
+
**What this means for you.** Behaviour-shaped conclusions travel between lanes: whether your skill
|
|
238
|
+
triggers, how it sequences tools, which questions it asks, whether it respects a permission gate.
|
|
239
|
+
Environment-shaped conclusions do not: anything asserting a path, a mount, or a delivery mechanism is a
|
|
240
|
+
statement about the **local** lane specifically. Scope your claims accordingly, and if you are probing
|
|
241
|
+
real Cowork to compare, turn "Only on this computer" **on** first or you will be measuring a lane this
|
|
242
|
+
tool does not model.
|
|
243
|
+
|
|
218
244
|
**Decision guide** — `fidelity:` takes exactly one of these five values (`protocol`/`container`/`microvm` vary isolation strength; `hostloop`/`cowork` are overlays that instead pick *where the loop runs* — there's no combining the two groups):
|
|
219
245
|
|
|
220
246
|
| Question | Choose |
|
|
@@ -281,6 +307,7 @@ See [DESIGN.md](./DESIGN.md) for the full parity matrix, the known deltas vs. re
|
|
|
281
307
|
|
|
282
308
|
## Limitations
|
|
283
309
|
|
|
310
|
+
- **One lane, deliberately.** Every tier models Cowork's desktop-local lane. The remote (cloud) lane runs server-side with a different filesystem, shell tool, delivery mechanism and a server-authored prompt; no tier reproduces it, and `lane: remote` exists to make the resulting blind spots refuse to grade rather than pass. See [Which Cowork lane this models](#which-cowork-lane-this-models--read-this-before-trusting-an-environment-assertion).
|
|
284
311
|
- **Not the full Desktop network transport.** L1 is a container, not a VM; L2 *is* a real Apple-VZ microVM but still does not reproduce Cowork's gVisor netstack — its egress is the same allowlist proxy as L1 (with a guest iptables firewall in front). If your skill depends on VM-kernel specifics, validate at L2; if it depends on packet-level gVisor behavior, no tier reproduces it.
|
|
285
312
|
- **Cowork in-guest context is partial.** Desktop supplies host-loop staging, runtime `mountPath` RPC, and the bridge. We reproduce the *filesystem and cowork mode*, not those host-side services. Skills that call Desktop-only host RPCs won't run here (they wouldn't be portable anyway).
|
|
286
313
|
- **The agent binary is the staged ELF** (`claude-code-vm/<ver>/claude`), **bind-mounted** from your own Claude Desktop install — nothing Anthropic-owned is bundled or installed. There is **no npm path**; override the path with `COWORK_AGENT_BINARY`. Check licensing/ToS for your use.
|
package/dist/sync/cowork-sync.js
CHANGED
|
@@ -518,13 +518,19 @@ export function checkSubagentOverrideGate(gates) {
|
|
|
518
518
|
"<vmCwd>/mnt/, and shell starting in <vmCwd> with non-mnt writes reaching neither the user nor the " +
|
|
519
519
|
"file tools — so NO override was reaching that account and the committed paraphrase is faithful. " +
|
|
520
520
|
"That is EVIDENCE, NOT PROOF: one account, one session, and a server rule can be segment-targeted. " +
|
|
521
|
-
"
|
|
522
|
-
"
|
|
523
|
-
"
|
|
524
|
-
"
|
|
525
|
-
"
|
|
521
|
+
"RE-PROBED 2026-09-05 against the 1.46388.3 composition, and it holds: a real host-loop sub-agent " +
|
|
522
|
+
"received all three composed parts byte-identical to this build's fallback text — the overridable " +
|
|
523
|
+
"section, the folder manifest, and the trailing skills sentence. The two composed parts are the " +
|
|
524
|
+
"CONTROL that makes this conclusive: they are appended AFTER resolveSection and cannot be " +
|
|
525
|
+
"server-replaced, so their being verbatim proves the sub-agent quoted faithfully rather than " +
|
|
526
|
+
"paraphrasing something that merely looked right. So the gate is ON with no payload, and the " +
|
|
527
|
+
"hardcoded fallback is what reaches the model. " +
|
|
528
|
+
"Still one account, one session, and still segment-targetable. " +
|
|
526
529
|
"If the sub-agent append matters to what you are about to ship, re-probe (dispatch a sub-agent, ask " +
|
|
527
|
-
"for its environment section verbatim, diff the three composed parts) rather than trusting this note."
|
|
530
|
+
"for its environment section verbatim, diff the three composed parts) rather than trusting this note. " +
|
|
531
|
+
"NOTE the probe now has a PRECONDITION: Cowork's `Only on this computer` setting (localAgentMode) is " +
|
|
532
|
+
"OFF by default, and with it off a session runs server-side with a server-authored prompt that has no " +
|
|
533
|
+
"`## Cowork environment` section at all. Turn it ON, or you will probe a lane this harness does not model.",
|
|
528
534
|
];
|
|
529
535
|
}
|
|
530
536
|
/** Read `network.allowDomains` from the NEWEST committed baseline — the pinned, hand-curated egress
|
package/docs/ci.md
CHANGED
|
@@ -97,7 +97,7 @@ jobs:
|
|
|
97
97
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
98
98
|
```
|
|
99
99
|
|
|
100
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.4.
|
|
100
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.4.1` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
101
101
|
|
|
102
102
|
CI uses `ANTHROPIC_API_KEY` specifically because there's no interactive browser available to run
|
|
103
103
|
`claude setup-token`'s OAuth flow in a GitHub Actions runner; locally, the OAuth token is preferred because
|
package/docs/cli.md
CHANGED
|
@@ -18,7 +18,7 @@ companion skill, CI). This page is the CLI one.
|
|
|
18
18
|
**Install from npm:**
|
|
19
19
|
|
|
20
20
|
```bash
|
|
21
|
-
npm install -g "cowork-harness@^3.4.
|
|
21
|
+
npm install -g "cowork-harness@^3.4.1" # puts the `cowork-harness` command on your PATH
|
|
22
22
|
```
|
|
23
23
|
|
|
24
24
|
**Or build from source:**
|
|
@@ -38,7 +38,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
38
38
|
|
|
39
39
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
40
40
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
41
|
-
> From a global install (`npm i -g "cowork-harness@^3.4.
|
|
41
|
+
> From a global install (`npm i -g "cowork-harness@^3.4.1"`), point at the package root instead:
|
|
42
42
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
43
43
|
> (or copy the cassette into your own project and pass that path).
|
|
44
44
|
|
|
@@ -48,7 +48,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
48
48
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
49
49
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
50
50
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
51
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.4.
|
|
51
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.4.1"`.
|
|
52
52
|
|
|
53
53
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
54
54
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
package/docs/companion-skill.md
CHANGED
|
@@ -28,7 +28,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
28
28
|
claude plugin install cowork-harness@cowork-harness
|
|
29
29
|
```
|
|
30
30
|
|
|
31
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.4.
|
|
31
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.4.1"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
32
32
|
|
|
33
33
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
34
34
|
|
|
@@ -38,7 +38,7 @@ npx skills add yaniv-golan/cowork-harness --skill cowork-harness
|
|
|
38
38
|
|
|
39
39
|
(Working *inside* this repo, the skill auto-loads as a project skill — no install needed.)
|
|
40
40
|
|
|
41
|
-
| What ships | npm global (`npm install -g "cowork-harness@^3.4.
|
|
41
|
+
| What ships | npm global (`npm install -g "cowork-harness@^3.4.1"`) | Source checkout (`git clone` + `npm ci`) |
|
|
42
42
|
|---|---|---|
|
|
43
43
|
| CLI, `scenario.py` + assertion keys (enough for `lint` in CI) | ✓ | ✓ |
|
|
44
44
|
| `SKILL.md`, all of `docs/`, `SPEC.md`/`DESIGN.md`/`AGENTS.md` | ✓ | ✓ |
|
|
@@ -53,6 +53,6 @@ global install puts nothing in your working directory. The matrix, answer-policy
|
|
|
53
53
|
ones that still need a source checkout. (The marketplace
|
|
54
54
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
55
55
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
56
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^3.4.
|
|
56
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^3.4.1"` — see
|
|
57
57
|
[above](#install) — which pulls the same npm package as the global-install row.)
|
|
58
58
|
|
package/docs/fidelity-gaps.md
CHANGED
|
@@ -8,6 +8,37 @@ For how the harness *enforces* the limitations it does reproduce (sealed filesys
|
|
|
8
8
|
|
|
9
9
|
---
|
|
10
10
|
|
|
11
|
+
## Which Cowork LANE this harness models — read first, it scopes everything below
|
|
12
|
+
|
|
13
|
+
Every fidelity tier reproduces Cowork's **desktop-local** lane. Cowork also runs sessions on a
|
|
14
|
+
**remote** lane, server-side in a cloud container that reaches the user's machine over a device
|
|
15
|
+
bridge, and **which lane a real session gets is a Cowork setting** — "Only on this computer"
|
|
16
|
+
(Settings → Cowork). Measured 2026-09-05 on a current install: that setting was **off**, so an
|
|
17
|
+
otherwise-default Cowork session ran remote.
|
|
18
|
+
|
|
19
|
+
That matters for how you read the rest of this file. Gaps documented here are gaps against the
|
|
20
|
+
*local* lane. On the remote lane the environment is different in kind, not degree: the cloud
|
|
21
|
+
container's cwd is `/home/claude` with no `mnt/` tree, delivery is `SendUserFile` rather than
|
|
22
|
+
`present_files`, and the session reaches the user's disk through the `device_*` tools into a local
|
|
23
|
+
VM (see "File delivery" and the device-tool section below). Its environment prompt is **authored by
|
|
24
|
+
the server**, not by Desktop — the heading and markers a remote sub-agent reports are 0 occurrences
|
|
25
|
+
in both the app bundle and the agent binary.
|
|
26
|
+
|
|
27
|
+
**The harness cannot execute the remote lane and does not pretend to.** That container is
|
|
28
|
+
Anthropic's; standing up a local imitation would be authoring an environment rather than reproducing
|
|
29
|
+
one, with no production to verify it against. What exists instead is `lane: remote` on a scenario,
|
|
30
|
+
which makes the affected assertions **refuse to grade** rather than pass — `file_absent` reports
|
|
31
|
+
evidence-unavailable, delivery is reported unobservable.
|
|
32
|
+
|
|
33
|
+
**Practical consequence.** Behaviour-shaped conclusions travel between lanes: whether a skill
|
|
34
|
+
triggers, how it sequences tools, which questions it asks, whether it honours a permission gate.
|
|
35
|
+
Environment-shaped conclusions do not: any assertion about a path, a mount, or a delivery mechanism
|
|
36
|
+
is a claim about the local lane only. And if you are probing real Cowork to compare against this
|
|
37
|
+
harness, **turn "Only on this computer" on first** — with it off you are measuring a lane this tool
|
|
38
|
+
does not model, which has already cost one wasted probe.
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
11
42
|
## Mid-session skill/plugin re-sync
|
|
12
43
|
|
|
13
44
|
Cowork re-syncs skills and plugins from the host into the session **while the session is running** —
|
package/docs/maintenance.md
CHANGED
|
@@ -236,6 +236,13 @@ committed baseline and say why in that baseline's `$comment`.
|
|
|
236
236
|
`vm` always mandatory; `manifest` and `suffix` mandatory together once either is recorded — a partial
|
|
237
237
|
entry is itself a hard-fail), then re-run `cowork-harness sync`.
|
|
238
238
|
|
|
239
|
+
> **PRECONDITION for any live probe of real Cowork.** Cowork's "Only on this computer" setting
|
|
240
|
+
> (Settings → Cowork) selects the lane. With it **off** — observed to be the default state on a
|
|
241
|
+
> current install — a session runs server-side under a server-authored prompt with no
|
|
242
|
+
> `## Cowork environment` section at all, and you will be diffing a lane this harness does not
|
|
243
|
+
> model. Turn it on and start a FRESH session before probing. This cost one wasted probe on
|
|
244
|
+
> 2026-09-05.
|
|
245
|
+
|
|
239
246
|
> **Then REPOINT the baseline at the new asset** — `spawn.subagentAppendHostLoop` (and/or
|
|
240
247
|
> `spawn.subagentAppend`) in the freshly written `baselines/desktop-<new>.json`. These pointers are
|
|
241
248
|
> hand-authored, so `sync` carries the PREVIOUS release's value forward untouched; writing the
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.4.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.4.1"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "3.4.
|
|
3
|
+
"version": "3.4.1",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|