cowork-harness 3.5.0 → 3.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +106 -0
- package/DESIGN.md +3 -3
- package/README.md +4 -4
- package/baselines/desktop-2.2553.1.json +1018 -0
- package/dist/baseline.js +5 -0
- package/dist/decide/llm-transport.js +31 -10
- package/dist/runtime/argv.js +27 -1
- package/dist/scan.js +12 -0
- package/dist/sync/cowork-sync.js +199 -10
- package/docs/ci.md +2 -2
- package/docs/cli.md +3 -3
- package/docs/companion-skill.md +3 -3
- package/docs/fidelity-gaps.md +9 -0
- package/docs/invariants.md +1 -1
- package/docs/maintenance.md +1 -1
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +65 -43
- package/examples/replays/example-pdf-skill.cassette.json +95 -94
- package/examples/replays/hostloop-computer-links.cassette.json +75 -63
- package/package.json +2 -1
- package/scripts/check-claims.ts +103 -0
- package/scripts/check-versions.ts +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.6.0
|
|
7
|
+
tracks-harness: cowork-harness 3.6.0 (baseline desktop-2.2553.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
29
|
-
> `desktop-
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.6.0` (baseline
|
|
29
|
+
> `desktop-2.2553.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.6.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.6.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.6.0"`. **Pin `@^3.6.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.6.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.275 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
|
|
41
41
|
# served only from .../claude-code-releases/rc/<commit>/ — the stable path 404s for those, and
|
|
42
42
|
# 2.1.255 is one. Take B from your pinned baseline's agentBinary.releaseBaseUrl; baselines
|
|
@@ -73,7 +73,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
73
73
|
GitHub-hosted runners, no token/Docker/agent:
|
|
74
74
|
|
|
75
75
|
```yaml
|
|
76
|
-
- run: npm i -g "cowork-harness@^3.
|
|
76
|
+
- run: npm i -g "cowork-harness@^3.6.0"
|
|
77
77
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
78
78
|
# no silent false-greens. WITHOUT --strict this
|
|
79
79
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -350,7 +350,7 @@ jobs:
|
|
|
350
350
|
with: { node-version: '24' }
|
|
351
351
|
- uses: actions/setup-python@v5
|
|
352
352
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
353
|
-
- run: npm i -g "cowork-harness@^3.
|
|
353
|
+
- run: npm i -g "cowork-harness@^3.6.0"
|
|
354
354
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
355
355
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
356
356
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -379,7 +379,7 @@ jobs:
|
|
|
379
379
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
380
380
|
fi
|
|
381
381
|
- if: steps.guard.outputs.live == 'true'
|
|
382
|
-
run: npm i -g "cowork-harness@^3.
|
|
382
|
+
run: npm i -g "cowork-harness@^3.6.0"
|
|
383
383
|
- if: steps.guard.outputs.live == 'true'
|
|
384
384
|
run: cowork-harness run scenarios/ --output-format json
|
|
385
385
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
4
|
-
(baseline `desktop-
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.6.0`
|
|
4
|
+
(baseline `desktop-2.2553.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,112 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.6.0] — 2026-09-18
|
|
10
|
+
|
|
11
|
+
### Upgrade notes
|
|
12
|
+
|
|
13
|
+
- **If you run against Claude Desktop 2.2553.1 (agent 2.1.275), upgrade — `critique` and `--decider-llm`
|
|
14
|
+
are broken on 3.5.0 there.** That agent makes an auxiliary Haiku call in `-p` mode, and 3.5.0's LLM
|
|
15
|
+
transport hard-fails on the two-model envelope it produces (`critique` exits 2 with no error text).
|
|
16
|
+
3.6.0 identifies the primary model instead of counting keys. Nothing changes on older agents.
|
|
17
|
+
- **`latest` now resolves to `desktop-2.2553.1`.** A cassette you recorded against `1.46388.4` with
|
|
18
|
+
`baseline: latest` reports `baseline` staleness on replay (warn by default; `--strict` fails). Re-record
|
|
19
|
+
it, or pin the scenario to `desktop-1.46388.4` if you are not ready to move.
|
|
20
|
+
- **The spawned agent's env gains `CLAUDE_CODE_DESKTOP_APP_VERSION`** on baselines from 2.2553.1 on. It
|
|
21
|
+
is the value the agent uses for the `anthropic-client-version` request header; older baselines are
|
|
22
|
+
unaffected. If you snapshot the spawn env, expect the new key.
|
|
23
|
+
- **Every `-p` call and every run on agent 2.1.275 now carries a ~$0.001 Haiku entry** in `modelUsage`
|
|
24
|
+
and in the run result's cost. Cost comparisons across the 2.1.260→2.1.275 bump will show it — it is
|
|
25
|
+
the agent's spend, not the harness's.
|
|
26
|
+
|
|
27
|
+
### Added
|
|
28
|
+
- **Parity: baseline `desktop-2.2553.1` (agent 2.1.275)** — the first `2.x` Claude Desktop. `sync` refused
|
|
29
|
+
to write with **10 unknown deltas**; all are resolved and the baseline is clean. Two of the ten turned
|
|
30
|
+
out to be defects in this repo's own extractor rather than changes in Desktop:
|
|
31
|
+
- **The S6c Artifact-gate flag was a false alarm.** The frame-artifacts predicate is byte-identical; it
|
|
32
|
+
merely stopped being its own statement (it now shares a declaration with the Artifact host-grant
|
|
33
|
+
binding). The sentinel's value capture ran past the top-level comma and swallowed the sibling binding,
|
|
34
|
+
so an anchored whole-expression match rejected an unchanged predicate — and the message it printed
|
|
35
|
+
("cached-arm/HIPAA/trailing-term change") was simply wrong. The value is now sliced brace/paren/quote
|
|
36
|
+
aware to the first top-level `,` or `;`. Both directions are pinned by tests: the sibling-binding shape
|
|
37
|
+
stays clean, and a real widening hidden before the comma still fires.
|
|
38
|
+
- **The path-gate tool set and path keys were refactored, not removed.** Desktop replaced two array
|
|
39
|
+
literals with one tool→path-key map plus `Object.keys` / `[...new Set(Object.values(…))]` derivations,
|
|
40
|
+
which accounted for 4 of the 10 deltas at once. Same five tools, same two keys — no contract change.
|
|
41
|
+
The extractor now accepts either form. The map is matched **exactly and in order**, because
|
|
42
|
+
`Object.keys` order reaches the sub-agent prompt through `.join(", ")` while the manifest fingerprint
|
|
43
|
+
hashes generator source — a reordered map would otherwise change what the model reads with every check
|
|
44
|
+
still green. Keys and values must resolve to the *same* map, and an ambiguous binding flags rather than
|
|
45
|
+
taking the first match: the bundle carries two more same-shaped maps that add Bash/NotebookEdit/MultiEdit.
|
|
46
|
+
|
|
47
|
+
- **`npm run check:claims` — a staleness report for this repo's "binary-verified" claims.** It lists every
|
|
48
|
+
version-stamped claim in `src/`, `scripts/` and `docs/` that is behind the currently pinned agent and
|
|
49
|
+
`app.asar` versions. First run: **42 of 49 claims behind the pin**, the oldest about 34 baselines back.
|
|
50
|
+
It **exits 0 by design and is not a gate** — a stale stamp is not a wrong claim, and hard-failing would
|
|
51
|
+
force a version bump every sync that anyone could satisfy by editing the digit without re-reading the
|
|
52
|
+
binary. Maintainers get it as step 3b of the parity-sync ritual; contributors need not run it.
|
|
53
|
+
- **`test/subagent-model-precedence-elf.test.ts`** — pins the sub-agent model resolution order against the
|
|
54
|
+
agent binary itself, so the precedence corrected in 3.5.0 cannot silently drift back. It reads the
|
|
55
|
+
binary's own telemetry labels rather than a minified symbol (the gate accessor beside that code renamed
|
|
56
|
+
between two Desktop patch builds). Its second assertion is **not** gated on a staged binary: it fails if
|
|
57
|
+
`docs/session.md`, `docs/subagents.md` or `src/session.ts` reintroduces the reversed order, which is the
|
|
58
|
+
half CI can enforce and the way that claim went wrong in three places at once.
|
|
59
|
+
|
|
60
|
+
**Why these two:** the repo carries ~49 version-stamped claims about the agent binary and,
|
|
61
|
+
before 3.5.0, exactly one was re-derived from the binary by a test. Both claims spot-checked during that
|
|
62
|
+
release turned out wrong — the hook-event list and the model precedence. Two for two is not a sample that
|
|
63
|
+
justifies leaving the rest unexamined, but it also does not justify pretending a report verifies them: it
|
|
64
|
+
shows the population and its age, and a human decides what to re-read.
|
|
65
|
+
|
|
66
|
+
### Fixed
|
|
67
|
+
|
|
68
|
+
- **The LLM decider transport no longer hard-fails on agent 2.1.275's two-model envelope.** The agent that
|
|
69
|
+
ships with Desktop 2.2553.1 makes an auxiliary Haiku call in `-p` mode, so `claude -p --output-format json`
|
|
70
|
+
now reports two `modelUsage` keys where 2.1.260 reported one — measured with the same prompt and flags
|
|
71
|
+
against both native binaries. The transport asserted exactly one key, which turned **every** critique
|
|
72
|
+
evaluator pass and every `--decider-llm` gate into an instrument failure on the new agent (`critique`
|
|
73
|
+
exited 2 with no error text; found by the live lane, which had this test gated off in the previous pass).
|
|
74
|
+
The primary model is now *identified* as the key that resolves the requested `--model` (exact id, or the
|
|
75
|
+
id carrying a floating alias like `sonnet` as a dash-separated segment) rather than *assumed* from the
|
|
76
|
+
count. Zero or several keys resolving the request still fails closed — that ambiguity is the contract
|
|
77
|
+
break the check exists to catch. The whole usage map is still passed through, so the auxiliary call's
|
|
78
|
+
cost is not lost.
|
|
79
|
+
- **The spawned agent now sends the client-identity headers production sends.** Desktop 2.2553.1 sets
|
|
80
|
+
`CLAUDE_CODE_DESKTOP_APP_VERSION` unconditionally on first-party sessions, and the agent reads it on the
|
|
81
|
+
`local-agent` entrypoint — which the harness pins — as the fallback source of the `anthropic-client-version`
|
|
82
|
+
header (its companion `anthropic-client-platform` is the hard-coded literal `desktop_app`) whenever
|
|
83
|
+
`ANTHROPIC_CUSTOM_HEADERS` carries none, which is the harness's case. The key is host-derived (an Electron `app.getVersion()` call), so it is allowlisted in the sync
|
|
84
|
+
**and** injected from the baseline's `appVersion` in both the container and native spawn envs — **version-gated** to baselines from 2.2553.1 on, since injecting it on an older baseline would hand the agent a key that baseline's Desktop never set (not symmetric with `CLAUDE_CODE_HOST_PLATFORM`, which every asar on record sets). Allowlisting
|
|
85
|
+
it alone would have been silent: the allowlist is consulted before the pin list and the key is not
|
|
86
|
+
required, so the harness would simply have stopped sending those headers with nothing failing.
|
|
87
|
+
- **`CLAUDE_ARTIFACT_HOST_GRANT` is guarded, not merely allowlisted.** The new key is allowlisted on the
|
|
88
|
+
grounds that a default session never receives it — a claim that rests entirely on its guard. A new
|
|
89
|
+
sentinel requires it to stay gated on the *same* predicate as the Artifact tool spread, and fires if it
|
|
90
|
+
is ever constructed unconditionally or re-keyed. (The existing frame-artifacts assertion could not reach
|
|
91
|
+
it: the two spreads have different shapes.)
|
|
92
|
+
- **`CLAUDE_CODE_DISABLE_CRON` gained a second disjunct** (a managed-settings scheduled-tasks switch). The
|
|
93
|
+
pinned value is unchanged at `"1"`, and it is *earned* rather than assumed — the spawn window still passes
|
|
94
|
+
`disableCron:!0`, which short-circuits. The resolver and its anchor admit the new shape; the disjunct is
|
|
95
|
+
not inert in general, only under that short-circuit.
|
|
96
|
+
|
|
97
|
+
### Changed
|
|
98
|
+
|
|
99
|
+
|
|
100
|
+
- `CLAUDE_CODE_MODEL_CATALOG` (new, third-party-only branch) is allowlisted, matching the standing rule for
|
|
101
|
+
third-party-only keys.
|
|
102
|
+
- **`design` added to the host-inventory scan's known-built-in skill roster.** It surfaced as a finding on
|
|
103
|
+
the first fresh `container` recording after this sync, on a cassette whose scenario declares no skills.
|
|
104
|
+
It qualifies under the roster's existing three criteria: the recording was sealed (`container`, so
|
|
105
|
+
`HOME=/tmp` and no host `~/.claude`), `"design"` is a bare literal in both the staged agent ELF and the
|
|
106
|
+
host CLI, and five personal skill names from the same machine are absent from that binary. It is not new
|
|
107
|
+
to this agent — the `design-consent` / `design-revoke` slash commands were already in the previously
|
|
108
|
+
shipped cassette; what changed is that the feature now also registers in `skills[]`, an axis the scan
|
|
109
|
+
treats more strictly.
|
|
110
|
+
- **All three committed cassettes in `examples/replays/` are re-recorded against `desktop-2.2553.1`**, each reporting no behavioural change versus the recording it replaced. The `protocol` fixture was recorded on the hermetic managed config dir (`ANTHROPIC_API_KEY` path) and the `container` one in a sealed container; `verify-cassettes` reports zero host-inventory findings on all three.
|
|
111
|
+
- **A full live pass was run against `desktop-2.2553.1` / agent 2.1.275**, all four suites and all four tiers: `boundary-check` 6/6; e2e self-tests 9/9 including `smoke-l2-microvm` in a real VM and `smoke-multiselect-deciderdir` through the `--decider-llm` path; `npm run test:live` 19 tests, 18 passed, 1 failed, **0 skipped** (the previous pass had one skip — the hostloop `critique` case — which this pass exercised for the first time and which found the transport defect fixed above); `run examples/scenarios/` 7/7. The one live red is a pre-existing `live-matrix` case on old baselines where the model sometimes answers as text instead of calling `AskUserQuestion` — model variance, re-run and flipped, logged for hardening.
|
|
112
|
+
- **`test/model-provenance.test.ts`'s pre-coverage-note test now builds its own fixture.** Every committed cassette now carries `model` coverage, so no shipped fixture emits the note the test reads. Rather than asserting the note's shape only when one happens to be present — a test that could not fail — it rewrites a real cassette's session fingerprint to the pre-`model` hash in a temp tree that preserves the relative session layout.
|
|
113
|
+
|
|
114
|
+
|
|
9
115
|
## [3.5.0] — 2026-09-06
|
|
10
116
|
|
|
11
117
|
**Live verification for this release** (macOS arm64, agent **2.1.260**, agent image `cowork-agent-base:2`,
|
package/DESIGN.md
CHANGED
|
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
47
47
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
48
48
|
|
|
49
49
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
50
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
50
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.275**, per `baselines/desktop-2.2553.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
|
|
51
51
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
52
52
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
53
53
|
|
|
@@ -174,9 +174,9 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
174
174
|
|
|
175
175
|
> **Machine-readable form:** the five shapes below are schema'd as `schema/protocol.v1.json`, with a golden vector pack at `fixtures/protocol/v1/` — see [docs/protocol.md](./docs/protocol.md) for scope, versioning, and how to conformance-test against them.
|
|
176
176
|
|
|
177
|
-
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS.
|
|
177
|
+
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. **This heading deliberately carries no version or baseline figures**: it used to restate the agent version and baseline of the last live pass, and those went stale independently of the note below it — at one point naming three different agent versions across two adjacent sentences. The note directly below is the single authority for which baseline and agent were actually exercised, and what the pass did and did not cover. Read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
178
178
|
|
|
179
|
-
> **Scope of that claim
|
|
179
|
+
> **Scope of that claim.** `2026-09-18 / desktop-2.2553.1` is the baseline carrying the latest live pass, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.275** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 tests — 18 passed, 1 failed, 0 skipped** (the fail is `live-matrix`'s 2-cell baseline-axis case on the OLD baselines 1.17377.2/1.18286.0, where `claude-opus-4-8` sometimes answers the prompt's question as text instead of calling `AskUserQuestion`; it passed 2 of 4 attempts on the 1.18286.0 cell and never wavered on the other — model variance on a test that predates this sync, logged for hardening); and `run examples/scenarios/` **7/7 success** on its first run (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe), then **6/7** on an accidental second run where `subagent-manifest-probe`'s sub-agent chose an absolute host path outside the connected folders and the hostloop path gate correctly denied it — the gate working as production does, on a model choice. and the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files) — `smoke-multiselect-deciderdir` runs the `--decider-llm` path, so it is live confirmation that the transport fix holds there as well as in `critique`. **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — and the previous pass HAD one skip, the hostloop `critique` case, which this pass exercised for the first time and which found a real defect (the LLM decider transport hard-failed on agent 2.1.275's two-model `-p` envelope; fixed in this release). Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`** (`smoke-l2-microvm` passed in a real VM with its own kernel, 151.9s). **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise — both reds above were re-run and both flipped. (c) `run examples/scenarios/` was accidentally run twice; the first run is the one counted, and the second's single red is described above for what it is. Separately and not a live matter: all three committed cassettes in `examples/replays/` were re-recorded against `desktop-2.2553.1` on 2026-09-18 (each reporting "no behavioural change — transcript wording only" versus the recording it replaced); the `protocol` fixture on the hermetic managed config dir and the `container` one in a sealed container; `verify-cassettes` reports `privacyScanned:true` and zero host-inventory findings on all three. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
180
180
|
|
|
181
181
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
182
182
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^3.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^3.6.0"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.6.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.6.0"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Not sure a harness is what you need? The next two sections are the argument.
|
|
@@ -380,6 +380,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
380
380
|
## Status
|
|
381
381
|
|
|
382
382
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
383
|
-
**`desktop-
|
|
383
|
+
**`desktop-2.2553.1`**. Release-by-release verification notes (what was re-verified against
|
|
384
384
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
385
385
|
this section would otherwise duplicate lives in the sections above.
|