cowork-harness 1.0.2 → 1.0.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +6 -6
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/CHANGELOG.md +26 -0
- package/README.md +6 -6
- package/RELEASING.md +3 -1
- package/baselines/desktop-1.20186.9.json +380 -0
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/package.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.0.
|
|
7
|
-
tracks-harness: cowork-harness 1.0.
|
|
6
|
+
version: 1.0.4
|
|
7
|
+
tracks-harness: cowork-harness 1.0.4 (baseline desktop-1.20186.9)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0.
|
|
26
|
-
> `desktop-1.20186.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0.4` (baseline
|
|
26
|
+
> `desktop-1.20186.9`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0.4**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.0.4" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.0.4"`. **Pin `@>=1.0.4`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
|
-
What the ≥ 1.0.
|
|
44
|
+
What the ≥ 1.0.4 floor gates, by release:
|
|
45
45
|
|
|
46
46
|
- **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
|
|
47
47
|
- **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.0.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.0.4` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.0.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.0.4"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.0.
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.0.4"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.0.
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.0.4"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.0.
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.0.4"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0.4`
|
|
4
4
|
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,32 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.0.4] — 2026-07-15
|
|
10
|
+
|
|
11
|
+
Patch: release workflows no longer trigger on the floating Marketplace alias tags. No runtime/API change.
|
|
12
|
+
|
|
13
|
+
### Fixed
|
|
14
|
+
|
|
15
|
+
- `.github/workflows/release.yml` and `.github/workflows/publish-image.yml` now trigger only on
|
|
16
|
+
full semver tags (`v[0-9]+.[0-9]+.[0-9]+*`, prereleases included) instead of `v*`. Moving the
|
|
17
|
+
packaged Action's floating alias tags (`v1`, `v1.0`) after a release previously kicked off both
|
|
18
|
+
workflows, which then correctly died at the tag-vs-`package.json` version guard — four spurious
|
|
19
|
+
failed runs per release. The aliases point at an already-published release commit, so nothing
|
|
20
|
+
should (or now does) run.
|
|
21
|
+
|
|
22
|
+
## [1.0.3] — 2026-07-14
|
|
23
|
+
|
|
24
|
+
Patch: parity sync to Claude Desktop `1.20186.9`. No runtime/API change.
|
|
25
|
+
|
|
26
|
+
### Changed
|
|
27
|
+
|
|
28
|
+
- Synced the platform baseline to Claude Desktop `1.20186.9`
|
|
29
|
+
(`baselines/desktop-1.20186.9.json`, now what `baseline: latest` resolves to). A routine
|
|
30
|
+
per-release parity refresh: the app version, the native agent staging path, and the asar
|
|
31
|
+
fingerprint moved; the Cowork system prompt, egress allowlist, gate states, and agent (VM)
|
|
32
|
+
version are unchanged from `1.20186.1`. README and the companion skill's baseline pointer were
|
|
33
|
+
updated to match.
|
|
34
|
+
|
|
9
35
|
## [1.0.2] — 2026-07-14
|
|
10
36
|
|
|
11
37
|
Patch: shorten the Action's Marketplace tagline. No runtime/API change.
|
package/README.md
CHANGED
|
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
91
91
|
|
|
92
92
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
93
93
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
94
|
-
> From a global install (`npm i -g "cowork-harness@>=1.0.
|
|
94
|
+
> From a global install (`npm i -g "cowork-harness@>=1.0.4"`), point at the package root instead:
|
|
95
95
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
96
96
|
> (or copy the cassette into your own project and pass that path).
|
|
97
97
|
|
|
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
101
101
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
102
102
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
103
103
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
104
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.0.
|
|
104
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.0.4"`.
|
|
105
105
|
|
|
106
106
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
107
107
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
126
126
|
claude plugin install cowork-harness@cowork-harness
|
|
127
127
|
```
|
|
128
128
|
|
|
129
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.0.
|
|
129
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.0.4"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
130
130
|
|
|
131
131
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
132
132
|
|
|
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
|
|
|
147
147
|
To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
|
|
148
148
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
149
149
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
150
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.0.
|
|
150
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.0.4"` — see
|
|
151
151
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
152
152
|
|
|
153
153
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -665,7 +665,7 @@ jobs:
|
|
|
665
665
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
666
666
|
```
|
|
667
667
|
|
|
668
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.0.
|
|
668
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.0.4` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
|
|
669
669
|
|
|
670
670
|
The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
671
671
|
|
|
@@ -826,6 +826,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
826
826
|
## Status
|
|
827
827
|
|
|
828
828
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
829
|
-
**`desktop-1.20186.
|
|
829
|
+
**`desktop-1.20186.9`**. Release-by-release verification notes (what was re-verified against
|
|
830
830
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
831
831
|
this section used to duplicate lives in the sections above.
|
package/RELEASING.md
CHANGED
|
@@ -184,7 +184,9 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
|
|
|
184
184
|
git tag -f vX vX.Y.Z && git tag -f vX.Y vX.Y.Z # e.g. v1 and v1.0 → v1.2.3
|
|
185
185
|
git push -f origin vX vX.Y
|
|
186
186
|
```
|
|
187
|
-
(Force-moving these ALIAS tags is expected; never force-move the immutable `vX.Y.Z` release tag.
|
|
187
|
+
(Force-moving these ALIAS tags is expected; never force-move the immutable `vX.Y.Z` release tag.
|
|
188
|
+
As of 1.0.4 the alias tags do NOT trigger `release.yml` / `publish-image.yml` — their `on.push.tags`
|
|
189
|
+
globs match full `vX.Y.Z` semver only — so pushing them produces no workflow runs at all.)
|
|
188
190
|
- [ ] Smoke the published artifact: `npx cowork-harness@X.Y.Z --version` and
|
|
189
191
|
`npx cowork-harness@X.Y.Z doctor --tier protocol`.
|
|
190
192
|
|
|
@@ -0,0 +1,380 @@
|
|
|
1
|
+
{
|
|
2
|
+
"baselineVersion": 1,
|
|
3
|
+
"appVersion": "1.20186.9",
|
|
4
|
+
"agentVersion": "2.1.205",
|
|
5
|
+
"agentBinary": {
|
|
6
|
+
"stagedPath": "~/Library/Application Support/Claude/claude-code-vm/2.1.205/claude",
|
|
7
|
+
"format": "elf-aarch64",
|
|
8
|
+
"nativeStagedPath": "~/Library/Application Support/Claude/claude-code/2.1.209/claude.app/Contents/MacOS/claude",
|
|
9
|
+
"sha256": "c1874c85bcd3a88b70439fd50ff5910b7e6ac5371c14dd49d4ccc2878a592d09",
|
|
10
|
+
"shaProvenance": "measured-local",
|
|
11
|
+
"manifestChecksumMatch": true
|
|
12
|
+
},
|
|
13
|
+
"guest": {
|
|
14
|
+
"os": "linux",
|
|
15
|
+
"arch": "arm64",
|
|
16
|
+
"baseImage": "ubuntu:22.04"
|
|
17
|
+
},
|
|
18
|
+
"spawn": {
|
|
19
|
+
"configDirInGuest": "mnt/.claude",
|
|
20
|
+
"settingSources": [
|
|
21
|
+
"user"
|
|
22
|
+
],
|
|
23
|
+
"permissionMode": "default",
|
|
24
|
+
"maxThinkingTokens": 31999,
|
|
25
|
+
"effortDefault": "medium",
|
|
26
|
+
"effortByModel": {
|
|
27
|
+
"claude-haiku-4-5": {
|
|
28
|
+
"modes": [
|
|
29
|
+
"extended"
|
|
30
|
+
]
|
|
31
|
+
},
|
|
32
|
+
"claude-sonnet-4-5": {
|
|
33
|
+
"modes": [
|
|
34
|
+
"extended"
|
|
35
|
+
]
|
|
36
|
+
},
|
|
37
|
+
"claude-sonnet-4-6": {
|
|
38
|
+
"effortLevels": [
|
|
39
|
+
"low",
|
|
40
|
+
"medium",
|
|
41
|
+
"high",
|
|
42
|
+
"max"
|
|
43
|
+
],
|
|
44
|
+
"recommended": "low",
|
|
45
|
+
"modes": [
|
|
46
|
+
"auto"
|
|
47
|
+
]
|
|
48
|
+
},
|
|
49
|
+
"claude-opus-4-6": {
|
|
50
|
+
"effortLevels": [
|
|
51
|
+
"low",
|
|
52
|
+
"medium",
|
|
53
|
+
"high",
|
|
54
|
+
"max"
|
|
55
|
+
],
|
|
56
|
+
"recommended": "medium",
|
|
57
|
+
"modes": [
|
|
58
|
+
"extended"
|
|
59
|
+
]
|
|
60
|
+
},
|
|
61
|
+
"claude-opus-4-7": {
|
|
62
|
+
"effortLevels": [
|
|
63
|
+
"low",
|
|
64
|
+
"medium",
|
|
65
|
+
"high",
|
|
66
|
+
"xhigh",
|
|
67
|
+
"max"
|
|
68
|
+
],
|
|
69
|
+
"recommended": "xhigh",
|
|
70
|
+
"modes": [
|
|
71
|
+
"auto"
|
|
72
|
+
]
|
|
73
|
+
},
|
|
74
|
+
"claude-opus-4-8": {
|
|
75
|
+
"effortLevels": [
|
|
76
|
+
"low",
|
|
77
|
+
"medium",
|
|
78
|
+
"high",
|
|
79
|
+
"xhigh",
|
|
80
|
+
"max"
|
|
81
|
+
],
|
|
82
|
+
"recommended": "high",
|
|
83
|
+
"modes": [
|
|
84
|
+
"auto"
|
|
85
|
+
]
|
|
86
|
+
}
|
|
87
|
+
},
|
|
88
|
+
"effortRegexDefault": {
|
|
89
|
+
"pattern": "^(?:claude-)?(?:fable|mythos)(?:-|$)",
|
|
90
|
+
"effortLevels": [
|
|
91
|
+
"low",
|
|
92
|
+
"medium",
|
|
93
|
+
"high",
|
|
94
|
+
"xhigh",
|
|
95
|
+
"max"
|
|
96
|
+
],
|
|
97
|
+
"recommended": "high",
|
|
98
|
+
"modes": [
|
|
99
|
+
"auto"
|
|
100
|
+
],
|
|
101
|
+
"disallowThinkingDisabled": true
|
|
102
|
+
},
|
|
103
|
+
"tools": [
|
|
104
|
+
"Task",
|
|
105
|
+
"Bash",
|
|
106
|
+
"Glob",
|
|
107
|
+
"Grep",
|
|
108
|
+
"Read",
|
|
109
|
+
"Edit",
|
|
110
|
+
"Write",
|
|
111
|
+
"NotebookEdit",
|
|
112
|
+
"WebFetch",
|
|
113
|
+
"TaskCreate",
|
|
114
|
+
"TaskUpdate",
|
|
115
|
+
"TaskGet",
|
|
116
|
+
"TaskList",
|
|
117
|
+
"TaskStop",
|
|
118
|
+
"WebSearch",
|
|
119
|
+
"Skill",
|
|
120
|
+
"REPL",
|
|
121
|
+
"JavaScript",
|
|
122
|
+
"AskUserQuestion",
|
|
123
|
+
"ToolSearch"
|
|
124
|
+
],
|
|
125
|
+
"allowedTools": [
|
|
126
|
+
"Task",
|
|
127
|
+
"Bash",
|
|
128
|
+
"Glob",
|
|
129
|
+
"Grep",
|
|
130
|
+
"Read",
|
|
131
|
+
"Edit",
|
|
132
|
+
"Write",
|
|
133
|
+
"NotebookEdit",
|
|
134
|
+
"WebFetch",
|
|
135
|
+
"TaskCreate",
|
|
136
|
+
"TaskUpdate",
|
|
137
|
+
"TaskGet",
|
|
138
|
+
"TaskList",
|
|
139
|
+
"TaskStop",
|
|
140
|
+
"WebSearch",
|
|
141
|
+
"Skill",
|
|
142
|
+
"REPL",
|
|
143
|
+
"JavaScript",
|
|
144
|
+
"ToolSearch"
|
|
145
|
+
],
|
|
146
|
+
"env": {
|
|
147
|
+
"CLAUDE_CODE_IS_COWORK": "1",
|
|
148
|
+
"CLAUDE_CODE_ENTRYPOINT": "local-agent",
|
|
149
|
+
"CLAUDE_CODE_TAGS": "lam_session_type:chat",
|
|
150
|
+
"CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST": "1",
|
|
151
|
+
"CLAUDE_CODE_ENABLE_ASK_USER_QUESTION_TOOL": "true",
|
|
152
|
+
"CLAUDE_CODE_DISABLE_CRON": "1",
|
|
153
|
+
"CLAUDE_CODE_DISABLE_BACKGROUND_TASKS": "1",
|
|
154
|
+
"CLAUDE_CODE_DISABLE_AGENTS_FLEET": "1",
|
|
155
|
+
"CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT": "1",
|
|
156
|
+
"CLAUDE_CODE_ENABLE_TASKS": "true",
|
|
157
|
+
"CLAUDE_CODE_DISABLE_TERMINAL_TITLE": "1",
|
|
158
|
+
"ENABLE_PROMPT_CACHING_1H": "1",
|
|
159
|
+
"DISABLE_MICROCOMPACT": "1",
|
|
160
|
+
"MCP_CONNECTION_NONBLOCKING": "true",
|
|
161
|
+
"API_TIMEOUT_MS": "900000",
|
|
162
|
+
"CLAUDE_CODE_EMIT_TOOL_USE_SUMMARIES": "",
|
|
163
|
+
"CLAUDE_CODE_ENABLE_FINE_GRAINED_TOOL_STREAMING": "1",
|
|
164
|
+
"DISABLE_AUTOUPDATER": "1",
|
|
165
|
+
"MCP_TOOL_TIMEOUT": "60000",
|
|
166
|
+
"USE_LOCAL_OAUTH": "",
|
|
167
|
+
"USE_STAGING_OAUTH": ""
|
|
168
|
+
},
|
|
169
|
+
"promptTemplate": "prompts/desktop-1.18286.0/system-prompt-append.md",
|
|
170
|
+
"subagentAppend": "prompts/desktop-1.15200.0/subagent-append-vm.md",
|
|
171
|
+
"subagentAppendHostLoop": "prompts/desktop-1.18286.2/subagent-append-hl.md",
|
|
172
|
+
"$comment": "Binary-verified Desktop->agent spawn contract, re-derived per release. spawn.env is GENERATED by deriveSpawnEnv() in src/sync/cowork-sync.ts (windowed enumeration of the asar env construction + gate/const value resolution); the scalar options, tools/allowedTools, and prompt-asset pointers are sentinel-guarded by checkSpawnContractFacts(). Do not hand-edit spawn.env — re-run sync.",
|
|
173
|
+
"$comment_handPinned": "Why the NON-env spawn fields stay hand-pinned: each is built in the asar as a non-literal expression (a session-path template, a session-type ternary, a const indirection, or a head+spread+tail array), so the windowed-enumeration generator that derives spawn.env cannot construct their VALUES without a full JS evaluator; instead each value was binary-verified once and is drift-guarded by a checkSpawnContractFacts() sentinel (cowork-sync.ts) that re-asserts the asar-side FACT at every sync. Scope caveat: the sentinels make DESKTOP-side drift loud; they do not validate this committed JSON itself — an erroneous hand-edit here is invisible to them.",
|
|
174
|
+
"$comment_configDirInGuest": "Hand-pinned: the asar builds it as a per-session path template (/sessions/${id}/mnt/.claude), not a constructable literal. Sentinel S1 pins the template shape.",
|
|
175
|
+
"$comment_settingSources": "Hand-pinned: sentinel S2 pins the settingSources:[\"user\"] literal.",
|
|
176
|
+
"$comment_permissionMode": "Hand-pinned: the asar computes it via a session ternary whose chat-session branch resolves to \"default\". Sentinel S3 pins the ternary shape.",
|
|
177
|
+
"$comment_maxThinkingTokens": "Hand-pinned: the asar reaches the value through const indirection. Sentinel S4 VALUE-pins the resolved const to 31999.",
|
|
178
|
+
"$comment_effortDefault": "Hand-pinned: sentinel S5 pins the .effort … :\"medium\" default.",
|
|
179
|
+
"$comment_tools": "Hand-pinned: the asar builds tools[] as head-list + Task-tools spread + session-type tail, not one literal. Sentinels S6 (head), S7 (the TaskCreate…TaskStop spread), S8 (tail-guard after ToolSearch) pin all three parts.",
|
|
180
|
+
"$comment_allowedTools": "Hand-pinned: same head+spread+tail construction as tools[], minus AskUserQuestion (tools-only by design). Sentinels S9 (head) and S10 (the built-in→mcp__ boundary tail-guard) pin it.",
|
|
181
|
+
"$comment_promptTemplate": "Hand-pinned pointer to a RECONSTRUCTED asset (see $comment_prompts) — the generator cannot extract prose. Sentinel S15 pins the claude_code preset-append delivery site.",
|
|
182
|
+
"$comment_subagentAppend": "Hand-pinned pointer to a reconstructed asset (see $comment_prompts). Sentinel S16 pins the per-session appendSubagentSystemPrompt generator call shape.",
|
|
183
|
+
"$comment_subagentAppendHostLoop": "Hand-pinned pointer to a reconstructed (paraphrased) asset for the HOST-LOOP branch of the per-session sub-agent append (section key subagent_env_hl; selected purely on hostLoopMode). Backfilled only for release families whose hl text is binary-verified byte-identical (1.18286.2+). Sentinel: checkSubagentPromptFacts (two-branch fingerprint + substitution-value proofs). A hostloop run on a baseline lacking this pointer fails loud rather than falling back to the VM text.",
|
|
184
|
+
"$comment_notSet": "Deliberately NOT set: CLAUDE_CODE_USE_COWORK_PLUGINS (Desktop never sets it; would flip the agent to cowork_settings.json/cowork_plugins — asserted absent by the S17 negative invariant). Enumerated-but-not-pinned keys are enforced by SPAWN_ENV_ALLOWLIST in src/sync/cowork-sync.ts, each with a reason; categories: host-derived (CLAUDE_CONFIG_DIR, TZ, HOST_PLATFORM, OAUTH_TOKEN/BASE_URL/CUSTOM_HEADERS, account UUIDs, WORKSPACE_HOST_PATHS, OTEL), constructed-then-deleted (ANTHROPIC_API_KEY/AUTH_TOKEN via FnA), gate-conditional-off (MCP_CONNECT_TIMEOUT_MS, ENABLE_TOOL_SEARCH, SKIP_PRECOMPACT_LOAD), non-chat/project-session (BRIEF*, PROJECT*), user-settings (SUBAGENT_MODEL, AUTO_COMPACT_WINDOW, ...), and 3p-provider-only branches. The opaque ...g.env/...l session spreads are the known static-extraction blind spot (runtime lane backstop).",
|
|
185
|
+
"$comment_prompts": "Reconstructed cowork-specific sections, re-paraphrased from asar 1.18286.0 const aui (system prompt; RESTRUCTURED at this release — see the asset header) and 1.15200.0 generator CVr (subagent; verified unchanged in the 1.18286.0 asar, generator Zgn). Not the full base prompt (not cleanly extractable); generic refusal/safety policy elided. Delivered via --append-system-prompt (layered on the agent's built-in base prompt), NOT the initialize handshake; only the subagent append goes over initialize (appendSubagentSystemPrompt), gated on CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT."
|
|
186
|
+
},
|
|
187
|
+
"mountLayout": {
|
|
188
|
+
"sessionRoot": "/sessions/{sessionId}",
|
|
189
|
+
"cwd": "/sessions/{sessionId}",
|
|
190
|
+
"mntRoot": "/sessions/{sessionId}/mnt",
|
|
191
|
+
"mounts": [
|
|
192
|
+
{
|
|
193
|
+
"name": "uploads",
|
|
194
|
+
"mountPath": "uploads",
|
|
195
|
+
"mode": "r",
|
|
196
|
+
"purpose": "user-uploaded files (read-only — asar 'ro')"
|
|
197
|
+
},
|
|
198
|
+
{
|
|
199
|
+
"name": "projects",
|
|
200
|
+
"mountPath": ".projects/{projectId}",
|
|
201
|
+
"mode": "rw",
|
|
202
|
+
"purpose": "RESERVED namespace (and the separate UUID project-sync feature) — NOT the work-folder path. From Desktop 1.14271.0 selected work folders mount at mnt/<collision-resolved-basename> (dynamic, derived per session by buildLaunchPlan; see MOUNT_BARE_NAME_MIN_VERSION). This decorative row is not consumed for binding (staged paths come from plan.mounts)."
|
|
203
|
+
},
|
|
204
|
+
{
|
|
205
|
+
"name": "local-plugins",
|
|
206
|
+
"mountPath": ".local-plugins/marketplaces",
|
|
207
|
+
"mode": "r",
|
|
208
|
+
"purpose": "marketplace skills/plugins, runtime-discovered"
|
|
209
|
+
},
|
|
210
|
+
{
|
|
211
|
+
"name": "remote-plugins",
|
|
212
|
+
"mountPath": ".remote-plugins",
|
|
213
|
+
"mode": "r",
|
|
214
|
+
"purpose": "org-remote plugins, runtime-discovered"
|
|
215
|
+
},
|
|
216
|
+
{
|
|
217
|
+
"name": "outputs",
|
|
218
|
+
"mountPath": "outputs",
|
|
219
|
+
"mode": "rw",
|
|
220
|
+
"purpose": "session outputs/artifacts — delete denied by default (asar IX); rwd only when approved"
|
|
221
|
+
},
|
|
222
|
+
{
|
|
223
|
+
"name": "skills",
|
|
224
|
+
"mountPath": ".claude/skills",
|
|
225
|
+
"mode": "r",
|
|
226
|
+
"purpose": "personal/saved skill doc bodies (NOT plugin-bundled skills, which live under local-plugins/remote-plugins above) — real VM confirmed via systemd unit sessions-<name>-mnt-.claude-skills.mount in vm_bundles/claudevm.bundle/rootfs.img. Decorative row like 'projects' above (not consumed for binding — resolveMounts()'s mounts[] is destructured away at every call site); the harness reproduces this via CLAUDE_CONFIG_DIR staging (session.ts skill copy + stage.ts cpSync), not a plan.mounts bind — see hostloop-prompt.ts's asar-verified skills bullet."
|
|
227
|
+
}
|
|
228
|
+
]
|
|
229
|
+
},
|
|
230
|
+
"network": {
|
|
231
|
+
"mode": "gvisor",
|
|
232
|
+
"allowKind": "allowlist",
|
|
233
|
+
"allowDomains": [
|
|
234
|
+
"preview.claude.ai",
|
|
235
|
+
"downloads.claude.ai",
|
|
236
|
+
"api.anthropic.com",
|
|
237
|
+
"a-cdn.anthropic.com",
|
|
238
|
+
"a-api.anthropic.com",
|
|
239
|
+
"assets.claude.ai",
|
|
240
|
+
"sentry.io",
|
|
241
|
+
"console.anthropic.com",
|
|
242
|
+
"api-staging.anthropic.com",
|
|
243
|
+
"www.anthropic.com",
|
|
244
|
+
"api.claude.ai",
|
|
245
|
+
"support.anthropic.com",
|
|
246
|
+
"docs.anthropic.com",
|
|
247
|
+
"mcp-proxy.anthropic.com",
|
|
248
|
+
"pivot.claude.ai"
|
|
249
|
+
]
|
|
250
|
+
},
|
|
251
|
+
"bgEnvStrip": {
|
|
252
|
+
"knownVars": [
|
|
253
|
+
"CLAUDE_CODE_OAUTH_TOKEN",
|
|
254
|
+
"CLAUDE_CODE_SESSION_KIND",
|
|
255
|
+
"CLAUDE_CODE_SESSION_ID",
|
|
256
|
+
"CLAUDE_CODE_SESSION_NAME",
|
|
257
|
+
"CLAUDE_CODE_SESSION_LOG"
|
|
258
|
+
]
|
|
259
|
+
},
|
|
260
|
+
"$comment": "Platform baseline auto-derived by `cowork-harness sync` from a live Claude Desktop install + app.asar. VOLATILE per-release facts only. Regenerate per release; review the diff. Captured 2026-07-14 on macOS arm64.",
|
|
261
|
+
"capturedAt": "2026-07-14",
|
|
262
|
+
"platform": "darwin-arm64",
|
|
263
|
+
"settings": {
|
|
264
|
+
"autoMountFolders": {
|
|
265
|
+
"key": "autoMountFolders",
|
|
266
|
+
"default": false
|
|
267
|
+
},
|
|
268
|
+
"localAgentModeTrustedFolders": {
|
|
269
|
+
"key": "localAgentModeTrustedFolders",
|
|
270
|
+
"default": []
|
|
271
|
+
}
|
|
272
|
+
},
|
|
273
|
+
"provenance": {
|
|
274
|
+
"asarPath": "/Applications/Claude.app/Contents/Resources/app.asar",
|
|
275
|
+
"asarFingerprint": "fdd33387498c9880",
|
|
276
|
+
"gates": {
|
|
277
|
+
"$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift.",
|
|
278
|
+
"emitToolUseSummaries:66187241": {
|
|
279
|
+
"on": false,
|
|
280
|
+
"source": "defaultValue",
|
|
281
|
+
"value": false
|
|
282
|
+
},
|
|
283
|
+
"autoMemoryStandardSessions:123929380": {
|
|
284
|
+
"on": false,
|
|
285
|
+
"source": "defaultValue",
|
|
286
|
+
"value": false
|
|
287
|
+
},
|
|
288
|
+
"subagentPromptServerOverride:124685897": {
|
|
289
|
+
"on": false,
|
|
290
|
+
"source": "defaultValue",
|
|
291
|
+
"value": false
|
|
292
|
+
},
|
|
293
|
+
"mcpConnectionNonblockingOff:434204418": {
|
|
294
|
+
"on": false,
|
|
295
|
+
"source": "defaultValue",
|
|
296
|
+
"value": false
|
|
297
|
+
},
|
|
298
|
+
"bridgeSdkTransport:583857784": {
|
|
299
|
+
"on": true,
|
|
300
|
+
"source": "force",
|
|
301
|
+
"value": true,
|
|
302
|
+
"note": "— Cowork uses the SDK-based transport (control protocol), confirming the harness's sdkMcpServers/mcp_message path is the production transport."
|
|
303
|
+
},
|
|
304
|
+
"fineGrainedToolStreaming:714014285": {
|
|
305
|
+
"on": true,
|
|
306
|
+
"source": "force",
|
|
307
|
+
"value": true
|
|
308
|
+
},
|
|
309
|
+
"enableToolSearchAuto:1129419822": {
|
|
310
|
+
"on": false,
|
|
311
|
+
"source": "absent"
|
|
312
|
+
},
|
|
313
|
+
"hostLoop:1143815894": {
|
|
314
|
+
"on": true,
|
|
315
|
+
"source": "force",
|
|
316
|
+
"value": true
|
|
317
|
+
},
|
|
318
|
+
"scheduledTaskSessionLimiter:1648655587": {
|
|
319
|
+
"on": true,
|
|
320
|
+
"source": "force",
|
|
321
|
+
"value": {
|
|
322
|
+
"global": 3,
|
|
323
|
+
"perTask": 1
|
|
324
|
+
},
|
|
325
|
+
"note": "SCHEDULED-TASK (cron) session limiter — NOT an in-conversation Task-tool cap (binary-verified 2026-07-04, asar 1.18286.0 class L9t [ScheduledTasks]). perTask=1: <=1 concurrent session PER SCHEDULED TASK; global=3: <=3 concurrent scheduled-task sessions globally (+_pendingTaskDispatches). Host-side SKIP (recordSkipAndEmit/PerTaskLimit|GlobalLimit — NOT queue/deny). Cowork imposes no cap on Task-tool sub-agent fan-out; the harness has no scheduled-task scheduler, so this gate has no applicable surface — pinned as a sync drift-sentinel only."
|
|
326
|
+
},
|
|
327
|
+
"memoryGuidelinesEnv:1696890383": {
|
|
328
|
+
"on": false,
|
|
329
|
+
"source": "defaultValue",
|
|
330
|
+
"value": false
|
|
331
|
+
},
|
|
332
|
+
"oauthScopesEnv:1936081873": {
|
|
333
|
+
"on": true,
|
|
334
|
+
"source": "force",
|
|
335
|
+
"value": true
|
|
336
|
+
},
|
|
337
|
+
"coworkRuntimeConfig:1978029737": {
|
|
338
|
+
"on": true,
|
|
339
|
+
"source": "force",
|
|
340
|
+
"value": {
|
|
341
|
+
"coworkNativeFilePreview": true,
|
|
342
|
+
"coworkWebFetchPrompt": true,
|
|
343
|
+
"coworkWebFetchViaApi": true,
|
|
344
|
+
"sessionsBridgePollBlockMs": 30,
|
|
345
|
+
"workspaceBashWaitLonger": true
|
|
346
|
+
},
|
|
347
|
+
"note": "coworkWebFetchViaApi=true coworkWebFetchPrompt=true workspaceBashWaitLonger=true sessionsBridgePollBlockMs=30 — web_fetch is host/API-routed (POST /api/organizations/<org>/cowork/web_fetch), NOT container egress; gated by a separate web-fetch hostname allowlist + URL provenance."
|
|
348
|
+
},
|
|
349
|
+
"cliPlugin:2307090146": {
|
|
350
|
+
"on": false,
|
|
351
|
+
"source": "defaultValue",
|
|
352
|
+
"value": false,
|
|
353
|
+
"note": "— the CLI-plugin credential broker is dark-launched off for standard interactive accounts (Ch23/L106)."
|
|
354
|
+
},
|
|
355
|
+
"pluginSyncSparkplug:2340532315": {
|
|
356
|
+
"on": true,
|
|
357
|
+
"source": "force",
|
|
358
|
+
"value": true,
|
|
359
|
+
"note": "— startup syncPlugins(); plugins load via --plugin-dir (registry inert in-VM)."
|
|
360
|
+
},
|
|
361
|
+
"skeletonHome:2614807392": {
|
|
362
|
+
"on": false,
|
|
363
|
+
"source": "absent"
|
|
364
|
+
},
|
|
365
|
+
"memoryExtraGuidelines:2860753854": {
|
|
366
|
+
"on": true,
|
|
367
|
+
"source": "defaultValue",
|
|
368
|
+
"value": "## Sensitive personal information\n\nDo not save the following to memory unless the user explicitly asks you to remember it:\n\n- Protected attributes: race, ethnicity, national origin, religion, age, sex, sexual orientation, gender identity, immigration status, disability, serious illness, union membership\n- Government identifiers: Social Security numbers, driver's license numbers, passport numbers, government ID numbers\n- Financial account details: credit card numbers, bank account numbers\n- Health information: medical conditions, diagnoses, lab results, mental health details, therapy or counseling\n- Home or personal mailing addresses (work addresses are fine)\n- Account passwords, secret tokens, or secret keys\n\nIf any of the above appears in conversation context, complete the task but do not persist it to a memory file. If the user explicitly says \"remember my address is X\", saving it is acceptable — they've given consent."
|
|
369
|
+
},
|
|
370
|
+
"skipPrecompactLoad:4153934152": {
|
|
371
|
+
"on": false,
|
|
372
|
+
"source": "defaultValue",
|
|
373
|
+
"value": false
|
|
374
|
+
}
|
|
375
|
+
},
|
|
376
|
+
"eipcChannelUuid": "4f426349-8d6f-45f3-ae22-280fef323564",
|
|
377
|
+
"$comment": "eipcChannelUuid is per-build; recorded for provenance only — the harness does not use Desktop IPC."
|
|
378
|
+
},
|
|
379
|
+
"requireFullVmSandbox": null
|
|
380
|
+
}
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.0.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.0.4"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
|
@@ -111,7 +111,7 @@
|
|
|
111
111
|
},
|
|
112
112
|
"scenarioSource": "../scenarios/example-pdf-skill.yaml",
|
|
113
113
|
"fingerprint": {
|
|
114
|
-
"baseline": "1.20186.
|
|
114
|
+
"baseline": "1.20186.9",
|
|
115
115
|
"skillHash": "b760ea90682777367fbd44d866ea66d316452da1abb18bf8cb61f1cec8e67806",
|
|
116
116
|
"contentSig": "ddf68cf6d74e0599729589f8be4f0a9eec440093c99c1060cd6f47dd671a0dfd",
|
|
117
117
|
"skillSources": [
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "1.0.
|
|
3
|
+
"version": "1.0.4",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios — same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|