cowork-harness 1.0.2 → 1.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.0.2
7
- tracks-harness: cowork-harness 1.0.2 (baseline desktop-1.20186.1)
6
+ version: 1.0.3
7
+ tracks-harness: cowork-harness 1.0.3 (baseline desktop-1.20186.9)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0.2` (baseline
26
- > `desktop-1.20186.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0.3` (baseline
26
+ > `desktop-1.20186.9`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0.2**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.0.2" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.0.2"`. **Pin `@>=1.0.2`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0.3**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.0.3" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.0.3"`. **Pin `@>=1.0.3`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
- What the ≥ 1.0.2 floor gates, by release:
44
+ What the ≥ 1.0.3 floor gates, by release:
45
45
 
46
46
  - **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
47
47
  - **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.0.2` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.0.3` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.0.2"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.0.3"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.0.2"
60
+ - run: npm i -g "cowork-harness@>=1.0.3"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.0.2"
200
+ - run: npm i -g "cowork-harness@>=1.0.3"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.0.2"
229
+ run: npm i -g "cowork-harness@>=1.0.3"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.0.2` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.0.3` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0.2`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0.3`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,19 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.0.3] — 2026-07-14
10
+
11
+ Patch: parity sync to Claude Desktop `1.20186.9`. No runtime/API change.
12
+
13
+ ### Changed
14
+
15
+ - Synced the platform baseline to Claude Desktop `1.20186.9`
16
+ (`baselines/desktop-1.20186.9.json`, now what `baseline: latest` resolves to). A routine
17
+ per-release parity refresh: the app version, the native agent staging path, and the asar
18
+ fingerprint moved; the Cowork system prompt, egress allowlist, gate states, and agent (VM)
19
+ version are unchanged from `1.20186.1`. README and the companion skill's baseline pointer were
20
+ updated to match.
21
+
9
22
  ## [1.0.2] — 2026-07-14
10
23
 
11
24
  Patch: shorten the Action's Marketplace tagline. No runtime/API change.
package/README.md CHANGED
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
91
91
 
92
92
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
93
93
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
94
- > From a global install (`npm i -g "cowork-harness@>=1.0.2"`), point at the package root instead:
94
+ > From a global install (`npm i -g "cowork-harness@>=1.0.3"`), point at the package root instead:
95
95
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
96
96
  > (or copy the cassette into your own project and pass that path).
97
97
 
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
101
101
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
102
102
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
103
103
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
104
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.0.2"`.
104
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.0.3"`.
105
105
 
106
106
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
107
107
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
126
126
  claude plugin install cowork-harness@cowork-harness
127
127
  ```
128
128
 
129
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.0.2"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
129
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.0.3"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
130
130
 
131
131
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
132
132
 
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
147
147
  To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
148
148
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
149
149
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
150
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.0.2"` — see
150
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.0.3"` — see
151
151
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
152
152
 
153
153
  ### Prerequisites for anything above `protocol` fidelity
@@ -665,7 +665,7 @@ jobs:
665
665
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
666
666
  ```
667
667
 
668
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.0.2` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
668
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.0.3` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
669
669
 
670
670
  The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
671
671
 
@@ -826,6 +826,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
826
826
  ## Status
827
827
 
828
828
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
829
- **`desktop-1.20186.1`**. Release-by-release verification notes (what was re-verified against
829
+ **`desktop-1.20186.9`**. Release-by-release verification notes (what was re-verified against
830
830
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
831
831
  this section used to duplicate lives in the sections above.
@@ -0,0 +1,380 @@
1
+ {
2
+ "baselineVersion": 1,
3
+ "appVersion": "1.20186.9",
4
+ "agentVersion": "2.1.205",
5
+ "agentBinary": {
6
+ "stagedPath": "~/Library/Application Support/Claude/claude-code-vm/2.1.205/claude",
7
+ "format": "elf-aarch64",
8
+ "nativeStagedPath": "~/Library/Application Support/Claude/claude-code/2.1.209/claude.app/Contents/MacOS/claude",
9
+ "sha256": "c1874c85bcd3a88b70439fd50ff5910b7e6ac5371c14dd49d4ccc2878a592d09",
10
+ "shaProvenance": "measured-local",
11
+ "manifestChecksumMatch": true
12
+ },
13
+ "guest": {
14
+ "os": "linux",
15
+ "arch": "arm64",
16
+ "baseImage": "ubuntu:22.04"
17
+ },
18
+ "spawn": {
19
+ "configDirInGuest": "mnt/.claude",
20
+ "settingSources": [
21
+ "user"
22
+ ],
23
+ "permissionMode": "default",
24
+ "maxThinkingTokens": 31999,
25
+ "effortDefault": "medium",
26
+ "effortByModel": {
27
+ "claude-haiku-4-5": {
28
+ "modes": [
29
+ "extended"
30
+ ]
31
+ },
32
+ "claude-sonnet-4-5": {
33
+ "modes": [
34
+ "extended"
35
+ ]
36
+ },
37
+ "claude-sonnet-4-6": {
38
+ "effortLevels": [
39
+ "low",
40
+ "medium",
41
+ "high",
42
+ "max"
43
+ ],
44
+ "recommended": "low",
45
+ "modes": [
46
+ "auto"
47
+ ]
48
+ },
49
+ "claude-opus-4-6": {
50
+ "effortLevels": [
51
+ "low",
52
+ "medium",
53
+ "high",
54
+ "max"
55
+ ],
56
+ "recommended": "medium",
57
+ "modes": [
58
+ "extended"
59
+ ]
60
+ },
61
+ "claude-opus-4-7": {
62
+ "effortLevels": [
63
+ "low",
64
+ "medium",
65
+ "high",
66
+ "xhigh",
67
+ "max"
68
+ ],
69
+ "recommended": "xhigh",
70
+ "modes": [
71
+ "auto"
72
+ ]
73
+ },
74
+ "claude-opus-4-8": {
75
+ "effortLevels": [
76
+ "low",
77
+ "medium",
78
+ "high",
79
+ "xhigh",
80
+ "max"
81
+ ],
82
+ "recommended": "high",
83
+ "modes": [
84
+ "auto"
85
+ ]
86
+ }
87
+ },
88
+ "effortRegexDefault": {
89
+ "pattern": "^(?:claude-)?(?:fable|mythos)(?:-|$)",
90
+ "effortLevels": [
91
+ "low",
92
+ "medium",
93
+ "high",
94
+ "xhigh",
95
+ "max"
96
+ ],
97
+ "recommended": "high",
98
+ "modes": [
99
+ "auto"
100
+ ],
101
+ "disallowThinkingDisabled": true
102
+ },
103
+ "tools": [
104
+ "Task",
105
+ "Bash",
106
+ "Glob",
107
+ "Grep",
108
+ "Read",
109
+ "Edit",
110
+ "Write",
111
+ "NotebookEdit",
112
+ "WebFetch",
113
+ "TaskCreate",
114
+ "TaskUpdate",
115
+ "TaskGet",
116
+ "TaskList",
117
+ "TaskStop",
118
+ "WebSearch",
119
+ "Skill",
120
+ "REPL",
121
+ "JavaScript",
122
+ "AskUserQuestion",
123
+ "ToolSearch"
124
+ ],
125
+ "allowedTools": [
126
+ "Task",
127
+ "Bash",
128
+ "Glob",
129
+ "Grep",
130
+ "Read",
131
+ "Edit",
132
+ "Write",
133
+ "NotebookEdit",
134
+ "WebFetch",
135
+ "TaskCreate",
136
+ "TaskUpdate",
137
+ "TaskGet",
138
+ "TaskList",
139
+ "TaskStop",
140
+ "WebSearch",
141
+ "Skill",
142
+ "REPL",
143
+ "JavaScript",
144
+ "ToolSearch"
145
+ ],
146
+ "env": {
147
+ "CLAUDE_CODE_IS_COWORK": "1",
148
+ "CLAUDE_CODE_ENTRYPOINT": "local-agent",
149
+ "CLAUDE_CODE_TAGS": "lam_session_type:chat",
150
+ "CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST": "1",
151
+ "CLAUDE_CODE_ENABLE_ASK_USER_QUESTION_TOOL": "true",
152
+ "CLAUDE_CODE_DISABLE_CRON": "1",
153
+ "CLAUDE_CODE_DISABLE_BACKGROUND_TASKS": "1",
154
+ "CLAUDE_CODE_DISABLE_AGENTS_FLEET": "1",
155
+ "CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT": "1",
156
+ "CLAUDE_CODE_ENABLE_TASKS": "true",
157
+ "CLAUDE_CODE_DISABLE_TERMINAL_TITLE": "1",
158
+ "ENABLE_PROMPT_CACHING_1H": "1",
159
+ "DISABLE_MICROCOMPACT": "1",
160
+ "MCP_CONNECTION_NONBLOCKING": "true",
161
+ "API_TIMEOUT_MS": "900000",
162
+ "CLAUDE_CODE_EMIT_TOOL_USE_SUMMARIES": "",
163
+ "CLAUDE_CODE_ENABLE_FINE_GRAINED_TOOL_STREAMING": "1",
164
+ "DISABLE_AUTOUPDATER": "1",
165
+ "MCP_TOOL_TIMEOUT": "60000",
166
+ "USE_LOCAL_OAUTH": "",
167
+ "USE_STAGING_OAUTH": ""
168
+ },
169
+ "promptTemplate": "prompts/desktop-1.18286.0/system-prompt-append.md",
170
+ "subagentAppend": "prompts/desktop-1.15200.0/subagent-append-vm.md",
171
+ "subagentAppendHostLoop": "prompts/desktop-1.18286.2/subagent-append-hl.md",
172
+ "$comment": "Binary-verified Desktop->agent spawn contract, re-derived per release. spawn.env is GENERATED by deriveSpawnEnv() in src/sync/cowork-sync.ts (windowed enumeration of the asar env construction + gate/const value resolution); the scalar options, tools/allowedTools, and prompt-asset pointers are sentinel-guarded by checkSpawnContractFacts(). Do not hand-edit spawn.env — re-run sync.",
173
+ "$comment_handPinned": "Why the NON-env spawn fields stay hand-pinned: each is built in the asar as a non-literal expression (a session-path template, a session-type ternary, a const indirection, or a head+spread+tail array), so the windowed-enumeration generator that derives spawn.env cannot construct their VALUES without a full JS evaluator; instead each value was binary-verified once and is drift-guarded by a checkSpawnContractFacts() sentinel (cowork-sync.ts) that re-asserts the asar-side FACT at every sync. Scope caveat: the sentinels make DESKTOP-side drift loud; they do not validate this committed JSON itself — an erroneous hand-edit here is invisible to them.",
174
+ "$comment_configDirInGuest": "Hand-pinned: the asar builds it as a per-session path template (/sessions/${id}/mnt/.claude), not a constructable literal. Sentinel S1 pins the template shape.",
175
+ "$comment_settingSources": "Hand-pinned: sentinel S2 pins the settingSources:[\"user\"] literal.",
176
+ "$comment_permissionMode": "Hand-pinned: the asar computes it via a session ternary whose chat-session branch resolves to \"default\". Sentinel S3 pins the ternary shape.",
177
+ "$comment_maxThinkingTokens": "Hand-pinned: the asar reaches the value through const indirection. Sentinel S4 VALUE-pins the resolved const to 31999.",
178
+ "$comment_effortDefault": "Hand-pinned: sentinel S5 pins the .effort … :\"medium\" default.",
179
+ "$comment_tools": "Hand-pinned: the asar builds tools[] as head-list + Task-tools spread + session-type tail, not one literal. Sentinels S6 (head), S7 (the TaskCreate…TaskStop spread), S8 (tail-guard after ToolSearch) pin all three parts.",
180
+ "$comment_allowedTools": "Hand-pinned: same head+spread+tail construction as tools[], minus AskUserQuestion (tools-only by design). Sentinels S9 (head) and S10 (the built-in→mcp__ boundary tail-guard) pin it.",
181
+ "$comment_promptTemplate": "Hand-pinned pointer to a RECONSTRUCTED asset (see $comment_prompts) — the generator cannot extract prose. Sentinel S15 pins the claude_code preset-append delivery site.",
182
+ "$comment_subagentAppend": "Hand-pinned pointer to a reconstructed asset (see $comment_prompts). Sentinel S16 pins the per-session appendSubagentSystemPrompt generator call shape.",
183
+ "$comment_subagentAppendHostLoop": "Hand-pinned pointer to a reconstructed (paraphrased) asset for the HOST-LOOP branch of the per-session sub-agent append (section key subagent_env_hl; selected purely on hostLoopMode). Backfilled only for release families whose hl text is binary-verified byte-identical (1.18286.2+). Sentinel: checkSubagentPromptFacts (two-branch fingerprint + substitution-value proofs). A hostloop run on a baseline lacking this pointer fails loud rather than falling back to the VM text.",
184
+ "$comment_notSet": "Deliberately NOT set: CLAUDE_CODE_USE_COWORK_PLUGINS (Desktop never sets it; would flip the agent to cowork_settings.json/cowork_plugins — asserted absent by the S17 negative invariant). Enumerated-but-not-pinned keys are enforced by SPAWN_ENV_ALLOWLIST in src/sync/cowork-sync.ts, each with a reason; categories: host-derived (CLAUDE_CONFIG_DIR, TZ, HOST_PLATFORM, OAUTH_TOKEN/BASE_URL/CUSTOM_HEADERS, account UUIDs, WORKSPACE_HOST_PATHS, OTEL), constructed-then-deleted (ANTHROPIC_API_KEY/AUTH_TOKEN via FnA), gate-conditional-off (MCP_CONNECT_TIMEOUT_MS, ENABLE_TOOL_SEARCH, SKIP_PRECOMPACT_LOAD), non-chat/project-session (BRIEF*, PROJECT*), user-settings (SUBAGENT_MODEL, AUTO_COMPACT_WINDOW, ...), and 3p-provider-only branches. The opaque ...g.env/...l session spreads are the known static-extraction blind spot (runtime lane backstop).",
185
+ "$comment_prompts": "Reconstructed cowork-specific sections, re-paraphrased from asar 1.18286.0 const aui (system prompt; RESTRUCTURED at this release — see the asset header) and 1.15200.0 generator CVr (subagent; verified unchanged in the 1.18286.0 asar, generator Zgn). Not the full base prompt (not cleanly extractable); generic refusal/safety policy elided. Delivered via --append-system-prompt (layered on the agent's built-in base prompt), NOT the initialize handshake; only the subagent append goes over initialize (appendSubagentSystemPrompt), gated on CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT."
186
+ },
187
+ "mountLayout": {
188
+ "sessionRoot": "/sessions/{sessionId}",
189
+ "cwd": "/sessions/{sessionId}",
190
+ "mntRoot": "/sessions/{sessionId}/mnt",
191
+ "mounts": [
192
+ {
193
+ "name": "uploads",
194
+ "mountPath": "uploads",
195
+ "mode": "r",
196
+ "purpose": "user-uploaded files (read-only — asar 'ro')"
197
+ },
198
+ {
199
+ "name": "projects",
200
+ "mountPath": ".projects/{projectId}",
201
+ "mode": "rw",
202
+ "purpose": "RESERVED namespace (and the separate UUID project-sync feature) — NOT the work-folder path. From Desktop 1.14271.0 selected work folders mount at mnt/<collision-resolved-basename> (dynamic, derived per session by buildLaunchPlan; see MOUNT_BARE_NAME_MIN_VERSION). This decorative row is not consumed for binding (staged paths come from plan.mounts)."
203
+ },
204
+ {
205
+ "name": "local-plugins",
206
+ "mountPath": ".local-plugins/marketplaces",
207
+ "mode": "r",
208
+ "purpose": "marketplace skills/plugins, runtime-discovered"
209
+ },
210
+ {
211
+ "name": "remote-plugins",
212
+ "mountPath": ".remote-plugins",
213
+ "mode": "r",
214
+ "purpose": "org-remote plugins, runtime-discovered"
215
+ },
216
+ {
217
+ "name": "outputs",
218
+ "mountPath": "outputs",
219
+ "mode": "rw",
220
+ "purpose": "session outputs/artifacts — delete denied by default (asar IX); rwd only when approved"
221
+ },
222
+ {
223
+ "name": "skills",
224
+ "mountPath": ".claude/skills",
225
+ "mode": "r",
226
+ "purpose": "personal/saved skill doc bodies (NOT plugin-bundled skills, which live under local-plugins/remote-plugins above) — real VM confirmed via systemd unit sessions-<name>-mnt-.claude-skills.mount in vm_bundles/claudevm.bundle/rootfs.img. Decorative row like 'projects' above (not consumed for binding — resolveMounts()'s mounts[] is destructured away at every call site); the harness reproduces this via CLAUDE_CONFIG_DIR staging (session.ts skill copy + stage.ts cpSync), not a plan.mounts bind — see hostloop-prompt.ts's asar-verified skills bullet."
227
+ }
228
+ ]
229
+ },
230
+ "network": {
231
+ "mode": "gvisor",
232
+ "allowKind": "allowlist",
233
+ "allowDomains": [
234
+ "preview.claude.ai",
235
+ "downloads.claude.ai",
236
+ "api.anthropic.com",
237
+ "a-cdn.anthropic.com",
238
+ "a-api.anthropic.com",
239
+ "assets.claude.ai",
240
+ "sentry.io",
241
+ "console.anthropic.com",
242
+ "api-staging.anthropic.com",
243
+ "www.anthropic.com",
244
+ "api.claude.ai",
245
+ "support.anthropic.com",
246
+ "docs.anthropic.com",
247
+ "mcp-proxy.anthropic.com",
248
+ "pivot.claude.ai"
249
+ ]
250
+ },
251
+ "bgEnvStrip": {
252
+ "knownVars": [
253
+ "CLAUDE_CODE_OAUTH_TOKEN",
254
+ "CLAUDE_CODE_SESSION_KIND",
255
+ "CLAUDE_CODE_SESSION_ID",
256
+ "CLAUDE_CODE_SESSION_NAME",
257
+ "CLAUDE_CODE_SESSION_LOG"
258
+ ]
259
+ },
260
+ "$comment": "Platform baseline auto-derived by `cowork-harness sync` from a live Claude Desktop install + app.asar. VOLATILE per-release facts only. Regenerate per release; review the diff. Captured 2026-07-14 on macOS arm64.",
261
+ "capturedAt": "2026-07-14",
262
+ "platform": "darwin-arm64",
263
+ "settings": {
264
+ "autoMountFolders": {
265
+ "key": "autoMountFolders",
266
+ "default": false
267
+ },
268
+ "localAgentModeTrustedFolders": {
269
+ "key": "localAgentModeTrustedFolders",
270
+ "default": []
271
+ }
272
+ },
273
+ "provenance": {
274
+ "asarPath": "/Applications/Claude.app/Contents/Resources/app.asar",
275
+ "asarFingerprint": "fdd33387498c9880",
276
+ "gates": {
277
+ "$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift.",
278
+ "emitToolUseSummaries:66187241": {
279
+ "on": false,
280
+ "source": "defaultValue",
281
+ "value": false
282
+ },
283
+ "autoMemoryStandardSessions:123929380": {
284
+ "on": false,
285
+ "source": "defaultValue",
286
+ "value": false
287
+ },
288
+ "subagentPromptServerOverride:124685897": {
289
+ "on": false,
290
+ "source": "defaultValue",
291
+ "value": false
292
+ },
293
+ "mcpConnectionNonblockingOff:434204418": {
294
+ "on": false,
295
+ "source": "defaultValue",
296
+ "value": false
297
+ },
298
+ "bridgeSdkTransport:583857784": {
299
+ "on": true,
300
+ "source": "force",
301
+ "value": true,
302
+ "note": "— Cowork uses the SDK-based transport (control protocol), confirming the harness's sdkMcpServers/mcp_message path is the production transport."
303
+ },
304
+ "fineGrainedToolStreaming:714014285": {
305
+ "on": true,
306
+ "source": "force",
307
+ "value": true
308
+ },
309
+ "enableToolSearchAuto:1129419822": {
310
+ "on": false,
311
+ "source": "absent"
312
+ },
313
+ "hostLoop:1143815894": {
314
+ "on": true,
315
+ "source": "force",
316
+ "value": true
317
+ },
318
+ "scheduledTaskSessionLimiter:1648655587": {
319
+ "on": true,
320
+ "source": "force",
321
+ "value": {
322
+ "global": 3,
323
+ "perTask": 1
324
+ },
325
+ "note": "SCHEDULED-TASK (cron) session limiter — NOT an in-conversation Task-tool cap (binary-verified 2026-07-04, asar 1.18286.0 class L9t [ScheduledTasks]). perTask=1: <=1 concurrent session PER SCHEDULED TASK; global=3: <=3 concurrent scheduled-task sessions globally (+_pendingTaskDispatches). Host-side SKIP (recordSkipAndEmit/PerTaskLimit|GlobalLimit — NOT queue/deny). Cowork imposes no cap on Task-tool sub-agent fan-out; the harness has no scheduled-task scheduler, so this gate has no applicable surface — pinned as a sync drift-sentinel only."
326
+ },
327
+ "memoryGuidelinesEnv:1696890383": {
328
+ "on": false,
329
+ "source": "defaultValue",
330
+ "value": false
331
+ },
332
+ "oauthScopesEnv:1936081873": {
333
+ "on": true,
334
+ "source": "force",
335
+ "value": true
336
+ },
337
+ "coworkRuntimeConfig:1978029737": {
338
+ "on": true,
339
+ "source": "force",
340
+ "value": {
341
+ "coworkNativeFilePreview": true,
342
+ "coworkWebFetchPrompt": true,
343
+ "coworkWebFetchViaApi": true,
344
+ "sessionsBridgePollBlockMs": 30,
345
+ "workspaceBashWaitLonger": true
346
+ },
347
+ "note": "coworkWebFetchViaApi=true coworkWebFetchPrompt=true workspaceBashWaitLonger=true sessionsBridgePollBlockMs=30 — web_fetch is host/API-routed (POST /api/organizations/<org>/cowork/web_fetch), NOT container egress; gated by a separate web-fetch hostname allowlist + URL provenance."
348
+ },
349
+ "cliPlugin:2307090146": {
350
+ "on": false,
351
+ "source": "defaultValue",
352
+ "value": false,
353
+ "note": "— the CLI-plugin credential broker is dark-launched off for standard interactive accounts (Ch23/L106)."
354
+ },
355
+ "pluginSyncSparkplug:2340532315": {
356
+ "on": true,
357
+ "source": "force",
358
+ "value": true,
359
+ "note": "— startup syncPlugins(); plugins load via --plugin-dir (registry inert in-VM)."
360
+ },
361
+ "skeletonHome:2614807392": {
362
+ "on": false,
363
+ "source": "absent"
364
+ },
365
+ "memoryExtraGuidelines:2860753854": {
366
+ "on": true,
367
+ "source": "defaultValue",
368
+ "value": "## Sensitive personal information\n\nDo not save the following to memory unless the user explicitly asks you to remember it:\n\n- Protected attributes: race, ethnicity, national origin, religion, age, sex, sexual orientation, gender identity, immigration status, disability, serious illness, union membership\n- Government identifiers: Social Security numbers, driver's license numbers, passport numbers, government ID numbers\n- Financial account details: credit card numbers, bank account numbers\n- Health information: medical conditions, diagnoses, lab results, mental health details, therapy or counseling\n- Home or personal mailing addresses (work addresses are fine)\n- Account passwords, secret tokens, or secret keys\n\nIf any of the above appears in conversation context, complete the task but do not persist it to a memory file. If the user explicitly says \"remember my address is X\", saving it is acceptable — they've given consent."
369
+ },
370
+ "skipPrecompactLoad:4153934152": {
371
+ "on": false,
372
+ "source": "defaultValue",
373
+ "value": false
374
+ }
375
+ },
376
+ "eipcChannelUuid": "4f426349-8d6f-45f3-ae22-280fef323564",
377
+ "$comment": "eipcChannelUuid is per-build; recorded for provenance only — the harness does not use Desktop IPC."
378
+ },
379
+ "requireFullVmSandbox": null
380
+ }
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
16
16
 
17
17
  Run it with:
18
18
 
19
- > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.0.2"`. (`replay` itself needs nothing else — no token, no Docker.)
19
+ > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.0.3"`. (`replay` itself needs nothing else — no token, no Docker.)
20
20
 
21
21
  ```sh
22
22
  cowork-harness replay examples/replays/example-pdf-skill.cassette.json
@@ -100,7 +100,7 @@
100
100
  "preRunHashes": {},
101
101
  "scenarioSource": "../../e2e/scenarios/smoke-multiselect.yaml",
102
102
  "fingerprint": {
103
- "baseline": "1.20186.1"
103
+ "baseline": "1.20186.9"
104
104
  },
105
105
  "timeline": [
106
106
  {
@@ -111,7 +111,7 @@
111
111
  },
112
112
  "scenarioSource": "../scenarios/example-pdf-skill.yaml",
113
113
  "fingerprint": {
114
- "baseline": "1.20186.1",
114
+ "baseline": "1.20186.9",
115
115
  "skillHash": "b760ea90682777367fbd44d866ea66d316452da1abb18bf8cb61f1cec8e67806",
116
116
  "contentSig": "ddf68cf6d74e0599729589f8be4f0a9eec440093c99c1060cd6f47dd671a0dfd",
117
117
  "skillSources": [
@@ -78,7 +78,7 @@
78
78
  "preRunHashes": {},
79
79
  "scenarioSource": "../scenarios/hostloop-computer-links.yaml",
80
80
  "fingerprint": {
81
- "baseline": "1.20186.1"
81
+ "baseline": "1.20186.9"
82
82
  },
83
83
  "timeline": [
84
84
  {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cowork-harness",
3
- "version": "1.0.2",
3
+ "version": "1.0.3",
4
4
  "description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios — same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
5
5
  "license": "MIT",
6
6
  "type": "module",