cowork-harness 3.8.0 → 3.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.8.0
7
- tracks-harness: cowork-harness 3.8.0 (baseline desktop-2.7032.0)
6
+ version: 3.8.1
7
+ tracks-harness: cowork-harness 3.8.1 (baseline desktop-2.7032.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.8.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.8.1` (baseline
29
29
  > `desktop-2.7032.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.8.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.8.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.8.0"`. **Pin `@^3.8.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.8.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.8.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.8.1"`. **Pin `@^3.8.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.8.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.8.1"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -77,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
77
77
  GitHub-hosted runners, no token/Docker/agent:
78
78
 
79
79
  ```yaml
80
- - run: npm i -g "cowork-harness@^3.8.0"
80
+ - run: npm i -g "cowork-harness@^3.8.1"
81
81
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
82
82
  # no silent false-greens. WITHOUT --strict this
83
83
  # step cannot fail on a WARN-class rule (e.g.
@@ -364,7 +364,7 @@ jobs:
364
364
  with: { node-version: '24' }
365
365
  - uses: actions/setup-python@v5
366
366
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
367
- - run: npm i -g "cowork-harness@^3.8.0"
367
+ - run: npm i -g "cowork-harness@^3.8.1"
368
368
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
369
369
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
370
370
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -393,7 +393,7 @@ jobs:
393
393
  echo "live=true" >> "$GITHUB_OUTPUT"
394
394
  fi
395
395
  - if: steps.guard.outputs.live == 'true'
396
- run: npm i -g "cowork-harness@^3.8.0"
396
+ run: npm i -g "cowork-harness@^3.8.1"
397
397
  - if: steps.guard.outputs.live == 'true'
398
398
  run: cowork-harness run scenarios/ --output-format json
399
399
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.8.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.8.1`
4
4
  (baseline `desktop-2.7032.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -192,7 +192,7 @@ plugins:
192
192
  skills:
193
193
  local: [] # extra host skill dirs
194
194
  suggest_enabled: true # gate 245679952 override — `mcp__skills__suggest_skills` on/off (default true)
195
- proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = synced baseline gate (ON from 1.24012.11)
195
+ proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = always on from the 1.46388.3 baseline (where `false` models a surface production does not ship there), synced gate before it
196
196
  mcp:
197
197
  config: null # --mcp-config file (standard mcpServers map)
198
198
  enabled: []
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,56 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.8.1] — 2026-09-24
10
+
11
+ ### Upgrade notes
12
+
13
+ - **Cassettes: no re-record needed.** Nothing in this release changes what a run declares or spawns for
14
+ any committed baseline: the proactive-mode fix resolves to the same value on every baseline it touches
15
+ (all from 1.46388.3 on already carry the gate on), and the other edits under `src/runtime`,
16
+ `src/hostloop`, `src/session.ts` and `baselines/` are comments, schema description text and one
17
+ baseline annotation string. `verify-cassettes examples/replays` reports all three committed cassettes
18
+ clean.
19
+ - **Live-validated against `desktop-2.7032.0`**, which 3.8.0 shipped without. On agent 2.1.280, all four
20
+ tiers: `boundary-check` 6/6, e2e self-tests 9/9, `test:live` 19/20 on the first run, and
21
+ `run examples/scenarios/` 6/7. The `test:live` red was model variance in `live-matrix` (the model asked
22
+ in plain text instead of calling `AskUserQuestion`); the file passed on two re-runs. The seventh example,
23
+ `subagent-manifest-probe`, stopped when the operator account hit its usage limit and was not re-run.
24
+ Details in `DESIGN.md`'s scope note.
25
+
26
+ ### Fixed
27
+
28
+ - **Proactive `suggest_skills` mode follows the Desktop version, not a gate Desktop no longer reads.**
29
+ From Desktop 1.46388.3 the asar has no reference to gate `1598976391` (`proactiveSkillSuggestEnabled`):
30
+ `suggest_skills` always carries the proactive description and `trigger` param when it is declared. The
31
+ server still serves the gate, and the harness read it for every baseline, so a server-side flip to off
32
+ would have switched proactive mode off here while production kept it on. For a baseline at 1.46388.3 or
33
+ later the row is now ignored and proactive mode is on; older baselines still read the gate, and
34
+ `skills.proactive_suggest_enabled` still overrides both. No change for any committed baseline: all of
35
+ them from 1.46388.3 on carry the gate on.
36
+
37
+ ### Documentation
38
+
39
+ - **`docs/fidelity-gaps.md` — what the plugin-MCP shadow file carries depends on enforcement.** The page
40
+ said the whole rewritten server set goes into `cowork-plugin-mcp-shadow.json`. Read in the 2.7032.0
41
+ asar, that holds only when Desktop enforces remote shadowing (gate `2529235968` on and no enterprise
42
+ managed-configuration override); otherwise the file names only the remote servers a stand-in replaced,
43
+ policy stubs travel in the in-process SDK server map alone, and a failed write is logged rather than
44
+ refusing the session. The section also states that any failure building the stubs refuses the session
45
+ while an MCP policy is active, that a stub never appears as a `LocalMcpServerManager` connection, and
46
+ which `main.log` lines record the remote arm.
47
+ - **Two pinned gate rows are records, not sentinels.** `canSaveSkill` (`3246569822`) has no reference in
48
+ any Desktop asar from 1.44121.1 on, and `proactiveSkillSuggestEnabled` (`1598976391`) none from
49
+ 1.46388.3 on; the server still serves both, so `sync` keeps recording them. The gates `$comment` in the
50
+ 2.7032.0 baseline, which `sync` carries forward, says so, and a flip of either row changes no run
51
+ against a current baseline.
52
+ - **The unmodeled proactive skills-prompt line applies to every current session.** From Desktop 1.46388.3,
53
+ Desktop's generated `<skills_instructions>` block always carries its proactive suggestion guidance when
54
+ `suggest_skills` and `search_plugins` are available; the harness renders no such block, and
55
+ `docs/fidelity-gaps.md` states the gap at that scope. `skills.proactive_suggest_enabled: false` on such
56
+ a baseline builds a non-proactive `suggest_skills` that production does not ship there —
57
+ `docs/session.md`, the session schema's description and the companion skill's schema reference say so.
58
+
9
59
  ## [3.8.0] — 2026-09-22
10
60
 
11
61
  ### Upgrade notes
package/DESIGN.md CHANGED
@@ -203,7 +203,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
203
203
  > Cowork system-prompt fingerprint all unchanged from `1.20186.0`, with the staged VM ELF re-synced
204
204
  > 2.1.202 → 2.1.205 — and the live pass of that era was deliberately **not** restamped onto it.)
205
205
 
206
- > **Scope of that claim.** `2026-09-20 / desktop-2.2553.1` is the baseline carrying the latest live pass, and it was the newest committed baseline when that pass ran. **One baseline has shipped since (`2.7032.0`), one of which moved the agent ELF most recently to **2.1.280**.** **No live pass has been run against that newer baseline — it was skipped, not blocked.** The agent it pins (2.1.280) IS the staged one, so nothing prevented it; the decision was to ship the sync without re-running the suites. What the baseline rests on instead: a sync with zero unknown deltas, and a prompt-text change derived from the generator source of both asars rather than from a run. The pass below ran against the then-staged agent 2.1.275 (its VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, on harness **3.7.0**. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 tests — 19 passed, 0 failed, 0 skipped**; the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files); and `run examples/scenarios/` **7/7 success** on its first run (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe). **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously. `smoke-multiselect-deciderdir` runs the `--decider-llm` path, and `smoke-l2-microvm` passed in a real VM with its own kernel. Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`**. Both credential paths were exercised: the `.env` OAuth token for the container/protocol tiers, and the agent's own macOS-Keychain self-sourcing at `hostloop`. **One fixture was repaired mid-pass, and the repair is the interesting part.** `smoke-semantic-evidence-files` failed twice with the same reasoning — its deliverable asked the agent to write a factual claim about a migration that never happened, and current models refuse to ship that unflagged. Fixture rot, not a regression: no commit had touched the file since v3.6.0 and its pass rate had already decayed to 56% over 9 runs. The payload now reports what the run itself did. Its falsifiability was re-verified rather than assumed — the unscoped twin still refuses with `evidence_incomplete` naming the dropped files, which is what makes the scoped form's green mean anything. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise — and the one red here was re-run, reproduced, and traced to the fixture rather than to the code. (c) The first `examples/scenarios/` run is the one counted; there was no second. Separately and not a live matter: all three committed cassettes in `examples/replays/` remain those re-recorded against `desktop-2.2553.1` on 2026-09-18; `verify-cassettes` exits 0 with all three clean and one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
206
+ > **Scope of that claim.** `2026-09-24 / desktop-2.7032.0` is the baseline carrying the latest live pass, run against agent **2.1.280**, and it is the newest committed baseline — no baselines have shipped since. The pass ran against the staged agent (its VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, on harness **3.8.1**. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **5 files, 20 tests — 19 passed, 1 failed, 0 skipped** on the first run; the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files); and `run examples/scenarios/` **6/7 success** on its first run (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads). **The two non-greens, stated rather than rounded away.** (1) The `test:live` red is `live-matrix`'s second cell (protocol tier, `desktop-1.18286.0`): the model asked its A-or-B question in plain text instead of calling `AskUserQuestion`, so the run ended on an unanswered question. The file then passed on two consecutive re-runs (2/2 each) — model variance, not a regression; nothing in this release touches the protocol tier or that baseline. (2) `subagent-manifest-probe` (hostloop) did not complete: the operator account's usage limit was reached mid-run, and the agent reported the limit instead of a result. That is a quota stop, not a verdict on the scenario; it was not re-run, and hostloop is still covered by `canary-hostloop` and `hostloop-computer-links`, both green. **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously. `smoke-multiselect-deciderdir` runs the `--decider-llm` path, and `smoke-l2-microvm` passed in a real VM with its own kernel. Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`**. Both credential paths were exercised: the `.env` credentials for the container/protocol tiers, and the agent's own macOS-Keychain self-sourcing at `hostloop`. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (c) The first `examples/scenarios/` run is the one counted; there was no second. Separately and not a live matter: all three committed cassettes in `examples/replays/` are those re-recorded against `desktop-2.7032.0` on 2026-09-23; `verify-cassettes` exits 0 with all three clean and one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
207
207
 
208
208
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
209
209
 
package/README.md CHANGED
@@ -36,7 +36,7 @@ npm ci && npm run build
36
36
  node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
37
37
  ```
38
38
 
39
- (Installing globally — `npm install -g "cowork-harness@^3.8.0"` — gives you the `cowork-harness` CLI for your own
39
+ (Installing globally — `npm install -g "cowork-harness@^3.8.1"` — gives you the `cowork-harness` CLI for your own
40
40
  scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
41
41
 
42
42
  Full setup → [Quick start](./docs/cli.md#quick-start).
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
49
49
 
50
50
  | I want to… | Start here | Needs |
51
51
  |---|---|---|
52
- | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.8.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
- | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.8.0"` |
52
+ | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.8.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
+ | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.8.1"` |
54
54
  | **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
55
55
 
56
56
  Not sure a harness is what you need? The next two sections are the argument.
@@ -348,7 +348,7 @@
348
348
  "asarPath": "/Applications/Claude.app/Contents/Resources/app.asar",
349
349
  "asarFingerprint": "7c4da429c2aae02c",
350
350
  "gates": {
351
- "$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift. The skill-family gates are pinned on the same principle but are NOT all dormant: 245679952 (suggestSkillsEnabled) and 1598976391 (proactiveSkillSuggestEnabled) ARE modeled — they gate the skills SDK-MCP tool surface. 3246569822 (canSaveSkill) is served ON/force by a server-side rollout (independent of Desktop version) and is deliberately NOT modeled: ON adds an mcp__cowork__save_skill tool and changes the skills system prompt, so this pin records a known fidelity gap rather than a modeled surface — see docs/fidelity-gaps.md. 1824824999 (canProposeSkills) is present-but-off, pinned so the same class of silent widening cannot land unnoticed. 1598976391 (proactiveSkillSuggestEnabled) flipped off/defaultValue -> ON/force by a SERVER-SIDE rollout observed 2026-08-04 — NOT a Desktop change: the gate id occurs exactly once in both the 1.24012.9 and 1.24012.11 asars and would read ON on .9 today, so this value is Desktop-version-INDEPENDENT despite living in a version-named file (same provenance class as canSaveSkill above). It is read from a SINGLE account's fcache; force rules are server-evaluated and can be segment-targeted, so whether the rollout is global is not determinable from anything on disk. Unlike canSaveSkill this one IS modeled (both branches exist), so the pin changes the emulated suggest_skills surface: proactive description + an optional trigger param + a chained empty-catalog note. Override per-session with skills.proactive_suggest_enabled. 4074604942 (1p-direct-mcp) was NEW in 1.24012.11 and DARK (absent from a standard fcache, hence its DARK_GATES entry); it arms a Desktop-side direct-MCP pool for MDM-managed 1P servers, inert for an unmanaged account, pinned as a sentinel only. Observed 2026-08-05 SERVED rather than absent (source \"force\", value false) — the rollout reached this account with the gate OFF, so nothing it arms is reachable and no modeled surface changes. Same provenance class as canSaveSkill: read from a SINGLE account's fcache, and force rules are server-evaluated and segment-targetable, so its DARK_GATES entry is retained deliberately — another account may still see it absent, and that must stay tolerated rather than hard-failing their sync.",
351
+ "$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift. The skill-family gates are pinned on the same principle but are NOT all dormant: 245679952 (suggestSkillsEnabled) is modeled — it gates whether suggest_skills is declared. 3246569822 (canSaveSkill) is DEAD in Desktop code from 1.44121.1 (0 occurrences in every asar from 1.44121.1 through this one) although the fcache still serves it ON/force: its row is a record of what the server sends, not a sentinel, and a flip of it changes nothing. The mcp__cowork__save_skill tool it once gated is present in real sessions and is not modeled — see docs/fidelity-gaps.md. 1824824999 (canProposeSkills) is present-but-off, pinned so the same class of silent widening cannot land unnoticed. 1598976391 (proactiveSkillSuggestEnabled) flipped off/defaultValue -> ON/force by a SERVER-SIDE rollout observed 2026-08-04 — NOT a Desktop change: the gate id occurs exactly once in both the 1.24012.9 and 1.24012.11 asars and would read ON on .9 today, so this value is Desktop-version-INDEPENDENT despite living in a version-named file (same provenance class as canSaveSkill above). It is read from a SINGLE account's fcache; force rules are server-evaluated and can be segment-targeted, so whether the rollout is global is not determinable from anything on disk. Like canSaveSkill, this id is DEAD in Desktop code: 0 occurrences in every asar from 1.46388.3 through this one, where the proactive suggest_skills description and the trigger param are unconditional whenever suggest_skills is declared. The harness reads this row only for a baseline older than 1.46388.3; from there proactive mode is always on, as in production, so a flip of this row does not change a run. Override per-session with skills.proactive_suggest_enabled. 4074604942 (1p-direct-mcp) was NEW in 1.24012.11 and DARK (absent from a standard fcache, hence its DARK_GATES entry); it arms a Desktop-side direct-MCP pool for MDM-managed 1P servers, inert for an unmanaged account, pinned as a sentinel only. Observed 2026-08-05 SERVED rather than absent (source \"force\", value false) — the rollout reached this account with the gate OFF, so nothing it arms is reachable and no modeled surface changes. Same provenance class as canSaveSkill: read from a SINGLE account's fcache, and force rules are server-evaluated and segment-targetable, so its DARK_GATES entry is retained deliberately — another account may still see it absent, and that must stay tolerated rather than hard-failing their sync.",
352
352
  "emitToolUseSummaries:66187241": {
353
353
  "on": false,
354
354
  "source": "defaultValue",
@@ -1,3 +1,4 @@
1
+ import { cmpVersionStrings } from "./baseline.js";
1
2
  /**
2
3
  * Read a GrowthBook gate sub-flag (e.g. `coworkWebFetchPrompt`) from the baseline's `provenance.gates`.
3
4
  *
@@ -129,13 +130,33 @@ export function readGateBool(baseline, id) {
129
130
  *
130
131
  * Defaults when the gate is absent from the baseline: `suggestSkills` → true, `proactiveSuggest` → false
131
132
  * (the documented production state).
133
+ *
134
+ * Proactive mode is version-dependent. Up to Desktop 1.44121.1 it is gate `1598976391`. From
135
+ * {@link PROACTIVE_SUGGEST_UNCONDITIONAL_FROM} the gate id is gone from the asar and `suggest_skills` always
136
+ * carries the proactive description and `trigger` param whenever it is declared — so for those baselines
137
+ * the gate row is IGNORED even though the fcache still serves it. Reading it there would let a server-side
138
+ * flip of a gate Desktop no longer reads turn the harness's proactive mode off while production keeps it on.
139
+ * The session knob still wins either way.
132
140
  */
133
141
  export function resolveSkillDiscoveryGates(baseline, knobs = {}) {
134
142
  return {
135
143
  suggestSkillsEnabled: knobs.suggest_enabled ?? readGateBool(baseline, "245679952") ?? true,
136
- proactiveSkillSuggestEnabled: knobs.proactive_suggest_enabled ?? readGateBool(baseline, "1598976391") ?? false,
144
+ proactiveSkillSuggestEnabled: knobs.proactive_suggest_enabled ?? proactiveFromBaseline(baseline),
137
145
  };
138
146
  }
147
+ /** First backed-up Desktop build whose asar has no reference to gate `1598976391` (0 occurrences from
148
+ * here through 2.7032.0; 2 in 1.44121.1, the previous backup — builds in between are unobserved, and a
149
+ * baseline for one would take the gate path, the conservative side). From this build proactive suggest
150
+ * mode is unconditional. */
151
+ export const PROACTIVE_SUGGEST_UNCONDITIONAL_FROM = "1.46388.3";
152
+ function proactiveFromBaseline(baseline) {
153
+ // `appVersion` is required by the schema, but synthetic test baselines omit it; a missing version
154
+ // compares as 0.0.0 and takes the gate path, which is the pre-1.46388.3 behaviour.
155
+ const v = baseline.appVersion;
156
+ if (typeof v === "string" && cmpVersionStrings(v, PROACTIVE_SUGGEST_UNCONDITIONAL_FROM) >= 0)
157
+ return true;
158
+ return readGateBool(baseline, "1598976391") ?? false;
159
+ }
139
160
  export function decideLoop(inputs) {
140
161
  if (inputs.requireFullVmSandbox === true)
141
162
  return "vm"; // HeA()
package/dist/session.js CHANGED
@@ -142,9 +142,10 @@ export const SessionConfig = z.strictObject({
142
142
  .optional()
143
143
  .describe("override for gate 1598976391 (proactiveSkillSuggestEnabled): when true (and suggest_enabled is not " +
144
144
  "false), suggest_skills swaps to the proactive description, gains an optional `trigger` enum param, " +
145
- "and chains its empty-catalog note into search_plugins. Omit to use the synced baseline gate — on " +
146
- "from the 1.24012.11 baseline, matching what production serves; the fallback for a baseline that " +
147
- "predates the gate is false."),
145
+ "and chains its empty-catalog note into search_plugins. Omit to follow the baseline: from the 1.46388.3 " +
146
+ "baseline Desktop does not read the gate and proactive mode is always on; for an older baseline the " +
147
+ "synced gate decides (on from 1.24012.11; false for a baseline that predates the gate). On a 1.46388.3+ " +
148
+ "baseline, false builds a non-proactive surface production does not ship there."),
148
149
  })
149
150
  .default({ local: [] }),
150
151
  mcp: z
@@ -108,24 +108,38 @@ export const PINNED_GATES = {
108
108
  // init.tools of 8 real sessions) render, and in what mode. None was pinned before, so 245679952
109
109
  // being live on/force was invisible to the drift guard. Present in the live fcache (NOT dark), so
110
110
  // they are read at their real state — no DARK_GATES entry. BEHAVIORALLY MODELED since A2: the harness
111
- // declares the skills/plugins SDK-MCP servers and reads BOTH gates at spawn
112
- // (`resolveSkillDiscoveryGates`), so a flip here CHANGES the declared tool set on container/hostloop
113
- // (see `src/hostloop/skills-handler.ts`) — it is NOT inert. A pinned drift alone WARNS + still writes.
111
+ // declares the skills/plugins SDK-MCP servers and reads 245679952 at spawn (`resolveSkillDiscoveryGates`),
112
+ // so a flip of it CHANGES the declared tool set on container/hostloop (see
113
+ // `src/hostloop/skills-handler.ts`) — it is NOT inert. 1598976391 is read only for a baseline older than
114
+ // 1.46388.3 (see its entry below). A pinned drift alone WARNS + still writes.
114
115
  "245679952": "suggestSkillsEnabled", // live on/force — gates whether suggest_skills renders at all
115
- // proactive (unprompted) suggest mode. (A prior note speculated this widens at agent >=2.1.217 to gate
116
- // the whole discovery-tool family; REFUTED — 2.1.205-2.1.217 sessions carry the full skills family with
117
- // this gate OFF.) It has THREE effects, not two: it swaps suggest_skills's description, adds `trigger`,
118
- // AND is passed into generateSkillsSystemPrompt, where it swaps the suggest-guidance line inside the
119
- // dynamically-generated `<skills_instructions>` block (plus a once-per-conversation sentence). A prior
120
- // version of this comment claimed "only swaps the description and adds `trigger`" — that was wrong about
121
- // the product. The harness models the first two and renders no `<skills_instructions>` section at all,
122
- // so the prompt effect is a disclosed gap, not a modelled surface (same shape as canSaveSkill below).
116
+ // proactive (unprompted) suggest mode. DEAD IN DESKTOP CODE FROM 1.46388.3: 2 occurrences in the 1.40609.1
117
+ // and 1.44121.1 asars, 0 in every asar from 1.46388.3 through 2.7032.0, and the `proactiveSkillSuggestEnabled`
118
+ // name is gone too. From that build Desktop gives suggest_skills its proactive description and `trigger`
119
+ // param unconditionally whenever the tool is declared. Up to 1.44121.1 the gate was live, with three effects:
120
+ // the description swap, `trigger`, and a swapped suggest-guidance line in the generated
121
+ // `<skills_instructions>` block (a prompt effect the harness never models).
122
+ // The harness reads this row at spawn ONLY for a baseline older than 1.46388.3; from there
123
+ // `resolveSkillDiscoveryGates` ignores it and proactive mode is always on, as in production. So a
124
+ // server-side flip here is a sync diff with no effect on runs against a current baseline. Like
125
+ // canSaveSkill below, the row is a record of what the server sends, not a sentinel.
123
126
  "1598976391": "proactiveSkillSuggestEnabled",
124
- // Flipped off/defaultValue -> ON/force server-side (fcache) as of 2026-07-25, i.e. for current users on
125
- // any Desktop version — NOT a Desktop change; the machinery already shipped in 1.24012.1 gated off. ON
126
- // adds a `save_skill` tool to the session's SDK-MCP inventory AND is passed into
127
- // generateSkillsSystemPrompt, so it changes both the tool set and the skills prompt. The harness models
128
- // NEITHER yet, so this is a known fidelity gap, not a modelled surface.
127
+ // DEAD IN DESKTOP CODE SINCE 1.44121.1 — THIS PIN PROVIDES NO COVERAGE OF `save_skill`. The id has 3
128
+ // occurrences in the 1.40609.1 asar and 0 in every asar from 1.44121.1 through 2.7032.0 (count with
129
+ // `grep -ao <id> app.asar | wc -l`; `grep -c` counts lines of a minified file and undercounts). The
130
+ // fcache still serves it on/force, so sync round-trips it and reports "no change": a green light on a
131
+ // disconnected sensor. Up to 1.40609.1 the gate was live, and its off -> on flip on 2026-07-25 is what
132
+ // put `mcp__cowork__save_skill` into real sessions. The row is kept only as a record of what the server
133
+ // sends: baselines from 1.24012.1 on carry it (from 1.44121.1 on with a `note` saying it is unread), and
134
+ // docs/fidelity-gaps.md names it. Do not read it as a tripwire.
135
+ // At 2.7032.0 the tool follows `skillsEnabled`, the managed `skillCreationEnabled` setting and two gates
136
+ // that first appear in that build: 3656976882 (org skills off) and 3469616823 (bypass for the org skill-
137
+ // creation block). The org check reads a `skill_creation` status from an account access list Desktop
138
+ // fetches at runtime (`/api/bootstrap/<org>/current_user_access`, held in memory, absent from the fcache)
139
+ // and can only turn the tool OFF — an unloaded list allows it. Pinning the two newer gates would NOT make
140
+ // this a sentinel: the access list sits outside the fcache. The observable outcome is whether a REAL
141
+ // Desktop session's init tool list carries `mcp__cowork__save_skill` — never a harness run's, which lists
142
+ // only what the harness itself declares and does not declare `save_skill` at any tier.
129
143
  "3246569822": "canSaveSkill",
130
144
  // off/defaultValue and PRESENT in the fcache (so NOT dark — no DARK_GATES entry) — the `propose_skills`
131
145
  // render-only sibling. Pinned so a production flip surfaces as a sync diff instead of silently widening
package/docs/ci.md CHANGED
@@ -107,7 +107,7 @@ jobs:
107
107
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
108
108
  ```
109
109
 
110
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.8.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
110
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.8.1` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
111
111
 
112
112
  CI uses `ANTHROPIC_API_KEY` specifically because there's no interactive browser available to run
113
113
  `claude setup-token`'s OAuth flow in a GitHub Actions runner; locally, the OAuth token is preferred because
package/docs/cli.md CHANGED
@@ -18,7 +18,7 @@ companion skill, CI). This page is the CLI one.
18
18
  **Install from npm:**
19
19
 
20
20
  ```bash
21
- npm install -g "cowork-harness@^3.8.0" # puts the `cowork-harness` command on your PATH
21
+ npm install -g "cowork-harness@^3.8.1" # puts the `cowork-harness` command on your PATH
22
22
  ```
23
23
 
24
24
  **Or build from source:**
@@ -38,7 +38,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
38
38
 
39
39
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
40
40
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
41
- > From a global install (`npm i -g "cowork-harness@^3.8.0"`), point at the package root instead:
41
+ > From a global install (`npm i -g "cowork-harness@^3.8.1"`), point at the package root instead:
42
42
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
43
43
  > (or copy the cassette into your own project and pass that path).
44
44
 
@@ -48,7 +48,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
48
48
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
49
49
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
50
50
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
51
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.8.0"`.
51
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.8.1"`.
52
52
 
53
53
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
54
54
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -124,7 +124,7 @@ The parts people ask about. **`package.json`'s `files[]` is the exhaustive, mach
124
124
  this table is the readable summary of it, and deliberately omits the infrastructure that always ships
125
125
  (`baselines/`, `schema/`, `fixtures/`, `scripts/`, `docker/`).
126
126
 
127
- | What ships | npm global (`npm install -g "cowork-harness@^3.8.0"`) | Source checkout (`git clone` + `npm ci`) |
127
+ | What ships | npm global (`npm install -g "cowork-harness@^3.8.1"`) | Source checkout (`git clone` + `npm ci`) |
128
128
  |---|---|---|
129
129
  | CLI, `scenario.py` + assertion keys (enough for `lint` in CI) | ✓ | ✓ |
130
130
  | `SKILL.md`, all of `docs/`, `SPEC.md`/`DESIGN.md`/`AGENTS.md` | ✓ | ✓ |
@@ -139,7 +139,7 @@ since a global install puts nothing in your working directory. The matrix, answe
139
139
  examples are the only ones that still need a source checkout. The **marketplace skill install** is
140
140
  narrower again — it pulls only `.claude/skills/cowork-harness/` (SKILL.md + `references/` +
141
141
  `scenario.py`/assertion keys, per `.claude-plugin/marketplace.json`'s `source`); everything in the npm column
142
- arrives when the skill's first command self-bootstraps `npx "cowork-harness@^3.8.0"` — the last row stays
142
+ arrives when the skill's first command self-bootstraps `npx "cowork-harness@^3.8.1"` — the last row stays
143
143
  ✗ either way, since `matrices/`, `answer-policies/` and `probes/` are not published at all. See
144
144
  [docs/companion-skill.md](./companion-skill.md) for that install path.
145
145
 
@@ -28,7 +28,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
28
28
  claude plugin install cowork-harness@cowork-harness
29
29
  ```
30
30
 
31
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.8.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
31
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.8.1"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
32
32
 
33
33
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
34
34
 
@@ -41,7 +41,7 @@ npx skills add yaniv-golan/cowork-harness --skill cowork-harness
41
41
  **What the marketplace install actually pulls:** only `.claude/skills/cowork-harness/` — SKILL.md +
42
42
  `references/` + `scenario.py`/assertion keys, per `.claude-plugin/marketplace.json`'s `source`. Everything
43
43
  else (the CLI, `docs/`, the worked examples, the pytest lane) arrives when the skill's first command
44
- self-bootstraps `npx "cowork-harness@^3.8.0"`, which pulls the same npm package as a global install.
44
+ self-bootstraps `npx "cowork-harness@^3.8.1"`, which pulls the same npm package as a global install.
45
45
 
46
46
  For the full package-contents table — what a global install gives you versus a source checkout — see
47
47
  [docs/cli.md → What ships](./cli.md#what-ships). It is maintained there, once.
@@ -327,7 +327,7 @@ ordinary session**:
327
327
  session shadows remote servers at all (gate `2529235968`, or at least one third-party direct MCP server
328
328
  present) **and** a stand-in already provides that server, matched by URL hostname or by name. The
329
329
  stand-ins are the session's claude.ai connectors that carry at least one enabled tool, plus those direct
330
- servers. Logged *"Plugin `<name>` declares remote MCP servers (…). Overriding with no-ops so the CLI does
330
+ servers. Logged *"Plugin `<id>` declares remote MCP servers (…). Overriding with no-ops so the CLI does
331
331
  not open its own client."* With the gate off and no such server, the remote arm never runs, and a
332
332
  plugin's remote servers reach the CLI intact. Earlier builds (1.37937.0 through 1.46388.x) stubbed
333
333
  **every** remote plugin server once the gate was on; the stand-in narrowing arrives with 2.2553.1.
@@ -339,9 +339,28 @@ A plugin Desktop treats as official is exempt from both arms unless a policy is
339
339
 
340
340
  A replaced server is renamed `plugin:<pluginName>:<serverName>` and constructed as
341
341
  `createSdkMcpServer({name, tools: []})` — **the server name is present in the session's inventory and offers
342
- zero tools.** The rewritten set is written to `cowork-plugin-mcp-shadow.json` and delivered to the agent as
343
- an `--mcp-config` payload; when the session shadows remote servers *and* an MCP policy is active, a failed
344
- delivery makes Desktop refuse to start the session rather than launch with unenforced plugin servers.
342
+ zero tools.** Every stub reaches the agent in Desktop's in-process SDK server map. Which stubs are _also_
343
+ named in `cowork-plugin-mcp-shadow.json`, delivered as an `--mcp-config` payload, depends on whether
344
+ Desktop enforces remote shadowing for the session: gate `2529235968` on **and** no enterprise
345
+ managed-configuration override in force. When it does, the file names every replaced server, policy stubs
346
+ included. When it does not, the file names only the remote servers a stand-in replaced, and a session with
347
+ none writes no file at all; local and `.mcpb` policy stubs then travel in the SDK map alone.
348
+
349
+ While an MCP policy is active, Desktop refuses to start the session rather than launch with unenforced
350
+ plugin servers whenever building the overrides fails, whatever the gate: a failed plugin scan (*"Plugin MCP
351
+ scan failed while an MCP policy is active"*) or any other error in that step. It also refuses when the
352
+ shadow file cannot be written, but only while it enforces remote shadowing; otherwise a failed write is
353
+ logged and the spawn goes ahead without the file.
354
+
355
+ Desktop's own MCP pool, `LocalMcpServerManager`, takes only a plugin's local stdio and `.mcpb` servers, and
356
+ skips any a policy blocks. Remote plugin servers that no stand-in replaced are left for the CLI to open
357
+ itself. **A stub therefore never appears as a `LocalMcpServerManager` connection**, which makes Desktop's
358
+ `main.log` an unambiguous record of the remote arm: each replacement logs
359
+ `Replacing plugin "<plugin>" MCP server "<server>": "<connector>" already provides it (matched by url)`,
360
+ then `Plugin "<id>" declares remote MCP servers (…). Overriding with no-ops …`. On a Desktop whose
361
+ connected Slack, Notion or Airtable connector matches a plugin's declared server by URL, both lines appear
362
+ for every session that loads the plugin, and the plugin servers that connect through
363
+ `LocalMcpServerManager` are only local stdio ones.
345
364
 
346
365
  **What the harness does:** nothing. Plugins are staged with `--plugin-dir` and the CLI reads each plugin's
347
366
  own declaration, so a plugin under test gets **working** MCP servers with their real tools.
@@ -852,19 +871,23 @@ over the control protocol (`sdkMcpServers` in `initialize`, tunneled as `mcp_mes
852
871
  `list_skills` returns no match, and the result renders an "Add" card. The call has **no side effect**
853
872
  (nothing installs; the user's Add click happens out of band). Ground truth: these tools appear in the
854
873
  `system/init` `tools` array of real on-disk sessions (`local-agent-mode-sessions/**/audit.jsonl`); the
855
- `suggestSkillsEnabled` gate `245679952` is on. A proactive-suggestion mode sits behind a second gate
856
- `1598976391` (`proactiveSkillSuggestEnabled`), which is **served ON** for a standard account as of the
857
- `1.24012.11` baseline — by a server-side rollout, not a Desktop change (the gate reads ON on earlier
858
- Desktop versions too). With it off, `suggest_skills` keeps its base description and the model suggests
859
- only when the conversation invites it. With it on, the tool gains an optional `trigger` parameter
874
+ `suggestSkillsEnabled` gate `245679952` is on. `suggest_skills` also has a proactive-suggestion mode.
875
+ From Desktop `1.46388.3` (the first backed-up build after `1.44121.1`) that mode is **unconditional** whenever `suggest_skills` is declared: no gate
876
+ is read for it. Up to `1.44121.1` it sits behind gate `1598976391`
877
+ (`proactiveSkillSuggestEnabled`), which the server serves **on** for a standard account from the
878
+ `1.24012.11` baseline; the server still serves that gate, but Desktop from `1.46388.3` never reads it.
879
+ With the mode off, `suggest_skills` keeps its base description and the model suggests only when the
880
+ conversation invites it. With it on, the tool gains an optional `trigger` parameter
860
881
  (`user_asked` | `proactive`), a proactive description that also carries production's *constraints* (a
861
882
  do-not-call list, a suggest-at-most-once-per-conversation rule, a no-lead-in rule, and forwarding the
862
883
  same keywords **and trigger** to `search_plugins`), and an empty-catalog `note` that chains into
863
884
  `search_plugins` for every trigger state — silence is only the `proactive` tail, and a trigger the model
864
- never supplied is never forwarded back to it. One production effect is **not** modeled: the flag is also
865
- passed into Desktop's `generateSkillsSystemPrompt`, where it swaps a guidance line inside the generated
866
- `<skills_instructions>` block. The harness renders no such section at all, so that effect lands in an
867
- already-unmodeled surface.
885
+ never supplied is never forwarded back to it. One production effect is **not** modeled: a proactive
886
+ suggest-guidance line (with a suggest-at-most-once-per-conversation sentence) inside the generated
887
+ `<skills_instructions>` block. Up to `1.44121.1` the gate selects it; from `1.46388.3` Desktop emits it
888
+ whenever `suggest_skills` and `search_plugins` are available, so on a current baseline this gap applies to
889
+ every session that declares `suggest_skills`. The harness renders no such section at all, so the effect
890
+ lands in an already-unmodeled surface.
868
891
 
869
892
  **Harness behaviour:** `container` and `hostloop` (and `cowork`, which resolves to one of those) now
870
893
  declare a `skills` and a `plugins` SDK-MCP server alongside `cowork`/`workspace` (`combineSdkMcp`,
@@ -878,9 +901,11 @@ synced baseline (`readGateBool`, bare-boolean shape — distinct from the sub-fl
878
901
  reads) with a session-level override (`skills.suggest_enabled` / `skills.proactive_suggest_enabled`, see
879
902
  [session.md](./session.md)). Precedence is knob ▸ baseline gate ▸ hardcoded fallback, and the three are
880
903
  distinct: omit the knob and the value comes from the **synced baseline** (on `latest` that is
881
- `suggestSkillsEnabled` on and `proactiveSkillSuggestEnabled` **on**, mirroring what production serves);
882
- the hardcoded fallback, which applies only to a baseline old enough to predate the gate entirely, stays
883
- on for `suggestSkillsEnabled` and **off** for `proactiveSkillSuggestEnabled`.
904
+ `suggestSkillsEnabled` on); the hardcoded fallback, which applies only to a baseline old enough to predate
905
+ the gate entirely, is on. Proactive mode follows the Desktop version the same way production does: from
906
+ the `1.46388.3` baseline it is **always on** and the `1598976391` row is ignored, so a server-side flip of
907
+ a gate Desktop does not read cannot change a run; for an older baseline the synced gate decides, with a
908
+ hardcoded fallback of **off** for a baseline that predates it.
884
909
 
885
910
  **How exact is the model?** Not uniformly — and the difference matters, so it is stated plainly. The
886
911
  tool **inventory** (which five tools exist), their **inputSchemas**, the **gating** semantics, and the
package/docs/session.md CHANGED
@@ -95,11 +95,11 @@ plugins:
95
95
  skills:
96
96
  local: [] # extra host skill dirs → CLAUDE_CONFIG_DIR/skills
97
97
  # Overrides for the skill-discovery SDK-MCP gates (container/hostloop only — see fidelity-gaps.md
98
- # "Skill/plugin discovery SDK-MCP servers"). Precedence: this knob → the synced baseline gate → the
99
- # documented default. Both omit-able; the harness resolves gate 245679952/1598976391 from the baseline
100
- # when unset.
98
+ # "Skill/plugin discovery SDK-MCP servers"). Precedence: this knob → the baseline → the documented
99
+ # default. Both omit-able; unset, suggest_enabled follows gate 245679952, and proactive_suggest_enabled is
100
+ # always on from the 1.46388.3 baseline and follows gate 1598976391 before it.
101
101
  suggest_enabled: true # gate 245679952 (suggestSkillsEnabled) override; default true when unset
102
- proactive_suggest_enabled: false # gate 1598976391 (proactiveSkillSuggestEnabled) override; unset = the synced baseline gate (ON from 1.24012.11)
102
+ proactive_suggest_enabled: false # gate 1598976391 (proactiveSkillSuggestEnabled) override; unset = always on from the 1.46388.3 baseline, the synced gate before it
103
103
  mcp:
104
104
  config: null # --mcp-config file (standard mcpServers map), e.g. ../data/mcp.json
105
105
  enabled: [] # enabledMcpjsonServers
@@ -229,7 +229,7 @@ See [discovery.md](./discovery.md) for the full model. In short: the harness bui
229
229
  | `plugins.local_plugins[]` / `remote_plugins[]` | Cowork plugin mounts | → `mnt/.local-plugins/marketplaces/<marketplace>/<plugin>` (≥1.14271.0; older baselines use `.local-plugins/cache`) / `mnt/.remote-plugins/plugin_<id>` (migrated-Cowork uploaded/org-remote shape; the id is a stable hash of the declared source). A skill that references these via `${CLAUDE_PLUGIN_ROOT}` must mind [the two-namespace resolution model](./plugin-root.md) — the token is unset in host-loop VM bash. |
230
230
  | `skills.local[]` | `CLAUDE_CONFIG_DIR/skills` | extra host **skill** dirs (a folder *without* `.claude-plugin/plugin.json`) staged into the config dir's `skills/`. Use this for a single-skill folder; use `plugins.local_plugins` for a plugin root. |
231
231
  | `skills.suggest_enabled` | gate `245679952` (`suggestSkillsEnabled`) override | `container`/`hostloop` (and `cowork`) only. `true` (or unset, if the synced baseline gate is on/absent) declares the `skills` SDK-MCP server's `suggest_skills` tool; `false` omits it and drops `list_skills`' fallback-to-`suggest_skills` clause. See [fidelity-gaps.md](./fidelity-gaps.md). |
232
- | `skills.proactive_suggest_enabled` | gate `1598976391` (`proactiveSkillSuggestEnabled`) override | Only consulted when `suggest_enabled` (effective) is true. `true` swaps `suggest_skills` to the proactive description, adds an optional `trigger` enum param (`user_asked` \| `proactive`), and chains the empty-catalog `note` into `search_plugins`. Omit to use the synced baseline gate — **on** from the `1.24012.11` baseline, which is what production serves; the fallback for a baseline predating the gate is `false`. |
232
+ | `skills.proactive_suggest_enabled` | gate `1598976391` (`proactiveSkillSuggestEnabled`) override | Only consulted when `suggest_enabled` (effective) is true. `true` swaps `suggest_skills` to the proactive description, adds an optional `trigger` enum param (`user_asked` \| `proactive`), and chains the empty-catalog `note` into `search_plugins`. Omit to follow the baseline: from the `1.46388.3` baseline Desktop does not read the gate and proactive mode is **always on**, so the gate row is ignored; for an older baseline the synced gate decides (**on** from `1.24012.11`; `false` for a baseline predating the gate). On a baseline from `1.46388.3`, `false` builds a non-proactive `suggest_skills` that production does not ship there; use it only to exercise the older surface. |
233
233
  | `mcp.config` / `mcp.enabled[]` | `--mcp-config` / `enabledMcpjsonServers` | the supported way to attach an MCP server to a session under test. |
234
234
 
235
235
  > Inside a git repo, `folders[]` and `skills.local[]` stage only **git-tracked** files into the mount (matching
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
16
16
 
17
17
  Run it with:
18
18
 
19
- > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.8.0"`. (`replay` itself needs nothing else — no token, no Docker.)
19
+ > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.8.1"`. (`replay` itself needs nothing else — no token, no Docker.)
20
20
 
21
21
  ```sh
22
22
  cowork-harness replay examples/replays/example-pdf-skill.cassette.json
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cowork-harness",
3
- "version": "3.8.0",
3
+ "version": "3.8.1",
4
4
  "description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -200,7 +200,7 @@
200
200
  "type": "boolean"
201
201
  },
202
202
  "proactive_suggest_enabled": {
203
- "description": "override for gate 1598976391 (proactiveSkillSuggestEnabled): when true (and suggest_enabled is not false), suggest_skills swaps to the proactive description, gains an optional `trigger` enum param, and chains its empty-catalog note into search_plugins. Omit to use the synced baseline gate — on from the 1.24012.11 baseline, matching what production serves; the fallback for a baseline that predates the gate is false.",
203
+ "description": "override for gate 1598976391 (proactiveSkillSuggestEnabled): when true (and suggest_enabled is not false), suggest_skills swaps to the proactive description, gains an optional `trigger` enum param, and chains its empty-catalog note into search_plugins. Omit to follow the baseline: from the 1.46388.3 baseline Desktop does not read the gate and proactive mode is always on; for an older baseline the synced gate decides (on from 1.24012.11; false for a baseline that predates the gate). On a 1.46388.3+ baseline, false builds a non-proactive surface production does not ship there.",
204
204
  "type": "boolean"
205
205
  }
206
206
  },