cowork-harness 3.8.0 → 3.8.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +4 -4
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +50 -0
- package/DESIGN.md +1 -1
- package/README.md +3 -3
- package/baselines/desktop-2.7032.0.json +1 -1
- package/dist/loop-decision.js +22 -1
- package/dist/session.js +4 -3
- package/dist/sync/cowork-sync.js +30 -16
- package/docs/ci.md +1 -1
- package/docs/cli.md +5 -5
- package/docs/companion-skill.md +2 -2
- package/docs/fidelity-gaps.md +41 -16
- package/docs/session.md +5 -5
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
- package/schema/session.schema.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.8.
|
|
7
|
-
tracks-harness: cowork-harness 3.8.
|
|
6
|
+
version: 3.8.1
|
|
7
|
+
tracks-harness: cowork-harness 3.8.1 (baseline desktop-2.7032.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.8.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.8.1` (baseline
|
|
29
29
|
> `desktop-2.7032.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.8.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.8.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.8.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.8.1"`. **Pin `@^3.8.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.8.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.8.
|
|
20
|
+
(e.g. `version: "3.8.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -77,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
77
77
|
GitHub-hosted runners, no token/Docker/agent:
|
|
78
78
|
|
|
79
79
|
```yaml
|
|
80
|
-
- run: npm i -g "cowork-harness@^3.8.
|
|
80
|
+
- run: npm i -g "cowork-harness@^3.8.1"
|
|
81
81
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
82
82
|
# no silent false-greens. WITHOUT --strict this
|
|
83
83
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -364,7 +364,7 @@ jobs:
|
|
|
364
364
|
with: { node-version: '24' }
|
|
365
365
|
- uses: actions/setup-python@v5
|
|
366
366
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
367
|
-
- run: npm i -g "cowork-harness@^3.8.
|
|
367
|
+
- run: npm i -g "cowork-harness@^3.8.1"
|
|
368
368
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
369
369
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
370
370
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -393,7 +393,7 @@ jobs:
|
|
|
393
393
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
394
394
|
fi
|
|
395
395
|
- if: steps.guard.outputs.live == 'true'
|
|
396
|
-
run: npm i -g "cowork-harness@^3.8.
|
|
396
|
+
run: npm i -g "cowork-harness@^3.8.1"
|
|
397
397
|
- if: steps.guard.outputs.live == 'true'
|
|
398
398
|
run: cowork-harness run scenarios/ --output-format json
|
|
399
399
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.8.
|
|
3
|
+
Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.8.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.8.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.8.1`
|
|
4
4
|
(baseline `desktop-2.7032.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -192,7 +192,7 @@ plugins:
|
|
|
192
192
|
skills:
|
|
193
193
|
local: [] # extra host skill dirs
|
|
194
194
|
suggest_enabled: true # gate 245679952 override — `mcp__skills__suggest_skills` on/off (default true)
|
|
195
|
-
proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset =
|
|
195
|
+
proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = always on from the 1.46388.3 baseline (where `false` models a surface production does not ship there), synced gate before it
|
|
196
196
|
mcp:
|
|
197
197
|
config: null # --mcp-config file (standard mcpServers map)
|
|
198
198
|
enabled: []
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.8.
|
|
5
|
+
Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,56 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.8.1] — 2026-09-24
|
|
10
|
+
|
|
11
|
+
### Upgrade notes
|
|
12
|
+
|
|
13
|
+
- **Cassettes: no re-record needed.** Nothing in this release changes what a run declares or spawns for
|
|
14
|
+
any committed baseline: the proactive-mode fix resolves to the same value on every baseline it touches
|
|
15
|
+
(all from 1.46388.3 on already carry the gate on), and the other edits under `src/runtime`,
|
|
16
|
+
`src/hostloop`, `src/session.ts` and `baselines/` are comments, schema description text and one
|
|
17
|
+
baseline annotation string. `verify-cassettes examples/replays` reports all three committed cassettes
|
|
18
|
+
clean.
|
|
19
|
+
- **Live-validated against `desktop-2.7032.0`**, which 3.8.0 shipped without. On agent 2.1.280, all four
|
|
20
|
+
tiers: `boundary-check` 6/6, e2e self-tests 9/9, `test:live` 19/20 on the first run, and
|
|
21
|
+
`run examples/scenarios/` 6/7. The `test:live` red was model variance in `live-matrix` (the model asked
|
|
22
|
+
in plain text instead of calling `AskUserQuestion`); the file passed on two re-runs. The seventh example,
|
|
23
|
+
`subagent-manifest-probe`, stopped when the operator account hit its usage limit and was not re-run.
|
|
24
|
+
Details in `DESIGN.md`'s scope note.
|
|
25
|
+
|
|
26
|
+
### Fixed
|
|
27
|
+
|
|
28
|
+
- **Proactive `suggest_skills` mode follows the Desktop version, not a gate Desktop no longer reads.**
|
|
29
|
+
From Desktop 1.46388.3 the asar has no reference to gate `1598976391` (`proactiveSkillSuggestEnabled`):
|
|
30
|
+
`suggest_skills` always carries the proactive description and `trigger` param when it is declared. The
|
|
31
|
+
server still serves the gate, and the harness read it for every baseline, so a server-side flip to off
|
|
32
|
+
would have switched proactive mode off here while production kept it on. For a baseline at 1.46388.3 or
|
|
33
|
+
later the row is now ignored and proactive mode is on; older baselines still read the gate, and
|
|
34
|
+
`skills.proactive_suggest_enabled` still overrides both. No change for any committed baseline: all of
|
|
35
|
+
them from 1.46388.3 on carry the gate on.
|
|
36
|
+
|
|
37
|
+
### Documentation
|
|
38
|
+
|
|
39
|
+
- **`docs/fidelity-gaps.md` — what the plugin-MCP shadow file carries depends on enforcement.** The page
|
|
40
|
+
said the whole rewritten server set goes into `cowork-plugin-mcp-shadow.json`. Read in the 2.7032.0
|
|
41
|
+
asar, that holds only when Desktop enforces remote shadowing (gate `2529235968` on and no enterprise
|
|
42
|
+
managed-configuration override); otherwise the file names only the remote servers a stand-in replaced,
|
|
43
|
+
policy stubs travel in the in-process SDK server map alone, and a failed write is logged rather than
|
|
44
|
+
refusing the session. The section also states that any failure building the stubs refuses the session
|
|
45
|
+
while an MCP policy is active, that a stub never appears as a `LocalMcpServerManager` connection, and
|
|
46
|
+
which `main.log` lines record the remote arm.
|
|
47
|
+
- **Two pinned gate rows are records, not sentinels.** `canSaveSkill` (`3246569822`) has no reference in
|
|
48
|
+
any Desktop asar from 1.44121.1 on, and `proactiveSkillSuggestEnabled` (`1598976391`) none from
|
|
49
|
+
1.46388.3 on; the server still serves both, so `sync` keeps recording them. The gates `$comment` in the
|
|
50
|
+
2.7032.0 baseline, which `sync` carries forward, says so, and a flip of either row changes no run
|
|
51
|
+
against a current baseline.
|
|
52
|
+
- **The unmodeled proactive skills-prompt line applies to every current session.** From Desktop 1.46388.3,
|
|
53
|
+
Desktop's generated `<skills_instructions>` block always carries its proactive suggestion guidance when
|
|
54
|
+
`suggest_skills` and `search_plugins` are available; the harness renders no such block, and
|
|
55
|
+
`docs/fidelity-gaps.md` states the gap at that scope. `skills.proactive_suggest_enabled: false` on such
|
|
56
|
+
a baseline builds a non-proactive `suggest_skills` that production does not ship there —
|
|
57
|
+
`docs/session.md`, the session schema's description and the companion skill's schema reference say so.
|
|
58
|
+
|
|
9
59
|
## [3.8.0] — 2026-09-22
|
|
10
60
|
|
|
11
61
|
### Upgrade notes
|
package/DESIGN.md
CHANGED
|
@@ -203,7 +203,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
203
203
|
> Cowork system-prompt fingerprint all unchanged from `1.20186.0`, with the staged VM ELF re-synced
|
|
204
204
|
> 2.1.202 → 2.1.205 — and the live pass of that era was deliberately **not** restamped onto it.)
|
|
205
205
|
|
|
206
|
-
> **Scope of that claim.** `2026-09-
|
|
206
|
+
> **Scope of that claim.** `2026-09-24 / desktop-2.7032.0` is the baseline carrying the latest live pass, run against agent **2.1.280**, and it is the newest committed baseline — no baselines have shipped since. The pass ran against the staged agent (its VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, on harness **3.8.1**. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **5 files, 20 tests — 19 passed, 1 failed, 0 skipped** on the first run; the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files); and `run examples/scenarios/` **6/7 success** on its first run (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads). **The two non-greens, stated rather than rounded away.** (1) The `test:live` red is `live-matrix`'s second cell (protocol tier, `desktop-1.18286.0`): the model asked its A-or-B question in plain text instead of calling `AskUserQuestion`, so the run ended on an unanswered question. The file then passed on two consecutive re-runs (2/2 each) — model variance, not a regression; nothing in this release touches the protocol tier or that baseline. (2) `subagent-manifest-probe` (hostloop) did not complete: the operator account's usage limit was reached mid-run, and the agent reported the limit instead of a result. That is a quota stop, not a verdict on the scenario; it was not re-run, and hostloop is still covered by `canary-hostloop` and `hostloop-computer-links`, both green. **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously. `smoke-multiselect-deciderdir` runs the `--decider-llm` path, and `smoke-l2-microvm` passed in a real VM with its own kernel. Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`**. Both credential paths were exercised: the `.env` credentials for the container/protocol tiers, and the agent's own macOS-Keychain self-sourcing at `hostloop`. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (c) The first `examples/scenarios/` run is the one counted; there was no second. Separately and not a live matter: all three committed cassettes in `examples/replays/` are those re-recorded against `desktop-2.7032.0` on 2026-09-23; `verify-cassettes` exits 0 with all three clean and one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
207
207
|
|
|
208
208
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
209
209
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^3.8.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^3.8.1"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.8.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.8.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.8.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.8.1"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Not sure a harness is what you need? The next two sections are the argument.
|
|
@@ -348,7 +348,7 @@
|
|
|
348
348
|
"asarPath": "/Applications/Claude.app/Contents/Resources/app.asar",
|
|
349
349
|
"asarFingerprint": "7c4da429c2aae02c",
|
|
350
350
|
"gates": {
|
|
351
|
-
"$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift. The skill-family gates are pinned on the same principle but are NOT all dormant: 245679952 (suggestSkillsEnabled)
|
|
351
|
+
"$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift. The skill-family gates are pinned on the same principle but are NOT all dormant: 245679952 (suggestSkillsEnabled) is modeled — it gates whether suggest_skills is declared. 3246569822 (canSaveSkill) is DEAD in Desktop code from 1.44121.1 (0 occurrences in every asar from 1.44121.1 through this one) although the fcache still serves it ON/force: its row is a record of what the server sends, not a sentinel, and a flip of it changes nothing. The mcp__cowork__save_skill tool it once gated is present in real sessions and is not modeled — see docs/fidelity-gaps.md. 1824824999 (canProposeSkills) is present-but-off, pinned so the same class of silent widening cannot land unnoticed. 1598976391 (proactiveSkillSuggestEnabled) flipped off/defaultValue -> ON/force by a SERVER-SIDE rollout observed 2026-08-04 — NOT a Desktop change: the gate id occurs exactly once in both the 1.24012.9 and 1.24012.11 asars and would read ON on .9 today, so this value is Desktop-version-INDEPENDENT despite living in a version-named file (same provenance class as canSaveSkill above). It is read from a SINGLE account's fcache; force rules are server-evaluated and can be segment-targeted, so whether the rollout is global is not determinable from anything on disk. Like canSaveSkill, this id is DEAD in Desktop code: 0 occurrences in every asar from 1.46388.3 through this one, where the proactive suggest_skills description and the trigger param are unconditional whenever suggest_skills is declared. The harness reads this row only for a baseline older than 1.46388.3; from there proactive mode is always on, as in production, so a flip of this row does not change a run. Override per-session with skills.proactive_suggest_enabled. 4074604942 (1p-direct-mcp) was NEW in 1.24012.11 and DARK (absent from a standard fcache, hence its DARK_GATES entry); it arms a Desktop-side direct-MCP pool for MDM-managed 1P servers, inert for an unmanaged account, pinned as a sentinel only. Observed 2026-08-05 SERVED rather than absent (source \"force\", value false) — the rollout reached this account with the gate OFF, so nothing it arms is reachable and no modeled surface changes. Same provenance class as canSaveSkill: read from a SINGLE account's fcache, and force rules are server-evaluated and segment-targetable, so its DARK_GATES entry is retained deliberately — another account may still see it absent, and that must stay tolerated rather than hard-failing their sync.",
|
|
352
352
|
"emitToolUseSummaries:66187241": {
|
|
353
353
|
"on": false,
|
|
354
354
|
"source": "defaultValue",
|
package/dist/loop-decision.js
CHANGED
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
import { cmpVersionStrings } from "./baseline.js";
|
|
1
2
|
/**
|
|
2
3
|
* Read a GrowthBook gate sub-flag (e.g. `coworkWebFetchPrompt`) from the baseline's `provenance.gates`.
|
|
3
4
|
*
|
|
@@ -129,13 +130,33 @@ export function readGateBool(baseline, id) {
|
|
|
129
130
|
*
|
|
130
131
|
* Defaults when the gate is absent from the baseline: `suggestSkills` → true, `proactiveSuggest` → false
|
|
131
132
|
* (the documented production state).
|
|
133
|
+
*
|
|
134
|
+
* Proactive mode is version-dependent. Up to Desktop 1.44121.1 it is gate `1598976391`. From
|
|
135
|
+
* {@link PROACTIVE_SUGGEST_UNCONDITIONAL_FROM} the gate id is gone from the asar and `suggest_skills` always
|
|
136
|
+
* carries the proactive description and `trigger` param whenever it is declared — so for those baselines
|
|
137
|
+
* the gate row is IGNORED even though the fcache still serves it. Reading it there would let a server-side
|
|
138
|
+
* flip of a gate Desktop no longer reads turn the harness's proactive mode off while production keeps it on.
|
|
139
|
+
* The session knob still wins either way.
|
|
132
140
|
*/
|
|
133
141
|
export function resolveSkillDiscoveryGates(baseline, knobs = {}) {
|
|
134
142
|
return {
|
|
135
143
|
suggestSkillsEnabled: knobs.suggest_enabled ?? readGateBool(baseline, "245679952") ?? true,
|
|
136
|
-
proactiveSkillSuggestEnabled: knobs.proactive_suggest_enabled ??
|
|
144
|
+
proactiveSkillSuggestEnabled: knobs.proactive_suggest_enabled ?? proactiveFromBaseline(baseline),
|
|
137
145
|
};
|
|
138
146
|
}
|
|
147
|
+
/** First backed-up Desktop build whose asar has no reference to gate `1598976391` (0 occurrences from
|
|
148
|
+
* here through 2.7032.0; 2 in 1.44121.1, the previous backup — builds in between are unobserved, and a
|
|
149
|
+
* baseline for one would take the gate path, the conservative side). From this build proactive suggest
|
|
150
|
+
* mode is unconditional. */
|
|
151
|
+
export const PROACTIVE_SUGGEST_UNCONDITIONAL_FROM = "1.46388.3";
|
|
152
|
+
function proactiveFromBaseline(baseline) {
|
|
153
|
+
// `appVersion` is required by the schema, but synthetic test baselines omit it; a missing version
|
|
154
|
+
// compares as 0.0.0 and takes the gate path, which is the pre-1.46388.3 behaviour.
|
|
155
|
+
const v = baseline.appVersion;
|
|
156
|
+
if (typeof v === "string" && cmpVersionStrings(v, PROACTIVE_SUGGEST_UNCONDITIONAL_FROM) >= 0)
|
|
157
|
+
return true;
|
|
158
|
+
return readGateBool(baseline, "1598976391") ?? false;
|
|
159
|
+
}
|
|
139
160
|
export function decideLoop(inputs) {
|
|
140
161
|
if (inputs.requireFullVmSandbox === true)
|
|
141
162
|
return "vm"; // HeA()
|
package/dist/session.js
CHANGED
|
@@ -142,9 +142,10 @@ export const SessionConfig = z.strictObject({
|
|
|
142
142
|
.optional()
|
|
143
143
|
.describe("override for gate 1598976391 (proactiveSkillSuggestEnabled): when true (and suggest_enabled is not " +
|
|
144
144
|
"false), suggest_skills swaps to the proactive description, gains an optional `trigger` enum param, " +
|
|
145
|
-
"and chains its empty-catalog note into search_plugins. Omit to
|
|
146
|
-
"
|
|
147
|
-
"predates the gate
|
|
145
|
+
"and chains its empty-catalog note into search_plugins. Omit to follow the baseline: from the 1.46388.3 " +
|
|
146
|
+
"baseline Desktop does not read the gate and proactive mode is always on; for an older baseline the " +
|
|
147
|
+
"synced gate decides (on from 1.24012.11; false for a baseline that predates the gate). On a 1.46388.3+ " +
|
|
148
|
+
"baseline, false builds a non-proactive surface production does not ship there."),
|
|
148
149
|
})
|
|
149
150
|
.default({ local: [] }),
|
|
150
151
|
mcp: z
|
package/dist/sync/cowork-sync.js
CHANGED
|
@@ -108,24 +108,38 @@ export const PINNED_GATES = {
|
|
|
108
108
|
// init.tools of 8 real sessions) render, and in what mode. None was pinned before, so 245679952
|
|
109
109
|
// being live on/force was invisible to the drift guard. Present in the live fcache (NOT dark), so
|
|
110
110
|
// they are read at their real state — no DARK_GATES entry. BEHAVIORALLY MODELED since A2: the harness
|
|
111
|
-
// declares the skills/plugins SDK-MCP servers and reads
|
|
112
|
-
//
|
|
113
|
-
//
|
|
111
|
+
// declares the skills/plugins SDK-MCP servers and reads 245679952 at spawn (`resolveSkillDiscoveryGates`),
|
|
112
|
+
// so a flip of it CHANGES the declared tool set on container/hostloop (see
|
|
113
|
+
// `src/hostloop/skills-handler.ts`) — it is NOT inert. 1598976391 is read only for a baseline older than
|
|
114
|
+
// 1.46388.3 (see its entry below). A pinned drift alone WARNS + still writes.
|
|
114
115
|
"245679952": "suggestSkillsEnabled", // live on/force — gates whether suggest_skills renders at all
|
|
115
|
-
// proactive (unprompted) suggest mode.
|
|
116
|
-
//
|
|
117
|
-
//
|
|
118
|
-
//
|
|
119
|
-
//
|
|
120
|
-
//
|
|
121
|
-
//
|
|
122
|
-
//
|
|
116
|
+
// proactive (unprompted) suggest mode. DEAD IN DESKTOP CODE FROM 1.46388.3: 2 occurrences in the 1.40609.1
|
|
117
|
+
// and 1.44121.1 asars, 0 in every asar from 1.46388.3 through 2.7032.0, and the `proactiveSkillSuggestEnabled`
|
|
118
|
+
// name is gone too. From that build Desktop gives suggest_skills its proactive description and `trigger`
|
|
119
|
+
// param unconditionally whenever the tool is declared. Up to 1.44121.1 the gate was live, with three effects:
|
|
120
|
+
// the description swap, `trigger`, and a swapped suggest-guidance line in the generated
|
|
121
|
+
// `<skills_instructions>` block (a prompt effect the harness never models).
|
|
122
|
+
// The harness reads this row at spawn ONLY for a baseline older than 1.46388.3; from there
|
|
123
|
+
// `resolveSkillDiscoveryGates` ignores it and proactive mode is always on, as in production. So a
|
|
124
|
+
// server-side flip here is a sync diff with no effect on runs against a current baseline. Like
|
|
125
|
+
// canSaveSkill below, the row is a record of what the server sends, not a sentinel.
|
|
123
126
|
"1598976391": "proactiveSkillSuggestEnabled",
|
|
124
|
-
//
|
|
125
|
-
//
|
|
126
|
-
//
|
|
127
|
-
//
|
|
128
|
-
//
|
|
127
|
+
// DEAD IN DESKTOP CODE SINCE 1.44121.1 — THIS PIN PROVIDES NO COVERAGE OF `save_skill`. The id has 3
|
|
128
|
+
// occurrences in the 1.40609.1 asar and 0 in every asar from 1.44121.1 through 2.7032.0 (count with
|
|
129
|
+
// `grep -ao <id> app.asar | wc -l`; `grep -c` counts lines of a minified file and undercounts). The
|
|
130
|
+
// fcache still serves it on/force, so sync round-trips it and reports "no change": a green light on a
|
|
131
|
+
// disconnected sensor. Up to 1.40609.1 the gate was live, and its off -> on flip on 2026-07-25 is what
|
|
132
|
+
// put `mcp__cowork__save_skill` into real sessions. The row is kept only as a record of what the server
|
|
133
|
+
// sends: baselines from 1.24012.1 on carry it (from 1.44121.1 on with a `note` saying it is unread), and
|
|
134
|
+
// docs/fidelity-gaps.md names it. Do not read it as a tripwire.
|
|
135
|
+
// At 2.7032.0 the tool follows `skillsEnabled`, the managed `skillCreationEnabled` setting and two gates
|
|
136
|
+
// that first appear in that build: 3656976882 (org skills off) and 3469616823 (bypass for the org skill-
|
|
137
|
+
// creation block). The org check reads a `skill_creation` status from an account access list Desktop
|
|
138
|
+
// fetches at runtime (`/api/bootstrap/<org>/current_user_access`, held in memory, absent from the fcache)
|
|
139
|
+
// and can only turn the tool OFF — an unloaded list allows it. Pinning the two newer gates would NOT make
|
|
140
|
+
// this a sentinel: the access list sits outside the fcache. The observable outcome is whether a REAL
|
|
141
|
+
// Desktop session's init tool list carries `mcp__cowork__save_skill` — never a harness run's, which lists
|
|
142
|
+
// only what the harness itself declares and does not declare `save_skill` at any tier.
|
|
129
143
|
"3246569822": "canSaveSkill",
|
|
130
144
|
// off/defaultValue and PRESENT in the fcache (so NOT dark — no DARK_GATES entry) — the `propose_skills`
|
|
131
145
|
// render-only sibling. Pinned so a production flip surfaces as a sync diff instead of silently widening
|
package/docs/ci.md
CHANGED
|
@@ -107,7 +107,7 @@ jobs:
|
|
|
107
107
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
108
108
|
```
|
|
109
109
|
|
|
110
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.8.
|
|
110
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.8.1` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
111
111
|
|
|
112
112
|
CI uses `ANTHROPIC_API_KEY` specifically because there's no interactive browser available to run
|
|
113
113
|
`claude setup-token`'s OAuth flow in a GitHub Actions runner; locally, the OAuth token is preferred because
|
package/docs/cli.md
CHANGED
|
@@ -18,7 +18,7 @@ companion skill, CI). This page is the CLI one.
|
|
|
18
18
|
**Install from npm:**
|
|
19
19
|
|
|
20
20
|
```bash
|
|
21
|
-
npm install -g "cowork-harness@^3.8.
|
|
21
|
+
npm install -g "cowork-harness@^3.8.1" # puts the `cowork-harness` command on your PATH
|
|
22
22
|
```
|
|
23
23
|
|
|
24
24
|
**Or build from source:**
|
|
@@ -38,7 +38,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
38
38
|
|
|
39
39
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
40
40
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
41
|
-
> From a global install (`npm i -g "cowork-harness@^3.8.
|
|
41
|
+
> From a global install (`npm i -g "cowork-harness@^3.8.1"`), point at the package root instead:
|
|
42
42
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
43
43
|
> (or copy the cassette into your own project and pass that path).
|
|
44
44
|
|
|
@@ -48,7 +48,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
48
48
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
49
49
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
50
50
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
51
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.8.
|
|
51
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.8.1"`.
|
|
52
52
|
|
|
53
53
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
54
54
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -124,7 +124,7 @@ The parts people ask about. **`package.json`'s `files[]` is the exhaustive, mach
|
|
|
124
124
|
this table is the readable summary of it, and deliberately omits the infrastructure that always ships
|
|
125
125
|
(`baselines/`, `schema/`, `fixtures/`, `scripts/`, `docker/`).
|
|
126
126
|
|
|
127
|
-
| What ships | npm global (`npm install -g "cowork-harness@^3.8.
|
|
127
|
+
| What ships | npm global (`npm install -g "cowork-harness@^3.8.1"`) | Source checkout (`git clone` + `npm ci`) |
|
|
128
128
|
|---|---|---|
|
|
129
129
|
| CLI, `scenario.py` + assertion keys (enough for `lint` in CI) | ✓ | ✓ |
|
|
130
130
|
| `SKILL.md`, all of `docs/`, `SPEC.md`/`DESIGN.md`/`AGENTS.md` | ✓ | ✓ |
|
|
@@ -139,7 +139,7 @@ since a global install puts nothing in your working directory. The matrix, answe
|
|
|
139
139
|
examples are the only ones that still need a source checkout. The **marketplace skill install** is
|
|
140
140
|
narrower again — it pulls only `.claude/skills/cowork-harness/` (SKILL.md + `references/` +
|
|
141
141
|
`scenario.py`/assertion keys, per `.claude-plugin/marketplace.json`'s `source`); everything in the npm column
|
|
142
|
-
arrives when the skill's first command self-bootstraps `npx "cowork-harness@^3.8.
|
|
142
|
+
arrives when the skill's first command self-bootstraps `npx "cowork-harness@^3.8.1"` — the last row stays
|
|
143
143
|
✗ either way, since `matrices/`, `answer-policies/` and `probes/` are not published at all. See
|
|
144
144
|
[docs/companion-skill.md](./companion-skill.md) for that install path.
|
|
145
145
|
|
package/docs/companion-skill.md
CHANGED
|
@@ -28,7 +28,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
28
28
|
claude plugin install cowork-harness@cowork-harness
|
|
29
29
|
```
|
|
30
30
|
|
|
31
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.8.
|
|
31
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.8.1"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
32
32
|
|
|
33
33
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
34
34
|
|
|
@@ -41,7 +41,7 @@ npx skills add yaniv-golan/cowork-harness --skill cowork-harness
|
|
|
41
41
|
**What the marketplace install actually pulls:** only `.claude/skills/cowork-harness/` — SKILL.md +
|
|
42
42
|
`references/` + `scenario.py`/assertion keys, per `.claude-plugin/marketplace.json`'s `source`. Everything
|
|
43
43
|
else (the CLI, `docs/`, the worked examples, the pytest lane) arrives when the skill's first command
|
|
44
|
-
self-bootstraps `npx "cowork-harness@^3.8.
|
|
44
|
+
self-bootstraps `npx "cowork-harness@^3.8.1"`, which pulls the same npm package as a global install.
|
|
45
45
|
|
|
46
46
|
For the full package-contents table — what a global install gives you versus a source checkout — see
|
|
47
47
|
[docs/cli.md → What ships](./cli.md#what-ships). It is maintained there, once.
|
package/docs/fidelity-gaps.md
CHANGED
|
@@ -327,7 +327,7 @@ ordinary session**:
|
|
|
327
327
|
session shadows remote servers at all (gate `2529235968`, or at least one third-party direct MCP server
|
|
328
328
|
present) **and** a stand-in already provides that server, matched by URL hostname or by name. The
|
|
329
329
|
stand-ins are the session's claude.ai connectors that carry at least one enabled tool, plus those direct
|
|
330
|
-
servers. Logged *"Plugin `<
|
|
330
|
+
servers. Logged *"Plugin `<id>` declares remote MCP servers (…). Overriding with no-ops so the CLI does
|
|
331
331
|
not open its own client."* With the gate off and no such server, the remote arm never runs, and a
|
|
332
332
|
plugin's remote servers reach the CLI intact. Earlier builds (1.37937.0 through 1.46388.x) stubbed
|
|
333
333
|
**every** remote plugin server once the gate was on; the stand-in narrowing arrives with 2.2553.1.
|
|
@@ -339,9 +339,28 @@ A plugin Desktop treats as official is exempt from both arms unless a policy is
|
|
|
339
339
|
|
|
340
340
|
A replaced server is renamed `plugin:<pluginName>:<serverName>` and constructed as
|
|
341
341
|
`createSdkMcpServer({name, tools: []})` — **the server name is present in the session's inventory and offers
|
|
342
|
-
zero tools.**
|
|
343
|
-
an `--mcp-config` payload
|
|
344
|
-
|
|
342
|
+
zero tools.** Every stub reaches the agent in Desktop's in-process SDK server map. Which stubs are _also_
|
|
343
|
+
named in `cowork-plugin-mcp-shadow.json`, delivered as an `--mcp-config` payload, depends on whether
|
|
344
|
+
Desktop enforces remote shadowing for the session: gate `2529235968` on **and** no enterprise
|
|
345
|
+
managed-configuration override in force. When it does, the file names every replaced server, policy stubs
|
|
346
|
+
included. When it does not, the file names only the remote servers a stand-in replaced, and a session with
|
|
347
|
+
none writes no file at all; local and `.mcpb` policy stubs then travel in the SDK map alone.
|
|
348
|
+
|
|
349
|
+
While an MCP policy is active, Desktop refuses to start the session rather than launch with unenforced
|
|
350
|
+
plugin servers whenever building the overrides fails, whatever the gate: a failed plugin scan (*"Plugin MCP
|
|
351
|
+
scan failed while an MCP policy is active"*) or any other error in that step. It also refuses when the
|
|
352
|
+
shadow file cannot be written, but only while it enforces remote shadowing; otherwise a failed write is
|
|
353
|
+
logged and the spawn goes ahead without the file.
|
|
354
|
+
|
|
355
|
+
Desktop's own MCP pool, `LocalMcpServerManager`, takes only a plugin's local stdio and `.mcpb` servers, and
|
|
356
|
+
skips any a policy blocks. Remote plugin servers that no stand-in replaced are left for the CLI to open
|
|
357
|
+
itself. **A stub therefore never appears as a `LocalMcpServerManager` connection**, which makes Desktop's
|
|
358
|
+
`main.log` an unambiguous record of the remote arm: each replacement logs
|
|
359
|
+
`Replacing plugin "<plugin>" MCP server "<server>": "<connector>" already provides it (matched by url)`,
|
|
360
|
+
then `Plugin "<id>" declares remote MCP servers (…). Overriding with no-ops …`. On a Desktop whose
|
|
361
|
+
connected Slack, Notion or Airtable connector matches a plugin's declared server by URL, both lines appear
|
|
362
|
+
for every session that loads the plugin, and the plugin servers that connect through
|
|
363
|
+
`LocalMcpServerManager` are only local stdio ones.
|
|
345
364
|
|
|
346
365
|
**What the harness does:** nothing. Plugins are staged with `--plugin-dir` and the CLI reads each plugin's
|
|
347
366
|
own declaration, so a plugin under test gets **working** MCP servers with their real tools.
|
|
@@ -852,19 +871,23 @@ over the control protocol (`sdkMcpServers` in `initialize`, tunneled as `mcp_mes
|
|
|
852
871
|
`list_skills` returns no match, and the result renders an "Add" card. The call has **no side effect**
|
|
853
872
|
(nothing installs; the user's Add click happens out of band). Ground truth: these tools appear in the
|
|
854
873
|
`system/init` `tools` array of real on-disk sessions (`local-agent-mode-sessions/**/audit.jsonl`); the
|
|
855
|
-
`suggestSkillsEnabled` gate `245679952` is on.
|
|
856
|
-
`
|
|
857
|
-
|
|
858
|
-
|
|
859
|
-
|
|
874
|
+
`suggestSkillsEnabled` gate `245679952` is on. `suggest_skills` also has a proactive-suggestion mode.
|
|
875
|
+
From Desktop `1.46388.3` (the first backed-up build after `1.44121.1`) that mode is **unconditional** whenever `suggest_skills` is declared: no gate
|
|
876
|
+
is read for it. Up to `1.44121.1` it sits behind gate `1598976391`
|
|
877
|
+
(`proactiveSkillSuggestEnabled`), which the server serves **on** for a standard account from the
|
|
878
|
+
`1.24012.11` baseline; the server still serves that gate, but Desktop from `1.46388.3` never reads it.
|
|
879
|
+
With the mode off, `suggest_skills` keeps its base description and the model suggests only when the
|
|
880
|
+
conversation invites it. With it on, the tool gains an optional `trigger` parameter
|
|
860
881
|
(`user_asked` | `proactive`), a proactive description that also carries production's *constraints* (a
|
|
861
882
|
do-not-call list, a suggest-at-most-once-per-conversation rule, a no-lead-in rule, and forwarding the
|
|
862
883
|
same keywords **and trigger** to `search_plugins`), and an empty-catalog `note` that chains into
|
|
863
884
|
`search_plugins` for every trigger state — silence is only the `proactive` tail, and a trigger the model
|
|
864
|
-
never supplied is never forwarded back to it. One production effect is **not** modeled:
|
|
865
|
-
|
|
866
|
-
`<skills_instructions>` block.
|
|
867
|
-
|
|
885
|
+
never supplied is never forwarded back to it. One production effect is **not** modeled: a proactive
|
|
886
|
+
suggest-guidance line (with a suggest-at-most-once-per-conversation sentence) inside the generated
|
|
887
|
+
`<skills_instructions>` block. Up to `1.44121.1` the gate selects it; from `1.46388.3` Desktop emits it
|
|
888
|
+
whenever `suggest_skills` and `search_plugins` are available, so on a current baseline this gap applies to
|
|
889
|
+
every session that declares `suggest_skills`. The harness renders no such section at all, so the effect
|
|
890
|
+
lands in an already-unmodeled surface.
|
|
868
891
|
|
|
869
892
|
**Harness behaviour:** `container` and `hostloop` (and `cowork`, which resolves to one of those) now
|
|
870
893
|
declare a `skills` and a `plugins` SDK-MCP server alongside `cowork`/`workspace` (`combineSdkMcp`,
|
|
@@ -878,9 +901,11 @@ synced baseline (`readGateBool`, bare-boolean shape — distinct from the sub-fl
|
|
|
878
901
|
reads) with a session-level override (`skills.suggest_enabled` / `skills.proactive_suggest_enabled`, see
|
|
879
902
|
[session.md](./session.md)). Precedence is knob ▸ baseline gate ▸ hardcoded fallback, and the three are
|
|
880
903
|
distinct: omit the knob and the value comes from the **synced baseline** (on `latest` that is
|
|
881
|
-
`suggestSkillsEnabled` on
|
|
882
|
-
the
|
|
883
|
-
|
|
904
|
+
`suggestSkillsEnabled` on); the hardcoded fallback, which applies only to a baseline old enough to predate
|
|
905
|
+
the gate entirely, is on. Proactive mode follows the Desktop version the same way production does: from
|
|
906
|
+
the `1.46388.3` baseline it is **always on** and the `1598976391` row is ignored, so a server-side flip of
|
|
907
|
+
a gate Desktop does not read cannot change a run; for an older baseline the synced gate decides, with a
|
|
908
|
+
hardcoded fallback of **off** for a baseline that predates it.
|
|
884
909
|
|
|
885
910
|
**How exact is the model?** Not uniformly — and the difference matters, so it is stated plainly. The
|
|
886
911
|
tool **inventory** (which five tools exist), their **inputSchemas**, the **gating** semantics, and the
|
package/docs/session.md
CHANGED
|
@@ -95,11 +95,11 @@ plugins:
|
|
|
95
95
|
skills:
|
|
96
96
|
local: [] # extra host skill dirs → CLAUDE_CONFIG_DIR/skills
|
|
97
97
|
# Overrides for the skill-discovery SDK-MCP gates (container/hostloop only — see fidelity-gaps.md
|
|
98
|
-
# "Skill/plugin discovery SDK-MCP servers"). Precedence: this knob → the
|
|
99
|
-
#
|
|
100
|
-
#
|
|
98
|
+
# "Skill/plugin discovery SDK-MCP servers"). Precedence: this knob → the baseline → the documented
|
|
99
|
+
# default. Both omit-able; unset, suggest_enabled follows gate 245679952, and proactive_suggest_enabled is
|
|
100
|
+
# always on from the 1.46388.3 baseline and follows gate 1598976391 before it.
|
|
101
101
|
suggest_enabled: true # gate 245679952 (suggestSkillsEnabled) override; default true when unset
|
|
102
|
-
proactive_suggest_enabled: false # gate 1598976391 (proactiveSkillSuggestEnabled) override; unset = the
|
|
102
|
+
proactive_suggest_enabled: false # gate 1598976391 (proactiveSkillSuggestEnabled) override; unset = always on from the 1.46388.3 baseline, the synced gate before it
|
|
103
103
|
mcp:
|
|
104
104
|
config: null # --mcp-config file (standard mcpServers map), e.g. ../data/mcp.json
|
|
105
105
|
enabled: [] # enabledMcpjsonServers
|
|
@@ -229,7 +229,7 @@ See [discovery.md](./discovery.md) for the full model. In short: the harness bui
|
|
|
229
229
|
| `plugins.local_plugins[]` / `remote_plugins[]` | Cowork plugin mounts | → `mnt/.local-plugins/marketplaces/<marketplace>/<plugin>` (≥1.14271.0; older baselines use `.local-plugins/cache`) / `mnt/.remote-plugins/plugin_<id>` (migrated-Cowork uploaded/org-remote shape; the id is a stable hash of the declared source). A skill that references these via `${CLAUDE_PLUGIN_ROOT}` must mind [the two-namespace resolution model](./plugin-root.md) — the token is unset in host-loop VM bash. |
|
|
230
230
|
| `skills.local[]` | `CLAUDE_CONFIG_DIR/skills` | extra host **skill** dirs (a folder *without* `.claude-plugin/plugin.json`) staged into the config dir's `skills/`. Use this for a single-skill folder; use `plugins.local_plugins` for a plugin root. |
|
|
231
231
|
| `skills.suggest_enabled` | gate `245679952` (`suggestSkillsEnabled`) override | `container`/`hostloop` (and `cowork`) only. `true` (or unset, if the synced baseline gate is on/absent) declares the `skills` SDK-MCP server's `suggest_skills` tool; `false` omits it and drops `list_skills`' fallback-to-`suggest_skills` clause. See [fidelity-gaps.md](./fidelity-gaps.md). |
|
|
232
|
-
| `skills.proactive_suggest_enabled` | gate `1598976391` (`proactiveSkillSuggestEnabled`) override | Only consulted when `suggest_enabled` (effective) is true. `true` swaps `suggest_skills` to the proactive description, adds an optional `trigger` enum param (`user_asked` \| `proactive`), and chains the empty-catalog `note` into `search_plugins`. Omit to
|
|
232
|
+
| `skills.proactive_suggest_enabled` | gate `1598976391` (`proactiveSkillSuggestEnabled`) override | Only consulted when `suggest_enabled` (effective) is true. `true` swaps `suggest_skills` to the proactive description, adds an optional `trigger` enum param (`user_asked` \| `proactive`), and chains the empty-catalog `note` into `search_plugins`. Omit to follow the baseline: from the `1.46388.3` baseline Desktop does not read the gate and proactive mode is **always on**, so the gate row is ignored; for an older baseline the synced gate decides (**on** from `1.24012.11`; `false` for a baseline predating the gate). On a baseline from `1.46388.3`, `false` builds a non-proactive `suggest_skills` that production does not ship there; use it only to exercise the older surface. |
|
|
233
233
|
| `mcp.config` / `mcp.enabled[]` | `--mcp-config` / `enabledMcpjsonServers` | the supported way to attach an MCP server to a session under test. |
|
|
234
234
|
|
|
235
235
|
> Inside a git repo, `folders[]` and `skills.local[]` stage only **git-tracked** files into the mount (matching
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.8.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.8.1"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "3.8.
|
|
3
|
+
"version": "3.8.1",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -200,7 +200,7 @@
|
|
|
200
200
|
"type": "boolean"
|
|
201
201
|
},
|
|
202
202
|
"proactive_suggest_enabled": {
|
|
203
|
-
"description": "override for gate 1598976391 (proactiveSkillSuggestEnabled): when true (and suggest_enabled is not false), suggest_skills swaps to the proactive description, gains an optional `trigger` enum param, and chains its empty-catalog note into search_plugins. Omit to
|
|
203
|
+
"description": "override for gate 1598976391 (proactiveSkillSuggestEnabled): when true (and suggest_enabled is not false), suggest_skills swaps to the proactive description, gains an optional `trigger` enum param, and chains its empty-catalog note into search_plugins. Omit to follow the baseline: from the 1.46388.3 baseline Desktop does not read the gate and proactive mode is always on; for an older baseline the synced gate decides (on from 1.24012.11; false for a baseline that predates the gate). On a 1.46388.3+ baseline, false builds a non-proactive surface production does not ship there.",
|
|
204
204
|
"type": "boolean"
|
|
205
205
|
}
|
|
206
206
|
},
|