cowork-harness 3.2.0 → 3.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +36 -0
- package/DESIGN.md +2 -2
- package/README.md +4 -4
- package/RELEASING.md +14 -4
- package/baselines/desktop-1.40609.1.json +878 -0
- package/docs/ci.md +2 -2
- package/docs/cli.md +3 -3
- package/docs/companion-skill.md +3 -3
- package/docs/maintenance.md +1 -1
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/package.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.2.
|
|
7
|
-
tracks-harness: cowork-harness 3.2.
|
|
6
|
+
version: 3.2.1
|
|
7
|
+
tracks-harness: cowork-harness 3.2.1 (baseline desktop-1.40609.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.2.
|
|
29
|
-
> `desktop-1.40609.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.2.1` (baseline
|
|
29
|
+
> `desktop-1.40609.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.2.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.2.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.2.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.2.1"`. **Pin `@^3.2.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.2.1` (baseline `desktop-1.40609.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.2.
|
|
20
|
+
(e.g. `version: "3.2.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.255 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
41
41
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
42
42
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^3.2.
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.2.1"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -344,7 +344,7 @@ jobs:
|
|
|
344
344
|
with: { node-version: '24' }
|
|
345
345
|
- uses: actions/setup-python@v5
|
|
346
346
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
347
|
-
- run: npm i -g "cowork-harness@^3.2.
|
|
347
|
+
- run: npm i -g "cowork-harness@^3.2.1"
|
|
348
348
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
349
349
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
350
350
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -373,7 +373,7 @@ jobs:
|
|
|
373
373
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
374
374
|
fi
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
|
-
run: npm i -g "cowork-harness@^3.2.
|
|
376
|
+
run: npm i -g "cowork-harness@^3.2.1"
|
|
377
377
|
- if: steps.guard.outputs.live == 'true'
|
|
378
378
|
run: cowork-harness run scenarios/ --output-format json
|
|
379
379
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.2.
|
|
3
|
+
Tracks `cowork-harness 3.2.1` (baseline `desktop-1.40609.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.2.
|
|
4
|
-
(baseline `desktop-1.40609.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.2.1`
|
|
4
|
+
(baseline `desktop-1.40609.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.2.
|
|
5
|
+
Tracks `cowork-harness 3.2.1` (baseline `desktop-1.40609.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,42 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.2.1] — 2026-09-02
|
|
10
|
+
|
|
11
|
+
### Parity
|
|
12
|
+
|
|
13
|
+
- **Baseline `desktop-1.40609.1` (agent `2.1.255`).** `sync` wrote on the first attempt — no unknown
|
|
14
|
+
deltas. The **`spawn` and `network` blocks are byte-identical** to `desktop-1.40609.0`, as are all 29
|
|
15
|
+
recorded gate rows, the asar's gate-id set (301, none added or removed), the VM rootfs hash, the Cowork
|
|
16
|
+
system prompt and both sub-agent appends. The three committed example cassettes were **re-stamped**, not
|
|
17
|
+
re-recorded: `promptAssetsHash` resolves to the same `491afe2862dc67ea` under both baselines, and with
|
|
18
|
+
the spawn contract, egress policy, prompt and tool surface all measured unchanged, a re-record would
|
|
19
|
+
have bought nothing.
|
|
20
|
+
|
|
21
|
+
- **`agentBinary.manifestChecksumMatch` reads `"unknown"` for this baseline, and that is a limitation of
|
|
22
|
+
the check, not a finding about the binary.** Agent `2.1.255` is served from a release-candidate path
|
|
23
|
+
rather than the stable versioned one that `sync` queries, so the cross-check 404s and — by design — is
|
|
24
|
+
swallowed rather than failing the sync. The recorded `sha256` is still `measured-local`, hashed from the
|
|
25
|
+
staged ELF, and it does match the checksum the release candidate's own manifest publishes; the baseline
|
|
26
|
+
simply cannot say so through this field yet. Run-time ELF verification against the recorded hash is
|
|
27
|
+
unaffected.
|
|
28
|
+
|
|
29
|
+
- **Agent `2.1.247` → `2.1.255`.** The agent's `CLAUDE_*` env-flag table moved 569 → 588 (+24, −5).
|
|
30
|
+
**None of the new flags is set by the Cowork spawn**, and no flag the spawn does set changed from or to
|
|
31
|
+
zero consumers — including `CLAUDE_PREVIEW_CLASSIFIER_FLOOR`, which remains inert agent-side (the
|
|
32
|
+
rename to `CLAUDE_CHROME_CLASSIFIER_FLOOR` recorded in 2.3.0 still stands, and Desktop still has not
|
|
33
|
+
followed it). No harness change follows from the bump.
|
|
34
|
+
|
|
35
|
+
### Documentation
|
|
36
|
+
|
|
37
|
+
- **`RELEASING.md`'s smoke step: `--package=` pins the fetch, but cwd decides which binary runs.**
|
|
38
|
+
Observed releasing 3.2.0 using the invocation this file prescribed: from the repo root,
|
|
39
|
+
`npx -y --package=cowork-harness@3.2.0 -- cowork-harness --version` printed **3.1.0** — a stale global on
|
|
40
|
+
PATH — while the identical command from `/tmp` printed 3.2.0. The step now takes the `--version` reading
|
|
41
|
+
from outside the repo, then `npm i -g`s the published version so the repo-root `doctor`/`replay` checks
|
|
42
|
+
run against the artifact under test.
|
|
43
|
+
|
|
44
|
+
|
|
9
45
|
## [3.2.0] — 2026-09-01
|
|
10
46
|
|
|
11
47
|
### Changed
|
package/DESIGN.md
CHANGED
|
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
47
47
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
48
48
|
|
|
49
49
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
50
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
50
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.255**, per `baselines/desktop-1.40609.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
|
|
51
51
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
52
52
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
53
53
|
|
|
@@ -176,7 +176,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
176
176
|
|
|
177
177
|
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
178
178
|
|
|
179
|
-
> **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is no longer the newest committed baseline: **
|
|
179
|
+
> **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is no longer the newest committed baseline: **two** baselines have shipped since (`1.40609.0`, `1.40609.1`), **two** of which moved the agent ELF, most recently to **2.1.255** — so the newest baseline is **not** live-verified, and this paragraph describes the 1.37937.1 pass only. The pass ran on agent `2.1.246` (the staged VM ELF and the native `.app` were both at that version) and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe, resume continuity on the native binary, critique at the unpinned tier, uploads readability, sub-agent WebSearch capture, and the discovery-server declaration check). It was one invocation of `npm run test:live` — **4 suites, 19 assertions, 19 green / 0 skipped**. **Nothing was gated out**, which is the part worth stating: every `describe` in this lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported zero skips, and the 19 that ran are exactly the 19 the suite enumerates. **Scope-out, so this is not read as more than it is.** (a) The assertion population is 19 here against the 24 recorded for the 1.32885.1 pass; cases have been retired and consolidated since (the `live-outputs-delete` whole-line-`#`-comment case was retired in 1.25.0 after its pinned command stopped being executed by the model), so the two counts are not comparable and the drop is not coverage lost in this pass. (b) The `boundary-check` sandbox proof and the example-scenario suite were part of the 1.32885.1 stamp and were **NOT** run here — this paragraph claims `npm run test:live` only. (c) A live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: the three committed example cassettes (`example-pdf-skill` at `container`, `example-multiselect-gate` at `protocol`, `hostloop-computer-links` at `hostloop`) were re-recorded against this baseline in the same change, so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
180
180
|
|
|
181
181
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
182
182
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^3.2.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^3.2.1"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.2.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.2.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.2.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.2.1"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Not sure a harness is what you need? The next two sections are the argument.
|
|
@@ -353,6 +353,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
353
353
|
## Status
|
|
354
354
|
|
|
355
355
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
356
|
-
**`desktop-1.40609.
|
|
356
|
+
**`desktop-1.40609.1`**. Release-by-release verification notes (what was re-verified against
|
|
357
357
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
358
358
|
this section would otherwise duplicate lives in the sections above.
|
package/RELEASING.md
CHANGED
|
@@ -239,11 +239,21 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
|
|
|
239
239
|
globs match full `vX.Y.Z` semver only — so pushing them produces no workflow runs at all.)
|
|
240
240
|
- [ ] **Smoke the published artifact — and use THIS invocation, not a bare `npx`:**
|
|
241
241
|
```
|
|
242
|
-
|
|
243
|
-
npx -y --package=cowork-harness@X.Y.Z -- cowork-harness --version
|
|
244
|
-
|
|
245
|
-
|
|
242
|
+
# 1. version — from OUTSIDE the repo (see "cwd matters", below)
|
|
243
|
+
(cd /tmp && npx -y --package=cowork-harness@X.Y.Z -- cowork-harness --version) # must print X.Y.Z
|
|
244
|
+
# 2. install the published artifact, then smoke it from the repo root
|
|
245
|
+
npm i -g cowork-harness@X.Y.Z && cowork-harness --version # must print X.Y.Z
|
|
246
|
+
cd <repo root>
|
|
247
|
+
cowork-harness doctor --tier protocol
|
|
248
|
+
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
|
246
249
|
```
|
|
250
|
+
**`--package=` is necessary but NOT sufficient — cwd matters too.** Observed releasing 3.2.0:
|
|
251
|
+
run from the repo root, `npx -y --package=cowork-harness@3.2.0 -- cowork-harness --version` printed
|
|
252
|
+
**3.1.0** (the stale Homebrew global); the identical command from `/tmp` printed 3.2.0. So inside
|
|
253
|
+
the repo, npx still resolves the bin off PATH and the `--package=` pin buys you nothing — the same
|
|
254
|
+
false smoke `--package=` was added to prevent, one release later. Hence the split above: take the
|
|
255
|
+
`--version` reading from outside the repo, then `npm i -g` the published version so the binary on
|
|
256
|
+
PATH *is* the artifact under test, and run the repo-root checks against that.
|
|
247
257
|
**`npx cowork-harness@X.Y.Z …` is NOT good enough, and fails silently.** If a `cowork-harness` is
|
|
248
258
|
already on PATH (a global `npm i -g`, Homebrew shim, …), npx runs THAT binary and ignores the
|
|
249
259
|
`@X.Y.Z` spec entirely — no warning. Observed 2026-08-31 releasing 3.1.0: the smoke printed `3.0.1`
|
|
@@ -0,0 +1,878 @@
|
|
|
1
|
+
{
|
|
2
|
+
"baselineVersion": 1,
|
|
3
|
+
"appVersion": "1.40609.1",
|
|
4
|
+
"agentVersion": "2.1.255",
|
|
5
|
+
"agentBinary": {
|
|
6
|
+
"stagedPath": "~/Library/Application Support/Claude/claude-code-vm/2.1.255/claude",
|
|
7
|
+
"format": "elf-aarch64",
|
|
8
|
+
"nativeStagedPath": "~/Library/Application Support/Claude/claude-code/2.1.255/claude.app/Contents/MacOS/claude",
|
|
9
|
+
"sha256": "31822816a0d92b4b0324d84b99d349f25d3f252aca06edf964d7324f23e1bec7",
|
|
10
|
+
"shaProvenance": "measured-local",
|
|
11
|
+
"manifestChecksumMatch": "unknown",
|
|
12
|
+
"stringSentinels": {
|
|
13
|
+
"tengu_saddle_lantern": 2
|
|
14
|
+
}
|
|
15
|
+
},
|
|
16
|
+
"guest": {
|
|
17
|
+
"os": "linux",
|
|
18
|
+
"arch": "arm64",
|
|
19
|
+
"baseImage": "ubuntu:22.04"
|
|
20
|
+
},
|
|
21
|
+
"spawn": {
|
|
22
|
+
"configDirInGuest": "mnt/.claude",
|
|
23
|
+
"settingSources": [
|
|
24
|
+
"user"
|
|
25
|
+
],
|
|
26
|
+
"permissionMode": "default",
|
|
27
|
+
"maxThinkingTokens": 31999,
|
|
28
|
+
"effortDefault": "medium",
|
|
29
|
+
"effortByModel": {
|
|
30
|
+
"claude-haiku-4-5": {
|
|
31
|
+
"modes": [
|
|
32
|
+
"extended"
|
|
33
|
+
]
|
|
34
|
+
},
|
|
35
|
+
"claude-sonnet-4-5": {
|
|
36
|
+
"modes": [
|
|
37
|
+
"extended"
|
|
38
|
+
]
|
|
39
|
+
},
|
|
40
|
+
"claude-sonnet-4-6": {
|
|
41
|
+
"effortLevels": [
|
|
42
|
+
"low",
|
|
43
|
+
"medium",
|
|
44
|
+
"high",
|
|
45
|
+
"max"
|
|
46
|
+
],
|
|
47
|
+
"recommended": "low",
|
|
48
|
+
"modes": [
|
|
49
|
+
"auto"
|
|
50
|
+
]
|
|
51
|
+
},
|
|
52
|
+
"claude-sonnet-5": {
|
|
53
|
+
"effortLevels": [
|
|
54
|
+
"low",
|
|
55
|
+
"medium",
|
|
56
|
+
"high",
|
|
57
|
+
"xhigh",
|
|
58
|
+
"max"
|
|
59
|
+
],
|
|
60
|
+
"recommended": "medium",
|
|
61
|
+
"modes": [
|
|
62
|
+
"auto"
|
|
63
|
+
]
|
|
64
|
+
},
|
|
65
|
+
"claude-opus-4-6": {
|
|
66
|
+
"effortLevels": [
|
|
67
|
+
"low",
|
|
68
|
+
"medium",
|
|
69
|
+
"high",
|
|
70
|
+
"max"
|
|
71
|
+
],
|
|
72
|
+
"recommended": "medium",
|
|
73
|
+
"modes": [
|
|
74
|
+
"extended"
|
|
75
|
+
]
|
|
76
|
+
},
|
|
77
|
+
"claude-opus-4-7": {
|
|
78
|
+
"effortLevels": [
|
|
79
|
+
"low",
|
|
80
|
+
"medium",
|
|
81
|
+
"high",
|
|
82
|
+
"xhigh",
|
|
83
|
+
"max"
|
|
84
|
+
],
|
|
85
|
+
"recommended": "xhigh",
|
|
86
|
+
"modes": [
|
|
87
|
+
"auto"
|
|
88
|
+
]
|
|
89
|
+
},
|
|
90
|
+
"claude-opus-4-8": {
|
|
91
|
+
"effortLevels": [
|
|
92
|
+
"low",
|
|
93
|
+
"medium",
|
|
94
|
+
"high",
|
|
95
|
+
"xhigh",
|
|
96
|
+
"max"
|
|
97
|
+
],
|
|
98
|
+
"recommended": "high",
|
|
99
|
+
"modes": [
|
|
100
|
+
"auto"
|
|
101
|
+
]
|
|
102
|
+
},
|
|
103
|
+
"claude-opus-5": {
|
|
104
|
+
"effortLevels": [
|
|
105
|
+
"low",
|
|
106
|
+
"medium",
|
|
107
|
+
"high",
|
|
108
|
+
"xhigh",
|
|
109
|
+
"max"
|
|
110
|
+
],
|
|
111
|
+
"recommended": "high",
|
|
112
|
+
"modes": [
|
|
113
|
+
"auto"
|
|
114
|
+
],
|
|
115
|
+
"disallowThinkingDisabled": true
|
|
116
|
+
}
|
|
117
|
+
},
|
|
118
|
+
"effortRegexDefault": {
|
|
119
|
+
"pattern": "^(?:claude-)?(?:fable|mythos)(?:-|$)",
|
|
120
|
+
"effortLevels": [
|
|
121
|
+
"low",
|
|
122
|
+
"medium",
|
|
123
|
+
"high",
|
|
124
|
+
"xhigh",
|
|
125
|
+
"max"
|
|
126
|
+
],
|
|
127
|
+
"recommended": "high",
|
|
128
|
+
"modes": [
|
|
129
|
+
"auto"
|
|
130
|
+
],
|
|
131
|
+
"disallowThinkingDisabled": true
|
|
132
|
+
},
|
|
133
|
+
"tools": [
|
|
134
|
+
"Task",
|
|
135
|
+
"Bash",
|
|
136
|
+
"Glob",
|
|
137
|
+
"Grep",
|
|
138
|
+
"Read",
|
|
139
|
+
"Edit",
|
|
140
|
+
"Write",
|
|
141
|
+
"NotebookEdit",
|
|
142
|
+
"WebFetch",
|
|
143
|
+
"TaskCreate",
|
|
144
|
+
"TaskUpdate",
|
|
145
|
+
"TaskGet",
|
|
146
|
+
"TaskList",
|
|
147
|
+
"TaskStop",
|
|
148
|
+
"WebSearch",
|
|
149
|
+
"Skill",
|
|
150
|
+
"REPL",
|
|
151
|
+
"JavaScript",
|
|
152
|
+
"AskUserQuestion",
|
|
153
|
+
"ToolSearch"
|
|
154
|
+
],
|
|
155
|
+
"allowedTools": [
|
|
156
|
+
"Task",
|
|
157
|
+
"Bash",
|
|
158
|
+
"Glob",
|
|
159
|
+
"Grep",
|
|
160
|
+
"Read",
|
|
161
|
+
"Edit",
|
|
162
|
+
"Write",
|
|
163
|
+
"NotebookEdit",
|
|
164
|
+
"WebFetch",
|
|
165
|
+
"TaskCreate",
|
|
166
|
+
"TaskUpdate",
|
|
167
|
+
"TaskGet",
|
|
168
|
+
"TaskList",
|
|
169
|
+
"TaskStop",
|
|
170
|
+
"WebSearch",
|
|
171
|
+
"Skill",
|
|
172
|
+
"REPL",
|
|
173
|
+
"JavaScript",
|
|
174
|
+
"ToolSearch"
|
|
175
|
+
],
|
|
176
|
+
"env": {
|
|
177
|
+
"CLAUDE_CODE_IS_COWORK": "1",
|
|
178
|
+
"CLAUDE_CODE_ENTRYPOINT": "local-agent",
|
|
179
|
+
"CLAUDE_CODE_TAGS": "lam_session_type:chat",
|
|
180
|
+
"CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST": "1",
|
|
181
|
+
"CLAUDE_CODE_ENABLE_ASK_USER_QUESTION_TOOL": "true",
|
|
182
|
+
"CLAUDE_CODE_DISABLE_CRON": "1",
|
|
183
|
+
"CLAUDE_CODE_DISABLE_BACKGROUND_TASKS": "1",
|
|
184
|
+
"CLAUDE_CODE_DISABLE_AGENTS_FLEET": "1",
|
|
185
|
+
"CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT": "1",
|
|
186
|
+
"CLAUDE_CODE_ENABLE_TASKS": "true",
|
|
187
|
+
"CLAUDE_CODE_DISABLE_TERMINAL_TITLE": "1",
|
|
188
|
+
"ENABLE_PROMPT_CACHING_1H": "1",
|
|
189
|
+
"DISABLE_MICROCOMPACT": "1",
|
|
190
|
+
"MCP_CONNECTION_NONBLOCKING": "true",
|
|
191
|
+
"API_TIMEOUT_MS": "900000",
|
|
192
|
+
"CLAUDE_CODE_EMIT_TOOL_USE_SUMMARIES": "",
|
|
193
|
+
"CLAUDE_CODE_ENABLE_FINE_GRAINED_TOOL_STREAMING": "1",
|
|
194
|
+
"DISABLE_AUTOUPDATER": "1",
|
|
195
|
+
"MCP_TOOL_TIMEOUT": "180000",
|
|
196
|
+
"USE_LOCAL_OAUTH": "",
|
|
197
|
+
"USE_STAGING_OAUTH": "",
|
|
198
|
+
"CLAUDE_PREVIEW_CLASSIFIER_FLOOR": "1",
|
|
199
|
+
"CLAUDE_CODE_PROMPT_CACHE_TTL": "1h",
|
|
200
|
+
"CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL": "5m"
|
|
201
|
+
},
|
|
202
|
+
"promptTemplate": "prompts/desktop-1.18286.0/system-prompt-append.md",
|
|
203
|
+
"subagentAppend": "prompts/desktop-1.15200.0/subagent-append-vm.md",
|
|
204
|
+
"subagentAppendHostLoop": "prompts/desktop-1.32885.1/subagent-append-hl.md",
|
|
205
|
+
"hooks": {
|
|
206
|
+
"PreToolUse": [
|
|
207
|
+
{
|
|
208
|
+
"matcher": "Task",
|
|
209
|
+
"note": "blocks run_in_background ('Background agents disabled'); emits subagent_invoked telemetry",
|
|
210
|
+
"served": true
|
|
211
|
+
},
|
|
212
|
+
{
|
|
213
|
+
"matcher": "Skill",
|
|
214
|
+
"note": "emits skill_invoked telemetry (plugin name/source/id/marketplace); logs cowork_consolidate_memory_called; injects per-skill additionalContext",
|
|
215
|
+
"served": false
|
|
216
|
+
},
|
|
217
|
+
{
|
|
218
|
+
"matcher": "COWORK_FORCE_ASK_TOOLS|MCP_{CREATE,UPDATE,DELETE}_SCHEDULED_TASK|MCP_{START,STOP}_WATCHING",
|
|
219
|
+
"note": "permissionDecision:'ask' regardless of permission mode",
|
|
220
|
+
"served": false
|
|
221
|
+
},
|
|
222
|
+
{
|
|
223
|
+
"matcher": "mcp__.*",
|
|
224
|
+
"note": "remote-MCP deny hook (evaluateRemoteMcpDenyHook) -> decision:'block'",
|
|
225
|
+
"served": false
|
|
226
|
+
}
|
|
227
|
+
],
|
|
228
|
+
"PostToolUse": [
|
|
229
|
+
{
|
|
230
|
+
"matcher": "WebSearch",
|
|
231
|
+
"note": "seeds session.webFetchAllowedUrls from search results (ingestWebSearchResultForProvenance) - a WebSearch WIDENS the web_fetch allowlist",
|
|
232
|
+
"served": false
|
|
233
|
+
}
|
|
234
|
+
],
|
|
235
|
+
"UserPromptSubmit": [
|
|
236
|
+
{
|
|
237
|
+
"matcher": null,
|
|
238
|
+
"note": "expands a leading /slash command into hookSpecificOutput.additionalContext",
|
|
239
|
+
"served": false
|
|
240
|
+
}
|
|
241
|
+
]
|
|
242
|
+
},
|
|
243
|
+
"$comment": "Binary-verified Desktop->agent spawn contract, re-derived per release. spawn.env is GENERATED by deriveSpawnEnv() in src/sync/cowork-sync.ts (windowed enumeration of the asar env construction + gate/const value resolution); the scalar options, tools/allowedTools, and prompt-asset pointers are sentinel-guarded by checkSpawnContractFacts(). Do not hand-edit spawn.env — re-run sync.",
|
|
244
|
+
"$comment_handPinned": "Why the NON-env spawn fields stay hand-pinned: each is built in the asar as a non-literal expression (a session-path template, a session-type ternary, a const indirection, or a head+spread+tail array), so the windowed-enumeration generator that derives spawn.env cannot construct their VALUES without a full JS evaluator; instead each value was binary-verified once and is drift-guarded by a checkSpawnContractFacts() sentinel (cowork-sync.ts) that re-asserts the asar-side FACT at every sync. Scope caveat: the sentinels make DESKTOP-side drift loud; they do not validate this committed JSON itself — an erroneous hand-edit here is invisible to them.",
|
|
245
|
+
"$comment_configDirInGuest": "Hand-pinned: the asar builds it as a per-session path template (/sessions/${id}/mnt/.claude), not a constructable literal. Sentinel S1 pins the template shape.",
|
|
246
|
+
"$comment_settingSources": "Hand-pinned: sentinel S2 pins the settingSources:[\"user\"] literal.",
|
|
247
|
+
"$comment_permissionMode": "Hand-pinned: the asar computes it via a session ternary whose chat-session branch resolves to \"default\". Sentinel S3 pins the ternary shape.",
|
|
248
|
+
"$comment_maxThinkingTokens": "Hand-pinned: the asar reaches the value through const indirection. Sentinel S4 VALUE-pins the resolved const to 31999.",
|
|
249
|
+
"$comment_effortDefault": "Hand-pinned: sentinel S5 pins the .effort … :\"medium\" default.",
|
|
250
|
+
"$comment_tools": "Hand-pinned: the asar builds tools[] as head-list + Task-tools spread + session-type tail, not one literal. Sentinels S6 (head), S7 (the TaskCreate…TaskStop spread), S8 (tail-guard after ToolSearch) pin all three parts. As of desktop-1.21459.0 the asar head also carries an INERT `...CLAUDE_DESIGN_TOOLS` spread between Task and Bash that resolves to [] (deployment-gated off on first-party), so the rendered list — and this pin — stay 20 entries; S6b asserts it empty and fails loud if a build ever populates it.",
|
|
251
|
+
"$comment_allowedTools": "Hand-pinned: same head+spread+tail construction as tools[], minus AskUserQuestion (tools-only by design). Sentinels S9 (head) and S10 (the built-in→mcp__ boundary tail-guard) pin it.",
|
|
252
|
+
"$comment_promptTemplate": "Hand-pinned pointer to a RECONSTRUCTED asset (see $comment_prompts) — the generator cannot extract prose. Sentinel S15 pins the claude_code preset-append delivery site.",
|
|
253
|
+
"$comment_subagentAppend": "Hand-pinned pointer to a reconstructed asset (see $comment_prompts). Sentinel S16 pins the per-session appendSubagentSystemPrompt generator call shape.",
|
|
254
|
+
"$comment_subagentAppendHostLoop": "Hand-pinned pointer to a reconstructed (paraphrased) asset for the HOST-LOOP branch of the per-session sub-agent append (section key subagent_env_hl; selected purely on hostLoopMode). Backfilled only for release families whose hl text is binary-verified byte-identical (1.18286.2+). Sentinel: checkSubagentPromptFacts (two-branch fingerprint + substitution-value proofs). A hostloop run on a baseline lacking this pointer fails loud rather than falling back to the VM text. Repointed at 1.32885.1: the hl branch text changed at this release (fingerprint 8fe5583c5996776a -> 71e028bfa7ce596d, one appended sentence on shell-command cwd and non-mnt writes), so the 1.18286.2 asset is no longer faithful for this family. This pointer is hand-authored and is NOT re-derived by sync — it carries forward untouched, so it must be repointed by hand whenever the branch fingerprint moves.",
|
|
255
|
+
"$comment_notSet": "Deliberately NOT set: CLAUDE_CODE_USE_COWORK_PLUGINS (Desktop never sets it; would flip the agent to cowork_settings.json/cowork_plugins — asserted absent by the S17 negative invariant). Enumerated-but-not-pinned keys are enforced by SPAWN_ENV_ALLOWLIST in src/sync/cowork-sync.ts, each with a reason; categories: host-derived (CLAUDE_CONFIG_DIR, TZ, HOST_PLATFORM, OAUTH_TOKEN/BASE_URL/CUSTOM_HEADERS, account UUIDs, WORKSPACE_HOST_PATHS, OTEL), constructed-then-deleted (ANTHROPIC_API_KEY/AUTH_TOKEN via FnA), gate-conditional-off (MCP_CONNECT_TIMEOUT_MS, ENABLE_TOOL_SEARCH, SKIP_PRECOMPACT_LOAD), non-chat/project-session (BRIEF*, PROJECT*), user-settings (SUBAGENT_MODEL, AUTO_COMPACT_WINDOW, ...), and 3p-provider-only branches. The opaque ...g.env/...l session spreads are the known static-extraction blind spot (runtime lane backstop).",
|
|
256
|
+
"$comment_prompts": "Reconstructed cowork-specific sections, re-paraphrased from asar 1.18286.0 const aui (system prompt; RESTRUCTURED at this release — see the asset header) and 1.15200.0 generator CVr (subagent; verified unchanged in the 1.18286.0 asar, generator Zgn). Not the full base prompt (not cleanly extractable); generic refusal/safety policy elided. Delivered via --append-system-prompt (layered on the agent's built-in base prompt), NOT the initialize handshake; only the subagent append goes over initialize (appendSubagentSystemPrompt), gated on CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT.",
|
|
257
|
+
"$comment_hooks": "Binary-verified against app.asar 1.24012.9: the Cowork spawn's hooks object, identified by the env block immediately following it (CLAUDE_CODE_IS_COWORK=1, CLAUDE_CODE_ENTRYPOINT=local-agent). Recorded as a DRIFT TRIPWIRE, not an emulation source - `served` marks what this harness actually installs (PreToolUse:Task only). NOT installed at the Cowork spawn: SessionStart, SessionEnd, SubagentStop, PreCompact, Notification, Stop (their asar occurrences are an event-name validation list, a config-error hint, and a UI filter that skips SessionStart hook traffic - that filter implies plugin-supplied hooks.json events DO arrive, via the agent binary's own --plugin-dir channel, which is separate from this spawn bundle)."
|
|
258
|
+
},
|
|
259
|
+
"mountLayout": {
|
|
260
|
+
"sessionRoot": "/sessions/{sessionId}",
|
|
261
|
+
"cwd": "/sessions/{sessionId}",
|
|
262
|
+
"mntRoot": "/sessions/{sessionId}/mnt",
|
|
263
|
+
"mounts": [
|
|
264
|
+
{
|
|
265
|
+
"name": "uploads",
|
|
266
|
+
"mountPath": "uploads",
|
|
267
|
+
"mode": "r",
|
|
268
|
+
"purpose": "user-uploaded files (read-only — asar 'ro')"
|
|
269
|
+
},
|
|
270
|
+
{
|
|
271
|
+
"name": "projects",
|
|
272
|
+
"mountPath": ".projects/{projectId}",
|
|
273
|
+
"mode": "r",
|
|
274
|
+
"purpose": "RESERVED namespace (and the separate UUID project-sync feature) — NOT the work-folder path. From Desktop 1.14271.0 selected work folders mount at mnt/<collision-resolved-basename> (dynamic, derived per session by buildLaunchPlan; see MOUNT_BARE_NAME_MIN_VERSION). This decorative row is not consumed for binding (staged paths come from plan.mounts). MODE CORRECTED rw -> r on 2026-08-05, first-party from the asar's mount-set builder (VM-loop runs it at spawn; host-loop recomputes it per bash call): a project ATTACHMENT mounts at `.projects/<uuid>` with `mode:\"ro\"`, and the sibling `.claude/projects` is `\"ro\"` too. Neither passes through the delete-deny resolver, so a project mount is NOT writable-but-delete-denied — it is not writable at all, and therefore correctly outside deleteDeniedRootsFromPlan. checkMountModeFacts pins both. Older baselines still carry `rw` here and are deliberately NOT back-edited: below MOUNT_BARE_NAME_MIN_VERSION (1.14271.0) `.projects/<name>` WAS the connected-folder namespace, which is resolver-driven `rw` — so `rw` is correct there, and a blanket correction would have replaced a right fact with a wrong one. Between that boundary and 1.25927.0 the row is stale rather than meaningful (the namespace is reserved); those files are frozen per-release snapshots and the mode is consumed by nothing — only cwd/sessionRoot/mntRoot are read from mountLayout."
|
|
275
|
+
},
|
|
276
|
+
{
|
|
277
|
+
"name": "local-plugins",
|
|
278
|
+
"mountPath": ".local-plugins/marketplaces",
|
|
279
|
+
"mode": "r",
|
|
280
|
+
"purpose": "marketplace skills/plugins, runtime-discovered"
|
|
281
|
+
},
|
|
282
|
+
{
|
|
283
|
+
"name": "remote-plugins",
|
|
284
|
+
"mountPath": ".remote-plugins",
|
|
285
|
+
"mode": "r",
|
|
286
|
+
"purpose": "org-remote plugins, runtime-discovered"
|
|
287
|
+
},
|
|
288
|
+
{
|
|
289
|
+
"name": "outputs",
|
|
290
|
+
"mountPath": "outputs",
|
|
291
|
+
"mode": "rw",
|
|
292
|
+
"purpose": "session outputs/artifacts — delete denied by default (asar IX); rwd only when approved"
|
|
293
|
+
},
|
|
294
|
+
{
|
|
295
|
+
"name": "skills",
|
|
296
|
+
"mountPath": ".claude/skills",
|
|
297
|
+
"mode": "r",
|
|
298
|
+
"purpose": "personal/saved skill doc bodies (NOT plugin-bundled skills, which live under local-plugins/remote-plugins above) — real VM confirmed via systemd unit sessions-<name>-mnt-.claude-skills.mount in vm_bundles/claudevm.bundle/rootfs.img. Decorative row like 'projects' above (not consumed for binding — resolveMounts()'s mounts[] is destructured away at every call site); the harness reproduces this via CLAUDE_CONFIG_DIR staging (session.ts skill copy + stage.ts cpSync), not a plan.mounts bind — see hostloop-prompt.ts's asar-verified skills bullet."
|
|
299
|
+
}
|
|
300
|
+
]
|
|
301
|
+
},
|
|
302
|
+
"network": {
|
|
303
|
+
"mode": "gvisor",
|
|
304
|
+
"allowKind": "allowlist",
|
|
305
|
+
"allowDomains": [
|
|
306
|
+
"downloads.claude.ai",
|
|
307
|
+
"api.anthropic.com",
|
|
308
|
+
"a-cdn.anthropic.com",
|
|
309
|
+
"a-api.anthropic.com",
|
|
310
|
+
"assets.claude.ai",
|
|
311
|
+
"sentry.io",
|
|
312
|
+
"api-staging.anthropic.com",
|
|
313
|
+
"api.claude.ai",
|
|
314
|
+
"preview.claude.ai",
|
|
315
|
+
"www.anthropic.com",
|
|
316
|
+
"console.anthropic.com",
|
|
317
|
+
"support.anthropic.com",
|
|
318
|
+
"docs.anthropic.com",
|
|
319
|
+
"mcp-proxy.anthropic.com",
|
|
320
|
+
"pivot.claude.ai"
|
|
321
|
+
],
|
|
322
|
+
"$comment": "network.allowDomains is a PINNED, hand-curated list — `sync` carries it forward and never re-derives it. On the first-party deployment this harness models, the VM egress allowlist is NOT in the app bundle: the 1p deployment class returns `vmEgressPolicy(){return null}`, so `resolveVmAllowedDomains` falls through to the session's SERVER-DELIVERED `egressAllowedDomains`, and the only host the bundle contributes is the OTLP endpoint appended by the augmenter. The entries below are therefore a curated RECONSTRUCTION, not an extraction — and this field is the allowlist the harness ENFORCES (boundaryAllowList + the session egress plan). UNVERIFIED as VM egress, retained deliberately rather than pruned on a guess because the true 1p list cannot be read from the asar: www.anthropic.com, console.anthropic.com, support.anthropic.com, docs.anthropic.com — these read as Desktop UI/help links swept in by the predecessor's bundle-wide domain regex. Removing one is a deliberate, reviewed act. checkEgressContractFacts in src/sync/cowork-sync.ts fails closed if the construction that justifies pinning moves."
|
|
323
|
+
},
|
|
324
|
+
"bgEnvStrip": {
|
|
325
|
+
"knownVars": [
|
|
326
|
+
"CLAUDE_CODE_OAUTH_TOKEN",
|
|
327
|
+
"CLAUDE_CODE_SESSION_KIND",
|
|
328
|
+
"CLAUDE_CODE_SESSION_ID",
|
|
329
|
+
"CLAUDE_CODE_SESSION_NAME",
|
|
330
|
+
"CLAUDE_CODE_SESSION_LOG"
|
|
331
|
+
]
|
|
332
|
+
},
|
|
333
|
+
"$comment": "Platform baseline auto-derived by `cowork-harness sync` from a live Claude Desktop install + app.asar. VOLATILE per-release facts only. Regenerate per release; review the diff. Captured 2026-09-02 on macOS arm64.",
|
|
334
|
+
"capturedAt": "2026-09-02",
|
|
335
|
+
"platform": "darwin-arm64",
|
|
336
|
+
"settings": {
|
|
337
|
+
"autoMountFolders": {
|
|
338
|
+
"key": "autoMountFolders",
|
|
339
|
+
"default": false
|
|
340
|
+
},
|
|
341
|
+
"localAgentModeTrustedFolders": {
|
|
342
|
+
"key": "localAgentModeTrustedFolders",
|
|
343
|
+
"default": []
|
|
344
|
+
}
|
|
345
|
+
},
|
|
346
|
+
"provenance": {
|
|
347
|
+
"asarPath": "/Applications/Claude.app/Contents/Resources/app.asar",
|
|
348
|
+
"asarFingerprint": "87c18c1a17873ba6",
|
|
349
|
+
"gates": {
|
|
350
|
+
"$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift. The skill-family gates are pinned on the same principle but are NOT all dormant: 245679952 (suggestSkillsEnabled) and 1598976391 (proactiveSkillSuggestEnabled) ARE modeled — they gate the skills SDK-MCP tool surface. 3246569822 (canSaveSkill) is served ON/force by a server-side rollout (independent of Desktop version) and is deliberately NOT modeled: ON adds an mcp__cowork__save_skill tool and changes the skills system prompt, so this pin records a known fidelity gap rather than a modeled surface — see docs/fidelity-gaps.md. 1824824999 (canProposeSkills) is present-but-off, pinned so the same class of silent widening cannot land unnoticed. 1598976391 (proactiveSkillSuggestEnabled) flipped off/defaultValue -> ON/force by a SERVER-SIDE rollout observed 2026-08-04 — NOT a Desktop change: the gate id occurs exactly once in both the 1.24012.9 and 1.24012.11 asars and would read ON on .9 today, so this value is Desktop-version-INDEPENDENT despite living in a version-named file (same provenance class as canSaveSkill above). It is read from a SINGLE account's fcache; force rules are server-evaluated and can be segment-targeted, so whether the rollout is global is not determinable from anything on disk. Unlike canSaveSkill this one IS modeled (both branches exist), so the pin changes the emulated suggest_skills surface: proactive description + an optional trigger param + a chained empty-catalog note. Override per-session with skills.proactive_suggest_enabled. 4074604942 (1p-direct-mcp) was NEW in 1.24012.11 and DARK (absent from a standard fcache, hence its DARK_GATES entry); it arms a Desktop-side direct-MCP pool for MDM-managed 1P servers, inert for an unmanaged account, pinned as a sentinel only. Observed 2026-08-05 SERVED rather than absent (source \"force\", value false) — the rollout reached this account with the gate OFF, so nothing it arms is reachable and no modeled surface changes. Same provenance class as canSaveSkill: read from a SINGLE account's fcache, and force rules are server-evaluated and segment-targetable, so its DARK_GATES entry is retained deliberately — another account may still see it absent, and that must stay tolerated rather than hard-failing their sync.",
|
|
351
|
+
"emitToolUseSummaries:66187241": {
|
|
352
|
+
"on": false,
|
|
353
|
+
"source": "defaultValue",
|
|
354
|
+
"value": false
|
|
355
|
+
},
|
|
356
|
+
"autoMemoryStandardSessions:123929380": {
|
|
357
|
+
"on": false,
|
|
358
|
+
"source": "defaultValue",
|
|
359
|
+
"value": false
|
|
360
|
+
},
|
|
361
|
+
"subagentPromptServerOverride:124685897": {
|
|
362
|
+
"on": true,
|
|
363
|
+
"source": "defaultValue",
|
|
364
|
+
"value": true
|
|
365
|
+
},
|
|
366
|
+
"suggestSkillsEnabled:245679952": {
|
|
367
|
+
"on": true,
|
|
368
|
+
"source": "force",
|
|
369
|
+
"value": true
|
|
370
|
+
},
|
|
371
|
+
"skill-arg-elicitation:286376943": {
|
|
372
|
+
"on": true,
|
|
373
|
+
"source": "force",
|
|
374
|
+
"value": true
|
|
375
|
+
},
|
|
376
|
+
"mcpConnectionNonblockingOff:434204418": {
|
|
377
|
+
"on": false,
|
|
378
|
+
"source": "defaultValue",
|
|
379
|
+
"value": false
|
|
380
|
+
},
|
|
381
|
+
"bridgeSdkTransport:583857784": {
|
|
382
|
+
"on": true,
|
|
383
|
+
"source": "force",
|
|
384
|
+
"value": true,
|
|
385
|
+
"note": "— Cowork uses the SDK-based transport (control protocol), confirming the harness's sdkMcpServers/mcp_message path is the production transport."
|
|
386
|
+
},
|
|
387
|
+
"fineGrainedToolStreaming:714014285": {
|
|
388
|
+
"on": true,
|
|
389
|
+
"source": "force",
|
|
390
|
+
"value": true
|
|
391
|
+
},
|
|
392
|
+
"enableToolSearchAuto:1129419822": {
|
|
393
|
+
"on": false,
|
|
394
|
+
"source": "absent"
|
|
395
|
+
},
|
|
396
|
+
"hostLoop:1143815894": {
|
|
397
|
+
"on": true,
|
|
398
|
+
"source": "force",
|
|
399
|
+
"value": true
|
|
400
|
+
},
|
|
401
|
+
"scheduledTaskToolsApprovableByAutoMode:1447478638": {
|
|
402
|
+
"on": true,
|
|
403
|
+
"source": "force",
|
|
404
|
+
"value": true
|
|
405
|
+
},
|
|
406
|
+
"proactiveSkillSuggestEnabled:1598976391": {
|
|
407
|
+
"on": true,
|
|
408
|
+
"source": "force",
|
|
409
|
+
"value": true
|
|
410
|
+
},
|
|
411
|
+
"scheduledTaskSessionLimiter:1648655587": {
|
|
412
|
+
"on": true,
|
|
413
|
+
"source": "force",
|
|
414
|
+
"value": {
|
|
415
|
+
"global": 3,
|
|
416
|
+
"perTask": 1
|
|
417
|
+
},
|
|
418
|
+
"note": "SCHEDULED-TASK (cron) session limiter — NOT an in-conversation Task-tool cap (binary-verified 2026-07-04, asar 1.18286.0 class L9t [ScheduledTasks]). perTask=1: <=1 concurrent session PER SCHEDULED TASK; global=3: <=3 concurrent scheduled-task sessions globally (+_pendingTaskDispatches). Host-side SKIP (recordSkipAndEmit/PerTaskLimit|GlobalLimit — NOT queue/deny). Cowork imposes no cap on Task-tool sub-agent fan-out; the harness has no scheduled-task scheduler, so this gate has no applicable surface — pinned as a sync drift-sentinel only."
|
|
419
|
+
},
|
|
420
|
+
"memoryGuidelinesEnv:1696890383": {
|
|
421
|
+
"on": false,
|
|
422
|
+
"source": "defaultValue",
|
|
423
|
+
"value": false
|
|
424
|
+
},
|
|
425
|
+
"canProposeSkills:1824824999": {
|
|
426
|
+
"on": false,
|
|
427
|
+
"source": "defaultValue",
|
|
428
|
+
"value": false
|
|
429
|
+
},
|
|
430
|
+
"oauthScopesEnv:1936081873": {
|
|
431
|
+
"on": true,
|
|
432
|
+
"source": "force",
|
|
433
|
+
"value": true
|
|
434
|
+
},
|
|
435
|
+
"coworkRuntimeConfig:1978029737": {
|
|
436
|
+
"on": true,
|
|
437
|
+
"source": "experiment",
|
|
438
|
+
"value": {
|
|
439
|
+
"coworkNativeFilePreview": true,
|
|
440
|
+
"coworkWebFetchDedup": true,
|
|
441
|
+
"coworkWebFetchDedupMaxEntries": 100,
|
|
442
|
+
"coworkWebFetchDedupTtlMs": 3600000,
|
|
443
|
+
"coworkWebFetchPrompt": true,
|
|
444
|
+
"coworkWebFetchViaApi": true,
|
|
445
|
+
"pluginsFullSyncStalenessMs": 3600000,
|
|
446
|
+
"pluginsSyncIntervalMs": 1200000,
|
|
447
|
+
"sessionsBridgePollBlockMs": 30,
|
|
448
|
+
"skillsSyncIntervalMs": 1200000,
|
|
449
|
+
"workspaceBashWaitLonger": true
|
|
450
|
+
},
|
|
451
|
+
"note": "coworkWebFetchViaApi=true coworkWebFetchPrompt=true workspaceBashWaitLonger=true sessionsBridgePollBlockMs=30 — web_fetch is host/API-routed (POST /api/organizations/<org>/cowork/web_fetch), NOT container egress; gated by a separate web-fetch hostname allowlist + URL provenance."
|
|
452
|
+
},
|
|
453
|
+
"cicCanUseToolEnabled:2051942385": {
|
|
454
|
+
"on": true,
|
|
455
|
+
"source": "force",
|
|
456
|
+
"value": true
|
|
457
|
+
},
|
|
458
|
+
"cliPlugin:2307090146": {
|
|
459
|
+
"on": false,
|
|
460
|
+
"source": "defaultValue",
|
|
461
|
+
"value": false,
|
|
462
|
+
"note": "— the CLI-plugin credential broker is dark-launched off for standard interactive accounts (Ch23/L106)."
|
|
463
|
+
},
|
|
464
|
+
"pluginSyncSparkplug:2340532315": {
|
|
465
|
+
"on": true,
|
|
466
|
+
"source": "force",
|
|
467
|
+
"value": true,
|
|
468
|
+
"note": "— startup syncPlugins(); plugins load via --plugin-dir (registry inert in-VM)."
|
|
469
|
+
},
|
|
470
|
+
"skeletonHome:2614807392": {
|
|
471
|
+
"on": false,
|
|
472
|
+
"source": "absent"
|
|
473
|
+
},
|
|
474
|
+
"memoryExtraGuidelines:2860753854": {
|
|
475
|
+
"on": true,
|
|
476
|
+
"source": "defaultValue",
|
|
477
|
+
"value": "## Sensitive personal information\n\nDo not save the following to memory unless the user explicitly asks you to remember it:\n\n- Protected attributes: race, ethnicity, national origin, religion, age, sex, sexual orientation, gender identity, immigration status, disability, serious illness, union membership\n- Government identifiers: Social Security numbers, driver's license numbers, passport numbers, government ID numbers\n- Financial account details: credit card numbers, bank account numbers\n- Health information: medical conditions, diagnoses, lab results, mental health details, therapy or counseling\n- Home or personal mailing addresses (work addresses are fine)\n- Account passwords, secret tokens, or secret keys\n\nIf any of the above appears in conversation context, complete the task but do not persist it to a memory file. If the user explicitly says \"remember my address is X\", saving it is acceptable — they've given consent."
|
|
478
|
+
},
|
|
479
|
+
"coworkArtifacts:2940196192": {
|
|
480
|
+
"on": true,
|
|
481
|
+
"source": "force",
|
|
482
|
+
"value": true
|
|
483
|
+
},
|
|
484
|
+
"canSaveSkill:3246569822": {
|
|
485
|
+
"on": true,
|
|
486
|
+
"source": "force",
|
|
487
|
+
"value": true
|
|
488
|
+
},
|
|
489
|
+
"automode-permission-rubric:3424551112": {
|
|
490
|
+
"on": true,
|
|
491
|
+
"source": "force",
|
|
492
|
+
"value": true
|
|
493
|
+
},
|
|
494
|
+
"1p-direct-mcp:4074604942": {
|
|
495
|
+
"on": false,
|
|
496
|
+
"source": "force",
|
|
497
|
+
"value": false
|
|
498
|
+
},
|
|
499
|
+
"skipPrecompactLoad:4153934152": {
|
|
500
|
+
"on": false,
|
|
501
|
+
"source": "defaultValue",
|
|
502
|
+
"value": false
|
|
503
|
+
},
|
|
504
|
+
"autoModeOverridesAlwaysAllow:4200321681": {
|
|
505
|
+
"on": true,
|
|
506
|
+
"source": "force",
|
|
507
|
+
"value": true
|
|
508
|
+
}
|
|
509
|
+
},
|
|
510
|
+
"spawnEnvKeys": [
|
|
511
|
+
"ANTHROPIC_API_KEY",
|
|
512
|
+
"ANTHROPIC_AUTH_TOKEN",
|
|
513
|
+
"ANTHROPIC_BASE_URL",
|
|
514
|
+
"ANTHROPIC_CUSTOM_HEADERS",
|
|
515
|
+
"API_TIMEOUT_MS",
|
|
516
|
+
"CLAUDE_CODE_ACCOUNT_TAGGED_ID",
|
|
517
|
+
"CLAUDE_CODE_ACCOUNT_UUID",
|
|
518
|
+
"CLAUDE_CODE_ADDITIONAL_DIRECTORIES_CLAUDE_MD",
|
|
519
|
+
"CLAUDE_CODE_ATTRIBUTION_HEADER",
|
|
520
|
+
"CLAUDE_CODE_AUTO_COMPACT_WINDOW",
|
|
521
|
+
"CLAUDE_CODE_BRIEF",
|
|
522
|
+
"CLAUDE_CODE_BRIEF_UPLOAD",
|
|
523
|
+
"CLAUDE_CODE_COWORK_FRAME_ARTIFACTS",
|
|
524
|
+
"CLAUDE_CODE_DIAGNOSTICS_FILE",
|
|
525
|
+
"CLAUDE_CODE_DISABLE_AGENTS_FLEET",
|
|
526
|
+
"CLAUDE_CODE_DISABLE_BACKGROUND_TASKS",
|
|
527
|
+
"CLAUDE_CODE_DISABLE_BUNDLED_SKILLS",
|
|
528
|
+
"CLAUDE_CODE_DISABLE_CRON",
|
|
529
|
+
"CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS",
|
|
530
|
+
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC",
|
|
531
|
+
"CLAUDE_CODE_DISABLE_OFFICIAL_MARKETPLACE_AUTOINSTALL",
|
|
532
|
+
"CLAUDE_CODE_DISABLE_REFUSAL_FALLBACK",
|
|
533
|
+
"CLAUDE_CODE_DISABLE_TERMINAL_TITLE",
|
|
534
|
+
"CLAUDE_CODE_EMIT_TOOL_USE_SUMMARIES",
|
|
535
|
+
"CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT",
|
|
536
|
+
"CLAUDE_CODE_ENABLE_ASK_USER_QUESTION_TOOL",
|
|
537
|
+
"CLAUDE_CODE_ENABLE_AUTO_MODE",
|
|
538
|
+
"CLAUDE_CODE_ENABLE_FINE_GRAINED_TOOL_STREAMING",
|
|
539
|
+
"CLAUDE_CODE_ENABLE_TASKS",
|
|
540
|
+
"CLAUDE_CODE_ENTRYPOINT",
|
|
541
|
+
"CLAUDE_CODE_HOST_AUTH_ENV_VAR",
|
|
542
|
+
"CLAUDE_CODE_HOST_PLATFORM",
|
|
543
|
+
"CLAUDE_CODE_IS_COWORK",
|
|
544
|
+
"CLAUDE_CODE_OAUTH_SCOPES",
|
|
545
|
+
"CLAUDE_CODE_OAUTH_TOKEN",
|
|
546
|
+
"CLAUDE_CODE_ORGANIZATION_UUID",
|
|
547
|
+
"CLAUDE_CODE_PROMPT_CACHE_TTL",
|
|
548
|
+
"CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST",
|
|
549
|
+
"CLAUDE_CODE_RATE_LIMIT_TIER",
|
|
550
|
+
"CLAUDE_CODE_SKIP_PRECOMPACT_LOAD",
|
|
551
|
+
"CLAUDE_CODE_SUBAGENT_MODEL",
|
|
552
|
+
"CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL",
|
|
553
|
+
"CLAUDE_CODE_SUBSCRIPTION_TYPE",
|
|
554
|
+
"CLAUDE_CODE_TAGS",
|
|
555
|
+
"CLAUDE_CODE_USER_EMAIL",
|
|
556
|
+
"CLAUDE_CODE_WORKSPACE_HOST_PATHS",
|
|
557
|
+
"CLAUDE_CONFIG_DIR",
|
|
558
|
+
"CLAUDE_PREVIEW_CLASSIFIER_FLOOR",
|
|
559
|
+
"CLAUDE_PROJECT_TOOL",
|
|
560
|
+
"CLAUDE_PROJECT_UUID",
|
|
561
|
+
"DISABLE_AUTOUPDATER",
|
|
562
|
+
"DISABLE_BRIEF_MODE_STOP_HOOK",
|
|
563
|
+
"DISABLE_ERROR_REPORTING",
|
|
564
|
+
"DISABLE_FEEDBACK_COMMAND",
|
|
565
|
+
"DISABLE_GROWTHBOOK",
|
|
566
|
+
"DISABLE_MICROCOMPACT",
|
|
567
|
+
"DISABLE_TELEMETRY",
|
|
568
|
+
"ENABLE_PROMPT_CACHING_1H",
|
|
569
|
+
"ENABLE_TOOL_SEARCH",
|
|
570
|
+
"MCP_CONNECTION_NONBLOCKING",
|
|
571
|
+
"MCP_CONNECT_TIMEOUT_MS",
|
|
572
|
+
"MCP_TOOL_TIMEOUT",
|
|
573
|
+
"TZ",
|
|
574
|
+
"USE_LOCAL_OAUTH",
|
|
575
|
+
"USE_STAGING_OAUTH"
|
|
576
|
+
],
|
|
577
|
+
"spawnEnvSpreadCount": 32,
|
|
578
|
+
"fcache": {
|
|
579
|
+
"content16": "7cba90e619996290",
|
|
580
|
+
"embeddedTimestamp": 1788347009597,
|
|
581
|
+
"featureCount": 316
|
|
582
|
+
},
|
|
583
|
+
"asarGateIds": [
|
|
584
|
+
"17519066",
|
|
585
|
+
"36693946",
|
|
586
|
+
"40173473",
|
|
587
|
+
"49458538",
|
|
588
|
+
"66187241",
|
|
589
|
+
"98041341",
|
|
590
|
+
"108465228",
|
|
591
|
+
"123929380",
|
|
592
|
+
"124685897",
|
|
593
|
+
"133902057",
|
|
594
|
+
"144158705",
|
|
595
|
+
"162211072",
|
|
596
|
+
"180602792",
|
|
597
|
+
"206750215",
|
|
598
|
+
"227459766",
|
|
599
|
+
"245679952",
|
|
600
|
+
"254738541",
|
|
601
|
+
"262787483",
|
|
602
|
+
"278625510",
|
|
603
|
+
"286376943",
|
|
604
|
+
"291584251",
|
|
605
|
+
"304458538",
|
|
606
|
+
"326791966",
|
|
607
|
+
"371539023",
|
|
608
|
+
"397125142",
|
|
609
|
+
"416245092",
|
|
610
|
+
"434204418",
|
|
611
|
+
"451382573",
|
|
612
|
+
"476513332",
|
|
613
|
+
"505512513",
|
|
614
|
+
"541746109",
|
|
615
|
+
"552157343",
|
|
616
|
+
"554317356",
|
|
617
|
+
"574905726",
|
|
618
|
+
"581786799",
|
|
619
|
+
"583857784",
|
|
620
|
+
"607406988",
|
|
621
|
+
"629684104",
|
|
622
|
+
"642265585",
|
|
623
|
+
"657187776",
|
|
624
|
+
"700397605",
|
|
625
|
+
"714014285",
|
|
626
|
+
"717163759",
|
|
627
|
+
"720735283",
|
|
628
|
+
"721728391",
|
|
629
|
+
"732695530",
|
|
630
|
+
"733405693",
|
|
631
|
+
"743194442",
|
|
632
|
+
"748063099",
|
|
633
|
+
"751369921",
|
|
634
|
+
"762798616",
|
|
635
|
+
"763725229",
|
|
636
|
+
"769234850",
|
|
637
|
+
"790863764",
|
|
638
|
+
"816856638",
|
|
639
|
+
"822923030",
|
|
640
|
+
"850702611",
|
|
641
|
+
"873030668",
|
|
642
|
+
"879583975",
|
|
643
|
+
"884132720",
|
|
644
|
+
"885442848",
|
|
645
|
+
"919579692",
|
|
646
|
+
"922442190",
|
|
647
|
+
"939257113",
|
|
648
|
+
"942840715",
|
|
649
|
+
"954625922",
|
|
650
|
+
"959099749",
|
|
651
|
+
"975112542",
|
|
652
|
+
"976614668",
|
|
653
|
+
"982691970",
|
|
654
|
+
"986399546",
|
|
655
|
+
"999999999",
|
|
656
|
+
"1004628546",
|
|
657
|
+
"1032963206",
|
|
658
|
+
"1061122496",
|
|
659
|
+
"1101873029",
|
|
660
|
+
"1109029378",
|
|
661
|
+
"1111144847",
|
|
662
|
+
"1126577245",
|
|
663
|
+
"1129419822",
|
|
664
|
+
"1131777852",
|
|
665
|
+
"1143815894",
|
|
666
|
+
"1183214304",
|
|
667
|
+
"1197768857",
|
|
668
|
+
"1215221242",
|
|
669
|
+
"1263782781",
|
|
670
|
+
"1265511872",
|
|
671
|
+
"1278239935",
|
|
672
|
+
"1284392461",
|
|
673
|
+
"1291166712",
|
|
674
|
+
"1294710626",
|
|
675
|
+
"1294982044",
|
|
676
|
+
"1310358601",
|
|
677
|
+
"1315974108",
|
|
678
|
+
"1323782925",
|
|
679
|
+
"1340622498",
|
|
680
|
+
"1346958739",
|
|
681
|
+
"1403324732",
|
|
682
|
+
"1410651677",
|
|
683
|
+
"1412563253",
|
|
684
|
+
"1434290056",
|
|
685
|
+
"1441987769",
|
|
686
|
+
"1447478638",
|
|
687
|
+
"1477483922",
|
|
688
|
+
"1480778051",
|
|
689
|
+
"1544796833",
|
|
690
|
+
"1549258603",
|
|
691
|
+
"1569828280",
|
|
692
|
+
"1588340729",
|
|
693
|
+
"1598976391",
|
|
694
|
+
"1629866860",
|
|
695
|
+
"1648655587",
|
|
696
|
+
"1677081600",
|
|
697
|
+
"1695017395",
|
|
698
|
+
"1696890383",
|
|
699
|
+
"1703762832",
|
|
700
|
+
"1707927936",
|
|
701
|
+
"1710585411",
|
|
702
|
+
"1743527783",
|
|
703
|
+
"1748356779",
|
|
704
|
+
"1754160972",
|
|
705
|
+
"1824824999",
|
|
706
|
+
"1825995196",
|
|
707
|
+
"1836754949",
|
|
708
|
+
"1868684740",
|
|
709
|
+
"1875770773",
|
|
710
|
+
"1893165035",
|
|
711
|
+
"1904135164",
|
|
712
|
+
"1915174500",
|
|
713
|
+
"1924247864",
|
|
714
|
+
"1928275548",
|
|
715
|
+
"1936081873",
|
|
716
|
+
"1942337209",
|
|
717
|
+
"1942781881",
|
|
718
|
+
"1947305033",
|
|
719
|
+
"1972091654",
|
|
720
|
+
"1978029737",
|
|
721
|
+
"1992087837",
|
|
722
|
+
"2004571505",
|
|
723
|
+
"2016258596",
|
|
724
|
+
"2023768496",
|
|
725
|
+
"2039376689",
|
|
726
|
+
"2048589233",
|
|
727
|
+
"2049450122",
|
|
728
|
+
"2051751800",
|
|
729
|
+
"2051942385",
|
|
730
|
+
"2062156710",
|
|
731
|
+
"2067027393",
|
|
732
|
+
"2099281725",
|
|
733
|
+
"2114777685",
|
|
734
|
+
"2115990222",
|
|
735
|
+
"2129861473",
|
|
736
|
+
"2140326016",
|
|
737
|
+
"2143883161",
|
|
738
|
+
"2192324205",
|
|
739
|
+
"2214981414",
|
|
740
|
+
"2216414644",
|
|
741
|
+
"2216766246",
|
|
742
|
+
"2216901299",
|
|
743
|
+
"2220415149",
|
|
744
|
+
"2229805612",
|
|
745
|
+
"2294160313",
|
|
746
|
+
"2307090146",
|
|
747
|
+
"2309422447",
|
|
748
|
+
"2339084909",
|
|
749
|
+
"2340532315",
|
|
750
|
+
"2345107588",
|
|
751
|
+
"2345515473",
|
|
752
|
+
"2358734848",
|
|
753
|
+
"2369776764",
|
|
754
|
+
"2392971184",
|
|
755
|
+
"2393677837",
|
|
756
|
+
"2427043945",
|
|
757
|
+
"2431502897",
|
|
758
|
+
"2464731750",
|
|
759
|
+
"2486083521",
|
|
760
|
+
"2529235968",
|
|
761
|
+
"2547348043",
|
|
762
|
+
"2576868839",
|
|
763
|
+
"2605355193",
|
|
764
|
+
"2614807392",
|
|
765
|
+
"2654621331",
|
|
766
|
+
"2685067074",
|
|
767
|
+
"2688060585",
|
|
768
|
+
"2719700143",
|
|
769
|
+
"2720310975",
|
|
770
|
+
"2722545484",
|
|
771
|
+
"2724639973",
|
|
772
|
+
"2725876754",
|
|
773
|
+
"2726556121",
|
|
774
|
+
"2745857735",
|
|
775
|
+
"2768844978",
|
|
776
|
+
"2795002549",
|
|
777
|
+
"2795595714",
|
|
778
|
+
"2800354941",
|
|
779
|
+
"2806360886",
|
|
780
|
+
"2833632524",
|
|
781
|
+
"2848557028",
|
|
782
|
+
"2857785401",
|
|
783
|
+
"2860753854",
|
|
784
|
+
"2864556627",
|
|
785
|
+
"2893011886",
|
|
786
|
+
"2895944283",
|
|
787
|
+
"2906430762",
|
|
788
|
+
"2938421209",
|
|
789
|
+
"2940196192",
|
|
790
|
+
"2961849615",
|
|
791
|
+
"2973881027",
|
|
792
|
+
"2979038612",
|
|
793
|
+
"3007887412",
|
|
794
|
+
"3018088575",
|
|
795
|
+
"3023518717",
|
|
796
|
+
"3045399524",
|
|
797
|
+
"3046457088",
|
|
798
|
+
"3046702961",
|
|
799
|
+
"3070110303",
|
|
800
|
+
"3089387226",
|
|
801
|
+
"3093186863",
|
|
802
|
+
"3123045134",
|
|
803
|
+
"3142047527",
|
|
804
|
+
"3150971238",
|
|
805
|
+
"3163246478",
|
|
806
|
+
"3183093548",
|
|
807
|
+
"3229517805",
|
|
808
|
+
"3246569822",
|
|
809
|
+
"3269331205",
|
|
810
|
+
"3300773012",
|
|
811
|
+
"3302457740",
|
|
812
|
+
"3326082604",
|
|
813
|
+
"3353525254",
|
|
814
|
+
"3356268835",
|
|
815
|
+
"3366735351",
|
|
816
|
+
"3368286709",
|
|
817
|
+
"3371831021",
|
|
818
|
+
"3377630395",
|
|
819
|
+
"3414805749",
|
|
820
|
+
"3424551112",
|
|
821
|
+
"3431784271",
|
|
822
|
+
"3436441689",
|
|
823
|
+
"3444158716",
|
|
824
|
+
"3448679706",
|
|
825
|
+
"3491600236",
|
|
826
|
+
"3516166472",
|
|
827
|
+
"3531779070",
|
|
828
|
+
"3547093683",
|
|
829
|
+
"3555657854",
|
|
830
|
+
"3558849738",
|
|
831
|
+
"3559681707",
|
|
832
|
+
"3572572142",
|
|
833
|
+
"3577536076",
|
|
834
|
+
"3586389629",
|
|
835
|
+
"3602629573",
|
|
836
|
+
"3633961296",
|
|
837
|
+
"3640318556",
|
|
838
|
+
"3646818354",
|
|
839
|
+
"3671534883",
|
|
840
|
+
"3691521536",
|
|
841
|
+
"3705360580",
|
|
842
|
+
"3723845789",
|
|
843
|
+
"3728132896",
|
|
844
|
+
"3758515526",
|
|
845
|
+
"3764441751",
|
|
846
|
+
"3778159589",
|
|
847
|
+
"3796647113",
|
|
848
|
+
"3807767338",
|
|
849
|
+
"3920548810",
|
|
850
|
+
"3927880029",
|
|
851
|
+
"3946462706",
|
|
852
|
+
"3961433847",
|
|
853
|
+
"3976799455",
|
|
854
|
+
"3982397363",
|
|
855
|
+
"3990395613",
|
|
856
|
+
"4034153053",
|
|
857
|
+
"4055864154",
|
|
858
|
+
"4066504968",
|
|
859
|
+
"4074604942",
|
|
860
|
+
"4085357330",
|
|
861
|
+
"4108768567",
|
|
862
|
+
"4114957886",
|
|
863
|
+
"4116586025",
|
|
864
|
+
"4141490266",
|
|
865
|
+
"4153934152",
|
|
866
|
+
"4156472024",
|
|
867
|
+
"4160352601",
|
|
868
|
+
"4185841952",
|
|
869
|
+
"4200321681",
|
|
870
|
+
"4202409342",
|
|
871
|
+
"4217215889",
|
|
872
|
+
"4272200640",
|
|
873
|
+
"4282876673",
|
|
874
|
+
"4293378213"
|
|
875
|
+
]
|
|
876
|
+
},
|
|
877
|
+
"requireFullVmSandbox": null
|
|
878
|
+
}
|
package/docs/ci.md
CHANGED
|
@@ -74,7 +74,7 @@ jobs:
|
|
|
74
74
|
- uses: actions/checkout@v4
|
|
75
75
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
76
76
|
run: |
|
|
77
|
-
V=2.1.
|
|
77
|
+
V=2.1.255 # match your scenario's pinned baseline's agentVersion
|
|
78
78
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here,
|
|
79
79
|
# or read it with jq if you vendor the baseline. An unverified download is an unverified
|
|
80
80
|
# agent: this step FAILS rather than staging one, which is the point of calling it verified.
|
|
@@ -91,7 +91,7 @@ jobs:
|
|
|
91
91
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
92
92
|
```
|
|
93
93
|
|
|
94
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.2.
|
|
94
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^3` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^3.2.1` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
95
95
|
|
|
96
96
|
CI uses `ANTHROPIC_API_KEY` specifically because there's no interactive browser available to run
|
|
97
97
|
`claude setup-token`'s OAuth flow in a GitHub Actions runner; locally, the OAuth token is preferred because
|
package/docs/cli.md
CHANGED
|
@@ -18,7 +18,7 @@ companion skill, CI). This page is the CLI one.
|
|
|
18
18
|
**Install from npm:**
|
|
19
19
|
|
|
20
20
|
```bash
|
|
21
|
-
npm install -g "cowork-harness@^3.2.
|
|
21
|
+
npm install -g "cowork-harness@^3.2.1" # puts the `cowork-harness` command on your PATH
|
|
22
22
|
```
|
|
23
23
|
|
|
24
24
|
**Or build from source:**
|
|
@@ -38,7 +38,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
38
38
|
|
|
39
39
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
40
40
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
41
|
-
> From a global install (`npm i -g "cowork-harness@^3.2.
|
|
41
|
+
> From a global install (`npm i -g "cowork-harness@^3.2.1"`), point at the package root instead:
|
|
42
42
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
43
43
|
> (or copy the cassette into your own project and pass that path).
|
|
44
44
|
|
|
@@ -48,7 +48,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
48
48
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
49
49
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
50
50
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
51
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.2.
|
|
51
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^3.2.1"`.
|
|
52
52
|
|
|
53
53
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
54
54
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
package/docs/companion-skill.md
CHANGED
|
@@ -28,7 +28,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
28
28
|
claude plugin install cowork-harness@cowork-harness
|
|
29
29
|
```
|
|
30
30
|
|
|
31
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.2.
|
|
31
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^3.2.1"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
32
32
|
|
|
33
33
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
34
34
|
|
|
@@ -38,7 +38,7 @@ npx skills add yaniv-golan/cowork-harness --skill cowork-harness
|
|
|
38
38
|
|
|
39
39
|
(Working *inside* this repo, the skill auto-loads as a project skill — no install needed.)
|
|
40
40
|
|
|
41
|
-
| What ships | npm global (`npm install -g "cowork-harness@^3.2.
|
|
41
|
+
| What ships | npm global (`npm install -g "cowork-harness@^3.2.1"`) | Source checkout (`git clone` + `npm ci`) |
|
|
42
42
|
|---|---|---|
|
|
43
43
|
| CLI, `scenario.py` + assertion keys (enough for `lint` in CI) | ✓ | ✓ |
|
|
44
44
|
| `SKILL.md`, all of `docs/`, `SPEC.md`/`DESIGN.md`/`AGENTS.md` | ✓ | ✓ |
|
|
@@ -53,6 +53,6 @@ global install puts nothing in your working directory. The matrix, answer-policy
|
|
|
53
53
|
ones that still need a source checkout. (The marketplace
|
|
54
54
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
55
55
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
56
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^3.2.
|
|
56
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^3.2.1"` — see
|
|
57
57
|
[above](#install) — which pulls the same npm package as the global-install row.)
|
|
58
58
|
|
package/docs/maintenance.md
CHANGED
|
@@ -111,7 +111,7 @@ Another runtime knob in the same family: `COWORK_HARNESS_RESOURCE_INTERVAL_MS` s
|
|
|
111
111
|
Old staged binaries are re-downloadable from Anthropic's own release channel. For the **container/microvm** tiers the harness needs the **Linux/arm64 ELF**, so download it directly and point the resolver at it:
|
|
112
112
|
|
|
113
113
|
```bash
|
|
114
|
-
V=2.1.
|
|
114
|
+
V=2.1.255 # your baseline's agentVersion (read it from baselines/desktop-<latest>.json)
|
|
115
115
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "claude-$V"
|
|
116
116
|
# verify against the committed baseline sha256 (== manifest platforms["linux-arm64"].checksum):
|
|
117
117
|
shasum -a 256 "claude-$V"
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.2.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^3.2.1"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
|
@@ -139,7 +139,7 @@
|
|
|
139
139
|
},
|
|
140
140
|
"scenarioSource": "../scenarios/example-pdf-skill.yaml",
|
|
141
141
|
"fingerprint": {
|
|
142
|
-
"baseline": "1.40609.
|
|
142
|
+
"baseline": "1.40609.1",
|
|
143
143
|
"hashFormat": "jcs1",
|
|
144
144
|
"skillHash": "b760ea90682777367fbd44d866ea66d316452da1abb18bf8cb61f1cec8e67806",
|
|
145
145
|
"promptAssetsHash": "491afe2862dc67ea",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "3.2.
|
|
3
|
+
"version": "3.2.1",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|