cowork-harness 1.3.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.3.0
7
- tracks-harness: cowork-harness 1.3.0 (baseline desktop-1.21459.0)
6
+ version: 1.4.0
7
+ tracks-harness: cowork-harness 1.4.0 (baseline desktop-1.22209.3)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.3.0` (baseline
26
- > `desktop-1.21459.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.4.0` (baseline
26
+ > `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.3.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.3.0"`. **Pin `@>=1.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.4.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.4.0"`. **Pin `@>=1.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
- What the ≥ 1.3.0 floor gates, by release:
44
+ What the ≥ 1.4.0 floor gates, by release:
45
45
 
46
46
  - **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
47
47
  - **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
@@ -55,6 +55,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
55
55
  - **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
56
56
  - **1.2.0:** three new assertion keys — `no_lost_write_back: true` (the write-back detector above wired as a per-scenario gate over the run's authored files, live-only) and the regex siblings `tool_result_matches`/`tool_result_not_matches` (case-insensitive per-result, for an error-signature *family* a literal substring can't express). **`microvm` outputs are now observable** — its session tree is snapshotted from the VM into the run dir, so `file_exists`/`artifact_json`/`user_visible_artifact`/`no_unexpected_files`/`input_unmodified`/`no_lost_write_back`/`semantic_matches` all work there, no longer `container`/`hostloop`-only. `status <dir>` also resolves the newest session under a `--run-dir` root; the completion footer prints a `→ result: …/result.json` pointer; the `on_unanswered=fail` error also points to `on_unanswered: llm`. `analyze-skill` hardening: a phantom `<script>` prose block no longer sinks a real verdict, a delete/remove flow claiming success classifies as lost (error) not just suspect, and write-back detection widened (optional-call `?.` spellings, member-spelled/aliased `fetch`/`sendBeacon`, axios instance/config forms).
57
57
  - **1.3.0:** run-identity for the iterate-across-fixes loop — `skill`/`run` take `--label <tag>` (a generation tag surfaced in `result.json` `runLabel`, the run-index row, `inspect`, and `status.json`) alongside an auto-recorded `skillCommit`, layered on the **authoritative** content-exact `fingerprint.skillHash` a harvest step should pair critiques by (`inspect`/index surface a short `skillHash` prefix). `trace --full-results` captures the full input+result of **every** tool call — successful ones too, not just errors — so an external grader can ground a self-critique finding against the call it cites. `verify-run` now **warns** on skillHash drift for answer-less scenarios (was silent unless the scenario declared scripted `answers`). Plus `skill --allow-missing-capability` (the open-ended-run opt-out for a capability FALSE-NEGATIVE on the lean `core` image), a new **warn-severity `ended_with_question`** verdict signal (the agent's final answer contains a question and the run wrote no `outputs/` deliverable — the lenient sibling of the strict `stalled`), and LLM-decider `OTHER:` free-text answers now marked `[via Other free-text]` in gate provenance.
58
+ - **1.4.0:** platform baseline synced to Desktop **1.22209.3** (agent 2.1.215; no prompt/spawn/egress drift), and **`coworkWebFetchDedup` is now enacted** on the hostloop `web_fetch` path — a repeat fetch of the same normalized URL within the TTL (15 min, cap 100) makes **no network request** and returns a marker to re-use the earlier result, with no egress event. Baseline-gated (Desktop ≥ 1.22209.3), keyed under the request + terminal `destination_url`, never caching errors/empty; a hit is assertable via `tool_result_contains: "Already fetched"`.
58
59
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
59
60
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
60
61
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.3.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.4.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.3.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.4.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -32,7 +32,7 @@ jobs:
32
32
  - uses: actions/checkout@v4
33
33
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
34
34
  run: |
35
- V=2.1.209 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.215 # match your scenario's pinned baseline's agentVersion
36
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
37
37
  chmod +x "$RUNNER_TEMP/claude-$V"
38
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.3.0"
60
+ - run: npm i -g "cowork-harness@>=1.4.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.3.0"
200
+ - run: npm i -g "cowork-harness@>=1.4.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.3.0"
229
+ run: npm i -g "cowork-harness@>=1.4.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.3.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.4.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.3.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.4.0`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,26 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.4.0] — 2026-07-19
10
+
11
+ ### Added
12
+
13
+ - **`coworkWebFetchDedup` enacted (hostloop `web_fetch`).** Real Cowork keeps a per-session negative-work
14
+ cache: a repeat `web_fetch` of the same normalized URL within a TTL (default 15 min; cap 100; FIFO
15
+ eviction; a hit does not refresh recency) makes **no network request** and returns a marker telling the
16
+ model to re-use the earlier result. The harness now reproduces this on the host-API (`coworkWebFetchViaApi`)
17
+ path — **baseline-gated** (only when the resolved baseline's `coworkWebFetchDedup` gate is on, i.e. Desktop
18
+ ≥ 1.22209.3), keyed under both the request URL and the terminal `destination_url`, never caching errors /
19
+ empty / non-2xx responses, and emitting **no egress event** on a hit (matching production's zero-network
20
+ dedup). A hit is observable via the marker text (`tool_result_contains: "Already fetched"`).
21
+
22
+ ### Changed
23
+
24
+ - **Platform baseline synced to Desktop 1.22209.3** (agent `2.1.215`). No prompt / spawn-env / egress-allowlist
25
+ drift vs 1.21459.0; the sync captured the new `coworkWebFetchDedup` runtime config (enacted above) plus a
26
+ few new (off) GrowthBook gates. The skill/README/reference version floors and agent-binary pins track the
27
+ new baseline.
28
+
9
29
  ## [1.3.0] — 2026-07-19
10
30
 
11
31
  ### Added
package/README.md CHANGED
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
91
91
 
92
92
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
93
93
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
94
- > From a global install (`npm i -g "cowork-harness@>=1.3.0"`), point at the package root instead:
94
+ > From a global install (`npm i -g "cowork-harness@>=1.4.0"`), point at the package root instead:
95
95
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
96
96
  > (or copy the cassette into your own project and pass that path).
97
97
 
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
101
101
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
102
102
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
103
103
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
104
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.3.0"`.
104
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.4.0"`.
105
105
 
106
106
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
107
107
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
126
126
  claude plugin install cowork-harness@cowork-harness
127
127
  ```
128
128
 
129
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.3.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
129
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.4.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
130
130
 
131
131
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
132
132
 
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
147
147
  To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
148
148
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
149
149
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
150
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.3.0"` — see
150
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.4.0"` — see
151
151
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
152
152
 
153
153
  ### Prerequisites for anything above `protocol` fidelity
@@ -659,7 +659,7 @@ jobs:
659
659
  - uses: actions/checkout@v4
660
660
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
661
661
  run: |
662
- V=2.1.209 # match your scenario's pinned baseline's agentVersion
662
+ V=2.1.215 # match your scenario's pinned baseline's agentVersion
663
663
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
664
664
  chmod +x "$RUNNER_TEMP/claude-$V"
665
665
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
@@ -670,7 +670,7 @@ jobs:
670
670
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
671
671
  ```
672
672
 
673
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.3.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
673
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.4.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
674
674
 
675
675
  The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
676
676
 
@@ -831,6 +831,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
831
831
  ## Status
832
832
 
833
833
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
834
- **`desktop-1.21459.0`**. Release-by-release verification notes (what was re-verified against
834
+ **`desktop-1.22209.3`**. Release-by-release verification notes (what was re-verified against
835
835
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
836
836
  this section used to duplicate lives in the sections above.
package/SPEC.md CHANGED
@@ -309,7 +309,7 @@ Payload sits under an **inner** `response`. Missing the nesting ⇒ `ZodError: e
309
309
  > Channels 2-3 = CLI `--mcp-config` / `.mcp.json` servers (honored in plain cowork mode; dropped only in
310
310
  > hermetic mode). Details per channel below.
311
311
 
312
- 1. **SDK servers over the control protocol** — declared via `sdkMcpServers` in `initialize`; tool calls tunnel as `mcp_message`. This is how the **desktop host** bridges its own servers (incl. `claude_desktop_config.json` `mcpServers`, spawned host-side with full host env) and how the harness delivers the workspace shell. The workspace handler (`src/hostloop/workspace-handler.ts`) implements `initialize`/`tools/list`/`tools/call`; `bash`→`docker exec -w <mntRoot> <container> sh -c <cmd>` (container-egress-gated). **CB-8:** `makeWorkspaceHandler` accepts an `onInfraError?: (message: string) => void` callback at parameter position 6 (after `onEgress`); on infrastructure errors (ETIMEDOUT / killed / no code+stdout+stderr) the handler calls `onInfraError?.(e.message)`, returns a textResult with `"[infrastructure error: …]"`, and `spawnHostLoop` wires `onInfraError` to append `{type:"infra_error", ts, message}` to `events.jsonl`. **`web_fetch` is NOT container-egress-gated:** real Cowork routes it through the host API (gate `1978029737` `coworkWebFetchViaApi:true` → `POST /api/organizations/<org>/cowork/web_fetch`), gated by a **separate web-fetch hostname allowlist** (`getWebFetchAllowedUrls`, `*`=unrestricted) + a **URL-provenance** rule (URL must have appeared in a prior message/result). The harness mirrors this with the **two-path model** (`src/hostloop/workspace-handler.ts`, binary-verified `G1t`/`U1t`): **Path A** (provenance engaged — `coworkWebFetchViaApi` on) gates on the **exact-URL provenance set** ONLY (seeded from user-turn + tool-result URLs; `src/hostloop/provenance.ts`), with **no** hostname allowlist — but still an `http(s)`-scheme + private-address **SSRF backstop re-checked on every redirect hop** (a manual redirect loop, not `curl -L`) — a miss raises a per-domain approval (`webfetch:<domain>` permission with options `Allow once | Allow all for website | Deny`) routed through the Decider; "Allow all for website" approves the host for the rest of the run (`Run.approvedDomains`, per-run/ephemeral). **Path B** (gate off) is a direct host fetch gated by the egress domain list via the same `wen()`/`compile()` matcher container egress uses, with `redirect:"manual"` re-checking `U1t` (scheme + private-address SSRF + allowlist) on **every** redirect hop. web_fetch is thus **decoupled from `plan.egressAllow`** on Path A (egress applies to `bash`/Path B only). An unanswered cold miss is fail-closed; scenarios answer via `--answer "webfetch:<domain>=allow"` (with `grant`), `web_fetch.approved_domains`, or an LLM/external terminal. `bash` stays container-egress-sandboxed.
312
+ 1. **SDK servers over the control protocol** — declared via `sdkMcpServers` in `initialize`; tool calls tunnel as `mcp_message`. This is how the **desktop host** bridges its own servers (incl. `claude_desktop_config.json` `mcpServers`, spawned host-side with full host env) and how the harness delivers the workspace shell. The workspace handler (`src/hostloop/workspace-handler.ts`) implements `initialize`/`tools/list`/`tools/call`; `bash`→`docker exec -w <mntRoot> <container> sh -c <cmd>` (container-egress-gated). **CB-8:** `makeWorkspaceHandler` accepts an `onInfraError?: (message: string) => void` callback at parameter position 6 (after `onEgress`); on infrastructure errors (ETIMEDOUT / killed / no code+stdout+stderr) the handler calls `onInfraError?.(e.message)`, returns a textResult with `"[infrastructure error: …]"`, and `spawnHostLoop` wires `onInfraError` to append `{type:"infra_error", ts, message}` to `events.jsonl`. **`web_fetch` is NOT container-egress-gated:** real Cowork routes it through the host API (gate `1978029737` `coworkWebFetchViaApi:true` → `POST /api/organizations/<org>/cowork/web_fetch`), gated by a **separate web-fetch hostname allowlist** (`getWebFetchAllowedUrls`, `*`=unrestricted) + a **URL-provenance** rule (URL must have appeared in a prior message/result). The harness mirrors this with the **two-path model** (`src/hostloop/workspace-handler.ts`, binary-verified `G1t`/`U1t`): **Path A** (provenance engaged — `coworkWebFetchViaApi` on) gates on the **exact-URL provenance set** ONLY (seeded from user-turn + tool-result URLs; `src/hostloop/provenance.ts`), with **no** hostname allowlist — but still an `http(s)`-scheme + private-address **SSRF backstop re-checked on every redirect hop** (a manual redirect loop, not `curl -L`) — a miss raises a per-domain approval (`webfetch:<domain>` permission with options `Allow once | Allow all for website | Deny`) routed through the Decider; "Allow all for website" approves the host for the rest of the run (`Run.approvedDomains`, per-run/ephemeral). On Path A, **AFTER** the provenance/approval gate, **`coworkWebFetchDedup`** (gate `1978029737`, on for baselines ≥ `1.22209.3`) applies a per-session **negative-work cache** (`src/hostloop/webfetch-dedup.ts`, binary-verified): a repeat `web_fetch` of the same **normalized** URL (`normalizeUrl`) within the TTL (baseline-sourced, default `900000` ms; cap `100`, FIFO eviction; a hit does NOT refresh recency) returns a marker telling the model to re-use the earlier result, with **no network request and no egress event** — matching production's zero-network dedup. Only successful (HTTP-2xx, trimmed-nonempty) fetches are cached, keyed under both the request URL and the terminal `destination_url`; errors/empty/non-2xx are never cached. Dedup is baseline-gated (an older baseline lacking the flag never dedups) and Path-A-only. **Path B** (gate off) is a direct host fetch gated by the egress domain list via the same `wen()`/`compile()` matcher container egress uses, with `redirect:"manual"` re-checking `U1t` (scheme + private-address SSRF + allowlist) on **every** redirect hop. web_fetch is thus **decoupled from `plan.egressAllow`** on Path A (egress applies to `bash`/Path B only). An unanswered cold miss is fail-closed; scenarios answer via `--answer "webfetch:<domain>=allow"` (with `grant`), `web_fetch.approved_domains`, or an LLM/external terminal. `bash` stays container-egress-sandboxed.
313
313
  2. **CLI-spawned `--mcp-config` / `.mcp.json` servers — HONORED in plain cowork mode** (NOT ignored). **Verified:** a valid `--mcp-config` populates `mcp_servers` (`[{name,status:"pending"|"connected"}]`); these run in-sandbox with the env-allowlist `CLAUDE_CODE_MCP_ALLOWLIST_ENV` (`RW8`/`oG8`/`LU5` = {HOME, LOGNAME, PATH, SHELL, TERM, USER}). The harness MAY use this as a convenience injection path.
314
314
  3. **The drop is SAFE/HERMETIC-mode-gated, not cowork-gated.** `--mcp-config` is filtered to SDK-only (`ap5()`) **only when** safe mode (`I5()`) or `xB8()` is true, and `xB8()` requires **both** `CLAUDE_CODE_REMOTE` **and** `CLAUDE_CODE_REMOTE_HERMETIC_MODE`. **Verified:** with both set, `mcp_servers:[]`; without them (plain `SESSION_KIND=bg`), the config is honored. The earlier "cowork ignores `--mcp-config`" was a hermetic-session observation over-generalized.
315
315
 
@@ -0,0 +1,392 @@
1
+ {
2
+ "baselineVersion": 1,
3
+ "appVersion": "1.22209.3",
4
+ "agentVersion": "2.1.215",
5
+ "agentBinary": {
6
+ "stagedPath": "~/Library/Application Support/Claude/claude-code-vm/2.1.215/claude",
7
+ "format": "elf-aarch64",
8
+ "nativeStagedPath": "~/Library/Application Support/Claude/claude-code/2.1.215/claude.app/Contents/MacOS/claude",
9
+ "sha256": "2b43a3d5b0787217e5d7381fad42c7314292546fe9db9eb8b9b379de90509b30",
10
+ "shaProvenance": "measured-local",
11
+ "manifestChecksumMatch": true
12
+ },
13
+ "guest": {
14
+ "os": "linux",
15
+ "arch": "arm64",
16
+ "baseImage": "ubuntu:22.04"
17
+ },
18
+ "spawn": {
19
+ "configDirInGuest": "mnt/.claude",
20
+ "settingSources": [
21
+ "user"
22
+ ],
23
+ "permissionMode": "default",
24
+ "maxThinkingTokens": 31999,
25
+ "effortDefault": "medium",
26
+ "effortByModel": {
27
+ "claude-haiku-4-5": {
28
+ "modes": [
29
+ "extended"
30
+ ]
31
+ },
32
+ "claude-sonnet-4-5": {
33
+ "modes": [
34
+ "extended"
35
+ ]
36
+ },
37
+ "claude-sonnet-4-6": {
38
+ "effortLevels": [
39
+ "low",
40
+ "medium",
41
+ "high",
42
+ "max"
43
+ ],
44
+ "recommended": "low",
45
+ "modes": [
46
+ "auto"
47
+ ]
48
+ },
49
+ "claude-opus-4-6": {
50
+ "effortLevels": [
51
+ "low",
52
+ "medium",
53
+ "high",
54
+ "max"
55
+ ],
56
+ "recommended": "medium",
57
+ "modes": [
58
+ "extended"
59
+ ]
60
+ },
61
+ "claude-opus-4-7": {
62
+ "effortLevels": [
63
+ "low",
64
+ "medium",
65
+ "high",
66
+ "xhigh",
67
+ "max"
68
+ ],
69
+ "recommended": "xhigh",
70
+ "modes": [
71
+ "auto"
72
+ ]
73
+ },
74
+ "claude-opus-4-8": {
75
+ "effortLevels": [
76
+ "low",
77
+ "medium",
78
+ "high",
79
+ "xhigh",
80
+ "max"
81
+ ],
82
+ "recommended": "high",
83
+ "modes": [
84
+ "auto"
85
+ ]
86
+ }
87
+ },
88
+ "effortRegexDefault": {
89
+ "pattern": "^(?:claude-)?(?:fable|mythos)(?:-|$)",
90
+ "effortLevels": [
91
+ "low",
92
+ "medium",
93
+ "high",
94
+ "xhigh",
95
+ "max"
96
+ ],
97
+ "recommended": "high",
98
+ "modes": [
99
+ "auto"
100
+ ],
101
+ "disallowThinkingDisabled": true
102
+ },
103
+ "tools": [
104
+ "Task",
105
+ "Bash",
106
+ "Glob",
107
+ "Grep",
108
+ "Read",
109
+ "Edit",
110
+ "Write",
111
+ "NotebookEdit",
112
+ "WebFetch",
113
+ "TaskCreate",
114
+ "TaskUpdate",
115
+ "TaskGet",
116
+ "TaskList",
117
+ "TaskStop",
118
+ "WebSearch",
119
+ "Skill",
120
+ "REPL",
121
+ "JavaScript",
122
+ "AskUserQuestion",
123
+ "ToolSearch"
124
+ ],
125
+ "allowedTools": [
126
+ "Task",
127
+ "Bash",
128
+ "Glob",
129
+ "Grep",
130
+ "Read",
131
+ "Edit",
132
+ "Write",
133
+ "NotebookEdit",
134
+ "WebFetch",
135
+ "TaskCreate",
136
+ "TaskUpdate",
137
+ "TaskGet",
138
+ "TaskList",
139
+ "TaskStop",
140
+ "WebSearch",
141
+ "Skill",
142
+ "REPL",
143
+ "JavaScript",
144
+ "ToolSearch"
145
+ ],
146
+ "env": {
147
+ "CLAUDE_CODE_IS_COWORK": "1",
148
+ "CLAUDE_CODE_ENTRYPOINT": "local-agent",
149
+ "CLAUDE_CODE_TAGS": "lam_session_type:chat",
150
+ "CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST": "1",
151
+ "CLAUDE_CODE_ENABLE_ASK_USER_QUESTION_TOOL": "true",
152
+ "CLAUDE_CODE_DISABLE_CRON": "1",
153
+ "CLAUDE_CODE_DISABLE_BACKGROUND_TASKS": "1",
154
+ "CLAUDE_CODE_DISABLE_AGENTS_FLEET": "1",
155
+ "CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT": "1",
156
+ "CLAUDE_CODE_ENABLE_TASKS": "true",
157
+ "CLAUDE_CODE_DISABLE_TERMINAL_TITLE": "1",
158
+ "ENABLE_PROMPT_CACHING_1H": "1",
159
+ "DISABLE_MICROCOMPACT": "1",
160
+ "MCP_CONNECTION_NONBLOCKING": "true",
161
+ "API_TIMEOUT_MS": "900000",
162
+ "CLAUDE_CODE_EMIT_TOOL_USE_SUMMARIES": "",
163
+ "CLAUDE_CODE_ENABLE_FINE_GRAINED_TOOL_STREAMING": "1",
164
+ "DISABLE_AUTOUPDATER": "1",
165
+ "MCP_TOOL_TIMEOUT": "60000",
166
+ "USE_LOCAL_OAUTH": "",
167
+ "USE_STAGING_OAUTH": ""
168
+ },
169
+ "promptTemplate": "prompts/desktop-1.18286.0/system-prompt-append.md",
170
+ "subagentAppend": "prompts/desktop-1.15200.0/subagent-append-vm.md",
171
+ "subagentAppendHostLoop": "prompts/desktop-1.18286.2/subagent-append-hl.md",
172
+ "$comment": "Binary-verified Desktop->agent spawn contract, re-derived per release. spawn.env is GENERATED by deriveSpawnEnv() in src/sync/cowork-sync.ts (windowed enumeration of the asar env construction + gate/const value resolution); the scalar options, tools/allowedTools, and prompt-asset pointers are sentinel-guarded by checkSpawnContractFacts(). Do not hand-edit spawn.env — re-run sync.",
173
+ "$comment_handPinned": "Why the NON-env spawn fields stay hand-pinned: each is built in the asar as a non-literal expression (a session-path template, a session-type ternary, a const indirection, or a head+spread+tail array), so the windowed-enumeration generator that derives spawn.env cannot construct their VALUES without a full JS evaluator; instead each value was binary-verified once and is drift-guarded by a checkSpawnContractFacts() sentinel (cowork-sync.ts) that re-asserts the asar-side FACT at every sync. Scope caveat: the sentinels make DESKTOP-side drift loud; they do not validate this committed JSON itself — an erroneous hand-edit here is invisible to them.",
174
+ "$comment_configDirInGuest": "Hand-pinned: the asar builds it as a per-session path template (/sessions/${id}/mnt/.claude), not a constructable literal. Sentinel S1 pins the template shape.",
175
+ "$comment_settingSources": "Hand-pinned: sentinel S2 pins the settingSources:[\"user\"] literal.",
176
+ "$comment_permissionMode": "Hand-pinned: the asar computes it via a session ternary whose chat-session branch resolves to \"default\". Sentinel S3 pins the ternary shape.",
177
+ "$comment_maxThinkingTokens": "Hand-pinned: the asar reaches the value through const indirection. Sentinel S4 VALUE-pins the resolved const to 31999.",
178
+ "$comment_effortDefault": "Hand-pinned: sentinel S5 pins the .effort … :\"medium\" default.",
179
+ "$comment_tools": "Hand-pinned: the asar builds tools[] as head-list + Task-tools spread + session-type tail, not one literal. Sentinels S6 (head), S7 (the TaskCreate…TaskStop spread), S8 (tail-guard after ToolSearch) pin all three parts. As of desktop-1.21459.0 the asar head also carries an INERT `...CLAUDE_DESIGN_TOOLS` spread between Task and Bash that resolves to [] (deployment-gated off on first-party), so the rendered list — and this pin — stay 20 entries; S6b asserts it empty and fails loud if a build ever populates it.",
180
+ "$comment_allowedTools": "Hand-pinned: same head+spread+tail construction as tools[], minus AskUserQuestion (tools-only by design). Sentinels S9 (head) and S10 (the built-in→mcp__ boundary tail-guard) pin it.",
181
+ "$comment_promptTemplate": "Hand-pinned pointer to a RECONSTRUCTED asset (see $comment_prompts) — the generator cannot extract prose. Sentinel S15 pins the claude_code preset-append delivery site.",
182
+ "$comment_subagentAppend": "Hand-pinned pointer to a reconstructed asset (see $comment_prompts). Sentinel S16 pins the per-session appendSubagentSystemPrompt generator call shape.",
183
+ "$comment_subagentAppendHostLoop": "Hand-pinned pointer to a reconstructed (paraphrased) asset for the HOST-LOOP branch of the per-session sub-agent append (section key subagent_env_hl; selected purely on hostLoopMode). Backfilled only for release families whose hl text is binary-verified byte-identical (1.18286.2+). Sentinel: checkSubagentPromptFacts (two-branch fingerprint + substitution-value proofs). A hostloop run on a baseline lacking this pointer fails loud rather than falling back to the VM text.",
184
+ "$comment_notSet": "Deliberately NOT set: CLAUDE_CODE_USE_COWORK_PLUGINS (Desktop never sets it; would flip the agent to cowork_settings.json/cowork_plugins — asserted absent by the S17 negative invariant). Enumerated-but-not-pinned keys are enforced by SPAWN_ENV_ALLOWLIST in src/sync/cowork-sync.ts, each with a reason; categories: host-derived (CLAUDE_CONFIG_DIR, TZ, HOST_PLATFORM, OAUTH_TOKEN/BASE_URL/CUSTOM_HEADERS, account UUIDs, WORKSPACE_HOST_PATHS, OTEL), constructed-then-deleted (ANTHROPIC_API_KEY/AUTH_TOKEN via FnA), gate-conditional-off (MCP_CONNECT_TIMEOUT_MS, ENABLE_TOOL_SEARCH, SKIP_PRECOMPACT_LOAD), non-chat/project-session (BRIEF*, PROJECT*), user-settings (SUBAGENT_MODEL, AUTO_COMPACT_WINDOW, ...), and 3p-provider-only branches. The opaque ...g.env/...l session spreads are the known static-extraction blind spot (runtime lane backstop).",
185
+ "$comment_prompts": "Reconstructed cowork-specific sections, re-paraphrased from asar 1.18286.0 const aui (system prompt; RESTRUCTURED at this release — see the asset header) and 1.15200.0 generator CVr (subagent; verified unchanged in the 1.18286.0 asar, generator Zgn). Not the full base prompt (not cleanly extractable); generic refusal/safety policy elided. Delivered via --append-system-prompt (layered on the agent's built-in base prompt), NOT the initialize handshake; only the subagent append goes over initialize (appendSubagentSystemPrompt), gated on CLAUDE_CODE_ENABLE_APPEND_SUBAGENT_PROMPT."
186
+ },
187
+ "mountLayout": {
188
+ "sessionRoot": "/sessions/{sessionId}",
189
+ "cwd": "/sessions/{sessionId}",
190
+ "mntRoot": "/sessions/{sessionId}/mnt",
191
+ "mounts": [
192
+ {
193
+ "name": "uploads",
194
+ "mountPath": "uploads",
195
+ "mode": "r",
196
+ "purpose": "user-uploaded files (read-only — asar 'ro')"
197
+ },
198
+ {
199
+ "name": "projects",
200
+ "mountPath": ".projects/{projectId}",
201
+ "mode": "rw",
202
+ "purpose": "RESERVED namespace (and the separate UUID project-sync feature) — NOT the work-folder path. From Desktop 1.14271.0 selected work folders mount at mnt/<collision-resolved-basename> (dynamic, derived per session by buildLaunchPlan; see MOUNT_BARE_NAME_MIN_VERSION). This decorative row is not consumed for binding (staged paths come from plan.mounts)."
203
+ },
204
+ {
205
+ "name": "local-plugins",
206
+ "mountPath": ".local-plugins/marketplaces",
207
+ "mode": "r",
208
+ "purpose": "marketplace skills/plugins, runtime-discovered"
209
+ },
210
+ {
211
+ "name": "remote-plugins",
212
+ "mountPath": ".remote-plugins",
213
+ "mode": "r",
214
+ "purpose": "org-remote plugins, runtime-discovered"
215
+ },
216
+ {
217
+ "name": "outputs",
218
+ "mountPath": "outputs",
219
+ "mode": "rw",
220
+ "purpose": "session outputs/artifacts — delete denied by default (asar IX); rwd only when approved"
221
+ },
222
+ {
223
+ "name": "skills",
224
+ "mountPath": ".claude/skills",
225
+ "mode": "r",
226
+ "purpose": "personal/saved skill doc bodies (NOT plugin-bundled skills, which live under local-plugins/remote-plugins above) — real VM confirmed via systemd unit sessions-<name>-mnt-.claude-skills.mount in vm_bundles/claudevm.bundle/rootfs.img. Decorative row like 'projects' above (not consumed for binding — resolveMounts()'s mounts[] is destructured away at every call site); the harness reproduces this via CLAUDE_CONFIG_DIR staging (session.ts skill copy + stage.ts cpSync), not a plan.mounts bind — see hostloop-prompt.ts's asar-verified skills bullet."
227
+ }
228
+ ]
229
+ },
230
+ "network": {
231
+ "mode": "gvisor",
232
+ "allowKind": "allowlist",
233
+ "allowDomains": [
234
+ "preview.claude.ai",
235
+ "downloads.claude.ai",
236
+ "api.anthropic.com",
237
+ "a-cdn.anthropic.com",
238
+ "a-api.anthropic.com",
239
+ "assets.claude.ai",
240
+ "sentry.io",
241
+ "console.anthropic.com",
242
+ "api-staging.anthropic.com",
243
+ "www.anthropic.com",
244
+ "api.claude.ai",
245
+ "support.anthropic.com",
246
+ "docs.anthropic.com",
247
+ "mcp-proxy.anthropic.com",
248
+ "pivot.claude.ai"
249
+ ]
250
+ },
251
+ "bgEnvStrip": {
252
+ "knownVars": [
253
+ "CLAUDE_CODE_OAUTH_TOKEN",
254
+ "CLAUDE_CODE_SESSION_KIND",
255
+ "CLAUDE_CODE_SESSION_ID",
256
+ "CLAUDE_CODE_SESSION_NAME",
257
+ "CLAUDE_CODE_SESSION_LOG"
258
+ ]
259
+ },
260
+ "$comment": "Platform baseline auto-derived by `cowork-harness sync` from a live Claude Desktop install + app.asar. VOLATILE per-release facts only. Regenerate per release; review the diff. Captured 2026-07-19 on macOS arm64.",
261
+ "capturedAt": "2026-07-19",
262
+ "platform": "darwin-arm64",
263
+ "settings": {
264
+ "autoMountFolders": {
265
+ "key": "autoMountFolders",
266
+ "default": false
267
+ },
268
+ "localAgentModeTrustedFolders": {
269
+ "key": "localAgentModeTrustedFolders",
270
+ "default": []
271
+ }
272
+ },
273
+ "provenance": {
274
+ "asarPath": "/Applications/Claude.app/Contents/Resources/app.asar",
275
+ "asarFingerprint": "a8cf7b198d479a2b",
276
+ "gates": {
277
+ "$comment": "Production GrowthBook gate states decoded from ~/Library/Application Support/Claude/fcache (standard interactive Anthropic account, 2026-06-13; binary-verified app.asar 1.12603.1). Pin per release. Behavior-affecting gates the harness models: 1143815894 (loop), 1648655587 (dispatch cap), 1978029737 (web_fetch routing). Telemetry/auth-internal gates omitted. Also pinned: 2614807392 (skeletonHome), 123929380 (autoMemoryStandardSessions), 1696890383 (memoryGuidelinesEnv), 2860753854 (memoryExtraGuidelines) — dormant drift-sentinels for dark-launched features (host-fs skeleton, auto-memory) the harness deliberately models as OFF (or, for memoryExtraGuidelines, as inert-default: on in production but its served value equals the hardcoded default); pinned so a production flip surfaces as a sync diff instead of silent drift.",
278
+ "emitToolUseSummaries:66187241": {
279
+ "on": false,
280
+ "source": "defaultValue",
281
+ "value": false
282
+ },
283
+ "autoMemoryStandardSessions:123929380": {
284
+ "on": false,
285
+ "source": "defaultValue",
286
+ "value": false
287
+ },
288
+ "subagentPromptServerOverride:124685897": {
289
+ "on": false,
290
+ "source": "defaultValue",
291
+ "value": false
292
+ },
293
+ "mcpConnectionNonblockingOff:434204418": {
294
+ "on": false,
295
+ "source": "defaultValue",
296
+ "value": false
297
+ },
298
+ "bridgeSdkTransport:583857784": {
299
+ "on": true,
300
+ "source": "force",
301
+ "value": true,
302
+ "note": "— Cowork uses the SDK-based transport (control protocol), confirming the harness's sdkMcpServers/mcp_message path is the production transport."
303
+ },
304
+ "fineGrainedToolStreaming:714014285": {
305
+ "on": true,
306
+ "source": "force",
307
+ "value": true
308
+ },
309
+ "enableToolSearchAuto:1129419822": {
310
+ "on": false,
311
+ "source": "absent"
312
+ },
313
+ "hostLoop:1143815894": {
314
+ "on": true,
315
+ "source": "force",
316
+ "value": true
317
+ },
318
+ "scheduledTaskToolsApprovableByAutoMode:1447478638": {
319
+ "on": false,
320
+ "source": "absent"
321
+ },
322
+ "scheduledTaskSessionLimiter:1648655587": {
323
+ "on": true,
324
+ "source": "force",
325
+ "value": {
326
+ "global": 3,
327
+ "perTask": 1
328
+ },
329
+ "note": "SCHEDULED-TASK (cron) session limiter — NOT an in-conversation Task-tool cap (binary-verified 2026-07-04, asar 1.18286.0 class L9t [ScheduledTasks]). perTask=1: <=1 concurrent session PER SCHEDULED TASK; global=3: <=3 concurrent scheduled-task sessions globally (+_pendingTaskDispatches). Host-side SKIP (recordSkipAndEmit/PerTaskLimit|GlobalLimit — NOT queue/deny). Cowork imposes no cap on Task-tool sub-agent fan-out; the harness has no scheduled-task scheduler, so this gate has no applicable surface — pinned as a sync drift-sentinel only."
330
+ },
331
+ "memoryGuidelinesEnv:1696890383": {
332
+ "on": false,
333
+ "source": "defaultValue",
334
+ "value": false
335
+ },
336
+ "oauthScopesEnv:1936081873": {
337
+ "on": true,
338
+ "source": "force",
339
+ "value": true
340
+ },
341
+ "coworkRuntimeConfig:1978029737": {
342
+ "on": true,
343
+ "source": "experiment",
344
+ "value": {
345
+ "coworkNativeFilePreview": true,
346
+ "coworkWebFetchDedup": true,
347
+ "coworkWebFetchDedupMaxEntries": 100,
348
+ "coworkWebFetchDedupTtlMs": 900000,
349
+ "coworkWebFetchPrompt": true,
350
+ "coworkWebFetchViaApi": true,
351
+ "pluginsFullSyncStalenessMs": 0,
352
+ "sessionsBridgePollBlockMs": 30,
353
+ "workspaceBashWaitLonger": true
354
+ },
355
+ "note": "coworkWebFetchViaApi=true coworkWebFetchPrompt=true workspaceBashWaitLonger=true sessionsBridgePollBlockMs=30 — web_fetch is host/API-routed (POST /api/organizations/<org>/cowork/web_fetch), NOT container egress; gated by a separate web-fetch hostname allowlist + URL provenance."
356
+ },
357
+ "cliPlugin:2307090146": {
358
+ "on": false,
359
+ "source": "defaultValue",
360
+ "value": false,
361
+ "note": "— the CLI-plugin credential broker is dark-launched off for standard interactive accounts (Ch23/L106)."
362
+ },
363
+ "pluginSyncSparkplug:2340532315": {
364
+ "on": true,
365
+ "source": "force",
366
+ "value": true,
367
+ "note": "— startup syncPlugins(); plugins load via --plugin-dir (registry inert in-VM)."
368
+ },
369
+ "skeletonHome:2614807392": {
370
+ "on": false,
371
+ "source": "absent"
372
+ },
373
+ "memoryExtraGuidelines:2860753854": {
374
+ "on": true,
375
+ "source": "defaultValue",
376
+ "value": "## Sensitive personal information\n\nDo not save the following to memory unless the user explicitly asks you to remember it:\n\n- Protected attributes: race, ethnicity, national origin, religion, age, sex, sexual orientation, gender identity, immigration status, disability, serious illness, union membership\n- Government identifiers: Social Security numbers, driver's license numbers, passport numbers, government ID numbers\n- Financial account details: credit card numbers, bank account numbers\n- Health information: medical conditions, diagnoses, lab results, mental health details, therapy or counseling\n- Home or personal mailing addresses (work addresses are fine)\n- Account passwords, secret tokens, or secret keys\n\nIf any of the above appears in conversation context, complete the task but do not persist it to a memory file. If the user explicitly says \"remember my address is X\", saving it is acceptable — they've given consent."
377
+ },
378
+ "skipPrecompactLoad:4153934152": {
379
+ "on": false,
380
+ "source": "defaultValue",
381
+ "value": false
382
+ },
383
+ "autoModeOverridesAlwaysAllow:4200321681": {
384
+ "on": false,
385
+ "source": "absent"
386
+ }
387
+ },
388
+ "eipcChannelUuid": "4f426349-8d6f-45f3-ae22-280fef323564",
389
+ "$comment": "eipcChannelUuid is per-build; recorded for provenance only — the harness does not use Desktop IPC."
390
+ },
391
+ "requireFullVmSandbox": null
392
+ }
@@ -0,0 +1,50 @@
1
+ /**
2
+ * `coworkWebFetchDedup` — a per-Cowork-session NEGATIVE-WORK cache for hostloop `web_fetch` (binary-verified
3
+ * port of Claude Desktop 1.22209.3's `Ne` write/evict + the lookup in `.vite/build/index.chunk-CYQPQGee.js`).
4
+ *
5
+ * It stores only `{ts, size, hits}` (never content). A repeat fetch of the same normalized URL within the
6
+ * TTL is a HIT: real Cowork skips the network entirely and returns a marker telling the model to re-use the
7
+ * earlier result. See the workspace-handler seam for the marker + the no-egress semantics.
8
+ *
9
+ * Faithful details (do NOT "improve"): FIFO (insertion-order) eviction — NOT LRU; a hit does NOT refresh the
10
+ * entry's `ts`/recency (`hits` is incremented but is behaviorally inert), so an entry hard-expires exactly
11
+ * `ttlMs` after the fetch that created it. Hit boundary is INCLUSIVE (`now - ts <= ttlMs`); the write-time
12
+ * TTL sweep is STRICT (`now - ts > ttlMs`). Clock is injected (`nowMs`) — no `Date.now()` here, so tests
13
+ * pin the boundaries deterministically.
14
+ */
15
+ export function makeWebFetchDedupCache(opts) {
16
+ const { ttlMs, maxEntries } = opts;
17
+ const map = new Map();
18
+ return {
19
+ lookup(normUrl, nowMs) {
20
+ const e = map.get(normUrl);
21
+ if (e === undefined)
22
+ return null;
23
+ if (nowMs - e.ts <= ttlMs) {
24
+ e.hits += 1; // structural parity only — does NOT refresh recency/ts (hard-expiry stands)
25
+ return { ageS: Math.round((nowMs - e.ts) / 1000) };
26
+ }
27
+ map.delete(normUrl); // expired → delete + miss
28
+ return null;
29
+ },
30
+ record(normUrl, size, nowMs) {
31
+ // 1) TTL sweep (STRICT >), deleting entries older than the window. Deleting during Map iteration is safe.
32
+ for (const [k, v] of map)
33
+ if (nowMs - v.ts > ttlMs)
34
+ map.delete(k);
35
+ // 2) delete-then-set moves the key to the tail so recency == insertion/re-insertion order (FIFO).
36
+ map.delete(normUrl);
37
+ map.set(normUrl, { ts: nowMs, size, hits: 0 });
38
+ // 3) FIFO-evict the oldest-inserted entries until at the cap.
39
+ while (map.size > maxEntries) {
40
+ const oldest = map.keys().next().value;
41
+ if (oldest === undefined)
42
+ break;
43
+ map.delete(oldest);
44
+ }
45
+ },
46
+ ttlSeconds() {
47
+ return Math.round(ttlMs / 1000);
48
+ },
49
+ };
50
+ }
@@ -6,6 +6,7 @@ import http from "node:http";
6
6
  import https from "node:https";
7
7
  import { lookup } from "node:dns/promises";
8
8
  import { compile } from "../egress/proxy.js";
9
+ import { normalizeUrl } from "./provenance.js";
9
10
  const pexec = promisify(execFile);
10
11
  const MAX_REDIRECTS = 5; // Cowork's RZe redirect cap (Path B re-checks U1t per hop)
11
12
  /** Is a dotted-quad's four octets in a loopback/this-host/private/link-local range? */
@@ -227,7 +228,7 @@ const defaultRawFetch = async (url, pinnedAddresses) => {
227
228
  const BASH_DESC = "Run a shell command in the session's isolated Linux workspace. Your connected folders are mounted under {{mnt}}/ — the Shell access section of your system prompt lists the exact path for each folder. Each bash call is independent (no cwd/env carryover). Use absolute paths.";
228
229
  const FETCH_DESC = "Fetch a URL from the session network (subject to the egress allowlist). web_fetch can only retrieve URLs that appeared in a user message or a prior result.";
229
230
  export function makeWorkspaceHandler(opts) {
230
- const { containerName, vmMnt, runner = "docker", webFetchAllow = ["*"], onEgress, onInfraError, provenanceRef, rawFetch = defaultRawFetch, resolve = defaultResolver, execCwd = vmMnt, } = opts;
231
+ const { containerName, vmMnt, runner = "docker", webFetchAllow = ["*"], onEgress, onInfraError, provenanceRef, dedup, rawFetch = defaultRawFetch, resolve = defaultResolver, execCwd = vmMnt, } = opts;
231
232
  // Per-handler (per-spawn) latch for the provenance-unenforced warning — was module-level, which
232
233
  // silenced the gap after the first run in a long-lived process. Each fresh handler warns once.
233
234
  const provWarned = { value: false };
@@ -264,7 +265,7 @@ export function makeWorkspaceHandler(opts) {
264
265
  };
265
266
  if (name === "web_fetch")
266
267
  return {
267
- result: await fetchViaHost(String(a.url ?? ""), webFetchAllow, onEgress, provenanceRef?.current, provWarned, rawFetch, resolve),
268
+ result: await fetchViaHost(String(a.url ?? ""), webFetchAllow, onEgress, provenanceRef?.current, provWarned, rawFetch, resolve, dedup),
268
269
  };
269
270
  return { error: { code: -32602, message: `unknown tool: ${name}` } };
270
271
  }
@@ -277,6 +278,13 @@ function textResult(text, isError = false) {
277
278
  r.isError = true;
278
279
  return r;
279
280
  }
281
+ /** The `coworkWebFetchDedup` HIT marker — verbatim port of Claude Desktop 1.22209.3 (asar chunk
282
+ * `index.chunk-CYQPQGee.js`). `url` is the raw model-sent argument; `ageS`/`ttlS` are already Math.round-ed. */
283
+ function dedupMarker(url, ageS, ttlS) {
284
+ return (`Already fetched ${url} ${ageS}s ago in this session. Re-use the content from that earlier web_fetch ` +
285
+ `result instead of re-reading it. Fetch again only if the page is likely to have changed ` +
286
+ `(deduplicated for up to ${ttlS}s).`);
287
+ }
280
288
  // Clamp a model-requested bash timeout into a sane range. Guards NaN/negative/missing → the
281
289
  // 120s default; floors at 1s and caps at 10min. This is INFERRED-parity for the bash tool — Cowork's
282
290
  // binary-verified timeout_ms honoring is for web_fetch, not bash — but the bash inputSchema advertises
@@ -368,7 +376,11 @@ function schemePrivateGate(u) {
368
376
  * protection that `curl -L` lacked. Emits one egress allow/deny on the terminal decision. Shared by
369
377
  * both web_fetch paths; the gate is the only difference (Path A: scheme+private; Path B: full u1t).
370
378
  */
371
- async function followWithRedirects(startUrl, rawFetch, gate, resolve, onEgress) {
379
+ async function followWithRedirects(startUrl, rawFetch, gate, resolve, onEgress,
380
+ // coworkWebFetchDedup record hook — called ONLY on a terminal HTTP-2xx, trimmed-nonempty response, with
381
+ // the post-redirect terminal URL (`cur.href` = destination_url) + the untrimmed emitted-text length. Not
382
+ // called on errors, non-2xx bodies, empty/whitespace bodies, or redirect hops (so those are never cached).
383
+ onSuccess) {
372
384
  let cur;
373
385
  try {
374
386
  cur = new URL(startUrl);
@@ -442,16 +454,21 @@ async function followWithRedirects(startUrl, rawFetch, gate, resolve, onEgress)
442
454
  }
443
455
  const decoder = new TextDecoder();
444
456
  const text = chunks.map((c) => decoder.decode(c, { stream: true })).join("") + decoder.decode();
457
+ // dedup record: 2xx + trimmed-nonempty only; untrimmed emitted length (pre-`[truncated]` suffix).
458
+ if (resp.status >= 200 && resp.status < 300 && text.trim().length > 0)
459
+ onSuccess?.(cur.href, text.length);
445
460
  return textResult(truncated ? text + "\n[truncated]" : text);
446
461
  }
447
462
  const raw = await resp.text();
448
463
  const sliced = raw.slice(0, LIMIT);
449
464
  const wasTruncated = resp.truncated || sliced.length < raw.length;
465
+ if (resp.status >= 200 && resp.status < 300 && sliced.trim().length > 0)
466
+ onSuccess?.(cur.href, sliced.length);
450
467
  return textResult(wasTruncated ? sliced + "\n[truncated]" : sliced);
451
468
  }
452
469
  return textResult(`web_fetch failed: too many redirects (> ${MAX_REDIRECTS}).`, true);
453
470
  }
454
- async function fetchViaHost(url, allow, onEgress, prov, warned, rawFetch = defaultRawFetch, resolve = defaultResolver) {
471
+ async function fetchViaHost(url, allow, onEgress, prov, warned, rawFetch = defaultRawFetch, resolve = defaultResolver, dedup) {
455
472
  if (!url)
456
473
  return textResult("error: missing 'url'", true);
457
474
  let host;
@@ -488,6 +505,27 @@ async function fetchViaHost(url, allow, onEgress, prov, warned, rawFetch = defau
488
505
  // Provenance satisfied. Cowork fetches server-side (host API); the hostname allowlist does NOT apply
489
506
  // here (decoupled from egress). But scheme + private-address ARE enforced per
490
507
  // hop: follow redirects manually instead of `curl -L`, blocking file:// / SSRF targets.
508
+ //
509
+ // coworkWebFetchDedup (host-API path only): AFTER the provenance gate, a repeat of the same normalized
510
+ // URL within the TTL returns a marker with NO network request and NO egress event — matching
511
+ // production's zero-network dedup. On a miss, record the request key + the terminal destination_url on
512
+ // a 2xx, non-empty success (the guards live in followWithRedirects's onSuccess call).
513
+ if (dedup) {
514
+ const key = normalizeUrl(url);
515
+ if (key) {
516
+ const cache = dedup;
517
+ const hit = cache.lookup(key, Date.now());
518
+ if (hit)
519
+ return textResult(dedupMarker(url, hit.ageS, cache.ttlSeconds()), false);
520
+ const onSuccess = (terminalUrl, size) => {
521
+ cache.record(key, size, Date.now());
522
+ const dk = normalizeUrl(terminalUrl); // terminal destination_url key
523
+ if (dk && dk !== key)
524
+ cache.record(dk, size, Date.now());
525
+ };
526
+ return followWithRedirects(url, rawFetch, schemePrivateGate, resolve, onEgress, onSuccess);
527
+ }
528
+ }
491
529
  return followWithRedirects(url, rawFetch, schemePrivateGate, resolve, onEgress);
492
530
  }
493
531
  // PATH B (provenance not enforced — coworkWebFetchViaApi off). Faithful port of U1t re-checked on EVERY
@@ -28,6 +28,36 @@ export function readGateFlag(baseline, id, flag) {
28
28
  }
29
29
  return false;
30
30
  }
31
+ /**
32
+ * Numeric sibling of `readGateFlag` (e.g. `coworkWebFetchDedupTtlMs`). Same prefixed-then-bare key lookup;
33
+ * reads the number from the structured `value[flag]` (or the bare-entry `[flag]`), else parses `flag=<n>`
34
+ * from the committed prose-string shape. Returns `undefined` when absent — the caller supplies the default
35
+ * (so an older baseline that carries the flag without the numbers still gets the binary default).
36
+ */
37
+ export function readGateNumber(baseline, id, flag) {
38
+ const gates = baseline.provenance?.gates ?? {};
39
+ let entry = gates[id];
40
+ if (entry === undefined) {
41
+ for (const k of Object.keys(gates)) {
42
+ if (k.endsWith(":" + id)) {
43
+ entry = gates[k];
44
+ break;
45
+ }
46
+ }
47
+ }
48
+ if (entry == null)
49
+ return undefined;
50
+ if (typeof entry === "string") {
51
+ const m = new RegExp(`\\b${flag}=(\\d+)\\b`).exec(entry);
52
+ return m ? Number(m[1]) : undefined;
53
+ }
54
+ if (typeof entry === "object") {
55
+ const v = entry.value;
56
+ const raw = v && typeof v === "object" ? v[flag] : entry[flag];
57
+ return typeof raw === "number" ? raw : undefined;
58
+ }
59
+ return undefined;
60
+ }
31
61
  export function decideLoop(inputs) {
32
62
  if (inputs.requireFullVmSandbox === true)
33
63
  return "vm"; // HeA()
package/dist/run/chat.js CHANGED
@@ -23,7 +23,8 @@ import { writeTrace, scrubRawRunLogs } from "./execute.js";
23
23
  import { appendIndexRow, indexRowFromResult } from "./run-index.js";
24
24
  import { scrub, collectSecrets } from "../secrets.js";
25
25
  import { Chain, ScriptedDecider, PermissionDefaultDecider, PromptDecider } from "../decide/decider.js";
26
- import { readGateFlag } from "../loop-decision.js";
26
+ import { readGateFlag, readGateNumber } from "../loop-decision.js";
27
+ import { makeWebFetchDedupCache } from "../hostloop/webfetch-dedup.js";
27
28
  import { checkHostLoopWriteConsent, logHostWriteNotice } from "../hostloop/safety.js";
28
29
  import { PATH_GATE_TOOL_NAMES } from "../hostloop/pretooluse-path-hook.js";
29
30
  import { makeHostLoopCanUseToolGate } from "../hostloop/canusetool-gate.js";
@@ -280,6 +281,13 @@ export async function cmdChat(args) {
280
281
  const viaApiOn = readGateFlag(baseline, "1978029737", "coworkWebFetchViaApi");
281
282
  const promptGateOn = readGateFlag(baseline, "1978029737", "coworkWebFetchPrompt");
282
283
  const provenanceRef = {};
284
+ // coworkWebFetchDedup — per-session cache; kept for the chat REPL's lifetime (= one Cowork session).
285
+ const dedup = viaApiOn && readGateFlag(baseline, "1978029737", "coworkWebFetchDedup")
286
+ ? makeWebFetchDedupCache({
287
+ ttlMs: readGateNumber(baseline, "1978029737", "coworkWebFetchDedupTtlMs") ?? 900000,
288
+ maxEntries: readGateNumber(baseline, "1978029737", "coworkWebFetchDedupMaxEntries") ?? 100,
289
+ })
290
+ : undefined;
283
291
  // ONE readline interface on process.stdin, shared by the turn reader (ttyTurns) and the gate
284
292
  // prompter (PromptDecider). Two interfaces would race for the same stdin → undefined input routing.
285
293
  const rl = readline.createInterface({ input: process.stdin, output: process.stderr });
@@ -326,6 +334,7 @@ export async function cmdChat(args) {
326
334
  dockerNetwork: sidecar.network,
327
335
  provenanceRef,
328
336
  webFetchViaApi: viaApiOn,
337
+ dedup,
329
338
  });
330
339
  child = hl.child;
331
340
  containerName = hl.containerName;
package/dist/run/diff.js CHANGED
@@ -14,6 +14,8 @@ const MASKS = [
14
14
  { re: /\blocal_[a-z0-9]+/gi, token: "<SESSION>" },
15
15
  { re: /\bsess-[A-Za-z0-9-]+/g, token: "<SESSION>" },
16
16
  { re: /\b\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:\.\d+)?Z?\b/g, token: "<TIMESTAMP>" },
17
+ // coworkWebFetchDedup marker's "…{N}s ago…" — a live-run age that varies run-to-run (advisory diff view only).
18
+ { re: /\b\d+s ago\b/g, token: "<AGE>" },
17
19
  ];
18
20
  /** Replaces every volatile-but-not-meaningful span (tool-use ids, UUIDs, session-dir markers,
19
21
  * ISO-8601 timestamps, host paths) with a stable placeholder token, so two runs of the SAME scenario
@@ -27,7 +27,8 @@ import { spawnMicroVm, snapshotMicroVmWorkspace } from "../runtime/microvm.js";
27
27
  import { probeImageOmitted, probeMicrovmOmitted, detectCapabilityUse, capabilityPreflightDecision, CAPABILITY_FAMILIES, } from "../runtime/image-capabilities.js";
28
28
  import { instanceName, VM_WORK_HOST } from "../runtime/lima.js";
29
29
  import { ResourceSampler, makeSampleOnce, foldResources, resolveIntervalMs } from "../runtime/resource-sampler.js";
30
- import { decideLoopFromBaseline, readGateFlag } from "../loop-decision.js";
30
+ import { decideLoopFromBaseline, readGateFlag, readGateNumber } from "../loop-decision.js";
31
+ import { makeWebFetchDedupCache } from "../hostloop/webfetch-dedup.js";
31
32
  import { startEgressSidecar, registerCleanup } from "../egress/sidecar.js";
32
33
  import { startEgressProxy } from "../egress/proxy.js";
33
34
  import { evaluate, hostMatches, budgetFields, runSemanticJudges } from "../assert.js";
@@ -414,6 +415,14 @@ export async function executeScenario(scenario, opts = {}) {
414
415
  const viaApiOn = readGateFlag(baseline, "1978029737", "coworkWebFetchViaApi");
415
416
  const promptGateOn = readGateFlag(baseline, "1978029737", "coworkWebFetchPrompt");
416
417
  const provenanceRef = {};
418
+ // coworkWebFetchDedup (host-API path only): a per-session negative-work cache. Built only when the gate is
419
+ // on (an older baseline that lacks it ⇒ undefined ⇒ no behavior change); 100/900000 come from the baseline.
420
+ const dedup = viaApiOn && readGateFlag(baseline, "1978029737", "coworkWebFetchDedup")
421
+ ? makeWebFetchDedupCache({
422
+ ttlMs: readGateNumber(baseline, "1978029737", "coworkWebFetchDedupTtlMs") ?? 900000,
423
+ maxEntries: readGateNumber(baseline, "1978029737", "coworkWebFetchDedupMaxEntries") ?? 100,
424
+ })
425
+ : undefined;
417
426
  // Pre-flight: if the skill DECLARES required capabilities and the image provably omits one, FAIL FAST here
418
427
  // — before any paid agent run — instead of burning ~12 min to reach a verdict the post-run guard already
419
428
  // knows. The author can opt out with `allow_missing_capability: true` (the fallback is equivalent), which
@@ -524,6 +533,7 @@ export async function executeScenario(scenario, opts = {}) {
524
533
  dockerNetwork: sidecar?.network,
525
534
  provenanceRef,
526
535
  webFetchViaApi: viaApiOn,
536
+ dedup,
527
537
  });
528
538
  child = hl.child;
529
539
  sdkMcp = hl.sdkMcp;
@@ -261,6 +261,7 @@ export function spawnHostLoop(_scenario, baseline, plan, outDir, sessionId, opts
261
261
  onEgress: (e) => hostEgress.push(e),
262
262
  onInfraError: logInfra,
263
263
  provenanceRef: opts.provenanceRef,
264
+ dedup: opts.dedup,
264
265
  execCwd,
265
266
  });
266
267
  const sdkMcp = { servers: ["workspace"], handle: workspaceHandle };
@@ -71,7 +71,7 @@ Another runtime knob in the same family: `COWORK_HARNESS_RESOURCE_INTERVAL_MS` s
71
71
  Old staged binaries are re-downloadable from Anthropic's own release channel. For the **container/microvm** tiers the harness needs the **Linux/arm64 ELF**, so download it directly and point the resolver at it:
72
72
 
73
73
  ```bash
74
- V=2.1.209 # your baseline's agentVersion (read it from baselines/desktop-<latest>.json)
74
+ V=2.1.215 # your baseline's agentVersion (read it from baselines/desktop-<latest>.json)
75
75
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "claude-$V"
76
76
  # verify against the committed baseline sha256 (== manifest platforms["linux-arm64"].checksum):
77
77
  shasum -a 256 "claude-$V"
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
16
16
 
17
17
  Run it with:
18
18
 
19
- > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.3.0"`. (`replay` itself needs nothing else — no token, no Docker.)
19
+ > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.4.0"`. (`replay` itself needs nothing else — no token, no Docker.)
20
20
 
21
21
  ```sh
22
22
  cowork-harness replay examples/replays/example-pdf-skill.cassette.json
@@ -100,7 +100,7 @@
100
100
  "preRunHashes": {},
101
101
  "scenarioSource": "../../e2e/scenarios/smoke-multiselect.yaml",
102
102
  "fingerprint": {
103
- "baseline": "1.21459.0"
103
+ "baseline": "1.22209.3"
104
104
  },
105
105
  "timeline": [
106
106
  {
@@ -111,7 +111,7 @@
111
111
  },
112
112
  "scenarioSource": "../scenarios/example-pdf-skill.yaml",
113
113
  "fingerprint": {
114
- "baseline": "1.21459.0",
114
+ "baseline": "1.22209.3",
115
115
  "skillHash": "b760ea90682777367fbd44d866ea66d316452da1abb18bf8cb61f1cec8e67806",
116
116
  "contentSig": "ddf68cf6d74e0599729589f8be4f0a9eec440093c99c1060cd6f47dd671a0dfd",
117
117
  "skillSources": [
@@ -78,7 +78,7 @@
78
78
  "preRunHashes": {},
79
79
  "scenarioSource": "../scenarios/hostloop-computer-links.yaml",
80
80
  "fingerprint": {
81
- "baseline": "1.21459.0"
81
+ "baseline": "1.22209.3"
82
82
  },
83
83
  "timeline": [
84
84
  {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cowork-harness",
3
- "version": "1.3.0",
3
+ "version": "1.4.0",
4
4
  "description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios — same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
5
5
  "license": "MIT",
6
6
  "type": "module",