cowork-harness 4.2.0 → 4.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did the change make its answers worse? (`eval`: paired, interleaved A/B of two plugin versions, pinned models). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 4.2.0
7
- tracks-harness: cowork-harness 4.2.0 (baseline desktop-2.16120.0)
6
+ version: 4.2.1
7
+ tracks-harness: cowork-harness 4.2.1 (baseline desktop-2.16120.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
26
26
  full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
27
  Read them.
28
28
 
29
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.2.0` (baseline
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.2.1` (baseline
30
30
  > `desktop-2.16120.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
31
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
32
32
 
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
43
43
 
44
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
45
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
46
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.2.0"`. **Pin `@^4.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.2.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.2.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.2.1"`. **Pin `@^4.2.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
47
47
 
48
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
49
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Every `assert:` key with its semantics, and the
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Every `assert:` key with its semantics, and the
4
4
  verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
5
5
  the scenario and session YAML fields are there too.
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Assertions guide
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
4
4
 
5
5
  ### Assertions: two orthogonal axes
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Authoring a scenario
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
4
4
 
5
5
  ## Part I — AUTHOR a scenario
6
6
 
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`).
3
+ Self-contained reference. Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
82
82
  GitHub-hosted runners, no token/Docker/agent:
83
83
 
84
84
  ```yaml
85
- - run: npm i -g "cowork-harness@^4.2.0"
85
+ - run: npm i -g "cowork-harness@^4.2.1"
86
86
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
87
87
  # no silent false-greens. WITHOUT --strict this
88
88
  # step cannot fail on a WARN-class rule (e.g.
@@ -395,7 +395,7 @@ jobs:
395
395
  with: { node-version: '24' }
396
396
  - uses: actions/setup-python@v5
397
397
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
398
- - run: npm i -g "cowork-harness@^4.2.0"
398
+ - run: npm i -g "cowork-harness@^4.2.1"
399
399
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
400
400
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
401
401
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -424,7 +424,7 @@ jobs:
424
424
  echo "live=true" >> "$GITHUB_OUTPUT"
425
425
  fi
426
426
  - if: steps.guard.outputs.live == 'true'
427
- run: npm i -g "cowork-harness@^4.2.0"
427
+ run: npm i -g "cowork-harness@^4.2.1"
428
428
  - if: steps.guard.outputs.live == 'true'
429
429
  run: cowork-harness run scenarios/ --output-format json
430
430
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Debugging a run
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
4
4
 
5
5
  ## Part III — Debug
6
6
 
@@ -1,6 +1,6 @@
1
1
  # `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). The full guide is
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). The full guide is
4
4
  [docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
5
5
  need while running it.
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`).
3
+ Self-contained reference. Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -1,6 +1,6 @@
1
1
  # Gotchas
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). The full "✓ passed ≠ correct" landmine catalog.
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). The full "✓ passed ≠ correct" landmine catalog.
4
4
 
5
5
  ## Gotchas — the "✓ passed ≠ correct" landmines
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Measurement
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
4
4
 
5
5
  ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Run, record and lock
2
2
 
3
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
3
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
4
4
 
5
5
  ## Part II — RUN, RECORD & LOCK
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, replay class, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.2.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.2.1`
4
4
  (baseline `desktop-2.16120.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,42 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [4.2.1] — 2026-10-01
10
+
11
+ A security fix: the judge, the LLM decider and the `critique` evaluator run tool-less, isolated from your own Claude
12
+ Code setup. **Needs Claude Code 2.1.197 or later on the host** for any run that uses them.
13
+
14
+ ### Upgrade notes
15
+
16
+ - **Cassettes: no re-record needed.** Nothing under `src/runtime`, `src/hostloop`, `src/staging`, `src/session.ts`,
17
+ `baselines/` or `docker/` changed, and `CASSETTE_VERSION` is still 13. A committed cassette's recorded content and
18
+ fingerprints are unchanged; `verify-cassettes` and `replay --strict` pass on the bundled ones.
19
+
20
+ ### Security
21
+
22
+ - **The LLM judge, the LLM decider and the `critique` evaluator call the host `claude` with no tools, isolated
23
+ from your own Claude Code setup.** All three read untrusted agent output. Up to 4.2.0 the call ran with Claude
24
+ Code's default tools: read-only tools (Read, Glob, Grep) always ran, and Write, Bash and web tools ran whenever
25
+ your settings allowed them — allow rules, an `auto`/`acceptEdits`/`bypassPermissions` default mode, or a
26
+ `PermissionRequest` hook — so text in a judged output could steer a tool call on the machine running the harness.
27
+ The call also loaded your CLAUDE.md, skills, plugins, hooks and MCP servers, and the project settings of the
28
+ directory the harness ran in, and saved a transcript into your session history. Every call passes `--safe-mode`,
29
+ `--strict-mcp-config`, `--no-session-persistence`, `--setting-sources user` and `--tools ""`. Your user settings
30
+ still apply (their `env`, `apiKeyHelper` and model settings), as do managed and policy settings; auth configured
31
+ only in a project's `.claude/settings.json` is not read. On a machine with an enterprise MCP config — a
32
+ `managed-mcp.json` in Claude Code's managed-settings directory
33
+ (`/Library/Application Support/ClaudeCode/` on macOS, `/etc/claude-code/` on Linux, `C:\Program Files\ClaudeCode\`
34
+ on Windows) — the call leaves out `--strict-mcp-config`, which Claude Code refuses beside one; `--safe-mode` still
35
+ keeps every MCP server out, your organisation's managed ones included, while its managed hooks and policy
36
+ settings still apply. A managed config at another path makes Claude Code refuse `--strict-mcp-config`; the call is
37
+ then retried once without it, before any model call.
38
+ **This needs Claude Code 2.1.197 or later on the host.** An older CLI is refused before any model call, saying
39
+ why and what to do. `run`, `record` and `skill` make that check before the agent spends when a scenario
40
+ has a `semantic_matches` assert graded by the host `claude`, or `on_unanswered: llm` / `--decider-llm` with no
41
+ external decider channel; `critique` makes it before its task turn and `decide --decider-llm` before its model
42
+ call — all exit 2. `eval` refuses each such run before its agent spends; the eval still writes its report, with
43
+ every run errored. Runs that use none of them are unaffected.
44
+
9
45
  ## [4.2.0] — 2026-09-30
10
46
 
11
47
  Groundwork for skill hillclimbing: `eval` for paired before/after comparisons of a skill edit, plus
package/README.md CHANGED
@@ -36,7 +36,7 @@ npm ci && npm run build
36
36
  node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
37
37
  ```
38
38
 
39
- (Installing globally — `npm install -g "cowork-harness@^4.2.0"` — gives you the `cowork-harness` CLI for your own
39
+ (Installing globally — `npm install -g "cowork-harness@^4.2.1"` — gives you the `cowork-harness` CLI for your own
40
40
  scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
41
41
 
42
42
  Full setup → [Quick start](./docs/cli.md#quick-start).
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
49
49
 
50
50
  | I want to… | Start here | Needs |
51
51
  |---|---|---|
52
- | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.2.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
- | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.2.0"` |
52
+ | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.2.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
+ | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.2.1"` |
54
54
  | **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v4`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
55
55
 
56
56
  **In short:** three ways in (the table above). Five `fidelity:` tiers — `protocol`, `container`, `microvm`, `hostloop`,
@@ -186,6 +186,7 @@ run, under the constraints it will meet in production".
186
186
  > - **Claude Desktop, opened once** — stages the agent; nothing is bundled.
187
187
  > - **A Claude token** — real per-run cost, runs take minutes; mint one with `claude setup-token` (needs the **`claude` CLI**: `npm i -g @anthropic-ai/claude-code`).
188
188
  > - **A runtime** — **Docker (arm64)** for `container` / `hostloop`, or **Lima (Apple-VZ)** for `microvm`.
189
+ > - **Claude Code 2.1.197 or later** on the host, if a scenario uses the `semantic_matches` judge or the LLM decider, or you run `critique` — those calls run isolated and tool-less, and an older CLI is refused up front.
189
190
  > - The `protocol` tier skips the runtime + the staged agent but still calls a real model, so it still needs the token. Run `cowork-harness doctor --tier <t>` to check exactly what a given tier needs.
190
191
  > - **Platform:** best on **macOS Apple Silicon**; **Windows is not supported** for the live tiers (use the token-free `replay`); `sync` and `microvm` are **macOS-arm64 only**. Full detail in [Prerequisites](./docs/cli.md#prerequisites-for-anything-above-protocol-fidelity) on the CLI page.
191
192
 
package/dist/cli.js CHANGED
@@ -11,7 +11,7 @@ import { loadSession, resolveSessionPaths, applySessionOverrides, resolveLaunchS
11
11
  import { executeScenario, parseScenarioFile, loadSessionFromFile, firstLine, unresolvedModelPreflight, scenarioInputRefusal, UnansweredError, BoundaryError, UsageError, SessionFileError, LegacyRunDirError, effectiveTier, } from "./run/execute.js";
12
12
  import { unresolvedModelRefusal, envModelDefault } from "./run/model-provenance.js";
13
13
  import { ScriptedDecider, ExternalDecider, LlmDecider, ABSTAIN, coerceLabel, } from "./decide/decider.js";
14
- import { claudeCliComplete } from "./decide/llm-transport.js";
14
+ import { claudeCliComplete, isolationRefusal } from "./decide/llm-transport.js";
15
15
  import { vmInit, vmDelete, vmStatus, vmPrune, instanceName, vmProvisioned } from "./runtime/lima.js";
16
16
  import { resolveVmBaselineArg } from "./runtime/vm-baseline-arg.js";
17
17
  import { sync, canonicalizeEnv, syncedNetworkBlock } from "./sync/cowork-sync.js";
@@ -3368,6 +3368,12 @@ async function cmdDecide(args) {
3368
3368
  fail("decide", "usage", "--decider-llm conflicts with --answer/--answer-policy (one terminal decider — the scripted rules would never be used).", undefined, json);
3369
3369
  if (policy)
3370
3370
  rules.push(...loadAnswerPolicy("decide", policy, json));
3371
+ // The LLM decider calls the host `claude`; one that cannot run isolated is a usage refusal (exit 2), like the run's.
3372
+ if (deciderLlm) {
3373
+ const refusal = isolationRefusal();
3374
+ if (refusal)
3375
+ fail("decide", "usage", refusal, undefined, json);
3376
+ }
3371
3377
  const opts = options.length ? options : ["Looks right", "Change it", "Correct or add data"];
3372
3378
  const req = { id: "check", kind: "question", questions: [{ question, options: opts.map((label) => ({ label })) }] };
3373
3379
  const ctx = { task: "", transcript: () => "(sample transcript context)", toolLog: () => [], runId: "decide-check" };
@@ -14,6 +14,7 @@
14
14
  // microvm/protocol stay refused (see ./limitations.ts for why each).
15
15
  // A cross-tier resume is blocked fail-loud by the session-manifest fidelity stamp (src/run/execute.ts),
16
16
  // and at hostloop a writable connected folder requires --allow-host-writes (forwarded to both turns).
17
+ import { isolationRefusal } from "../decide/llm-transport.js";
17
18
  import { spawn } from "node:child_process";
18
19
  import { fileURLToPath } from "node:url";
19
20
  import { lookupSkillFlag } from "../run/skill-flag-surface.js";
@@ -151,7 +152,8 @@ COST AND PREREQUISITES — read before running:
151
152
  total; on a real document-analysis run the ratio INVERTS (measured: task turn ~61%, evaluator ~30%).
152
153
  Read the per-run split off the cost line / costUsd rather than assuming either — a cheaper
153
154
  --evaluator-model buys you at most the evaluator's share, and it voids the armor's
154
- injection-resistance verification, which covers the DEFAULT evaluator only. When the task turn
155
+ injection-resistance verification, which covers the DEFAULT evaluator only (a probe not yet repeated
156
+ with the evaluator tool-less and isolated). When the task turn
155
157
  dominates, the levers are --model, --timeout and probe scope.
156
158
  * container needs Docker/Lima; hostloop needs Docker (the bash/web_fetch sidecar) PLUS the staged native
157
159
  agent binary, and writes to the real host FS (a writable --folder requires --allow-host-writes). Both
@@ -1786,6 +1788,11 @@ async function main(argv = process.argv.slice(2)) {
1786
1788
  const forwardsModel = opts.forwardBoth.some((a) => a === "--model" || a.startsWith("--model="));
1787
1789
  if (!forwardsModel && envModelDefault() === undefined)
1788
1790
  return refuse("usage", `critique: ${unresolvedModelRefusal("this critique's task and reflection turns")}`);
1791
+ // The evaluator runs the host `claude` isolated and tool-less, which needs a CLI that accepts the isolation flags.
1792
+ // Checked here, before the task and reflection turns spend, not at the evaluator's first call after them.
1793
+ const iso = isolationRefusal();
1794
+ if (iso)
1795
+ return refuse("usage", `critique: ${iso}`);
1789
1796
  // Past this point a turn WILL run. parseArgs guarantees a probe on every non-corpus-only line; the
1790
1797
  // narrowing is for the type, not a second validation.
1791
1798
  const prompt = opts.prompt;
@@ -3,7 +3,8 @@ import { extractAllJsonObjects } from "../decide/semantic-judge.js";
3
3
  import { validateCitations } from "./evidence.js";
4
4
  import { armorEvidence, headTag, evidenceOpen, evidenceClose } from "./armor.js";
5
5
  import { ROOT_REFERENCE_SECTION_PREFIX, AGENT_SECTION_PREFIX } from "./package-evidence.js";
6
- // The two-pass, tool-less evaluator. Reuses the shared `claude -p` transport (same reasoning as
6
+ // The two-pass, tool-less evaluator (the transport passes `--tools ""` — see ISOLATION_ARGS in
7
+ // src/decide/llm-transport.ts). Reuses the shared `claude -p` transport (same reasoning as
7
8
  // `semantic-judge.ts`: the harness process itself is not behind the egress proxy, so a direct API call
8
9
  // would bypass the very allowlist the harness enforces; `claude -p` is egress-consistent). Unlike the
9
10
  // judge, this evaluator's output isn't a fixed indexed rubric — it's an open-ended set of findings — so
@@ -1,12 +1,144 @@
1
- import { spawn } from "node:child_process";
1
+ import { spawn, spawnSync } from "node:child_process";
2
+ import { existsSync } from "node:fs";
2
3
  import { assertSpawnAllowed } from "../spawn-guard.js";
3
4
  import { warn, envPositiveNumber } from "../io.js";
4
5
  import { isUsageLimit } from "../usage-limit.js";
6
+ /** Every `claude -p` the harness runs (the LLM judge, the LLM decider, the critique evaluator) runs ISOLATED from the
7
+ * operator's own setup. The model reads untrusted agent output, so it must be able to call no tool (`--tools ""`;
8
+ * without it, read-only tools always run and Write/Bash/Web run whenever the operator's settings allow them), and
9
+ * nothing of the operator's environment may shape its answer: no CLAUDE.md, skills, plugins, hooks or MCP servers
10
+ * (`--safe-mode`, `--strict-mcp-config`), no project or local settings from the harness's working directory
11
+ * (`--setting-sources user` — user settings stay, so `apiKeyHelper` auth keeps working), and no transcript written
12
+ * into the operator's session history (`--no-session-persistence`). Every flag exists in Claude Code 2.1.197 and
13
+ * later; `assertIsolationSupported` refuses an older CLI before any model call.
14
+ *
15
+ * `--tools` takes a variadic value, so it goes LAST: an empty value followed by more flags parses correctly on the
16
+ * CLIs measured, but nothing after it can be swallowed if a future parser reads the variadic list greedily.
17
+ *
18
+ * On a machine with an enterprise MCP config, `isolationArgs` leaves `--strict-mcp-config` out (see there). */
19
+ export const ISOLATION_ARGS = [
20
+ "--safe-mode",
21
+ "--strict-mcp-config",
22
+ "--no-session-persistence",
23
+ "--setting-sources",
24
+ "user",
25
+ "--tools",
26
+ "",
27
+ ];
28
+ /** The flags `ISOLATION_ARGS` needs the host CLI to accept, matched against its `--help`. */
29
+ const ISOLATION_FLAGS = ["--safe-mode", "--strict-mcp-config", "--no-session-persistence", "--setting-sources", "--tools"];
30
+ /** The oldest Claude Code version verified to accept every isolation flag. */
31
+ export const ISOLATION_MIN_CLI = "2.1.197";
32
+ /** True when `help` (a `claude --help` text) declares `flag` as an option: an option line starts with exactly two
33
+ * spaces and the flag, then whitespace, `<`, `=` or the end. Anchored so the flag named inside another option's
34
+ * wrapped description (`--strict-mcp-config` and `--tools` both appear there) does not count. An unrecognised help
35
+ * format matches nothing, so the probe fails CLOSED. A plain line scan, so nothing in `flag` is read as a pattern. */
36
+ export function helpDeclaresFlag(help, flag) {
37
+ const lead = ` ${flag}`;
38
+ return help.split("\n").some((line) => {
39
+ if (!line.startsWith(lead))
40
+ return false;
41
+ const next = line.charAt(lead.length);
42
+ return next === "" || next === "<" || next === "=" || /\s/.test(next);
43
+ });
44
+ }
45
+ /** Where Claude Code reads an enterprise MCP config (`managed-mcp.json` in its managed-settings directory). */
46
+ export function defaultManagedMcpPath(platform = process.platform) {
47
+ if (platform === "darwin")
48
+ return "/Library/Application Support/ClaudeCode/managed-mcp.json";
49
+ if (platform === "win32")
50
+ return "C:\\Program Files\\ClaudeCode\\managed-mcp.json";
51
+ return "/etc/claude-code/managed-mcp.json";
52
+ }
53
+ const isolationChecked = new Map();
54
+ let managedMcpPathOverride;
55
+ /** Binaries whose Claude Code refused `--strict-mcp-config` for an enterprise MCP config at a path this harness does
56
+ * not check: their calls leave the flag out from then on, as for one found at the managed-settings path. */
57
+ const strictMcpRefused = new Set();
58
+ /** The isolation flags for this machine. Claude Code refuses `--strict-mcp-config` outright while an enterprise MCP
59
+ * config is present ("You cannot use --strict-mcp-config when an enterprise MCP config is present"), so there the
60
+ * flag is left out. `--safe-mode` already keeps every MCP server out, the organisation's managed ones included: in
61
+ * safe mode Claude Code's MCP loader returns no servers before it reads the enterprise or managed scope (verified in
62
+ * the 2.1.197 and 2.1.286 binaries), and its safe-mode notice says managed MCP servers do not apply. A managed config
63
+ * elsewhere (Claude Code can be pointed at another managed-settings path) is learnt from the CLI's own refusal; see
64
+ * `claudeCliComplete`. */
65
+ export function isolationArgs(bin, managedMcpPath = managedMcpPathOverride ?? defaultManagedMcpPath()) {
66
+ return existsSync(managedMcpPath) || strictMcpRefused.has(bin)
67
+ ? ISOLATION_ARGS.filter((a) => a !== "--strict-mcp-config")
68
+ : ISOLATION_ARGS;
69
+ }
70
+ /** Refuse a host `claude` that does not accept every isolation flag with a clear, actionable error before any model
71
+ * call, instead of an unknown-option exit retried as if it were transient. One `--help` probe per binary per process
72
+ * (no model call, no stdin). A probe that could not run (missing binary, timeout) is not cached. */
73
+ export function assertIsolationSupported(bin) {
74
+ const cached = isolationChecked.get(bin);
75
+ if (cached !== undefined) {
76
+ if (cached !== true)
77
+ throw cached;
78
+ return;
79
+ }
80
+ const help = spawnSync(bin, ["--help"], { stdio: ["ignore", "pipe", "pipe"], encoding: "utf8", timeout: 15_000 });
81
+ if (help.error) {
82
+ const code = help.error.code;
83
+ // Not a version problem, and not a verdict to remember: the probe itself did not complete.
84
+ if (code === "ETIMEDOUT")
85
+ throw new Error(`the host \`claude\` (${bin}) did not answer \`--help\` within 15s, so the harness cannot confirm it runs judge, decider and ` +
86
+ `critique-evaluator calls isolated — refusing rather than guessing. Retry, or check that ${bin} starts.`);
87
+ throw new Error(`LLM decider transport (${bin} -p) failed to spawn: ${help.error.message} — ensure 'claude' is installed and on PATH, or set COWORK_HARNESS_CLAUDE_BIN to its path`);
88
+ }
89
+ const text = `${help.stdout ?? ""}\n${help.stderr ?? ""}`;
90
+ // A `--help` that crashed or was killed and printed nothing says nothing about the version: refuse, uncached.
91
+ if ((help.status !== 0 || help.signal) && !text.trim())
92
+ throw new Error(`the host \`claude\` (${bin}) printed nothing for \`--help\` (${help.signal ? `killed by ${help.signal}` : `exit ${help.status}`}), so the ` +
93
+ `harness cannot confirm it runs judge, decider and critique-evaluator calls isolated — refusing rather than guessing. ` +
94
+ `Check that ${bin} --help works.`);
95
+ const missing = ISOLATION_FLAGS.filter((f) => !helpDeclaresFlag(text, f));
96
+ let verdict = true;
97
+ if (missing.length) {
98
+ const version = spawnSync(bin, ["--version"], { stdio: ["ignore", "pipe", "pipe"], encoding: "utf8", timeout: 15_000 });
99
+ const v = (version.stdout ?? "").trim() || "unknown version";
100
+ verdict = new Error(`the host \`claude\` (${bin}, ${v}) does not accept ${missing.join(", ")} — ` +
101
+ `the harness runs every judge, decider and critique-evaluator call isolated from your own Claude Code setup, which ` +
102
+ `needs Claude Code ${ISOLATION_MIN_CLI} or later. To fix: upgrade the host \`claude\` (or point COWORK_HARNESS_CLAUDE_BIN at a newer one).`);
103
+ }
104
+ isolationChecked.set(bin, verdict);
105
+ if (verdict !== true)
106
+ throw verdict;
107
+ }
108
+ /** The pre-spend form of `assertIsolationSupported`: the refusal message for the configured host `claude`, or
109
+ * undefined when it can run isolated. For commands that spend (an agent run, critique task turns) before their first
110
+ * judge, decider or evaluator call — they refuse up front (exit 2) instead of after the spend. */
111
+ export function isolationRefusal() {
112
+ // Under the spawn guard nothing may be launched — not even a `--help` probe; the later model call is refused by
113
+ // the guard itself, so there is nothing to pre-empt here.
114
+ try {
115
+ assertSpawnAllowed("the host `claude` isolation probe");
116
+ }
117
+ catch {
118
+ return undefined;
119
+ }
120
+ try {
121
+ assertIsolationSupported(process.env.COWORK_HARNESS_CLAUDE_BIN || "claude");
122
+ return undefined;
123
+ }
124
+ catch (e) {
125
+ return e.message;
126
+ }
127
+ }
128
+ /** Test seam: forget every cached probe, and (optionally) point the enterprise-MCP check at another path so a
129
+ * test controls whether the machine "has" a managed config. */
130
+ export function resetIsolationPreflight(managedMcpPath) {
131
+ managedMcpPathOverride = managedMcpPath;
132
+ strictMcpRefused.clear();
133
+ isolationChecked.clear();
134
+ }
5
135
  /** A spawn rejection the retry wrapper may re-attempt: a TRANSIENT non-zero exit. Timeout / maxBytes /
6
136
  * spawn-ENOENT failures leave this false so they fail loud on the first attempt (see claudeCliComplete).
7
137
  * A usage-limit exit also sets it false — retrying into a spent quota just burns the batch. */
8
138
  class TransportExit extends Error {
9
139
  retryable;
140
+ /** Claude Code refused `--strict-mcp-config` for an enterprise MCP config: retry without that flag. */
141
+ strictMcpRefused = false;
10
142
  constructor(message, retryable = true) {
11
143
  super(message);
12
144
  this.retryable = retryable;
@@ -89,6 +221,14 @@ function tryExtractResultText(raw) {
89
221
  * envelope itself doesn't parse (e.g. a failure that never reached the CLI's own JSON emitter).
90
222
  */
91
223
  function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
224
+ // The backstop for every host-`claude` call, whichever caller reaches it: never spawn the model call on a CLI
225
+ // that would drop (or reject) an isolation flag. Cached per binary, so a retry or a batch probes once.
226
+ try {
227
+ assertIsolationSupported(bin);
228
+ }
229
+ catch (e) {
230
+ return Promise.reject(e);
231
+ }
92
232
  return new Promise((resolve, reject) => {
93
233
  // bound the `claude -p` spawn — a hung-but-alive child would otherwise block the harness forever.
94
234
  // On expiry SIGKILL the child and reject LOUD; clear the timer on close/error so a fast call never leaks it.
@@ -96,7 +236,8 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
96
236
  // The prompt is delivered on STDIN, not argv: an argv prompt is world-readable via `ps` for the life of
97
237
  // the child (verified: `echo '...' | claude -p --output-format json` with no positional prompt reads
98
238
  // from stdin and returns the identical success envelope) — stdin is process-private.
99
- const child = spawn(bin, ["-p", "--model", model, "--output-format", "json"], { stdio: ["pipe", "pipe", "pipe"] });
239
+ const args = isolationArgs(bin);
240
+ const child = spawn(bin, ["-p", "--model", model, "--output-format", "json", ...args], { stdio: ["pipe", "pipe", "pipe"] });
100
241
  // A child that exits/errors before consuming stdin (e.g. ENOENT, or a fake bin that exits immediately)
101
242
  // delivers EPIPE asynchronously as an `error` event on stdin — without a listener Node escalates it to
102
243
  // an uncaughtException. The child's own "error"/"close" handlers below already reject loud, so swallow
@@ -172,6 +313,14 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
172
313
  const o = tail(resultText ?? raw);
173
314
  const e = tail(err);
174
315
  const diag = [o && `stdout: ${o}`, e && `stderr: ${e}`].filter(Boolean).join(" | ");
316
+ // An enterprise MCP config somewhere other than the managed-settings path `isolationArgs` checks: Claude Code
317
+ // refused --strict-mcp-config before any model call. `claudeCliComplete` retries once without the flag.
318
+ if (/cannot use --strict-mcp-config when an enterprise MCP config is present/i.test(`${raw}\n${err}`)) {
319
+ const refused = new TransportExit(`LLM decider transport (${bin} -p): Claude Code refused --strict-mcp-config because an enterprise MCP config is present`, false);
320
+ refused.strictMcpRefused = args.includes("--strict-mcp-config");
321
+ reject(refused);
322
+ return;
323
+ }
175
324
  // Usage/quota limit: don't retry into a spent quota — fail loud & fast so a batch halts.
176
325
  if (resultText && isUsageLimit(resultText, tryExtractApiErrorStatus(raw))) {
177
326
  reject(new TransportExit(`LLM decider transport (${bin} -p): usage/quota limit hit — retry after the limit resets. ${diag}`, false));
@@ -187,8 +336,8 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
187
336
  * (only the spawned agent child is), so a direct API call would bypass the very allowlist the harness
188
337
  * enforces. `claude -p` reuses the run's own auth path and is dogfood-consistent. Requests
189
338
  * `--output-format json` so the resolved model (`modelUsage`) can be recorded for provenance
190
- * even when `model` is a floating alias like `"sonnet"`. One short, tool-less
191
- * call per gate (one call, no recursion into the harness; model is the decider default or --decider-model).
339
+ * even when `model` is a floating alias like `"sonnet"`. One short, isolated, tool-less call per gate (see
340
+ * `ISOLATION_ARGS`) (one call, no recursion into the harness; model is the decider default or --decider-model).
192
341
  *
193
342
  * Non-zero-exit retry: a single `claude -p` spawn can exit non-zero on a TRANSIENT upstream hiccup
194
343
  * (rate-limit/overload/network) during a long back-to-back batch — observed live, not reproducible on demand.
@@ -199,8 +348,8 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
199
348
  * captured stdout names the cause, so we accept it rather than brittle stdout pattern-matching.
200
349
  *
201
350
  * Retry never double-answers: this transport has NO harness side effects (it returns a string; the gate is
202
- * answered exactly once, downstream of a SUCCESSFUL call — a non-zero exit delivers no string), and `claude
203
- * -p` runs headless with no tool approval, so the model call itself is read-only in practice. Only the
351
+ * answered exactly once, downstream of a SUCCESSFUL call — a non-zero exit delivers no string), and the call runs
352
+ * with no tools at all (`--tools ""`, `ISOLATION_ARGS`), so a retried call cannot act twice. Only the
204
353
  * non-zero-exit class retries; timeout / maxBytes-overflow / spawn-ENOENT are not transient and fail loud on
205
354
  * the first attempt. Set `COWORK_HARNESS_LLM_RETRIES=0` to disable (e.g. deterministic CI).
206
355
  */
@@ -229,6 +378,14 @@ export const claudeCliComplete = async (prompt, model) => {
229
378
  catch (e) {
230
379
  const err = e;
231
380
  lastErr = err;
381
+ // Once per binary, and not counted against the retries: the refusal comes before any model call, and the flag
382
+ // only adds to what --safe-mode already keeps out (see `isolationArgs`).
383
+ if (err.strictMcpRefused && !strictMcpRefused.has(bin)) {
384
+ strictMcpRefused.add(bin);
385
+ warn(`${err.message} — running without --strict-mcp-config; --safe-mode keeps every MCP server out`);
386
+ attempt--;
387
+ continue;
388
+ }
232
389
  if (!err.retryable || attempt === retries)
233
390
  throw err;
234
391
  warn(`${err.message} — retrying (attempt ${attempt + 2}/${retries + 1})`);
@@ -55,7 +55,7 @@ import { readTimeline } from "../agent/timeline.js";
55
55
  import { toolDurationFields, foldSkillActivity, attributeSubagentSkills } from "./timeline-fold.js";
56
56
  import { captureSubagentReasoning } from "./subagent-reasoning.js";
57
57
  import { buildDecider, Chain, ExternalDecider, LlmDecider, UnansweredError } from "../decide/decider.js";
58
- import { claudeCliComplete } from "../decide/llm-transport.js";
58
+ import { claudeCliComplete, isolationRefusal } from "../decide/llm-transport.js";
59
59
  import { Run, infraErrorsForResult, evidenceErrorsForResult, unionReferenceAccesses } from "./run.js";
60
60
  import { runsWriteRoot } from "./trace-view.js";
61
61
  import { summarizeGateProvenance } from "./gate-provenance.js";
@@ -703,6 +703,17 @@ export async function executeScenario(scenario, opts = {}) {
703
703
  // and spawns the agent. The unit lane sets COWORK_HARNESS_FORBID_SPAWN so a scenario a refusal should
704
704
  // have caught fails red instead of launching a real agent.
705
705
  assertSpawnAllowed(`scenario "${scenario.name}"`);
706
+ // A run that will call the host `claude` after the agent (the semantic_matches judge, the LLM decider) refuses one
707
+ // that cannot run isolated HERE, before the agent spends anything: those calls run isolated and tool-less, which
708
+ // needs a CLI that accepts the isolation flags. An injected judge
709
+ // (a test or library caller) never reaches the host CLI.
710
+ // An external channel replaces the LLM decider as the terminal, so `on_unanswered: llm` then never calls it.
711
+ if ((scenario.assert.some((a) => a.semantic_matches !== undefined) && !opts.semanticJudge) ||
712
+ (onUnanswered === "llm" && !opts.externalChannel)) {
713
+ const iso = isolationRefusal();
714
+ if (iso)
715
+ throw new UsageError(iso);
716
+ }
706
717
  // Pre-flight: if the skill DECLARES required capabilities and the image provably omits one, FAIL FAST here
707
718
  // — before any paid agent run — instead of burning ~12 min to reach a verdict the post-run guard already
708
719
  // knows. The author can opt out with `allow_missing_capability: true` (the fallback is equivalent), which
package/docs/ci.md CHANGED
@@ -144,7 +144,7 @@ jobs:
144
144
  model: claude-sonnet-5 # used only where a scenario's session sets no `model:`
145
145
  ```
146
146
 
147
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^4` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^4.2.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN; INFO never gates unless `lint` gets `--min-severity INFO` via `extra-args`; a reviewed `lint-skill` finding you accept, such as a size cap, is suppressed with `extra-args: --ignore-rule <rule>[=<glob>]`, which keeps every other WARN gating — see [cli.md](./cli.md#flags-worth-knowing)), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only), `model` (live lane only: exported as `COWORK_HARNESS_MODEL`, which fills in where a session sets no `model:` and no `--model` is passed; a `run` that resolves no model is refused with exit 2). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
147
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^4` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^4.2.1` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN; INFO never gates unless `lint` gets `--min-severity INFO` via `extra-args`; a reviewed `lint-skill` finding you accept, such as a size cap, is suppressed with `extra-args: --ignore-rule <rule>[=<glob>]`, which keeps every other WARN gating — see [cli.md](./cli.md#flags-worth-knowing)), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only), `model` (live lane only: exported as `COWORK_HARNESS_MODEL`, which fills in where a session sets no `model:` and no `--model` is passed; a `run` that resolves no model is refused with exit 2). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
148
148
 
149
149
  CI uses `ANTHROPIC_API_KEY` specifically because there's no interactive browser available to run
150
150
  `claude setup-token`'s OAuth flow in a GitHub Actions runner; locally, the OAuth token is preferred because
package/docs/cli.md CHANGED
@@ -18,7 +18,7 @@ companion skill, CI). This page is the CLI one.
18
18
  **Install from npm:**
19
19
 
20
20
  ```bash
21
- npm install -g "cowork-harness@^4.2.0" # puts the `cowork-harness` command on your PATH
21
+ npm install -g "cowork-harness@^4.2.1" # puts the `cowork-harness` command on your PATH
22
22
  ```
23
23
 
24
24
  **Or build from source:**
@@ -38,7 +38,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
38
38
 
39
39
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
40
40
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
41
- > From a global install (`npm i -g "cowork-harness@^4.2.0"`), point at the package root instead:
41
+ > From a global install (`npm i -g "cowork-harness@^4.2.1"`), point at the package root instead:
42
42
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
43
43
  > (or copy the cassette into your own project and pass that path).
44
44
 
@@ -48,7 +48,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
48
48
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
49
49
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
50
50
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
51
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^4.2.0"`.
51
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^4.2.1"`.
52
52
 
53
53
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
54
54
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -116,6 +116,7 @@ So **Linux live == `container` only**: `microvm` is Apple-VZ (macOS), and `hostl
116
116
  - **Which file supplied your credential:** loading is silent, with one exception. When `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` comes from `<install>/.env` and you are running from another directory, stderr gets one line naming the variables and the file (never the values): `[env] using CLAUDE_CODE_OAUTH_TOKEN from ~/code/cowork-harness/.env (the install's .env, not this directory's)`. A run started with `node <clone>/dist/cli.js` from elsewhere is billed to that clone's credential; to use another, export it, put it in `./.env` or pass `--dotenv`. A `--dotenv` given after the subcommand is read later, so when it replaces that credential a second line names it: `[env] CLAUDE_CODE_OAUTH_TOKEN from ~/work/my.env (replacing the install's .env)`. `COWORK_HARNESS_DEBUG=1` lists every loaded key.
117
117
  - **Placement:** keep `.env` at a working-dir or install root, never inside a mounted skill/project folder.
118
118
  - **Global install:** find the package root with `` `$(npm root -g)/cowork-harness` `` (e.g. `$(npm root -g)/cowork-harness/.env`) — or simpler, just use `--dotenv <path>` / `./.env` in your working directory, which take priority over the package root anyway.
119
+ 4. **Claude Code 2.1.197 or later on the host, for the judge, the LLM decider and `critique`.** A `semantic_matches` assert, `on_unanswered: llm` / `--decider-llm` and `critique` call the host `claude`, isolated from your own setup (see `COWORK_HARNESS_CLAUDE_BIN` under [Advanced / internal escape hatches](#advanced--internal-escape-hatches)); an older CLI is refused before the run spends anything. Runs that use none of them do not need it.
119
120
 
120
121
  > `sync` (below) is **optional for a first run** — the repo ships `baselines/desktop-*.json`, so `baseline: latest` already resolves. Run `sync` only to refresh the platform baseline after Claude Desktop updates. (`sync` is **macOS-only** today; on Linux/Windows use the committed baselines — they work cross-platform.)
121
122
 
@@ -125,7 +126,7 @@ The parts people ask about. **`package.json`'s `files[]` is the exhaustive, mach
125
126
  this table is the readable summary of it, and deliberately omits the infrastructure that always ships
126
127
  (`baselines/`, `schema/`, `fixtures/`, `scripts/`, `docker/`).
127
128
 
128
- | What ships | npm global (`npm install -g "cowork-harness@^4.2.0"`) | Source checkout (`git clone` + `npm ci`) |
129
+ | What ships | npm global (`npm install -g "cowork-harness@^4.2.1"`) | Source checkout (`git clone` + `npm ci`) |
129
130
  |---|---|---|
130
131
  | CLI, `scenario.py` + assertion keys (enough for `lint` in CI) | ✓ | ✓ |
131
132
  | `SKILL.md`, all of `docs/`, `SPEC.md`/`DESIGN.md`/`AGENTS.md` | ✓ | ✓ |
@@ -140,7 +141,7 @@ since a global install puts nothing in your working directory. The matrix, answe
140
141
  examples are the only ones that still need a source checkout. The **marketplace skill install** is
141
142
  narrower again — it pulls only `.claude/skills/cowork-harness/` (SKILL.md + `references/` +
142
143
  `scenario.py`/assertion keys, per `.claude-plugin/marketplace.json`'s `source`); everything in the npm column
143
- arrives when the skill's first command self-bootstraps `npx "cowork-harness@^4.2.0"` — the last row stays
144
+ arrives when the skill's first command self-bootstraps `npx "cowork-harness@^4.2.1"` — the last row stays
144
145
  ✗ either way, since `matrices/`, `answer-policies/` and `probes/` are not published at all. See
145
146
  [docs/companion-skill.md](./companion-skill.md) for that install path.
146
147
 
@@ -677,7 +678,20 @@ Rarely needed.
677
678
 
678
679
  - `PYTHON` — overrides the interpreter for `lint` / scenario tooling (default `python3`).
679
680
  - `COWORK_HARNESS_DEBUG=1` — surfaces which `.env` files were loaded.
680
- - `COWORK_HARNESS_CLAUDE_BIN=<path>` — points the `--decider-llm` transport at a specific `claude` binary.
681
+ - `COWORK_HARNESS_CLAUDE_BIN=<path>` — points the host `claude` calls at a specific binary. Every model call the
682
+ harness makes through it — the `semantic_matches` judge, the LLM decider and the `critique` evaluator — runs
683
+ **isolated from your own Claude Code setup**: no tools (`--tools ""`), no CLAUDE.md, skills, plugins, hooks or MCP
684
+ servers (`--safe-mode`, `--strict-mcp-config`), no project or local settings from the directory the harness runs
685
+ in (`--setting-sources user`), and no session saved (`--no-session-persistence`). Your user settings still apply
686
+ (their `env`, `apiKeyHelper` and model settings), as do managed and policy settings; auth configured only in a
687
+ project's `.claude/settings.json` is not read. On a machine with an enterprise MCP config — a
688
+ `managed-mcp.json` in Claude Code's managed-settings directory
689
+ (`/Library/Application Support/ClaudeCode/` on macOS, `/etc/claude-code/` on Linux, `C:\Program Files\ClaudeCode\`
690
+ on Windows) — the call leaves out `--strict-mcp-config`, which Claude Code refuses beside one; `--safe-mode` still
691
+ keeps every MCP server out, your organisation's managed ones included, while its managed hooks and policy
692
+ settings still apply. A managed config at another path makes Claude Code refuse `--strict-mcp-config`; the call is
693
+ then retried once without it, before any model call. This needs **Claude Code 2.1.197 or later** on
694
+ the host; an older one is refused before any model call, saying why and what to do.
681
695
  - `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` — lets the harness use the newest sibling agent binary when the baseline-pinned version is missing (a fidelity compromise — off by default). A same-major.minor **patch** bump of the staged NATIVE binary is auto-accepted without this flag (the native binary carries no sha256 pin, so a patch drift is safe by default; it prints a loud stderr note naming the pinned and substituted versions). At `hostloop`, and at `cowork` **only when it resolves to host-loop** on the synced baseline, the staged **VM ELF** is auto-accepted on a patch bump too, because on that path it's a non-executed parity mount into the bash sidecar; a `cowork` baseline that resolves to VM-loop instead executes the ELF directly, so it keeps the strict sha-pinned exact-version match, same as `container`/`microvm`, which always keep it (the ELF is the executed agent there), so the flag remains required for any ELF drift on those tiers/paths, and for a major/minor gap everywhere.
682
696
  - `COWORK_MANAGED_CONFIG=1` — forces the managed-config path on `protocol`, and `=0` suppresses the token-derived managed branch there (leaving the `ANTHROPIC_API_KEY` CI path intact); any other value is rejected rather than silently picking a branch.
683
697
  - `COWORK_HARNESS_ALLOW_MISSING_PROMPT=1` — downgrades a missing prompt asset to a warning.
@@ -28,7 +28,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
28
28
  claude plugin install cowork-harness@cowork-harness
29
29
  ```
30
30
 
31
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^4.2.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
31
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^4.2.1"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
32
32
 
33
33
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
34
34
 
@@ -43,7 +43,7 @@ npx skills add yaniv-golan/cowork-harness --skill cowork-harness
43
43
  short entrypoint (routing, the false-green invariants, short workflows); the detail it routes to lives in
44
44
  `references/`, which the agent reads on demand. Everything
45
45
  else (the CLI, `docs/`, the worked examples, the pytest lane) arrives when the skill's first command
46
- self-bootstraps `npx "cowork-harness@^4.2.0"`, which pulls the same npm package as a global install.
46
+ self-bootstraps `npx "cowork-harness@^4.2.1"`, which pulls the same npm package as a global install.
47
47
 
48
48
  For the full package-contents table — what a global install gives you versus a source checkout — see
49
49
  [docs/cli.md → What ships](./cli.md#what-ships). It is maintained there, once.
package/docs/critique.md CHANGED
@@ -283,7 +283,9 @@ It does **not** record their contents — see Known limitations.
283
283
  (gate on `costUsd.complete` — `false` means the total undercounts), then decide. Caveat: the armor's
284
284
  injection-resistance is verified for the
285
285
  shipped **default** evaluator model only — changing it voids that specific verification (matters when
286
- critiquing skills you did not write).
286
+ critiquing skills you did not write). That verification was measured with an evaluator that had tools; it has not
287
+ been repeated with the tool-less evaluator, isolated from your own Claude Code setup (see
288
+ `COWORK_HARNESS_CLAUDE_BIN` in the CLI guide).
287
289
  - **Trending spend across critiques: use the run index, not the reports.** Each critique appends a
288
290
  roll-up row (`critiqueRole:"rollup"`) carrying `critiqueTotalUsd`; its `costUsd` is the evaluator
289
291
  passes only, so `sum(costUsd)` over every row is exactly true spend — the two graded turns already
@@ -576,7 +578,9 @@ planes; it cannot make a reader immune to persuasion. Treat critique output on a
576
578
  which is how you should treat it anyway.
577
579
 
578
580
  Resistance is also **per-model and perishable**: it is verified for the shipped default evaluator model.
579
- Changing the evaluator model invalidates that verification.
581
+ Changing the evaluator model invalidates that verification. It was also measured before the evaluator ran
582
+ tool-less and isolated from your own Claude Code setup; the isolation removes what an injection could reach (no
583
+ tool can run), but the probe has not been repeated under it.
580
584
 
581
585
  This is the same "advisory, not an attestation" property named under [Known limitations](#known-limitations):
582
586
  a skill you did not write can steer the grade, so its output is a lead to run down — never proof.
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
16
16
 
17
17
  Run it with:
18
18
 
19
- > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^4.2.0"`. (`replay` itself needs nothing else — no token, no Docker.)
19
+ > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^4.2.1"`. (`replay` itself needs nothing else — no token, no Docker.)
20
20
 
21
21
  ```sh
22
22
  cowork-harness replay examples/replays/example-pdf-skill.cassette.json
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cowork-harness",
3
- "version": "4.2.0",
3
+ "version": "4.2.1",
4
4
  "description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
5
5
  "license": "MIT",
6
6
  "type": "module",