cowork-harness 4.2.0 → 4.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +4 -4
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +1 -1
- package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
- package/.claude/skills/cowork-harness/references/authoring.md +1 -1
- package/.claude/skills/cowork-harness/references/ci-recipe.md +4 -4
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/debugging.md +1 -1
- package/.claude/skills/cowork-harness/references/eval.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/gotchas.md +1 -1
- package/.claude/skills/cowork-harness/references/measurement.md +1 -1
- package/.claude/skills/cowork-harness/references/run-record-replay.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +36 -0
- package/README.md +4 -3
- package/dist/cli.js +7 -1
- package/dist/critique/command.js +8 -1
- package/dist/critique/evaluator.js +2 -1
- package/dist/decide/llm-transport.js +163 -6
- package/dist/run/execute.js +12 -1
- package/docs/ci.md +1 -1
- package/docs/cli.md +20 -6
- package/docs/companion-skill.md +2 -2
- package/docs/critique.md +6 -2
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did the change make its answers worse? (`eval`: paired, interleaved A/B of two plugin versions, pinned models). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 4.2.
|
|
7
|
-
tracks-harness: cowork-harness 4.2.
|
|
6
|
+
version: 4.2.1
|
|
7
|
+
tracks-harness: cowork-harness 4.2.1 (baseline desktop-2.16120.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
|
|
|
26
26
|
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
27
|
Read them.
|
|
28
28
|
|
|
29
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.2.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.2.1` (baseline
|
|
30
30
|
> `desktop-2.16120.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
31
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
32
32
|
|
|
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
43
43
|
|
|
44
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
45
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
46
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.2.
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.2.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.2.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.2.1"`. **Pin `@^4.2.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
47
47
|
|
|
48
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
49
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertion catalog
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Every `assert:` key with its semantics, and the
|
|
4
4
|
verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
|
|
5
5
|
the scenario and session YAML fields are there too.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertions guide
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
|
|
4
4
|
|
|
5
5
|
### Assertions: two orthogonal axes
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Authoring a scenario
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
|
|
4
4
|
|
|
5
5
|
## Part I — AUTHOR a scenario
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
82
82
|
GitHub-hosted runners, no token/Docker/agent:
|
|
83
83
|
|
|
84
84
|
```yaml
|
|
85
|
-
- run: npm i -g "cowork-harness@^4.2.
|
|
85
|
+
- run: npm i -g "cowork-harness@^4.2.1"
|
|
86
86
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
87
87
|
# no silent false-greens. WITHOUT --strict this
|
|
88
88
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -395,7 +395,7 @@ jobs:
|
|
|
395
395
|
with: { node-version: '24' }
|
|
396
396
|
- uses: actions/setup-python@v5
|
|
397
397
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
398
|
-
- run: npm i -g "cowork-harness@^4.2.
|
|
398
|
+
- run: npm i -g "cowork-harness@^4.2.1"
|
|
399
399
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
400
400
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
401
401
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -424,7 +424,7 @@ jobs:
|
|
|
424
424
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
425
425
|
fi
|
|
426
426
|
- if: steps.guard.outputs.live == 'true'
|
|
427
|
-
run: npm i -g "cowork-harness@^4.2.
|
|
427
|
+
run: npm i -g "cowork-harness@^4.2.1"
|
|
428
428
|
- if: steps.guard.outputs.live == 'true'
|
|
429
429
|
run: cowork-harness run scenarios/ --output-format json
|
|
430
430
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Debugging a run
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
|
|
4
4
|
|
|
5
5
|
## Part III — Debug
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). The full guide is
|
|
4
4
|
[docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
|
|
5
5
|
need while running it.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Gotchas
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). The full "✓ passed ≠ correct" landmine catalog.
|
|
4
4
|
|
|
5
5
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Measurement
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
|
|
4
4
|
|
|
5
5
|
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Run, record and lock
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.2.
|
|
3
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
|
|
4
4
|
|
|
5
5
|
## Part II — RUN, RECORD & LOCK
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, replay class, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.2.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.2.1`
|
|
4
4
|
(baseline `desktop-2.16120.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 4.2.
|
|
5
|
+
Tracks `cowork-harness 4.2.1` (baseline `desktop-2.16120.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,42 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [4.2.1] — 2026-10-01
|
|
10
|
+
|
|
11
|
+
A security fix: the judge, the LLM decider and the `critique` evaluator run tool-less, isolated from your own Claude
|
|
12
|
+
Code setup. **Needs Claude Code 2.1.197 or later on the host** for any run that uses them.
|
|
13
|
+
|
|
14
|
+
### Upgrade notes
|
|
15
|
+
|
|
16
|
+
- **Cassettes: no re-record needed.** Nothing under `src/runtime`, `src/hostloop`, `src/staging`, `src/session.ts`,
|
|
17
|
+
`baselines/` or `docker/` changed, and `CASSETTE_VERSION` is still 13. A committed cassette's recorded content and
|
|
18
|
+
fingerprints are unchanged; `verify-cassettes` and `replay --strict` pass on the bundled ones.
|
|
19
|
+
|
|
20
|
+
### Security
|
|
21
|
+
|
|
22
|
+
- **The LLM judge, the LLM decider and the `critique` evaluator call the host `claude` with no tools, isolated
|
|
23
|
+
from your own Claude Code setup.** All three read untrusted agent output. Up to 4.2.0 the call ran with Claude
|
|
24
|
+
Code's default tools: read-only tools (Read, Glob, Grep) always ran, and Write, Bash and web tools ran whenever
|
|
25
|
+
your settings allowed them — allow rules, an `auto`/`acceptEdits`/`bypassPermissions` default mode, or a
|
|
26
|
+
`PermissionRequest` hook — so text in a judged output could steer a tool call on the machine running the harness.
|
|
27
|
+
The call also loaded your CLAUDE.md, skills, plugins, hooks and MCP servers, and the project settings of the
|
|
28
|
+
directory the harness ran in, and saved a transcript into your session history. Every call passes `--safe-mode`,
|
|
29
|
+
`--strict-mcp-config`, `--no-session-persistence`, `--setting-sources user` and `--tools ""`. Your user settings
|
|
30
|
+
still apply (their `env`, `apiKeyHelper` and model settings), as do managed and policy settings; auth configured
|
|
31
|
+
only in a project's `.claude/settings.json` is not read. On a machine with an enterprise MCP config — a
|
|
32
|
+
`managed-mcp.json` in Claude Code's managed-settings directory
|
|
33
|
+
(`/Library/Application Support/ClaudeCode/` on macOS, `/etc/claude-code/` on Linux, `C:\Program Files\ClaudeCode\`
|
|
34
|
+
on Windows) — the call leaves out `--strict-mcp-config`, which Claude Code refuses beside one; `--safe-mode` still
|
|
35
|
+
keeps every MCP server out, your organisation's managed ones included, while its managed hooks and policy
|
|
36
|
+
settings still apply. A managed config at another path makes Claude Code refuse `--strict-mcp-config`; the call is
|
|
37
|
+
then retried once without it, before any model call.
|
|
38
|
+
**This needs Claude Code 2.1.197 or later on the host.** An older CLI is refused before any model call, saying
|
|
39
|
+
why and what to do. `run`, `record` and `skill` make that check before the agent spends when a scenario
|
|
40
|
+
has a `semantic_matches` assert graded by the host `claude`, or `on_unanswered: llm` / `--decider-llm` with no
|
|
41
|
+
external decider channel; `critique` makes it before its task turn and `decide --decider-llm` before its model
|
|
42
|
+
call — all exit 2. `eval` refuses each such run before its agent spends; the eval still writes its report, with
|
|
43
|
+
every run errored. Runs that use none of them are unaffected.
|
|
44
|
+
|
|
9
45
|
## [4.2.0] — 2026-09-30
|
|
10
46
|
|
|
11
47
|
Groundwork for skill hillclimbing: `eval` for paired before/after comparisons of a skill edit, plus
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^4.2.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^4.2.1"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.2.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.2.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.2.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.2.1"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v4`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
**In short:** three ways in (the table above). Five `fidelity:` tiers — `protocol`, `container`, `microvm`, `hostloop`,
|
|
@@ -186,6 +186,7 @@ run, under the constraints it will meet in production".
|
|
|
186
186
|
> - **Claude Desktop, opened once** — stages the agent; nothing is bundled.
|
|
187
187
|
> - **A Claude token** — real per-run cost, runs take minutes; mint one with `claude setup-token` (needs the **`claude` CLI**: `npm i -g @anthropic-ai/claude-code`).
|
|
188
188
|
> - **A runtime** — **Docker (arm64)** for `container` / `hostloop`, or **Lima (Apple-VZ)** for `microvm`.
|
|
189
|
+
> - **Claude Code 2.1.197 or later** on the host, if a scenario uses the `semantic_matches` judge or the LLM decider, or you run `critique` — those calls run isolated and tool-less, and an older CLI is refused up front.
|
|
189
190
|
> - The `protocol` tier skips the runtime + the staged agent but still calls a real model, so it still needs the token. Run `cowork-harness doctor --tier <t>` to check exactly what a given tier needs.
|
|
190
191
|
> - **Platform:** best on **macOS Apple Silicon**; **Windows is not supported** for the live tiers (use the token-free `replay`); `sync` and `microvm` are **macOS-arm64 only**. Full detail in [Prerequisites](./docs/cli.md#prerequisites-for-anything-above-protocol-fidelity) on the CLI page.
|
|
191
192
|
|
package/dist/cli.js
CHANGED
|
@@ -11,7 +11,7 @@ import { loadSession, resolveSessionPaths, applySessionOverrides, resolveLaunchS
|
|
|
11
11
|
import { executeScenario, parseScenarioFile, loadSessionFromFile, firstLine, unresolvedModelPreflight, scenarioInputRefusal, UnansweredError, BoundaryError, UsageError, SessionFileError, LegacyRunDirError, effectiveTier, } from "./run/execute.js";
|
|
12
12
|
import { unresolvedModelRefusal, envModelDefault } from "./run/model-provenance.js";
|
|
13
13
|
import { ScriptedDecider, ExternalDecider, LlmDecider, ABSTAIN, coerceLabel, } from "./decide/decider.js";
|
|
14
|
-
import { claudeCliComplete } from "./decide/llm-transport.js";
|
|
14
|
+
import { claudeCliComplete, isolationRefusal } from "./decide/llm-transport.js";
|
|
15
15
|
import { vmInit, vmDelete, vmStatus, vmPrune, instanceName, vmProvisioned } from "./runtime/lima.js";
|
|
16
16
|
import { resolveVmBaselineArg } from "./runtime/vm-baseline-arg.js";
|
|
17
17
|
import { sync, canonicalizeEnv, syncedNetworkBlock } from "./sync/cowork-sync.js";
|
|
@@ -3368,6 +3368,12 @@ async function cmdDecide(args) {
|
|
|
3368
3368
|
fail("decide", "usage", "--decider-llm conflicts with --answer/--answer-policy (one terminal decider — the scripted rules would never be used).", undefined, json);
|
|
3369
3369
|
if (policy)
|
|
3370
3370
|
rules.push(...loadAnswerPolicy("decide", policy, json));
|
|
3371
|
+
// The LLM decider calls the host `claude`; one that cannot run isolated is a usage refusal (exit 2), like the run's.
|
|
3372
|
+
if (deciderLlm) {
|
|
3373
|
+
const refusal = isolationRefusal();
|
|
3374
|
+
if (refusal)
|
|
3375
|
+
fail("decide", "usage", refusal, undefined, json);
|
|
3376
|
+
}
|
|
3371
3377
|
const opts = options.length ? options : ["Looks right", "Change it", "Correct or add data"];
|
|
3372
3378
|
const req = { id: "check", kind: "question", questions: [{ question, options: opts.map((label) => ({ label })) }] };
|
|
3373
3379
|
const ctx = { task: "", transcript: () => "(sample transcript context)", toolLog: () => [], runId: "decide-check" };
|
package/dist/critique/command.js
CHANGED
|
@@ -14,6 +14,7 @@
|
|
|
14
14
|
// microvm/protocol stay refused (see ./limitations.ts for why each).
|
|
15
15
|
// A cross-tier resume is blocked fail-loud by the session-manifest fidelity stamp (src/run/execute.ts),
|
|
16
16
|
// and at hostloop a writable connected folder requires --allow-host-writes (forwarded to both turns).
|
|
17
|
+
import { isolationRefusal } from "../decide/llm-transport.js";
|
|
17
18
|
import { spawn } from "node:child_process";
|
|
18
19
|
import { fileURLToPath } from "node:url";
|
|
19
20
|
import { lookupSkillFlag } from "../run/skill-flag-surface.js";
|
|
@@ -151,7 +152,8 @@ COST AND PREREQUISITES — read before running:
|
|
|
151
152
|
total; on a real document-analysis run the ratio INVERTS (measured: task turn ~61%, evaluator ~30%).
|
|
152
153
|
Read the per-run split off the cost line / costUsd rather than assuming either — a cheaper
|
|
153
154
|
--evaluator-model buys you at most the evaluator's share, and it voids the armor's
|
|
154
|
-
injection-resistance verification, which covers the DEFAULT evaluator only
|
|
155
|
+
injection-resistance verification, which covers the DEFAULT evaluator only (a probe not yet repeated
|
|
156
|
+
with the evaluator tool-less and isolated). When the task turn
|
|
155
157
|
dominates, the levers are --model, --timeout and probe scope.
|
|
156
158
|
* container needs Docker/Lima; hostloop needs Docker (the bash/web_fetch sidecar) PLUS the staged native
|
|
157
159
|
agent binary, and writes to the real host FS (a writable --folder requires --allow-host-writes). Both
|
|
@@ -1786,6 +1788,11 @@ async function main(argv = process.argv.slice(2)) {
|
|
|
1786
1788
|
const forwardsModel = opts.forwardBoth.some((a) => a === "--model" || a.startsWith("--model="));
|
|
1787
1789
|
if (!forwardsModel && envModelDefault() === undefined)
|
|
1788
1790
|
return refuse("usage", `critique: ${unresolvedModelRefusal("this critique's task and reflection turns")}`);
|
|
1791
|
+
// The evaluator runs the host `claude` isolated and tool-less, which needs a CLI that accepts the isolation flags.
|
|
1792
|
+
// Checked here, before the task and reflection turns spend, not at the evaluator's first call after them.
|
|
1793
|
+
const iso = isolationRefusal();
|
|
1794
|
+
if (iso)
|
|
1795
|
+
return refuse("usage", `critique: ${iso}`);
|
|
1789
1796
|
// Past this point a turn WILL run. parseArgs guarantees a probe on every non-corpus-only line; the
|
|
1790
1797
|
// narrowing is for the type, not a second validation.
|
|
1791
1798
|
const prompt = opts.prompt;
|
|
@@ -3,7 +3,8 @@ import { extractAllJsonObjects } from "../decide/semantic-judge.js";
|
|
|
3
3
|
import { validateCitations } from "./evidence.js";
|
|
4
4
|
import { armorEvidence, headTag, evidenceOpen, evidenceClose } from "./armor.js";
|
|
5
5
|
import { ROOT_REFERENCE_SECTION_PREFIX, AGENT_SECTION_PREFIX } from "./package-evidence.js";
|
|
6
|
-
// The two-pass, tool-less evaluator
|
|
6
|
+
// The two-pass, tool-less evaluator (the transport passes `--tools ""` — see ISOLATION_ARGS in
|
|
7
|
+
// src/decide/llm-transport.ts). Reuses the shared `claude -p` transport (same reasoning as
|
|
7
8
|
// `semantic-judge.ts`: the harness process itself is not behind the egress proxy, so a direct API call
|
|
8
9
|
// would bypass the very allowlist the harness enforces; `claude -p` is egress-consistent). Unlike the
|
|
9
10
|
// judge, this evaluator's output isn't a fixed indexed rubric — it's an open-ended set of findings — so
|
|
@@ -1,12 +1,144 @@
|
|
|
1
|
-
import { spawn } from "node:child_process";
|
|
1
|
+
import { spawn, spawnSync } from "node:child_process";
|
|
2
|
+
import { existsSync } from "node:fs";
|
|
2
3
|
import { assertSpawnAllowed } from "../spawn-guard.js";
|
|
3
4
|
import { warn, envPositiveNumber } from "../io.js";
|
|
4
5
|
import { isUsageLimit } from "../usage-limit.js";
|
|
6
|
+
/** Every `claude -p` the harness runs (the LLM judge, the LLM decider, the critique evaluator) runs ISOLATED from the
|
|
7
|
+
* operator's own setup. The model reads untrusted agent output, so it must be able to call no tool (`--tools ""`;
|
|
8
|
+
* without it, read-only tools always run and Write/Bash/Web run whenever the operator's settings allow them), and
|
|
9
|
+
* nothing of the operator's environment may shape its answer: no CLAUDE.md, skills, plugins, hooks or MCP servers
|
|
10
|
+
* (`--safe-mode`, `--strict-mcp-config`), no project or local settings from the harness's working directory
|
|
11
|
+
* (`--setting-sources user` — user settings stay, so `apiKeyHelper` auth keeps working), and no transcript written
|
|
12
|
+
* into the operator's session history (`--no-session-persistence`). Every flag exists in Claude Code 2.1.197 and
|
|
13
|
+
* later; `assertIsolationSupported` refuses an older CLI before any model call.
|
|
14
|
+
*
|
|
15
|
+
* `--tools` takes a variadic value, so it goes LAST: an empty value followed by more flags parses correctly on the
|
|
16
|
+
* CLIs measured, but nothing after it can be swallowed if a future parser reads the variadic list greedily.
|
|
17
|
+
*
|
|
18
|
+
* On a machine with an enterprise MCP config, `isolationArgs` leaves `--strict-mcp-config` out (see there). */
|
|
19
|
+
export const ISOLATION_ARGS = [
|
|
20
|
+
"--safe-mode",
|
|
21
|
+
"--strict-mcp-config",
|
|
22
|
+
"--no-session-persistence",
|
|
23
|
+
"--setting-sources",
|
|
24
|
+
"user",
|
|
25
|
+
"--tools",
|
|
26
|
+
"",
|
|
27
|
+
];
|
|
28
|
+
/** The flags `ISOLATION_ARGS` needs the host CLI to accept, matched against its `--help`. */
|
|
29
|
+
const ISOLATION_FLAGS = ["--safe-mode", "--strict-mcp-config", "--no-session-persistence", "--setting-sources", "--tools"];
|
|
30
|
+
/** The oldest Claude Code version verified to accept every isolation flag. */
|
|
31
|
+
export const ISOLATION_MIN_CLI = "2.1.197";
|
|
32
|
+
/** True when `help` (a `claude --help` text) declares `flag` as an option: an option line starts with exactly two
|
|
33
|
+
* spaces and the flag, then whitespace, `<`, `=` or the end. Anchored so the flag named inside another option's
|
|
34
|
+
* wrapped description (`--strict-mcp-config` and `--tools` both appear there) does not count. An unrecognised help
|
|
35
|
+
* format matches nothing, so the probe fails CLOSED. A plain line scan, so nothing in `flag` is read as a pattern. */
|
|
36
|
+
export function helpDeclaresFlag(help, flag) {
|
|
37
|
+
const lead = ` ${flag}`;
|
|
38
|
+
return help.split("\n").some((line) => {
|
|
39
|
+
if (!line.startsWith(lead))
|
|
40
|
+
return false;
|
|
41
|
+
const next = line.charAt(lead.length);
|
|
42
|
+
return next === "" || next === "<" || next === "=" || /\s/.test(next);
|
|
43
|
+
});
|
|
44
|
+
}
|
|
45
|
+
/** Where Claude Code reads an enterprise MCP config (`managed-mcp.json` in its managed-settings directory). */
|
|
46
|
+
export function defaultManagedMcpPath(platform = process.platform) {
|
|
47
|
+
if (platform === "darwin")
|
|
48
|
+
return "/Library/Application Support/ClaudeCode/managed-mcp.json";
|
|
49
|
+
if (platform === "win32")
|
|
50
|
+
return "C:\\Program Files\\ClaudeCode\\managed-mcp.json";
|
|
51
|
+
return "/etc/claude-code/managed-mcp.json";
|
|
52
|
+
}
|
|
53
|
+
const isolationChecked = new Map();
|
|
54
|
+
let managedMcpPathOverride;
|
|
55
|
+
/** Binaries whose Claude Code refused `--strict-mcp-config` for an enterprise MCP config at a path this harness does
|
|
56
|
+
* not check: their calls leave the flag out from then on, as for one found at the managed-settings path. */
|
|
57
|
+
const strictMcpRefused = new Set();
|
|
58
|
+
/** The isolation flags for this machine. Claude Code refuses `--strict-mcp-config` outright while an enterprise MCP
|
|
59
|
+
* config is present ("You cannot use --strict-mcp-config when an enterprise MCP config is present"), so there the
|
|
60
|
+
* flag is left out. `--safe-mode` already keeps every MCP server out, the organisation's managed ones included: in
|
|
61
|
+
* safe mode Claude Code's MCP loader returns no servers before it reads the enterprise or managed scope (verified in
|
|
62
|
+
* the 2.1.197 and 2.1.286 binaries), and its safe-mode notice says managed MCP servers do not apply. A managed config
|
|
63
|
+
* elsewhere (Claude Code can be pointed at another managed-settings path) is learnt from the CLI's own refusal; see
|
|
64
|
+
* `claudeCliComplete`. */
|
|
65
|
+
export function isolationArgs(bin, managedMcpPath = managedMcpPathOverride ?? defaultManagedMcpPath()) {
|
|
66
|
+
return existsSync(managedMcpPath) || strictMcpRefused.has(bin)
|
|
67
|
+
? ISOLATION_ARGS.filter((a) => a !== "--strict-mcp-config")
|
|
68
|
+
: ISOLATION_ARGS;
|
|
69
|
+
}
|
|
70
|
+
/** Refuse a host `claude` that does not accept every isolation flag with a clear, actionable error before any model
|
|
71
|
+
* call, instead of an unknown-option exit retried as if it were transient. One `--help` probe per binary per process
|
|
72
|
+
* (no model call, no stdin). A probe that could not run (missing binary, timeout) is not cached. */
|
|
73
|
+
export function assertIsolationSupported(bin) {
|
|
74
|
+
const cached = isolationChecked.get(bin);
|
|
75
|
+
if (cached !== undefined) {
|
|
76
|
+
if (cached !== true)
|
|
77
|
+
throw cached;
|
|
78
|
+
return;
|
|
79
|
+
}
|
|
80
|
+
const help = spawnSync(bin, ["--help"], { stdio: ["ignore", "pipe", "pipe"], encoding: "utf8", timeout: 15_000 });
|
|
81
|
+
if (help.error) {
|
|
82
|
+
const code = help.error.code;
|
|
83
|
+
// Not a version problem, and not a verdict to remember: the probe itself did not complete.
|
|
84
|
+
if (code === "ETIMEDOUT")
|
|
85
|
+
throw new Error(`the host \`claude\` (${bin}) did not answer \`--help\` within 15s, so the harness cannot confirm it runs judge, decider and ` +
|
|
86
|
+
`critique-evaluator calls isolated — refusing rather than guessing. Retry, or check that ${bin} starts.`);
|
|
87
|
+
throw new Error(`LLM decider transport (${bin} -p) failed to spawn: ${help.error.message} — ensure 'claude' is installed and on PATH, or set COWORK_HARNESS_CLAUDE_BIN to its path`);
|
|
88
|
+
}
|
|
89
|
+
const text = `${help.stdout ?? ""}\n${help.stderr ?? ""}`;
|
|
90
|
+
// A `--help` that crashed or was killed and printed nothing says nothing about the version: refuse, uncached.
|
|
91
|
+
if ((help.status !== 0 || help.signal) && !text.trim())
|
|
92
|
+
throw new Error(`the host \`claude\` (${bin}) printed nothing for \`--help\` (${help.signal ? `killed by ${help.signal}` : `exit ${help.status}`}), so the ` +
|
|
93
|
+
`harness cannot confirm it runs judge, decider and critique-evaluator calls isolated — refusing rather than guessing. ` +
|
|
94
|
+
`Check that ${bin} --help works.`);
|
|
95
|
+
const missing = ISOLATION_FLAGS.filter((f) => !helpDeclaresFlag(text, f));
|
|
96
|
+
let verdict = true;
|
|
97
|
+
if (missing.length) {
|
|
98
|
+
const version = spawnSync(bin, ["--version"], { stdio: ["ignore", "pipe", "pipe"], encoding: "utf8", timeout: 15_000 });
|
|
99
|
+
const v = (version.stdout ?? "").trim() || "unknown version";
|
|
100
|
+
verdict = new Error(`the host \`claude\` (${bin}, ${v}) does not accept ${missing.join(", ")} — ` +
|
|
101
|
+
`the harness runs every judge, decider and critique-evaluator call isolated from your own Claude Code setup, which ` +
|
|
102
|
+
`needs Claude Code ${ISOLATION_MIN_CLI} or later. To fix: upgrade the host \`claude\` (or point COWORK_HARNESS_CLAUDE_BIN at a newer one).`);
|
|
103
|
+
}
|
|
104
|
+
isolationChecked.set(bin, verdict);
|
|
105
|
+
if (verdict !== true)
|
|
106
|
+
throw verdict;
|
|
107
|
+
}
|
|
108
|
+
/** The pre-spend form of `assertIsolationSupported`: the refusal message for the configured host `claude`, or
|
|
109
|
+
* undefined when it can run isolated. For commands that spend (an agent run, critique task turns) before their first
|
|
110
|
+
* judge, decider or evaluator call — they refuse up front (exit 2) instead of after the spend. */
|
|
111
|
+
export function isolationRefusal() {
|
|
112
|
+
// Under the spawn guard nothing may be launched — not even a `--help` probe; the later model call is refused by
|
|
113
|
+
// the guard itself, so there is nothing to pre-empt here.
|
|
114
|
+
try {
|
|
115
|
+
assertSpawnAllowed("the host `claude` isolation probe");
|
|
116
|
+
}
|
|
117
|
+
catch {
|
|
118
|
+
return undefined;
|
|
119
|
+
}
|
|
120
|
+
try {
|
|
121
|
+
assertIsolationSupported(process.env.COWORK_HARNESS_CLAUDE_BIN || "claude");
|
|
122
|
+
return undefined;
|
|
123
|
+
}
|
|
124
|
+
catch (e) {
|
|
125
|
+
return e.message;
|
|
126
|
+
}
|
|
127
|
+
}
|
|
128
|
+
/** Test seam: forget every cached probe, and (optionally) point the enterprise-MCP check at another path so a
|
|
129
|
+
* test controls whether the machine "has" a managed config. */
|
|
130
|
+
export function resetIsolationPreflight(managedMcpPath) {
|
|
131
|
+
managedMcpPathOverride = managedMcpPath;
|
|
132
|
+
strictMcpRefused.clear();
|
|
133
|
+
isolationChecked.clear();
|
|
134
|
+
}
|
|
5
135
|
/** A spawn rejection the retry wrapper may re-attempt: a TRANSIENT non-zero exit. Timeout / maxBytes /
|
|
6
136
|
* spawn-ENOENT failures leave this false so they fail loud on the first attempt (see claudeCliComplete).
|
|
7
137
|
* A usage-limit exit also sets it false — retrying into a spent quota just burns the batch. */
|
|
8
138
|
class TransportExit extends Error {
|
|
9
139
|
retryable;
|
|
140
|
+
/** Claude Code refused `--strict-mcp-config` for an enterprise MCP config: retry without that flag. */
|
|
141
|
+
strictMcpRefused = false;
|
|
10
142
|
constructor(message, retryable = true) {
|
|
11
143
|
super(message);
|
|
12
144
|
this.retryable = retryable;
|
|
@@ -89,6 +221,14 @@ function tryExtractResultText(raw) {
|
|
|
89
221
|
* envelope itself doesn't parse (e.g. a failure that never reached the CLI's own JSON emitter).
|
|
90
222
|
*/
|
|
91
223
|
function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
|
|
224
|
+
// The backstop for every host-`claude` call, whichever caller reaches it: never spawn the model call on a CLI
|
|
225
|
+
// that would drop (or reject) an isolation flag. Cached per binary, so a retry or a batch probes once.
|
|
226
|
+
try {
|
|
227
|
+
assertIsolationSupported(bin);
|
|
228
|
+
}
|
|
229
|
+
catch (e) {
|
|
230
|
+
return Promise.reject(e);
|
|
231
|
+
}
|
|
92
232
|
return new Promise((resolve, reject) => {
|
|
93
233
|
// bound the `claude -p` spawn — a hung-but-alive child would otherwise block the harness forever.
|
|
94
234
|
// On expiry SIGKILL the child and reject LOUD; clear the timer on close/error so a fast call never leaks it.
|
|
@@ -96,7 +236,8 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
|
|
|
96
236
|
// The prompt is delivered on STDIN, not argv: an argv prompt is world-readable via `ps` for the life of
|
|
97
237
|
// the child (verified: `echo '...' | claude -p --output-format json` with no positional prompt reads
|
|
98
238
|
// from stdin and returns the identical success envelope) — stdin is process-private.
|
|
99
|
-
const
|
|
239
|
+
const args = isolationArgs(bin);
|
|
240
|
+
const child = spawn(bin, ["-p", "--model", model, "--output-format", "json", ...args], { stdio: ["pipe", "pipe", "pipe"] });
|
|
100
241
|
// A child that exits/errors before consuming stdin (e.g. ENOENT, or a fake bin that exits immediately)
|
|
101
242
|
// delivers EPIPE asynchronously as an `error` event on stdin — without a listener Node escalates it to
|
|
102
243
|
// an uncaughtException. The child's own "error"/"close" handlers below already reject loud, so swallow
|
|
@@ -172,6 +313,14 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
|
|
|
172
313
|
const o = tail(resultText ?? raw);
|
|
173
314
|
const e = tail(err);
|
|
174
315
|
const diag = [o && `stdout: ${o}`, e && `stderr: ${e}`].filter(Boolean).join(" | ");
|
|
316
|
+
// An enterprise MCP config somewhere other than the managed-settings path `isolationArgs` checks: Claude Code
|
|
317
|
+
// refused --strict-mcp-config before any model call. `claudeCliComplete` retries once without the flag.
|
|
318
|
+
if (/cannot use --strict-mcp-config when an enterprise MCP config is present/i.test(`${raw}\n${err}`)) {
|
|
319
|
+
const refused = new TransportExit(`LLM decider transport (${bin} -p): Claude Code refused --strict-mcp-config because an enterprise MCP config is present`, false);
|
|
320
|
+
refused.strictMcpRefused = args.includes("--strict-mcp-config");
|
|
321
|
+
reject(refused);
|
|
322
|
+
return;
|
|
323
|
+
}
|
|
175
324
|
// Usage/quota limit: don't retry into a spent quota — fail loud & fast so a batch halts.
|
|
176
325
|
if (resultText && isUsageLimit(resultText, tryExtractApiErrorStatus(raw))) {
|
|
177
326
|
reject(new TransportExit(`LLM decider transport (${bin} -p): usage/quota limit hit — retry after the limit resets. ${diag}`, false));
|
|
@@ -187,8 +336,8 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
|
|
|
187
336
|
* (only the spawned agent child is), so a direct API call would bypass the very allowlist the harness
|
|
188
337
|
* enforces. `claude -p` reuses the run's own auth path and is dogfood-consistent. Requests
|
|
189
338
|
* `--output-format json` so the resolved model (`modelUsage`) can be recorded for provenance
|
|
190
|
-
* even when `model` is a floating alias like `"sonnet"`. One short, tool-less
|
|
191
|
-
*
|
|
339
|
+
* even when `model` is a floating alias like `"sonnet"`. One short, isolated, tool-less call per gate (see
|
|
340
|
+
* `ISOLATION_ARGS`) (one call, no recursion into the harness; model is the decider default or --decider-model).
|
|
192
341
|
*
|
|
193
342
|
* Non-zero-exit retry: a single `claude -p` spawn can exit non-zero on a TRANSIENT upstream hiccup
|
|
194
343
|
* (rate-limit/overload/network) during a long back-to-back batch — observed live, not reproducible on demand.
|
|
@@ -199,8 +348,8 @@ function spawnOnce(bin, prompt, model, timeoutMs, maxBytes) {
|
|
|
199
348
|
* captured stdout names the cause, so we accept it rather than brittle stdout pattern-matching.
|
|
200
349
|
*
|
|
201
350
|
* Retry never double-answers: this transport has NO harness side effects (it returns a string; the gate is
|
|
202
|
-
* answered exactly once, downstream of a SUCCESSFUL call — a non-zero exit delivers no string), and
|
|
203
|
-
*
|
|
351
|
+
* answered exactly once, downstream of a SUCCESSFUL call — a non-zero exit delivers no string), and the call runs
|
|
352
|
+
* with no tools at all (`--tools ""`, `ISOLATION_ARGS`), so a retried call cannot act twice. Only the
|
|
204
353
|
* non-zero-exit class retries; timeout / maxBytes-overflow / spawn-ENOENT are not transient and fail loud on
|
|
205
354
|
* the first attempt. Set `COWORK_HARNESS_LLM_RETRIES=0` to disable (e.g. deterministic CI).
|
|
206
355
|
*/
|
|
@@ -229,6 +378,14 @@ export const claudeCliComplete = async (prompt, model) => {
|
|
|
229
378
|
catch (e) {
|
|
230
379
|
const err = e;
|
|
231
380
|
lastErr = err;
|
|
381
|
+
// Once per binary, and not counted against the retries: the refusal comes before any model call, and the flag
|
|
382
|
+
// only adds to what --safe-mode already keeps out (see `isolationArgs`).
|
|
383
|
+
if (err.strictMcpRefused && !strictMcpRefused.has(bin)) {
|
|
384
|
+
strictMcpRefused.add(bin);
|
|
385
|
+
warn(`${err.message} — running without --strict-mcp-config; --safe-mode keeps every MCP server out`);
|
|
386
|
+
attempt--;
|
|
387
|
+
continue;
|
|
388
|
+
}
|
|
232
389
|
if (!err.retryable || attempt === retries)
|
|
233
390
|
throw err;
|
|
234
391
|
warn(`${err.message} — retrying (attempt ${attempt + 2}/${retries + 1})`);
|
package/dist/run/execute.js
CHANGED
|
@@ -55,7 +55,7 @@ import { readTimeline } from "../agent/timeline.js";
|
|
|
55
55
|
import { toolDurationFields, foldSkillActivity, attributeSubagentSkills } from "./timeline-fold.js";
|
|
56
56
|
import { captureSubagentReasoning } from "./subagent-reasoning.js";
|
|
57
57
|
import { buildDecider, Chain, ExternalDecider, LlmDecider, UnansweredError } from "../decide/decider.js";
|
|
58
|
-
import { claudeCliComplete } from "../decide/llm-transport.js";
|
|
58
|
+
import { claudeCliComplete, isolationRefusal } from "../decide/llm-transport.js";
|
|
59
59
|
import { Run, infraErrorsForResult, evidenceErrorsForResult, unionReferenceAccesses } from "./run.js";
|
|
60
60
|
import { runsWriteRoot } from "./trace-view.js";
|
|
61
61
|
import { summarizeGateProvenance } from "./gate-provenance.js";
|
|
@@ -703,6 +703,17 @@ export async function executeScenario(scenario, opts = {}) {
|
|
|
703
703
|
// and spawns the agent. The unit lane sets COWORK_HARNESS_FORBID_SPAWN so a scenario a refusal should
|
|
704
704
|
// have caught fails red instead of launching a real agent.
|
|
705
705
|
assertSpawnAllowed(`scenario "${scenario.name}"`);
|
|
706
|
+
// A run that will call the host `claude` after the agent (the semantic_matches judge, the LLM decider) refuses one
|
|
707
|
+
// that cannot run isolated HERE, before the agent spends anything: those calls run isolated and tool-less, which
|
|
708
|
+
// needs a CLI that accepts the isolation flags. An injected judge
|
|
709
|
+
// (a test or library caller) never reaches the host CLI.
|
|
710
|
+
// An external channel replaces the LLM decider as the terminal, so `on_unanswered: llm` then never calls it.
|
|
711
|
+
if ((scenario.assert.some((a) => a.semantic_matches !== undefined) && !opts.semanticJudge) ||
|
|
712
|
+
(onUnanswered === "llm" && !opts.externalChannel)) {
|
|
713
|
+
const iso = isolationRefusal();
|
|
714
|
+
if (iso)
|
|
715
|
+
throw new UsageError(iso);
|
|
716
|
+
}
|
|
706
717
|
// Pre-flight: if the skill DECLARES required capabilities and the image provably omits one, FAIL FAST here
|
|
707
718
|
// — before any paid agent run — instead of burning ~12 min to reach a verdict the post-run guard already
|
|
708
719
|
// knows. The author can opt out with `allow_missing_capability: true` (the fallback is equivalent), which
|
package/docs/ci.md
CHANGED
|
@@ -144,7 +144,7 @@ jobs:
|
|
|
144
144
|
model: claude-sonnet-5 # used only where a scenario's session sets no `model:`
|
|
145
145
|
```
|
|
146
146
|
|
|
147
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^4` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^4.2.
|
|
147
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^4` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^4.2.1` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN; INFO never gates unless `lint` gets `--min-severity INFO` via `extra-args`; a reviewed `lint-skill` finding you accept, such as a size cap, is suppressed with `extra-args: --ignore-rule <rule>[=<glob>]`, which keeps every other WARN gating — see [cli.md](./cli.md#flags-worth-knowing)), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only), `model` (live lane only: exported as `COWORK_HARNESS_MODEL`, which fills in where a session sets no `model:` and no `--model` is passed; a `run` that resolves no model is refused with exit 2). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
148
148
|
|
|
149
149
|
CI uses `ANTHROPIC_API_KEY` specifically because there's no interactive browser available to run
|
|
150
150
|
`claude setup-token`'s OAuth flow in a GitHub Actions runner; locally, the OAuth token is preferred because
|
package/docs/cli.md
CHANGED
|
@@ -18,7 +18,7 @@ companion skill, CI). This page is the CLI one.
|
|
|
18
18
|
**Install from npm:**
|
|
19
19
|
|
|
20
20
|
```bash
|
|
21
|
-
npm install -g "cowork-harness@^4.2.
|
|
21
|
+
npm install -g "cowork-harness@^4.2.1" # puts the `cowork-harness` command on your PATH
|
|
22
22
|
```
|
|
23
23
|
|
|
24
24
|
**Or build from source:**
|
|
@@ -38,7 +38,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
38
38
|
|
|
39
39
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
40
40
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
41
|
-
> From a global install (`npm i -g "cowork-harness@^4.2.
|
|
41
|
+
> From a global install (`npm i -g "cowork-harness@^4.2.1"`), point at the package root instead:
|
|
42
42
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
43
43
|
> (or copy the cassette into your own project and pass that path).
|
|
44
44
|
|
|
@@ -48,7 +48,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
48
48
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
49
49
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
50
50
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
51
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^4.2.
|
|
51
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^4.2.1"`.
|
|
52
52
|
|
|
53
53
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
54
54
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -116,6 +116,7 @@ So **Linux live == `container` only**: `microvm` is Apple-VZ (macOS), and `hostl
|
|
|
116
116
|
- **Which file supplied your credential:** loading is silent, with one exception. When `CLAUDE_CODE_OAUTH_TOKEN`, `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` comes from `<install>/.env` and you are running from another directory, stderr gets one line naming the variables and the file (never the values): `[env] using CLAUDE_CODE_OAUTH_TOKEN from ~/code/cowork-harness/.env (the install's .env, not this directory's)`. A run started with `node <clone>/dist/cli.js` from elsewhere is billed to that clone's credential; to use another, export it, put it in `./.env` or pass `--dotenv`. A `--dotenv` given after the subcommand is read later, so when it replaces that credential a second line names it: `[env] CLAUDE_CODE_OAUTH_TOKEN from ~/work/my.env (replacing the install's .env)`. `COWORK_HARNESS_DEBUG=1` lists every loaded key.
|
|
117
117
|
- **Placement:** keep `.env` at a working-dir or install root, never inside a mounted skill/project folder.
|
|
118
118
|
- **Global install:** find the package root with `` `$(npm root -g)/cowork-harness` `` (e.g. `$(npm root -g)/cowork-harness/.env`) — or simpler, just use `--dotenv <path>` / `./.env` in your working directory, which take priority over the package root anyway.
|
|
119
|
+
4. **Claude Code 2.1.197 or later on the host, for the judge, the LLM decider and `critique`.** A `semantic_matches` assert, `on_unanswered: llm` / `--decider-llm` and `critique` call the host `claude`, isolated from your own setup (see `COWORK_HARNESS_CLAUDE_BIN` under [Advanced / internal escape hatches](#advanced--internal-escape-hatches)); an older CLI is refused before the run spends anything. Runs that use none of them do not need it.
|
|
119
120
|
|
|
120
121
|
> `sync` (below) is **optional for a first run** — the repo ships `baselines/desktop-*.json`, so `baseline: latest` already resolves. Run `sync` only to refresh the platform baseline after Claude Desktop updates. (`sync` is **macOS-only** today; on Linux/Windows use the committed baselines — they work cross-platform.)
|
|
121
122
|
|
|
@@ -125,7 +126,7 @@ The parts people ask about. **`package.json`'s `files[]` is the exhaustive, mach
|
|
|
125
126
|
this table is the readable summary of it, and deliberately omits the infrastructure that always ships
|
|
126
127
|
(`baselines/`, `schema/`, `fixtures/`, `scripts/`, `docker/`).
|
|
127
128
|
|
|
128
|
-
| What ships | npm global (`npm install -g "cowork-harness@^4.2.
|
|
129
|
+
| What ships | npm global (`npm install -g "cowork-harness@^4.2.1"`) | Source checkout (`git clone` + `npm ci`) |
|
|
129
130
|
|---|---|---|
|
|
130
131
|
| CLI, `scenario.py` + assertion keys (enough for `lint` in CI) | ✓ | ✓ |
|
|
131
132
|
| `SKILL.md`, all of `docs/`, `SPEC.md`/`DESIGN.md`/`AGENTS.md` | ✓ | ✓ |
|
|
@@ -140,7 +141,7 @@ since a global install puts nothing in your working directory. The matrix, answe
|
|
|
140
141
|
examples are the only ones that still need a source checkout. The **marketplace skill install** is
|
|
141
142
|
narrower again — it pulls only `.claude/skills/cowork-harness/` (SKILL.md + `references/` +
|
|
142
143
|
`scenario.py`/assertion keys, per `.claude-plugin/marketplace.json`'s `source`); everything in the npm column
|
|
143
|
-
arrives when the skill's first command self-bootstraps `npx "cowork-harness@^4.2.
|
|
144
|
+
arrives when the skill's first command self-bootstraps `npx "cowork-harness@^4.2.1"` — the last row stays
|
|
144
145
|
✗ either way, since `matrices/`, `answer-policies/` and `probes/` are not published at all. See
|
|
145
146
|
[docs/companion-skill.md](./companion-skill.md) for that install path.
|
|
146
147
|
|
|
@@ -677,7 +678,20 @@ Rarely needed.
|
|
|
677
678
|
|
|
678
679
|
- `PYTHON` — overrides the interpreter for `lint` / scenario tooling (default `python3`).
|
|
679
680
|
- `COWORK_HARNESS_DEBUG=1` — surfaces which `.env` files were loaded.
|
|
680
|
-
- `COWORK_HARNESS_CLAUDE_BIN=<path>` — points the
|
|
681
|
+
- `COWORK_HARNESS_CLAUDE_BIN=<path>` — points the host `claude` calls at a specific binary. Every model call the
|
|
682
|
+
harness makes through it — the `semantic_matches` judge, the LLM decider and the `critique` evaluator — runs
|
|
683
|
+
**isolated from your own Claude Code setup**: no tools (`--tools ""`), no CLAUDE.md, skills, plugins, hooks or MCP
|
|
684
|
+
servers (`--safe-mode`, `--strict-mcp-config`), no project or local settings from the directory the harness runs
|
|
685
|
+
in (`--setting-sources user`), and no session saved (`--no-session-persistence`). Your user settings still apply
|
|
686
|
+
(their `env`, `apiKeyHelper` and model settings), as do managed and policy settings; auth configured only in a
|
|
687
|
+
project's `.claude/settings.json` is not read. On a machine with an enterprise MCP config — a
|
|
688
|
+
`managed-mcp.json` in Claude Code's managed-settings directory
|
|
689
|
+
(`/Library/Application Support/ClaudeCode/` on macOS, `/etc/claude-code/` on Linux, `C:\Program Files\ClaudeCode\`
|
|
690
|
+
on Windows) — the call leaves out `--strict-mcp-config`, which Claude Code refuses beside one; `--safe-mode` still
|
|
691
|
+
keeps every MCP server out, your organisation's managed ones included, while its managed hooks and policy
|
|
692
|
+
settings still apply. A managed config at another path makes Claude Code refuse `--strict-mcp-config`; the call is
|
|
693
|
+
then retried once without it, before any model call. This needs **Claude Code 2.1.197 or later** on
|
|
694
|
+
the host; an older one is refused before any model call, saying why and what to do.
|
|
681
695
|
- `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` — lets the harness use the newest sibling agent binary when the baseline-pinned version is missing (a fidelity compromise — off by default). A same-major.minor **patch** bump of the staged NATIVE binary is auto-accepted without this flag (the native binary carries no sha256 pin, so a patch drift is safe by default; it prints a loud stderr note naming the pinned and substituted versions). At `hostloop`, and at `cowork` **only when it resolves to host-loop** on the synced baseline, the staged **VM ELF** is auto-accepted on a patch bump too, because on that path it's a non-executed parity mount into the bash sidecar; a `cowork` baseline that resolves to VM-loop instead executes the ELF directly, so it keeps the strict sha-pinned exact-version match, same as `container`/`microvm`, which always keep it (the ELF is the executed agent there), so the flag remains required for any ELF drift on those tiers/paths, and for a major/minor gap everywhere.
|
|
682
696
|
- `COWORK_MANAGED_CONFIG=1` — forces the managed-config path on `protocol`, and `=0` suppresses the token-derived managed branch there (leaving the `ANTHROPIC_API_KEY` CI path intact); any other value is rejected rather than silently picking a branch.
|
|
683
697
|
- `COWORK_HARNESS_ALLOW_MISSING_PROMPT=1` — downgrades a missing prompt asset to a warning.
|
package/docs/companion-skill.md
CHANGED
|
@@ -28,7 +28,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
28
28
|
claude plugin install cowork-harness@cowork-harness
|
|
29
29
|
```
|
|
30
30
|
|
|
31
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^4.2.
|
|
31
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^4.2.1"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
32
32
|
|
|
33
33
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
34
34
|
|
|
@@ -43,7 +43,7 @@ npx skills add yaniv-golan/cowork-harness --skill cowork-harness
|
|
|
43
43
|
short entrypoint (routing, the false-green invariants, short workflows); the detail it routes to lives in
|
|
44
44
|
`references/`, which the agent reads on demand. Everything
|
|
45
45
|
else (the CLI, `docs/`, the worked examples, the pytest lane) arrives when the skill's first command
|
|
46
|
-
self-bootstraps `npx "cowork-harness@^4.2.
|
|
46
|
+
self-bootstraps `npx "cowork-harness@^4.2.1"`, which pulls the same npm package as a global install.
|
|
47
47
|
|
|
48
48
|
For the full package-contents table — what a global install gives you versus a source checkout — see
|
|
49
49
|
[docs/cli.md → What ships](./cli.md#what-ships). It is maintained there, once.
|
package/docs/critique.md
CHANGED
|
@@ -283,7 +283,9 @@ It does **not** record their contents — see Known limitations.
|
|
|
283
283
|
(gate on `costUsd.complete` — `false` means the total undercounts), then decide. Caveat: the armor's
|
|
284
284
|
injection-resistance is verified for the
|
|
285
285
|
shipped **default** evaluator model only — changing it voids that specific verification (matters when
|
|
286
|
-
critiquing skills you did not write).
|
|
286
|
+
critiquing skills you did not write). That verification was measured with an evaluator that had tools; it has not
|
|
287
|
+
been repeated with the tool-less evaluator, isolated from your own Claude Code setup (see
|
|
288
|
+
`COWORK_HARNESS_CLAUDE_BIN` in the CLI guide).
|
|
287
289
|
- **Trending spend across critiques: use the run index, not the reports.** Each critique appends a
|
|
288
290
|
roll-up row (`critiqueRole:"rollup"`) carrying `critiqueTotalUsd`; its `costUsd` is the evaluator
|
|
289
291
|
passes only, so `sum(costUsd)` over every row is exactly true spend — the two graded turns already
|
|
@@ -576,7 +578,9 @@ planes; it cannot make a reader immune to persuasion. Treat critique output on a
|
|
|
576
578
|
which is how you should treat it anyway.
|
|
577
579
|
|
|
578
580
|
Resistance is also **per-model and perishable**: it is verified for the shipped default evaluator model.
|
|
579
|
-
Changing the evaluator model invalidates that verification.
|
|
581
|
+
Changing the evaluator model invalidates that verification. It was also measured before the evaluator ran
|
|
582
|
+
tool-less and isolated from your own Claude Code setup; the isolation removes what an injection could reach (no
|
|
583
|
+
tool can run), but the probe has not been repeated under it.
|
|
580
584
|
|
|
581
585
|
This is the same "advisory, not an attestation" property named under [Known limitations](#known-limitations):
|
|
582
586
|
a skill you did not write can steer the grade, so its output is a lead to run down — never proof.
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^4.2.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^4.2.1"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "4.2.
|
|
3
|
+
"version": "4.2.1",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|