cowork-harness 1.3.0 → 1.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +30 -11
- package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +16 -3
- package/CHANGELOG.md +93 -0
- package/README.md +9 -8
- package/SPEC.md +2 -2
- package/baselines/desktop-1.22209.3.json +392 -0
- package/dist/cli.js +122 -75
- package/dist/critique/armor.js +44 -0
- package/dist/critique/command.js +676 -0
- package/dist/critique/evaluator.js +388 -0
- package/dist/critique/evidence.js +175 -0
- package/dist/critique/package-evidence.js +220 -0
- package/dist/hostloop/webfetch-dedup.js +50 -0
- package/dist/hostloop/workspace-handler.js +42 -4
- package/dist/loop-decision.js +30 -0
- package/dist/run/cassette.js +30 -14
- package/dist/run/chat-result.js +1 -0
- package/dist/run/chat.js +10 -1
- package/dist/run/diff.js +13 -1
- package/dist/run/envelope.js +5 -1
- package/dist/run/execute.js +20 -5
- package/dist/run/outcome.js +15 -0
- package/dist/run/repeat-flags.js +70 -0
- package/dist/runtime/hostloop.js +1 -0
- package/docs/README.md +1 -0
- package/docs/critique.md +117 -0
- package/docs/debugging.md +62 -4
- package/docs/gotchas.md +3 -2
- package/docs/maintenance.md +1 -1
- package/docs/scenario.md +3 -2
- package/docs/stats.md +43 -1
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/llms.txt +2 -1
- package/package.json +3 -4
- package/schema/run-result.json +160 -54
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.5.0
|
|
7
|
+
tracks-harness: cowork-harness 1.5.0 (baseline desktop-1.22209.3)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
26
|
-
> `desktop-1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.5.0` (baseline
|
|
26
|
+
> `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,10 +39,20 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
43
|
-
|
|
44
|
-
What the ≥ 1.
|
|
45
|
-
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.5.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.5.0"`. **Pin `@>=1.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
|
+
|
|
44
|
+
What the ≥ 1.5.0 floor gates, by release:
|
|
45
|
+
|
|
46
|
+
- **1.5.0:** `critique <skill-folder> --prompt "<probe>"` (EXPERIMENTAL) — runs a skill, asks the agent
|
|
47
|
+
what confused it, then grades that self-report against a frozen record of the run: a blinded evaluator
|
|
48
|
+
plus mechanical citation checking. Findings NEVER gate (exit 0); exit 2 means no critique was produced.
|
|
49
|
+
`skill --repeat N` (2-100) with `--min-pass-rate`/`--stop-on-diverge`/`--max-budget-usd`/
|
|
50
|
+
`--allow-budget-stop` brings the variance rollup to the exploratory lane. `result.json` gains
|
|
51
|
+
`outcome` (`errored`/`no_deliverable`/`delivered_with_verdict_fail`/`delivered_clean`). The
|
|
52
|
+
`skill`/`probe-dispatch` lanes now emit `fingerprint.skillHash`/`skillCommit` at all — before 1.5.0
|
|
53
|
+
they emitted NEITHER, so a generation-pairing step over those lanes silently grouped on an absent key.
|
|
54
|
+
`diff`'s exit code now honours its own documented gateable signal (tools/artifacts/meta; a
|
|
55
|
+
transcript-only difference no longer fails it).
|
|
46
56
|
- **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
|
|
47
57
|
- **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
|
|
48
58
|
- **0.22.0:** `computer_links_resolve`.
|
|
@@ -55,6 +65,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
55
65
|
- **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
|
|
56
66
|
- **1.2.0:** three new assertion keys — `no_lost_write_back: true` (the write-back detector above wired as a per-scenario gate over the run's authored files, live-only) and the regex siblings `tool_result_matches`/`tool_result_not_matches` (case-insensitive per-result, for an error-signature *family* a literal substring can't express). **`microvm` outputs are now observable** — its session tree is snapshotted from the VM into the run dir, so `file_exists`/`artifact_json`/`user_visible_artifact`/`no_unexpected_files`/`input_unmodified`/`no_lost_write_back`/`semantic_matches` all work there, no longer `container`/`hostloop`-only. `status <dir>` also resolves the newest session under a `--run-dir` root; the completion footer prints a `→ result: …/result.json` pointer; the `on_unanswered=fail` error also points to `on_unanswered: llm`. `analyze-skill` hardening: a phantom `<script>` prose block no longer sinks a real verdict, a delete/remove flow claiming success classifies as lost (error) not just suspect, and write-back detection widened (optional-call `?.` spellings, member-spelled/aliased `fetch`/`sendBeacon`, axios instance/config forms).
|
|
57
67
|
- **1.3.0:** run-identity for the iterate-across-fixes loop — `skill`/`run` take `--label <tag>` (a generation tag surfaced in `result.json` `runLabel`, the run-index row, `inspect`, and `status.json`) alongside an auto-recorded `skillCommit`, layered on the **authoritative** content-exact `fingerprint.skillHash` a harvest step should pair critiques by (`inspect`/index surface a short `skillHash` prefix). `trace --full-results` captures the full input+result of **every** tool call — successful ones too, not just errors — so an external grader can ground a self-critique finding against the call it cites. `verify-run` now **warns** on skillHash drift for answer-less scenarios (was silent unless the scenario declared scripted `answers`). Plus `skill --allow-missing-capability` (the open-ended-run opt-out for a capability FALSE-NEGATIVE on the lean `core` image), a new **warn-severity `ended_with_question`** verdict signal (the agent's final answer contains a question and the run wrote no `outputs/` deliverable — the lenient sibling of the strict `stalled`), and LLM-decider `OTHER:` free-text answers now marked `[via Other free-text]` in gate provenance.
|
|
68
|
+
- **1.4.0:** platform baseline synced to Desktop **1.22209.3** (agent 2.1.215; no prompt/spawn/egress drift), and **`coworkWebFetchDedup` is now enacted** on the hostloop `web_fetch` path — a repeat fetch of the same normalized URL within the TTL (15 min, cap 100) makes **no network request** and returns a marker to re-use the earlier result, with no egress event. Baseline-gated (Desktop ≥ 1.22209.3), keyed under the request + terminal `destination_url`, never caching errors/empty; a hit is assertable via `tool_result_contains: "Already fetched"`.
|
|
58
69
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
59
70
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
60
71
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
@@ -85,9 +96,14 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
85
96
|
- **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
|
|
86
97
|
TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
|
|
87
98
|
|
|
99
|
+
> **"repo-only" in this skill means "not bundled with the installed SKILL"** — not "unavailable". An
|
|
100
|
+
> **npm** install ships `docs/`, `README.md` and `SPEC.md` in the tarball, so try
|
|
101
|
+
> `node_modules/cowork-harness/docs/<name>.md` before assuming a pointer dangles. A **plugin**
|
|
102
|
+
> install loads a trimmed source-only cache where those pointers genuinely do dangle.
|
|
103
|
+
|
|
88
104
|
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
|
|
89
105
|
lint-skill · analyze-skill · probe-dispatch ·
|
|
90
|
-
verify-run · trace · inspect · diff · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
106
|
+
verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
91
107
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
92
108
|
|
|
93
109
|
**Two different `scaffold` tools — don't confuse them.** The native `cowork-harness scaffold <run-id>`
|
|
@@ -365,7 +381,10 @@ cassette — has its own recipe:
|
|
|
365
381
|
input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
|
|
366
382
|
pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
|
|
367
383
|
produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
|
|
368
|
-
kept run predates the current skill).
|
|
384
|
+
kept run predates the current skill). **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
|
|
385
|
+
pairing step there silently groups on an absent key instead of erroring — check the field is present, or
|
|
386
|
+
require ≥ 1.5.0. See `docs/debugging.md`
|
|
387
|
+
(repo-only) for the full loop.
|
|
369
388
|
|
|
370
389
|
#### Interpreting verdict signals
|
|
371
390
|
|
|
@@ -461,7 +480,7 @@ Docker, no re-record.
|
|
|
461
480
|
| Situation | Symptom | Reach for (in order) |
|
|
462
481
|
|---|---|---|
|
|
463
482
|
| **The skill misbehaved** | wrong output, an unexpected gate, a denied tool, an opaque crash | `inspect` — what did it produce? · `trace <run-dir> --view <view>` — what did it actually do (tools, gates, sub-agent tree)? · `verify-run` — re-assert cheaply when only an assertion is wrong · `diff <old-run> <new-run>` — what changed since it worked · `chat` — reproduce it by hand |
|
|
464
|
-
| **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
|
|
483
|
+
| **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` / `skill --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
|
|
465
484
|
|
|
466
485
|
A failed run also records `errorSource` (where the failure originated) and `stderrLogPath` (the captured
|
|
467
486
|
agent stderr) — read those before re-running; a re-record rarely tells you more than the captured stderr
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.5.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.5.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -32,7 +32,7 @@ jobs:
|
|
|
32
32
|
- uses: actions/checkout@v4
|
|
33
33
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
34
34
|
run: |
|
|
35
|
-
V=2.1.
|
|
35
|
+
V=2.1.215 # match your scenario's pinned baseline's agentVersion
|
|
36
36
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
37
37
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
38
38
|
# verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.5.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.5.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.5.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.5.0`
|
|
4
4
|
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
|
|
5
|
-
front-matter (currently
|
|
5
|
+
front-matter (currently 1.5.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -170,7 +170,17 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
170
170
|
|
|
171
171
|
1. **Verify before you trust.** A green run is not a correct run, and a skill's self-reported finding (a
|
|
172
172
|
self-critique appendix, "I extracted X") is not real until its cited evidence is found in the run's own
|
|
173
|
-
output.
|
|
173
|
+
output. **Reproduce before acting on a finding:** `cowork-harness skill <folder> "<prompt>" --repeat 5 --label gen-1`
|
|
174
|
+
runs the same skill+prompt N times (2-100) and prints a variance rollup instead of a single pass/fail —
|
|
175
|
+
`--repeat` works on the `skill` lane, not just `run`. A single green run proves it passed *once*.
|
|
176
|
+
Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`. It rejects `--session-id`/
|
|
177
|
+
`--resume` (both pin one run dir) and `--decider-cmd`/`--decider-dir` (a driving agent x N is not a
|
|
178
|
+
measurement).
|
|
179
|
+
The full loop (harvest -> reproduce -> fix -> prove freshness -> compare) is written out end-to-end in
|
|
180
|
+
docs/debugging.md under "The whole loop, end to end". The harness now SHIPS a grader — `cowork-harness critique <skill-folder> --prompt "<probe>"` runs the
|
|
181
|
+
skill, asks the agent what confused it, and grades that self-report against a frozen record of the run
|
|
182
|
+
(blinded evaluator + mechanical citation checking). See docs/critique.md for cost and limits. If you
|
|
183
|
+
prefer to build your own grader, the substrate is still here:
|
|
174
184
|
- `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
|
|
175
185
|
- `cowork-harness trace <run-dir> --output-format json` → the tool-call stream. Add `--full-results` so
|
|
176
186
|
a **successful** call's full input + result are captured (the default view slices them to ~100/120
|
|
@@ -181,7 +191,10 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
181
191
|
verdict.
|
|
182
192
|
2. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
|
|
183
193
|
`result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
|
|
184
|
-
content-exact, on every live run
|
|
194
|
+
content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
|
|
195
|
+
no `skillHash` on the `skill` lane at all, so verify the field is present before pairing on it** (a run that mounts nothing
|
|
196
|
+
records none; the `chat` lane records no fingerprint), changes on any tracked edit. **Group/pair on it**
|
|
197
|
+
(`inspect` and the
|
|
185
198
|
run-index row surface a short prefix). Add `--label <tag>` for a human-readable generation name
|
|
186
199
|
(skillHash is the correctness key; the label is ergonomics). `cowork-harness verify-run <run-dir>
|
|
187
200
|
<scenario.yaml>` is the native staleness guard: it **warns** when a kept run predates the current
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,99 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.5.0] — 2026-07-20
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`cowork-harness critique <skill-folder> --prompt "<probe>"` (EXPERIMENTAL).** Runs a skill, asks the
|
|
14
|
+
agent what confused it, then grades that self-report against a frozen record of what actually happened —
|
|
15
|
+
a byte-boundary evidence snapshot taken *before* the reflection turn, a first evaluator pass that is
|
|
16
|
+
structurally blind to the self-report (the text is never put in its prompt, not merely ignored), and
|
|
17
|
+
mechanical citation checking that drops any claim not quoting the evidence verbatim. **A discovery
|
|
18
|
+
instrument, never a gate:** findings of any classification exit 0 (including a task run that itself
|
|
19
|
+
errored — that is a finding about the skill); exit 2 is reserved for a usage error or an instrument
|
|
20
|
+
failure, where no critique was produced. Previously a maintainer-only script that could not run from an installed package at all.
|
|
21
|
+
Costs four model workloads per critique — see [docs/critique.md](./docs/critique.md).
|
|
22
|
+
- **Evidence-package armoring.** The self-report was already fenced, but the evidence package — which
|
|
23
|
+
carries a third-party SKILL.md verbatim into both evaluator prompts — was not, so hostile skill content
|
|
24
|
+
could steer the grader directly. Untrusted content now sits inside per-run nonce markers and only
|
|
25
|
+
nonce-tagged headings outside them count as instructions; a skill cannot pre-author the nonce. Verified
|
|
26
|
+
by a red-team probe across three models: the structural-forgery payload that steered all three now
|
|
27
|
+
matches control. Content that merely *argues* is a documented residual — fencing separates planes, it
|
|
28
|
+
cannot stop persuasion.
|
|
29
|
+
- **`skill --repeat <N>`** (2–100), with `--min-pass-rate` / `--stop-on-diverge` / `--max-budget-usd` /
|
|
30
|
+
`--allow-budget-stop`. The variance rollup already existed but was `run`-only, because the flags were
|
|
31
|
+
parsed inline in the `run` command — the exploratory lane, where an iterate-across-fixes loop actually
|
|
32
|
+
lives, rejected them as unknown. Both lanes now share one parse (`run/repeat-flags.ts`) and one batch
|
|
33
|
+
engine, so rollup shape, JSON envelope, and batch verdict match. `skill --repeat` additionally rejects
|
|
34
|
+
`--session-id`/`--resume` (both pin a single run dir, so iterations would overwrite each other rather
|
|
35
|
+
than produce N independent samples) and `--decider-dir`/`--decider-cmd` (the reproducibility invariant
|
|
36
|
+
`run` already enforced).
|
|
37
|
+
- **`result.json` `outcome`** — a one-field rollup of the `result` × `verdict.pass` × exit-code matrix:
|
|
38
|
+
`errored` / `no_deliverable` / `delivered_with_verdict_fail` / `delivered_clean`. Those three signals
|
|
39
|
+
legitimately disagree (a fail-severity signal flips the verdict while `result` stays `"success"`), and a
|
|
40
|
+
consumer driving a loop had to reconstruct "did this iteration deliver something usable?" from all
|
|
41
|
+
three. A pure function of fields the run already carries, so it cannot disagree with them; the granular
|
|
42
|
+
fields stay authoritative. Absent whenever `verdict` is absent. Note `delivered_*` means "no
|
|
43
|
+
stall/question signal fired", not positive evidence a deliverable exists — check `artifacts`.
|
|
44
|
+
- **Generation-pairing `jq` recipes** in [docs/stats.md](./docs/stats.md) — pass-rate/cost per generation,
|
|
45
|
+
a per-generation verdict-signal histogram, and the before/after of a single fix, grouped on
|
|
46
|
+
`skillHash`/`runLabel` over `index.jsonl`.
|
|
47
|
+
|
|
48
|
+
### Changed
|
|
49
|
+
|
|
50
|
+
- **`diff`'s exit code now honors its own documented contract.** `--help` has always said "transcript is
|
|
51
|
+
advisory … tools/artifacts/meta are the gateable signal", but `identical` conjoined all four views, so
|
|
52
|
+
two live runs of the SAME skill exited 1 and the signal could not separate "behaviour changed" from
|
|
53
|
+
model stochasticity. `identical` now means the gateable views agree; transcript drift is reported
|
|
54
|
+
separately (`transcriptDiffers` in JSON, rendered in text when you ask for that view) so it stays
|
|
55
|
+
visible. `diff` also carries `skillHash` in the meta view now — a diff across a fix could not previously
|
|
56
|
+
name which two generations it compared.
|
|
57
|
+
|
|
58
|
+
### Fixed
|
|
59
|
+
|
|
60
|
+
- **`fingerprint.skillHash` and `skillCommit` are now recorded on the `skill` and `probe-dispatch` lanes.**
|
|
61
|
+
Both resolved their skill dirs by re-reading the session *file*, so the lanes that mount via an in-memory
|
|
62
|
+
session — and pass the `"(inline)"` sentinel as the session path — emitted no `skillHash` and a null
|
|
63
|
+
`skillCommit`, even though the mounts were already in scope at the call site. The resolved session object
|
|
64
|
+
is now threaded through instead. Consequences, all previously broken on those lanes: `result.json` carries
|
|
65
|
+
the content-exact generation key, the run-index `skillHash` column populates (so a harvest step can group
|
|
66
|
+
runs by generation), and the run's own "pair critiques by `fingerprint.skillHash`" tip — which is printed
|
|
67
|
+
*only* on the `skill` lane — is no longer advertising a field that lane could not emit. Scoped to the
|
|
68
|
+
sentinel branch: the file-based path is untouched, so recorded cassettes and the
|
|
69
|
+
staleness / `verify-run` recomputes are unaffected. A session that mounts nothing still yields no hash —
|
|
70
|
+
there is nothing to hash.
|
|
71
|
+
|
|
72
|
+
### Changed
|
|
73
|
+
|
|
74
|
+
- **Reflective skill-critique prompt v2** (`REFLECTION_PROMPT_VERSION` 1 → 2; maintainer instrument, not a
|
|
75
|
+
shipped surface). Adds a sub-agent question (were any dispatched, and was the skill clear about when to
|
|
76
|
+
dispatch, what context to hand them, and what to expect back); replaces the "change ONE thing" cap with
|
|
77
|
+
exhaustive solicitation, since a separate evaluator already triages and drops ungrounded findings, so
|
|
78
|
+
capping at the source loses signal for no quality gain; drops a "fidelity tier" example that is
|
|
79
|
+
cowork-harness vocabulary a third-party skill's agent never encountered; and bounds the pass-2 self-report
|
|
80
|
+
now that the prompt invites longer replies.
|
|
81
|
+
|
|
82
|
+
## [1.4.0] — 2026-07-19
|
|
83
|
+
|
|
84
|
+
### Added
|
|
85
|
+
|
|
86
|
+
- **`coworkWebFetchDedup` enacted (hostloop `web_fetch`).** Real Cowork keeps a per-session negative-work
|
|
87
|
+
cache: a repeat `web_fetch` of the same normalized URL within a TTL (default 15 min; cap 100; FIFO
|
|
88
|
+
eviction; a hit does not refresh recency) makes **no network request** and returns a marker telling the
|
|
89
|
+
model to re-use the earlier result. The harness now reproduces this on the host-API (`coworkWebFetchViaApi`)
|
|
90
|
+
path — **baseline-gated** (only when the resolved baseline's `coworkWebFetchDedup` gate is on, i.e. Desktop
|
|
91
|
+
≥ 1.22209.3), keyed under both the request URL and the terminal `destination_url`, never caching errors /
|
|
92
|
+
empty / non-2xx responses, and emitting **no egress event** on a hit (matching production's zero-network
|
|
93
|
+
dedup). A hit is observable via the marker text (`tool_result_contains: "Already fetched"`).
|
|
94
|
+
|
|
95
|
+
### Changed
|
|
96
|
+
|
|
97
|
+
- **Platform baseline synced to Desktop 1.22209.3** (agent `2.1.215`). No prompt / spawn-env / egress-allowlist
|
|
98
|
+
drift vs 1.21459.0; the sync captured the new `coworkWebFetchDedup` runtime config (enacted above) plus a
|
|
99
|
+
few new (off) GrowthBook gates. The skill/README/reference version floors and agent-binary pins track the
|
|
100
|
+
new baseline.
|
|
101
|
+
|
|
9
102
|
## [1.3.0] — 2026-07-19
|
|
10
103
|
|
|
11
104
|
### Added
|
package/README.md
CHANGED
|
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
91
91
|
|
|
92
92
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
93
93
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
94
|
-
> From a global install (`npm i -g "cowork-harness@>=1.
|
|
94
|
+
> From a global install (`npm i -g "cowork-harness@>=1.5.0"`), point at the package root instead:
|
|
95
95
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
96
96
|
> (or copy the cassette into your own project and pass that path).
|
|
97
97
|
|
|
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
101
101
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
102
102
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
103
103
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
104
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.
|
|
104
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.5.0"`.
|
|
105
105
|
|
|
106
106
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
107
107
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
126
126
|
claude plugin install cowork-harness@cowork-harness
|
|
127
127
|
```
|
|
128
128
|
|
|
129
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.
|
|
129
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.5.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
130
130
|
|
|
131
131
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
132
132
|
|
|
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
|
|
|
147
147
|
To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
|
|
148
148
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
149
149
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
150
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.
|
|
150
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.5.0"` — see
|
|
151
151
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
152
152
|
|
|
153
153
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -392,7 +392,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
392
392
|
|
|
393
393
|
| Command | What it does | Reach for it when… |
|
|
394
394
|
|---|---|---|
|
|
395
|
-
| `skill <folder> "<prompt>"` | Run a local skill/plugin folder once against the staged agent; `--compact`/`--demo` trim output for shareable screenshots/GIFs; `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift); `--label <tag>` stamps a generation name for the iterate-across-fixes loop (surfaced in `result.json`/index/`inspect`); `--allow-missing-capability` stops a capability false-negative on the lean `core` image from failing the verdict (open-ended equivalent of asserting `allow_missing_capability: true`) | ad-hoc "is the skill alive / does it do X?" — the fast inner loop |
|
|
395
|
+
| `skill <folder> "<prompt>"` | Run a local skill/plugin folder once against the staged agent; `--compact`/`--demo` trim output for shareable screenshots/GIFs; `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift); `--label <tag>` stamps a generation name for the iterate-across-fixes loop (surfaced in `result.json`/index/`inspect`); `--repeat N` (2-100) runs the same skill+prompt N times and aggregates a variance rollup — "did this finding reproduce, or did it pass once?"; `--allow-missing-capability` stops a capability false-negative on the lean `core` image from failing the verdict (open-ended equivalent of asserting `allow_missing_capability: true`) | ad-hoc "is the skill alive / does it do X?" — the fast inner loop |
|
|
396
396
|
| `run <scenario.yaml \| dir/>` | Run authored scenarios with `assert:` + a CI-ready exit code; a decider can answer unscripted gates; `--repeat`/`--matrix` add variance runs / a compatibility matrix (detail below); `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift) | you want a repeatable, **asserted regression test** — to **measure flakiness** instead of trusting one green — or to **test a compatibility matrix** (multiple baselines/models/skill variants) in one run |
|
|
397
397
|
| `chat <folder> [prompt]` | Interactive multi-turn REPL against a skill (TTY); optional seed prompt is sent as the first turn. `--upload <file>` / `--folder <dir>` (repeatable) attach files/project folders; `--verbose` shows thinking blocks + tool inputs; `--fidelity protocol\|container\|hostloop` (no `microvm`/`cowork` in the REPL); `--allow-host-writes` consents to a writable `hostloop`-fidelity connected folder (same consent as clicking "connect folder" in Desktop); `--raw` skips the control protocol for native `docker run -it` (rejects `--upload`/`--folder`/`--plugin`/`--fidelity`) | debugging a multi-turn flow by hand |
|
|
398
398
|
| `record` / `replay` | **Record a live run once → replay it token-free, Docker-free thereafter** (key flags below; `replay --explain` prints the evidence behind every passing assert) | **token-free, Docker-free CI** from a once-recorded run |
|
|
@@ -415,6 +415,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
415
415
|
| `boundary-check [baseline] [--session <file>]` | Prove the **L1 Docker** sandbox enforces Cowork's limitations (sealed FS + default-deny egress; `container`/`hostloop` share this sandbox — `microvm`'s guest firewall is not probed here); `--session` folds a session's `egress.extra_allow` into the probe allowlist | verifying the harness's own fidelity |
|
|
416
416
|
| `sync` / `list` | Derive/refresh (`sync [--diff] [--allow-empty\|--force]`) & list platform baselines from the Desktop install | after Claude Desktop updates (baselines ship, so it's optional otherwise) |
|
|
417
417
|
| `diff <a> <b>` | Compare two baselines, two runs, two cassettes, or a run+cassette — kind auto-detected by content (view/normalization flags below). Token-free, no live Desktop/Docker needed | "what changed between two runs/cassettes/baselines?" |
|
|
418
|
+
| `critique <skill-folder>` | **EXPERIMENTAL.** Run a skill, ask the agent what confused it, then grade that self-report against a frozen record of the run — blinded evaluator, mechanical citation checking. Discovery instrument, never a gate (findings always exit 0). Costs four model workloads; see [docs/critique.md](./docs/critique.md) | "what confused the agent about my skill?" |
|
|
418
419
|
| `doctor [--tier <t>]` | Read-only prerequisite check, per tier (Docker + agent image for `container`/`hostloop`/`cowork`; **Lima** for `microvm`; plus staged agent, token, baseline); prints the exact `docker build` line if the agent image is missing | "can I run the live tiers — what's missing?" before a first live run |
|
|
419
420
|
| `prune [<runs-dir>] [--keep-last <n>] [--pinned-older-than <N>d\|h\|m]` | Prune accumulated run dirs (keeps the N most recent per scenario; pinned `--session-id` runs are never pruned unless `--pinned-older-than` opts in to reclaiming stale ones by last-activity age); the optional positional overrides the runs root; `--dry-run` | the machine-global runs root has grown and you want space back |
|
|
420
421
|
| `rehash <dir/>` | Migrate cassette fingerprints to the current format version when the content is provably unchanged (`--dry-run`); no re-record needed | a cassette-format bump flagged committed fixtures as stale |
|
|
@@ -659,7 +660,7 @@ jobs:
|
|
|
659
660
|
- uses: actions/checkout@v4
|
|
660
661
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
661
662
|
run: |
|
|
662
|
-
V=2.1.
|
|
663
|
+
V=2.1.215 # match your scenario's pinned baseline's agentVersion
|
|
663
664
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
664
665
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
665
666
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
@@ -670,7 +671,7 @@ jobs:
|
|
|
670
671
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
671
672
|
```
|
|
672
673
|
|
|
673
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.
|
|
674
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.5.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
|
|
674
675
|
|
|
675
676
|
The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
676
677
|
|
|
@@ -831,6 +832,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
831
832
|
## Status
|
|
832
833
|
|
|
833
834
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
834
|
-
**`desktop-1.
|
|
835
|
+
**`desktop-1.22209.3`**. Release-by-release verification notes (what was re-verified against
|
|
835
836
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
836
837
|
this section used to duplicate lives in the sections above.
|
package/SPEC.md
CHANGED
|
@@ -309,7 +309,7 @@ Payload sits under an **inner** `response`. Missing the nesting ⇒ `ZodError: e
|
|
|
309
309
|
> Channels 2-3 = CLI `--mcp-config` / `.mcp.json` servers (honored in plain cowork mode; dropped only in
|
|
310
310
|
> hermetic mode). Details per channel below.
|
|
311
311
|
|
|
312
|
-
1. **SDK servers over the control protocol** — declared via `sdkMcpServers` in `initialize`; tool calls tunnel as `mcp_message`. This is how the **desktop host** bridges its own servers (incl. `claude_desktop_config.json` `mcpServers`, spawned host-side with full host env) and how the harness delivers the workspace shell. The workspace handler (`src/hostloop/workspace-handler.ts`) implements `initialize`/`tools/list`/`tools/call`; `bash`→`docker exec -w <mntRoot> <container> sh -c <cmd>` (container-egress-gated). **CB-8:** `makeWorkspaceHandler` accepts an `onInfraError?: (message: string) => void` callback at parameter position 6 (after `onEgress`); on infrastructure errors (ETIMEDOUT / killed / no code+stdout+stderr) the handler calls `onInfraError?.(e.message)`, returns a textResult with `"[infrastructure error: …]"`, and `spawnHostLoop` wires `onInfraError` to append `{type:"infra_error", ts, message}` to `events.jsonl`. **`web_fetch` is NOT container-egress-gated:** real Cowork routes it through the host API (gate `1978029737` `coworkWebFetchViaApi:true` → `POST /api/organizations/<org>/cowork/web_fetch`), gated by a **separate web-fetch hostname allowlist** (`getWebFetchAllowedUrls`, `*`=unrestricted) + a **URL-provenance** rule (URL must have appeared in a prior message/result). The harness mirrors this with the **two-path model** (`src/hostloop/workspace-handler.ts`, binary-verified `G1t`/`U1t`): **Path A** (provenance engaged — `coworkWebFetchViaApi` on) gates on the **exact-URL provenance set** ONLY (seeded from user-turn + tool-result URLs; `src/hostloop/provenance.ts`), with **no** hostname allowlist — but still an `http(s)`-scheme + private-address **SSRF backstop re-checked on every redirect hop** (a manual redirect loop, not `curl -L`) — a miss raises a per-domain approval (`webfetch:<domain>` permission with options `Allow once | Allow all for website | Deny`) routed through the Decider; "Allow all for website" approves the host for the rest of the run (`Run.approvedDomains`, per-run/ephemeral). **Path B** (gate off) is a direct host fetch gated by the egress domain list via the same `wen()`/`compile()` matcher container egress uses, with `redirect:"manual"` re-checking `U1t` (scheme + private-address SSRF + allowlist) on **every** redirect hop. web_fetch is thus **decoupled from `plan.egressAllow`** on Path A (egress applies to `bash`/Path B only). An unanswered cold miss is fail-closed; scenarios answer via `--answer "webfetch:<domain>=allow"` (with `grant`), `web_fetch.approved_domains`, or an LLM/external terminal. `bash` stays container-egress-sandboxed.
|
|
312
|
+
1. **SDK servers over the control protocol** — declared via `sdkMcpServers` in `initialize`; tool calls tunnel as `mcp_message`. This is how the **desktop host** bridges its own servers (incl. `claude_desktop_config.json` `mcpServers`, spawned host-side with full host env) and how the harness delivers the workspace shell. The workspace handler (`src/hostloop/workspace-handler.ts`) implements `initialize`/`tools/list`/`tools/call`; `bash`→`docker exec -w <mntRoot> <container> sh -c <cmd>` (container-egress-gated). **CB-8:** `makeWorkspaceHandler` accepts an `onInfraError?: (message: string) => void` callback at parameter position 6 (after `onEgress`); on infrastructure errors (ETIMEDOUT / killed / no code+stdout+stderr) the handler calls `onInfraError?.(e.message)`, returns a textResult with `"[infrastructure error: …]"`, and `spawnHostLoop` wires `onInfraError` to append `{type:"infra_error", ts, message}` to `events.jsonl`. **`web_fetch` is NOT container-egress-gated:** real Cowork routes it through the host API (gate `1978029737` `coworkWebFetchViaApi:true` → `POST /api/organizations/<org>/cowork/web_fetch`), gated by a **separate web-fetch hostname allowlist** (`getWebFetchAllowedUrls`, `*`=unrestricted) + a **URL-provenance** rule (URL must have appeared in a prior message/result). The harness mirrors this with the **two-path model** (`src/hostloop/workspace-handler.ts`, binary-verified `G1t`/`U1t`): **Path A** (provenance engaged — `coworkWebFetchViaApi` on) gates on the **exact-URL provenance set** ONLY (seeded from user-turn + tool-result URLs; `src/hostloop/provenance.ts`), with **no** hostname allowlist — but still an `http(s)`-scheme + private-address **SSRF backstop re-checked on every redirect hop** (a manual redirect loop, not `curl -L`) — a miss raises a per-domain approval (`webfetch:<domain>` permission with options `Allow once | Allow all for website | Deny`) routed through the Decider; "Allow all for website" approves the host for the rest of the run (`Run.approvedDomains`, per-run/ephemeral). On Path A, **AFTER** the provenance/approval gate, **`coworkWebFetchDedup`** (gate `1978029737`, on for baselines ≥ `1.22209.3`) applies a per-session **negative-work cache** (`src/hostloop/webfetch-dedup.ts`, binary-verified): a repeat `web_fetch` of the same **normalized** URL (`normalizeUrl`) within the TTL (baseline-sourced, default `900000` ms; cap `100`, FIFO eviction; a hit does NOT refresh recency) returns a marker telling the model to re-use the earlier result, with **no network request and no egress event** — matching production's zero-network dedup. Only successful (HTTP-2xx, trimmed-nonempty) fetches are cached, keyed under both the request URL and the terminal `destination_url`; errors/empty/non-2xx are never cached. Dedup is baseline-gated (an older baseline lacking the flag never dedups) and Path-A-only. **Path B** (gate off) is a direct host fetch gated by the egress domain list via the same `wen()`/`compile()` matcher container egress uses, with `redirect:"manual"` re-checking `U1t` (scheme + private-address SSRF + allowlist) on **every** redirect hop. web_fetch is thus **decoupled from `plan.egressAllow`** on Path A (egress applies to `bash`/Path B only). An unanswered cold miss is fail-closed; scenarios answer via `--answer "webfetch:<domain>=allow"` (with `grant`), `web_fetch.approved_domains`, or an LLM/external terminal. `bash` stays container-egress-sandboxed.
|
|
313
313
|
2. **CLI-spawned `--mcp-config` / `.mcp.json` servers — HONORED in plain cowork mode** (NOT ignored). **Verified:** a valid `--mcp-config` populates `mcp_servers` (`[{name,status:"pending"|"connected"}]`); these run in-sandbox with the env-allowlist `CLAUDE_CODE_MCP_ALLOWLIST_ENV` (`RW8`/`oG8`/`LU5` = {HOME, LOGNAME, PATH, SHELL, TERM, USER}). The harness MAY use this as a convenience injection path.
|
|
314
314
|
3. **The drop is SAFE/HERMETIC-mode-gated, not cowork-gated.** `--mcp-config` is filtered to SDK-only (`ap5()`) **only when** safe mode (`I5()`) or `xB8()` is true, and `xB8()` requires **both** `CLAUDE_CODE_REMOTE` **and** `CLAUDE_CODE_REMOTE_HERMETIC_MODE`. **Verified:** with both set, `mcp_servers:[]`; without them (plain `SESSION_KIND=bg`), the config is honored. The earlier "cowork ignores `--mcp-config`" was a hermetic-session observation over-generalized.
|
|
315
315
|
|
|
@@ -603,7 +603,7 @@ unused — reserving it now keeps a later addition additive rather than a renumb
|
|
|
603
603
|
`0`/`1`/`2`/`3` space. Exit-code space is **per-command**, not global (`status` uses `0`/`1`/`2`/`3`
|
|
604
604
|
with its own meanings); this reservation applies only to the `run`/`skill` family.
|
|
605
605
|
|
|
606
|
-
**Per-command exceptions:** `lint` exits `127` when `python3` is missing (spawn error); `replay` exits
|
|
606
|
+
**Per-command exceptions:** `critique` **never gates on findings** — it exits `0` for any finding of any classification, and even when the task run it graded ERRORED (that is a finding about the skill, not a broken instrument). It exits `2` only for a usage error or an **instrument failure**: the turn was killed, the reflection protocol broke, or the evaluator was never invoked *or threw* — i.e. no critique was produced. Do not gate CI on `critique`; that inverts its design. `lint` exits `127` when `python3` is missing (spawn error); `replay` exits
|
|
607
607
|
`2` on a **whole-cassette operational failure** — anything `readCassette` rejects (unreadable, invalid
|
|
608
608
|
shape, unsupported version, unrecognized assertion key) or any per-file throw, plus the batch loop's
|
|
609
609
|
own source-resolution failures (`--assert-from`/`--reassert` drift, scenario-parse errors, `--write`
|