cowork-harness 2.5.0 → 3.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +7 -7
- package/.claude/skills/cowork-harness/references/ci-recipe.md +17 -17
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +61 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +17 -5
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -2
- package/.claude/skills/cowork-harness/scripts/scenario.py +3 -2
- package/AGENTS.md +1 -1
- package/CHANGELOG.md +218 -0
- package/CONTRIBUTING.md +35 -0
- package/DESIGN.md +2 -2
- package/README.md +87 -704
- package/SPEC.md +2 -2
- package/baselines/desktop-1.40609.0.json +878 -0
- package/baselines/provisioning/rootfs-provisioning.json +32 -40
- package/dist/cli.js +8 -1
- package/dist/run/cassette.js +4 -3
- package/dist/run/chat-result.js +1 -1
- package/dist/run/chat.js +15 -1
- package/dist/run/execute.js +15 -6
- package/dist/run/hook-events.js +84 -0
- package/dist/run/skill-flag-surface.js +11 -0
- package/dist/run/verdict.js +17 -8
- package/dist/runtime/argv.js +16 -1
- package/dist/runtime/lima.js +37 -3
- package/dist/runtime/protocol.js +85 -13
- package/dist/scan.js +1 -0
- package/dist/sync/cowork-sync.js +39 -2
- package/dist/types.js +11 -3
- package/docs/README.md +9 -6
- package/docs/boundary.md +7 -0
- package/docs/cassette.md +3 -3
- package/docs/ci.md +114 -0
- package/docs/cli.md +532 -0
- package/docs/companion-skill.md +58 -0
- package/docs/debugging.md +3 -3
- package/docs/discovery.md +1 -1
- package/docs/fidelity-gaps.md +88 -7
- package/docs/gotchas.md +27 -2
- package/docs/maintenance.md +29 -1
- package/docs/scenario.md +6 -3
- package/docs/session.md +1 -1
- package/docs/stats.md +1 -1
- package/examples/README.md +3 -3
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/examples/sessions/l0-plugin-delivery.yaml +6 -0
- package/llms.txt +3 -0
- package/package.json +1 -1
- package/python/test_scenario_lint.py +1 -1
- package/schema/run-result.json +2 -2
- package/schema/scenario.schema.json +6 -2
- package/scripts/bump-version.ts +12 -1
- package/scripts/check-versions.ts +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version:
|
|
7
|
-
tracks-harness: cowork-harness
|
|
6
|
+
version: 3.0.1
|
|
7
|
+
tracks-harness: cowork-harness 3.0.1 (baseline desktop-1.40609.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness
|
|
29
|
-
> `desktop-1.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.1` (baseline
|
|
29
|
+
> `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,13 +42,13 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.1"`. **Pin `@^3.0.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
49
49
|
— upgrade rather than work around it, since this skill's `file:line` pointers and flag names track the floor.
|
|
50
50
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
51
|
-
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
51
|
+
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
|
|
52
52
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
53
53
|
- **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
|
|
54
54
|
|
|
@@ -679,7 +679,7 @@ than a stuck `"running"`.)
|
|
|
679
679
|
### Place assertions in the right CI lane
|
|
680
680
|
|
|
681
681
|
CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
|
|
682
|
-
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@
|
|
682
|
+
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
|
|
683
683
|
PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
|
|
684
684
|
the four-stage pipeline.
|
|
685
685
|
|
|
@@ -1,23 +1,23 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
7
7
|
|
|
8
8
|
```yaml
|
|
9
|
-
- uses: yaniv-golan/cowork-harness@
|
|
9
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
10
10
|
with:
|
|
11
11
|
command: replay
|
|
12
12
|
path: cassettes/
|
|
13
|
-
version: "^
|
|
13
|
+
version: "^3" # hold the major; see below
|
|
14
14
|
```
|
|
15
15
|
|
|
16
|
-
**These recipes pin `version: "^
|
|
16
|
+
**These recipes pin `version: "^3"`.** The Action's `version` input *defaults* to `latest`, which means a
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "
|
|
20
|
+
(e.g. `version: "3.0.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.247 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
41
41
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
42
42
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -47,11 +47,11 @@ jobs:
|
|
|
47
47
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
48
48
|
# Background on the provenance chain: the "Agent-binary provenance" section of
|
|
49
49
|
# https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
|
|
50
|
-
- uses: yaniv-golan/cowork-harness@
|
|
50
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
51
51
|
with:
|
|
52
52
|
command: run
|
|
53
53
|
path: scenarios/
|
|
54
|
-
version: "^
|
|
54
|
+
version: "^3"
|
|
55
55
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
56
56
|
```
|
|
57
57
|
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.0.1"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -119,18 +119,18 @@ Action has no input for, and it creates a coupling nothing checks:
|
|
|
119
119
|
you have a reason:
|
|
120
120
|
|
|
121
121
|
```yaml
|
|
122
|
-
- uses: yaniv-golan/cowork-harness@
|
|
122
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
123
123
|
with:
|
|
124
124
|
command: lint
|
|
125
125
|
path: scenarios/
|
|
126
|
-
version: "^
|
|
127
|
-
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any
|
|
126
|
+
version: "^3" # holds the major
|
|
127
|
+
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 3.x satisfies that
|
|
128
128
|
```
|
|
129
129
|
|
|
130
130
|
**If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
|
|
131
131
|
floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
|
|
132
132
|
recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
|
|
133
|
-
major instead — `version: "^
|
|
133
|
+
major instead — `version: "^3"`, which is what the steps above use — keeping the floor's intent while
|
|
134
134
|
stopping at the major boundary. An exact
|
|
135
135
|
pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
|
|
136
136
|
moment a recipe adopts a newer flag.
|
|
@@ -147,7 +147,7 @@ The split is not just about tokens — it decides **where each lane can run**:
|
|
|
147
147
|
Docker, no agent binary** — runs on a stock GitHub Actions runner. Evaluates **content** assertions —
|
|
148
148
|
`transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
|
|
149
149
|
`max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
|
|
150
|
-
`allow_permissive_auto_allow` / `allow_missing_capability` / `
|
|
150
|
+
`allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
|
|
151
151
|
`allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
|
|
152
152
|
`question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
153
153
|
(`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
|
|
@@ -322,7 +322,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
322
322
|
|
|
323
323
|
## GitHub Actions sketch
|
|
324
324
|
|
|
325
|
-
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@
|
|
325
|
+
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v3` does
|
|
326
326
|
in one step (see the top of this doc) — reach for this form when you need independent per-command
|
|
327
327
|
gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
|
|
328
328
|
equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
|
|
@@ -342,7 +342,7 @@ jobs:
|
|
|
342
342
|
with: { node-version: '24' }
|
|
343
343
|
- uses: actions/setup-python@v5
|
|
344
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^
|
|
345
|
+
- run: npm i -g "cowork-harness@^3.0.1"
|
|
346
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +371,7 @@ jobs:
|
|
|
371
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
372
|
fi
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^
|
|
374
|
+
run: npm i -g "cowork-harness@^3.0.1"
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness
|
|
3
|
+
Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -20,6 +20,12 @@ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.379
|
|
|
20
20
|
- A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
|
|
21
21
|
true` — with no container around the native file tools, that combination gives the agent genuine,
|
|
22
22
|
software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
|
|
23
|
+
- A `protocol` scenario staging a plugin that declares runnable hooks — in `<plugin>/hooks/hooks.json`
|
|
24
|
+
or the plugin manifest's `hooks` key, both live — needs `allow_host_hooks: true`
|
|
25
|
+
(`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
|
|
26
|
+
hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
|
|
27
|
+
Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
|
|
28
|
+
does not fall back to the default.
|
|
23
29
|
- **Set the tier in the scenario's `fidelity:` field — not a flag.** `--fidelity` is accepted only by
|
|
24
30
|
`skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
|
|
25
31
|
positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
|
|
@@ -150,6 +156,60 @@ hands `fn` exactly this dict.
|
|
|
150
156
|
- `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
|
|
151
157
|
`["fail", "prompt", "llm", "first"]`.
|
|
152
158
|
|
|
159
|
+
### A script-running skill trips `permissive_auto_allow` — at `protocol` specifically
|
|
160
|
+
|
|
161
|
+
The default `permission_parity: cowork` auto-allows an unscripted, off-registry tool ask — and records
|
|
162
|
+
it, because real Cowork would have BLOCKED for the user. `computeVerdict` then FAILS the run: a green
|
|
163
|
+
carrying one would not be a faithful pass. Only `Read`, `Glob` and `Grep` are default-allow; **`Bash` is
|
|
164
|
+
off-registry**.
|
|
165
|
+
|
|
166
|
+
Two things decide whether you ever see it, and neither is obvious. The harness only decides how to
|
|
167
|
+
ANSWER a permission ask; whether the agent ASKS is the agent binary's own logic, and that varies by
|
|
168
|
+
both the command and the tier. All rows below are measured, same scenario shape, same session:
|
|
169
|
+
|
|
170
|
+
| Tier | Bash command | tool used | ask | verdict |
|
|
171
|
+
|---|---|---|---|---|
|
|
172
|
+
| `protocol` | `echo hello` | `Bash` | none | ✓ green |
|
|
173
|
+
| `protocol` | `python3 -c "print(42)"` | `Bash` | **one** | ✗ `permissive_auto_allow` |
|
|
174
|
+
| `container` | `python3 -c "print(42)"` | `Bash` | none | ✓ green |
|
|
175
|
+
| `hostloop` | `python3 -c "print(42)"` | `mcp__workspace__bash` | none | ✓ green |
|
|
176
|
+
|
|
177
|
+
So the guard is **not** a general hazard for script-running skills — it is a `protocol` one. The host
|
|
178
|
+
CLI that L0 spawns asks for a command like `python3 -c …` and not for `echo`; the staged agent at
|
|
179
|
+
`container` asks for neither, and `hostloop` routes shell through `mcp__workspace__bash` and so is not
|
|
180
|
+
even the same tool. A skill that runs `python3 ${CLAUDE_SKILL_DIR}/scripts/…` therefore goes red at
|
|
181
|
+
`protocol` on an otherwise-correct run, while the same skill is green on the sandboxed tiers — and a
|
|
182
|
+
hello-world probe is green everywhere and teaches the wrong expectation.
|
|
183
|
+
|
|
184
|
+
Three ways through, in preference order: script the gate (`--answer` / `answers:`), which is what the
|
|
185
|
+
warning tells you and keeps the run deterministic; set `permission_parity: strict` to deny instead of
|
|
186
|
+
allow, if refusal is what you want to test; or assert `allow_permissive_auto_allow: true` when the
|
|
187
|
+
permissive behaviour is deliberately what the scenario is about.
|
|
188
|
+
|
|
189
|
+
> **What is NOT established.** Why the host CLI asks for one command and not another was observed, not
|
|
190
|
+
> traced — do not infer a rule for which commands ask from these rows. `microvm` was not measured.
|
|
191
|
+
|
|
192
|
+
### `baseline agent binary not found` — a Desktop update pruned the pinned ELF
|
|
193
|
+
|
|
194
|
+
A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
|
|
195
|
+
a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
|
|
196
|
+
current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
|
|
197
|
+
|
|
198
|
+
Prefer **repinning `baseline:` to an installed version** for anything
|
|
199
|
+
reproducibility-bound — you keep an exact pin and move it deliberately. `baseline: latest` never rots
|
|
200
|
+
but silently drifts, so two runs weeks apart are not comparable; a pin and `latest` have opposite
|
|
201
|
+
failure modes.
|
|
202
|
+
|
|
203
|
+
To find which versions you actually have, do NOT use `cowork-harness list` — it enumerates the baseline
|
|
204
|
+
definitions shipped with the harness, which are present regardless of what Desktop pruned from this
|
|
205
|
+
machine, so a pruned pin lists as healthy. Test the staged binary: read `agentBinary.stagedPath` from
|
|
206
|
+
the baseline JSON, expand the leading `~`, and check it is a FILE — the pruned case leaves the empty
|
|
207
|
+
directory behind, so a directory test passes on exactly the case that fails.
|
|
208
|
+
|
|
209
|
+
`COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` runs the newest sibling instead of the pinned
|
|
210
|
+
binary and downgrades the sha check to advisory — that is the substitution the hard failure exists to
|
|
211
|
+
prevent, so use it to unblock once, never in CI.
|
|
212
|
+
|
|
153
213
|
### Determinism contract
|
|
154
214
|
|
|
155
215
|
- `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.1`
|
|
4
|
+
(baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -98,6 +98,18 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
|
|
|
98
98
|
# around hostloop's native file tools, that combination gives the
|
|
99
99
|
# agent genuine, software-checked-only host filesystem access.
|
|
100
100
|
# Read-only folders and folder-less runs need no opt-in.
|
|
101
|
+
|
|
102
|
+
allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
|
|
103
|
+
# declares runnable hooks — either `<plugin>/hooks/hooks.json` OR
|
|
104
|
+
# the manifest's `hooks` key: L0 passes
|
|
105
|
+
# --plugin-dir, so the CLI executes those hooks as NATIVE HOST
|
|
106
|
+
# processes under your account, with no container sandbox. A plugin
|
|
107
|
+
# that declares no hooks needs no opt-in, and a misplaced root-level
|
|
108
|
+
# `hooks.json` cannot execute so it does not trigger the gate.
|
|
109
|
+
# Use `--fidelity container` to run them sandboxed instead.
|
|
110
|
+
# NEEDS cowork-harness >= 3.0.0. The loader is a strict object, so an
|
|
111
|
+
# OLDER CLI does not default it — it hard-errors
|
|
112
|
+
# `Unrecognized key: "allow_host_hooks"` and exits 2.
|
|
101
113
|
```
|
|
102
114
|
|
|
103
115
|
Relative paths resolve from the file's own directory, so a scenario + session + referenced files
|
|
@@ -362,7 +374,7 @@ same set live from the schema.
|
|
|
362
374
|
| `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
|
|
363
375
|
| `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
|
|
364
376
|
| `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
|
|
365
|
-
| `
|
|
377
|
+
| `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
|
|
366
378
|
| `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
|
|
367
379
|
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
368
380
|
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
@@ -404,7 +416,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
404
416
|
| `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
|
|
405
417
|
| `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
|
|
406
418
|
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
407
|
-
| `
|
|
419
|
+
| `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
|
|
408
420
|
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
409
421
|
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
410
422
|
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
@@ -443,7 +455,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
|
|
|
443
455
|
`skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
|
|
444
456
|
`max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
|
|
445
457
|
(`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
|
|
446
|
-
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `
|
|
458
|
+
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
|
|
447
459
|
`allow_stall` are also kept on replay, evaluated as no-op passes.
|
|
448
460
|
|
|
449
461
|
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness
|
|
5
|
+
Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"keys": [
|
|
4
4
|
"all_tasks_completed",
|
|
5
5
|
"allow_delete_in",
|
|
6
|
-
"
|
|
6
|
+
"allow_l0_host_config_contamination",
|
|
7
7
|
"allow_missing_capability",
|
|
8
8
|
"allow_outputs_delete",
|
|
9
9
|
"allow_permissive_auto_allow",
|
|
@@ -83,6 +83,7 @@
|
|
|
83
83
|
"vm_path_denied"
|
|
84
84
|
],
|
|
85
85
|
"topLevelKeys": [
|
|
86
|
+
"allow_host_hooks",
|
|
86
87
|
"allow_host_writes",
|
|
87
88
|
"answers",
|
|
88
89
|
"assert",
|
|
@@ -101,7 +102,7 @@
|
|
|
101
102
|
],
|
|
102
103
|
"verdictModifierKeys": [
|
|
103
104
|
"allow_delete_in",
|
|
104
|
-
"
|
|
105
|
+
"allow_l0_host_config_contamination",
|
|
105
106
|
"allow_missing_capability",
|
|
106
107
|
"allow_outputs_delete",
|
|
107
108
|
"allow_permissive_auto_allow",
|
|
@@ -173,7 +173,7 @@ LANE_REMOTE_INCOMPATIBLE_KEYS = {"present_files_called", "no_scratchpad_leak", "
|
|
|
173
173
|
# verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
|
|
174
174
|
VERDICT_MODIFIER_KEYS = {
|
|
175
175
|
"allow_permissive_auto_allow",
|
|
176
|
-
"
|
|
176
|
+
"allow_l0_host_config_contamination",
|
|
177
177
|
"allow_missing_capability",
|
|
178
178
|
"allow_stall",
|
|
179
179
|
"allow_undelivered_deliverables",
|
|
@@ -262,7 +262,8 @@ _EMBEDDED_TOP_LEVEL_KEYS = {
|
|
|
262
262
|
"assert",
|
|
263
263
|
"skills", # opt-in skill-staleness hash scope
|
|
264
264
|
"requires_capabilities", # Fix 4b: scenario-level required-capability declaration (pre-flight gate)
|
|
265
|
-
"allow_host_writes",
|
|
265
|
+
"allow_host_writes",
|
|
266
|
+
"allow_host_hooks", # protocol consent: a staged plugin's hooks run as NATIVE HOST processes # hostloop native-split: consent for a writable connected folder (pre-run gate)
|
|
266
267
|
}
|
|
267
268
|
|
|
268
269
|
|
package/AGENTS.md
CHANGED
|
@@ -22,7 +22,7 @@ protocol layer or run-loop bookkeeping in the CLI.
|
|
|
22
22
|
Don't add a test that needs a live model or Docker to the default suite; that's the `pytest -m cowork` /
|
|
23
23
|
`npm run test:live` lane. Python fast lane (from `python/`): `pytest -m 'not cowork'`.
|
|
24
24
|
- CLI binary `cowork-harness`; env vars `COWORK_HARNESS_*` (+ `COWORK_AGENT_BINARY` / `COWORK_AGENT_IMAGE`) —
|
|
25
|
-
see README's [Reproducibility knobs](./
|
|
25
|
+
see README's [Reproducibility knobs](./docs/cli.md#reproducibility-knobs) for the full env-var list.
|
|
26
26
|
Node ≥ 22.
|
|
27
27
|
- `cowork-harness sync` is **local-only** (needs Desktop + `app.asar`; not on CI). The committed
|
|
28
28
|
`baselines/*.json` are CI's source of truth — never hand-edit release facts into source; they come from
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,224 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.0.1] — 2026-08-30
|
|
10
|
+
|
|
11
|
+
### Changed
|
|
12
|
+
|
|
13
|
+
- **The README now shows the product, not just the argument for it.** Three independent reviews found
|
|
14
|
+
the same gap: for a test harness, the file a user authors is the most persuasive thing the project
|
|
15
|
+
owns, and the router split had left it with zero examples. Adds a worked scenario (prompt, scripted
|
|
16
|
+
answers, assertions) with its verdict output — the scenario is extracted and linted in place, so it is
|
|
17
|
+
a real file rather than plausible-looking YAML — plus a "What that catches" section built from an
|
|
18
|
+
actual run where a named skill was offered, declined, and the green answer looked fine anyway
|
|
19
|
+
(`skillsInvoked: []`, `toolCounts: {}`) — framed as a closed loop, since the author fixed and re-verified
|
|
20
|
+
it with the same instrument roughly ninety minutes later. An outcome example is a claim about a MOMENT;
|
|
21
|
+
the better the tool works the faster its own examples get fixed out from under it, so it needs a tense. Also names two things the docs asserted but never taught:
|
|
22
|
+
diffing the same scenario across two fidelity tiers as a discovery technique, and proving an assertion
|
|
23
|
+
can fail before trusting it. The requirements block moves below the `claude -p` argument — three
|
|
24
|
+
reviewers independently reported reading install prerequisites before any reason to want them.
|
|
25
|
+
|
|
26
|
+
- **The README is now a router, not the whole manual.** It was 975 lines carrying three unrelated
|
|
27
|
+
audiences at once; it is now **282** and branches to per-audience pages: **`docs/cli.md`** (install,
|
|
28
|
+
prerequisites, commands, the two files, run output, env knobs), **`docs/companion-skill.md`** (install
|
|
29
|
+
+ orientation; usage stays in `SKILL.md`), and **`docs/ci.md`** (the token-free gate, the packaged
|
|
30
|
+
Action, the live lane — written to stand alone). The harness's own pipeline and contributor suite moved
|
|
31
|
+
to `CONTRIBUTING.md`, which keeps `docs/ci.md` about *consuming* the Action rather than about this repo.
|
|
32
|
+
README keeps what all three audiences share: the fidelity-tier vocabulary, architecture, limitations,
|
|
33
|
+
the docs index, and versioning. Content moved verbatim — this is a relocation, not a rewrite.
|
|
34
|
+
|
|
35
|
+
Nothing is dropped: `Sandboxing` folded into `docs/boundary.md` and `Maintenance` into
|
|
36
|
+
`docs/maintenance.md`; `Discovery` was already covered in `docs/discovery.md`, so it was deleted rather
|
|
37
|
+
than duplicated. Every moved relative link was rewritten for its new depth, and every deep link into a
|
|
38
|
+
moved section was repointed across `docs/`, `examples/README.md`, `llms.txt` and the docs index.
|
|
39
|
+
|
|
40
|
+
### Fixed
|
|
41
|
+
|
|
42
|
+
- **`bump-version` did not know about the router-split pages.** `docs/cli.md`, `docs/ci.md` and
|
|
43
|
+
`docs/companion-skill.md` carry `cowork-harness@^X.Y.Z` install floors that `check:versions` enforces,
|
|
44
|
+
but the bump tool's target list predated them — so `npm run bump` rewrote every other floor and then
|
|
45
|
+
failed its own post-write lockstep check. Caught at release time by that self-check rather than
|
|
46
|
+
shipping a repo with mismatched floors. All three are now registered, and the test pinning that list
|
|
47
|
+
carries the reason.
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
- **22 dead anchor links introduced by the README router split, and the guard gap that let them ship.**
|
|
51
|
+
Moving 11 `##` sections out of README left every `](#slug)` link pointing at them dangling — 19 in
|
|
52
|
+
README (including two badges and the two nav lines the page opened with), plus one each in `AGENTS.md`,
|
|
53
|
+
`docs/ci.md` and `docs/companion-skill.md`. GitHub fails these silently: the page simply does not
|
|
54
|
+
scroll. Every file-path link still resolved and every existing guard stayed green, because the anchor
|
|
55
|
+
suite validated links *into* README and *into* sibling docs but never a page's links to its **own**
|
|
56
|
+
headings. `test/repo-docs-anchors.test.ts` now covers that third case across root pages and `docs/`,
|
|
57
|
+
with a canary against a checker that extracts nothing, and was confirmed to fail against the exact
|
|
58
|
+
regression that shipped. The two nav lines are deleted rather than repointed — they were a
|
|
59
|
+
table-of-contents for a 975-line page, and "Pick your path" plus the Documentation table now do that job.
|
|
60
|
+
|
|
61
|
+
Found by two independent reviewers; the second caught the three non-README instances a README-scoped
|
|
62
|
+
fix would have missed.
|
|
63
|
+
|
|
64
|
+
Follow-up from the same review: three cross-page links still read "…\[Prerequisites](./docs/cli.md#…)
|
|
65
|
+
**below**" — resolving correctly while the sentence lied, because Prerequisites had moved to another
|
|
66
|
+
page. Reworded, and guarded: a cross-page link followed closely by "below"/"above" now fails the suite.
|
|
67
|
+
This half is the nastier one — the link works, so every link checker stays green forever.
|
|
68
|
+
|
|
69
|
+
- **The `protocol` host-hook consent gate was defeatable by a spelling choice.** `allow_host_hooks`
|
|
70
|
+
(3.0.0) refuses a spawn until the operator consents to a staged plugin's hooks running as native host
|
|
71
|
+
processes — but detection keyed only on a file named `hooks.json`, and a plugin may equally declare its
|
|
72
|
+
hooks in the manifest's `hooks` key (`.claude-plugin/plugin.json`, or a bare `plugin.json`), which is
|
|
73
|
+
the more common spelling. Such a plugin sailed past the gate: reproduced end-to-end, the hook executed
|
|
74
|
+
under the operator's own account while the run went green, no consent was asked and no disclosure
|
|
75
|
+
printed. The same blind predicate feeds the tier-independent disclosure, so those hooks also ran
|
|
76
|
+
unannounced at `hostloop`. Detection now returns the UNION of both channels, which fixes the gate and
|
|
77
|
+
the disclosure together since both call through it. The `hooks/hooks.json` placement carve-out is
|
|
78
|
+
deliberately NOT extended to the manifest — that carve-out exists because a misplaced `hooks.json` is
|
|
79
|
+
inert, and a manifest-declared hook is live wherever the manifest sits.
|
|
80
|
+
|
|
81
|
+
Reported by a consumer during a 3.0.0 adoption pass, with the defect proven from committed cassettes:
|
|
82
|
+
10 of theirs carried a `SessionStart:startup` hook_started/hook_response pair from a manifest-only
|
|
83
|
+
declaration, including one recorded at `container`.
|
|
84
|
+
|
|
85
|
+
### Documentation
|
|
86
|
+
|
|
87
|
+
- **The pruned-agent-binary failure now has a documented remedy, with the tradeoffs named.** A Desktop
|
|
88
|
+
update deletes the prior version's staged ELF while often leaving an empty version directory, so a
|
|
89
|
+
scenario pinning that agent version dies with `baseline agent binary not found`. The code anticipates
|
|
90
|
+
this by name, but the remedy appeared in no skill surface at all — and an installed companion skill
|
|
91
|
+
ships without repo `docs/`, so the agent most likely to hit it had nothing to read. Now in
|
|
92
|
+
`docs/gotchas.md` and the skill's own reference, stating that the three remedies are NOT equivalent:
|
|
93
|
+
repin for a reproducibility-bound suite, `latest` for a one-off (a pin rots silently, `latest` drifts
|
|
94
|
+
silently — opposite failure modes), and `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` only to unblock once,
|
|
95
|
+
never in CI, since it runs the newest sibling and downgrades the sha check to advisory.
|
|
96
|
+
|
|
97
|
+
- **`permissive_auto_allow` and script-running skills is a `protocol`-specific hazard, not a general one.**
|
|
98
|
+
`Bash` is off-registry (only `Read`/`Glob`/`Grep` are default-allow), but the harness only decides how to
|
|
99
|
+
ANSWER a permission ask — whether the agent ASKS varies by command AND tier. Measured: at `protocol`, `echo
|
|
100
|
+
hello` produces no ask (green) while `python3 -c "print(42)"` produces one and fails the guard; the
|
|
101
|
+
identical command at `container` produces no ask (green), and `hostloop` routes shell through
|
|
102
|
+
`mcp__workspace__bash` entirely. So a skill running `python3 ${CLAUDE_SKILL_DIR}/scripts/…` goes red at L0
|
|
103
|
+
and green on the sandboxed tiers, while a hello-world probe is green everywhere and teaches the wrong
|
|
104
|
+
expectation. `references/fidelity-and-answers.md` carries the measured table and the three ways through; why
|
|
105
|
+
the host CLI asks for one command and not another is explicitly left untraced, and `microvm` is named as
|
|
106
|
+
unmeasured.
|
|
107
|
+
|
|
108
|
+
## [3.0.0] — 2026-08-29
|
|
109
|
+
|
|
110
|
+
### Breaking
|
|
111
|
+
|
|
112
|
+
- **`l0_plugin_divergence` is renamed `l0_host_config_contamination`**, and its modifier
|
|
113
|
+
`allow_l0_plugin_divergence` is renamed `allow_l0_host_config_contamination`. The `RunResult` field
|
|
114
|
+
`l0PluginDivergence` becomes `l0HostConfigContamination`. Verdict-signal codes and the scenario schema
|
|
115
|
+
are covered surfaces ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)), so this is
|
|
116
|
+
a MAJOR bump. The old name described plugin *delivery* diverging at L0; delivery now works, and what the
|
|
117
|
+
signal actually reports is that the run read the operator's real config dir. A scenario asserting the old
|
|
118
|
+
key must rename it; a consumer keying on the old code must too.
|
|
119
|
+
|
|
120
|
+
- **The signal's firing conditions changed with it.** It fired when a session declared plugin dirs; it now
|
|
121
|
+
fires when `protocol` reads the operator's real config dir — tested on the dir the agent will actually
|
|
122
|
+
read, so a pinned `plugins.config_dir` is caught even with the managed branch nominally active. A
|
|
123
|
+
`skills.local`-only protocol run that previously passed can now fail, and a sealed protocol run with
|
|
124
|
+
plugins that previously failed now passes.
|
|
125
|
+
|
|
126
|
+
|
|
127
|
+
### Added
|
|
128
|
+
|
|
129
|
+
- **`allow_host_hooks` (scenario key) and `--allow-host-hooks` (`chat` / `skill`).** Consent to
|
|
130
|
+
running a staged plugin's hooks as native host processes at `protocol`. Top-level like
|
|
131
|
+
`allow_host_writes` rather than a verdict modifier — it gates the SPAWN, it does not suppress a signal.
|
|
132
|
+
|
|
133
|
+
**Version floor: `allow_host_hooks` needs cowork-harness ≥ 3.0.0.** The scenario loader is a `z.strictObject`, so an older CLI does NOT fall back to the default — it hard-errors `Unrecognized key: "allow_host_hooks"` and exits 2 (verified). Adopting the key is a floor bump for every consumer of that scenario.
|
|
134
|
+
|
|
135
|
+
- **A committed live probe for L0 plugin delivery** — `examples/probes/l0-plugin-delivery.scenario.yaml`
|
|
136
|
+
plus its fixture and session. It asserts the plugin under test reaches the agent's init inventory at
|
|
137
|
+
`protocol`, which no unit test can see: the argv builder was correct in isolation and the defect was a
|
|
138
|
+
missing call site. It pins the sealed config-dir branch deliberately — off it, the operator's own
|
|
139
|
+
installed plugins are in the inventory and a name collision would satisfy the assertion with or without
|
|
140
|
+
`--plugin-dir`, measuring the machine instead of the argv. The fixture also ships an agent, deliberately
|
|
141
|
+
unasserted: there is no `agent_available` assertion key (tool / skill / connector have one, agents do
|
|
142
|
+
not), and naming that gap is better than inventing a key to hide it.
|
|
143
|
+
|
|
144
|
+
### Changed
|
|
145
|
+
|
|
146
|
+
- **`protocol` (L0) now passes `--plugin-dir`, so a declared plugin or skill dir is actually delivered.**
|
|
147
|
+
It previously passed no `--plugin-dir` and `local_plugins` never reached the generated `settings.json`,
|
|
148
|
+
so **the positional argument was silently inert**: `cowork-harness chat <dir> --fidelity protocol` (and
|
|
149
|
+
the `skill` / `probe-dispatch` equivalents) measured whatever the operator had installed rather than the
|
|
150
|
+
tree they passed. Live-verified: a bare skill dir and a plugin root both register at L0 now, with their
|
|
151
|
+
skills and their declared agents. Expect scenarios that previously passed *vacuously* — asserting against
|
|
152
|
+
an agent that never had the plugin — to start failing honestly.
|
|
153
|
+
|
|
154
|
+
- **A plugin's hooks and MCP servers now run at `protocol`, and hooks require consent.** Loading a plugin
|
|
155
|
+
means the CLI executes its `<plugin>/hooks/hooks.json` as **native host processes** — the operator's
|
|
156
|
+
account and environment, no container sandbox — and opens its declared MCP servers. `protocol` therefore
|
|
157
|
+
refuses to spawn when a staged plugin declares runnable hooks unless the scenario sets
|
|
158
|
+
`allow_host_hooks: true` (or `--allow-host-hooks` for `chat`/`skill`); a plugin without hooks
|
|
159
|
+
needs no opt-in, and a misplaced root-level `hooks.json` (which cannot execute) does not trigger it. A
|
|
160
|
+
per-run disclosure is printed even when consent was given. Mirrors the existing `allow_host_writes` gate.
|
|
161
|
+
|
|
162
|
+
- **`protocol` takes the managed config dir when any credential is in the environment.**
|
|
163
|
+
`CLAUDE_CODE_OAUTH_TOKEN` and `ANTHROPIC_AUTH_TOKEN` now select it alongside `ANTHROPIC_API_KEY`, and the
|
|
164
|
+
token is injected into the agent's env (a managed config dir with no credential yields "Not logged in").
|
|
165
|
+
This severs host plugin/skill/MCP discovery — verified: a token-only L0 run delivered the plugin under
|
|
166
|
+
test and **zero** host plugins. `COWORK_MANAGED_CONFIG=0` suppresses only the token-derived branch,
|
|
167
|
+
never the `ANTHROPIC_API_KEY` CI path; an unrecognized value is now rejected rather than silently
|
|
168
|
+
selecting the managed branch (`COWORK_MANAGED_CONFIG=false` previously did).
|
|
169
|
+
|
|
170
|
+
- **`l0_plugin_divergence` now reports contamination rather than the `--plugin-dir` layout.** Delivery is
|
|
171
|
+
fixed, so the signal fires when L0 runs against the operator's **real** config dir, and it tests the dir
|
|
172
|
+
the agent will actually read — a pinned `plugins.config_dir` reaches host discovery with the managed
|
|
173
|
+
branch nominally active. It stays a hard fail: nothing else catches this (the host-inventory scan runs
|
|
174
|
+
only at cassette-record time, and `host_path_leak`'s default-fail is skipped at this tier).
|
|
175
|
+
|
|
176
|
+
- **`workflow-authoring` added to the built-in skill roster** (`KNOWN_BUILTIN_SKILLS`), measured against a
|
|
177
|
+
sealed `protocol` run rather than a contaminated one.
|
|
178
|
+
|
|
179
|
+
- **Platform baseline `desktop-1.40609.0` (agent `2.1.247`).** The Cowork system prompt, both sub-agent
|
|
180
|
+
appends, the egress allowlist, `spawn.env`, `tools[]`/`allowedTools` and `mountLayout` are all re-derived
|
|
181
|
+
unchanged. The one `sync` delta was a pure refactor: W2's `CLAUDE_CODE_ENTRYPOINT` deployment ternary is
|
|
182
|
+
now a hoisted helper, and W1 hard-sets `local-agent` after the spread either way, so the derived value
|
|
183
|
+
does not move. The VM rootfs is re-captured — node `v22.22.3` → `v22.23.2`, pip 144 → 136 packages
|
|
184
|
+
(`xlrd` added; the nine removals are Ubuntu system packages, no analysis library dropped).
|
|
185
|
+
|
|
186
|
+
### Fixed
|
|
187
|
+
|
|
188
|
+
- **`microvm` resolves its agent binary through the shared resolver instead of deriving the path itself.**
|
|
189
|
+
It read `agentBinary.stagedPath` raw and handed it to the guest mount, so a pin that Claude Desktop had
|
|
190
|
+
pruned surfaced as `env: 'claude': No such file or directory` and exit 127 — the least informative
|
|
191
|
+
message possible for a condition the other three tiers name precisely. The quieter half mattered more:
|
|
192
|
+
the raw path also skipped `verifiedElf`, leaving the one tier that actually **executes** the ELF in a VM
|
|
193
|
+
as the only one not verifying it against the baseline pin, while `container` hard-fails on the same
|
|
194
|
+
mismatch. Routing it through `resolveAgentBinary` restores all three safeguards at once — existence
|
|
195
|
+
check, sha verification and the pruned-binary fallback — and removes a fourth derivation of a rule
|
|
196
|
+
`baseline.ts` already documented `microvm` as following. Resolution happens **before** the
|
|
197
|
+
already-Running reuse short-circuit, since a VM created while the binary was present keeps a mount at
|
|
198
|
+
the pruned path — the originally reported state. `vm status` / `vm prune` / `doctor` keep working when
|
|
199
|
+
the binary is missing (that is when an operator reaches for them) and degrade to the pinned path rather
|
|
200
|
+
than throwing.
|
|
201
|
+
|
|
202
|
+
- **`sync` resolves a spawn-env value expression that is a hoisted one-line helper** (`uH(n.type)`), where
|
|
203
|
+
it previously refused the baseline. The branch is narrow by construction: the argument must be a `.type`
|
|
204
|
+
member and the callee body must be exactly the literal deployment ternary over its own parameter, resolved
|
|
205
|
+
in the window's own chunk — a two-character minified name hopped against the joined bundle lands on an
|
|
206
|
+
unrelated helper. Anything else stays unresolvable rather than guessed.
|
|
207
|
+
|
|
208
|
+
- **`provenance.asarGateIds` now includes the gate-defaults map's bare-numeric keys.** The scan matched
|
|
209
|
+
quoted literals only, so any gate that is keyed in that map and never read through a quoted id was
|
|
210
|
+
missing — `provenance.asarGateIds` 252 → 291, and the 1.40609.0 delta corrects from +27/−4 to +30/−4.
|
|
211
|
+
This is not the bare-*number* scan the extractor deliberately rejects: that matches any numeric and adds
|
|
212
|
+
1687 ids over the same bundle, while the defaults-map entry shape adds 39.
|
|
213
|
+
|
|
214
|
+
### Documentation
|
|
215
|
+
|
|
216
|
+
- **A plugin's declared MCP servers are a documented fidelity gap.** Production replaces them with
|
|
217
|
+
zero-tool SDK stubs named `plugin:<plugin>:<server>` — remote (`url` + http/sse) unconditionally, local
|
|
218
|
+
and `.mcpb` under an MCP policy — while the harness stages plugins with `--plugin-dir` and lets the CLI
|
|
219
|
+
open the real ones. A plugin under test therefore sees a tool surface production would not give it, in
|
|
220
|
+
both directions and silently. Not modeled: the remote rule is mechanically reproducible, but the
|
|
221
|
+
local/`.mcpb` rule is conditioned on Desktop policy state the harness has no source for.
|
|
222
|
+
|
|
223
|
+
- **The auto-mode permission rubric gap is tier-independent.** `docs/fidelity-gaps.md` scopes it to both
|
|
224
|
+
loops rather than VM-loop only, and records that the rubric's Filesystem section is tier-dependent
|
|
225
|
+
model-visible text — a different kind of divergence from a permission verdict.
|
|
226
|
+
|
|
9
227
|
## [2.5.0] — 2026-08-28
|
|
10
228
|
|
|
11
229
|
### Added
|