cowork-harness 2.0.1 → 2.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +29 -16
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -3
- package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
- package/CHANGELOG.md +366 -0
- package/README.md +25 -8
- package/RELEASING.md +7 -2
- package/SPEC.md +24 -4
- package/dist/assert.js +18 -3
- package/dist/baseline.js +37 -15
- package/dist/cli.js +5 -1
- package/dist/prompt.js +5 -2
- package/dist/redact.js +44 -0
- package/dist/run/budget.js +21 -6
- package/dist/run/cassette.js +76 -14
- package/dist/run/chat-result.js +1 -0
- package/dist/run/chat.js +2 -0
- package/dist/run/execute.js +21 -6
- package/dist/run/run.js +19 -0
- package/dist/runtime/argv.js +20 -1
- package/dist/runtime/container.js +11 -3
- package/dist/runtime/hostloop.js +5 -1
- package/dist/runtime/lima.js +7 -1
- package/dist/runtime/microvm.js +22 -4
- package/dist/session.js +4 -0
- package/dist/types.js +1 -1
- package/docs/cassette.md +18 -6
- package/docs/invariants.md +5 -2
- package/docs/maintenance.md +5 -3
- package/docs/protocol.md +23 -5
- package/docs/scenario.md +9 -6
- package/examples/replays/README.md +1 -1
- package/fixtures/protocol/v1/dialog-response.json +10 -0
- package/fixtures/protocol/v1/elicit-response.json +10 -0
- package/fixtures/protocol/v1/elicitation-request.json +17 -0
- package/fixtures/protocol/v1/error-response.json +8 -0
- package/fixtures/protocol/v1/user-dialog-request.json +11 -0
- package/package.json +1 -1
- package/schema/cassette.v12.json +2 -2
- package/schema/protocol.v1.json +447 -79
- package/schema/run-result.json +4 -0
- package/schema/scenario.schema.json +1 -1
- package/scripts/check-versions.ts +134 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 2.0
|
|
7
|
-
tracks-harness: cowork-harness 2.0
|
|
6
|
+
version: 2.2.0
|
|
7
|
+
tracks-harness: cowork-harness 2.2.0 (baseline desktop-1.34493.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.0
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.2.0` (baseline
|
|
26
26
|
> `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.0
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.2.0"`. **Pin `@^2.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -656,7 +656,7 @@ than a stuck `"running"`.)
|
|
|
656
656
|
### Place assertions in the right CI lane
|
|
657
657
|
|
|
658
658
|
CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
|
|
659
|
-
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@
|
|
659
|
+
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v2` (a packaged GitHub Action with a
|
|
660
660
|
PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
|
|
661
661
|
the four-stage pipeline.
|
|
662
662
|
|
|
@@ -1,19 +1,23 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 2.0
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
7
7
|
|
|
8
8
|
```yaml
|
|
9
|
-
- uses: yaniv-golan/cowork-harness@
|
|
9
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
10
10
|
with:
|
|
11
11
|
command: replay
|
|
12
12
|
path: cassettes/
|
|
13
|
+
version: "^2" # hold the major; see below
|
|
13
14
|
```
|
|
14
15
|
|
|
15
|
-
The Action's `version` input defaults to `latest
|
|
16
|
-
|
|
16
|
+
**These recipes pin `version: "^2"`.** The Action's `version` input *defaults* to `latest`, which means a
|
|
17
|
+
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
|
+
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
|
+
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
+
(e.g. `version: "2.2.0"`) instead when you want byte-reproducible CI.
|
|
17
21
|
|
|
18
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -43,10 +47,11 @@ jobs:
|
|
|
43
47
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
44
48
|
# Background on the provenance chain: the "Agent-binary provenance" section of
|
|
45
49
|
# https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
|
|
46
|
-
- uses: yaniv-golan/cowork-harness@
|
|
50
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
47
51
|
with:
|
|
48
52
|
command: run
|
|
49
53
|
path: scenarios/
|
|
54
|
+
version: "^2"
|
|
50
55
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
51
56
|
```
|
|
52
57
|
|
|
@@ -62,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
62
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
63
68
|
|
|
64
69
|
```yaml
|
|
65
|
-
- run: npm i -g "cowork-harness@^2.0
|
|
70
|
+
- run: npm i -g "cowork-harness@^2.2.0"
|
|
66
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
67
72
|
# no silent false-greens. WITHOUT --strict this
|
|
68
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -110,20 +115,28 @@ Action has no input for, and it creates a coupling nothing checks:
|
|
|
110
115
|
|
|
111
116
|
> **If a flag in `extra-args` was added in release X, floor that step's `version` to `>=X`.**
|
|
112
117
|
|
|
113
|
-
`version` defaults to `latest`, and accepts **any npm range** — not just an exact pin.
|
|
118
|
+
`version` defaults to `latest`, and accepts **any npm range** — not just an exact pin. Leave it off unless
|
|
119
|
+
you have a reason:
|
|
114
120
|
|
|
115
121
|
```yaml
|
|
116
|
-
- uses: yaniv-golan/cowork-harness@
|
|
122
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
117
123
|
with:
|
|
118
124
|
command: lint
|
|
119
125
|
path: scenarios/
|
|
120
|
-
version: "
|
|
121
|
-
extra-args: --min-severity WARN
|
|
126
|
+
version: "^2" # holds the major
|
|
127
|
+
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 2.x satisfies that
|
|
122
128
|
```
|
|
123
129
|
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
130
|
+
**If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
|
|
131
|
+
floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
|
|
132
|
+
recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
|
|
133
|
+
major instead — `version: "^2"`, which is what the steps above use — keeping the floor's intent while
|
|
134
|
+
stopping at the major boundary. An exact
|
|
135
|
+
pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
|
|
136
|
+
moment a recipe adopts a newer flag.
|
|
137
|
+
|
|
138
|
+
Without a satisfied floor, an older CLI fails the step with `unrecognized arguments: --min-severity WARN`
|
|
139
|
+
(exit 2, wrapped in an `ok:false` envelope) — it does **not** degrade gracefully.
|
|
127
140
|
|
|
128
141
|
|
|
129
142
|
The harness has two execution lanes with different cost, coverage, AND infrastructure requirements.
|
|
@@ -307,7 +320,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
307
320
|
|
|
308
321
|
## GitHub Actions sketch
|
|
309
322
|
|
|
310
|
-
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@
|
|
323
|
+
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v2` does
|
|
311
324
|
in one step (see the top of this doc) — reach for this form when you need independent per-command
|
|
312
325
|
gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
|
|
313
326
|
equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
|
|
@@ -327,7 +340,7 @@ jobs:
|
|
|
327
340
|
with: { node-version: '24' }
|
|
328
341
|
- uses: actions/setup-python@v5
|
|
329
342
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
330
|
-
- run: npm i -g "cowork-harness@^2.0
|
|
343
|
+
- run: npm i -g "cowork-harness@^2.2.0"
|
|
331
344
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
332
345
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
333
346
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -356,7 +369,7 @@ jobs:
|
|
|
356
369
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
357
370
|
fi
|
|
358
371
|
- if: steps.guard.outputs.live == 'true'
|
|
359
|
-
run: npm i -g "cowork-harness@^2.0
|
|
372
|
+
run: npm i -g "cowork-harness@^2.2.0"
|
|
360
373
|
- if: steps.guard.outputs.live == 'true'
|
|
361
374
|
run: cowork-harness run scenarios/ --output-format json
|
|
362
375
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 2.0
|
|
3
|
+
Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.0
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.2.0`
|
|
4
4
|
(baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -342,8 +342,8 @@ same set live from the schema.
|
|
|
342
342
|
| `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
|
|
343
343
|
| `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
|
|
344
344
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
345
|
-
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives
|
|
346
|
-
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.
|
|
345
|
+
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
346
|
+
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
347
347
|
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
|
|
348
348
|
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered — what the user was actually SHOWN. `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable |
|
|
349
349
|
| `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 2.0
|
|
5
|
+
Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -66,7 +66,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](htt
|
|
|
66
66
|
| `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
|
|
67
67
|
| `scenarioSource` | Relative path to the authored YAML this was recorded from |
|
|
68
68
|
| `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
|
|
69
|
-
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
69
|
+
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
70
70
|
| `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
|
|
71
71
|
| `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
|
|
72
72
|
| `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
|
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,372 @@ All notable changes to this project are documented here. The format is based on
|
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
|
|
5
5
|
[Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
|
|
6
6
|
|
|
7
|
+
## [2.2.0] — 2026-08-25
|
|
8
|
+
|
|
9
|
+
### Upgrade impact
|
|
10
|
+
|
|
11
|
+
Two behaviour changes can turn a previously-green run red. Neither breaks a covered surface
|
|
12
|
+
([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) — both make an assertion report
|
|
13
|
+
what it always documented — so they ship in a minor:
|
|
14
|
+
|
|
15
|
+
- **`no_scratchpad_leak` at `container` can now FAIL.** It was measuring containment against a root in a
|
|
16
|
+
different path space, so every presented file classified `leaked: false` and the check passed
|
|
17
|
+
vacuously. A scenario whose skill genuinely leaves a presented file in the scratchpad will now red —
|
|
18
|
+
that is the leak the key exists to catch.
|
|
19
|
+
- **`baseline: desktop-1.11847.5` is now refused at `container`/`hostloop`/`microvm`.** It carries no
|
|
20
|
+
`spawn` block, so those tiers cannot reproduce Cowork's toolset — a run on it launched an agent with no
|
|
21
|
+
file or bash tools and still reported a verdict. Use a `sync`-recorded baseline, or `fidelity: protocol`.
|
|
22
|
+
|
|
23
|
+
### Fixed
|
|
24
|
+
|
|
25
|
+
- **`present_files_called` no longer reports "the tool was never called" about a run that called it, and a
|
|
26
|
+
hostloop delivery scenario can be recorded with a host-path redaction policy at last.** The assertion
|
|
27
|
+
read presence off `RunResult.presentedFiles`, which is a *classification* of each presented path
|
|
28
|
+
(scratchpad? promoted? leaked?) and requires an absolute path to compute. At `hostloop` a presented path
|
|
29
|
+
is a real host path, so the shipped redaction policy rewrites it to
|
|
30
|
+
`[REDACTED:local-path:<hash>]/mnt/outputs/report.html` — correct, documented, and deliberately ordered so
|
|
31
|
+
the mount tail survives — and the classifier then drops every entry as un-normalizable. The list came
|
|
32
|
+
back empty and the assertion stated, as fact, that the tool had never been called.
|
|
33
|
+
|
|
34
|
+
Because `record` replays the base and redacted cassettes and refuses to write when the verdict differs,
|
|
35
|
+
this was not merely a wrong message: **no cassette asserting `present_files_called` could be recorded at
|
|
36
|
+
that tier at all**, so the assertion had never executed on the replay lane. The refusal was right — the
|
|
37
|
+
redacted cassette genuinely could not support the assert — but the defect it was reporting was in the
|
|
38
|
+
harness, not in the recording.
|
|
39
|
+
|
|
40
|
+
Presence now comes from `RunResult.presentFilesCalls`, a new count of the `present_files` invocations
|
|
41
|
+
that carried a well-formed `file_path`, taken from the tool_use input's shape and never from a path's
|
|
42
|
+
content, so redaction cannot alter it. Classification is untouched: `presentedFiles` still drops what it
|
|
43
|
+
cannot resolve, and `no_scratchpad_leak` still reads it. A run recorded before the field falls back to
|
|
44
|
+
the old `presentedFiles`-non-empty test, so no existing green changes.
|
|
45
|
+
|
|
46
|
+
A run whose every `present_files` call carried an unusable path now reports **cannot verify** rather than
|
|
47
|
+
"never called" — the tool *was* invoked, and the harness already knew it. (The malformed count is
|
|
48
|
+
deliberately kept out of that message: the record self-check normalizes `[REDACTED…]` tokens out of
|
|
49
|
+
failing messages but not digits, so an interpolated count that differed between the two replays would
|
|
50
|
+
refuse a cassette that is otherwise fine to write.)
|
|
51
|
+
|
|
52
|
+
Pinned as an invariant ([docs/invariants.md](./docs/invariants.md)) with the end-to-end record self-check
|
|
53
|
+
as its test anchor — the case that could not be recorded, and so had never run in CI.
|
|
54
|
+
|
|
55
|
+
- **Guest paths are built from the tree the harness stages, not from a baseline's recorded mount layout.**
|
|
56
|
+
`resolveMounts` returned `mountLayout.mntRoot` verbatim — or, when that field was absent and the recorded
|
|
57
|
+
`sessionRoot` already ended in `/mnt`, the session root itself. But the staged tree is always
|
|
58
|
+
`<sessionRoot>/mnt`: `stageWorkspace` creates it there and `dockerRunArgv` nests its read-only binds
|
|
59
|
+
there. On a baseline recording anything else, `--plugin-dir` therefore pointed one directory **above**
|
|
60
|
+
the staged plugin tree, so the plugin under test never loaded. `mntRoot` is now derived from the session
|
|
61
|
+
root, and a recorded layout the harness cannot stage is reported as a fidelity divergence at spawn
|
|
62
|
+
instead of silently composing a path no stager creates.
|
|
63
|
+
|
|
64
|
+
Guest paths now anchor on `sessionRoot` — the bind target — rather than `cwd`, which is only where the
|
|
65
|
+
agent's process starts (production's own working directory is a folder mount or `outputs`, so the two are
|
|
66
|
+
not interchangeable even though every synced baseline records them equal). `dockerRunArgv` takes an
|
|
67
|
+
explicit `agentCwd` for `-w`.
|
|
68
|
+
|
|
69
|
+
Two more derivations of the same rule are gone: `prompt.ts` had a private `/sessions/<id>` +`/mnt`
|
|
70
|
+
literal (so the prompt could describe a tree the runtimes had not staged), and **microvm** read the
|
|
71
|
+
agent's root from the baseline when lima structurally mounts it at `/sessions/<sessionId>` — which put
|
|
72
|
+
`CLAUDE_CONFIG_DIR` and `--mcp-config` at paths nothing stages. A baseline recording a cwd that tier
|
|
73
|
+
cannot honour is now **refused**, not warned about: the guest `cd` would otherwise succeed at the wrong
|
|
74
|
+
directory.
|
|
75
|
+
|
|
76
|
+
- **A baseline with no `spawn` block is refused at the sandbox tiers.** That block carries the tool set,
|
|
77
|
+
pre-approvals, effort default and config-dir location, and the `?? []` fallbacks meant a run would launch
|
|
78
|
+
an agent with **no Read/Write/Bash/Skill/Task at all** and still report a verdict. `fidelity: protocol`
|
|
79
|
+
builds its own argv and is unaffected.
|
|
80
|
+
|
|
81
|
+
- **`no_scratchpad_leak` can see a container leak again — the session root was in the wrong path space.** The
|
|
82
|
+
root that `presentedFiles`' promoted/leaked classification is measured from was derived by the caller as
|
|
83
|
+
`<run-dir>/work/session`, a HOST path, and handed to every non-`protocol` tier. But the path space the
|
|
84
|
+
agent reports in is per-tier: at `container` (the default) it runs inside the sandbox and reports
|
|
85
|
+
`/sessions/<id>/…`, and only at `hostloop` does it run natively and report host paths. Measured against a
|
|
86
|
+
host root, no VM path is ever inside the root, so every presented file classified `leaked: false` —
|
|
87
|
+
including the handler's copy-failure branch, which returns the source path unchanged when a file is
|
|
88
|
+
blocked-extension, a directory, or absent. `no_scratchpad_leak` evaluates at `container` and nowhere else,
|
|
89
|
+
so the assertion that exists to catch that leak could not catch it, and `verdict.ts`'s delivery check read
|
|
90
|
+
the same `leaked: false` as a successful delivery.
|
|
91
|
+
|
|
92
|
+
Each runtime now reports the session root it actually launched the agent with, and the run consumes that
|
|
93
|
+
instead of deriving a second one — the two can no longer drift into different spaces. `chat` sets it on
|
|
94
|
+
both serving tiers as well; it never did, so a hostloop chat's `presentedFiles` was inverted in its result
|
|
95
|
+
file. `protocol` and `microvm` serve no `present_files` and keep the cwd fallback.
|
|
96
|
+
|
|
97
|
+
Independently, the classifier now fails CLOSED on a space mismatch: if the agent's cwd is not at or inside
|
|
98
|
+
the session root, the presented batch counts as malformed (`no_scratchpad_leak` → cannot verify) rather
|
|
99
|
+
than being graded against a root it cannot be compared to. `leaked: true` is not derivable in that state
|
|
100
|
+
either, and a `leaked: false` verdict there is exactly the vacuous pass the key exists to prevent.
|
|
101
|
+
|
|
102
|
+
### Added
|
|
103
|
+
|
|
104
|
+
- **First live coverage for `present_files_called` / `no_scratchpad_leak`.** No scenario in the repo
|
|
105
|
+
asserted either key, which is how a session root in the wrong path space could record `leaked: false`
|
|
106
|
+
for every presented file without anything noticing. `e2e/scenarios/smoke-present-files.yaml` writes a
|
|
107
|
+
file OUTSIDE `mnt/` and delivers it, so the promotion is real and the pair is non-vacuous; it runs in
|
|
108
|
+
CI's live e2e loop. Measured on a live container run: `presentFilesCalls: 1`, promoted `true`, leaked
|
|
109
|
+
`false`.
|
|
110
|
+
|
|
111
|
+
- **`RunResult.presentFilesCalls`** — the count of `present_files` invocations that carried a well-formed
|
|
112
|
+
`file_path`, in `result.json` and [schema/run-result.json](./schema/run-result.json). Content-class, so a
|
|
113
|
+
replay re-drive reproduces it; read by `run`, `replay` and `verify-run` alike, and absent (not `0`) on a
|
|
114
|
+
result written before this release, which is what the assertion's fallback distinguishes. Use it to
|
|
115
|
+
answer "did the agent deliver anything?" from a result file without interpreting `presentedFiles`'
|
|
116
|
+
promoted/leaked classification.
|
|
117
|
+
|
|
118
|
+
## [2.1.0] — 2026-08-24
|
|
119
|
+
|
|
120
|
+
### Changed
|
|
121
|
+
|
|
122
|
+
- **`rehash` leads with the split, and PARTIAL success has its own exit code (`4`).** The summary used to
|
|
123
|
+
print last, after every per-file line had scrolled past — and those two counts *are* the decision:
|
|
124
|
+
commit what migrated and budget a re-record for the rest, versus nothing here is salvageable. It now
|
|
125
|
+
prints first, with the per-file lines as the detail behind it.
|
|
126
|
+
|
|
127
|
+
**`4 migrated, 18 failed` and `0 migrated, 22 failed` both exited `1`**, which made a shell consumer
|
|
128
|
+
unable to tell apart two situations demanding opposite responses. The JSON envelope always carried the
|
|
129
|
+
split as `migrated`/`skipped`/`errors`; a bare terminal run did not. `rehash` now exits `0` (all
|
|
130
|
+
migrated, or nothing needed migrating) / **`4`** (partial) / `1` (nothing migrated and at least one
|
|
131
|
+
could not) / `2` (usage).
|
|
132
|
+
|
|
133
|
+
**This is a behaviour change for anyone branching on `rehash`'s exit code** — a script testing
|
|
134
|
+
`rc == 1` for "something failed" will miss the partial case, though `rc != 0` is unaffected. It ships in
|
|
135
|
+
a minor deliberately: `rehash`'s codes were documented **nowhere** — zero mentions in its `--help`,
|
|
136
|
+
absent from SPEC §11 — so there was no published contract to break, and a consumer on `rc == 1` was
|
|
137
|
+
relying on observed behaviour. They are now documented in both places. `4` rather than `3` because the
|
|
138
|
+
code space is per-command and `3`'s "could not verify" meaning is load-bearing on `verify-cassettes`.
|
|
139
|
+
|
|
140
|
+
- **The batch cost estimate now reports its basis instead of claiming authority.** `estimatedCostUsd` is
|
|
141
|
+
`sum(max(local run history))` — a max over whatever *this machine* has run. The line already qualified
|
|
142
|
+
the partially-priced case with `— LOWER BOUND`, but the fully-priced case said
|
|
143
|
+
`(all N scenario(s) priced from prior runs)`, which is an active claim of completeness and was the one
|
|
144
|
+
case that said nothing qualifying. At a new baseline or a new agent binary the history describes a
|
|
145
|
+
materially different configuration, so the estimate can **under**-predict exactly when it is most
|
|
146
|
+
consulted — a consumer wrote "that is the ceiling, not the scope" into a plan off this line and had to
|
|
147
|
+
retract it.
|
|
148
|
+
|
|
149
|
+
The line now carries `basis: N prior run(s) on THIS machine, thinnest scenario has M; a max over that
|
|
150
|
+
history, NOT a bound`, and the JSON payload gains `estimateBasis`
|
|
151
|
+
(`{source, pricedRuns, thinnestScenarioRuns}`). `thinnest` is the useful half: a scenario with one prior
|
|
152
|
+
run contributes a single sample, not a worst case. Wired through `pricedRunCount`, which had been
|
|
153
|
+
exported and doc-commented *"for messages that report their own basis"* with zero callers.
|
|
154
|
+
|
|
155
|
+
- **Pre-epoch cassettes: we could report ordinary content drift, and chose not to.** When a cassette
|
|
156
|
+
recorded before the 2.0.0 hash-format epoch is read, `rehash` recomputes the **legacy** digest over the
|
|
157
|
+
current tree and compares it to the legacy digest in the cassette. On a mismatch that is a positive
|
|
158
|
+
determination of ordinary content drift — strictly more informative than `unverifiable-skill`, and the
|
|
159
|
+
same determination 1.25.0 reported as warn-only `skill` drift.
|
|
160
|
+
|
|
161
|
+
Replay does not report it, and that is a decision rather than a limit. Reporting it would require
|
|
162
|
+
keeping a fold of the retired hash algorithm alive in the replay path — the most-run lane — permanently,
|
|
163
|
+
to soften one release's migration; the population it would help shrinks with every re-record. The
|
|
164
|
+
runtime cost would be nil (both digests fold from one tree walk); the cost is code that could never be
|
|
165
|
+
deleted. **"We can compute this and chose not to report it, so the legacy fold can die" is a different
|
|
166
|
+
claim from "this cannot be known", and the earlier phrasing implied the latter.** The remedy for a
|
|
167
|
+
pre-epoch cassette is `rehash` where it can prove content unchanged, and a re-record where it cannot.
|
|
168
|
+
|
|
169
|
+
### Fixed
|
|
170
|
+
|
|
171
|
+
- **The fast test lane had no global timeout, so 93 subprocess-spawning test files inherited vitest's 5s
|
|
172
|
+
default.** They measure 167-888ms locally — a fine margin until you remember the lane runs 344 files in
|
|
173
|
+
parallel across every core, and a CI runner is ~3x slower again. That is how an 888ms test crosses 5s;
|
|
174
|
+
it cost two red CI runs on unrelated PRs before anyone looked at the failure rather than re-running it.
|
|
175
|
+
`vitest.config.ts` now sets `testTimeout: 30_000` — ~34x the slowest measured non-e2e case, while a
|
|
176
|
+
genuine hang still fails ~6x faster than in the live lane (which sets 180s). Per-test values still win.
|
|
177
|
+
|
|
178
|
+
- **Redaction pattern ORDER is load-bearing, and now says so — plus a warning when it is wrong.** Patterns
|
|
179
|
+
apply in sequence over the accumulating output, so a bare catch-all placed ahead of a lookahead-anchored
|
|
180
|
+
rule for the same prefix matches first and eats the `/mnt/` tail the lookahead exists to preserve. The
|
|
181
|
+
shipped policy depends on that order: with it, a run-dir path redacts to
|
|
182
|
+
`[REDACTED:local-path:…]/mnt/outputs/report.md` and still resolves; without it the whole path is consumed
|
|
183
|
+
and `normalizeHostShapedForReplay` returns `null` — **every `computer://` structural-marker resolution
|
|
184
|
+
silently stops working, with no error and no finding.**
|
|
185
|
+
|
|
186
|
+
`loadRedactionPolicy` now warns, naming both pattern indices, when a policy is in the hazardous order.
|
|
187
|
+
Detection is deliberately conservative — it fires only when a later pattern's source is exactly an
|
|
188
|
+
earlier one's plus a trailing lookahead (modulo lazy quantifiers) — because regex subsumption is
|
|
189
|
+
undecidable in general and a false positive would train authors to ignore the warning.
|
|
190
|
+
|
|
191
|
+
The remainder is matched by **shape**, never by parsing the lookahead's body. A first cut used
|
|
192
|
+
`\(\?=[^()]*\)`, whose `[^()]*` silently skipped every lookahead containing a group — so
|
|
193
|
+
`(?=/mnt(?:/|$|[\s"'\\)\]]))`, the natural way to write "slash, end, or delimiter" and arguably more
|
|
194
|
+
correct than a bare `(?=/mnt/)`, went unflagged while being just as dangerous. Caught by a consumer
|
|
195
|
+
running it against their own policy, which is now a regression fixture. Not looking inside also sidesteps
|
|
196
|
+
escape- and char-class-awareness, since that policy carries an escaped `\)` inside a character class.
|
|
197
|
+
|
|
198
|
+
`docs/cassette.md` states the requirement next to the existing "stop before `/mnt/`" guidance, which had
|
|
199
|
+
the shape of the rule but not the ordering half. The new test pins the runtime consequence, not just the
|
|
200
|
+
detector: a reorder must make the link fail to normalize AND be flagged, so the syntactic check cannot
|
|
201
|
+
drift away from what it is standing in for.
|
|
202
|
+
|
|
203
|
+
- **The copy-pasteable Action steps now pin `version: "^2"`, and a guard requires it.** Bounding every
|
|
204
|
+
published npm floor last release fixed the *form* of a floor (`>=1.11.0` reads as a bound and silently
|
|
205
|
+
means "and every future major too") but removed the input from the recipes rather than correcting it —
|
|
206
|
+
so the shipped snippets carried no `version:` at all and fell back to the input's `latest` default. That
|
|
207
|
+
reproduced the exact footgun `action.yml`'s own description warns about two lines earlier: a CLI major
|
|
208
|
+
reaches a workflow the moment it is promoted, even though the `uses:` ref never changed.
|
|
209
|
+
|
|
210
|
+
`^2` was not among the alternatives weighed at the time, and it is the form `action.yml` itself
|
|
211
|
+
recommends: it holds the major, needs no patch number to remember, and only wants a human decision at
|
|
212
|
+
the next major bump. **Five** steps were unpinned, not the three in the CI recipe — `README.md` carries
|
|
213
|
+
two more.
|
|
214
|
+
|
|
215
|
+
`action-docs-sync` now requires every copy-pasteable step (a `uses:` line inside a fenced block with a
|
|
216
|
+
`with:`) to pin `^<package major>`; inline prose mentions are excluded, since there is nothing to pin.
|
|
217
|
+
Verified by mutation: dropping one `version:`, regressing a pin to `^1`, and bumping the package major
|
|
218
|
+
each fail. The guard also asserts it found the steps at all, because a parser that matches nothing
|
|
219
|
+
passes every assertion after it. Guarding a floor's FORM does not guarantee a floor is PRESENT — this
|
|
220
|
+
pins the behaviour instead.
|
|
221
|
+
|
|
222
|
+
- **The `sessionFingerprint` field set is now stated completely, and a guard discovers the sites that
|
|
223
|
+
state it.** The hash covers eight session fields; every place that enumerated them named six or fewer.
|
|
224
|
+
`web_fetch` was missing everywhere, `agent_env` was missing everywhere, and `docs/invariants.md`
|
|
225
|
+
also omitted `skills`. Fourteen sites carried the claim while the working assumption was three: **twelve
|
|
226
|
+
enumerated the set** — four in prose, two in the current cassette schema, six in the retained v9-v11
|
|
227
|
+
schemas — and **two denied it existed at all**. `SPEC.md` and `docs/scenario.md` both said the session
|
|
228
|
+
is "not drift-checked or fingerprinted". `model` genuinely is not hashed; connected folders and plugin/skill/MCP discovery are,
|
|
229
|
+
which is the half a reader would have trusted.
|
|
230
|
+
|
|
231
|
+
The `verify-cassettes` staleness message — the only enumeration a user ever sees — omitted `projects`,
|
|
232
|
+
and `docs/cassette.md` quoted it with `projects` present. Both are fixed and now pinned to each other
|
|
233
|
+
by test, so the doc cannot drift from the string again.
|
|
234
|
+
|
|
235
|
+
New **invariant 14** in `check:versions` derives the field set from `buildSessionFingerprint`'s shape
|
|
236
|
+
and **discovers** the enumeration sites rather than reading a list, so a new one is covered the day it
|
|
237
|
+
lands. Two deliberate limitations are recorded as tests rather than left to look covered: it cannot see
|
|
238
|
+
a flat denial (there is no enumeration to check — the coverage floor is what notices), and it cannot
|
|
239
|
+
see a deleted "only when set" qualifier. `schema/cassette.v{9,10,11}.json` are allowlisted as frozen
|
|
240
|
+
history: a retained schema documents the format as it shipped, and rewriting its description would make
|
|
241
|
+
it describe a shape its own consumers never saw.
|
|
242
|
+
|
|
243
|
+
Verified by mutation: run against the previous revision the guard flags all six then-existing sites
|
|
244
|
+
with the correct missing fields; renaming the shape literal makes it error rather than silently pass;
|
|
245
|
+
a ninth field invalidates a previously-complete site; and a whole-file token check — which would have
|
|
246
|
+
passed today and then never failed again — is rejected in favour of span-scoped matching.
|
|
247
|
+
|
|
248
|
+
- **The documented Action ref is now `@v2`, not `@main`** — 7 references across `README.md`,
|
|
249
|
+
`SKILL.md` and `ci-recipe.md`. `@main` was right when it was written: no alias tag had ever been
|
|
250
|
+
published, so naming one would have sent a copy-pasting reader to a `uses:` that 404s, and the guard's
|
|
251
|
+
own note said to revisit "once 1.0.0 ships". Two things had to be true first, and now are — `v2`/`v2.0`
|
|
252
|
+
point at a real release, and `release.yml` moves them on every stable release rather than leaving it to a
|
|
253
|
+
checklist. Recommending a floating tag nobody remembers to move is worse than recommending `@main`; that
|
|
254
|
+
was the actual situation while `v1` sat at 1.24.0.
|
|
255
|
+
|
|
256
|
+
`action-docs-sync` now derives the expected ref from `package.json`'s major instead of hardcoding it, so
|
|
257
|
+
the next major forces these docs to move with it rather than silently pointing a reader at the previous
|
|
258
|
+
line. `@main` is deliberately no longer accepted there: permitting both would let the recommendation
|
|
259
|
+
drift back with nothing noticing. Verified by mutation — regressing one reference to `@main` fails, and
|
|
260
|
+
setting the package version to 3.0.0 fails all three files.
|
|
261
|
+
|
|
262
|
+
**This changes nothing about which CLI you get.** The ref selects the Action; the CLI still comes from the
|
|
263
|
+
`version:` input, which still defaults to `latest`. `@v2` looks more like a version pin than `@main` did,
|
|
264
|
+
so that distinction matters more now, not less — it is spelled out in `README.md`'s Action section and in
|
|
265
|
+
`action.yml`'s own input description.
|
|
266
|
+
|
|
267
|
+
- **The CI recipe no longer teaches a bare version floor.** It
|
|
268
|
+
carried `version: ">=1.11.0"` — which reads as "at least 1.11.0" and silently means "and every future
|
|
269
|
+
major too", so a copy-paster gets the next major with no say in it. It was **not** broken today
|
|
270
|
+
(`--min-severity` still exists in 2.x, and `lint` reads no cassette, so 2.0.0's hash-format epoch never
|
|
271
|
+
applied to that step) — the defect was latent and in the FORM.
|
|
272
|
+
|
|
273
|
+
The bare floor was first dropped rather than corrected — `^1.11.0` would have frozen every new
|
|
274
|
+
copy-paster on the previous major, and `^2.0.1` needs remembering at each release — so the guidance moved
|
|
275
|
+
to prose. **That went one step too far, and the same release corrects it** (see the `version: "^2"` entry
|
|
276
|
+
above): dropping the input entirely falls back to `action.yml`'s `latest` default, which is the one
|
|
277
|
+
remaining unbounded form and the exact footgun the input's own description warns about two lines earlier.
|
|
278
|
+
`^2` was never among the alternatives weighed at the time, and it is what the recipes now carry. Reach for
|
|
279
|
+
an exact pin only when you want byte-reproducible CI. [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml)'s own description stops offering `>=1.11.0` and
|
|
280
|
+
`^1.11.0` as interchangeable — they are not, and it had been recommending the unbounded one.
|
|
281
|
+
|
|
282
|
+
`check:versions` invariant 13 now covers the Action's `version:` input, which it could not see before:
|
|
283
|
+
it keys on `@>=`, and `version: ">=1.11.0"` has no `@` — the same defect in different syntax, with no
|
|
284
|
+
coverage. Only the **unbounded** floor is rejected; verified against each form, `>=1.11.0 <3`, `^2`,
|
|
285
|
+
`2.0.1` and `latest` all pass.
|
|
286
|
+
|
|
287
|
+
- **`projects[].from` was missing from two places, and the second was a false green.** A connected project
|
|
288
|
+
is a host path exactly like a connected folder, and it was:
|
|
289
|
+
- **not resolved against the session file.** [`docs/session.md`](./docs/session.md) promises, without
|
|
290
|
+
qualification, that *"relative paths resolve from the session file's own directory"* — and it was true
|
|
291
|
+
of every path field except this one, which resolved against the **process CWD**. So the same session
|
|
292
|
+
file mounted different content depending on which directory you invoked from. The doc was right; the
|
|
293
|
+
resolver had simply skipped the field.
|
|
294
|
+
- **not part of the session fingerprint.** Swapping which directory is mounted at `.projects/<uuid>`
|
|
295
|
+
changed the run's inputs and `verify-cassettes` reported nothing — a false green in the gate whose job
|
|
296
|
+
is to notice that inputs moved. Folded in on the same **non-empty-only** terms as `agent_env`, so a
|
|
297
|
+
session with no `projects:` (and one with an explicit `projects: []`) hashes byte-identically to
|
|
298
|
+
before; only sessions that use the feature move. No committed cassette does.
|
|
299
|
+
|
|
300
|
+
**A cassette recorded before this is reported `unverifiable`, not clean.** Its hash contains nothing
|
|
301
|
+
about `projects[]`, so it cannot distinguish "the field was never covered" from "the mount changed since
|
|
302
|
+
record time" — and reporting the mismatch as a benign migration would put the same false green back in
|
|
303
|
+
the remedy. When everything else matches exactly, `verify-cassettes` says so and asks for a re-record to
|
|
304
|
+
gain the coverage. `sessionFingerprintDrift` remains `verify-cassettes`-only: none of this can change a
|
|
305
|
+
`replay` verdict, even under `--strict`.
|
|
306
|
+
|
|
307
|
+
Three enumerations of the covered fields gained `projects` — [`docs/cassette.md`](./docs/cassette.md),
|
|
308
|
+
[`docs/invariants.md`](./docs/invariants.md) and the shipped skill's `task-recipes.md`. `invariants.md`
|
|
309
|
+
also described this as a hash of the **resolved** session; it is the **authored, pre-resolution** shape,
|
|
310
|
+
deliberately, so the digest survives a different checkout — the function's own comment says so, and a
|
|
311
|
+
resolved hash could never match on another clone.
|
|
312
|
+
|
|
313
|
+
- **The published control-protocol schema rejected five kinds of frame the harness sends and answers.**
|
|
314
|
+
`schema/protocol.v1.json` described four request subtypes; the harness has always answered **six** —
|
|
315
|
+
adding `request_user_dialog` and `elicitation`/`side_question` — and it also sends a fail-closed
|
|
316
|
+
`subtype:"error"` response envelope, whose payload is a *string* under `error` rather than an object
|
|
317
|
+
under `response`, so a validator that knew only the success envelope rejected every one. Measured
|
|
318
|
+
against the previous schema: all five representative frames **REJECT**; against the new one, all five
|
|
319
|
+
accept. Anyone validating real traffic was seeing failures on frames the harness handles correctly.
|
|
320
|
+
|
|
321
|
+
Added as new `oneOf`/`anyOf` branches plus one new top-level response shape — **111 insertions, zero
|
|
322
|
+
deletions** in the surface baseline, so nothing existing was narrowed and no frame the schema already
|
|
323
|
+
accepted is affected. Both spellings the parser accepts are admitted (`dialogKind`/`dialog_kind`,
|
|
324
|
+
`mcp_server_name`/`server`, `message`/`prompt`) rather than guessing which the agent sends.
|
|
325
|
+
|
|
326
|
+
The golden vector pack grew with it, generated from the **real** envelope builders rather than
|
|
327
|
+
hand-authored lookalikes. The existing lockstep test — every schema definition must be exercised by a
|
|
328
|
+
vector — caught all five additions immediately, which is what forced those vectors to exist. One of them
|
|
329
|
+
asserts the error envelope does **not** validate as a success envelope, so the two shapes cannot be
|
|
330
|
+
quietly conflated later.
|
|
331
|
+
|
|
332
|
+
[`SPEC.md`](./SPEC.md) §12 now states the additive latitude for this surface explicitly. Its silence had
|
|
333
|
+
read as a prohibition, which is plausibly why three subtypes went undescribed rather than added — while
|
|
334
|
+
[`docs/protocol.md`](./docs/protocol.md)'s own versioning policy had said all along that an additive
|
|
335
|
+
variant is a v1 minor note, which is where the dated entry now lives.
|
|
336
|
+
|
|
337
|
+
- **The Marketplace alias tags are moved by the release workflow instead of by remembering.** `v1` sat at
|
|
338
|
+
1.24.0 through two releases because moving it was a checklist line. `release.yml`'s last step now points
|
|
339
|
+
`vX` and `vX.Y` at the release it just published, with two guards a hand-run `git tag -f` skips: a
|
|
340
|
+
**prerelease** tag moves nothing (the trigger accepts `v1.0.4-rc.1`, and pointing `v1` at an rc would
|
|
341
|
+
hand every `@v1` consumer a prerelease), and an alias **never moves backwards** — re-releasing an older
|
|
342
|
+
patch on a line moves `vX.Y` and leaves `vX` alone. Verified by executing the logic against a synthetic
|
|
343
|
+
tag set rather than by reading it: releasing `v1.20.5` while `v1.25.0` exists correctly skips `v1` and
|
|
344
|
+
still moves `v1.20`. It runs last, after publish and the GitHub Release, so a failure there cannot
|
|
345
|
+
half-publish anything.
|
|
346
|
+
|
|
347
|
+
Alongside it, the tags are now correct: **`v2` and `v2.0` created** (they did not exist, so nothing
|
|
348
|
+
pointed at the 2.x Action), and **`v1` moved 1.24.0 → 1.25.0**. Worth recording what that move did and
|
|
349
|
+
did not fix: the Action's whole surface — `action.yml` plus the `render.js` it loads — is **byte-identical
|
|
350
|
+
from 1.24.0 through 2.0.1** apart from three lines of input *description*. So a stale `v1` was a promise
|
|
351
|
+
the repo had stopped keeping, not a functional gap, and the one public `@v1` consumer pins `version:` on
|
|
352
|
+
every step and was never exposed to the `latest` default at all.
|
|
353
|
+
|
|
354
|
+
### Documentation
|
|
355
|
+
|
|
356
|
+
- **The invariants index said `check-versions.ts` has "no dedicated vitest file — it's a standalone script,
|
|
357
|
+
not a unit-testable module boundary".** Three exist, for the invariants whose logic is an exported pure
|
|
358
|
+
function: `check-cassette-version-claims`, `check-fingerprint-field-claims` and
|
|
359
|
+
`check-design-scope-note`. The claim had already been stale before this release.
|
|
360
|
+
|
|
361
|
+
- **The `uses:` ref pins the Action; the `version:` input pins the CLI — and only the second one holds a
|
|
362
|
+
major.** Both are documented as if pinning `@v1` bounded what you install. It does not: they move
|
|
363
|
+
independently, and `version:` defaults to `latest`, so promoting a CLI major reaches a workflow whose
|
|
364
|
+
`uses:` ref has not changed in months. Measured at the time of writing — `v1` points at **1.24.0** and
|
|
365
|
+
has never been moved, yet an `@v1` workflow with no `version:` input installs **2.x**. `README.md`,
|
|
366
|
+
`action.yml` (the text GitHub Marketplace renders) and `RELEASING.md`'s alias-tag step now say so, and
|
|
367
|
+
name the fix: pin the **input** (`version: ^2`), not the ref. Crossing 1.x → 2.x this way means the
|
|
368
|
+
hash-format epoch, so pre-v12 cassettes need `cowork-harness rehash <dir/>`.
|
|
369
|
+
|
|
370
|
+
`RELEASING.md`'s "move the major/minor tags" step additionally records what moving `vX` does *not* do,
|
|
371
|
+
since that step reads as the thing that controls consumer upgrades and is not.
|
|
372
|
+
|
|
7
373
|
## [2.0.1] — 2026-08-23
|
|
8
374
|
|
|
9
375
|
### Added
|