cowork-harness 2.0.1 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +29 -16
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
- package/CHANGELOG.md +255 -0
- package/README.md +25 -8
- package/RELEASING.md +7 -2
- package/SPEC.md +24 -4
- package/dist/cli.js +4 -1
- package/dist/redact.js +44 -0
- package/dist/run/budget.js +21 -6
- package/dist/run/cassette.js +68 -12
- package/dist/session.js +4 -0
- package/docs/cassette.md +16 -4
- package/docs/invariants.md +2 -2
- package/docs/protocol.md +23 -5
- package/docs/scenario.md +7 -4
- package/examples/replays/README.md +1 -1
- package/fixtures/protocol/v1/dialog-response.json +10 -0
- package/fixtures/protocol/v1/elicit-response.json +10 -0
- package/fixtures/protocol/v1/elicitation-request.json +17 -0
- package/fixtures/protocol/v1/error-response.json +8 -0
- package/fixtures/protocol/v1/user-dialog-request.json +11 -0
- package/package.json +1 -1
- package/schema/cassette.v12.json +2 -2
- package/schema/protocol.v1.json +447 -79
- package/scripts/check-versions.ts +134 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 2.0
|
|
7
|
-
tracks-harness: cowork-harness 2.0
|
|
6
|
+
version: 2.1.0
|
|
7
|
+
tracks-harness: cowork-harness 2.1.0 (baseline desktop-1.34493.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.0
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.1.0` (baseline
|
|
26
26
|
> `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.0
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.1.0"`. **Pin `@^2.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -656,7 +656,7 @@ than a stuck `"running"`.)
|
|
|
656
656
|
### Place assertions in the right CI lane
|
|
657
657
|
|
|
658
658
|
CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
|
|
659
|
-
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@
|
|
659
|
+
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v2` (a packaged GitHub Action with a
|
|
660
660
|
PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
|
|
661
661
|
the four-stage pipeline.
|
|
662
662
|
|
|
@@ -1,19 +1,23 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 2.0
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
7
7
|
|
|
8
8
|
```yaml
|
|
9
|
-
- uses: yaniv-golan/cowork-harness@
|
|
9
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
10
10
|
with:
|
|
11
11
|
command: replay
|
|
12
12
|
path: cassettes/
|
|
13
|
+
version: "^2" # hold the major; see below
|
|
13
14
|
```
|
|
14
15
|
|
|
15
|
-
The Action's `version` input defaults to `latest
|
|
16
|
-
|
|
16
|
+
**These recipes pin `version: "^2"`.** The Action's `version` input *defaults* to `latest`, which means a
|
|
17
|
+
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
|
+
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
|
+
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
+
(e.g. `version: "2.1.0"`) instead when you want byte-reproducible CI.
|
|
17
21
|
|
|
18
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -43,10 +47,11 @@ jobs:
|
|
|
43
47
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
44
48
|
# Background on the provenance chain: the "Agent-binary provenance" section of
|
|
45
49
|
# https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
|
|
46
|
-
- uses: yaniv-golan/cowork-harness@
|
|
50
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
47
51
|
with:
|
|
48
52
|
command: run
|
|
49
53
|
path: scenarios/
|
|
54
|
+
version: "^2"
|
|
50
55
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
51
56
|
```
|
|
52
57
|
|
|
@@ -62,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
62
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
63
68
|
|
|
64
69
|
```yaml
|
|
65
|
-
- run: npm i -g "cowork-harness@^2.0
|
|
70
|
+
- run: npm i -g "cowork-harness@^2.1.0"
|
|
66
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
67
72
|
# no silent false-greens. WITHOUT --strict this
|
|
68
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -110,20 +115,28 @@ Action has no input for, and it creates a coupling nothing checks:
|
|
|
110
115
|
|
|
111
116
|
> **If a flag in `extra-args` was added in release X, floor that step's `version` to `>=X`.**
|
|
112
117
|
|
|
113
|
-
`version` defaults to `latest`, and accepts **any npm range** — not just an exact pin.
|
|
118
|
+
`version` defaults to `latest`, and accepts **any npm range** — not just an exact pin. Leave it off unless
|
|
119
|
+
you have a reason:
|
|
114
120
|
|
|
115
121
|
```yaml
|
|
116
|
-
- uses: yaniv-golan/cowork-harness@
|
|
122
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
117
123
|
with:
|
|
118
124
|
command: lint
|
|
119
125
|
path: scenarios/
|
|
120
|
-
version: "
|
|
121
|
-
extra-args: --min-severity WARN
|
|
126
|
+
version: "^2" # holds the major
|
|
127
|
+
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 2.x satisfies that
|
|
122
128
|
```
|
|
123
129
|
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
130
|
+
**If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
|
|
131
|
+
floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
|
|
132
|
+
recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
|
|
133
|
+
major instead — `version: "^2"`, which is what the steps above use — keeping the floor's intent while
|
|
134
|
+
stopping at the major boundary. An exact
|
|
135
|
+
pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
|
|
136
|
+
moment a recipe adopts a newer flag.
|
|
137
|
+
|
|
138
|
+
Without a satisfied floor, an older CLI fails the step with `unrecognized arguments: --min-severity WARN`
|
|
139
|
+
(exit 2, wrapped in an `ok:false` envelope) — it does **not** degrade gracefully.
|
|
127
140
|
|
|
128
141
|
|
|
129
142
|
The harness has two execution lanes with different cost, coverage, AND infrastructure requirements.
|
|
@@ -307,7 +320,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
307
320
|
|
|
308
321
|
## GitHub Actions sketch
|
|
309
322
|
|
|
310
|
-
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@
|
|
323
|
+
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v2` does
|
|
311
324
|
in one step (see the top of this doc) — reach for this form when you need independent per-command
|
|
312
325
|
gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
|
|
313
326
|
equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
|
|
@@ -327,7 +340,7 @@ jobs:
|
|
|
327
340
|
with: { node-version: '24' }
|
|
328
341
|
- uses: actions/setup-python@v5
|
|
329
342
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
330
|
-
- run: npm i -g "cowork-harness@^2.0
|
|
343
|
+
- run: npm i -g "cowork-harness@^2.1.0"
|
|
331
344
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
332
345
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
333
346
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -356,7 +369,7 @@ jobs:
|
|
|
356
369
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
357
370
|
fi
|
|
358
371
|
- if: steps.guard.outputs.live == 'true'
|
|
359
|
-
run: npm i -g "cowork-harness@^2.0
|
|
372
|
+
run: npm i -g "cowork-harness@^2.1.0"
|
|
360
373
|
- if: steps.guard.outputs.live == 'true'
|
|
361
374
|
run: cowork-harness run scenarios/ --output-format json
|
|
362
375
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 2.0
|
|
3
|
+
Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.0
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.1.0`
|
|
4
4
|
(baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 2.0
|
|
5
|
+
Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -66,7 +66,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](htt
|
|
|
66
66
|
| `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
|
|
67
67
|
| `scenarioSource` | Relative path to the authored YAML this was recorded from |
|
|
68
68
|
| `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
|
|
69
|
-
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
69
|
+
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
70
70
|
| `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
|
|
71
71
|
| `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
|
|
72
72
|
| `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
|
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,261 @@ All notable changes to this project are documented here. The format is based on
|
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
|
|
5
5
|
[Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
|
|
6
6
|
|
|
7
|
+
## [2.1.0] — 2026-08-24
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **`rehash` leads with the split, and PARTIAL success has its own exit code (`4`).** The summary used to
|
|
12
|
+
print last, after every per-file line had scrolled past — and those two counts *are* the decision:
|
|
13
|
+
commit what migrated and budget a re-record for the rest, versus nothing here is salvageable. It now
|
|
14
|
+
prints first, with the per-file lines as the detail behind it.
|
|
15
|
+
|
|
16
|
+
**`4 migrated, 18 failed` and `0 migrated, 22 failed` both exited `1`**, which made a shell consumer
|
|
17
|
+
unable to tell apart two situations demanding opposite responses. The JSON envelope always carried the
|
|
18
|
+
split as `migrated`/`skipped`/`errors`; a bare terminal run did not. `rehash` now exits `0` (all
|
|
19
|
+
migrated, or nothing needed migrating) / **`4`** (partial) / `1` (nothing migrated and at least one
|
|
20
|
+
could not) / `2` (usage).
|
|
21
|
+
|
|
22
|
+
**This is a behaviour change for anyone branching on `rehash`'s exit code** — a script testing
|
|
23
|
+
`rc == 1` for "something failed" will miss the partial case, though `rc != 0` is unaffected. It ships in
|
|
24
|
+
a minor deliberately: `rehash`'s codes were documented **nowhere** — zero mentions in its `--help`,
|
|
25
|
+
absent from SPEC §11 — so there was no published contract to break, and a consumer on `rc == 1` was
|
|
26
|
+
relying on observed behaviour. They are now documented in both places. `4` rather than `3` because the
|
|
27
|
+
code space is per-command and `3`'s "could not verify" meaning is load-bearing on `verify-cassettes`.
|
|
28
|
+
|
|
29
|
+
- **The batch cost estimate now reports its basis instead of claiming authority.** `estimatedCostUsd` is
|
|
30
|
+
`sum(max(local run history))` — a max over whatever *this machine* has run. The line already qualified
|
|
31
|
+
the partially-priced case with `— LOWER BOUND`, but the fully-priced case said
|
|
32
|
+
`(all N scenario(s) priced from prior runs)`, which is an active claim of completeness and was the one
|
|
33
|
+
case that said nothing qualifying. At a new baseline or a new agent binary the history describes a
|
|
34
|
+
materially different configuration, so the estimate can **under**-predict exactly when it is most
|
|
35
|
+
consulted — a consumer wrote "that is the ceiling, not the scope" into a plan off this line and had to
|
|
36
|
+
retract it.
|
|
37
|
+
|
|
38
|
+
The line now carries `basis: N prior run(s) on THIS machine, thinnest scenario has M; a max over that
|
|
39
|
+
history, NOT a bound`, and the JSON payload gains `estimateBasis`
|
|
40
|
+
(`{source, pricedRuns, thinnestScenarioRuns}`). `thinnest` is the useful half: a scenario with one prior
|
|
41
|
+
run contributes a single sample, not a worst case. Wired through `pricedRunCount`, which had been
|
|
42
|
+
exported and doc-commented *"for messages that report their own basis"* with zero callers.
|
|
43
|
+
|
|
44
|
+
- **Pre-epoch cassettes: we could report ordinary content drift, and chose not to.** When a cassette
|
|
45
|
+
recorded before the 2.0.0 hash-format epoch is read, `rehash` recomputes the **legacy** digest over the
|
|
46
|
+
current tree and compares it to the legacy digest in the cassette. On a mismatch that is a positive
|
|
47
|
+
determination of ordinary content drift — strictly more informative than `unverifiable-skill`, and the
|
|
48
|
+
same determination 1.25.0 reported as warn-only `skill` drift.
|
|
49
|
+
|
|
50
|
+
Replay does not report it, and that is a decision rather than a limit. Reporting it would require
|
|
51
|
+
keeping a fold of the retired hash algorithm alive in the replay path — the most-run lane — permanently,
|
|
52
|
+
to soften one release's migration; the population it would help shrinks with every re-record. The
|
|
53
|
+
runtime cost would be nil (both digests fold from one tree walk); the cost is code that could never be
|
|
54
|
+
deleted. **"We can compute this and chose not to report it, so the legacy fold can die" is a different
|
|
55
|
+
claim from "this cannot be known", and the earlier phrasing implied the latter.** The remedy for a
|
|
56
|
+
pre-epoch cassette is `rehash` where it can prove content unchanged, and a re-record where it cannot.
|
|
57
|
+
|
|
58
|
+
### Fixed
|
|
59
|
+
|
|
60
|
+
- **The fast test lane had no global timeout, so 93 subprocess-spawning test files inherited vitest's 5s
|
|
61
|
+
default.** They measure 167-888ms locally — a fine margin until you remember the lane runs 344 files in
|
|
62
|
+
parallel across every core, and a CI runner is ~3x slower again. That is how an 888ms test crosses 5s;
|
|
63
|
+
it cost two red CI runs on unrelated PRs before anyone looked at the failure rather than re-running it.
|
|
64
|
+
`vitest.config.ts` now sets `testTimeout: 30_000` — ~34x the slowest measured non-e2e case, while a
|
|
65
|
+
genuine hang still fails ~6x faster than in the live lane (which sets 180s). Per-test values still win.
|
|
66
|
+
|
|
67
|
+
- **Redaction pattern ORDER is load-bearing, and now says so — plus a warning when it is wrong.** Patterns
|
|
68
|
+
apply in sequence over the accumulating output, so a bare catch-all placed ahead of a lookahead-anchored
|
|
69
|
+
rule for the same prefix matches first and eats the `/mnt/` tail the lookahead exists to preserve. The
|
|
70
|
+
shipped policy depends on that order: with it, a run-dir path redacts to
|
|
71
|
+
`[REDACTED:local-path:…]/mnt/outputs/report.md` and still resolves; without it the whole path is consumed
|
|
72
|
+
and `normalizeHostShapedForReplay` returns `null` — **every `computer://` structural-marker resolution
|
|
73
|
+
silently stops working, with no error and no finding.**
|
|
74
|
+
|
|
75
|
+
`loadRedactionPolicy` now warns, naming both pattern indices, when a policy is in the hazardous order.
|
|
76
|
+
Detection is deliberately conservative — it fires only when a later pattern's source is exactly an
|
|
77
|
+
earlier one's plus a trailing lookahead (modulo lazy quantifiers) — because regex subsumption is
|
|
78
|
+
undecidable in general and a false positive would train authors to ignore the warning.
|
|
79
|
+
|
|
80
|
+
The remainder is matched by **shape**, never by parsing the lookahead's body. A first cut used
|
|
81
|
+
`\(\?=[^()]*\)`, whose `[^()]*` silently skipped every lookahead containing a group — so
|
|
82
|
+
`(?=/mnt(?:/|$|[\s"'\\)\]]))`, the natural way to write "slash, end, or delimiter" and arguably more
|
|
83
|
+
correct than a bare `(?=/mnt/)`, went unflagged while being just as dangerous. Caught by a consumer
|
|
84
|
+
running it against their own policy, which is now a regression fixture. Not looking inside also sidesteps
|
|
85
|
+
escape- and char-class-awareness, since that policy carries an escaped `\)` inside a character class.
|
|
86
|
+
|
|
87
|
+
`docs/cassette.md` states the requirement next to the existing "stop before `/mnt/`" guidance, which had
|
|
88
|
+
the shape of the rule but not the ordering half. The new test pins the runtime consequence, not just the
|
|
89
|
+
detector: a reorder must make the link fail to normalize AND be flagged, so the syntactic check cannot
|
|
90
|
+
drift away from what it is standing in for.
|
|
91
|
+
|
|
92
|
+
- **The copy-pasteable Action steps now pin `version: "^2"`, and a guard requires it.** Bounding every
|
|
93
|
+
published npm floor last release fixed the *form* of a floor (`>=1.11.0` reads as a bound and silently
|
|
94
|
+
means "and every future major too") but removed the input from the recipes rather than correcting it —
|
|
95
|
+
so the shipped snippets carried no `version:` at all and fell back to the input's `latest` default. That
|
|
96
|
+
reproduced the exact footgun `action.yml`'s own description warns about two lines earlier: a CLI major
|
|
97
|
+
reaches a workflow the moment it is promoted, even though the `uses:` ref never changed.
|
|
98
|
+
|
|
99
|
+
`^2` was not among the alternatives weighed at the time, and it is the form `action.yml` itself
|
|
100
|
+
recommends: it holds the major, needs no patch number to remember, and only wants a human decision at
|
|
101
|
+
the next major bump. **Five** steps were unpinned, not the three in the CI recipe — `README.md` carries
|
|
102
|
+
two more.
|
|
103
|
+
|
|
104
|
+
`action-docs-sync` now requires every copy-pasteable step (a `uses:` line inside a fenced block with a
|
|
105
|
+
`with:`) to pin `^<package major>`; inline prose mentions are excluded, since there is nothing to pin.
|
|
106
|
+
Verified by mutation: dropping one `version:`, regressing a pin to `^1`, and bumping the package major
|
|
107
|
+
each fail. The guard also asserts it found the steps at all, because a parser that matches nothing
|
|
108
|
+
passes every assertion after it. Guarding a floor's FORM does not guarantee a floor is PRESENT — this
|
|
109
|
+
pins the behaviour instead.
|
|
110
|
+
|
|
111
|
+
- **The `sessionFingerprint` field set is now stated completely, and a guard discovers the sites that
|
|
112
|
+
state it.** The hash covers eight session fields; every place that enumerated them named six or fewer.
|
|
113
|
+
`web_fetch` was missing everywhere, `agent_env` was missing everywhere, and `docs/invariants.md`
|
|
114
|
+
also omitted `skills`. Fourteen sites carried the claim while the working assumption was three: **twelve
|
|
115
|
+
enumerated the set** — four in prose, two in the current cassette schema, six in the retained v9-v11
|
|
116
|
+
schemas — and **two denied it existed at all**. `SPEC.md` and `docs/scenario.md` both said the session
|
|
117
|
+
is "not drift-checked or fingerprinted". `model` genuinely is not hashed; connected folders and plugin/skill/MCP discovery are,
|
|
118
|
+
which is the half a reader would have trusted.
|
|
119
|
+
|
|
120
|
+
The `verify-cassettes` staleness message — the only enumeration a user ever sees — omitted `projects`,
|
|
121
|
+
and `docs/cassette.md` quoted it with `projects` present. Both are fixed and now pinned to each other
|
|
122
|
+
by test, so the doc cannot drift from the string again.
|
|
123
|
+
|
|
124
|
+
New **invariant 14** in `check:versions` derives the field set from `buildSessionFingerprint`'s shape
|
|
125
|
+
and **discovers** the enumeration sites rather than reading a list, so a new one is covered the day it
|
|
126
|
+
lands. Two deliberate limitations are recorded as tests rather than left to look covered: it cannot see
|
|
127
|
+
a flat denial (there is no enumeration to check — the coverage floor is what notices), and it cannot
|
|
128
|
+
see a deleted "only when set" qualifier. `schema/cassette.v{9,10,11}.json` are allowlisted as frozen
|
|
129
|
+
history: a retained schema documents the format as it shipped, and rewriting its description would make
|
|
130
|
+
it describe a shape its own consumers never saw.
|
|
131
|
+
|
|
132
|
+
Verified by mutation: run against the previous revision the guard flags all six then-existing sites
|
|
133
|
+
with the correct missing fields; renaming the shape literal makes it error rather than silently pass;
|
|
134
|
+
a ninth field invalidates a previously-complete site; and a whole-file token check — which would have
|
|
135
|
+
passed today and then never failed again — is rejected in favour of span-scoped matching.
|
|
136
|
+
|
|
137
|
+
- **The documented Action ref is now `@v2`, not `@main`** — 7 references across `README.md`,
|
|
138
|
+
`SKILL.md` and `ci-recipe.md`. `@main` was right when it was written: no alias tag had ever been
|
|
139
|
+
published, so naming one would have sent a copy-pasting reader to a `uses:` that 404s, and the guard's
|
|
140
|
+
own note said to revisit "once 1.0.0 ships". Two things had to be true first, and now are — `v2`/`v2.0`
|
|
141
|
+
point at a real release, and `release.yml` moves them on every stable release rather than leaving it to a
|
|
142
|
+
checklist. Recommending a floating tag nobody remembers to move is worse than recommending `@main`; that
|
|
143
|
+
was the actual situation while `v1` sat at 1.24.0.
|
|
144
|
+
|
|
145
|
+
`action-docs-sync` now derives the expected ref from `package.json`'s major instead of hardcoding it, so
|
|
146
|
+
the next major forces these docs to move with it rather than silently pointing a reader at the previous
|
|
147
|
+
line. `@main` is deliberately no longer accepted there: permitting both would let the recommendation
|
|
148
|
+
drift back with nothing noticing. Verified by mutation — regressing one reference to `@main` fails, and
|
|
149
|
+
setting the package version to 3.0.0 fails all three files.
|
|
150
|
+
|
|
151
|
+
**This changes nothing about which CLI you get.** The ref selects the Action; the CLI still comes from the
|
|
152
|
+
`version:` input, which still defaults to `latest`. `@v2` looks more like a version pin than `@main` did,
|
|
153
|
+
so that distinction matters more now, not less — it is spelled out in `README.md`'s Action section and in
|
|
154
|
+
`action.yml`'s own input description.
|
|
155
|
+
|
|
156
|
+
- **The CI recipe no longer teaches a bare version floor.** It
|
|
157
|
+
carried `version: ">=1.11.0"` — which reads as "at least 1.11.0" and silently means "and every future
|
|
158
|
+
major too", so a copy-paster gets the next major with no say in it. It was **not** broken today
|
|
159
|
+
(`--min-severity` still exists in 2.x, and `lint` reads no cassette, so 2.0.0's hash-format epoch never
|
|
160
|
+
applied to that step) — the defect was latent and in the FORM.
|
|
161
|
+
|
|
162
|
+
The bare floor was first dropped rather than corrected — `^1.11.0` would have frozen every new
|
|
163
|
+
copy-paster on the previous major, and `^2.0.1` needs remembering at each release — so the guidance moved
|
|
164
|
+
to prose. **That went one step too far, and the same release corrects it** (see the `version: "^2"` entry
|
|
165
|
+
above): dropping the input entirely falls back to `action.yml`'s `latest` default, which is the one
|
|
166
|
+
remaining unbounded form and the exact footgun the input's own description warns about two lines earlier.
|
|
167
|
+
`^2` was never among the alternatives weighed at the time, and it is what the recipes now carry. Reach for
|
|
168
|
+
an exact pin only when you want byte-reproducible CI. [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml)'s own description stops offering `>=1.11.0` and
|
|
169
|
+
`^1.11.0` as interchangeable — they are not, and it had been recommending the unbounded one.
|
|
170
|
+
|
|
171
|
+
`check:versions` invariant 13 now covers the Action's `version:` input, which it could not see before:
|
|
172
|
+
it keys on `@>=`, and `version: ">=1.11.0"` has no `@` — the same defect in different syntax, with no
|
|
173
|
+
coverage. Only the **unbounded** floor is rejected; verified against each form, `>=1.11.0 <3`, `^2`,
|
|
174
|
+
`2.0.1` and `latest` all pass.
|
|
175
|
+
|
|
176
|
+
- **`projects[].from` was missing from two places, and the second was a false green.** A connected project
|
|
177
|
+
is a host path exactly like a connected folder, and it was:
|
|
178
|
+
- **not resolved against the session file.** [`docs/session.md`](./docs/session.md) promises, without
|
|
179
|
+
qualification, that *"relative paths resolve from the session file's own directory"* — and it was true
|
|
180
|
+
of every path field except this one, which resolved against the **process CWD**. So the same session
|
|
181
|
+
file mounted different content depending on which directory you invoked from. The doc was right; the
|
|
182
|
+
resolver had simply skipped the field.
|
|
183
|
+
- **not part of the session fingerprint.** Swapping which directory is mounted at `.projects/<uuid>`
|
|
184
|
+
changed the run's inputs and `verify-cassettes` reported nothing — a false green in the gate whose job
|
|
185
|
+
is to notice that inputs moved. Folded in on the same **non-empty-only** terms as `agent_env`, so a
|
|
186
|
+
session with no `projects:` (and one with an explicit `projects: []`) hashes byte-identically to
|
|
187
|
+
before; only sessions that use the feature move. No committed cassette does.
|
|
188
|
+
|
|
189
|
+
**A cassette recorded before this is reported `unverifiable`, not clean.** Its hash contains nothing
|
|
190
|
+
about `projects[]`, so it cannot distinguish "the field was never covered" from "the mount changed since
|
|
191
|
+
record time" — and reporting the mismatch as a benign migration would put the same false green back in
|
|
192
|
+
the remedy. When everything else matches exactly, `verify-cassettes` says so and asks for a re-record to
|
|
193
|
+
gain the coverage. `sessionFingerprintDrift` remains `verify-cassettes`-only: none of this can change a
|
|
194
|
+
`replay` verdict, even under `--strict`.
|
|
195
|
+
|
|
196
|
+
Three enumerations of the covered fields gained `projects` — [`docs/cassette.md`](./docs/cassette.md),
|
|
197
|
+
[`docs/invariants.md`](./docs/invariants.md) and the shipped skill's `task-recipes.md`. `invariants.md`
|
|
198
|
+
also described this as a hash of the **resolved** session; it is the **authored, pre-resolution** shape,
|
|
199
|
+
deliberately, so the digest survives a different checkout — the function's own comment says so, and a
|
|
200
|
+
resolved hash could never match on another clone.
|
|
201
|
+
|
|
202
|
+
- **The published control-protocol schema rejected five kinds of frame the harness sends and answers.**
|
|
203
|
+
`schema/protocol.v1.json` described four request subtypes; the harness has always answered **six** —
|
|
204
|
+
adding `request_user_dialog` and `elicitation`/`side_question` — and it also sends a fail-closed
|
|
205
|
+
`subtype:"error"` response envelope, whose payload is a *string* under `error` rather than an object
|
|
206
|
+
under `response`, so a validator that knew only the success envelope rejected every one. Measured
|
|
207
|
+
against the previous schema: all five representative frames **REJECT**; against the new one, all five
|
|
208
|
+
accept. Anyone validating real traffic was seeing failures on frames the harness handles correctly.
|
|
209
|
+
|
|
210
|
+
Added as new `oneOf`/`anyOf` branches plus one new top-level response shape — **111 insertions, zero
|
|
211
|
+
deletions** in the surface baseline, so nothing existing was narrowed and no frame the schema already
|
|
212
|
+
accepted is affected. Both spellings the parser accepts are admitted (`dialogKind`/`dialog_kind`,
|
|
213
|
+
`mcp_server_name`/`server`, `message`/`prompt`) rather than guessing which the agent sends.
|
|
214
|
+
|
|
215
|
+
The golden vector pack grew with it, generated from the **real** envelope builders rather than
|
|
216
|
+
hand-authored lookalikes. The existing lockstep test — every schema definition must be exercised by a
|
|
217
|
+
vector — caught all five additions immediately, which is what forced those vectors to exist. One of them
|
|
218
|
+
asserts the error envelope does **not** validate as a success envelope, so the two shapes cannot be
|
|
219
|
+
quietly conflated later.
|
|
220
|
+
|
|
221
|
+
[`SPEC.md`](./SPEC.md) §12 now states the additive latitude for this surface explicitly. Its silence had
|
|
222
|
+
read as a prohibition, which is plausibly why three subtypes went undescribed rather than added — while
|
|
223
|
+
[`docs/protocol.md`](./docs/protocol.md)'s own versioning policy had said all along that an additive
|
|
224
|
+
variant is a v1 minor note, which is where the dated entry now lives.
|
|
225
|
+
|
|
226
|
+
- **The Marketplace alias tags are moved by the release workflow instead of by remembering.** `v1` sat at
|
|
227
|
+
1.24.0 through two releases because moving it was a checklist line. `release.yml`'s last step now points
|
|
228
|
+
`vX` and `vX.Y` at the release it just published, with two guards a hand-run `git tag -f` skips: a
|
|
229
|
+
**prerelease** tag moves nothing (the trigger accepts `v1.0.4-rc.1`, and pointing `v1` at an rc would
|
|
230
|
+
hand every `@v1` consumer a prerelease), and an alias **never moves backwards** — re-releasing an older
|
|
231
|
+
patch on a line moves `vX.Y` and leaves `vX` alone. Verified by executing the logic against a synthetic
|
|
232
|
+
tag set rather than by reading it: releasing `v1.20.5` while `v1.25.0` exists correctly skips `v1` and
|
|
233
|
+
still moves `v1.20`. It runs last, after publish and the GitHub Release, so a failure there cannot
|
|
234
|
+
half-publish anything.
|
|
235
|
+
|
|
236
|
+
Alongside it, the tags are now correct: **`v2` and `v2.0` created** (they did not exist, so nothing
|
|
237
|
+
pointed at the 2.x Action), and **`v1` moved 1.24.0 → 1.25.0**. Worth recording what that move did and
|
|
238
|
+
did not fix: the Action's whole surface — `action.yml` plus the `render.js` it loads — is **byte-identical
|
|
239
|
+
from 1.24.0 through 2.0.1** apart from three lines of input *description*. So a stale `v1` was a promise
|
|
240
|
+
the repo had stopped keeping, not a functional gap, and the one public `@v1` consumer pins `version:` on
|
|
241
|
+
every step and was never exposed to the `latest` default at all.
|
|
242
|
+
|
|
243
|
+
### Documentation
|
|
244
|
+
|
|
245
|
+
- **The invariants index said `check-versions.ts` has "no dedicated vitest file — it's a standalone script,
|
|
246
|
+
not a unit-testable module boundary".** Three exist, for the invariants whose logic is an exported pure
|
|
247
|
+
function: `check-cassette-version-claims`, `check-fingerprint-field-claims` and
|
|
248
|
+
`check-design-scope-note`. The claim had already been stale before this release.
|
|
249
|
+
|
|
250
|
+
- **The `uses:` ref pins the Action; the `version:` input pins the CLI — and only the second one holds a
|
|
251
|
+
major.** Both are documented as if pinning `@v1` bounded what you install. It does not: they move
|
|
252
|
+
independently, and `version:` defaults to `latest`, so promoting a CLI major reaches a workflow whose
|
|
253
|
+
`uses:` ref has not changed in months. Measured at the time of writing — `v1` points at **1.24.0** and
|
|
254
|
+
has never been moved, yet an `@v1` workflow with no `version:` input installs **2.x**. `README.md`,
|
|
255
|
+
`action.yml` (the text GitHub Marketplace renders) and `RELEASING.md`'s alias-tag step now say so, and
|
|
256
|
+
name the fix: pin the **input** (`version: ^2`), not the ref. Crossing 1.x → 2.x this way means the
|
|
257
|
+
hash-format epoch, so pre-v12 cassettes need `cowork-harness rehash <dir/>`.
|
|
258
|
+
|
|
259
|
+
`RELEASING.md`'s "move the major/minor tags" step additionally records what moving `vX` does *not* do,
|
|
260
|
+
since that step reads as the thing that controls consumer upgrades and is not.
|
|
261
|
+
|
|
7
262
|
## [2.0.1] — 2026-08-23
|
|
8
263
|
|
|
9
264
|
### Added
|
package/README.md
CHANGED
|
@@ -115,7 +115,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
115
115
|
|
|
116
116
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
117
117
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
118
|
-
> From a global install (`npm i -g "cowork-harness@^2.0
|
|
118
|
+
> From a global install (`npm i -g "cowork-harness@^2.1.0"`), point at the package root instead:
|
|
119
119
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
120
120
|
> (or copy the cassette into your own project and pass that path).
|
|
121
121
|
|
|
@@ -125,7 +125,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
125
125
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
126
126
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
127
127
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
128
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.0
|
|
128
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.1.0"`.
|
|
129
129
|
|
|
130
130
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
131
131
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -150,7 +150,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
150
150
|
claude plugin install cowork-harness@cowork-harness
|
|
151
151
|
```
|
|
152
152
|
|
|
153
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.0
|
|
153
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.1.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
154
154
|
|
|
155
155
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
156
156
|
|
|
@@ -175,7 +175,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
|
|
|
175
175
|
ones that still need a source checkout. (The marketplace
|
|
176
176
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
177
177
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
178
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.0
|
|
178
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.1.0"` — see
|
|
179
179
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
180
180
|
|
|
181
181
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -358,7 +358,11 @@ Exceptions worth knowing:
|
|
|
358
358
|
|
|
359
359
|
- the `boundary-check` **command**'s own probe failures follow the assertion convention and exit `1` (see SPEC.md §11);
|
|
360
360
|
- `verify-cassettes` reuses exit `3` with a different meaning — "could not verify" (vs. a verified failure's `1`); see the command table;
|
|
361
|
-
- exit `4`
|
|
361
|
+
- `rehash` uses **exit `4` for PARTIAL success** — some cassettes migrated, some could not. That and a total
|
|
362
|
+
failure (`1`, nothing migrated) demand opposite responses, so they are separate codes; `0` means all
|
|
363
|
+
migrated or nothing needed migrating (SPEC §11);
|
|
364
|
+
- exit `4` is otherwise reserved on the `run`/`skill` family, where it is still unused (SPEC §11) — the code
|
|
365
|
+
space is per-command, so `rehash`'s `4` above does not touch that reservation.
|
|
362
366
|
|
|
363
367
|
After a run, the footer **echoes every auto-answered
|
|
364
368
|
question as a copy-pasteable `--answer "<q>=<choice>"` line** — run once exploratorily, then paste them
|
|
@@ -719,11 +723,23 @@ cowork-harness run scenarios/ # your repo's scenarios; runs every *.y
|
|
|
719
723
|
|
|
720
724
|
The fastest path to CI: a composite action wrapping the token-free lane, with a PR job-summary reporter.
|
|
721
725
|
|
|
726
|
+
> **The `uses:` ref pins the Action, not the CLI.** `@main`, `@v1` and a commit SHA all select which
|
|
727
|
+
> *Action* runs; which *CLI* it installs is the separate `version:` input below, which defaults to
|
|
728
|
+
> `latest`. The two move independently — so a workflow whose `uses:` ref has not changed in months still
|
|
729
|
+
> picks up a CLI **major** the moment one is promoted to `latest`. That is not hypothetical: `@v1` points
|
|
730
|
+
> at 1.24.0 and has not moved, yet an `@v1` workflow without a `version:` input installs 2.x today.
|
|
731
|
+
> **To hold a major, pin the input** — `version: ^2` (or `^1`) — not the `uses:` ref.
|
|
732
|
+
>
|
|
733
|
+
> Upgrading across the 1.x → 2.x boundary this way means the hash-format epoch: cassettes recorded before
|
|
734
|
+
> format v12 need `cowork-harness rehash <dir/>` (no re-record). See the 2.0.0 entry in
|
|
735
|
+
> [CHANGELOG.md](./CHANGELOG.md).
|
|
736
|
+
|
|
722
737
|
```yaml
|
|
723
|
-
- uses: yaniv-golan/cowork-harness@
|
|
738
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
724
739
|
with:
|
|
725
740
|
command: replay # replay | lint | lint-skill | analyze-skill | verify-cassettes | run
|
|
726
741
|
path: cassettes/my-skill.cassette.json
|
|
742
|
+
version: "^2" # hold the CLI major; the input defaults to `latest`
|
|
727
743
|
```
|
|
728
744
|
|
|
729
745
|
| Lane | Commands | Runner requirements | What you get |
|
|
@@ -748,14 +764,15 @@ jobs:
|
|
|
748
764
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
749
765
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
750
766
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
751
|
-
- uses: yaniv-golan/cowork-harness@
|
|
767
|
+
- uses: yaniv-golan/cowork-harness@v2
|
|
752
768
|
with:
|
|
753
769
|
command: run
|
|
754
770
|
path: scenarios/
|
|
771
|
+
version: "^2"
|
|
755
772
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
756
773
|
```
|
|
757
774
|
|
|
758
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` —
|
|
775
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.1.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
759
776
|
|
|
760
777
|
The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
761
778
|
|
package/RELEASING.md
CHANGED
|
@@ -223,8 +223,13 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
|
|
|
223
223
|
gh run watch $(gh run list --workflow=release.yml --limit 1 --json databaseId --jq '.[0].databaseId')
|
|
224
224
|
```
|
|
225
225
|
- [ ] **Clean up**: `git push origin --delete release/X.Y.Z && git branch -d release/X.Y.Z`
|
|
226
|
-
- [
|
|
227
|
-
|
|
226
|
+
- [x] **Move the major/minor tags** — **AUTOMATED** by `release.yml`'s last step, which points `vX` and
|
|
227
|
+
`vX.Y` at the release it just published. It skips a prerelease tag entirely, and never moves an alias
|
|
228
|
+
backwards (re-releasing an older patch on a line moves `vX.Y` but leaves `vX` where it is). Left as a
|
|
229
|
+
checklist item, it lapsed: `v1` sat at 1.24.0 through two releases. Note what these tags do NOT do:
|
|
230
|
+
they select the ACTION, never the CLI. A consumer who leaves the `version:` input at its `latest`
|
|
231
|
+
default already tracks CLI majors regardless of which alias they pin, so moving `vX` neither causes
|
|
232
|
+
nor prevents a cross-major CLI upgrade for them. Manual fallback, if the step ever needs redoing:
|
|
228
233
|
```
|
|
229
234
|
git tag -f vX vX.Y.Z && git tag -f vX.Y vX.Y.Z # e.g. v1 and v1.0 → v1.2.3
|
|
230
235
|
git push -f origin vX vX.Y
|