cowork-harness 4.4.0 → 4.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +4 -4
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +2 -2
- package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
- package/.claude/skills/cowork-harness/references/authoring.md +5 -1
- package/.claude/skills/cowork-harness/references/ci-recipe.md +20 -6
- package/.claude/skills/cowork-harness/references/critique.md +4 -2
- package/.claude/skills/cowork-harness/references/debugging.md +24 -4
- package/.claude/skills/cowork-harness/references/eval.md +4 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/gotchas.md +1 -1
- package/.claude/skills/cowork-harness/references/hillclimb-recipe.md +37 -6
- package/.claude/skills/cowork-harness/references/hillclimb.md +17 -8
- package/.claude/skills/cowork-harness/references/measurement.md +2 -2
- package/.claude/skills/cowork-harness/references/run-record-replay.md +36 -5
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/semantic-judging.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +36 -0
- package/DESIGN.md +1 -1
- package/README.md +3 -3
- package/RELEASING.md +17 -2
- package/docs/ci.md +1 -1
- package/docs/cli.md +5 -5
- package/docs/companion-skill.md +2 -2
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did it make answers worse? (`eval`: paired A/B, pinned models) — or improving one round by round (`hillclimb`, the `/claude-api hillclimb` runner). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval / hillclimb commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 4.4.
|
|
7
|
-
tracks-harness: cowork-harness 4.4.
|
|
6
|
+
version: 4.4.1
|
|
7
|
+
tracks-harness: cowork-harness 4.4.1 (baseline desktop-2.19675.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
|
|
|
26
26
|
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
27
|
Read them.
|
|
28
28
|
|
|
29
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.4.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.4.1` (baseline
|
|
30
30
|
> `desktop-2.19675.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
31
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
32
32
|
|
|
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
43
43
|
|
|
44
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
45
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
46
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.4.
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.4.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.4.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.4.1"`. **Pin `@^4.4.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
47
47
|
|
|
48
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
49
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertion catalog
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). Every `assert:` key with its semantics, and the
|
|
4
4
|
verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
|
|
5
5
|
the scenario and session YAML fields are there too.
|
|
6
6
|
|
|
@@ -110,7 +110,7 @@ same set live from the schema.
|
|
|
110
110
|
| `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
|
|
111
111
|
| `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
|
|
112
112
|
| `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer: the final result text, the transcript and the files the agent authored. The transcript is top-level `assistant_text` only — it excludes every `tool_use`/`tool_result` and sub-agent text unless opted in. Full semantics (evidence scope, fork results, refusals, judge provenance, cost): [semantic-judging.md](semantic-judging.md). |
|
|
113
|
-
| `semantic_pairwise: {refs?: [..], rubric?: [..], pass_if?, order?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge compares this run's judged document with a **frozen reference** — the same document an earlier run (usually the baseline) produced, written once by `ref freeze` and never regenerated — and answers win, tie, loss or `both_bad`. **The judged document is the one `semantic_matches` builds** (final message + transcript + authored files, same `evidence_files` / `include_subagent_text` / `include_fork_results` options), so the transcript **excludes every `tool_use`/`tool_result`** and a criterion about whether a tool was called cannot be judged from it. `refs` are reference stores relative to the scenario file; the run is judged against every one (under `hillclimb run`, against the flow's own references instead — only the baseline's gates `pass_if`, a later variant's is a metric recorded with `gate: false`), and `pass_if` (default `not_worse`) must hold against each gating one: `win`, `not_worse` (win or tie), or `any` (graded at all — a metric only). `both_bad` fails `win` and `not_worse`. Which output the judge sees first is a seeded coin per run, assert and reference (`order: both` judges both orders: a win in one and a loss in the other is position bias and scores as a tie; any other disagreement keeps the worse outcome; each order's own outcome is recorded as `pairwise[].orders.candidate_first` / `.ref_first`, absent for a single-order grade, and a judge that favours whichever output it sees first shows across runs as `candidate_first` winning more often than `ref_first`). The judge never sees the words "reference" or "baseline". A reference must have been frozen with the same evidence options and for the same prompt; a missing, damaged, differently-scoped or other-prompt one **refuses the run before it spends anything**, as does a reference store inside any mounted source (the agent could read its own answer key). Unavailable evidence refuses with `semantic_matches`' typed reasons. Per-reference outcomes land in `assertions[].pairwise`. **LIVE-ONLY** — skipped on replay. |
|
|
113
|
+
| `semantic_pairwise: {refs?: [..], rubric?: [..], pass_if?, order?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge compares this run's judged document with a **frozen reference** — the same document an earlier run (usually the baseline) produced, written once by `ref freeze` and never regenerated — and answers win, tie, loss or `both_bad`. **The judged document is the one `semantic_matches` builds** (final message + transcript + authored files, same `evidence_files` / `include_subagent_text` / `include_fork_results` options), so the transcript **excludes every `tool_use`/`tool_result`** and a criterion about whether a tool was called cannot be judged from it. `refs` are reference stores relative to the scenario file; the run is judged against every one (under `hillclimb run`, against the flow's own references instead — only the baseline's gates `pass_if`, a later variant's is a metric recorded with `gate: false`), and `pass_if` (default `not_worse`) must hold against each gating one: `win`, `not_worse` (win or tie), or `any` (graded at all — a metric only). `both_bad` fails `win` and `not_worse`. Which output the judge sees first is a seeded coin per run, assert and reference (`order: both` judges both orders: a win in one and a loss in the other is position bias and scores as a tie; any other disagreement keeps the worse outcome; each order's own outcome is recorded as `pairwise[].orders.candidate_first` / `.ref_first`, absent for a single-order grade, and a judge that favours whichever output it sees first shows across runs as `candidate_first` winning more often than `ref_first`). The judge never sees the words "reference" or "baseline", and the stored reference is scrubbed with the grading process's secret set before it reads it (`pairwise[].refRedactions`). A reference must have been frozen with the same evidence options and for the same prompt; a missing, damaged, differently-scoped or other-prompt one **refuses the run before it spends anything**, as does a reference store inside any mounted source (the agent could read its own answer key). Unavailable evidence refuses with `semantic_matches`' typed reasons. Per-reference outcomes land in `assertions[].pairwise`. **LIVE-ONLY** — skipped on replay. |
|
|
114
114
|
| `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
|
|
115
115
|
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
116
116
|
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertions guide
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
|
|
4
4
|
|
|
5
5
|
### Assertions: two orthogonal axes
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Authoring a scenario
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
|
|
4
4
|
|
|
5
5
|
## Part I — AUTHOR a scenario
|
|
6
6
|
|
|
@@ -213,6 +213,10 @@ cowork-harness scaffold --name report-check --skill ./skills/report-gen \
|
|
|
213
213
|
--egress-allowed api.weather.example.com --out scenarios/report-check.yaml
|
|
214
214
|
```
|
|
215
215
|
|
|
216
|
+
Each repeatable flag adds one item: `--content` a `transcript_matches`, `--tool` a `tool_called`, `--subagent` a
|
|
217
|
+
`subagent_dispatched`, `--file` a `file_exists`, `--artifact` a `user_visible_artifact`, `--egress-allowed` /
|
|
218
|
+
`--egress-denied` an `egress_allowed` / `egress_denied`, `--gate REGEX=CHOICE` a scripted answer and `--web-fetch`
|
|
219
|
+
an approval rule; `--no-delete` adds `no_delete_in_outputs: true`, and `--no-validate` skips the self-lint.
|
|
216
220
|
The flag-built form runs the bundled `scripts/scenario.py scaffold`, which also runs directly with the
|
|
217
221
|
same flags (installed as a plugin, `${CLAUDE_PLUGIN_ROOT}/scripts/scenario.py`).
|
|
218
222
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "4.4.
|
|
20
|
+
(e.g. `version: "4.4.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
82
82
|
GitHub-hosted runners, no token/Docker/agent:
|
|
83
83
|
|
|
84
84
|
```yaml
|
|
85
|
-
- run: npm i -g "cowork-harness@^4.4.
|
|
85
|
+
- run: npm i -g "cowork-harness@^4.4.1"
|
|
86
86
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
87
87
|
# no silent false-greens. WITHOUT --strict this
|
|
88
88
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -131,6 +131,16 @@ exits 0, so dropping the
|
|
|
131
131
|
`verify-cassettes` step means a skill edit silently stops being tested. (One command instead of two:
|
|
132
132
|
`replay --fail-on-skill-drift`.)
|
|
133
133
|
|
|
134
|
+
`lint --json` and `lint-skill --json` print their findings as JSON for a script to read. `analyze-skill --strict
|
|
135
|
+
<plugin-dir>/` is the static pre-flight for host-loop path fidelity and lost artifact write-backs (exit 3 when it
|
|
136
|
+
could not parse a candidate); `--runtime` adds a headless-DOM confirmation (needs `jsdom`) that never changes
|
|
137
|
+
the exit code.
|
|
138
|
+
|
|
139
|
+
Other `verify-cassettes` flags: `--skip-privacy` or `--skip-staleness` runs only one half of the gate;
|
|
140
|
+
`--skip-scenario-drift` drops the scenario-prompt drift check; `--allow-empty` lets an existing directory with no
|
|
141
|
+
cassettes exit 0 (a missing path still fails); `--margins` prints each count-bound assert's recorded value against
|
|
142
|
+
its budget (diagnostic, never the verdict).
|
|
143
|
+
|
|
134
144
|
The rest of this doc explains the lane split, recording, privacy, and the full pipeline + live job.
|
|
135
145
|
|
|
136
146
|
## The core split: token-free PR gate + live nightly (self-hosted)
|
|
@@ -398,7 +408,7 @@ jobs:
|
|
|
398
408
|
with: { node-version: '24' }
|
|
399
409
|
- uses: actions/setup-python@v5
|
|
400
410
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
401
|
-
- run: npm i -g "cowork-harness@^4.4.
|
|
411
|
+
- run: npm i -g "cowork-harness@^4.4.1"
|
|
402
412
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
403
413
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
404
414
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -427,7 +437,7 @@ jobs:
|
|
|
427
437
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
428
438
|
fi
|
|
429
439
|
- if: steps.guard.outputs.live == 'true'
|
|
430
|
-
run: npm i -g "cowork-harness@^4.4.
|
|
440
|
+
run: npm i -g "cowork-harness@^4.4.1"
|
|
431
441
|
- if: steps.guard.outputs.live == 'true'
|
|
432
442
|
run: cowork-harness run scenarios/ --output-format json
|
|
433
443
|
env:
|
|
@@ -520,7 +530,11 @@ set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path
|
|
|
520
530
|
artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl` and
|
|
521
531
|
`egress.log` at the root, plus each turn's `run.jsonl` / `trace.json` / `result.json` under `turns/<N>/`
|
|
522
532
|
(a single-turn run has just `turns/1/`; there is no root compat copy of any of these). Digest one with `cowork-harness trace <run-id | dir>`.
|
|
523
|
-
Secrets are scrubbed from every persisted log by value.
|
|
533
|
+
Secrets are scrubbed from every persisted log by value. The first live run under a runs root also creates its
|
|
534
|
+
scrub-set key, `scrubset.key`, beside the runs root (in the working directory for `--run-dir runs`): add it to
|
|
535
|
+
`.gitignore`, and never upload or commit it (an upload of `runs/` does not include it). A fresh runner creates a new
|
|
536
|
+
key, so `regrade` of a CI run on another machine cannot prove its scrub set covered: new or edited rubric text is
|
|
537
|
+
refused there; re-run the case instead.
|
|
524
538
|
|
|
525
539
|
## Don't assume a fixed assertion count across lanes
|
|
526
540
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -35,7 +35,9 @@ Two things a harvester needs: roll-ups are excluded from `stats` aggregation (th
|
|
|
35
35
|
counting them adds a phantom run and drags `passRate` toward 1), so **filter them out of any pass-rate
|
|
36
36
|
computed over raw rows** — that exclusion also governs `stats --group-by skill-hash`, so a per-generation
|
|
37
37
|
**total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
|
|
38
|
-
UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
|
|
38
|
+
UNDERCOUNT. The index is the only cost record that survives run-dir pruning. `--evaluator-model <id>` changes
|
|
39
|
+
only the evaluator passes' model (and the armor's injection-resistance check covers the default evaluator only);
|
|
40
|
+
when the task turn dominates cost, the levers are `--model`, `--timeout` and the probe's scope.
|
|
39
41
|
|
|
40
42
|
## An exit-2 report — which turn failed, and whether it was really infrastructure
|
|
41
43
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Debugging a run
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
|
|
4
4
|
|
|
5
5
|
## Part III — Debug
|
|
6
6
|
|
|
@@ -40,10 +40,30 @@ rubric changed (or you want another judge model) on a run you already paid for,
|
|
|
40
40
|
kept run without re-running the agent — unlike the tools above it is not token-free (the judge call is its
|
|
41
41
|
spend) — writes the grade beside the run, and says whether the judge read the same document the live judge did.
|
|
42
42
|
Content the live judge never read (a widened `evidence_files` / `include_subagent_text` / `include_fork_results` scope, or a larger
|
|
43
|
-
`--authored-total-bytes`) is refused unless you pass `--allow-unchecked
|
|
43
|
+
`--authored-total-bytes`) is refused unless you pass `--allow-unchecked` (`error.code` `unchecked_content`); a
|
|
44
|
+
judged document that differs from the one the live judge read is refused as `doc_drift` unless you pass
|
|
45
|
+
`--allow-doc-drift`, and a scenario with no judged assert is `no_semantic_asserts`. A `task_unverifiable` /
|
|
44
46
|
`rubric_unverifiable` / `evidence_unverifiable` / `reference_unverifiable` refusal means a part of the judge's input
|
|
45
|
-
cannot be proven scrubbed with the run's scrub set (typically
|
|
46
|
-
|
|
47
|
+
cannot be proven scrubbed with the run's scrub set (typically an edited rubric on a run recorded before the scrub-set
|
|
48
|
+
fingerprint); `refusals[].scrubSet` says why (`legacy`, `unrecorded`, `mangled`, `key`, `smaller`). Re-run the
|
|
49
|
+
case, regrade with the run's `COWORK_HARNESS_SCRUB_VALUES` / `_KEYS`, or pass `--allow-scrub-change` after checking
|
|
50
|
+
them. Content accepted with `--allow-unchecked`, and the document of an `unknown` or `live_refused` assert, are
|
|
51
|
+
outside that proof (scrubbed with this process's set only); `regrade` names both in a `::warning::` before judging.
|
|
52
|
+
|
|
53
|
+
**`::warning:: [scrub-set] no scrub-set fingerprint is recorded: the installation key …`** means the run could not
|
|
54
|
+
use its installation key, `scrubset.key` beside the runs root (`~/.cowork-harness/scrubset.key` for the default
|
|
55
|
+
root; the current directory for `--run-dir runs`). The warning names the path and the defect: not a regular file
|
|
56
|
+
(a symlink or a directory), readable or writable by others (it must be mode 600), owned by another user, empty (an
|
|
57
|
+
interrupted create), unreadable, not a 64-hex-digit key, or could not be created. Such a run records
|
|
58
|
+
`scrubSetUnavailable` instead of a `scrubSet`, so a later `regrade` of new or edited judge input refuses with
|
|
59
|
+
`scrubSet: "unrecorded"` ("this run recorded no scrub set (its key …)"). The harness never removes or replaces an
|
|
60
|
+
existing key file: fix it as the warning says (`chmod 600` a key you own, or delete a bad file so the next run
|
|
61
|
+
creates a key), then re-run the case to record a fingerprint. Keep `scrubset.key` out of version control.
|
|
62
|
+
|
|
63
|
+
Two more read-only views: `trace <run> --translate-paths` rewrites VM paths to host paths in the text
|
|
64
|
+
`tools`/default views (an effective `hostloop` run with its `mounts.json`), and `diff <a> <b>` masks per-run noise
|
|
65
|
+
(ids, timestamps, host paths) unless you pass `--no-normalize`; on two baselines, `--changelog` renders the
|
|
66
|
+
known fields as prose.
|
|
47
67
|
|
|
48
68
|
**microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
|
|
49
69
|
stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). The full guide is
|
|
4
4
|
[docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
|
|
5
5
|
need while running it.
|
|
6
6
|
|
|
@@ -31,6 +31,9 @@ cowork-harness eval report <eval-dir> # rebuild the report from the eval dir
|
|
|
31
31
|
worst observed runs × 2 × `--reps` exceed x — pre-flight only, agent cost only, never a mid-run stop.
|
|
32
32
|
- In a scenario directory, YAML with no `prompt:` (a session file) is skipped.
|
|
33
33
|
- Run an A/A first (`--allow-identical-arms`, the same source twice) to see your scenarios' noise.
|
|
34
|
+
- A directory arm's snapshot leaves out untracked files unless you pass `--include-untracked` (not with a
|
|
35
|
+
`git:` arm). `--correction bh|holm` picks the multiple-comparison correction behind `confirmed` (default
|
|
36
|
+
`bh` at q = 0.10; `holm` is stricter).
|
|
34
37
|
|
|
35
38
|
## Reading the labels
|
|
36
39
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Gotchas
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). The full "✓ passed ≠ correct" landmine catalog.
|
|
4
4
|
|
|
5
5
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Recipe 7 — Climb a skill with `/claude-api hillclimb` and the harness as its runner
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
|
|
4
4
|
`hillclimb --help` lists `--skill` (help goes to stderr). This page is the loop's procedure, step by step, in the order of the
|
|
5
5
|
`/claude-api hillclimb` guide. Every mechanic (flags, refusals, the gate, `regrade`, `freeze-ref`, exit codes,
|
|
6
6
|
row keys) is in [`hillclimb.md`](hillclimb.md); the setup and the full list of differences from the guide's own
|
|
@@ -59,6 +59,9 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
|
|
|
59
59
|
- **Never pass `--approve-harness`.** It records the harness sha and is the user's to run. A refusal before
|
|
60
60
|
spending prints `refusing to run: …` and exits 2: stop and show the user what it names. `freeze-ref` needs
|
|
61
61
|
no approval.
|
|
62
|
+
- **Never pass `--allow-scrub-change` unattended,** and never put it in an allowed prefix. It sends the judge a
|
|
63
|
+
part of its input that cannot be proven scrubbed with the run's scrub set, so it can disclose a secret the run
|
|
64
|
+
scrubbed. Pass it only after the user has checked the scrub settings and said yes for that listing.
|
|
62
65
|
- **`--skill` takes the bare skill name** (a `skills/<dir>` name or the registered name), never `plugin:name`
|
|
63
66
|
or a path.
|
|
64
67
|
|
|
@@ -76,10 +79,27 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
|
|
|
76
79
|
- **Recompute the headline** from `F/<variant>/results.jsonl`, never from `summary.json`.
|
|
77
80
|
- **Spot-check grading.** Read the lowest-scoring baseline rows' `explanation` and traces. If a rubric is wrong,
|
|
78
81
|
tell the user; after they edit it and approve the new sha, `cowork-harness hillclimb regrade T --flow F`
|
|
79
|
-
re-evaluates
|
|
82
|
+
re-evaluates the rows in place from their kept runs without running the agent: changed deterministic
|
|
80
83
|
assertions and every metric without a judge call, a judged assertion re-judged because its rubric changed. An
|
|
81
84
|
unchanged assertion whose evaluation a harness upgrade changed keeps its recorded outcome (noted); editing the
|
|
82
85
|
assertion, or re-running the case, re-grades it.
|
|
86
|
+
- **Rows listed instead of re-graded (exit 1).** A new or edited rubric line (or another part of the judge's input)
|
|
87
|
+
is sent only when it can be proven scrubbed with the run's scrub set. When it cannot, the row is listed, untouched
|
|
88
|
+
and still carrying its old grade, with no judge call; stderr names the parts and why the run's set is not proven,
|
|
89
|
+
and `F/<variant>/regrade.md` lists the row. The usual causes: a run recorded before the scrub-set fingerprint (no
|
|
90
|
+
`scrubSet` in its `result.json`; one stderr summary line counts them), a run whose key was unusable ("this run
|
|
91
|
+
recorded no scrub set"), a run from another machine or under a replaced `scrubset.key`, or a token the run
|
|
92
|
+
scrubbed that has rotated since. What to do, in order: if the run recorded no scrub set, first have the user fix
|
|
93
|
+
`scrubset.key` as its `::warning:: [scrub-set]` line names (see the debugging reference); the harness never
|
|
94
|
+
replaces an existing bad key, so a new run before that fix records no scrub set again and is listed again. If this
|
|
95
|
+
process lacks a scrub value the run had (`COWORK_HARNESS_SCRUB_VALUES` / `COWORK_HARNESS_SCRUB_KEYS`, a rotated
|
|
96
|
+
token), ask the user to set the run's settings and regrade again; otherwise the listed cases need new runs,
|
|
97
|
+
which record a fresh fingerprint once the key is usable. A pass
|
|
98
|
+
runs nothing for a slot that already has a row, so tell the user and propose running the change as a new variant,
|
|
99
|
+
or a fresh flow dir from the baseline when the baseline's rows are listed. Only if the user has checked the scrub
|
|
100
|
+
settings and asks for it, re-run the same `regrade` with `--allow-scrub-change` (the regrade file then records
|
|
101
|
+
`scrubAcceptedBy`). Never pass it unattended; `--allow-doc-drift`, `--allow-unchecked` and `--rejudge` never
|
|
102
|
+
imply it. Don't compare variants while rows are listed: their grades are from the old rubric.
|
|
83
103
|
- **Triage every zero.** An agent's own failure is a scored row (`meta.failure_class: "errored_agent"`, with
|
|
84
104
|
`meta.termination_rule`). Infrastructure, timeouts, a wrong served model and invalid judge grades are
|
|
85
105
|
`errors.jsonl` rows, never in the scored denominator. A variant that rewords or batches its questions can miss
|
|
@@ -182,7 +202,11 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
|
|
|
182
202
|
- **Grader drift.** If a rubric is wrong, the user edits it and approves the sha; then
|
|
183
203
|
`cowork-harness hillclimb regrade T --flow F` re-grades every variant's rows (default `--variant all`) and writes
|
|
184
204
|
`F/<variant>/regrade.md` when a row moved or was listed. A row whose judge evidence itself changed since it was graded is
|
|
185
|
-
listed instead: ask the user before re-running with `--rejudge`.
|
|
205
|
+
listed instead: ask the user before re-running with `--rejudge`. A row whose new rubric text cannot be proven
|
|
206
|
+
scrubbed with its run's scrub set is listed too, and so is one whose evidence would be less redacted than the
|
|
207
|
+
graded document: handle both as in Step 0.5 (*Rows listed instead of re-graded*), never with `--allow-doc-drift`.
|
|
208
|
+
For the less-redacted row the override, if the user asks for it, is `--rejudge --allow-scrub-change` together:
|
|
209
|
+
`--allow-scrub-change` alone does not re-judge changed evidence.
|
|
186
210
|
- **A new metric.** Add `metrics:` to the scenario, have the user approve the sha, re-run
|
|
187
211
|
`cowork-harness hillclimb state-template T --flow F` and merge only the new `metrics` entries into
|
|
188
212
|
`_state.json`. Rows written before it lack the key (`check` notes them); `hillclimb regrade T --flow F` fills it
|
|
@@ -192,9 +216,16 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
|
|
|
192
216
|
`cowork-harness hillclimb regrade T --flow F --fill-refs` (no `--case`: rows of cases with no pairwise
|
|
193
217
|
assertion need the column too, at no judge cost), then `state-template T --flow F` and merge the new
|
|
194
218
|
`win_vN` entries (it declares them only once no scored row lacks them). Only the baseline's reference
|
|
195
|
-
decides `pass`.
|
|
196
|
-
-
|
|
197
|
-
|
|
219
|
+
decides `pass`. When an entry already frozen lacks a compose key (an assertion added or re-scoped since),
|
|
220
|
+
`freeze-ref` and a baseline pass add it only when this process's scrub set provably covers the run it was frozen
|
|
221
|
+
from; otherwise the case is refused (exit 1). The refusal says to re-run the variant, but a pass never re-runs a
|
|
222
|
+
filled slot and a baseline pass refuses again each time: restore the run's scrub settings and re-run `freeze-ref`,
|
|
223
|
+
or start a fresh flow dir. A `--fill-refs` comparison against a
|
|
224
|
+
reference the row's run never judged is proven only when the run's scrub set is covered; otherwise the row is
|
|
225
|
+
listed (exit 1) and keeps no `win_vN` column. Handle it as in Step 0.5: same scrub settings, or new runs; ask
|
|
226
|
+
the user before any `--allow-scrub-change`.
|
|
227
|
+
- **After a rubric re-grade, compare the ranking** once no row is listed. If the order of the variants flipped,
|
|
228
|
+
tell the user and propose restarting the climb from the baseline.
|
|
198
229
|
|
|
199
230
|
## Step 5 — report and hand back
|
|
200
231
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `hillclimb` — the runner for a `/claude-api hillclimb` loop
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
|
|
4
4
|
`hillclimb --help` lists `--skill` (help goes to stderr). The command reference is
|
|
5
5
|
[docs/cli.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md); this is the part a loop needs
|
|
6
6
|
while it runs. It covers `run`, `check`, `state-template`, `freeze-ref` and `regrade`.
|
|
@@ -224,7 +224,9 @@ measured) whose run delivered an output, under the variant's lock: it refuses wh
|
|
|
224
224
|
its `meta.run_dir`, else by its run id under the current runs root (`--run-dir` / `COWORK_HARNESS_RUNS_DIR`). An
|
|
225
225
|
entry that is already complete is reported (`exists`), never rewritten. One that lacks a compose key (an assert
|
|
226
226
|
added or re-scoped) gains it from the run it was frozen from, marked `unchecked`, and is refused when that run is
|
|
227
|
-
gone: start a fresh flow dir then.
|
|
227
|
+
gone: start a fresh flow dir then. The new document is composed with this process's scrub set, so it is added
|
|
228
|
+
(by `freeze-ref` or a baseline pass) only when that set provably covers the one the run recorded (`scrubSet`);
|
|
229
|
+
otherwise the case is refused with the remedy: re-run the variant, or add it under the run's scrub settings.
|
|
228
230
|
|
|
229
231
|
Flags: `--variant ID` (required), `--flow DIR`, `--case ID` (repeatable), `--output-format text|json` (json
|
|
230
232
|
carries `{frozen, added, exists, refused}`), `--dotenv FILE`, `--run-dir DIR`.
|
|
@@ -280,10 +282,12 @@ every climb there is finished.
|
|
|
280
282
|
the judge's input is covered. Otherwise each part must equal the run's own scrubbed record: the pairwise `## Task`
|
|
281
283
|
line (the prompt), each rubric line, each evidence note, and each reference, by its bytes whatever it is called
|
|
282
284
|
(its send must hash to what a live comparison of the run under the same compose key was sent, `refSentSha256`; a
|
|
283
|
-
|
|
284
|
-
since, or a `--fill-refs` reference the run never judged, proves nothing. A row with a part proven
|
|
285
|
-
|
|
286
|
-
|
|
285
|
+
grade recorded before the reference-send hash, the stored text must be one that comparison recorded). A reference
|
|
286
|
+
re-frozen since, or a `--fill-refs` reference the run never judged, proves nothing. A row with a part proven
|
|
287
|
+
neither way is listed, with no judge call. A run recorded before the scrub-set fingerprint (no `scrubSet`) proves
|
|
288
|
+
no coverage, so its new or edited rubric text is listed (one stderr line counts such rows); so is a run whose key
|
|
289
|
+
was unusable (`scrubSetUnavailable`), a run from another machine or under a replaced key, or one whose token has
|
|
290
|
+
rotated since. Re-run the case, or pass `--allow-scrub-change` after checking the scrub settings: the
|
|
287
291
|
regrade file then records `scrubAcceptedBy`. Neither `--rejudge` nor `--allow-doc-drift` implies it.
|
|
288
292
|
- **`--rejudge`:** every judged assert of every selected row is re-judged, with the flow's references as they are
|
|
289
293
|
now. Use it after a judge change the triggers above do not see. Not with `--fill-refs`.
|
|
@@ -370,7 +374,11 @@ Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant
|
|
|
370
374
|
assert does not line up with the scenario or whose deterministic outcome changed (run a default `regrade` first;
|
|
371
375
|
an agent-failed row aside); one whose
|
|
372
376
|
re-grade is judge-invalid; in a fill one whose kept outcome was judged against a reference that has changed
|
|
373
|
-
since;
|
|
377
|
+
since; one with a part of the judge's input not proven scrubbed with its run's scrub set, one whose evidence
|
|
378
|
+
would be less redacted than the graded document, or one whose re-judge needs an assert whose scrubbed literal
|
|
379
|
+
this process cannot reproduce (each released only by `--allow-scrub-change`, with `--rejudge` for a less-redacted
|
|
380
|
+
row, never by `--allow-doc-drift`); one whose assert has an edit inside a scrubbed literal (only a re-run applies
|
|
381
|
+
it); and an open `judge_invalid` slot.
|
|
374
382
|
|
|
375
383
|
## Exit codes
|
|
376
384
|
|
|
@@ -384,7 +392,8 @@ Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant
|
|
|
384
392
|
- `state-template`: `0`, or `2` on usage or a refusal.
|
|
385
393
|
- `freeze-ref`: `0` no case refused (an entry already complete is reported, not refused); `1` a case refused (no
|
|
386
394
|
good row whose run is under the runs root and delivered its output, a damaged entry, a reference frozen for a
|
|
387
|
-
different prompt, a missing compose key whose run is gone
|
|
395
|
+
different prompt, a missing compose key whose run is gone or whose scrub set this process cannot prove it
|
|
396
|
+
covers, a store write that failed, or a run whose judged
|
|
388
397
|
document cannot be composed, differs from the one its live judge read, or has no live fingerprint); `2` usage
|
|
389
398
|
(a bad `--variant`, no flow dir, a variant with no `results.jsonl`, no selected case with `semantic_pairwise`, the variant's
|
|
390
399
|
lock held by a live run).
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Measurement
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
|
|
4
4
|
|
|
5
5
|
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
6
6
|
|
|
@@ -56,7 +56,7 @@ Compare timings between runs of the same tier and model only.
|
|
|
56
56
|
`fingerprint.skillHash` is content-exact but one-way, so an edit mid-batch silently splits your
|
|
57
57
|
dataset into two generations — and a hash whose source was never frozen identifies a generation that
|
|
58
58
|
is unrecoverable. `stats --group-by skill-hash` separates them after the fact; nothing recovers the
|
|
59
|
-
source.
|
|
59
|
+
source. (`stats --reindex` rebuilds the runs index from the run dirs when it is lost or predates it.)
|
|
60
60
|
3. **Check which arm you actually ran** before analysing anything: `ablated` and
|
|
61
61
|
`context.availableSkills` in each `result.json`.
|
|
62
62
|
4. **Classify each rep three ways**: invocation (`skillsInvoked`), observed source access (did it read
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Run, record and lock
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
|
|
4
4
|
|
|
5
5
|
## Part II — RUN, RECORD & LOCK
|
|
6
6
|
|
|
@@ -21,9 +21,20 @@ never calls the semantic judge, so a `semantic_matches` or `semantic_pairwise` a
|
|
|
21
21
|
`cowork-harness regrade <run-dir> --scenario <scenario.yaml>` re-grades those against the kept run (the judge call
|
|
22
22
|
is the only spend) and reports whether the judge read the same document the live judge did (widening the evidence
|
|
23
23
|
scope needs `--allow-unchecked`: content the live judge never read is refused otherwise). A part of the judge's
|
|
24
|
-
input that cannot be proven scrubbed with the run's scrub set is refused too
|
|
25
|
-
|
|
26
|
-
|
|
24
|
+
input that cannot be proven scrubbed with the run's scrub set is refused too (exit 2, `refusals[]`, naming the parts,
|
|
25
|
+
never their text). A run records a keyed fingerprint of its scrub set (`result.json` `scrubSet`); when this
|
|
26
|
+
process's set provably covers it, everything is sent. Otherwise only parts equal to the run's own scrubbed record
|
|
27
|
+
are: the pairwise task line, each rubric line, each evidence-health and scratch note, and each reference still
|
|
28
|
+
byte-identical to the one a live comparison of the run was sent. So a new or edited rubric line
|
|
29
|
+
(`rubric_unverifiable`), a changed prompt (`task_unverifiable`), a changed evidence note (`evidence_unverifiable`),
|
|
30
|
+
or a reference re-frozen since the run (`reference_unverifiable`) is refused on a run recorded before the
|
|
31
|
+
scrub-set fingerprint (no `scrubSet`), on one whose key was unusable (`scrubSetUnavailable`, see
|
|
32
|
+
`debugging.md`), from another machine, or after a token it scrubbed rotated. Re-run the case (a new run records a
|
|
33
|
+
fresh fingerprint), regrade with the run's scrub settings, or pass `--allow-scrub-change` after checking them (the
|
|
34
|
+
grade records `scrubAcceptedBy`); `--allow-doc-drift` and `--allow-unchecked` never imply it.
|
|
35
|
+
`ref freeze` applies the same proof: a document with no live fingerprint (`--allow-unchecked`) is composed with this
|
|
36
|
+
process's scrub set, so it is frozen only when that set provably covers the run's, or with `--allow-scrub-change`;
|
|
37
|
+
`ref verify` refuses both flags. A run dir moved or
|
|
27
38
|
downloaded from where it ran is read from where it is; a COPY beside its still-present original is refused (its
|
|
28
39
|
`result.json` names the original's files), so grade the original, or re-run the scenario. Or skip
|
|
29
40
|
the discovery/encode/record dance entirely and answer gates **live during the recording** with
|
|
@@ -352,7 +363,8 @@ modes exist to withhold; `status.json` is still written either way, so `status`
|
|
|
352
363
|
passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
|
|
353
364
|
newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
|
|
354
365
|
rather than hanging forever. (Fuller recipe in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) — repo-only, not in the installed
|
|
355
|
-
payload; `cowork-harness status --help` has the flags.)
|
|
366
|
+
payload; `cowork-harness status --help` has the flags.) To find a scenario's newest run dir, use
|
|
367
|
+
`cowork-harness status --latest-for <scenario>`, which orders by run time rather than directory mtime.
|
|
356
368
|
|
|
357
369
|
**Poll with `--follow`, not with a shell loop over `status`'s stdout.** The one-shot text form prints to
|
|
358
370
|
**stderr** and writes nothing to stdout; `--output-format json` (one envelope) and `--follow` (one JSON
|
|
@@ -373,6 +385,25 @@ in-flight run mid-artifact-write. Run it foreground, or detached from any proces
|
|
|
373
385
|
(The `status.json` liveness above is exactly what surfaces such a teardown as `"error"`/`stale` rather
|
|
374
386
|
than a stuck `"running"`.)
|
|
375
387
|
|
|
388
|
+
### Other flags worth knowing
|
|
389
|
+
|
|
390
|
+
- `skill` / `critique`: `--prompt-file <path>` reads the prompt verbatim (no shell parsing); `--marketplace <dir>
|
|
391
|
+
--enable name@mkt` loads skills through a marketplace; `--timeout <ms>` is the wall-clock budget;
|
|
392
|
+
`--allow-host-writes` consents to a writable `hostloop` connected folder; `--verbose` adds thinking, tool inputs
|
|
393
|
+
and the sub-agent tree to the output.
|
|
394
|
+
- `run --matrix`: `--max-cells <n>` caps the cross-product (default 16), and a truncated matrix fails unless
|
|
395
|
+
`--allow-truncated-matrix` judges only the cells that ran.
|
|
396
|
+
- `record`: `--max-artifact-bytes <n>` caps an inlined artifact body (default 65536); `--rerecord-stale
|
|
397
|
+
--from-embedded` re-records from the cassette's embedded scenario when no source file resolves.
|
|
398
|
+
- `replay --mutate`: `--mutate-include` / `--mutate-exclude <glob>` scope which artifact paths are perturbed, and
|
|
399
|
+
`--mutate-max-per-file` / `--mutate-max-total` raise the sample caps (default 10 / 50).
|
|
400
|
+
- `probe-dispatch --expect-write <suffix>` counts only a sub-agent write whose path ends with the suffix as delivered.
|
|
401
|
+
- `ref freeze --case-id <id>` overrides the store entry's name (default: the scenario's name).
|
|
402
|
+
- `fixture export <run-dir> --out <dir>` refuses a file holding a secret, and a text file or file name holding a host
|
|
403
|
+
path unless `--allow-host-paths` (a path into a run dir or a guest session is refused regardless).
|
|
404
|
+
- `prune [--keep-last <n>] [--pinned-older-than <N>d]` removes old run dirs (default `--keep-last 5`);
|
|
405
|
+
`--dry-run` previews it.
|
|
406
|
+
|
|
376
407
|
### Place assertions in the right CI lane
|
|
377
408
|
|
|
378
409
|
CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, replay class, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.4.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.4.1`
|
|
4
4
|
(baseline `desktop-2.19675.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `semantic_matches` — the full semantics
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.4.
|
|
3
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`). The one-line summary is in
|
|
4
4
|
[assertion-catalog.md](assertion-catalog.md); this is the whole contract, split out so the catalog stays within the
|
|
5
5
|
agent's single-read size.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 4.4.
|
|
5
|
+
Tracks `cowork-harness 4.4.1` (baseline `desktop-2.19675.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,42 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [4.4.1] — 2026-10-05
|
|
10
|
+
|
|
11
|
+
A companion-skill release: the skill now teaches 4.4.0's re-grade scrub-set proof, and a test keeps every CLI flag
|
|
12
|
+
and `error.code` named in it. No CLI behaviour changes.
|
|
13
|
+
|
|
14
|
+
### Upgrade notes
|
|
15
|
+
|
|
16
|
+
- **Cassettes: no re-record needed.** Nothing under `src/`, `baselines/` or `docker/` changed, `latest` still
|
|
17
|
+
resolves to `desktop-2.19675.0`, and `CASSETTE_VERSION` is still 14.
|
|
18
|
+
|
|
19
|
+
### Documentation
|
|
20
|
+
|
|
21
|
+
- **The companion skill covers the re-grade scrub-set proof.**
|
|
22
|
+
- The hillclimb recipe says what to do with rows `hillclimb regrade` lists (exit 1) because a part of the
|
|
23
|
+
judge's input cannot be proven scrubbed with the run's scrub set (a run recorded before the scrub-set
|
|
24
|
+
fingerprint, an unusable key, another machine, a rotated token): regrade with the run's scrub settings, or
|
|
25
|
+
re-run the cases, which records a fresh fingerprint. It never passes `--allow-scrub-change` unattended; that
|
|
26
|
+
flag is the user's call after checking the scrub settings. Its `freeze-ref` and `--fill-refs` steps say when
|
|
27
|
+
they refuse or list for the same reason.
|
|
28
|
+
- The `regrade` reference names every part refused when unproven: the pairwise task line, rubric lines,
|
|
29
|
+
evidence-health and scratch notes, and a reference re-frozen since the run. It covers `ref freeze`'s proof for
|
|
30
|
+
an unchecked document (`--allow-scrub-change`), which `ref verify` refuses, and what the proof does not cover.
|
|
31
|
+
- The debugging reference explains the unusable-key warning (`scrubset.key` a symlink, readable by others,
|
|
32
|
+
owned by another user, empty, unreadable or not a key): such a run records `scrubSetUnavailable`, so its
|
|
33
|
+
later re-grades refuse with `scrubSet: "unrecorded"`; fix or delete the key file, then re-run the case. The
|
|
34
|
+
CI recipe says to keep `scrubset.key` out of version control and artifact uploads.
|
|
35
|
+
- **The companion skill names every CLI long flag and every `error.code`.** A new test reads the flags from each
|
|
36
|
+
command's `--help` and the codes from `schema/*.json`, and fails on one the skill never mentions, unless it is
|
|
37
|
+
allowlisted with a reason. Its first run found 41 flags and 3 `regrade` codes (`doc_drift`,
|
|
38
|
+
`unchecked_content`, `no_semantic_asserts`) the skill never named. The skill now describes 39 of the flags and
|
|
39
|
+
all 3 codes; the other 2 are allowlisted (`--diff`, used only by `sync --diff` for baseline maintenance, and
|
|
40
|
+
`--strict-independent`, which is help prose, not a flag).
|
|
41
|
+
- [RELEASING.md](./RELEASING.md) gains a step that reconciles the companion skill against the new version's
|
|
42
|
+
CHANGELOG, recipes first, before the live gate: a passage that is stale or missing blocks the release. The pull
|
|
43
|
+
request template asks every PR to state whether the companion skill is affected.
|
|
44
|
+
|
|
9
45
|
## [4.4.0] — 2026-10-05
|
|
10
46
|
|
|
11
47
|
A security fix: `regrade` and `hillclimb regrade` no longer send a judge a value the run scrubbed. A run now records
|
package/DESIGN.md
CHANGED
|
@@ -208,7 +208,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
208
208
|
> Cowork system-prompt fingerprint all unchanged from `1.20186.0`, with the staged VM ELF re-synced
|
|
209
209
|
> 2.1.202 → 2.1.205 — and the live pass of that era was deliberately **not** restamped onto it.)
|
|
210
210
|
|
|
211
|
-
> **Scope of that claim.** `2026-10-03 / desktop-2.19675.0` is the baseline carrying the latest live pass, run against agent **2.1.286** (the staged Linux ELF for `container` and `microvm`, and the staged native macOS build `f2326db61802` for `hostloop`); the `protocol` tier runs the HOST `claude`, **2.1.289** for the final pass. No baselines have shipped since. **The 4.4.0 pass, 2026-10-05, on the release-prep commit `e9c516df`** (same baseline and staged agent; host `claude` 2.1.289 for `protocol`): (i) `vitest list --staticParse=false` listed all 23 live tests with no `SKIPPED` warning when the token was exported, and 6 with four warnings when none was resolvable; (ii) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip; (iii) the auto-memory check on all three tiers, `Tests 2 passed (2)` each; (iv) the companion skill's router check, three prompts on `container`: the authoring prompt read `authoring.md` and wrote a scenario that passes `lint`, the tool-timing prompt read `measurement.md` and answered with `trace … --view tool-durations`, and the debugging prompt read no reference and answered with correct clarifying questions, so `debugging.md` routing was not exercised in this pass; (v) the re-grade scrub-set check: a live run on the default runs root created `scrubset.key` beside the runs directory (mode 0600) and recorded `scrubSet` (`v: 1`) with the configured scrub literal absent from `result.json`; a re-grade of that run with an edited rubric line and a smaller scrub set was refused as `rubric_unverifiable` (exit 2) naming only the edited line, with no judge call and the run directory byte-identical; and a re-grade of a run recorded before 4.4 with an edited rubric line was refused with the `legacy` reason and the one-line remedy, again with no judge call and the run directory byte-identical. Live spend about **$2.83** recorded. **The 4.3.0 pass, on the release commit `8a5e479f`:** (1) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip (the auto-memory prerequisites check, which runs only under `COWORK_LIVE_REQUIRE=1`); before it ran, `vitest list --staticParse=false` listed all 23 with no `SKIPPED` warning, and 6 with four warnings when no token was resolvable, so a silent skip would have shown; (2) the auto-memory check on all three tiers, `Tests 2 passed (2)` each, every init frame without `memory_paths`; (3) the graders' pinned effort: every judged assert that called the judge in the live hillclimb acceptance runs (45, Opus 4.8: 25 `semantic_matches`, 20 `semantic_pairwise`; host `claude` 2.1.288 and 2.1.289) records `judgeTransport.effort: "high"`; the five pairwise asserts without a `judgeTransport` are the reference variant's own rows, which call no judge. **Earlier in the same release, on the release candidate's tree:** (4) a `microvm` smoke on a freshly booted VM (agent 2.1.286); (5) `boundary-check`, 6/6 constraints enforced; (6) an `eval` A/A at `--reps 4` on `container`, every row `no detectable change`, its report rebuilt byte-identical; (7) the companion skill's router check, three prompts on `container`: routing was correct on all three; one answer was wrong because a reference showed `artifact_text`'s `contains` without its list shape (fixed in the references, with a test that keeps every list-typed field shown as a list), and the same prompt's verdict went red on a false `host_path_leak`, because a reference spelled out the host roots that check looks for and the agent echoed it (fixed in the skill's files and in the check, which no longer counts a literal that comes from the plugin's or a local skill's own staged files); re-run after the fixes, the prompt routed correctly, wrote a scenario that passes `lint` and loads, and carried no `host_path_leak`; (8) the `hostloop` detached-agent kill check: the first pass found the workspace sidecar container, its `docker run` client and its network surviving SIGINT, which was fixed, and the re-run passed on a normal run, SIGINT to the harness process and SIGINT to the whole process group (exit 130, nothing left behind); (9) `vm delete`'s usage refusal; (10) `chat` under a real terminal, on `protocol`: Ctrl-C mid-turn exits 130 in about 2 s with nothing left running and no `result.json`; `/exit` followed by Ctrl-C exits 130 after writing the turn's `result.json`; Ctrl-C at the idle prompt ends the session normally (exit 0, `result.json` written); SIGHUP mid-turn exits 129 with nothing left running; and closing the terminal window mid-turn (run by hand) leaves no harness or agent process and no `result.json`. Total live spend about **$11.20**, including the re-runs. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction, and these checks are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (b) **CI does not live-validate anything.** Its "scenario suite (… live inference)" job is skipped as a whole without an `ANTHROPIC_API_KEY` repository secret, and none is set, so it shows as *skipped*. Before 2026-09 it instead ran with every real step skipped and reported *success*, so an older green check there means nothing ran. This note, not CI, is the live evidence. Separately and not a live matter: all four committed cassettes were **re-recorded** on 2026-10-03 against `2.19675.0` with auto-memory off, each on its original model: `example-pdf-skill` and `test/fixtures/tool-call-dispatch/dispatch-shell.cassette.json` (`container`, agent 2.1.286), `hostloop-computer-links` (`hostloop`, agent 2.1.286) and `example-multiselect-gate` (`protocol`, which runs the host CLI, here 2.1.287), about $0.75 in all. Their init frames carry no `memory_paths`. The recorder's own scrub (now applied to every recording) removed, across the four, the account's model menu, a subscription account's `rate_limit_info` (from `example-pdf-skill`, `dispatch-shell` and `hostloop-computer-links`) and the agent's sub-agent hand-back frame (from `dispatch-shell`). `verify-cassettes` exits 0 on all four, with one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
211
|
+
> **Scope of that claim.** `2026-10-03 / desktop-2.19675.0` is the baseline carrying the latest live pass, run against agent **2.1.286** (the staged Linux ELF for `container` and `microvm`, and the staged native macOS build `f2326db61802` for `hostloop`); the `protocol` tier runs the HOST `claude`, **2.1.289** for the final pass. No baselines have shipped since. **The 4.4.1 pass, 2026-10-05, on the release-prep commit `3e6fe443`** (a companion-skill-only release; same baseline and staged agent; host `claude` 2.1.289 for `protocol`): (i) `vitest list --staticParse=false` listed all 23 live tests with no `SKIPPED` warning when the token was exported, and 6 with four warnings when none was resolvable; (ii) the auto-memory check on all three tiers, `Tests 2 passed (2)` each, every init frame without `memory_paths`; (iii) the companion skill's router check, three prompts on `container`: the authoring prompt read `authoring.md` (and `scenario-schema.md`, searching `assertion-catalog.md`) and wrote a scenario that passes `lint` with no error or warning, the debugging prompt read `debugging.md` and answered with correct clarifying questions, and the tool-timing prompt read `measurement.md` (and `debugging.md`, which documents `trace`'s views) and answered with `trace … --view tool-durations`. The live suite itself was not re-run: this release changes no code that runs. Live spend about **$1.08**. **The 4.4.0 pass, 2026-10-05, on the release-prep commit `e9c516df`** (same baseline and staged agent; host `claude` 2.1.289 for `protocol`): (i) `vitest list --staticParse=false` listed all 23 live tests with no `SKIPPED` warning when the token was exported, and 6 with four warnings when none was resolvable; (ii) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip; (iii) the auto-memory check on all three tiers, `Tests 2 passed (2)` each; (iv) the companion skill's router check, three prompts on `container`: the authoring prompt read `authoring.md` and wrote a scenario that passes `lint`, the tool-timing prompt read `measurement.md` and answered with `trace … --view tool-durations`, and the debugging prompt read no reference and answered with correct clarifying questions, so `debugging.md` routing was not exercised in this pass; (v) the re-grade scrub-set check: a live run on the default runs root created `scrubset.key` beside the runs directory (mode 0600) and recorded `scrubSet` (`v: 1`) with the configured scrub literal absent from `result.json`; a re-grade of that run with an edited rubric line and a smaller scrub set was refused as `rubric_unverifiable` (exit 2) naming only the edited line, with no judge call and the run directory byte-identical; and a re-grade of a run recorded before 4.4 with an edited rubric line was refused with the `legacy` reason and the one-line remedy, again with no judge call and the run directory byte-identical. Live spend about **$2.83** recorded. **The 4.3.0 pass, on the release commit `8a5e479f`:** (1) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip (the auto-memory prerequisites check, which runs only under `COWORK_LIVE_REQUIRE=1`); before it ran, `vitest list --staticParse=false` listed all 23 with no `SKIPPED` warning, and 6 with four warnings when no token was resolvable, so a silent skip would have shown; (2) the auto-memory check on all three tiers, `Tests 2 passed (2)` each, every init frame without `memory_paths`; (3) the graders' pinned effort: every judged assert that called the judge in the live hillclimb acceptance runs (45, Opus 4.8: 25 `semantic_matches`, 20 `semantic_pairwise`; host `claude` 2.1.288 and 2.1.289) records `judgeTransport.effort: "high"`; the five pairwise asserts without a `judgeTransport` are the reference variant's own rows, which call no judge. **Earlier in the same release, on the release candidate's tree:** (4) a `microvm` smoke on a freshly booted VM (agent 2.1.286); (5) `boundary-check`, 6/6 constraints enforced; (6) an `eval` A/A at `--reps 4` on `container`, every row `no detectable change`, its report rebuilt byte-identical; (7) the companion skill's router check, three prompts on `container`: routing was correct on all three; one answer was wrong because a reference showed `artifact_text`'s `contains` without its list shape (fixed in the references, with a test that keeps every list-typed field shown as a list), and the same prompt's verdict went red on a false `host_path_leak`, because a reference spelled out the host roots that check looks for and the agent echoed it (fixed in the skill's files and in the check, which no longer counts a literal that comes from the plugin's or a local skill's own staged files); re-run after the fixes, the prompt routed correctly, wrote a scenario that passes `lint` and loads, and carried no `host_path_leak`; (8) the `hostloop` detached-agent kill check: the first pass found the workspace sidecar container, its `docker run` client and its network surviving SIGINT, which was fixed, and the re-run passed on a normal run, SIGINT to the harness process and SIGINT to the whole process group (exit 130, nothing left behind); (9) `vm delete`'s usage refusal; (10) `chat` under a real terminal, on `protocol`: Ctrl-C mid-turn exits 130 in about 2 s with nothing left running and no `result.json`; `/exit` followed by Ctrl-C exits 130 after writing the turn's `result.json`; Ctrl-C at the idle prompt ends the session normally (exit 0, `result.json` written); SIGHUP mid-turn exits 129 with nothing left running; and closing the terminal window mid-turn (run by hand) leaves no harness or agent process and no `result.json`. Total live spend about **$11.20**, including the re-runs. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction, and these checks are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (b) **CI does not live-validate anything.** Its "scenario suite (… live inference)" job is skipped as a whole without an `ANTHROPIC_API_KEY` repository secret, and none is set, so it shows as *skipped*. Before 2026-09 it instead ran with every real step skipped and reported *success*, so an older green check there means nothing ran. This note, not CI, is the live evidence. Separately and not a live matter: all four committed cassettes were **re-recorded** on 2026-10-03 against `2.19675.0` with auto-memory off, each on its original model: `example-pdf-skill` and `test/fixtures/tool-call-dispatch/dispatch-shell.cassette.json` (`container`, agent 2.1.286), `hostloop-computer-links` (`hostloop`, agent 2.1.286) and `example-multiselect-gate` (`protocol`, which runs the host CLI, here 2.1.287), about $0.75 in all. Their init frames carry no `memory_paths`. The recorder's own scrub (now applied to every recording) removed, across the four, the account's model menu, a subscription account's `rate_limit_info` (from `example-pdf-skill`, `dispatch-shell` and `hostloop-computer-links`) and the agent's sub-agent hand-back frame (from `dispatch-shell`). `verify-cassettes` exits 0 on all four, with one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
212
212
|
|
|
213
213
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
214
214
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^4.4.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^4.4.1"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.4.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.4.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.4.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.4.1"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v4`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Improving a skill round by round with Claude Code's `/claude-api hillclimb` loop? The harness is its runner, with
|
package/RELEASING.md
CHANGED
|
@@ -177,8 +177,11 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
|
|
|
177
177
|
- [ ] **Does this release add or change a user-facing CLI flag, assertion key, cassette field, message,
|
|
178
178
|
top-level scenario key, or version coupling?** If so update **CHANGELOG.md + README.md +
|
|
179
179
|
`.claude/skills/cowork-harness/SKILL.md` + `references/`** — a version bump is NOT documentation.
|
|
180
|
-
|
|
181
|
-
|
|
180
|
+
`test/skill-docs-sync.test.ts` and `test/skill-flag-coverage.test.ts` guard names only: every assertion
|
|
181
|
+
key, cassette schema field, CLI long flag and `error.code` value must appear somewhere in the skill. A
|
|
182
|
+
message is guarded by nothing. A name can appear in a passage that no longer describes what it does, so
|
|
183
|
+
those tests pass while the instructions go stale; a changed refusal, exit code, default or flag meaning is
|
|
184
|
+
caught only by the companion-skill reconcile step below.
|
|
182
185
|
Two consecutive consumer adoption reports spent ~40% of their findings on exactly this.
|
|
183
186
|
- [ ] **New top-level scenario key?** Then the docs above MUST also state **the version floor and what an
|
|
184
187
|
older CLI does with the key** — the loader is `z.strictObject`, so an unknown key is a hard error
|
|
@@ -211,6 +214,18 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
|
|
|
211
214
|
a release that reports nothing reads as "unchanged" whether or not anyone checked — the same
|
|
212
215
|
silence-is-not-a-verdict failure the 3.8.0 corpus preview exists to close. A consumer diffed two
|
|
213
216
|
tags' `src/` by hand and still had to ask (2026-09-21) because the 3.7.0 notes said nothing either way.
|
|
217
|
+
- [ ] **Reconcile the companion skill against the CHANGELOG.** Do it once the new version's CHANGELOG section is
|
|
218
|
+
final, and before any live gate. For every bullet in that section (every heading: Security, Upgrade notes,
|
|
219
|
+
Added, Changed, Fixed, Documentation), find each passage of `.claude/skills/cowork-harness/SKILL.md` and
|
|
220
|
+
`references/*.md` that describes the behaviour, and mark it **ok**, **stale** or **missing**. Search widely:
|
|
221
|
+
the command, its flags, its refusal and `error.code` names, and the nouns the bullet uses. Do the
|
|
222
|
+
step-by-step recipes first (`references/hillclimb-recipe.md`, `references/task-recipes.md`,
|
|
223
|
+
`references/authoring.md`): an agent follows a recipe literally, so a stale step there does the most harm. A
|
|
224
|
+
behaviour change is more than a new name: a changed refusal, exit code or default, or a change to what a flag
|
|
225
|
+
unlocks, makes every passage describing the old behaviour stale. Keep the table (bullet | skill file:line |
|
|
226
|
+
status | fix) with the release work; it is the evidence for this step. **Any stale or missing passage
|
|
227
|
+
blocks the release** until it is fixed, in present tense, with no per-release history in the skill. The
|
|
228
|
+
name guard above passes while a passage is stale, which is why this step is manual.
|
|
214
229
|
- [ ] Bump every version location (items 1–10) with **`npm run bump -- X.Y.Z --write`** — it rewrites all
|
|
215
230
|
of them via targeted patterns and updates the lockfile + self-checks `check:versions` (run without
|
|
216
231
|
`--write` first to preview the diff; dry-run is the default). It deliberately does **not** touch the
|
package/docs/ci.md
CHANGED
|
@@ -144,7 +144,7 @@ jobs:
|
|
|
144
144
|
model: claude-sonnet-5 # used only where a scenario's session sets no `model:`
|
|
145
145
|
```
|
|
146
146
|
|
|
147
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^4` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^4.4.
|
|
147
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^4` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^4.4.1` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN; INFO never gates unless `lint` gets `--min-severity INFO` via `extra-args`; a reviewed `lint-skill` finding you accept is suppressed with `extra-args: --suppressions <file> --strict-ignores` (one entry per accepted site, in a JSON file kept outside the plugin; a stale entry fails too) or `--ignore-rule <rule>[=<glob>]`, either of which keeps every other WARN gating — see [cli.md](./cli.md#flags-worth-knowing)), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only), `model` (live lane only: exported as `COWORK_HARNESS_MODEL`, which fills in where a session sets no `model:` and no `--model` is passed; a `run` that resolves no model is refused with exit 2). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
148
148
|
|
|
149
149
|
CI uses `ANTHROPIC_API_KEY` specifically because there's no interactive browser available to run
|
|
150
150
|
`claude setup-token`'s OAuth flow in a GitHub Actions runner; locally, the OAuth token is preferred because
|
package/docs/cli.md
CHANGED
|
@@ -18,7 +18,7 @@ companion skill, CI). This page is the CLI one.
|
|
|
18
18
|
**Install from npm:**
|
|
19
19
|
|
|
20
20
|
```bash
|
|
21
|
-
npm install -g "cowork-harness@^4.4.
|
|
21
|
+
npm install -g "cowork-harness@^4.4.1" # puts the `cowork-harness` command on your PATH
|
|
22
22
|
```
|
|
23
23
|
|
|
24
24
|
**Or build from source:**
|
|
@@ -38,7 +38,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
38
38
|
|
|
39
39
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
40
40
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
41
|
-
> From a global install (`npm i -g "cowork-harness@^4.4.
|
|
41
|
+
> From a global install (`npm i -g "cowork-harness@^4.4.1"`), point at the package root instead:
|
|
42
42
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
43
43
|
> (or copy the cassette into your own project and pass that path).
|
|
44
44
|
|
|
@@ -48,7 +48,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
48
48
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
49
49
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
50
50
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
51
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^4.4.
|
|
51
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^4.4.1"`.
|
|
52
52
|
|
|
53
53
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
54
54
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -126,7 +126,7 @@ The parts people ask about. **`package.json`'s `files[]` is the exhaustive, mach
|
|
|
126
126
|
this table is the readable summary of it, and deliberately omits the infrastructure that always ships
|
|
127
127
|
(`baselines/`, `schema/`, `fixtures/`, `scripts/`, `docker/`).
|
|
128
128
|
|
|
129
|
-
| What ships | npm global (`npm install -g "cowork-harness@^4.4.
|
|
129
|
+
| What ships | npm global (`npm install -g "cowork-harness@^4.4.1"`) | Source checkout (`git clone` + `npm ci`) |
|
|
130
130
|
|---|---|---|
|
|
131
131
|
| CLI, `scenario.py` + assertion keys (enough for `lint` in CI) | ✓ | ✓ |
|
|
132
132
|
| `SKILL.md`, all of `docs/`, `SPEC.md`/`DESIGN.md`/`AGENTS.md` | ✓ | ✓ |
|
|
@@ -141,7 +141,7 @@ since a global install puts nothing in your working directory. The matrix, answe
|
|
|
141
141
|
examples are the only ones that still need a source checkout. The **marketplace skill install** is
|
|
142
142
|
narrower again — it pulls only `.claude/skills/cowork-harness/` (SKILL.md + `references/` +
|
|
143
143
|
`scenario.py`/assertion keys, per `.claude-plugin/marketplace.json`'s `source`); everything in the npm column
|
|
144
|
-
arrives when the skill's first command self-bootstraps `npx "cowork-harness@^4.4.
|
|
144
|
+
arrives when the skill's first command self-bootstraps `npx "cowork-harness@^4.4.1"` — the last row stays
|
|
145
145
|
✗ either way, since `matrices/`, `answer-policies/` and `probes/` are not published at all. See
|
|
146
146
|
[docs/companion-skill.md](./companion-skill.md) for that install path.
|
|
147
147
|
|
package/docs/companion-skill.md
CHANGED
|
@@ -28,7 +28,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
28
28
|
claude plugin install cowork-harness@cowork-harness
|
|
29
29
|
```
|
|
30
30
|
|
|
31
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^4.4.
|
|
31
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^4.4.1"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
32
32
|
|
|
33
33
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
34
34
|
|
|
@@ -43,7 +43,7 @@ npx skills add yaniv-golan/cowork-harness --skill cowork-harness
|
|
|
43
43
|
short entrypoint (routing, the false-green invariants, short workflows); the detail it routes to lives in
|
|
44
44
|
`references/`, which the agent reads on demand. Everything
|
|
45
45
|
else (the CLI, `docs/`, the worked examples, the pytest lane) arrives when the skill's first command
|
|
46
|
-
self-bootstraps `npx "cowork-harness@^4.4.
|
|
46
|
+
self-bootstraps `npx "cowork-harness@^4.4.1"`, which pulls the same npm package as a global install.
|
|
47
47
|
|
|
48
48
|
For the full package-contents table — what a global install gives you versus a source checkout — see
|
|
49
49
|
[docs/cli.md → What ships](./cli.md#what-ships). It is maintained there, once.
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^4.4.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^4.4.1"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "4.4.
|
|
3
|
+
"version": "4.4.1",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|