cowork-harness 3.7.0 → 3.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (60) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +9 -7
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +25 -11
  3. package/.claude/skills/cowork-harness/references/critique.md +25 -4
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +6 -4
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +73 -0
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +28 -13
  9. package/CHANGELOG.md +331 -0
  10. package/DESIGN.md +2 -2
  11. package/README.md +4 -4
  12. package/RELEASING.md +8 -0
  13. package/SPEC.md +11 -1
  14. package/baselines/desktop-2.7032.0.json +1049 -0
  15. package/baselines/prompts/cowork-system-prompt-fingerprints.json +7 -0
  16. package/baselines/prompts/desktop-1.46388.3/subagent-append-hl.md +4 -1
  17. package/baselines/provisioning/rootfs-provisioning.json +17 -14
  18. package/dist/agent/session.js +4 -2
  19. package/dist/assert.js +36 -0
  20. package/dist/baseline.js +14 -0
  21. package/dist/cli.js +1 -1
  22. package/dist/critique/command.js +340 -56
  23. package/dist/critique/limitations.js +9 -0
  24. package/dist/critique/package-evidence.js +3 -1
  25. package/dist/critique/skill-invocation.js +140 -0
  26. package/dist/hostloop/workspace-handler.js +10 -3
  27. package/dist/loop-decision.js +22 -1
  28. package/dist/prompt/subagent-manifest.js +10 -3
  29. package/dist/run/cassette.js +2 -0
  30. package/dist/run/hook-events.js +10 -7
  31. package/dist/run/skill-flag-surface.js +4 -1
  32. package/dist/run/tool-name-canonicalization.js +39 -2
  33. package/dist/run/verdict.js +3 -1
  34. package/dist/runtime/argv.js +4 -0
  35. package/dist/session.js +39 -16
  36. package/dist/sync/cowork-sync.js +82 -31
  37. package/dist/types.js +9 -0
  38. package/docs/cassette.md +4 -1
  39. package/docs/ci.md +17 -7
  40. package/docs/cli.md +10 -10
  41. package/docs/companion-skill.md +2 -2
  42. package/docs/critique.md +113 -9
  43. package/docs/fidelity-gaps.md +164 -40
  44. package/docs/gotchas.md +3 -1
  45. package/docs/maintenance.md +1 -1
  46. package/docs/scenario.md +7 -2
  47. package/docs/session.md +8 -6
  48. package/docs/subagents.md +1 -1
  49. package/examples/replays/README.md +1 -1
  50. package/examples/replays/example-multiselect-gate.cassette.json +55 -62
  51. package/examples/replays/example-pdf-skill.cassette.json +122 -102
  52. package/examples/replays/hostloop-computer-links.cassette.json +62 -64
  53. package/examples/sessions/stop-hook-probe.yaml +5 -0
  54. package/llms.txt +1 -1
  55. package/package.json +1 -1
  56. package/schema/critique-report.json +9 -1
  57. package/schema/scenario.schema.json +78 -0
  58. package/schema/session.schema.json +1 -1
  59. package/scripts/capture-rootfs-manifest.ts +35 -0
  60. package/scripts/check-versions.ts +51 -0
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.7.0
7
- tracks-harness: cowork-harness 3.7.0 (baseline desktop-2.2553.1)
6
+ version: 3.8.1
7
+ tracks-harness: cowork-harness 3.8.1 (baseline desktop-2.7032.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.7.0` (baseline
29
- > `desktop-2.2553.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.8.1` (baseline
29
+ > `desktop-2.7032.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
32
32
  ## Preflight — make sure the harness can actually run
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.7.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.7.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.7.0"`. **Pin `@^3.7.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.8.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.8.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.8.1"`. **Pin `@^3.8.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -81,7 +81,7 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
81
81
  correct answers after you edit it?) → author `semantic_matches` scenarios and gate on the per-claim
82
82
  profile. See **Recipe 5** in `references/task-recipes.md` (validity, N≥3, discrimination — the traps).
83
83
  - **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
84
- <folder> --prompt "<probe>"`. Four model workloads and 10–20 minutes; budget from
84
+ <folder> --prompt "<probe>"`. Up to four model workloads (zero with `--corpus-only`; pass 2 is skipped with no self-report) and 10–20 minutes; budget from
85
85
  `report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
86
86
  **DO**" — routing, artifact location, narration — use `skill` instead**: no evaluator, a fraction of
87
87
  the cost, and it answers that question directly. Report and evidence-package shapes:
@@ -568,7 +568,9 @@ Recognize these before "fixing" a non-bug:
568
568
  - **`missing_capability`** — the lean `core` agent image is a deliberate partial mirror of real Cowork's
569
569
  rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
570
570
  `markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
571
- (`magick`) can trip this even though real Cowork **ships** those. The message says so ("likely a FALSE
571
+ (`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
572
+ Desktop `2.7032.0` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
573
+ behind that sentence). The message says so ("likely a FALSE
572
574
  NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
573
575
  `COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
574
576
  `allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.7.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.8.1"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,13 +36,17 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.275 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.280 # match your scenario's pinned baseline's agentVersion
40
40
  # The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
41
- # served only from .../claude-code-releases/rc/<commit>/ — the stable path 404s for those, and
42
- # 2.1.255 is one. Take B from your pinned baseline's agentBinary.releaseBaseUrl; baselines
43
- # written before that field existed were stable-staged, so their base is the plain
44
- # https://downloads.claude.ai/claude-code-releases.
45
- B=https://downloads.claude.ai/claude-code-releases
41
+ # served from .../claude-code-releases/rc/<commit>/. For some versions the stable path 404s
42
+ # (2.1.255); for others it returns 200 and serves a DIFFERENT BUILD UNDER THE SAME VERSION
43
+ # NUMBER — measured for 2.1.280 on 2026-09-23: stable linux-arm64 233,103,352 B / 92f2b4fd…
44
+ # (commit 80abbfe7) vs RC 233,037,816 B / a1b25d70… (commit bddba3ab), and the staged binary is
45
+ # the RC one. So "the stable URL works" is NOT evidence you have the right build: always take B
46
+ # from your pinned baseline's agentBinary.releaseBaseUrl. The checksum step fails closed if you
47
+ # don't, but it cannot tell you why. Baselines written before that field existed were
48
+ # stable-staged, so their base is the plain https://downloads.claude.ai/claude-code-releases.
49
+ B=https://downloads.claude.ai/claude-code-releases/rc/bddba3abd5da53d0c540cfc76a8d18b44633d568
46
50
  # The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
47
51
  # read it with jq if you vendor the baseline. An unverified download is an unverified agent:
48
52
  # this step FAILS rather than staging one, which is the whole point of naming it "verified".
@@ -73,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
73
77
  GitHub-hosted runners, no token/Docker/agent:
74
78
 
75
79
  ```yaml
76
- - run: npm i -g "cowork-harness@^3.7.0"
80
+ - run: npm i -g "cowork-harness@^3.8.1"
77
81
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
78
82
  # no silent false-greens. WITHOUT --strict this
79
83
  # step cannot fail on a WARN-class rule (e.g.
@@ -322,6 +326,16 @@ A typical skill repo runs four stages, fastest/cheapest first:
322
326
  literal only when your scenarios name a `fidelity:`; one still in the deprecation window prints one
323
327
  defaulted-fidelity notice per scenario.) A scenario that lints with only
324
328
  warnings can still be unloadable, so a green `lint` is not evidence the suite runs.
329
+
330
+ **If the repo pays for `critique`, gate the evidence corpus here first, for free:**
331
+
332
+ ```bash
333
+ cowork-harness critique <folder> [--skill <name>] --corpus-only --output-format json \
334
+ | jq -e '.corpus.corpusBytes <= .corpus.corpusCeiling'
335
+ ```
336
+
337
+ The `jq -e` IS the gate — `--corpus-only` exits 0 on a measurement even over the ceiling. Same
338
+ packager, same git filter as the paid run; the number is a floor (a run-time read can only add).
325
339
  3. **Scenarios (replay)** — `cowork-harness replay cassettes/` on every PR (the committed `*.cassette.json`).
326
340
  Token-free; content + structure + gate delivery.
327
341
  4. **Parity / live (nightly, self-hosted)** — `cowork-harness run scenarios/` with a token + Docker +
@@ -350,7 +364,7 @@ jobs:
350
364
  with: { node-version: '24' }
351
365
  - uses: actions/setup-python@v5
352
366
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
353
- - run: npm i -g "cowork-harness@^3.7.0"
367
+ - run: npm i -g "cowork-harness@^3.8.1"
354
368
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
355
369
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
356
370
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -379,7 +393,7 @@ jobs:
379
393
  echo "live=true" >> "$GITHUB_OUTPUT"
380
394
  fi
381
395
  - if: steps.guard.outputs.live == 'true'
382
- run: npm i -g "cowork-harness@^3.7.0"
396
+ run: npm i -g "cowork-harness@^3.8.1"
383
397
  - if: steps.guard.outputs.live == 'true'
384
398
  run: cowork-harness run scenarios/ --output-format json
385
399
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -22,9 +22,12 @@ finding's `evidence` excerpt must resolve verbatim against this package, or it l
22
22
 
23
23
  ## Cost across critiques — the index, not the reports
24
24
 
25
- A critique is FOUR model workloads but only TWO produce a run, so only two produce index rows; the
26
- evaluator passes produce none. Each critique therefore appends a **roll-up row** (`critiqueRole:"rollup"`)
27
- carrying `critiqueTotalUsd` — the whole four-workload spend. Its own `costUsd` is the **evaluator passes
25
+ A critique is **up to FOUR** model workloads — two graded turns and two evaluator passes — but only the
26
+ TWO graded turns produce a run, so only two produce index rows; the evaluator passes produce none. Each
27
+ critique therefore appends a **roll-up row** (`critiqueRole:"rollup"`) carrying `critiqueTotalUsd` — the
28
+ whole spend across whatever workloads actually ran. **Evaluator pass 2 is skipped entirely when no
29
+ self-report was captured** (nothing to verify), so a completed critique can be three workloads and the
30
+ roll-up covers three. Its own `costUsd` is the **evaluator passes
28
31
  only**, so `sum(costUsd)` over every row is exactly true spend with nothing double-counted or missed. The
29
32
  turn rows carry `critiqueRole:"task"` / `"reflection"`.
30
33
 
@@ -128,6 +131,24 @@ has no run to read, so its number can under-report there. It also does not apply
128
131
  so an untracked skill-local reference inflates the figure the other way; `corpusCuts`/`corpusOmitted`
129
132
  below stay the authority.
130
133
 
134
+ **Cheaper, and the packager's own floor: `critique <folder> --corpus-only`.** NO SPEND — no session, no spawn — it runs the
135
+ same `packageEvidence` call a paid critique makes, over an empty run dir, and prints these six fields
136
+ directly (`--output-format json` for the standard payload envelope; text mode prints the one-line
137
+ percent-of-ceiling summary + packaged-file count). `--prompt` becomes optional; every other flag is still
138
+ parsed and type-checked, but a run-shaping one is ignored and named in `ignoredFlags` (a path value such as
139
+ `--upload` is only checked when a turn stages). The number is a FLOOR — a plugin-root
140
+ reference the agent READS during the graded turn is added at critique time, so a paid run's `corpusBytes`
141
+ is `>=` this. Exit 0 = measured (even over the ceiling — gate yourself on `corpusBytes <= corpusCeiling`);
142
+ exit 2 = usage error, unresolvable target, no readable SKILL.md, or a work tree with 0 tracked files (a
143
+ non-git folder is measured raw, as staging copies it). Unlike `lint-skill`'s static count, it IS the
144
+ packager: untracked files and symlinks outside the plugin are excluded by the same filter and containment
145
+ rule a critique applies, bytes are the same UTF-8-decoded measurement, and the one clause no static
146
+ instrument can see (a plugin-root reference read at run time) is stated as the floor rather than guessed
147
+ at. Known gap: a skill that is a git submodule of its plugin (or any `--skill` subdirectory with nothing
148
+ tracked under it) is REFUSED by `--corpus-only` in staging's terms, but a live critique's packager still
149
+ accepts it from the directory's own index and grades a skill the mount never delivers — pre-existing,
150
+ rare, not fixed here.
151
+
131
152
  The report's `evidenceBudget` object says exactly what was shown — read it instead of inferring budgets
132
153
  from `dist/` source:
133
154
 
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.7.0`
4
- (baseline `desktop-2.2553.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.8.1`
4
+ (baseline `desktop-2.7032.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -192,7 +192,7 @@ plugins:
192
192
  skills:
193
193
  local: [] # extra host skill dirs
194
194
  suggest_enabled: true # gate 245679952 override — `mcp__skills__suggest_skills` on/off (default true)
195
- proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = synced baseline gate (ON from 1.24012.11)
195
+ proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = always on from the 1.46388.3 baseline (where `false` models a surface production does not ship there), synced gate before it
196
196
  mcp:
197
197
  config: null # --mcp-config file (standard mcpServers map)
198
198
  enabled: []
@@ -369,6 +369,8 @@ same set live from the schema.
369
369
  | `gate_answer_count_min: <N>` | at least N AskUserQuestion gates fired AND were delivered non-error — presence companion to `gate_answers_delivered`'s vacuous-pass. **`: 0` asserts nothing** and does not satisfy that pairing; `>= 1` is **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
370
370
  | `hook_blocked: <regex>` | a PreToolUse hook blocked a tool whose name matches the regex (`RunResult.hookEvents`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream) |
371
371
  | `no_hook_blocked: true` | no tool was hook-blocked during the run (distinguishes a real tool crash from an intentional hook block) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette. **Only `true` is valid** |
372
+ | `hook_event_fired: <HookEvent>` | a **command hook** for this event (a plugin's `hooks/hooks.json` or manifest hook — `Stop`, `SessionStart`, `PostToolUse`, …) ran: a `hook_response` system frame with that `hook_event` was recorded (`RunResult.contextEvents`). Any outcome counts. The harness passes `--include-hook-events` whenever a staged plugin declares hooks — that is what puts events other than SessionStart/Setup on the stream — so a recording made without it reports "never fired". Content-class, grades on replay. Recorded end-to-end for `Stop` ([stop-hook-probe.scenario.yaml](https://github.com/yaniv-golan/cowork-harness/blob/main/examples/probes/stop-hook-probe.scenario.yaml)); the other names match the same frame but have not each been recorded |
373
+ | `hook_event_blocked: <HookEvent>` | that command hook **blocked** at least once — a `hook_response` frame for the event carried `exit_code: 2`. Fails naming the exit codes seen when it fired without blocking (a frame with no `exit_code` is reported as such, never counted); fails "never fired" otherwise; cannot-verify when the run has no context events. Content-class |
372
374
  | `vm_path_denied: true` | **`fidelity: hostloop` only** — at least one recorded path denial (`RunResult.pathDenials`, any source) targeted a `/sessions` VM path — evidence-unavailable if path-denial telemetry is absent. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
373
375
  | `path_denied: {tool?, path_matches?, source?, agent_scope?}` | **`fidelity: hostloop` only** — a path denial matching ALL given matchers (`tool` glob, `path_matches` regex, `source` ∈ pretooluse/can_use_tool/permission_denied, `agent_scope` ∈ main/subagent/any) was recorded. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify" |
374
376
  | `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
@@ -454,7 +456,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
454
456
  `no_vm_path_file_op`, `dispatch_count_max`,
455
457
  `skill_triggered`, `no_skill_triggered`, `skill_available`, `connector_available`, `tool_available`,
456
458
  `skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
457
- `max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
459
+ `max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `hook_event_fired`, `hook_event_blocked`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
458
460
  (`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
459
461
  modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
460
462
  `allow_stall` are also kept on replay, evaluated as no-op passes.
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -23,6 +23,8 @@
23
23
  "gate_answer_count_min",
24
24
  "gate_answers_delivered",
25
25
  "hook_blocked",
26
+ "hook_event_blocked",
27
+ "hook_event_fired",
26
28
  "input_unmodified",
27
29
  "max_cost_usd",
28
30
  "max_peak_rss_bytes",
@@ -150,6 +152,7 @@
150
152
  "liveVerifiedHookEvents": [
151
153
  "PostToolUse",
152
154
  "SessionStart",
155
+ "Stop",
153
156
  "UserPromptSubmit"
154
157
  ],
155
158
  "enums": {
@@ -165,6 +168,76 @@
165
168
  "once",
166
169
  "domain"
167
170
  ],
171
+ "assert.hook_event_blocked": [
172
+ "PreToolUse",
173
+ "PostToolUse",
174
+ "PostToolUseFailure",
175
+ "PostToolBatch",
176
+ "Notification",
177
+ "UserPromptSubmit",
178
+ "UserPromptExpansion",
179
+ "SessionStart",
180
+ "SessionEnd",
181
+ "Stop",
182
+ "StopFailure",
183
+ "SubagentStart",
184
+ "SubagentStop",
185
+ "PreCompact",
186
+ "PostCompact",
187
+ "PreModelSwitch",
188
+ "PostModelSwitch",
189
+ "PermissionRequest",
190
+ "PermissionDenied",
191
+ "Setup",
192
+ "TeammateIdle",
193
+ "TaskCreated",
194
+ "TaskCompleted",
195
+ "Elicitation",
196
+ "ElicitationResult",
197
+ "ConfigChange",
198
+ "WorktreeCreate",
199
+ "WorktreeRemove",
200
+ "InstructionsLoaded",
201
+ "CwdChanged",
202
+ "FileChanged",
203
+ "DirectoryAdded",
204
+ "MessageDisplay"
205
+ ],
206
+ "assert.hook_event_fired": [
207
+ "PreToolUse",
208
+ "PostToolUse",
209
+ "PostToolUseFailure",
210
+ "PostToolBatch",
211
+ "Notification",
212
+ "UserPromptSubmit",
213
+ "UserPromptExpansion",
214
+ "SessionStart",
215
+ "SessionEnd",
216
+ "Stop",
217
+ "StopFailure",
218
+ "SubagentStart",
219
+ "SubagentStop",
220
+ "PreCompact",
221
+ "PostCompact",
222
+ "PreModelSwitch",
223
+ "PostModelSwitch",
224
+ "PermissionRequest",
225
+ "PermissionDenied",
226
+ "Setup",
227
+ "TeammateIdle",
228
+ "TaskCreated",
229
+ "TaskCompleted",
230
+ "Elicitation",
231
+ "ElicitationResult",
232
+ "ConfigChange",
233
+ "WorktreeCreate",
234
+ "WorktreeRemove",
235
+ "InstructionsLoaded",
236
+ "CwdChanged",
237
+ "FileChanged",
238
+ "DirectoryAdded",
239
+ "MessageDisplay"
240
+ ],
168
241
  "assert.path_denied.agent_scope": [
169
242
  "main",
170
243
  "subagent",
@@ -111,6 +111,8 @@ CONTENT_KEYS = {
111
111
  "max_redundant_tool_calls",
112
112
  "max_turns",
113
113
  "compaction_occurred",
114
+ "hook_event_fired",
115
+ "hook_event_blocked",
114
116
  "all_tasks_completed",
115
117
  "task_count_min",
116
118
  "task_status",
@@ -218,7 +220,9 @@ _FALLBACK_SERVED_HOOK_EVENTS = {"PreToolUse"}
218
220
  # Re-sourced 2026-09-06 from the agent's OWN hooks-config validator array (ELF 2.1.260), not from a grep
219
221
  # for event-name constants. The previous 9-name set reported the other 24 -- PostCompact and
220
222
  # MessageDisplay among them -- identically to a misspelling, at ERROR severity below.
221
- _FALLBACK_KNOWN_HOOK_EVENTS = {
223
+ # Ordered exactly like the TS `KNOWN_HOOK_EVENTS` array: the `hook_event_fired`/`hook_event_blocked` enum
224
+ # values in _EMBEDDED_ENUMS are derived from this list and must compare equal to the generated map.
225
+ _FALLBACK_KNOWN_HOOK_EVENTS_ORDERED = [
222
226
  "PreToolUse", "PostToolUse", "PostToolUseFailure", "PostToolBatch",
223
227
  "Notification", "UserPromptSubmit", "UserPromptExpansion", "SessionStart",
224
228
  "SessionEnd", "Stop", "StopFailure", "SubagentStart", "SubagentStop",
@@ -227,11 +231,12 @@ _FALLBACK_KNOWN_HOOK_EVENTS = {
227
231
  "TaskCreated", "TaskCompleted", "Elicitation", "ElicitationResult",
228
232
  "ConfigChange", "WorktreeCreate", "WorktreeRemove", "InstructionsLoaded",
229
233
  "CwdChanged", "FileChanged", "DirectoryAdded", "MessageDisplay",
230
- }
234
+ ]
235
+ _FALLBACK_KNOWN_HOOK_EVENTS = set(_FALLBACK_KNOWN_HOOK_EVENTS_ORDERED)
231
236
  # The subset a plugin hook has been OBSERVED to fire for here (live-verified 2026-08-01, container +
232
237
  # hostloop). Kept apart from the known set because the message wording depends on which claim we can
233
238
  # make: accepted-by-the-validator is not reached-by-a-run.
234
- _FALLBACK_LIVE_VERIFIED_HOOK_EVENTS = {"SessionStart", "UserPromptSubmit", "PostToolUse"}
239
+ _FALLBACK_LIVE_VERIFIED_HOOK_EVENTS = {"SessionStart", "UserPromptSubmit", "PostToolUse", "Stop"}
235
240
 
236
241
 
237
242
  def _load_hook_events():
@@ -350,6 +355,8 @@ _EMBEDDED_ENUMS = {
350
355
  "assert.path_denied.source": ["pretooluse", "can_use_tool", "permission_denied"],
351
356
  "assert.path_denied.agent_scope": ["main", "subagent", "any"],
352
357
  "assert.question_options.order": ["exact", "any"],
358
+ "assert.hook_event_fired": list(_FALLBACK_KNOWN_HOOK_EVENTS_ORDERED),
359
+ "assert.hook_event_blocked": list(_FALLBACK_KNOWN_HOOK_EVENTS_ORDERED),
353
360
  }
354
361
 
355
362
 
@@ -1569,14 +1576,15 @@ def _lint_hook_events(path):
1569
1576
  findings.append(Finding(
1570
1577
  "INFO", "hook-event-not-served",
1571
1578
  f"`{name}` {fires} — but cowork-harness "
1572
- f"itself installs only {', '.join(sorted(SERVED_HOOK_EVENTS))} on `initialize`. Two "
1573
- f"consequences: there is no assertion key for this event, so a scenario cannot GATE on it; "
1574
- f"and if real Cowork installs a `{name}` hook of its own, the harness does not reproduce it, "
1575
- f"so anything driven by that is absent here. (Cowork installs hooks of its own for "
1579
+ f"itself installs only {', '.join(sorted(SERVED_HOOK_EVENTS))} on `initialize`. "
1580
+ f"`hook_event_fired: {name}` / `hook_event_blocked: {name}` grade it from the agent's own "
1581
+ f"hook_response frames (the harness passes --include-hook-events because this plugin declares "
1582
+ f"hooks); but if real Cowork installs a `{name}` hook of its own, the harness does not reproduce "
1583
+ f"it, so anything driven by that is absent here. (Cowork installs hooks of its own for "
1576
1584
  f"PreToolUse, PostToolUse and UserPromptSubmit only.)",
1577
- "The harness does not block your hook — this is about assertability, not breakage. To gate "
1578
- "on its effect, assert the OBSERVABLE result instead (a file it writes, a tool it blocks), "
1579
- "not the hook itself.",
1585
+ "The harness does not block your hook — this is about what is reproduced, not breakage. Assert "
1586
+ "the hook with those keys, and its OBSERVABLE result as well (a file it writes, a tool it blocks) "
1587
+ "for anything Cowork's own hooks would have driven.",
1580
1588
  path, line_no,
1581
1589
  ))
1582
1590
  elif name.lower() in {e.lower() for e in KNOWN_HOOK_EVENTS}:
@@ -2269,9 +2277,16 @@ def _lint_skill_corpus_size(md_path):
2269
2277
  content the packager would cut. A proximity check that greens a corpus destined to be cut is worse
2270
2278
  than no check.
2271
2279
 
2272
- Still approximate in ONE direction only, and it now over- rather than under-counts: the packager
2273
- applies staging's git-tracked filter, so an untracked reference inflates this figure. That errs
2274
- toward warning early. The report's corpusCuts stays the authority."""
2280
+ It diverges from what a critique actually packages on four axes, two each way. OVER-counts: an untracked reference that staging would never deliver (the
2281
+ packager applies staging's git-tracked filter; this walk does not), and a symlink pointing outside
2282
+ the plugin, which the packager's containment rule refuses to follow. UNDER-counts: a plugin-root
2283
+ reference the graded agent only reaches by reading it during the run (added to the corpus at
2284
+ critique time -- invisible to any static count), and any byte that fails strict UTF-8 decoding, which
2285
+ the packager replaces with a 3-byte U+FFFD that st_size never sees (clean multibyte text round-trips
2286
+ byte-exact, so this axis is zero on ordinary markdown). `cowork-harness critique
2287
+ <folder> --corpus-only` runs the packager's own packageEvidence call over an empty run and prints the
2288
+ six corpus fields directly: the packager's own git filter, containment rule and byte measurement,
2289
+ and a stated FLOOR for the run-time-read clause (a read can only add to it)."""
2275
2290
  skill_dir = Path(md_path).parent
2276
2291
  total = 0
2277
2292
  files = [Path(md_path)]