cowork-harness 3.7.0 → 3.8.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +9 -7
- package/.claude/skills/cowork-harness/references/ci-recipe.md +25 -11
- package/.claude/skills/cowork-harness/references/critique.md +25 -4
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +6 -4
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +73 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +28 -13
- package/CHANGELOG.md +331 -0
- package/DESIGN.md +2 -2
- package/README.md +4 -4
- package/RELEASING.md +8 -0
- package/SPEC.md +11 -1
- package/baselines/desktop-2.7032.0.json +1049 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +7 -0
- package/baselines/prompts/desktop-1.46388.3/subagent-append-hl.md +4 -1
- package/baselines/provisioning/rootfs-provisioning.json +17 -14
- package/dist/agent/session.js +4 -2
- package/dist/assert.js +36 -0
- package/dist/baseline.js +14 -0
- package/dist/cli.js +1 -1
- package/dist/critique/command.js +340 -56
- package/dist/critique/limitations.js +9 -0
- package/dist/critique/package-evidence.js +3 -1
- package/dist/critique/skill-invocation.js +140 -0
- package/dist/hostloop/workspace-handler.js +10 -3
- package/dist/loop-decision.js +22 -1
- package/dist/prompt/subagent-manifest.js +10 -3
- package/dist/run/cassette.js +2 -0
- package/dist/run/hook-events.js +10 -7
- package/dist/run/skill-flag-surface.js +4 -1
- package/dist/run/tool-name-canonicalization.js +39 -2
- package/dist/run/verdict.js +3 -1
- package/dist/runtime/argv.js +4 -0
- package/dist/session.js +39 -16
- package/dist/sync/cowork-sync.js +82 -31
- package/dist/types.js +9 -0
- package/docs/cassette.md +4 -1
- package/docs/ci.md +17 -7
- package/docs/cli.md +10 -10
- package/docs/companion-skill.md +2 -2
- package/docs/critique.md +113 -9
- package/docs/fidelity-gaps.md +164 -40
- package/docs/gotchas.md +3 -1
- package/docs/maintenance.md +1 -1
- package/docs/scenario.md +7 -2
- package/docs/session.md +8 -6
- package/docs/subagents.md +1 -1
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +55 -62
- package/examples/replays/example-pdf-skill.cassette.json +122 -102
- package/examples/replays/hostloop-computer-links.cassette.json +62 -64
- package/examples/sessions/stop-hook-probe.yaml +5 -0
- package/llms.txt +1 -1
- package/package.json +1 -1
- package/schema/critique-report.json +9 -1
- package/schema/scenario.schema.json +78 -0
- package/schema/session.schema.json +1 -1
- package/scripts/capture-rootfs-manifest.ts +35 -0
- package/scripts/check-versions.ts +51 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.8.1
|
|
7
|
+
tracks-harness: cowork-harness 3.8.1 (baseline desktop-2.7032.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
29
|
-
> `desktop-2.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.8.1` (baseline
|
|
29
|
+
> `desktop-2.7032.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.8.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.8.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.8.1"`. **Pin `@^3.8.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -81,7 +81,7 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
|
|
|
81
81
|
correct answers after you edit it?) → author `semantic_matches` scenarios and gate on the per-claim
|
|
82
82
|
profile. See **Recipe 5** in `references/task-recipes.md` (validity, N≥3, discrimination — the traps).
|
|
83
83
|
- **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
|
|
84
|
-
<folder> --prompt "<probe>"`.
|
|
84
|
+
<folder> --prompt "<probe>"`. Up to four model workloads (zero with `--corpus-only`; pass 2 is skipped with no self-report) and 10–20 minutes; budget from
|
|
85
85
|
`report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
|
|
86
86
|
**DO**" — routing, artifact location, narration — use `skill` instead**: no evaluator, a fraction of
|
|
87
87
|
the cost, and it answers that question directly. Report and evidence-package shapes:
|
|
@@ -568,7 +568,9 @@ Recognize these before "fixing" a non-bug:
|
|
|
568
568
|
- **`missing_capability`** — the lean `core` agent image is a deliberate partial mirror of real Cowork's
|
|
569
569
|
rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
|
|
570
570
|
`markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
|
|
571
|
-
(`magick`) can trip this even though real Cowork **ships** those
|
|
571
|
+
(`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
|
|
572
|
+
Desktop `2.7032.0` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
|
|
573
|
+
behind that sentence). The message says so ("likely a FALSE
|
|
572
574
|
NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
|
|
573
575
|
`COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
|
|
574
576
|
`allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.8.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,13 +36,17 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.280 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
|
|
41
|
-
# served
|
|
42
|
-
# 2.1.255
|
|
43
|
-
#
|
|
44
|
-
#
|
|
45
|
-
B
|
|
41
|
+
# served from .../claude-code-releases/rc/<commit>/. For some versions the stable path 404s
|
|
42
|
+
# (2.1.255); for others it returns 200 and serves a DIFFERENT BUILD UNDER THE SAME VERSION
|
|
43
|
+
# NUMBER — measured for 2.1.280 on 2026-09-23: stable linux-arm64 233,103,352 B / 92f2b4fd…
|
|
44
|
+
# (commit 80abbfe7) vs RC 233,037,816 B / a1b25d70… (commit bddba3ab), and the staged binary is
|
|
45
|
+
# the RC one. So "the stable URL works" is NOT evidence you have the right build: always take B
|
|
46
|
+
# from your pinned baseline's agentBinary.releaseBaseUrl. The checksum step fails closed if you
|
|
47
|
+
# don't, but it cannot tell you why. Baselines written before that field existed were
|
|
48
|
+
# stable-staged, so their base is the plain https://downloads.claude.ai/claude-code-releases.
|
|
49
|
+
B=https://downloads.claude.ai/claude-code-releases/rc/bddba3abd5da53d0c540cfc76a8d18b44633d568
|
|
46
50
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
47
51
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
48
52
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -73,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
73
77
|
GitHub-hosted runners, no token/Docker/agent:
|
|
74
78
|
|
|
75
79
|
```yaml
|
|
76
|
-
- run: npm i -g "cowork-harness@^3.
|
|
80
|
+
- run: npm i -g "cowork-harness@^3.8.1"
|
|
77
81
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
78
82
|
# no silent false-greens. WITHOUT --strict this
|
|
79
83
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -322,6 +326,16 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
322
326
|
literal only when your scenarios name a `fidelity:`; one still in the deprecation window prints one
|
|
323
327
|
defaulted-fidelity notice per scenario.) A scenario that lints with only
|
|
324
328
|
warnings can still be unloadable, so a green `lint` is not evidence the suite runs.
|
|
329
|
+
|
|
330
|
+
**If the repo pays for `critique`, gate the evidence corpus here first, for free:**
|
|
331
|
+
|
|
332
|
+
```bash
|
|
333
|
+
cowork-harness critique <folder> [--skill <name>] --corpus-only --output-format json \
|
|
334
|
+
| jq -e '.corpus.corpusBytes <= .corpus.corpusCeiling'
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
The `jq -e` IS the gate — `--corpus-only` exits 0 on a measurement even over the ceiling. Same
|
|
338
|
+
packager, same git filter as the paid run; the number is a floor (a run-time read can only add).
|
|
325
339
|
3. **Scenarios (replay)** — `cowork-harness replay cassettes/` on every PR (the committed `*.cassette.json`).
|
|
326
340
|
Token-free; content + structure + gate delivery.
|
|
327
341
|
4. **Parity / live (nightly, self-hosted)** — `cowork-harness run scenarios/` with a token + Docker +
|
|
@@ -350,7 +364,7 @@ jobs:
|
|
|
350
364
|
with: { node-version: '24' }
|
|
351
365
|
- uses: actions/setup-python@v5
|
|
352
366
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
353
|
-
- run: npm i -g "cowork-harness@^3.
|
|
367
|
+
- run: npm i -g "cowork-harness@^3.8.1"
|
|
354
368
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
355
369
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
356
370
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -379,7 +393,7 @@ jobs:
|
|
|
379
393
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
380
394
|
fi
|
|
381
395
|
- if: steps.guard.outputs.live == 'true'
|
|
382
|
-
run: npm i -g "cowork-harness@^3.
|
|
396
|
+
run: npm i -g "cowork-harness@^3.8.1"
|
|
383
397
|
- if: steps.guard.outputs.live == 'true'
|
|
384
398
|
run: cowork-harness run scenarios/ --output-format json
|
|
385
399
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -22,9 +22,12 @@ finding's `evidence` excerpt must resolve verbatim against this package, or it l
|
|
|
22
22
|
|
|
23
23
|
## Cost across critiques — the index, not the reports
|
|
24
24
|
|
|
25
|
-
A critique is FOUR model workloads
|
|
26
|
-
|
|
27
|
-
|
|
25
|
+
A critique is **up to FOUR** model workloads — two graded turns and two evaluator passes — but only the
|
|
26
|
+
TWO graded turns produce a run, so only two produce index rows; the evaluator passes produce none. Each
|
|
27
|
+
critique therefore appends a **roll-up row** (`critiqueRole:"rollup"`) carrying `critiqueTotalUsd` — the
|
|
28
|
+
whole spend across whatever workloads actually ran. **Evaluator pass 2 is skipped entirely when no
|
|
29
|
+
self-report was captured** (nothing to verify), so a completed critique can be three workloads and the
|
|
30
|
+
roll-up covers three. Its own `costUsd` is the **evaluator passes
|
|
28
31
|
only**, so `sum(costUsd)` over every row is exactly true spend with nothing double-counted or missed. The
|
|
29
32
|
turn rows carry `critiqueRole:"task"` / `"reflection"`.
|
|
30
33
|
|
|
@@ -128,6 +131,24 @@ has no run to read, so its number can under-report there. It also does not apply
|
|
|
128
131
|
so an untracked skill-local reference inflates the figure the other way; `corpusCuts`/`corpusOmitted`
|
|
129
132
|
below stay the authority.
|
|
130
133
|
|
|
134
|
+
**Cheaper, and the packager's own floor: `critique <folder> --corpus-only`.** NO SPEND — no session, no spawn — it runs the
|
|
135
|
+
same `packageEvidence` call a paid critique makes, over an empty run dir, and prints these six fields
|
|
136
|
+
directly (`--output-format json` for the standard payload envelope; text mode prints the one-line
|
|
137
|
+
percent-of-ceiling summary + packaged-file count). `--prompt` becomes optional; every other flag is still
|
|
138
|
+
parsed and type-checked, but a run-shaping one is ignored and named in `ignoredFlags` (a path value such as
|
|
139
|
+
`--upload` is only checked when a turn stages). The number is a FLOOR — a plugin-root
|
|
140
|
+
reference the agent READS during the graded turn is added at critique time, so a paid run's `corpusBytes`
|
|
141
|
+
is `>=` this. Exit 0 = measured (even over the ceiling — gate yourself on `corpusBytes <= corpusCeiling`);
|
|
142
|
+
exit 2 = usage error, unresolvable target, no readable SKILL.md, or a work tree with 0 tracked files (a
|
|
143
|
+
non-git folder is measured raw, as staging copies it). Unlike `lint-skill`'s static count, it IS the
|
|
144
|
+
packager: untracked files and symlinks outside the plugin are excluded by the same filter and containment
|
|
145
|
+
rule a critique applies, bytes are the same UTF-8-decoded measurement, and the one clause no static
|
|
146
|
+
instrument can see (a plugin-root reference read at run time) is stated as the floor rather than guessed
|
|
147
|
+
at. Known gap: a skill that is a git submodule of its plugin (or any `--skill` subdirectory with nothing
|
|
148
|
+
tracked under it) is REFUSED by `--corpus-only` in staging's terms, but a live critique's packager still
|
|
149
|
+
accepts it from the directory's own index and grades a skill the mount never delivers — pre-existing,
|
|
150
|
+
rare, not fixed here.
|
|
151
|
+
|
|
131
152
|
The report's `evidenceBudget` object says exactly what was shown — read it instead of inferring budgets
|
|
132
153
|
from `dist/` source:
|
|
133
154
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
4
|
-
(baseline `desktop-2.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.8.1`
|
|
4
|
+
(baseline `desktop-2.7032.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -192,7 +192,7 @@ plugins:
|
|
|
192
192
|
skills:
|
|
193
193
|
local: [] # extra host skill dirs
|
|
194
194
|
suggest_enabled: true # gate 245679952 override — `mcp__skills__suggest_skills` on/off (default true)
|
|
195
|
-
proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset =
|
|
195
|
+
proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = always on from the 1.46388.3 baseline (where `false` models a surface production does not ship there), synced gate before it
|
|
196
196
|
mcp:
|
|
197
197
|
config: null # --mcp-config file (standard mcpServers map)
|
|
198
198
|
enabled: []
|
|
@@ -369,6 +369,8 @@ same set live from the schema.
|
|
|
369
369
|
| `gate_answer_count_min: <N>` | at least N AskUserQuestion gates fired AND were delivered non-error — presence companion to `gate_answers_delivered`'s vacuous-pass. **`: 0` asserts nothing** and does not satisfy that pairing; `>= 1` is **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
|
|
370
370
|
| `hook_blocked: <regex>` | a PreToolUse hook blocked a tool whose name matches the regex (`RunResult.hookEvents`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream) |
|
|
371
371
|
| `no_hook_blocked: true` | no tool was hook-blocked during the run (distinguishes a real tool crash from an intentional hook block) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette. **Only `true` is valid** |
|
|
372
|
+
| `hook_event_fired: <HookEvent>` | a **command hook** for this event (a plugin's `hooks/hooks.json` or manifest hook — `Stop`, `SessionStart`, `PostToolUse`, …) ran: a `hook_response` system frame with that `hook_event` was recorded (`RunResult.contextEvents`). Any outcome counts. The harness passes `--include-hook-events` whenever a staged plugin declares hooks — that is what puts events other than SessionStart/Setup on the stream — so a recording made without it reports "never fired". Content-class, grades on replay. Recorded end-to-end for `Stop` ([stop-hook-probe.scenario.yaml](https://github.com/yaniv-golan/cowork-harness/blob/main/examples/probes/stop-hook-probe.scenario.yaml)); the other names match the same frame but have not each been recorded |
|
|
373
|
+
| `hook_event_blocked: <HookEvent>` | that command hook **blocked** at least once — a `hook_response` frame for the event carried `exit_code: 2`. Fails naming the exit codes seen when it fired without blocking (a frame with no `exit_code` is reported as such, never counted); fails "never fired" otherwise; cannot-verify when the run has no context events. Content-class |
|
|
372
374
|
| `vm_path_denied: true` | **`fidelity: hostloop` only** — at least one recorded path denial (`RunResult.pathDenials`, any source) targeted a `/sessions` VM path — evidence-unavailable if path-denial telemetry is absent. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
|
|
373
375
|
| `path_denied: {tool?, path_matches?, source?, agent_scope?}` | **`fidelity: hostloop` only** — a path denial matching ALL given matchers (`tool` glob, `path_matches` regex, `source` ∈ pretooluse/can_use_tool/permission_denied, `agent_scope` ∈ main/subagent/any) was recorded. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify" |
|
|
374
376
|
| `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
|
|
@@ -454,7 +456,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
|
|
|
454
456
|
`no_vm_path_file_op`, `dispatch_count_max`,
|
|
455
457
|
`skill_triggered`, `no_skill_triggered`, `skill_available`, `connector_available`, `tool_available`,
|
|
456
458
|
`skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
|
|
457
|
-
`max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
|
|
459
|
+
`max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `hook_event_fired`, `hook_event_blocked`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
|
|
458
460
|
(`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
|
|
459
461
|
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
|
|
460
462
|
`allow_stall` are also kept on replay, evaluated as no-op passes.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.8.1` (baseline `desktop-2.7032.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -23,6 +23,8 @@
|
|
|
23
23
|
"gate_answer_count_min",
|
|
24
24
|
"gate_answers_delivered",
|
|
25
25
|
"hook_blocked",
|
|
26
|
+
"hook_event_blocked",
|
|
27
|
+
"hook_event_fired",
|
|
26
28
|
"input_unmodified",
|
|
27
29
|
"max_cost_usd",
|
|
28
30
|
"max_peak_rss_bytes",
|
|
@@ -150,6 +152,7 @@
|
|
|
150
152
|
"liveVerifiedHookEvents": [
|
|
151
153
|
"PostToolUse",
|
|
152
154
|
"SessionStart",
|
|
155
|
+
"Stop",
|
|
153
156
|
"UserPromptSubmit"
|
|
154
157
|
],
|
|
155
158
|
"enums": {
|
|
@@ -165,6 +168,76 @@
|
|
|
165
168
|
"once",
|
|
166
169
|
"domain"
|
|
167
170
|
],
|
|
171
|
+
"assert.hook_event_blocked": [
|
|
172
|
+
"PreToolUse",
|
|
173
|
+
"PostToolUse",
|
|
174
|
+
"PostToolUseFailure",
|
|
175
|
+
"PostToolBatch",
|
|
176
|
+
"Notification",
|
|
177
|
+
"UserPromptSubmit",
|
|
178
|
+
"UserPromptExpansion",
|
|
179
|
+
"SessionStart",
|
|
180
|
+
"SessionEnd",
|
|
181
|
+
"Stop",
|
|
182
|
+
"StopFailure",
|
|
183
|
+
"SubagentStart",
|
|
184
|
+
"SubagentStop",
|
|
185
|
+
"PreCompact",
|
|
186
|
+
"PostCompact",
|
|
187
|
+
"PreModelSwitch",
|
|
188
|
+
"PostModelSwitch",
|
|
189
|
+
"PermissionRequest",
|
|
190
|
+
"PermissionDenied",
|
|
191
|
+
"Setup",
|
|
192
|
+
"TeammateIdle",
|
|
193
|
+
"TaskCreated",
|
|
194
|
+
"TaskCompleted",
|
|
195
|
+
"Elicitation",
|
|
196
|
+
"ElicitationResult",
|
|
197
|
+
"ConfigChange",
|
|
198
|
+
"WorktreeCreate",
|
|
199
|
+
"WorktreeRemove",
|
|
200
|
+
"InstructionsLoaded",
|
|
201
|
+
"CwdChanged",
|
|
202
|
+
"FileChanged",
|
|
203
|
+
"DirectoryAdded",
|
|
204
|
+
"MessageDisplay"
|
|
205
|
+
],
|
|
206
|
+
"assert.hook_event_fired": [
|
|
207
|
+
"PreToolUse",
|
|
208
|
+
"PostToolUse",
|
|
209
|
+
"PostToolUseFailure",
|
|
210
|
+
"PostToolBatch",
|
|
211
|
+
"Notification",
|
|
212
|
+
"UserPromptSubmit",
|
|
213
|
+
"UserPromptExpansion",
|
|
214
|
+
"SessionStart",
|
|
215
|
+
"SessionEnd",
|
|
216
|
+
"Stop",
|
|
217
|
+
"StopFailure",
|
|
218
|
+
"SubagentStart",
|
|
219
|
+
"SubagentStop",
|
|
220
|
+
"PreCompact",
|
|
221
|
+
"PostCompact",
|
|
222
|
+
"PreModelSwitch",
|
|
223
|
+
"PostModelSwitch",
|
|
224
|
+
"PermissionRequest",
|
|
225
|
+
"PermissionDenied",
|
|
226
|
+
"Setup",
|
|
227
|
+
"TeammateIdle",
|
|
228
|
+
"TaskCreated",
|
|
229
|
+
"TaskCompleted",
|
|
230
|
+
"Elicitation",
|
|
231
|
+
"ElicitationResult",
|
|
232
|
+
"ConfigChange",
|
|
233
|
+
"WorktreeCreate",
|
|
234
|
+
"WorktreeRemove",
|
|
235
|
+
"InstructionsLoaded",
|
|
236
|
+
"CwdChanged",
|
|
237
|
+
"FileChanged",
|
|
238
|
+
"DirectoryAdded",
|
|
239
|
+
"MessageDisplay"
|
|
240
|
+
],
|
|
168
241
|
"assert.path_denied.agent_scope": [
|
|
169
242
|
"main",
|
|
170
243
|
"subagent",
|
|
@@ -111,6 +111,8 @@ CONTENT_KEYS = {
|
|
|
111
111
|
"max_redundant_tool_calls",
|
|
112
112
|
"max_turns",
|
|
113
113
|
"compaction_occurred",
|
|
114
|
+
"hook_event_fired",
|
|
115
|
+
"hook_event_blocked",
|
|
114
116
|
"all_tasks_completed",
|
|
115
117
|
"task_count_min",
|
|
116
118
|
"task_status",
|
|
@@ -218,7 +220,9 @@ _FALLBACK_SERVED_HOOK_EVENTS = {"PreToolUse"}
|
|
|
218
220
|
# Re-sourced 2026-09-06 from the agent's OWN hooks-config validator array (ELF 2.1.260), not from a grep
|
|
219
221
|
# for event-name constants. The previous 9-name set reported the other 24 -- PostCompact and
|
|
220
222
|
# MessageDisplay among them -- identically to a misspelling, at ERROR severity below.
|
|
221
|
-
|
|
223
|
+
# Ordered exactly like the TS `KNOWN_HOOK_EVENTS` array: the `hook_event_fired`/`hook_event_blocked` enum
|
|
224
|
+
# values in _EMBEDDED_ENUMS are derived from this list and must compare equal to the generated map.
|
|
225
|
+
_FALLBACK_KNOWN_HOOK_EVENTS_ORDERED = [
|
|
222
226
|
"PreToolUse", "PostToolUse", "PostToolUseFailure", "PostToolBatch",
|
|
223
227
|
"Notification", "UserPromptSubmit", "UserPromptExpansion", "SessionStart",
|
|
224
228
|
"SessionEnd", "Stop", "StopFailure", "SubagentStart", "SubagentStop",
|
|
@@ -227,11 +231,12 @@ _FALLBACK_KNOWN_HOOK_EVENTS = {
|
|
|
227
231
|
"TaskCreated", "TaskCompleted", "Elicitation", "ElicitationResult",
|
|
228
232
|
"ConfigChange", "WorktreeCreate", "WorktreeRemove", "InstructionsLoaded",
|
|
229
233
|
"CwdChanged", "FileChanged", "DirectoryAdded", "MessageDisplay",
|
|
230
|
-
|
|
234
|
+
]
|
|
235
|
+
_FALLBACK_KNOWN_HOOK_EVENTS = set(_FALLBACK_KNOWN_HOOK_EVENTS_ORDERED)
|
|
231
236
|
# The subset a plugin hook has been OBSERVED to fire for here (live-verified 2026-08-01, container +
|
|
232
237
|
# hostloop). Kept apart from the known set because the message wording depends on which claim we can
|
|
233
238
|
# make: accepted-by-the-validator is not reached-by-a-run.
|
|
234
|
-
_FALLBACK_LIVE_VERIFIED_HOOK_EVENTS = {"SessionStart", "UserPromptSubmit", "PostToolUse"}
|
|
239
|
+
_FALLBACK_LIVE_VERIFIED_HOOK_EVENTS = {"SessionStart", "UserPromptSubmit", "PostToolUse", "Stop"}
|
|
235
240
|
|
|
236
241
|
|
|
237
242
|
def _load_hook_events():
|
|
@@ -350,6 +355,8 @@ _EMBEDDED_ENUMS = {
|
|
|
350
355
|
"assert.path_denied.source": ["pretooluse", "can_use_tool", "permission_denied"],
|
|
351
356
|
"assert.path_denied.agent_scope": ["main", "subagent", "any"],
|
|
352
357
|
"assert.question_options.order": ["exact", "any"],
|
|
358
|
+
"assert.hook_event_fired": list(_FALLBACK_KNOWN_HOOK_EVENTS_ORDERED),
|
|
359
|
+
"assert.hook_event_blocked": list(_FALLBACK_KNOWN_HOOK_EVENTS_ORDERED),
|
|
353
360
|
}
|
|
354
361
|
|
|
355
362
|
|
|
@@ -1569,14 +1576,15 @@ def _lint_hook_events(path):
|
|
|
1569
1576
|
findings.append(Finding(
|
|
1570
1577
|
"INFO", "hook-event-not-served",
|
|
1571
1578
|
f"`{name}` {fires} — but cowork-harness "
|
|
1572
|
-
f"itself installs only {', '.join(sorted(SERVED_HOOK_EVENTS))} on `initialize`.
|
|
1573
|
-
f"
|
|
1574
|
-
f"
|
|
1575
|
-
f"
|
|
1579
|
+
f"itself installs only {', '.join(sorted(SERVED_HOOK_EVENTS))} on `initialize`. "
|
|
1580
|
+
f"`hook_event_fired: {name}` / `hook_event_blocked: {name}` grade it from the agent's own "
|
|
1581
|
+
f"hook_response frames (the harness passes --include-hook-events because this plugin declares "
|
|
1582
|
+
f"hooks); but if real Cowork installs a `{name}` hook of its own, the harness does not reproduce "
|
|
1583
|
+
f"it, so anything driven by that is absent here. (Cowork installs hooks of its own for "
|
|
1576
1584
|
f"PreToolUse, PostToolUse and UserPromptSubmit only.)",
|
|
1577
|
-
"The harness does not block your hook — this is about
|
|
1578
|
-
"
|
|
1579
|
-
"
|
|
1585
|
+
"The harness does not block your hook — this is about what is reproduced, not breakage. Assert "
|
|
1586
|
+
"the hook with those keys, and its OBSERVABLE result as well (a file it writes, a tool it blocks) "
|
|
1587
|
+
"for anything Cowork's own hooks would have driven.",
|
|
1580
1588
|
path, line_no,
|
|
1581
1589
|
))
|
|
1582
1590
|
elif name.lower() in {e.lower() for e in KNOWN_HOOK_EVENTS}:
|
|
@@ -2269,9 +2277,16 @@ def _lint_skill_corpus_size(md_path):
|
|
|
2269
2277
|
content the packager would cut. A proximity check that greens a corpus destined to be cut is worse
|
|
2270
2278
|
than no check.
|
|
2271
2279
|
|
|
2272
|
-
|
|
2273
|
-
applies staging's git-tracked filter
|
|
2274
|
-
|
|
2280
|
+
It diverges from what a critique actually packages on four axes, two each way. OVER-counts: an untracked reference that staging would never deliver (the
|
|
2281
|
+
packager applies staging's git-tracked filter; this walk does not), and a symlink pointing outside
|
|
2282
|
+
the plugin, which the packager's containment rule refuses to follow. UNDER-counts: a plugin-root
|
|
2283
|
+
reference the graded agent only reaches by reading it during the run (added to the corpus at
|
|
2284
|
+
critique time -- invisible to any static count), and any byte that fails strict UTF-8 decoding, which
|
|
2285
|
+
the packager replaces with a 3-byte U+FFFD that st_size never sees (clean multibyte text round-trips
|
|
2286
|
+
byte-exact, so this axis is zero on ordinary markdown). `cowork-harness critique
|
|
2287
|
+
<folder> --corpus-only` runs the packager's own packageEvidence call over an empty run and prints the
|
|
2288
|
+
six corpus fields directly: the packager's own git filter, containment rule and byte measurement,
|
|
2289
|
+
and a stated FLOOR for the run-time-read clause (a read can only add to it)."""
|
|
2275
2290
|
skill_dir = Path(md_path).parent
|
|
2276
2291
|
total = 0
|
|
2277
2292
|
files = [Path(md_path)]
|