cowork-harness 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (80) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +20 -11
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +45 -24
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +2 -2
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +4 -4
  7. package/.claude/skills/cowork-harness/scripts/scenario.py +80 -7
  8. package/CHANGELOG.md +596 -2
  9. package/DESIGN.md +8 -5
  10. package/README.md +44 -22
  11. package/RELEASING.md +35 -9
  12. package/SECURITY.md +10 -8
  13. package/SPEC.md +30 -9
  14. package/dist/answer-policy.js +42 -0
  15. package/dist/cli.js +10 -16
  16. package/dist/redact.js +44 -0
  17. package/dist/run/budget.js +21 -6
  18. package/dist/run/cassette.js +94 -17
  19. package/dist/session.js +17 -10
  20. package/dist/types.js +11 -11
  21. package/docs/boundary.md +2 -2
  22. package/docs/cassette.md +52 -16
  23. package/docs/decider-dir.md +16 -2
  24. package/docs/discovery.md +3 -1
  25. package/docs/fidelity-gaps.md +26 -6
  26. package/docs/invariants.md +2 -2
  27. package/docs/maintenance.md +2 -2
  28. package/docs/protocol.md +29 -8
  29. package/docs/run-status.md +2 -2
  30. package/docs/scenario.md +50 -14
  31. package/docs/session.md +11 -6
  32. package/docs/subagents.md +13 -6
  33. package/examples/README.md +9 -6
  34. package/examples/data/mcp.json +10 -0
  35. package/examples/data/project/notes.txt +1 -0
  36. package/examples/data/report.pdf +54 -0
  37. package/examples/data/sales.csv +6 -0
  38. package/examples/data/sales_eur.csv +6 -0
  39. package/examples/replays/README.md +1 -1
  40. package/examples/replays/example-pdf-skill.cassette.json +12 -12
  41. package/examples/scenarios/csv-fx-normalize.yaml +40 -0
  42. package/examples/scenarios/csv-metrics.yaml +41 -0
  43. package/examples/scenarios/example-pdf-skill.yaml +38 -0
  44. package/examples/scenarios/hostloop-computer-links.yaml +43 -0
  45. package/examples/scenarios/protocol-smoke.yaml +46 -0
  46. package/examples/scenarios/skill-loads.yaml +18 -0
  47. package/examples/scenarios/trigger-accuracy-sweep/negative-unrelated-request.yaml +18 -0
  48. package/examples/scenarios/trigger-accuracy-sweep/positive-clear-pdf-request.yaml +29 -0
  49. package/examples/sessions/csv-fx-normalize.yaml +15 -0
  50. package/examples/sessions/csv-metrics.yaml +15 -0
  51. package/examples/sessions/default.yaml +54 -0
  52. package/examples/sessions/hostloop-computer-links.yaml +9 -0
  53. package/examples/sessions/protocol-smoke.yaml +10 -0
  54. package/examples/sessions/skill.yaml +10 -0
  55. package/examples/skills/csv-fx-normalize/.claude-plugin/plugin.json +1 -0
  56. package/examples/skills/csv-fx-normalize/skills/csv-fx-normalize/SKILL.md +40 -0
  57. package/examples/skills/csv-fx-normalize/skills/csv-fx-normalize/scripts/normalize.py +118 -0
  58. package/examples/skills/csv-metrics/.claude-plugin/plugin.json +1 -0
  59. package/examples/skills/csv-metrics/skills/csv-metrics/SKILL.md +39 -0
  60. package/examples/skills/csv-metrics/skills/csv-metrics/scripts/metrics.py +143 -0
  61. package/examples/skills/my-pdf-skill/.claude-plugin/plugin.json +1 -0
  62. package/examples/skills/my-pdf-skill/skills/my-pdf-skill/SKILL.md +6 -0
  63. package/fixtures/protocol/v1/dialog-response.json +10 -0
  64. package/fixtures/protocol/v1/elicit-response.json +10 -0
  65. package/fixtures/protocol/v1/elicitation-request.json +17 -0
  66. package/fixtures/protocol/v1/error-response.json +8 -0
  67. package/fixtures/protocol/v1/user-dialog-request.json +11 -0
  68. package/llms.txt +1 -1
  69. package/package.json +6 -2
  70. package/python/README.md +12 -0
  71. package/python/conftest.py +15 -3
  72. package/python/cowork_harness.py +44 -0
  73. package/python/test_lane_optin.py +117 -0
  74. package/python/test_scenario_lint.py +69 -0
  75. package/schema/cassette.v12.json +2 -2
  76. package/schema/protocol.v1.json +447 -79
  77. package/schema/scenario.schema.json +1 -1
  78. package/scripts/bump-version.ts +6 -5
  79. package/scripts/check-versions.ts +332 -10
  80. package/scripts/gen-schema.ts +13 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.0.0
7
- tracks-harness: cowork-harness 2.0.0 (baseline desktop-1.34493.1)
6
+ version: 2.1.0
7
+ tracks-harness: cowork-harness 2.1.0 (baseline desktop-1.34493.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.0.0` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.1.0` (baseline
26
26
  > `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=2.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=2.0.0"`. **Pin `@>=2.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.1.0"`. **Pin `@^2.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
44
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
45
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -220,9 +220,14 @@ run's echoed `--answer "<q>=<choice>"` footer lines into the scenario's `answers
220
220
  being unattended. Skip the transcribe step only for one-off/exploratory runs.
221
221
  <!-- answer-channels:end -->
222
222
 
223
- **Never hand-write the `req-N.json`/`resp-N.json` files.** `gates` and `answer` wrap the protocol — the
224
- atomic temp+rename, the `{id, answers}` envelope, the multiSelect array shape. Hand-rolling a Monitor over
225
- the raw files is the single most common mistake on this channel.
223
+ **For a QUESTION gate, never hand-write the `req-N.json`/`resp-N.json` files.** `gates` and `answer` wrap
224
+ the protocol — the atomic temp+rename, the `{id, answers}` envelope, the multiSelect array shape.
225
+ Hand-rolling a Monitor over the raw files is the single most common mistake on this channel.
226
+
227
+ `answer` writes `{id, answers}` and nothing else, so it covers **question gates only**. The channel also
228
+ carries **permission**, **dialog** and **elicit** gates, whose replies need `{behavior}` / `{action}` — for
229
+ those, write `resp-N.json` yourself, following the `reply_with` template the gate's own `req-N.json`
230
+ advertises (it spells out the exact shape, e.g. `{"id":"…","behavior":"allow|deny"}`).
226
231
 
227
232
  Exact accepted values (teach precisely): `--on-unanswered` takes `fail|prompt|first` on `skill`,
228
233
  only `fail|first` on `run`. **`llm` is NOT an `--on-unanswered` value** — the bare flag
@@ -416,7 +421,8 @@ missing was anything saying so while you could still act. Related: recording at
416
421
  (`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25).
417
422
  The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
418
423
  redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
419
- for a permanently unverifiable cassette.
424
+ for a cassette that cannot verify staleness from its own location — recoverable only by passing
425
+ `--session <file>` on every invocation thereafter.
420
426
 
421
427
  **Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
422
428
  scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
@@ -650,7 +656,7 @@ than a stuck `"running"`.)
650
656
  ### Place assertions in the right CI lane
651
657
 
652
658
  CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
653
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@main` (a packaged GitHub Action with a
659
+ (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v2` (a packaged GitHub Action with a
654
660
  PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
655
661
  the four-stage pipeline.
656
662
 
@@ -917,8 +923,11 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
917
923
  `artifact_json` / `transcript_matches`), never just `result: success`.
918
924
  - `on_unanswered` governs **unanswered** `AskUserQuestion` gates; the `stalled` signal covers
919
925
  stalling *after* one is answered — two different failure modes.
920
- - **Free-text aside:** a "type-it-in-notes" option has **no scripted deterministic answer** today
921
- (the `OTHER:` directive works only on the LLM-decider path, not scripted `choose:`, and only on
926
+ - **Free-text aside:** the scripted key for a "type-it-in-notes" option is **`answer:`** — an
927
+ arbitrary string delivered verbatim, bypassing label validation by author intent (Cowork
928
+ auto-provides an "Other" free-text path on every gate). Mutually exclusive with `choose:`; setting
929
+ both fails loud. What has no scripted equivalent is the `OTHER:` *directive*
930
+ (it works only on the LLM-decider path, not scripted `choose:`, and only on
922
931
  **single-select** gates — a **multi-select** gate is index-only, so `OTHER:` fails loud there; on an
923
932
  options-bearing single-select gate a bare out-of-set LLM answer also fails loud (exit 2) — see the
924
933
  LLM-decider free-text note in `references/fidelity-and-answers.md`). An LLM decision answered via
@@ -1,19 +1,23 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
7
7
 
8
8
  ```yaml
9
- - uses: yaniv-golan/cowork-harness@main
9
+ - uses: yaniv-golan/cowork-harness@v2
10
10
  with:
11
11
  command: replay
12
12
  path: cassettes/
13
+ version: "^2" # hold the major; see below
13
14
  ```
14
15
 
15
- The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "2.0.0"`) for reproducible CI.
16
+ **These recipes pin `version: "^2"`.** The Action's `version` input *defaults* to `latest`, which means a
17
+ CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
+ copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
+ patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
+ (e.g. `version: "2.1.0"`) instead when you want byte-reproducible CI.
17
21
 
18
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -30,19 +34,24 @@ jobs:
30
34
  runs-on: [self-hosted, linux, arm64] # needs Docker + this staged ELF; not a stock GitHub-hosted runner
31
35
  steps:
32
36
  - uses: actions/checkout@v4
33
- - name: Stage the agent binary (official channel, sha256-verified — see https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md)
37
+ - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
34
38
  run: |
35
39
  V=2.1.237 # match your scenario's pinned baseline's agentVersion
40
+ # The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
41
+ # read it with jq if you vendor the baseline. An unverified download is an unverified agent:
42
+ # this step FAILS rather than staging one, which is the whole point of naming it "verified".
43
+ EXPECTED=<paste agentBinary.sha256 for $V>
36
44
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
45
+ echo "$EXPECTED $RUNNER_TEMP/claude-$V" | sha256sum -c -
37
46
  chmod +x "$RUNNER_TEMP/claude-$V"
38
- # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
39
- # before trusting it — see the "Agent-binary provenance" section of
40
- # https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
41
47
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
42
- - uses: yaniv-golan/cowork-harness@main
48
+ # Background on the provenance chain: the "Agent-binary provenance" section of
49
+ # https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
50
+ - uses: yaniv-golan/cowork-harness@v2
43
51
  with:
44
52
  command: run
45
53
  path: scenarios/
54
+ version: "^2"
46
55
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
47
56
  ```
48
57
 
@@ -58,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
58
67
  GitHub-hosted runners, no token/Docker/agent:
59
68
 
60
69
  ```yaml
61
- - run: npm i -g "cowork-harness@>=2.0.0"
70
+ - run: npm i -g "cowork-harness@^2.1.0"
62
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
63
72
  # no silent false-greens. WITHOUT --strict this
64
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -106,20 +115,28 @@ Action has no input for, and it creates a coupling nothing checks:
106
115
 
107
116
  > **If a flag in `extra-args` was added in release X, floor that step's `version` to `>=X`.**
108
117
 
109
- `version` defaults to `latest`, and accepts **any npm range** — not just an exact pin. So:
118
+ `version` defaults to `latest`, and accepts **any npm range** — not just an exact pin. Leave it off unless
119
+ you have a reason:
110
120
 
111
121
  ```yaml
112
- - uses: yaniv-golan/cowork-harness@main
122
+ - uses: yaniv-golan/cowork-harness@v2
113
123
  with:
114
124
  command: lint
115
125
  path: scenarios/
116
- version: ">=1.11.0" # --min-severity landed in 1.11.0
117
- extra-args: --min-severity WARN
126
+ version: "^2" # holds the major
127
+ extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 2.x satisfies that
118
128
  ```
119
129
 
120
- Without the floor, an older CLI fails the step with `unrecognized arguments: --min-severity WARN` (exit 2,
121
- wrapped in an `ok:false` envelope) — it does **not** degrade gracefully. An exact pin is fine for
122
- reproducibility, but it rots the moment a recipe adopts a newer flag; a floor expresses the real dependency.
130
+ **If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
131
+ floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
132
+ recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
133
+ major instead — `version: "^2"`, which is what the steps above use — keeping the floor's intent while
134
+ stopping at the major boundary. An exact
135
+ pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
136
+ moment a recipe adopts a newer flag.
137
+
138
+ Without a satisfied floor, an older CLI fails the step with `unrecognized arguments: --min-severity WARN`
139
+ (exit 2, wrapped in an `ok:false` envelope) — it does **not** degrade gracefully.
123
140
 
124
141
 
125
142
  The harness has two execution lanes with different cost, coverage, AND infrastructure requirements.
@@ -303,7 +320,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
303
320
 
304
321
  ## GitHub Actions sketch
305
322
 
306
- The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@main` does
323
+ The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v2` does
307
324
  in one step (see the top of this doc) — reach for this form when you need independent per-command
308
325
  gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
309
326
  equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
@@ -323,7 +340,7 @@ jobs:
323
340
  with: { node-version: '24' }
324
341
  - uses: actions/setup-python@v5
325
342
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
326
- - run: npm i -g "cowork-harness@>=2.0.0"
343
+ - run: npm i -g "cowork-harness@^2.1.0"
327
344
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
328
345
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
329
346
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -352,7 +369,7 @@ jobs:
352
369
  echo "live=true" >> "$GITHUB_OUTPUT"
353
370
  fi
354
371
  - if: steps.guard.outputs.live == 'true'
355
- run: npm i -g "cowork-harness@>=2.0.0"
372
+ run: npm i -g "cowork-harness@^2.1.0"
356
373
  - if: steps.guard.outputs.live == 'true'
357
374
  run: cowork-harness run scenarios/ --output-format json
358
375
  env:
@@ -374,10 +391,14 @@ sandbox).
374
391
  ## Reading results in CI
375
392
 
376
393
  `--output-format json` emits a machine envelope on stdout (human output goes to stderr):
377
- `{tool, version, command, ok, results[], error}` — one `RunResult` per scenario. Overall pass for a
378
- scenario = `result === "success" && assertions.every(pass)`. Exit code is non-zero if any assertion
379
- fails or a run errors, so a plain `cowork-harness run scenarios/` is already CI-ready without parsing
380
- JSON.
394
+ `{tool, version, command, ok, results[], error}` — one `RunResult` per scenario. **Overall pass for a
395
+ scenario is `verdict.pass`** (envelope-wide: `ok`), and it is strictly stronger than
396
+ `result === "success" && assertions.every(pass)`: the verdict also carries ~20 signal codes that fail a run
397
+ with no failing assertion at all — `stalled`, `outputs_delete`, `mount_delete`, `host_path_leak`,
398
+ `undelivered_deliverables`, `missing_capability`, `permissive_auto_allow`, `ended_with_question`,
399
+ `infra_error`, and more. A parser that reimplements the shorter formula greens through every one of them.
400
+ Read `ok` / `verdict.pass`, or just use the exit code — a plain `cowork-harness run scenarios/` is already
401
+ CI-ready without parsing JSON.
381
402
 
382
403
  **Telling *why* a run failed, without scraping stderr.** Each result carries a `verdict` whose
383
404
  `failures[]` is the one place every failure reason is enumerated, in one shape (the same object lands
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -210,7 +210,7 @@ up often enough to spell out:
210
210
  - **A green `replay` proves "same as when recorded," not "correct today."** `replay` never touches
211
211
  a filesystem or network — it re-evaluates assertions from the frozen cassette. A fixed set of
212
212
  keys is live-only and **skipped outright** on replay (absent from `assertions[]`, not vacuously
213
- passed): `file_absent`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
213
+ passed): `file_absent`, `no_delete_in_outputs`, `no_delete_in_mounts`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
214
214
  `egress_allowed`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back`, and `expect_denied`.
215
215
  Everything else that *is* evaluated is checked against the **recording**, not fresh behavior — a
216
216
  green replay says the skill produced these events when it was recorded, not that it still does
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.0.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.1.0`
4
4
  (baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -49,11 +49,11 @@ what production would do. Two layers of defense:
49
49
 
50
50
  ### Cassette anatomy (what you're looking at when you open one)
51
51
 
52
- Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v11.json`](https://github.com/yaniv-golan/cowork-harness/blob/main/schema/cassette.v11.json)):
52
+ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](https://github.com/yaniv-golan/cowork-harness/blob/main/schema/cassette.v12.json)):
53
53
 
54
54
  | Field | What it is |
55
55
  |---|---|
56
- | `$schema`, `generator`, `cassetteVersion` | Provenance: schema URL, producing tool, format version — the MINIMUM a reader needs for this scenario, not the recorder's version (current max: 11; a `lane: remote` scenario stamps 11, nearly everything else stays 10) |
56
+ | `$schema`, `generator`, `cassetteVersion` | Provenance: schema URL, producing tool, format version — the MINIMUM a reader needs for this scenario, not the recorder's version (current max: 12 — the hash-format epoch floors every stamp there, so a fresh recording stamps 12 whatever its `lane:`) |
57
57
  | `scenario` | The embedded scenario snapshot at record time |
58
58
  | `events` | The recorded agent event stream (the replay source) |
59
59
  | `controlOut` | Driver→agent control responses — presence unlocks gate asserts on replay |
@@ -66,7 +66,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v11.json`](htt
66
66
  | `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
67
67
  | `scenarioSource` | Relative path to the authored YAML this was recorded from |
68
68
  | `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
69
- | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
69
+ | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
70
70
  | `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
71
71
  | `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
72
72
  | `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
@@ -409,6 +409,67 @@ def _is_positional_choose(choose):
409
409
  return any(isinstance(v, str) and (v == "first" or v.isdigit()) for v in vals)
410
410
 
411
411
 
412
+ # Single-segment absolute paths that legitimately appear in prompt prose. A `/word` from this set is a
413
+ # path, not a slash command, so it never raises the not-leading warning below.
414
+ _SLASH_PATH_WORDS = frozenset(
415
+ {"outputs", "mnt", "tmp", "home", "users", "var", "etc", "usr", "bin", "dev", "opt", "workspace", "root", "srv"}
416
+ )
417
+ # A `/`-prefixed token at a word start (start-of-string or whitespace), captured WHOLE up to the next
418
+ # space. An opening bracket or quote also counts as a word start, so "(/deck-review)" is seen; `and/or`,
419
+ # `8/22` and `https://x` still cannot match, because their slash follows a letter, digit or colon. The
420
+ # token is then classified in Python rather than by a lookahead: an earlier lookahead-based pattern
421
+ # BACKTRACKED, matching `/mn` inside `/mnt/uploads` because a shorter prefix satisfied the lookahead.
422
+ _SLASH_TOKEN_RE = re.compile(r"""(?:^|[\s(\[{"'])/(\S+)""")
423
+ # What a slash command may look like once trailing sentence punctuation is stripped: a bare name, or a
424
+ # plugin-qualified `plugin:skill`. Anchored, so any residual `/` or `.` disqualifies it as a path/filename.
425
+ _SLASH_CMD_NAME_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_-]*(?::[A-Za-z][A-Za-z0-9_-]*)?$")
426
+
427
+
428
+ def _lint_prompt_slash(doc, path):
429
+ """W: `prompt:` names a slash command somewhere other than position 0.
430
+
431
+ The agent binary resolves a slash command only when the TRIMMED prompt starts with `/` — its parser
432
+ trims, then requires `startsWith("/")` (verified against agent 2.1.239 in the harness's own spawn
433
+ shape: `-p --input-format stream-json --output-format stream-json --setting-sources user`). A slash
434
+ named mid-sentence is never expanded. It reaches the model as ordinary prose, and the model may then
435
+ reach for the `Skill` tool on its own — the model-invocation path, i.e. exactly the unreliable
436
+ auto-trigger a slash is normally used to bypass. The scenario still runs and can still pass, so the
437
+ failure mode is a scenario that silently tests something other than what it reads as.
438
+
439
+ Deliberately silent when the prompt DOES start with `/`: that is the working case. Registration is not
440
+ checkable statically (it depends on how the skill is staged), so an unresolvable leading name is left
441
+ to the run itself, where it shows up as `Unknown command: /x` with `num_turns: 0`.
442
+ """
443
+ findings = []
444
+ prompt = doc.get("prompt")
445
+ if not isinstance(prompt, str) or prompt.lstrip().startswith("/"):
446
+ return findings
447
+ seen = []
448
+ for m in _SLASH_TOKEN_RE.finditer(prompt):
449
+ # Trailing sentence punctuation is not part of the name, so "use /deck-review." still counts.
450
+ name = m.group(1).rstrip(".,;:!?)]}\"'")
451
+ if not _SLASH_CMD_NAME_RE.match(name):
452
+ continue # a path (`/mnt/uploads`), a filename (`/deck.pdf`), or not command-shaped
453
+ if name.lower() in _SLASH_PATH_WORDS or name in seen:
454
+ continue
455
+ seen.append(name)
456
+ for name in seen:
457
+ findings.append(
458
+ Finding(
459
+ "WARN",
460
+ "prompt-slash-not-leading",
461
+ f"`prompt:` names `/{name}` but does not START with it. A slash command is expanded only "
462
+ "when the trimmed prompt begins with `/`, so here it reaches the model as ordinary prose "
463
+ "and the skill is NOT preloaded — the model may or may not reach for it on its own, which "
464
+ "is the auto-trigger path a slash is normally used to bypass.",
465
+ f'Put the command first — `prompt: "/{name} <args>"` — or drop the slash if the scenario '
466
+ "means to test auto-triggering from a natural request.",
467
+ path,
468
+ )
469
+ )
470
+ return findings
471
+
472
+
412
473
  def lint_doc(doc, path, raw_lines):
413
474
  findings = []
414
475
  if not isinstance(doc, dict):
@@ -456,6 +517,9 @@ def lint_doc(doc, path, raw_lines):
456
517
  )
457
518
  )
458
519
 
520
+ # W: a slash command named mid-prompt is never expanded — see _lint_prompt_slash.
521
+ findings.extend(_lint_prompt_slash(doc, path))
522
+
459
523
  # W: unknown assertion keys inside assert items (e.g. invented file_not_empty, kind, path)
460
524
  unknown_assert = sorted(assert_keys - ASSERT_KEYS)
461
525
  for k in unknown_assert:
@@ -1834,11 +1898,16 @@ def build_scenario(args):
1834
1898
  )
1835
1899
  content_lines.append(" - gate_answers_delivered: true # the steered answers actually reached the model")
1836
1900
 
1837
- live_lines = []
1901
+ # Two buckets, not one. `file_exists`/`user_visible_artifact` are MANIFEST_KEYS above — they DO
1902
+ # evaluate on replay whenever the cassette carries an artifacts manifest, which `record` has
1903
+ # snapshotted since 0.24. Filing them under a "LIVE-only" heading taught the reader the opposite of
1904
+ # what this file's own taxonomy says, and of what the `manifest-needs-snapshot` INFO tells them.
1905
+ manifest_lines = []
1838
1906
  for p in (args.file or []):
1839
- live_lines.append(f" - file_exists: {p}")
1907
+ manifest_lines.append(f" - file_exists: {p}")
1840
1908
  for p in (args.artifact or []):
1841
- live_lines.append(f" - user_visible_artifact: {p}")
1909
+ manifest_lines.append(f" - user_visible_artifact: {p}")
1910
+ live_lines = []
1842
1911
  if args.no_delete:
1843
1912
  live_lines.append(" - no_delete_in_outputs: true")
1844
1913
  for h in (args.egress_allowed or []):
@@ -1850,12 +1919,16 @@ def build_scenario(args):
1850
1919
  L.append("assert:")
1851
1920
  L.append(" # --- content / structure: evaluate on the token-free replay PR gate AND live ---")
1852
1921
  L.extend(content_lines)
1922
+ if manifest_lines:
1923
+ L.append(" # --- artifacts: replay-checkable WHEN the cassette carries an artifacts manifest ---")
1924
+ L.extend(manifest_lines)
1853
1925
  if live_lines:
1854
- L.append(" # --- filesystem / egress: LIVE-only (skipped on replay, with a loud warning) ---")
1926
+ L.append(" # --- egress / deletes: LIVE-only (skipped on replay, with a loud warning) ---")
1855
1927
  L.extend(live_lines)
1856
- else:
1857
- L.append(" # TODO add filesystem/egress checks (file_exists / user_visible_artifact /")
1858
- L.append(" # egress_denied / no_delete_in_outputs) — they run on the LIVE lane only.")
1928
+ if not manifest_lines and not live_lines:
1929
+ L.append(" # TODO add artifact checks (file_exists / user_visible_artifact — these replay from")
1930
+ L.append(" # the cassette's artifacts manifest) and filesystem/egress checks")
1931
+ L.append(" # (egress_denied / no_delete_in_outputs — LIVE lane only).")
1859
1932
 
1860
1933
  if args.web_fetch:
1861
1934
  notes.append(