cowork-harness 4.0.0 → 4.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +11 -14
  2. package/.claude/skills/cowork-harness/references/assertion-catalog.md +3 -3
  3. package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
  4. package/.claude/skills/cowork-harness/references/authoring.md +64 -14
  5. package/.claude/skills/cowork-harness/references/ci-recipe.md +41 -11
  6. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  7. package/.claude/skills/cowork-harness/references/debugging.md +16 -3
  8. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +5 -3
  9. package/.claude/skills/cowork-harness/references/gotchas.md +16 -11
  10. package/.claude/skills/cowork-harness/references/measurement.md +5 -2
  11. package/.claude/skills/cowork-harness/references/run-record-replay.md +18 -6
  12. package/.claude/skills/cowork-harness/references/scenario-schema.md +19 -6
  13. package/.claude/skills/cowork-harness/references/task-recipes.md +9 -4
  14. package/.claude/skills/cowork-harness/scripts/scenario.py +638 -52
  15. package/.cowork-redact.json +3 -3
  16. package/CHANGELOG.md +234 -3
  17. package/README.md +3 -3
  18. package/RELEASING.md +4 -0
  19. package/SPEC.md +32 -11
  20. package/baselines/prompts/cowork-system-prompt-fingerprints.json +90 -0
  21. package/dist/baseline.js +32 -1
  22. package/dist/cli.js +49 -14
  23. package/dist/errors.js +41 -0
  24. package/dist/redactable-literal.js +2 -0
  25. package/dist/run/cassette.js +111 -35
  26. package/dist/run/chat.js +6 -2
  27. package/dist/run/doctor.js +49 -11
  28. package/dist/run/execute.js +255 -68
  29. package/dist/run/host-path-tokens.js +73 -0
  30. package/dist/run/input-host-paths.js +138 -0
  31. package/dist/run/lint-load.js +12 -5
  32. package/dist/run/scenario-tool.js +8 -3
  33. package/dist/run/verdict.js +5 -0
  34. package/dist/runtime/container.js +4 -0
  35. package/dist/runtime/image-capabilities.js +21 -2
  36. package/dist/runtime/lima.js +122 -3
  37. package/dist/runtime/microvm.js +6 -0
  38. package/dist/scan.js +50 -2
  39. package/dist/session.js +5 -0
  40. package/dist/types.js +1 -1
  41. package/docs/cassette.md +7 -1
  42. package/docs/ci.md +13 -1
  43. package/docs/cli.md +19 -14
  44. package/docs/companion-skill.md +2 -2
  45. package/docs/debugging.md +5 -0
  46. package/docs/gotchas.md +5 -4
  47. package/docs/plugin-root.md +69 -29
  48. package/docs/scenario.md +40 -9
  49. package/docs/session.md +1 -1
  50. package/docs/subagents.md +3 -3
  51. package/examples/replays/README.md +1 -1
  52. package/examples/scenarios/csv-fx-normalize.yaml +1 -1
  53. package/examples/scenarios/csv-metrics.yaml +1 -1
  54. package/examples/skills/csv-fx-normalize/skills/csv-fx-normalize/SKILL.md +8 -1
  55. package/examples/skills/csv-metrics/skills/csv-metrics/SKILL.md +8 -1
  56. package/package.json +2 -2
  57. package/python/test_scenario_lint.py +244 -0
  58. package/schema/run-result.json +10 -0
  59. package/schema/scenario.schema.json +1 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 4.0.0
7
- tracks-harness: cowork-harness 4.0.0 (baseline desktop-2.9939.4)
6
+ version: 4.1.1
7
+ tracks-harness: cowork-harness 4.1.1 (baseline desktop-2.9939.4)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
26
26
  full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
27
  Read them.
28
28
 
29
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.0.0` (baseline
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.1.1` (baseline
30
30
  > `desktop-2.9939.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
31
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
32
32
 
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
43
43
 
44
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
45
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
46
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.0.0"`. **Pin `@^4.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.1.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.1.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.1.1"`. **Pin `@^4.1.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
47
47
 
48
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
49
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -67,14 +67,11 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
67
67
  no scenario file.
68
68
  - **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
69
69
  This is the CI-grade path and most of this skill.
70
- - **A run failed — or greened and you don't trust it** (the debugging loop) → don't re-run and hope.
71
- The run already wrote its evidence to a **kept run dir** (`~/.cowork-harness/runs/…`; `--keep` prints
72
- the path, `trace <run-id>` finds it). **Localize the failure post-hoc** from that evidence:
73
- `cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
74
- then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
75
- the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
76
- output* sections in [`references/debugging.md`](references/debugging.md) are the detail (the fuller human-facing map lives in
77
- [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only, not shipped with the installed skill).
70
+ - **A run failed — or greened and you don't trust it** (the debugging loop) → **Read
71
+ [`references/debugging.md`](references/debugging.md) before touching the run.** Its *Triage* table first
72
+ splits "the skill misbehaved" from "a green you don't trust" — they need different tools — then names the
73
+ tool for each, all token-free over the **kept run dir** (`~/.cowork-harness/runs/…`; `--keep` prints the
74
+ path, `trace <run-id>` finds it). Don't re-run and hope.
78
75
  **"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
79
76
  grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
80
77
  surface; see `references/critique.md`.
@@ -136,8 +133,8 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
136
133
  (a cassette cannot be moved) → `replay` on the PR gate.
137
134
  - **Fix answers without paying:** `--keep` one run → `trace <run-dir> --view questions` → edit
138
135
  `answers:` → `verify-run <run-dir> <scenario.yaml>` → record once.
139
- - **Debug:** the triage table in `references/debugging.md` → `inspect`, `trace --view …`,
140
- `verify-run`, `diff`. For a green you don't trust: `replay --explain`, then the gotchas.
136
+ - **Debug:** Read `references/debugging.md` first — its triage table picks the tool (`inspect`,
137
+ `trace --view …`, `verify-run`, `diff`, `replay --explain`).
141
138
  - **Measure:** `--repeat N` for flakiness. `--ablate-skill` runs the control arm only; run the treatment
142
139
  arm yourself, with the model pinned and a recoverable source frozen.
143
140
 
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). Every `assert:` key with its semantics, and the
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Every `assert:` key with its semantics, and the
4
4
  verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
5
5
  the scenario and session YAML fields are there too.
6
6
 
@@ -102,7 +102,7 @@ same set live from the schema.
102
102
  | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
103
103
  | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Silences `outputs_delete`, `outputs_delete_unconfirmed` and `outputs_diff_unavailable`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
104
104
  | `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
105
- | `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
105
+ | `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text (a path that came verbatim from the scenario's input files or prompt is exempt; `scan.hostPathsFromInputs` counts such paths) — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
106
106
  | `egress_denied: <host>` | the host was blocked by the egress proxy |
107
107
  | `egress_allowed: <host>` | the host was allowed through |
108
108
  | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
@@ -140,7 +140,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
140
140
  | `outputs_delete_unconfirmed` | warn | A delete-shaped command near `mnt/outputs` nothing confirms: no output present at turn start was deleted and no flagged delete has an outputs path as its own operand (e.g. a Python variable named `rm` next to an outputs path, quoted prose, a `sed`/`grep` pattern). A *statement* is a fragment split on newline/`;`/`&&`/`\|\|`, quote-blind, so some real deletes also land here — a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`), a `cd` then a relative path, chained variables (`A=…; B=$A/x; rm "$B"`), a Python path held in a variable set on another line (`p = …` then `os.remove(p)`, or `for p in …:` then `p.unlink()`), wrappers with flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`), and calls outside the modelled set such as Node's `fs.promises.rm(…)` — read the command. Raised even when `no_delete_in_outputs` is authored (that assertion passes on it). Kept false positives that still fail `outputs_delete`: quoted text in which a delete command with an outputs operand follows a shell separator, subshell or keyword — the classifier does not track quotes (`echo 'note; rm mnt/outputs/x'`, `echo "a & rm …/outputs/x"`), and a heredoc that *writes* a script rather than running it (`cat <<EOF > clean.sh` with an `rm …/outputs/x` line). Waive: `allow_outputs_delete` |
141
141
  | `outputs_diff_unavailable` | warn | The per-turn filesystem diff of `outputs/` could not verify this turn and the text scan flagged nothing — a delete made without a bash command would have gone undetected. With a text hit the turn fails `outputs_delete` instead |
142
142
  | `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
143
- | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
143
+ | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`). A path copied verbatim from the scenario's own uploads, connected folders or prompt is not a leak |
144
144
  | `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
145
145
  | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
146
146
  | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
@@ -1,6 +1,6 @@
1
1
  # Assertions guide
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
4
4
 
5
5
  ### Assertions: two orthogonal axes
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Authoring a scenario
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
4
4
 
5
5
  ## Part I — AUTHOR a scenario
6
6
 
@@ -220,25 +220,75 @@ so a file `run`/`record` would refuse fails lint too (running `scenario.py lint`
220
220
  cowork-harness lint scenarios/*.yaml
221
221
  ```
222
222
 
223
- `lint` flags: filesystem/egress-only assertions on a `replay` gate (silent no-op), bad regex
224
- quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `hostloop`/`protocol`
225
- (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
226
- baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
227
- `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
228
- unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
229
- `present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
230
- to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
231
- `protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
232
- and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
233
- (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
234
- emits a scenario `lint` would reject.
223
+ `lint` exits non-zero on any ERROR (CI-friendly); `--strict` also fails on WARN. Every rule it reports:
224
+
225
+ | Rule | Severity | Fires on |
226
+ |---|---|---|
227
+ | `assert-contradiction` | ERROR | assert items no single run can satisfy together |
228
+ | `assertions-key` | ERROR | `assertions:` instead of `assert:` — none of the checks would run |
229
+ | `authored-replay-fidelity` | ERROR | an authored `replay_protocol_fidelity` (only the replay lane synthesizes it) |
230
+ | `capabilities-on-protocol` | ERROR | non-empty `requires_capabilities` on `protocol` without `allow_missing_capability` — the probe cannot run there, so the run fails as unverifiable |
231
+ | `cassette-evidence-skipped` | INFO | with `--cassette-dir`: a cassette (or the directory) could not be read, so it cannot quiet replay-evidence advice |
232
+ | `container-only-key-off-container` | ERROR | `no_scratchpad_leak` off `container` — hostloop's `present_files` never promotes, so there is nothing to leak (WARN on `cowork`, whose tier resolves per the baseline gate) |
233
+ | `egress-on-protocol` | ERROR | an egress assertion (`egress_*` / `expect_denied`) on `protocol`, which enforces no egress |
234
+ | `enum-value-invalid` | ERROR | a field value outside its allowed set |
235
+ | `fidelity-missing` | ERROR | no `fidelity:` (required since 4.0.0) |
236
+ | `file-absent-contradiction` | ERROR | one path under both `file_exists` and `file_absent` |
237
+ | `gate-needs-controlout` | INFO | gate assertions, which evaluate on replay only when the cassette has `controlOut` |
238
+ | `host-path-assert-cowork` | WARN | `transcript_no_host_path` on `cowork` — it fails by design if the tier resolves to hostloop |
239
+ | `host-path-assert-tier` | ERROR | `transcript_no_host_path` on `hostloop` / `protocol`, where it fails by design |
240
+ | `lane-remote-incompatible-key` | ERROR | `present_files_called` / `no_scratchpad_leak` / `user_visible_artifact` on `lane: remote` (the runtime rejects them at load, so the tier rules are suppressed there) |
241
+ | `linter-extra-findings-invalid` | ERROR | the loader findings `cowork-harness lint` hands the linter could not be read |
242
+ | `linter-unclassified-key` | ERROR | a valid assertion key this linter cannot classify (the linter is out of date) |
243
+ | `manifest-needs-snapshot` | INFO | manifest-backed keys, which evaluate on replay only when the cassette carries an `artifacts` manifest |
244
+ | `mixed-assert-item` | WARN | one assert item mixing replay-checkable and live-only keys (replay drops the live-only half) |
245
+ | `no-scenarios` | ERROR | a linted directory with no `*.yaml` / `*.yml` |
246
+ | `not-found` | ERROR | a named file that does not exist |
247
+ | `parse` | ERROR | a file that is not YAML, or not a mapping |
248
+ | `positional-choose-order` | INFO | an answer rule with a positional `choose` (first / index), which option re-ordering can move |
249
+ | `present-files-key-off-tier` | ERROR | `present_files_called` on `protocol` / `microvm` (served only at `container` / `hostloop`) |
250
+ | `prompt-slash-not-leading` | WARN | a `prompt:` that names `/<skill>` without starting with it, so it is never expanded |
251
+ | `reference-access-contradiction` | ERROR | one reference under both `reference_read` and `no_observed_reference_access` |
252
+ | `regex-double-quoted` | WARN | a double-quoted regex with an unescaped backslash (YAML strips it) |
253
+ | `replay-noop` | WARN | every assertion is live-only or a verdict modifier, so a replay gate verifies nothing |
254
+ | `tool-called-always-passes` | INFO | `tool_called` with `count: {min: 0}` and no `max` — it asserts nothing |
255
+ | `tool-input-regex-redactable` | WARN | a `tool_not_called` input literal the redaction policy rewrites in the committed cassette (or a policy pattern it cannot check offline) |
256
+ | `tool-input-shell-tier` | INFO | the object form with `tool: Bash` and a `command` on `hostloop` / `cowork`, where shell runs as `mcp__workspace__bash` — list both |
257
+ | `tool-not-called-tier-vacuous` | WARN | `tool_not_called` / `subagent_tool_absent` naming a tool the tier never serves |
258
+ | `transcript-command-shaped` | WARN | a `transcript_*` value shaped like a shell command — those keys read prose only, never a tool call |
259
+ | `unknown-assert-key` | WARN | an assertion key not in the catalog (the loader rejects it) |
260
+ | `unknown-top-key` | WARN | a scenario key not in the schema |
261
+ | `vacuous-gate-assert` | WARN | `gate_answers_delivered` with no presence companion (zero gates passes it), or inert beside `questions_count_max: 0` |
262
+ | `scenario-invalid` | ERROR | the harness's scenario loader refuses the file (via `cowork-harness lint` only — see below) |
263
+ | `baseline-unknown` | ERROR | `baseline:` names no baseline this CLI ships (via `cowork-harness lint` only) |
264
+ | `lint-loader-internal` | ERROR | the wrapper could not run its loader check on a file — a harness bug; it never falls back to a lint that skipped the loader |
265
+
266
+ `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never emits a scenario `lint`
267
+ would reject.
268
+
269
+ **Lint the skill itself: `cowork-harness lint-skill <skill-dir>`.** It checks the skill, not a scenario:
270
+ Cowork host-loop footguns (`${CLAUDE_PLUGIN_ROOT}` in a VM bash step, hook events, a misplaced
271
+ `hooks.json`, an unresolvable `subagent_type`), the evidence corpus a `critique` can package
272
+ (`references/critique.md`), and two size caps. `skill-body-over-reattach-cap` (WARN) fires when the
273
+ `SKILL.md` body, frontmatter excluded, passes 19,000 B — after a compaction the agent re-attaches only the
274
+ start of an invoked skill — and `skill-body-near-reattach-cap` (INFO) from 80% of that;
275
+ `skill-reference-over-read-cap` (WARN) fires on a `references/**.md` over 60,000 B, past which a
276
+ whole-file Read returns a partial view. `--strict` fails on WARN, never on INFO. To accept a reviewed
277
+ judgement-call finding, pass `--ignore-rule <rule>[=<glob>]` (repeatable; the glob matches the finding's
278
+ file) or fence the text in `SKILL.md` with `<!-- lint-skill: ignore-start <rule>[,<rule>…]: <reason> -->`
279
+ … `<!-- lint-skill: ignore-end -->` (outside any code fence). A suppressed finding is still printed; it
280
+ stops gating. A provable rule (an ERROR, a misplaced `hooks.json`, a missing pinned agent) cannot be
281
+ suppressed: naming it, or an unknown rule, in `--ignore-rule` is a usage error (exit 2); in a marker it is
282
+ WARN `lint-skill-ignore-invalid`, as is any other malformed marker. An unclosed marker is WARN
283
+ `lint-skill-ignore-unclosed`, and one that suppresses nothing is INFO `lint-skill-ignore-unused`.
235
284
 
236
285
  **`cowork-harness lint` runs the loader: a file it calls clean is one `run`/`record` will load.** Anything
237
286
  the loader refuses — an unknown key, a wrong value type (a scalar `semantic_matches.rubric`), a bad regex,
238
287
  a reserved value — is ✗ ERROR `scenario-invalid` (exit 1, with or without `--strict`), and a `baseline:`
239
288
  naming no baseline this installed CLI ships is ✗ ERROR `baseline-unknown` (`latest` always resolves). It
240
289
  does not check what depends on the machine the run happens on (the session file and its mounts, an
241
- absolute `baseline:` path, environment variables). A session or matrix YAML in a linted directory is not
290
+ absolute `baseline:` path that does not exist here — one that exists is checked (4.1.1 and later) —
291
+ environment variables). A session or matrix YAML in a linted directory is not
242
292
  a scenario and is reported as one that does not load — keep those out of the linted set. `python3
243
293
  scenario.py lint` run directly stays offline and lenient: there an unknown key is only a ⚠ WARN (exit 0).
244
294
  `cowork-harness record <file.yaml> --dry-run` also runs the loader and adds the pre-spend refusals (exit 2
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`).
3
+ Self-contained reference. Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "4.0.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "4.1.1"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
82
82
  GitHub-hosted runners, no token/Docker/agent:
83
83
 
84
84
  ```yaml
85
- - run: npm i -g "cowork-harness@^4.0.0"
85
+ - run: npm i -g "cowork-harness@^4.1.1"
86
86
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
87
87
  # no silent false-greens. WITHOUT --strict this
88
88
  # step cannot fail on a WARN-class rule (e.g.
@@ -93,7 +93,16 @@ GitHub-hosted runners, no token/Docker/agent:
93
93
  # 3.x CLI bare `--strict` also fails on INFO, so keep
94
94
  # the pair explicit. `--min-severity INFO` gates on
95
95
  # INFO as well. (`lint-skill --strict` never fails
96
- # on INFO and has no floor to widen.)
96
+ # on INFO and has no floor to widen; to accept a
97
+ # reviewed lint-skill WARN, keep --strict and add
98
+ # `--ignore-rule <rule>[=<skill-dir>/<path>]` or an
99
+ # ignore-start/ignore-end marker in the SKILL.md.)
100
+ - run: cowork-harness lint scenarios/*.yaml --strict --min-severity INFO --cassette-dir cassettes/
101
+ # committed *.cassette.json evidence suppresses only
102
+ # replay advisories proven by every matching cassette;
103
+ # skipped or malformed cassettes stay visible.
104
+ # Staleness does not affect this existence check;
105
+ # verify-cassettes checks freshness and drift.
97
106
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness — FAILS on a stale recording
98
107
  # ALSO fails on a leaked host inventory: recording at
99
108
  # protocol/hostloop freezes YOUR machine's MCP servers,
@@ -232,7 +241,9 @@ cowork-harness replay cassettes/ # replay every *.cassette.jso
232
241
  Re-record whenever the protocol or your scenario's expected content changes. An old cassette without
233
242
  `controlOut` excludes the gate keys (with a loud warning) — re-record to enable them. `record` **refuses
234
243
  to freeze a failing live run** into a cassette (pass `--allow-failing` to override) — a committed red
235
- cassette is a latent false-signal.
244
+ cassette is a latent false-signal. An `--allow-failing` recording of a red run exits 0 with `ok: true`:
245
+ `record`'s `ok` means a cassette was written, and the run's own verdict is `results[0].verdict.pass`
246
+ (`items[].verdict` on `record <dir/>`).
236
247
 
237
248
  ## Privacy: cassettes are committed fixtures → record only against SYNTHETIC inputs
238
249
 
@@ -247,7 +258,9 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
247
258
  dir** (first file found per dir; env vars merge on top). `cowork-harness init-redact` copies the
248
259
  packaged reference template (local-path prefixes, incl. macOS temp roots and slugged home segments
249
260
  like `-Users-<user>-…`, + a generic email regex) into the cwd as a starting point — review and tailor
250
- it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` or add them. Redaction is **verdict-preserving** — `record` refuses to write if
261
+ it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` (without it
262
+ `init-redact` refuses to overwrite an existing policy; with it your tailoring is replaced — save and
263
+ re-apply it) or add them by hand. Redaction is **verdict-preserving** — `record` refuses to write if
251
264
  redaction would flip an assertion (a manufactured green). `--no-redact` skips it for known-synthetic
252
265
  inputs.
253
266
  - **Pre-spawn preflight**: `record` warns (`::warning::`, before the paid run starts — once per batch
@@ -332,7 +345,16 @@ A typical skill repo runs four stages, fastest/cheapest first:
332
345
 
333
346
  This is the shape a CI step wants: **silent on success (no output, exit 0), loud and specific on
334
347
  failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
335
- offending file *and* the rejected key, one line per file, and the step still exits 1. Point `lint`
348
+ offending file *and* the rejected key, one line per file, and the step still exits 1. It exits 0,
349
+ though, on an input the real record would refuse (a `session:` file that cannot be read — 4.1.1 and
350
+ later — a missing path, an unknown baseline name, a tier-vacuous `tool_not_called`): that prints a `⚠ input error:` line and lands in `inputErrors[]`. To
351
+ gate on those too, use the JSON form (4.1.0 and later):
352
+
353
+ ```bash
354
+ cowork-harness record scenarios/ --dry-run --output-format json | jq -e '.ok and (.inputErrors == [])'
355
+ ```
356
+
357
+ Point `lint`
336
358
  at scenarios only: a session or matrix YAML in the linted set is reported as a file that does not
337
359
  load.
338
360
 
@@ -373,7 +395,7 @@ jobs:
373
395
  with: { node-version: '24' }
374
396
  - uses: actions/setup-python@v5
375
397
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
376
- - run: npm i -g "cowork-harness@^4.0.0"
398
+ - run: npm i -g "cowork-harness@^4.1.1"
377
399
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
378
400
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
379
401
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -402,7 +424,7 @@ jobs:
402
424
  echo "live=true" >> "$GITHUB_OUTPUT"
403
425
  fi
404
426
  - if: steps.guard.outputs.live == 'true'
405
- run: npm i -g "cowork-harness@^4.0.0"
427
+ run: npm i -g "cowork-harness@^4.1.1"
406
428
  - if: steps.guard.outputs.live == 'true'
407
429
  run: cowork-harness run scenarios/ --output-format json
408
430
  env:
@@ -425,7 +447,8 @@ sandbox).
425
447
 
426
448
  `--output-format json` emits a machine envelope on stdout (human output goes to stderr):
427
449
  `{tool, version, command, ok, results[], error}` — one `RunResult` per scenario. **Overall pass for a
428
- scenario is `verdict.pass`** (envelope-wide: `ok`), and it is strictly stronger than
450
+ scenario is `verdict.pass`** (envelope-wide: `ok`; on `record`, `ok` only means the command exited 0 —
451
+ see below), and it is strictly stronger than
429
452
  `result === "success" && assertions.every(pass)`: the verdict also carries ~20 signal codes that fail a run
430
453
  with no failing assertion at all — `stalled`, `outputs_delete`, `mount_delete`, `host_path_leak`,
431
454
  `undelivered_deliverables`, `missing_capability`, `permissive_auto_allow`, `ended_with_question`,
@@ -459,6 +482,11 @@ Keep the `.results[]?` hop and the `?` operators. `.results[0]` silently ignores
459
482
  the first when you pass a directory, and a bare `.verdict` does not exist at the envelope root at all —
460
483
  both read as "no failures" against a run that failed.
461
484
 
485
+ **`record` is shaped differently.** Its `ok` is the exit code's verdict (a cassette was written), not the
486
+ run's: `record --allow-failing` on a red run is `ok: true`. The verdict is in `results[0].verdict` on
487
+ `record <file>` and in `items[].verdict` on `record <dir/>` and `--rerecord-stale`, which carry no
488
+ `results[]` — so `.results[]?` reads nothing there. Use `[.items[]? | .verdict.failures[]?]` on a batch.
489
+
462
490
  > **Do not filter on whether `assertion` is present.** That was the only discriminator before `kind`
463
491
  > existed and it never worked in both directions: `coverage` entries carry a key too (an internal
464
492
  > `answer_coverage` marker), so they read as authored asserts, while `guard`, `staleness` and
@@ -468,7 +496,9 @@ both read as "no failures" against a run that failed.
468
496
  `--output-format json`, `run` / `record` / `replay` / `verify-cassettes` / `status` write their whole
469
497
  human rendering — warnings, verdict, `status`'s summary line — to **stderr**, and stdout stays empty.
470
498
  A wrapper that captures only stdout gets an empty log and, if it greps that for a state, a silent false
471
- negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout.
499
+ negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout —
500
+ `COWORK_HARNESS_OUTPUT_FORMAT=json` makes JSON the default for every command that takes `--output-format`
501
+ (an explicit flag still wins).
472
502
  (Commands whose whole job is to print a value — `--version`, `assertions --list`, `scaffold`, `gates`,
473
503
  `skill --dry-run` — write it to stdout by design. Under `--output-format json` the value rides inside the
474
504
  envelope: `scaffold`'s YAML is `.scenario`, `skill --dry-run`'s preview is the envelope's own fields;
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,13 +1,15 @@
1
1
  # Debugging a run
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
4
4
 
5
5
  ## Part III — Debug
6
6
 
7
7
  A run misbehaved, or greened when you don't trust it. Debugging is a first-class loop, not an
8
8
  afterthought: the run already wrote its evidence, so you **localize the failure post-hoc** rather than
9
9
  re-run and hope. Start at the triage below, then use the observability output and, when you need to
10
- reproduce interactively, `chat`.
10
+ reproduce interactively, `chat`. (The fuller human-facing map is
11
+ [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only,
12
+ not shipped with the installed skill.)
11
13
 
12
14
  > **"Evidence" below means the run's own record** — events, trace, transcript; what `trace` / `inspect` /
13
15
  > `diff` / `verify-run` / `replay --explain` read. `critique`'s **evaluator** grades against a separate,
@@ -32,6 +34,15 @@ agent stderr) — read those before re-running; a re-record rarely tells you mor
32
34
  already does.
33
35
  <!-- END triage-canonical -->
34
36
 
37
+ **microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
38
+ stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
39
+ `cowork-harness vm status` — a `provisioning` other than `ready` confirms it — and if a run does not
40
+ recover it on its own, `cowork-harness vm delete` and retry. A run on such a VM can also have cached an
41
+ empty toolchain for it in the capability probe's cache: `vm delete` drops that VM's entry (`vm prune`
42
+ forgets only the orphaned VMs it deletes, never the current one); otherwise delete `capability-cache.json`
43
+ from the runs root (`~/.cowork-harness/runs/` unless
44
+ `COWORK_HARNESS_RUNS_DIR` is set) so the next run probes again.
45
+
35
46
  **Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
36
47
  **deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
37
48
  so a "bug" you see here that real Cowork also has isn't yours to fix. The tier semantics are in
@@ -73,7 +84,9 @@ decide which assertions from *Assertions: two orthogonal axes* in `assertions-gu
73
84
  walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
74
85
  `total_cost_usd` for the run — the authoritative single-run spend of the agent session, which leaves out the `semantic_matches` judge and the LLM decider calls; NOT the same source as summing
75
86
  `modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
76
- `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`, `toolErrors`,
87
+ `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`,
88
+ `toolCalls` (every tool call in stream order: `name`, top-level `input` fields each capped at 10 KB, and
89
+ `origin` `main`/`subagent`/`unknown` — what the object form of `tool_called`/`tool_not_called` reads), `toolErrors`,
77
90
  `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
78
91
  `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
79
92
  `context` (tools/mcpServers/availableSkills), `tasks`,
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`).
3
+ Self-contained reference. Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -73,7 +73,9 @@ The instance is `cowork-vm-<config-hash>` — a config or agent-version change y
73
73
  stale VM is never silently reused (the old one is orphaned until `vm prune`). Pin a fixed name with
74
74
  `COWORK_LIMA_INSTANCE`. The optional argument to every `vm` subcommand is a **baseline**
75
75
  (`desktop-<version>`, default `latest`), never the `cowork-vm-<hash>` VM name — passing a VM name is a
76
- usage error that names the baseline(s) deriving it.
76
+ usage error that names the baseline(s) deriving it. A Running VM is used only once its provisioning has
77
+ finished (a run waits up to `COWORK_VM_PROVISION_TIMEOUT_S`, default 900 s); `microvm <instance> never
78
+ finished provisioning (…)` means it will not — `cowork-harness vm delete` and retry.
77
79
 
78
80
  ## Answer paths (resolving gates: AskUserQuestion + tool-permission)
79
81
 
@@ -325,7 +327,7 @@ up often enough to spell out:
325
327
  - `COWORK_HARNESS_SCRUB_KEYS` / `COWORK_HARNESS_SCRUB_VALUES` — extra env names / literal values to
326
328
  redact from logs (beyond the auth tokens + `ANTHROPIC_CUSTOM_HEADERS`).
327
329
  - `COWORK_HARNESS_SOFT_MISSING` — downgrade a missing mount source from hard-error to warn-and-skip.
328
- - `COWORK_VM_GATEWAY` / `COWORK_VM_PROXY_PORT` / `COWORK_LIMA_INSTANCE` — L2 (microVM) knobs.
330
+ - `COWORK_VM_GATEWAY` / `COWORK_VM_PROXY_PORT` / `COWORK_LIMA_INSTANCE` / `COWORK_VM_PROVISION_TIMEOUT_S` — L2 (microVM) knobs.
329
331
  - `COWORK_LOCKDOWN` — default `on`: gates sandbox hardening on every isolated tier — **aborts loudly**
330
332
  if the L2 guest firewall fails to apply (no silent unprotected run), and also gates the
331
333
  `container`/`hostloop` Docker hardening. Set `=off` to opt out and run without isolation deliberately.
@@ -1,6 +1,6 @@
1
1
  # Gotchas
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). The full "✓ passed ≠ correct" landmine catalog.
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). The full "✓ passed ≠ correct" landmine catalog.
4
4
 
5
5
  ## Gotchas — the "✓ passed ≠ correct" landmines
6
6
 
@@ -175,9 +175,10 @@ authorable). Reach for this list when debugging a run's behavior, that one while
175
175
  `--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
176
176
  `prompt`/`answers`/`baseline`/`fidelity`/`lane`/`skills`/`requires_capabilities` or the skill content (when a
177
177
  fingerprint exists) drifted from the recording (re-record then), and `expect_denied`/filesystem/egress keys
178
- are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the `session`
179
- (model / data mounts / discovery) is NOT drift-checked or fingerprinted, so a **model change** between record
180
- and re-assert is undetected — the notice flags this; re-record if the session changed. `verify-run` reads
178
+ are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the session's
179
+ `model:` IS in the cassette's `sessionFingerprint` (a `--model`/env model is not; `environment.model`
180
+ records what ran). `verify-cassettes` reports a changed session as staleness (exit 1); `replay` never
181
+ checks it — plain, `--strict` or `--assert-from` — so re-record if the session changed. `verify-run` reads
181
182
  on-disk `assert:` against a kept *run dir*; `replay --assert-from` is the equivalent for a *cassette*.
182
183
 
183
184
  18. **`questions_count_max` counts sub-questions, not gates.** One `AskUserQuestion` tool call can
@@ -214,13 +215,17 @@ authorable). Reach for this list when debugging a run's behavior, that one while
214
215
 
215
216
  22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
216
217
  `manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
217
- assertion keys. The linter is **static**: it never reads your cassettes, so it cannot know whether
218
- yours already carry an `artifacts` manifest and `controlOut` (a current cassette does). On a healthy
219
- fleet every one of those lines is a false alarm. One exception: `manifest-needs-snapshot` is
220
- suppressed for `user_visible_artifact` on `lane: remote` — the only manifest-backed key that lane also
221
- rejects outright (`lane-remote-incompatible-key`, an ERROR), so the INFO would be redundant advice
222
- about a key the scenario can never even load with. `gate-needs-controlout` has no such exception. *Fix:*
223
- `lint --min-severity WARN` in CI (≥1.11.0) — the INFO advisories stay one flag away for interactive use.
218
+ assertion keys. The linter is **static** until you opt in to cassette evidence. With committed
219
+ recordings, pass `lint --cassette-dir <dir>` (or one cassette file): it scans the same `*.cassette.json`
220
+ shape as replay and verify-cassettes, resolves each exact `scenarioSource` relative to its cassette,
221
+ and suppresses an advisory only when **all** matching cassettes carry the evidence the replay lane
222
+ needs. A malformed, wrong-shaped, unsupported-version, or provenance-less cassette is reported as INFO
223
+ and keeps the advisory visible — even beside a healthy sibling. The two dedicated diff assertions also
224
+ require their respective `preRunPaths` / `preRunHashes` baselines, and a non-empty `controlOut` /
225
+ artifact manifest is required where replay needs it. For a strict CI gate that keeps actionable INFO
226
+ rules visible while suppressing only proven replay noise, use `lint --strict --min-severity INFO
227
+ --cassette-dir <dir>`. Without the opt-in path, use `lint --min-severity WARN` in CI (≥1.11.0) to hide
228
+ the INFO class.
224
229
  From 4.0.0 WARN is `--strict`'s default floor, so bare `lint --strict` hides and passes INFO; add
225
230
  `--min-severity INFO` to fail on it. `--strict --min-severity ERROR` behaves as a plain lint, not a
226
231
  contradiction.
@@ -1,6 +1,6 @@
1
1
  # Measurement
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
4
4
 
5
5
  ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
6
6
 
@@ -40,7 +40,10 @@ Compare timings between runs of the same tier and model only.
40
40
 
41
41
  1. **Pin the model in the session.** A run that resolves no model is refused, but one pinned only by
42
42
  `COWORK_HARNESS_MODEL` takes its model from the machine, so two shells can run two models. Set
43
- `model:` in the session (or pass the same `--model` on every `skill` run). Read `result.json`'s `models` back before believing any cross-run comparison — and when
43
+ `model:` in the session (or pass the same `--model` on every `skill` run). Adding `model:` to a session
44
+ that already has cassettes re-stales them (`verify-cassettes` exits 1 — the model is in the session
45
+ fingerprint); to pin without re-recording now, use `--model` or `COWORK_HARNESS_MODEL` and move the
46
+ pin into the session at the next re-record. Read `result.json`'s `models` back before believing any cross-run comparison — and when
44
47
  you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
45
48
  fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
46
49
  array purely by whether such a turn occurred.
@@ -1,6 +1,6 @@
1
1
  # Run, record and lock
2
2
 
3
- Tracks `cowork-harness 4.0.0` (baseline `desktop-2.9939.4`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
3
+ Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
4
4
 
5
5
  ## Part II — RUN, RECORD & LOCK
6
6
 
@@ -25,12 +25,17 @@ the discovery/encode/record dance entirely and answer gates **live during the re
25
25
  refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
26
26
  cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
27
27
  guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
28
- one a paid run would give. On a **directory** the path-dependent verdicts (host-inventory, cassette
28
+ one a paid run would give. Two inputs it reports rather than refuses, at exit 0 under `inputErrors[]` with a
29
+ `⚠ input error:` line: a `session:` file that cannot be read (missing, a directory, not valid YAML — 4.1.1
30
+ and later) and a `tool_not_called` the tier can never violate — `run` and the real `record` refuse both. On a **directory** the path-dependent verdicts (host-inventory, cassette
29
31
  portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
30
32
  the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
31
33
  takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
32
- contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
33
- real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
34
+ contradiction, duplicate cassette target) gate the batch. An input the real record would refuse — a
35
+ `session:` file that cannot be read (4.1.1 and later), a missing path, an unknown baseline name, a `tool_not_called` the tier can never violate — is listed under
36
+ `inputErrors[]` with a `⚠ input error:` line, also at exit 0. So a directory dry-run CAN exit 0 on a scenario the
37
+ real `record` would refuse; re-run that one file with its real flags for a binding answer, or gate on
38
+ `.ok and (.inputErrors == [])` in the JSON payload (4.1.0 and later). A directory also
34
39
  reports every offender and the batch cost estimate. `lint` checks the assertion invariants AND that each file loads (the same loader, plus a named `baseline:`), but not the pre-spend refusals.
35
40
 
36
41
  **Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
@@ -65,7 +70,10 @@ finished: a failing verdict without `--allow-failing`, an assert on an artifact
65
70
  quarantined inventory finding, or any other error before the cassette is written. Only the after-the-run
66
71
  refusals report the run: under `--output-format json` they carry it in `results[0]` (verdict and
67
72
  cost) with `error.category: "runtime"`; every pre-spend refusal, and a run that throws before returning a
68
- result (an unanswered gate), has `results: []`.
73
+ result (an unanswered gate), has `results: []`. The envelope's `ok` is the exit code (`ok` ⇔ exit 0, a
74
+ cassette was written), not the verdict: an `--allow-failing` recording of a red run is `ok: true` with
75
+ `results[0].verdict.pass: false`. `record <dir/>` reports per scenario under `items[]` (each with `status`,
76
+ and `verdict`/`result` once its run finished), not `results[]`.
69
77
 
70
78
  **Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
71
79
  Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
@@ -249,7 +257,11 @@ Recognize these before "fixing" a non-bug:
249
257
  `container`/`microvm`, but only *fires* on an actual scanned leak with no authored
250
258
  `transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
251
259
  run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
252
- it's valid.
260
+ it's valid. A host path that came verbatim from the scenario's own uploads, connected folders or prompt
261
+ is not a leak (whole-token match; a token cut short where the path goes on, or one naming a location the
262
+ harness created for the run, is never exempt); `scan.inputHostPathTokens` counts the host-path tokens the
263
+ inputs carried, `scan.hostPathsFromInputs` the ones exempted, and a non-zero exemption prints a
264
+ `::notice::` so a clean scan that relied on it is never silent.
253
265
  - **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
254
266
  infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
255
267
  rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike