cowork-harness 4.6.0 → 4.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (76) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +5 -5
  2. package/.claude/skills/cowork-harness/references/assertion-catalog-agents-skills-budgets.md +1 -1
  3. package/.claude/skills/cowork-harness/references/assertion-catalog-gates-hooks-modifiers.md +5 -5
  4. package/.claude/skills/cowork-harness/references/assertion-catalog-outcome-files-tools.md +7 -7
  5. package/.claude/skills/cowork-harness/references/assertion-catalog.md +1 -1
  6. package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
  7. package/.claude/skills/cowork-harness/references/authoring.md +9 -7
  8. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  9. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  10. package/.claude/skills/cowork-harness/references/debugging.md +6 -3
  11. package/.claude/skills/cowork-harness/references/eval.md +3 -1
  12. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  13. package/.claude/skills/cowork-harness/references/gotchas.md +1 -1
  14. package/.claude/skills/cowork-harness/references/hillclimb-recipe.md +1 -1
  15. package/.claude/skills/cowork-harness/references/hillclimb.md +1 -1
  16. package/.claude/skills/cowork-harness/references/measurement.md +1 -1
  17. package/.claude/skills/cowork-harness/references/run-record-replay.md +2 -2
  18. package/.claude/skills/cowork-harness/references/scenario-schema.md +10 -6
  19. package/.claude/skills/cowork-harness/references/semantic-judging.md +7 -1
  20. package/.claude/skills/cowork-harness/references/task-recipes.md +20 -10
  21. package/.claude/skills/cowork-harness/scripts/scenario.py +67 -47
  22. package/CHANGELOG.md +110 -0
  23. package/DESIGN.md +2 -2
  24. package/README.md +10 -9
  25. package/RELEASING.md +5 -3
  26. package/baselines/desktop-2.31226.0.json +1528 -0
  27. package/baselines/prompts/cowork-system-prompt-fingerprints.json +10 -1
  28. package/dist/assert.js +59 -13
  29. package/dist/cli.js +11 -1
  30. package/dist/critique/command.js +12 -1
  31. package/dist/decide/llm-transport.js +7 -0
  32. package/dist/eval/command.js +8 -1
  33. package/dist/eval/snapshot.js +24 -5
  34. package/dist/hillclimb/freeze-ref.js +2 -1
  35. package/dist/hillclimb/job.js +4 -0
  36. package/dist/hillclimb/regrade.js +6 -3
  37. package/dist/hillclimb/rows.js +1 -1
  38. package/dist/hillclimb/run-command.js +7 -2
  39. package/dist/hillclimb/runner.js +11 -0
  40. package/dist/hillclimb/snapshot.js +3 -0
  41. package/dist/refs/compose.js +2 -2
  42. package/dist/refs/preflight.js +8 -1
  43. package/dist/refs/store.js +4 -1
  44. package/dist/run/cassette.js +12 -0
  45. package/dist/run/execute.js +8 -22
  46. package/dist/run/lane-notice.js +49 -2
  47. package/dist/run/pairwise-prepass.js +6 -3
  48. package/dist/run/pre-run-manifest.js +4 -2
  49. package/dist/run/regrade.js +2 -2
  50. package/dist/runtime/agent-tree.js +15 -5
  51. package/dist/runtime/host-cli-probe.js +5 -0
  52. package/dist/sync/cowork-sync.js +265 -19
  53. package/dist/termination.js +34 -0
  54. package/dist/types.js +13 -12
  55. package/docs/cassette.md +2 -1
  56. package/docs/ci.md +1 -1
  57. package/docs/cli.md +13 -9
  58. package/docs/companion-skill.md +2 -2
  59. package/docs/eval.md +4 -1
  60. package/docs/fidelity-gaps.md +13 -9
  61. package/docs/hillclimb.md +7 -0
  62. package/docs/invariants.md +1 -1
  63. package/docs/scenario.md +35 -22
  64. package/examples/README.md +2 -1
  65. package/examples/data/form-replies/README.md +16 -0
  66. package/examples/data/form-replies/deck-review.txt +8 -0
  67. package/examples/data/form-replies/pitch-review.txt +12 -0
  68. package/examples/replays/README.md +1 -1
  69. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  70. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  71. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  72. package/package.json +1 -1
  73. package/python/README.md +1 -1
  74. package/python/test_scenario_lint.py +43 -28
  75. package/schema/scenario.schema.json +15 -15
  76. package/scripts/bump-version.ts +22 -5
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did it make answers worse? (`eval`: paired A/B, pinned models) — or improving one round by round (`hillclimb`, the `/claude-api hillclimb` runner). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval / hillclimb commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 4.6.0
7
- tracks-harness: cowork-harness 4.6.0 (baseline desktop-2.26454.2)
6
+ version: 4.6.1
7
+ tracks-harness: cowork-harness 4.6.1 (baseline desktop-2.31226.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -26,8 +26,8 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
26
26
  full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
27
  Read them.
28
28
 
29
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.6.0` (baseline
30
- > `desktop-2.26454.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.6.1` (baseline
30
+ > `desktop-2.31226.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
31
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
32
32
 
33
33
  ## Preflight — make sure the harness can actually run
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
43
43
 
44
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
45
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
46
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.6.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.6.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.6.0"`. **Pin `@^4.6.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.6.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.6.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.6.1"`. **Pin `@^4.6.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
47
47
 
48
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
49
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog: sub-agents, skills and connectors, budgets, tasks, delivery
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Part 2 of 3 of the per-key table. The conventions every key shares (an item with several
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Part 2 of 3 of the per-key table. The conventions every key shares (an item with several
4
4
  keys is an AND; glob vs regex name matching) and the verdict-signal table are in
5
5
  [`assertion-catalog.md`](./assertion-catalog.md); the other parts are [`assertion-catalog-outcome-files-tools.md`](./assertion-catalog-outcome-files-tools.md), [`assertion-catalog-gates-hooks-modifiers.md`](./assertion-catalog-gates-hooks-modifiers.md).
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog: gates, hooks, path denial, verdict modifiers, egress, judged keys
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Part 3 of 3 of the per-key table. The conventions every key shares (an item with several
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Part 3 of 3 of the per-key table. The conventions every key shares (an item with several
4
4
  keys is an AND; glob vs regex name matching) and the verdict-signal table are in
5
5
  [`assertion-catalog.md`](./assertion-catalog.md); the other parts are [`assertion-catalog-outcome-files-tools.md`](./assertion-catalog-outcome-files-tools.md), [`assertion-catalog-agents-skills-budgets.md`](./assertion-catalog-agents-skills-budgets.md).
6
6
 
@@ -40,11 +40,11 @@ refused at load: no question reaches the harness to grade.
40
40
  | `egress_allowed: <host>` | the host was allowed through |
41
41
  | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
42
42
  | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
43
- | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer: the final result text, the transcript and the files the agent authored. The transcript is top-level `assistant_text` only — it excludes every `tool_use`/`tool_result` and sub-agent text unless opted in. Full semantics (evidence scope, fork results, refusals, judge provenance, cost): [semantic-judging.md](semantic-judging.md). |
43
+ | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer: the final result text, the transcript and the files the agent authored. The transcript is top-level `assistant_text` only — it excludes every `tool_use`/`tool_result` and sub-agent text unless opted in. Full semantics (evidence scope, fork results, refusals, judge provenance, cost): [semantic-judging.md](semantic-judging.md). On `lane: remote`: judged on the transcript only; `evidence_files` is rejected at load. |
44
44
  | `semantic_pairwise: {refs?: [..], rubric?: [..], pass_if?, order?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge compares this run's judged document with a **frozen reference** — the same document an earlier run (usually the baseline) produced, written once by `ref freeze` and never regenerated — and answers win, tie, loss or `both_bad`. **The judged document is the one `semantic_matches` builds** (final message + transcript + authored files, same `evidence_files` / `include_subagent_text` / `include_fork_results` options), so the transcript **excludes every `tool_use`/`tool_result`** and a criterion about whether a tool was called cannot be judged from it. `refs` are reference stores relative to the scenario file; the run is judged against every one (under `hillclimb run`, against the flow's own references instead — only the baseline's gates `pass_if`, a later variant's is a metric recorded with `gate: false`), and `pass_if` (default `not_worse`) must hold against each gating one: `win`, `not_worse` (win or tie), or `any` (graded at all — a metric only). `both_bad` fails `win` and `not_worse`. Which output the judge sees first is a seeded coin per run, assert and reference (`order: both` judges both orders: a win in one and a loss in the other is position bias and scores as a tie; any other disagreement keeps the worse outcome; each order's own outcome is recorded as `pairwise[].orders.candidate_first` / `.ref_first`, absent for a single-order grade, and a judge that favours whichever output it sees first shows across runs as `candidate_first` winning more often than `ref_first`). The judge never sees the words "reference" or "baseline", and the stored reference is scrubbed with the grading process's secret set before it reads it (`pairwise[].refRedactions`). A reference must have been frozen with the same evidence options and for the same prompt; a missing, damaged, differently-scoped or other-prompt one **refuses the run before it spends anything**, as does a reference store inside any mounted source (the agent could read its own answer key). Unavailable evidence refuses with `semantic_matches`' typed reasons. Per-reference outcomes land in `assertions[].pairwise`. **LIVE-ONLY** — skipped on replay. |
45
- | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud). A glob `artifact` (`*`, `?`, `**`; `[` is literal) needs `match: each` (every match) or `match: any` (≥1), e.g. `{artifact: outputs/artifacts/runs/*/run_status.json, match: each, path: status, equals: complete}`. It matches files under the user-visible roots only; uploaded inputs are not matched. Zero matches fail; >200 matches, a link (a matched file, or a symlinked directory on the way to one), a body-less match or an incomplete walk (over 32 levels below the glob's fixed part, over 20,000 entries, or an unreadable directory) is evidence-unavailable. Names match exactly, with no case or Unicode folding. `authored: true` applies to every match. A glob that can reach no user-visible root (an `uploads/` glob) fails and says why. At `record`, a match stored hash-only over the body cap, or an artifact walk that could not see the whole tree, is refused, as for a literal path. Fails on `lane: remote`, in every form: that lane's container filesystem is not locally observable, so there is no body to parse |
46
- | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
47
- | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
45
+ | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud). A glob `artifact` (`*`, `?`, `**`; `[` is literal) needs `match: each` (every match) or `match: any` (≥1), e.g. `{artifact: outputs/artifacts/runs/*/run_status.json, match: each, path: status, equals: complete}`. It matches files under the user-visible roots only; uploaded inputs are not matched. Zero matches fail; >200 matches, a link (a matched file, or a symlinked directory on the way to one), a body-less match or an incomplete walk (over 32 levels below the glob's fixed part, over 20,000 entries, or an unreadable directory) is evidence-unavailable. Names match exactly, with no case or Unicode folding. `authored: true` applies to every match. A glob that can reach no user-visible root (an `uploads/` glob) fails and says why. At `record`, a match stored hash-only over the body cap, or an artifact walk that could not see the whole tree, is refused, as for a literal path. Rejected at load on `lane: remote`, in every form (that lane's container filesystem is not locally observable) |
46
+ | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. Rejected at load on `lane: remote`. |
47
+ | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. Rejected at load on `lane: remote`. |
48
48
 
49
49
  `expect_denied: [host, …]` adds one `egress_denied` per host. Run `cowork-harness assertions --list` for this
50
50
  table from the live schema. Example: `artifact_json: { artifact: outputs/cap.json, path: me.run_id, equals: "r1" }`.
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog: outcome, transcript, files and artifacts, tools, tool results
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Part 1 of 3 of the per-key table. The conventions every key shares (an item with several
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Part 1 of 3 of the per-key table. The conventions every key shares (an item with several
4
4
  keys is an AND; glob vs regex name matching) and the verdict-signal table are in
5
5
  [`assertion-catalog.md`](./assertion-catalog.md); the other parts are [`assertion-catalog-agents-skills-budgets.md`](./assertion-catalog-agents-skills-budgets.md), [`assertion-catalog-gates-hooks-modifiers.md`](./assertion-catalog-gates-hooks-modifiers.md).
6
6
 
@@ -11,16 +11,16 @@ keys is an AND; glob vs regex name matching) and the verdict-signal table are in
11
11
  | `transcript_not_contains: <str>` | it does not **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
12
12
  | `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
13
13
  | `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
14
- | `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it. Object form `{path, authored}`: `authored: true` also requires that THIS run created or rewrote the file (an untouched pre-run file — e.g. from a `workspace_fixture` — fails; no pre-run manifest ⇒ evidence-unavailable); `authored: false` states that inheriting it is fine. On a `workspace_fixture` path one of the two is REQUIRED (refused at load otherwise) |
14
+ | `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it. Object form `{path, authored}`: `authored: true` also requires that THIS run created or rewrote the file (an untouched pre-run file — e.g. from a `workspace_fixture` — fails; no pre-run manifest ⇒ evidence-unavailable); `authored: false` states that inheriting it is fine. On a `workspace_fixture` path one of the two is REQUIRED (refused at load otherwise). On `lane: remote` plain `file_exists` is the proxy for a written path; with `authored: true` it is rejected at load. |
15
15
  | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected. Takes the same `{path, authored}` object form as `file_exists` (and `artifact_text`/`artifact_json` take an `authored` field) |
16
16
  | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. Checks on **every** baseline. Omitting it: on a baseline recording outputs `rw` (Desktop before 2.16120.0) a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored (`allow_outputs_delete: true` accepts an intended one); on `rwd` (2.16120.0+, including `latest`, where Cowork allows deletes in outputs) nothing checks outputs deletes, so author this key to keep the check. Detection is a post-run bash-command scan plus a per-turn filesystem diff of `outputs/`, not mount enforcement, so a green means none was *detected*. Fails when the diff proves the delete, when a delete in command/call position has an `outputs/` path as its own operand, or when the diff could not verify; a hit resting only on the detector's inference (e.g. a Python variable named `rm`) with a clean diff passes, and the `outputs_delete_unconfirmed` warn is still raised in the run output — see that code for the classes of real delete that land there |
17
17
  | `no_delete_in_mounts: true` | no delete op touched `outputs` or any `rw` connected folder, except mounts waived by `allow_delete_in`. Covers `outputs` on **every** baseline, including those where Cowork allows an outputs delete (unless outputs is waived, authoring it arms the outputs check, so a filesystem-proven delete fails as the `outputs_delete` signal, which `allow_outputs_delete` waives). Production denies `unlink`/`rmdir` on a `rw` connected folder until approval, so `no_delete_in_outputs` covers only part of the rule. **only `true` is valid** |
18
- | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
19
- | `input_unmodified: <glob> \| [<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
18
+ | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file. Rejected at load on `lane: remote`. |
19
+ | `input_unmodified: <glob> \| [<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree. On `lane: remote` it still reads the local stand-in for the inputs (unconfirmed). |
20
20
  | `self_heal_ran: <bool>` | a bash command the model wrote did (not) name a plugin under `/sessions/<id>/mnt/.local-plugins/` or `/sessions/<id>/mnt/.remote-plugins/` — the plugin-root self-heal path. It reads the command as the model wrote it, so a host plugin path that `hostloop`'s bash tool rewrote to the VM mount does not count |
21
- | `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); evidence-unavailable on `lane: remote` / `preRunOrigin: remote-unavailable` |
22
- | `artifact_text: {artifact, contains?: [..], not_contains?: [..], matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read. Fails on `lane: remote`: that lane's container filesystem is not locally observable, so there is no body to read |
23
- | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
21
+ | `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); rejected at load on `lane: remote`; evidence-unavailable on `preRunOrigin: remote-unavailable` |
22
+ | `artifact_text: {artifact, contains?: [..], not_contains?: [..], matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read. Rejected at load on `lane: remote` (that lane's container filesystem is not locally observable) |
23
+ | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate. Rejected at load on `lane: remote`. |
24
24
  | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. `cowork-harness lint` reports it too (ERROR `scenario-invalid` — the wrapper runs the same loader); the bundled `scenario.py lint` run directly does NOT |
25
25
  | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
26
26
  | `tool_called: {tool: <glob> \| [..], input?, input_any?, result?, scope?, subagent_type?, count?}` | **object form** — asserts what a call CARRIED, where it RAN and what its paired result SAID, which no name glob and no `transcript_*` key can see (`transcript_*` reads top-level prose only, so it passes when the agent merely *says* it ran a command). `tool`: a glob or a list of globs (any-of; list both shells, `[Bash, mcp__workspace__bash]`, for a claim that must hold at hostloop too). `input: {<field>: <regex>}`: every named top-level input field must match (case-insensitive, unanchored; a missing field is no match; a non-string value is matched as JSON). `input_any: <regex>`: some top-level field matches. `result: {matches?, not_matches?, is_error?}`: predicates on the call's paired `tool_result` (by `toolUseId`, 10,240 characters of text — 32,768 for a top-level `Skill` result). `scope`: `main` (default — the main agent, including a `Skill`'s or `Agent(fork)`'s children, the set the string form reads), `subagent` (the parent is a sub-agent dispatch this run recorded, at any depth), or `any` (also calls whose parent is not a recorded dispatch, e.g. a `Skill` invoked inside a sub-agent). `subagent_type` (with `scope: subagent`): regex over the IMMEDIATE parent dispatch's type or description. `count: {min?, max?}` (default `min: 1`). **Fails closed:** an unpaired call never satisfies `result`; a truncated result or input that cannot settle a predicate, or a result.json without `toolCalls`, reports *evidence unavailable*, never a pass or a plain "not called". A red lists the calls it considered and names matches in another scope ("1 matching call in scope subagent — set `scope: any`"). `{tool: X}` alone is exactly `tool_called: X`. A cassette using this form stamps **v13**. |
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Every `assert:` key with its semantics, and the
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Every `assert:` key with its semantics, and the
4
4
  verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
5
5
  the scenario and session YAML fields are there too.
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Assertions guide
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog starts at `assertion-catalog.md`, which indexes the per-family files that hold every key's row.
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog starts at `assertion-catalog.md`, which indexes the per-family files that hold every key's row.
4
4
 
5
5
  ### Assertions: two orthogonal axes
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Authoring a scenario
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
4
4
 
5
5
  ## Part I — AUTHOR a scenario
6
6
 
@@ -65,10 +65,13 @@ before then no setting reliably decided the lane. Check the session's own lane
65
65
  So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
66
66
  asserting a **path, mount, delivery mechanism or egress rule** is a claim about the local lane only. Declare
67
67
  `lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
68
- rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
69
- in `run-record-replay.md`; among the keys that load, `file_absent`, `artifact_text` and `artifact_json` fail
70
- when graded, since that lane's container filesystem is not locally observable. Other keys that read the work
71
- tree, such as `no_unexpected_files` and `input_unmodified`, still grade the local tree on that lane).
68
+ rather than passing. A check that depends on files inside the agent's container can never pass there, so it is
69
+ rejected at LOAD time (`lane-remote-incompatible-key` in lint): delivery by location, file bodies
70
+ (`artifact_text`/`artifact_json`), absence (`file_absent`/`no_unexpected_files`), `computer://` links,
71
+ `no_lost_write_back`, `file_exists` with `authored: true`, and `semantic_*` with `evidence_files`. `semantic_*` without it is judged on the transcript
72
+ only. `input_unmodified` still reads the local stand-in for the user's inputs (unconfirmed for the cloud lane), and
73
+ plain `file_exists` is the proxy for a written path. Delivery itself is the `delivery_unobservable` WARN (see
74
+ `run-record-replay.md`).
72
75
 
73
76
  ### Choose an answer path (gates: AskUserQuestion + tool-permission)
74
77
 
@@ -253,8 +256,7 @@ cowork-harness lint scenarios/*.yaml
253
256
  | `hook-output-control-char` | ERROR | a `hook_output_contains` / `hook_output_not_contains` `text` or `matches` holding a control character (a double-quoted YAML `\b` is a backspace) — the harness refuses it at load |
254
257
  | `host-path-assert-cowork` | WARN | `transcript_no_host_path` on `cowork` — it fails by design if the tier resolves to hostloop |
255
258
  | `host-path-assert-tier` | ERROR | `transcript_no_host_path` on `hostloop` / `protocol`, where it fails by design |
256
- | `lane-remote-incompatible-key` | ERROR | `present_files_called` / `no_scratchpad_leak` / `user_visible_artifact` on `lane: remote` (the runtime rejects them at load, so the tier rules are suppressed there) |
257
- | `lane-remote-unobservable-key` | WARN | `artifact_json` / `artifact_text` / `file_absent` on `lane: remote`: they load but always fail when graded on a live run or verify-run (replay skips the live-only `file_absent`), since that lane's container filesystem is not locally observable. Instead, assert `file_exists` + `transcript_matches` for content, or `transcript_not_matches` for an absence |
259
+ | `lane-remote-incompatible-key` | ERROR | a key that can never pass on `lane: remote`, which the runtime refuses at load: `present_files_called` / `no_scratchpad_leak` / `user_visible_artifact` (delivery), `artifact_text` / `artifact_json` (file bodies), `file_absent` / `no_unexpected_files` (absence), `computer_links_resolve(_if_present)`, `no_lost_write_back`, `file_exists` with `authored: true`, and `semantic_matches` / `semantic_pairwise` with `evidence_files`. Tier rules are suppressed there |
258
260
  | `linter-extra-findings-invalid` | ERROR | the loader findings `cowork-harness lint` hands the linter could not be read |
259
261
  | `linter-unclassified-key` | ERROR | a valid assertion key this linter cannot classify (the linter is out of date) |
260
262
  | `manifest-needs-snapshot` | INFO | manifest-backed keys, which evaluate on replay only when the cassette carries an `artifacts` manifest (not reported on `lane: remote` for `user_visible_artifact` / `artifact_json` / `artifact_text`, which cannot pass there) |
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`).
3
+ Self-contained reference. Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "4.6.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "4.6.1"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
82
82
  GitHub-hosted runners, no token/Docker/agent:
83
83
 
84
84
  ```yaml
85
- - run: npm i -g "cowork-harness@^4.6.0"
85
+ - run: npm i -g "cowork-harness@^4.6.1"
86
86
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
87
87
  # no silent false-greens. WITHOUT --strict this
88
88
  # step cannot fail on a WARN-class rule (e.g.
@@ -408,7 +408,7 @@ jobs:
408
408
  with: { node-version: '24' }
409
409
  - uses: actions/setup-python@v5
410
410
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
411
- - run: npm i -g "cowork-harness@^4.6.0"
411
+ - run: npm i -g "cowork-harness@^4.6.1"
412
412
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
413
413
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
414
414
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -437,7 +437,7 @@ jobs:
437
437
  echo "live=true" >> "$GITHUB_OUTPUT"
438
438
  fi
439
439
  - if: steps.guard.outputs.live == 'true'
440
- run: npm i -g "cowork-harness@^4.6.0"
440
+ run: npm i -g "cowork-harness@^4.6.1"
441
441
  - if: steps.guard.outputs.live == 'true'
442
442
  run: cowork-harness run scenarios/ --output-format json
443
443
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Debugging a run
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
4
4
 
5
5
  ## Part III — Debug
6
6
 
@@ -101,8 +101,11 @@ decide which assertions from *Assertions: two orthogonal axes* in `assertions-gu
101
101
  so, once per process; `--compact`/`--demo`, `CI` and `COWORK_HARNESS_NO_LANE_NOTICE=1` silence it. On `remote`,
102
102
  location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
103
103
  session end), `present_files` is NOT served, and `user_visible_artifact` /
104
- `present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass;
105
- `file_absent`, `artifact_text` and `artifact_json` fail when graded (no observable filesystem). Reach for it
104
+ `present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass, and so is every
105
+ other key that reads files inside the agent's container (`artifact_text`/`artifact_json`, `file_absent`/
106
+ `no_unexpected_files`, `computer_links_resolve*`, `no_lost_write_back`, `file_exists` with `authored: true`,
107
+ `semantic_*` with `evidence_files`);
108
+ `semantic_*` without it is judged on the transcript only. Reach for it
106
109
  to check a skill's delivery survives the cloud lane, where new Pro and Max tasks run from 2026-10-06. Orthogonal to `fidelity` — a `lane: remote`
107
110
  scenario still runs locally.
108
111
  - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
@@ -1,6 +1,6 @@
1
1
  # `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). The full guide is
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). The full guide is
4
4
  [docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
5
5
  need while running it.
6
6
 
@@ -20,6 +20,8 @@ cowork-harness eval report <eval-dir> # rebuild the report from the eval dir
20
20
  not the `remote_plugins` path a UI-installed plugin has (`mnt/.remote-plugins/plugin_<id>`). A skill that
21
21
  locates its own files at runtime sees a different path under each, so the comparison holds for the
22
22
  `local_plugins` layout only.
23
+ - Every row is a pass or a fail per rep, and `eval` compares pass rates. A scenario's `metrics:` are not graded or
24
+ compared here; `hillclimb` records and compares them.
23
25
  - scenarios × 2 × `--reps` live runs (10 per scenario at the default `--reps 5`), interleaved, plus one judge
24
26
  call per `semantic_matches` assert per run. The start-up line prints the job count.
25
27
  - Before an A/B, plan it: `eval … --dry-run --target-effect 30pp` runs nothing and prints, from the runs
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`).
3
+ Self-contained reference. Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -1,6 +1,6 @@
1
1
  # Gotchas
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). The full "✓ passed ≠ correct" landmine catalog.
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). The full "✓ passed ≠ correct" landmine catalog.
4
4
 
5
5
  ## Gotchas — the "✓ passed ≠ correct" landmines
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Recipe 7 — Climb a skill with `/claude-api hillclimb` and the harness as its runner
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). It needs a `cowork-harness` whose
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). It needs a `cowork-harness` whose
4
4
  `hillclimb --help` lists `--skill` (help goes to stderr). This page is the loop's procedure, step by step, in the order of the
5
5
  `/claude-api hillclimb` guide. Every mechanic (flags, refusals, the gate, `regrade`, `freeze-ref`, exit codes,
6
6
  row keys) is in [`hillclimb.md`](hillclimb.md); the setup and the full list of differences from the guide's own
@@ -1,6 +1,6 @@
1
1
  # `hillclimb` — the runner for a `/claude-api hillclimb` loop
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). It needs a `cowork-harness` whose
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). It needs a `cowork-harness` whose
4
4
  `hillclimb --help` lists `--skill` (help goes to stderr). The command reference is
5
5
  [docs/cli.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md); this is the part a loop needs
6
6
  while it runs. It covers `run`, `check`, `state-template`, `freeze-ref` and `regrade`.
@@ -1,6 +1,6 @@
1
1
  # Measurement
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
4
4
 
5
5
  ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Run, record and lock
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
4
4
 
5
5
  ## Part II — RUN, RECORD & LOCK
6
6
 
@@ -224,7 +224,7 @@ Recognize these before "fixing" a non-bug:
224
224
  `markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
225
225
  (`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
226
226
  Desktop `2.9939.2` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
227
- behind that sentence; 6 baselines have shipped since without a re-capture). The message says so ("likely a FALSE
227
+ behind that sentence; 7 baselines have shipped since without a re-capture). The message says so ("likely a FALSE
228
228
  NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
229
229
  `COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
230
230
  `allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, replay class, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.6.0`
4
- (baseline `desktop-2.26454.2`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.6.1`
4
+ (baseline `desktop-2.31226.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` and `fidelity` are required:
@@ -46,10 +46,14 @@ lane: local # OPTIONAL — which Cowork lane's DELIVERY
46
46
  # can't reach a remote session) — so `user_visible_artifact` and
47
47
  # present_files_called/no_scratchpad_leak are REJECTED AT SCENARIO-LOAD
48
48
  # TIME (before the run starts, before any spend), not left to fail
49
- # unverifiable/can't-verify at assertion time. file_absent /
50
- # artifact_text / artifact_json load but FAIL when graded (the container
51
- # filesystem is not locally observable; assert file_exists +
52
- # transcript_matches instead). Orthogonal to fidelity
49
+ # unverifiable/can't-verify at assertion time. So is every key that
50
+ # reads files inside the agent's container: artifact_text /
51
+ # artifact_json / file_absent / no_unexpected_files /
52
+ # computer_links_resolve* / no_lost_write_back, file_exists with
53
+ # authored: true, and semantic_* with
54
+ # evidence_files (semantic_* without it is judged on the transcript
55
+ # only). Assert file_exists + transcript_matches instead; input_unmodified
56
+ # still reads the local stand-in (unconfirmed). Orthogonal to fidelity
53
57
  # and execution — a `lane: remote` scenario still runs locally.
54
58
  # Delivery semantics only; the remote device bridge is deliberately
55
59
  # unmodeled.
@@ -1,9 +1,15 @@
1
1
  # `semantic_matches` — the full semantics
2
2
 
3
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). The one-line summary is in
3
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`). The one-line summary is in
4
4
  [assertion-catalog-gates-hooks-modifiers.md](assertion-catalog-gates-hooks-modifiers.md); this is the whole contract, split out so the catalog stays within the
5
5
  agent's single-read size.
6
6
 
7
7
  `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}`
8
8
 
9
+ **On `lane: remote`** the judged document is the final answer and transcript only: no authored file, capture-health
10
+ path or sub-agent text is sent (in a cloud run they live inside the container), and a note at the top tells the judge
11
+ so. A graded assert reports `semanticEvidence: {reason: "graded", paths: []}` and says "judged on the transcript
12
+ only (lane: remote)". `evidence_files` is rejected at load on that lane (`lane_remote_evidence_files` if it reaches
13
+ the evaluator another way), and a `semantic_pairwise` reference is stored per lane.
14
+
9
15
  a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). **A `context: fork` skill's own answer is not graded by default**, even with `include_subagent_text: true`: it is not a dispatch (so it has no `subagents[]` entry to fold in), and it reaches the main agent as the `Skill` tool result, which the judged document excludes. **Opt in with `include_fork_results: true`** to grade a FOREGROUND fork's answer: every top-level `Skill` call is joined to its result by `toolUseId` and appended as a `Fork skill result: <skill>` section (`Skill result: <skill>` when the result lacks the agent's `completed (forked execution)` marker — an inline skill's result is only its launch line). A top-level `Skill` result is captured up to 32,768 characters (every other tool result keeps 10,240); one cut at that cap fails evidence-unavailable (`semanticEvidence.reason: "fork_result_truncated"`), a call with no paired result `fork_result_unpaired`, a fork launched in the background (its result is only the `launched (forked execution, running in the background)` line) `fork_result_background`, a result.json with no tool-call record `fork_calls_unrecorded` — `paths` then lists the `Skill` calls' skill names; never a grade over part of an answer. A fork invoked by a leading `/<skill>` prompt makes no `Skill` call, so there is nothing to join: no section is added, the fork's answer is already top-level transcript text, and a pass's evidence reads `graded Skill results: (none — no Skill call)`. Without the key the judge sees the fork's answer only if the main agent restates it; `tool_result_matches` checks it structurally (see `subagent_output_contains`). ⚠️ **Consequence: a rubric claim about whether a tool was called cannot grade true** — the evidence is not in the judged document (one exception: with `include_fork_results: true`, a `## Skill result: X` / `## Fork skill result: X` section lets the judge confirm that skill X ran, from its result). Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` instead, and `hook_event_blocked` / `hook_decision` for a plugin's hook (`hook_blocked` reads only the harness's own). Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document, and the judge is not called for it (no `semanticClaims`, no judge spend or `judgedDoc` — `semanticEvidence.reason` says why); this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass, rationale?}]`, so a consumer can diff the per-claim profile across runs; `rationale` is the judge's one-sentence reason, printed under each failed claim in the failure footer: untrusted model text that can quote the judged document, whose content never affects `pass`, absent when the judge gave none (a reply whose shape is broken, such as unparseable JSON, a malformed `{"results": …}` group beside a valid grade, or a partial restatement that contradicts it, is retried once and then marked `judgeInvalid`), and comparable only between runs that share `judgePromptHash`); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible — or let `eval` compare two skill versions per claim, with the judge pinned). Each graded assert records the judge's provenance: `RunResult.assertions[].judgeModel` (the resolved model), `judgeCostUsd` (judge spend over both attempts, reported beside `cost.usd` and never inside it; absent when unpriced) and `judgePromptHash` (the grading-prompt template — compare only runs that share it). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) **`evidence_files: [globs]` scopes which authored files are graded** — reach for it the moment a run authors more than a couple of files. The capture budget (64 KiB total by default) is spent prefix-major then alphabetically, so a pipeline that stages intermediates (`outputs/_work/*.json`) exhausts it before reaching its own deliverable and the verdict is refused evidence-unavailable over files no rubric mentions. Scoping also makes the capture spend the budget on the named files FIRST and exempts them from the per-file cap. Paths are `<root>/<rel>` (`outputs/report.md`, never a bare `report.md`; session-root writes are `scratchpad/<rel>`); globs are `*`/`?`/`**`, not regex. A glob matching nothing FAILS and the message lists every authored path — read it rather than guessing. Still too big? Raise `$COWORK_HARNESS_AUTHORED_TOTAL_BYTES`. The typed reason is on `RunResult.assertions[].semanticEvidence` — check `.reason` instead of parsing the message
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 4.6.1` (baseline `desktop-2.31226.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -365,15 +365,25 @@ cowork-harness eval scenarios/ --arm full=./my-plugin --arm nosec=/tmp/nosec --d
365
365
 
366
366
  ### Test a skill's parsing of a Desktop form reply
367
367
 
368
- Desktop's elicitation form sends its answers as the next user
369
- message, as one line. Send that line as a resumed turn: `cowork-harness skill ./my-plugin "<the reply line>" --session-id s --resume`. The format,
370
- from Desktop's own form guide:
371
-
372
- - one line: `<Title> details — Label: value · Label: value`, labels being the form's field names in sentence case;
373
- - a multi-select value comma-joined; a short multi-line value flattened with ` / `; a value of 81–200 characters
374
- in quotes;
375
- - a value over 200 characters shown as `Label: (N chars — see below)`, and repeated in full after a
376
- `--- Full content ---` line;
368
+ Desktop's elicitation form sends its answers as the next user message. Send that message as a resumed turn, from a
369
+ file so the shell cannot expand a `$` in it: `cowork-harness skill ./my-plugin --prompt-file reply.txt --session-id s
370
+ --resume`. Two real-shaped replies ship in
371
+ [`examples/data/form-replies/`](https://github.com/yaniv-golan/cowork-harness/tree/main/examples/data/form-replies). The format, as Desktop's serializer builds it:
372
+
373
+ - a compact line, `<Title> — Label: value · Label: value`. The title is the form's header, usually `<Topic> details`;
374
+ with no header there is no title prefix. Labels are the field names with underscores as spaces and the first
375
+ letter capitalised: `_text` is dropped, `_file` becomes ` file`, `_other` becomes ` (other)`. Fields come in the
376
+ order the form collects them: pill groups (each followed by its `(other)` text), then file groups, then the
377
+ remaining fields (text, dates, sliders) in page order; text values are trimmed;
378
+ - a multi-select value comma-joined; newlines in a value of up to 200 characters replaced by ` / `; such a value
379
+ over 80 characters (after that) in quotes;
380
+ - a value over 200 characters shown as `Label: (N chars — see below)` and repeated in full, newlines kept, after a
381
+ blank line and a `--- Full content ---` line, under a `[Label]` line (blank lines between folded values), so
382
+ such a reply spans several lines;
383
+ - a file field shown as `Label: <name> (attached)` (several files comma-joined); the file itself arrives as an
384
+ attachment on that message, not in its text. A resumed turn cannot add a file: declare it with `--upload <file>`
385
+ on the first turn and repeat the same `--upload` on the resumed one, so it is already in `uploads/`;
386
+ - an empty answer left out; a form with every answer empty arrives as `<Title> — proceeding with defaults.`;
377
387
  - a skipped form arrives as one fixed sentence saying it was skipped.
378
388
 
379
389
  *Does not prove:* that the model would choose the form (the harness serves no `visualize` tools; see