cowork-harness 1.6.0 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (90) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +80 -21
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +10 -9
  3. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +18 -2
  4. package/.claude/skills/cowork-harness/references/scenario-schema.md +13 -11
  5. package/.claude/skills/cowork-harness/references/task-recipes.md +3 -3
  6. package/CHANGELOG.md +420 -0
  7. package/CONTRIBUTING.md +10 -6
  8. package/DESIGN.md +1 -1
  9. package/README.md +50 -29
  10. package/RELEASING.md +7 -5
  11. package/SPEC.md +14 -4
  12. package/baselines/desktop-1.24012.0.json +405 -0
  13. package/baselines/desktop-1.24012.1.json +485 -0
  14. package/dist/agent/session.js +9 -3
  15. package/dist/agent/timeline.js +76 -3
  16. package/dist/assert.js +4 -2
  17. package/dist/baseline.js +20 -0
  18. package/dist/cli.js +216 -36
  19. package/dist/critique/command.js +450 -43
  20. package/dist/critique/evaluator.js +111 -43
  21. package/dist/critique/evidence.js +80 -43
  22. package/dist/critique/limitations.js +134 -0
  23. package/dist/critique/package-evidence.js +149 -30
  24. package/dist/decide/llm-transport.js +3 -1
  25. package/dist/dotenv.js +38 -3
  26. package/dist/errors.js +12 -0
  27. package/dist/hostloop/canusetool-gate.js +4 -4
  28. package/dist/hostloop/safety.js +6 -5
  29. package/dist/hostloop/workspace-handler.js +35 -1
  30. package/dist/io.js +15 -7
  31. package/dist/run/analyze-artifact-runtime.js +3 -4
  32. package/dist/run/analyze-artifact.js +1 -1
  33. package/dist/run/artifacts.js +12 -0
  34. package/dist/run/cassette.js +4 -0
  35. package/dist/run/chat-result.js +14 -7
  36. package/dist/run/chat.js +39 -6
  37. package/dist/run/diff.js +26 -5
  38. package/dist/run/doctor.js +67 -1
  39. package/dist/run/execute.js +206 -67
  40. package/dist/run/inspect-view.js +11 -4
  41. package/dist/run/latest-run.js +33 -7
  42. package/dist/run/migrate-run-dir.js +817 -0
  43. package/dist/run/renderer.js +11 -3
  44. package/dist/run/run-index.js +225 -45
  45. package/dist/run/run.js +13 -4
  46. package/dist/run/runs-gc.js +45 -8
  47. package/dist/run/scaffold.js +28 -4
  48. package/dist/run/skill-files.js +18 -1
  49. package/dist/run/skill-flag-surface.js +13 -2
  50. package/dist/run/subagent-reasoning.js +58 -3
  51. package/dist/run/trace-view.js +110 -12
  52. package/dist/run/turn-events.js +34 -0
  53. package/dist/run/turn-layout.js +189 -0
  54. package/dist/run/verdict.js +17 -2
  55. package/dist/runtime/host-env.js +8 -2
  56. package/dist/runtime/hostloop-prompt.js +23 -4
  57. package/dist/runtime/hostloop.js +67 -20
  58. package/dist/runtime/image-capabilities.js +22 -8
  59. package/dist/runtime/resource-sampler.js +20 -6
  60. package/dist/sync/cowork-sync.js +137 -17
  61. package/dist/types.js +17 -2
  62. package/docker/Dockerfile.agent +9 -0
  63. package/docs/README.md +8 -3
  64. package/docs/boundary.md +1 -1
  65. package/docs/cassette.md +5 -5
  66. package/docs/chat.md +5 -4
  67. package/docs/critique.md +168 -10
  68. package/docs/debugging.md +47 -3
  69. package/docs/fidelity-gaps.md +51 -11
  70. package/docs/gotchas.md +9 -1
  71. package/docs/maintenance.md +2 -2
  72. package/docs/run-status.md +4 -3
  73. package/docs/scenario.md +21 -5
  74. package/docs/session.md +1 -1
  75. package/docs/stats.md +29 -7
  76. package/docs/subagents.md +1 -1
  77. package/examples/README.md +1 -1
  78. package/examples/replays/README.md +1 -1
  79. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  80. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  81. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  82. package/llms.txt +2 -2
  83. package/package.json +1 -1
  84. package/python/README.md +7 -2
  85. package/python/cowork_harness.py +35 -11
  86. package/schema/run-result.json +37 -4
  87. package/schema/scenario.schema.json +1 -1
  88. package/scripts/bump-version.ts +5 -3
  89. package/scripts/check-versions.ts +45 -11
  90. package/scripts/release-preflight.ts +0 -2
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.6.0
7
- tracks-harness: cowork-harness 1.6.0 (baseline desktop-1.22209.3)
6
+ version: 1.8.0
7
+ tracks-harness: cowork-harness 1.8.0 (baseline desktop-1.24012.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.6.0` (baseline
26
- > `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.8.0` (baseline
26
+ > `desktop-1.24012.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,9 +39,52 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.6.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.6.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.6.0"`. **Pin `@>=1.6.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
-
44
- What the ≥ 1.6.0 floor gates, by release:
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.8.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.8.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.8.0"`. **Pin `@>=1.8.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
+
44
+ What the ≥ 1.8.0 floor gates, by release:
45
+
46
+ - **1.8.0 (critique hardening + hostloop unpin):** `critique --skill <name>` (multi-skill-plugin
47
+ grading — a plugin root with no `--skill` is refused pre-spend; the evidence package gains the
48
+ invoked skill's `agents/<name>.md` + bounded `references/*.md` CONTENT), run-dir artifacts
49
+ (`critique-report.json` always, `critique-evidence-package.txt` when the evaluator ran,
50
+ `critique-salvage.json` on exit 2) + `--out <path>`, per-critique `costUsd` across all four
51
+ workloads, `findingFingerprint` per item (cross-INPUT clustering; skillHash stays the cross-FIX
52
+ key), gate-answer `--answer` echo in the report, per-item-tolerant evaluator parse
53
+ (`droppedEvaluatorItems`), `skillMdTruncated` (a readable-but-oversized SKILL.md is flagged
54
+ "graded a cut copy" — distinct from missing/unreadable), `subagents[].webSearches` +
55
+ `trace --view subagent-research` (live/record lane), and the hostloop uploads-bullet fix (Read of
56
+ an uploaded file at the advertised path works — no copy-into-outputs workaround needed). Also:
57
+ **`critique --fidelity hostloop`** (the container-tier pin is lifted — resume-continuity proven
58
+ live on the native binary; `microvm`/`protocol`/`cowork` stay refused, each with a tagged reason),
59
+ **`skill --allow-host-writes`** (hostloop writable-folder consent; `critique` forwards it to both
60
+ turns), `verdictProvenance` stamped on every critique report (advisory self-run, not an
61
+ attestation), evidence caps 64KB SKILL.md / 32KB transcript / 144KB package (a flagship-sized
62
+ SKILL.md grades untruncated; ~2–2.5× evaluator cost on large skills), pinned sessions stamp their
63
+ fidelity and a cross-tier `--resume` fails loud pre-spawn, and `doctor`'s advisory
64
+ `image-freshness` check (pulled agent image behind the published GHCR one → warn + re-pull
65
+ remedy).
66
+
67
+ - **1.7.0 (per-turn run-directory layout, single shape):** every run dir — `run`/`skill`/`chat`,
68
+ single-turn or multi-turn (any `--session-id` + `--resume`, and **every `critique`**: task turn +
69
+ reflection turn) — writes each turn's `result.json` / `run.jsonl` / `trace.json` / `resources.jsonl`
70
+ into **`turns/<N>/`**, written once and never renamed. **There is no root compat copy of anything —
71
+ `<run-dir>/result.json` does not exist.** **When a user asks about a run, point them at
72
+ `turns/1/result.json`** (the only completion on a single-turn dir) — or on a `critique` dir, better yet
73
+ at `result.graded.json` / `trace.graded.json`, the role-stable aliases `critique` writes (the graded
74
+ task turn, never the reflection one). `events.jsonl` / `timeline.jsonl` stay cumulative at the run-dir
75
+ root, always. A dir written before this layout existed is refused LOUD, by name, naming the shape found
76
+ — never silently misread as if it were `turns/1/`. **Convert it in place with
77
+ `cowork-harness migrate-run-dir` (dry-run by default); do NOT tell the user to re-run or re-record,
78
+ which throws away history the migrator recovers.** Its `events.jsonl` still fully supports `trace`,
79
+ which is why every refusal points there. `verify-run` REFUSES a multi-turn dir rather than certifying
80
+ the wrong turn, and `trace` shows the latest turn with a `::notice::` when earlier turns exist.
81
+
82
+ - **1.7.0 (limitation provenance):** every `critique` limitation in `critique --help` and
83
+ docs/critique.md is tagged with WHY it exists — `[structural]` (permanent), `[unverified]` (unproven,
84
+ **not** known-impossible), `[deliberate]`, `[not-built]`. **Read the tag before telling a user to
85
+ design around a limitation.** Worked example of why the tag matters: `critique`'s container-tier pin
86
+ was tagged `[unverified]`, not permanent — and 1.8.0 lifted it (hostloop allowed, see above) exactly
87
+ as the tag predicted a proof could.
45
88
 
46
89
  - **1.6.0 (`critique`'s skill-flag parity):** `critique` accepts most `skill` flags under the same names,
47
90
  which is what makes a document-analysis skill (cap table, deck, financial model, transcript)
@@ -81,7 +124,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
81
124
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
82
125
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
83
126
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
84
- - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — UNRELEASED, see the floor list; `--run-dir` stays global-only everywhere.
127
+ - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
85
128
 
86
129
  ## Orient — the three loops
87
130
 
@@ -113,7 +156,7 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
113
156
  > `node_modules/cowork-harness/docs/<name>.md` before assuming a pointer dangles. A **plugin**
114
157
  > install loads a trimmed source-only cache where those pointers genuinely do dangle.
115
158
 
116
- Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
159
+ Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · migrate-run-dir · lint ·
117
160
  lint-skill · analyze-skill · probe-dispatch ·
118
161
  verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
119
162
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
@@ -155,7 +198,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
155
198
  > is your **working tree**, so an uncommitted edit to an already-tracked file *is* tested — you needn't
156
199
  > commit to iterate. Only brand-new (untracked) files must be `git add`-ed to appear. Commit before you
157
200
  > record the **locking cassette**, though: real Cowork ships the *committed* tree, so a green on
158
- > uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder used to mount *empty* and the agent reported "the skill isn't
201
+ > uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder mounts *empty* and the agent reports "the skill isn't
159
202
  > installed" then did the work itself — a green-looking run where the skill never loaded. That now
160
203
  > **hard-fails** (`BoundaryError`, exit 3) naming the dir, and a partially-tracked folder emits a loud
161
204
  > `::notice:: [stage]` listing the excluded files. Fix: `git add` the skill, or `COWORK_HARNESS_GITSET=0`
@@ -372,7 +415,7 @@ cassette — has its own recipe:
372
415
  "<q>=Israeli company"` binds whichever option starts with `Israeli company`. It is uniqueness-guarded and
373
416
  **fails loud** if the anchor ever matches two options (the documented trade: drift-tolerance, not strict
374
417
  CI reproducibility — for that, pin a full exact label or a free-text `answer:`).
375
- 3. **Budget ~1 re-run per file.** If a gate whiffs, the run no longer vanishes — it exits non-zero but
418
+ 3. **Budget ~1 re-run per file.** If a gate whiffs, the run does not vanish — it exits non-zero but
376
419
  **salvages a PARTIAL run** (the extraction the agent already did is written to disk). So the cost of a
377
420
  missed gate is one re-run with a better `--intent` or a scripted answer, not a lost paid run.
378
421
  4. **Inspect the outputs to judge correctness.** `cowork-harness inspect <run-dir>` shows what the run
@@ -435,6 +478,14 @@ Recognize these before "fixing" a non-bug:
435
478
  `transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
436
479
  run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
437
480
  it's valid.
481
+ - **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
482
+ infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
483
+ rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
484
+ the fail-severity `infra_error`, where a **supervising process** died and contaminated everything. Note
485
+ a model-requested `timeout_ms` expiry is *not* this: it returns the command's partial output with
486
+ `Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
487
+ failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
488
+ suspiciously empty.
438
489
  - **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
439
490
  `RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
440
491
  pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
@@ -638,19 +689,27 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
638
689
  `mcp__workspace__bash` alias, so a "sub-agent used no shell" check must glob **both** `Bash` and
639
690
  `mcp__workspace__*` to hold across every tier.
640
691
 
641
- 6. **`dispatch_count_max` is an author-chosen budget, not a production cap.** It's a post-hoc count
642
- assertion: passing means "happened to dispatch ≤N this run," nothing more. Cowork imposes **no**
643
- in-conversation `Task`-dispatch cap — gate `1648655587`'s `{perTask:1, global:3}` governs the
644
- separate scheduled/cron-task session scheduler, not the `Task` tool (binary-verified; details in
645
- `SPEC.md` §10 — repo-only). So there is no "skip-on-cap" for the harness to reproduce; use this key
646
- only to catch a fan-out you don't want.
692
+ 6. **`dispatch_count_max` is your author-chosen budget UNDER Cowork's production cap, not a
693
+ reproduction of it.** It's a post-hoc count assertion: passing means "happened to dispatch ≤N this
694
+ run." Cowork DOES cap `Task` fan-out **agent-side** (`taskRegistry`: concurrent **20** /
695
+ per-session **200**, landed 2.1.212/2.1.217) — SEPARATE from the scheduled-task session limiter
696
+ (gate `1648655587`'s `{perTask:1, global:3}`, a different mechanism; binary-verified, `SPEC.md` §10
697
+ — repo-only). The harness **inherits** the production cap by spawning the real agent binary, so a
698
+ `dispatch_count_max` pass means "your tighter budget held," not "near a real limit"; use it to catch
699
+ a fan-out you don't want.
647
700
 
648
701
  7. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
649
702
  assertions need a sandboxed tier (`container`+). Good: this one fails loud by design.
650
703
 
651
- 8. **Read-only mounts are enforced; delete-deny is not.** `mode:r` mounts get a real `:ro` bind
652
- (a write fails in-guest). But `rw` vs `rwd` (write-but-no-delete on `outputs/` / connected folders) is
653
- *not* mount-enforced — `rm` succeeds and is only caught post-hoc by `no_delete_in_outputs`.
704
+ 8. **Read-only mounts are enforced; delete-deny is a HARNESS gap — production DOES enforce it.**
705
+ `mode:r` mounts get a real `:ro` bind (a write fails in-guest). But `rw` vs `rwd`
706
+ (write-but-no-delete on `outputs/` / connected folders) is *not* mount-enforced **in the harness** —
707
+ `rm` succeeds and is only caught post-hoc by `no_delete_in_outputs`. **Real Cowork enforces it live:**
708
+ `rm` under `outputs/` fails `Operation not permitted`, and a skill must request approval via
709
+ `allow_cowork_file_delete` to delete. Two consequences: a skill should not stage disposable scratch
710
+ under `outputs/` (in production, cleanup there costs an approval prompt), and a skill's
711
+ "catch-EPERM-then-request-approval" branch cannot be exercised at any harness tier (the `rm` just
712
+ succeeds here). Do not read this gotcha as "delete-deny may not be real in production" — it is real.
654
713
 
655
714
  9. **Keep `.env` out of any mounted folder** — it is copied into the sandbox and the token could
656
715
  leak. Put it at a working-dir or install root (token resolution: env > `--dotenv` > `./.env` >
@@ -727,7 +786,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
727
786
 
728
787
  17. **Editing `scenarios/*.yaml` `assert:` does NOT change a plain `replay`.** *Why:* `replay` evaluates the
729
788
  assertions **frozen in the cassette** by default — it is byte-deterministic and ignores the working tree (so
730
- a committed cassette can't silently re-interpret against an uncommitted YAML). This used to be a *silent*
789
+ a committed cassette can't silently re-interpret against an uncommitted YAML). This is *loud* rather than a *silent*
731
790
  no-op; now plain `replay` prints a `::notice::` when a sibling's `assert:` differs and points you at the fix.
732
791
  *Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
733
792
  `--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.6.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.6.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.8.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -32,7 +32,7 @@ jobs:
32
32
  - uses: actions/checkout@v4
33
33
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
34
34
  run: |
35
- V=2.1.215 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.217 # match your scenario's pinned baseline's agentVersion
36
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
37
37
  chmod +x "$RUNNER_TEMP/claude-$V"
38
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.6.0"
60
+ - run: npm i -g "cowork-harness@>=1.8.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.6.0"
200
+ - run: npm i -g "cowork-harness@>=1.8.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.6.0"
229
+ run: npm i -g "cowork-harness@>=1.8.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -263,8 +263,9 @@ JSON, not the human-readable text (which is explicitly NOT stable).
263
263
 
264
264
  A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
265
265
  set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
266
- artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl`, `run.jsonl`,
267
- `trace.json`, `egress.log`, `result.json`. Digest one with `cowork-harness trace <run-id | dir>`.
266
+ artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl` and
267
+ `egress.log` at the root, plus each turn's `run.jsonl` / `trace.json` / `result.json` under `turns/<N>/`
268
+ (a single-turn run has just `turns/1/`; there is no root compat copy of any of these). Digest one with `cowork-harness trace <run-id | dir>`.
268
269
  Secrets are scrubbed from every persisted log by value.
269
270
 
270
271
  ## Don't assume a fixed assertion count across lanes
@@ -293,7 +294,7 @@ does **not** imply the recording is still valid. Each replay result carries `sta
293
294
  (A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
294
295
  check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
295
296
  `·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
296
- gate (`ok:false`) — but it's no longer class-blind on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
297
+ gate (`ok:false`) — but it is class-AWARE on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
297
298
  `format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
298
299
  while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
299
300
  fail it either way.)
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.6.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -24,6 +24,22 @@ Self-contained reference. Tracks `cowork-harness 1.6.0` (baseline `desktop-1.201
24
24
  `skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
25
25
  positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
26
26
 
27
+ ### hostloop paths: what Read sees vs what bash sees (uploads included)
28
+
29
+ At `hostloop` the native file tools (Read/Write/Edit/Glob/Grep) run **on the host** and are gated by a
30
+ path-containment hook (production's own model), while `mcp__workspace__bash` runs **in the VM sidecar**
31
+ and sees `/sessions/<id>/...` paths. Consequences an agent (or a debugging author) must know:
32
+
33
+ - A `/sessions/...` path handed to `Read` fails ("is a VM path") — that split is **production-faithful**,
34
+ not a harness bug. Translate via the prompt's path-mapping table.
35
+ - **Uploads ARE Read-able at hostloop**: the staged uploads dir on the host is in the containment
36
+ allowlist (read-only — writes/edits there are blocked, matching production). If a Read of an uploads
37
+ path fails "outside this session's connected folders", the path *form* is wrong (e.g. a relative
38
+ `mnt/uploads/...`, which resolves against the outputs cwd) — the file itself is reachable at the
39
+ staged host uploads dir; it does not need to be copied into `outputs/` first.
40
+ - In real Cowork an upload is a **hardlink** in the session's uploads dir; the harness **copies** it
41
+ instead. Read-reachability is the same; only the advertised parent dir can differ.
42
+
27
43
  ### `microvm` prerequisites & lifecycle
28
44
 
29
45
  Requires macOS on arm64 and Lima (`brew install lima`; binary expected at
@@ -122,7 +138,7 @@ hands `fn` exactly this dict.
122
138
 
123
139
  - On `skill`: `fail | prompt | first`.
124
140
  - On `run`: `fail | first` (`prompt` rejected — it would break determinism).
125
- - On `critique` (UNRELEASED): `fail | first` (`prompt` rejected — there is no TTY inside the spawned turn).
141
+ - On `critique`: `fail | first` (`prompt` rejected — there is no TTY inside the spawned turn).
126
142
  - **`llm` is NOT an `--on-unanswered` value.** The bare flag `--on-unanswered llm` is rejected; use
127
143
  `--decider-llm` (CLI) or `on_unanswered: llm` (YAML).
128
144
  - `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.6.0`
4
- (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.8.0`
4
+ (baseline `desktop-1.24012.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -269,7 +269,7 @@ same set live from the schema.
269
269
  | `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
270
270
  | `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
271
271
  | `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
272
- | `dispatch_count_max: <N>` | at most N sub-agents dispatched — an author-chosen budget (Cowork imposes no in-conversation Task-dispatch cap; records only, enforces nothing — see gotcha 12) |
272
+ | `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 |
273
273
  | `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked via the `Skill` tool — evidence-unavailable (not a normal fail) if the agent's init tools have no `Skill` tool |
274
274
  | `no_skill_triggered: <regex>` | no invoked skill id matched — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent or the `Skill` tool is unobservable |
275
275
  | `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
@@ -348,16 +348,17 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
348
348
  | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
349
349
  | `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
350
350
  | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
351
- | `infra_error` | fail | A VM/egress sidecar crashed mid-run — not author-suppressible |
351
+ | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
352
352
  | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
353
353
  | `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
354
354
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
355
355
  | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
356
356
  | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
357
+ | `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
357
358
 
358
359
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
359
360
  overall run verdict and exit code — `assert result: success` alone won't catch it; check
360
- `result.verdict.signals[].severity` or the run's exit code. Only the four **warn** codes are truly benign.
361
+ `result.verdict.signals[].severity` or the run's exit code. Only the five **warn** codes are truly benign.
361
362
 
362
363
  ## Replay class
363
364
 
@@ -417,7 +418,7 @@ skill-source drift; every replay result also reports it class-tagged in `stalene
417
418
  **Mixed assertions on replay:** before evaluating, `replay` strips each assertion to its replay-checkable
418
419
  keys and drops any left empty. So `{result, egress_denied}` evaluates on replay as `{result}` alone — its
419
420
  `egress_denied` half is removed (not AND-ed against an unreadable value); with a manifest, `file_exists`/
420
- `artifact_json` are no longer stripped. The harness is **loud in two classes**: a *full skip* (`::warning::`
421
+ `artifact_json` are not stripped. The harness is **loud in two classes**: a *full skip* (`::warning::`
421
422
  with the count of pure live-only assertions not evaluated) and a *partial skip* (`::warning::` when a mixed
422
423
  assertion's live-only half was dropped).
423
424
  Two CI consequences: skipped assertions are **absent** from `results[].assertions[]` (not
@@ -529,11 +530,12 @@ reference omits). Neither list is a strict superset of the other — reach for t
529
530
  11. **`subagent_declared_but_unused` fires on declared-but-didn't-use-THAT-tool**, even if the
530
531
  sub-agent used other tools.
531
532
 
532
- 12. **`dispatch_count_max` is an author-chosen budget, not a production cap.** It records the count
533
- and asserts on it; passing means "dispatched ≤N this run," not "the harness capped it." Cowork
534
- imposes no in-conversation Task-dispatch cap to reproduce — gate `1648655587`'s
535
- `{perTask:1, global:3}` is the scheduled/cron-task session limiter, a different mechanism
536
- (binary-verified; SPEC §10).
533
+ 12. **`dispatch_count_max` is your author-chosen budget under Cowork's production cap, not a
534
+ reproduction of it.** It records the count and asserts on it; passing means "dispatched ≤N this
535
+ run," not "the harness capped it." Cowork DOES cap `Task` fan-out agent-side (`taskRegistry`:
536
+ concurrent 20 / per-session 200, landed 2.1.212/2.1.217), which the harness inherits by spawning the
537
+ real agent binary — SEPARATE from gate `1648655587`'s `{perTask:1, global:3}` scheduled/cron-task
538
+ session limiter, a different mechanism (binary-verified; SPEC §10).
537
539
 
538
540
  13. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
539
541
  assertions need `container`+. Fails loud by design.
@@ -1,8 +1,8 @@
1
1
  # Task recipes — end-to-end paths for the jobs consumers actually do
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
- decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
5
- front-matter (currently 1.5.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
4
+ decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
+ Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -182,7 +182,7 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
182
182
  (blinded evaluator + mechanical citation checking). **Critiquing a document-analysis skill?** The probe
183
183
  attaches nothing on its own — pass `--upload <path>` (repeatable) or `--folder <dir>` exactly as you
184
184
  would to `skill`, or the graded run has no file and "there was no file attached" is the correct finding,
185
- not a skill defect. Source flags reach both spawned turns automatically. UNRELEASED — see SKILL.md's
185
+ not a skill defect. Source flags reach both spawned turns automatically — see SKILL.md's
186
186
  floor list. See docs/critique.md for the full flag table, cost and limits. If you
187
187
  prefer to build your own grader, the substrate is still here:
188
188
  - `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).