cowork-harness 1.13.2 → 1.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (62) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +71 -17
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
  3. package/.claude/skills/cowork-harness/references/critique.md +3 -2
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +13 -9
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +16 -3
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +6 -3
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +4 -1
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +92 -37
  9. package/AGENTS.md +1 -1
  10. package/CHANGELOG.md +321 -0
  11. package/CONTRIBUTING.md +4 -3
  12. package/DESIGN.md +10 -4
  13. package/README.md +27 -23
  14. package/RELEASING.md +1 -1
  15. package/SPEC.md +1 -1
  16. package/dist/assert.js +35 -16
  17. package/dist/cli.js +207 -26
  18. package/dist/critique/command.js +87 -12
  19. package/dist/critique/limitations.js +0 -9
  20. package/dist/dotenv.js +30 -5
  21. package/dist/egress/proxy.js +1 -1
  22. package/dist/egress/sidecar.js +8 -4
  23. package/dist/hostloop/cowork-handler-hostloop.js +140 -0
  24. package/dist/hostloop/cowork-handler.js +23 -15
  25. package/dist/run/analyze-skill.js +2 -2
  26. package/dist/run/artifacts.js +59 -10
  27. package/dist/run/cassette.js +9 -0
  28. package/dist/run/chat-result.js +2 -0
  29. package/dist/run/doctor.js +8 -4
  30. package/dist/run/execute.js +38 -4
  31. package/dist/run/renderer.js +27 -1
  32. package/dist/run/repeat-flags.js +5 -2
  33. package/dist/run/run-index.js +208 -25
  34. package/dist/run/run.js +28 -1
  35. package/dist/run/skill-flag-surface.js +15 -2
  36. package/dist/run/verdict.js +118 -1
  37. package/dist/runtime/container.js +21 -12
  38. package/dist/runtime/hostloop.js +44 -7
  39. package/dist/session.js +5 -1
  40. package/dist/types.js +21 -2
  41. package/docker/Dockerfile.proxy +4 -2
  42. package/docker/compose.yml +5 -0
  43. package/docs/README.md +2 -2
  44. package/docs/cassette.md +3 -2
  45. package/docs/critique.md +14 -15
  46. package/docs/debugging.md +8 -6
  47. package/docs/fidelity-gaps.md +117 -1
  48. package/docs/gotchas.md +1 -1
  49. package/docs/invariants.md +1 -0
  50. package/docs/run-status.md +3 -0
  51. package/docs/scenario.md +51 -6
  52. package/docs/stats.md +92 -13
  53. package/examples/README.md +7 -5
  54. package/examples/replays/README.md +1 -1
  55. package/llms.txt +3 -1
  56. package/package.json +3 -3
  57. package/python/README.md +3 -2
  58. package/python/test_scenario_lint.py +176 -34
  59. package/schema/critique-report.json +6 -1
  60. package/schema/run-result.json +15 -3
  61. package/schema/scenario.schema.json +16 -2
  62. package/scripts/bump-version.ts +1 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.13.2
7
- tracks-harness: cowork-harness 1.13.2 (baseline desktop-1.24012.9)
6
+ version: 1.14.0
7
+ tracks-harness: cowork-harness 1.14.0 (baseline desktop-1.24012.9)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.13.2` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.14.0` (baseline
26
26
  > `desktop-1.24012.9`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.13.2**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.13.2" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.13.2"`. **Pin `@>=1.13.2`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.14.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.14.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.14.0"`. **Pin `@>=1.14.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
44
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
45
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -284,8 +284,10 @@ quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `
284
284
  (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
285
285
  baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
286
286
  `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
287
- unverifiable), a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) off the
288
- `container` tier (ERROR on `protocol`/`microvm`/`hostloop`, WARN on `cowork`), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
287
+ unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
288
+ `present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
289
+ to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
290
+ `protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
289
291
  and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
290
292
  (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
291
293
  emits a scenario `lint` would reject.
@@ -368,7 +370,12 @@ cassette — has its own recipe:
368
370
  input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
369
371
  pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
370
372
  produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
371
- kept run predates the current skill). **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
373
+ kept run predates the current skill). **The hazard is general, not critique-specific:** repeated
374
+ `run`/`skill` invocations of one scenario accumulate in the SAME scenario directory regardless of skill
375
+ version, so a plain `stats <scenario>` silently averages pre-fix and post-fix runs together. Compare
376
+ generations with **`stats <scenario> --group-by skill-hash`** (or narrow with `--skill-hash <prefix>` /
377
+ `--label <tag>`); an un-split window spanning more than one generation now warns.
378
+ **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
372
379
  skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
373
380
  critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
374
381
  pairing step there silently groups on an absent key instead of erroring — check the field is present, or
@@ -406,6 +413,16 @@ Recognize these before "fixing" a non-bug:
406
413
  covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
407
414
  the final message before acting — a legitimate question-posing answer that wrote a file never fires.
408
415
  Assert `allow_stall: true` if ending on a question is the intended terminal state.
416
+ - **`undelivered_deliverables`** (`WARN`) — the skill produced file(s) **outside every user-visible root**
417
+ and never delivered them. On a **remote** Cowork session the workspace is reclaimed at session end, so
418
+ they are destroyed; on a **local** one they persist but stay invisible to the user. Either way the user
419
+ does not get them. It fires with no assertion written — `present_files_called` covers the positive case
420
+ only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
421
+ **Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
422
+ scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
423
+ Fix by writing deliverables under `outputs/` or a connected folder, or delivering them explicitly; assert
424
+ **`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
425
+ downloaded inputs) rather than a delivery gap.
409
426
  - **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
410
427
  paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
411
428
  `container`/`microvm`, but only *fires* on an actual scanned leak with no authored
@@ -424,7 +441,7 @@ Recognize these before "fixing" a non-bug:
424
441
  `RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
425
442
  pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
426
443
 
427
- The full 14-code signal table (severity + per-signal opt-out) is in
444
+ The full 17-code signal table (severity + per-signal opt-out) is in
428
445
  [`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
429
446
  the fuller narrative.
430
447
 
@@ -440,7 +457,10 @@ harness writes/updates throughout the run's lifecycle (including a crash-safety
440
457
  error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
441
458
  either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
442
459
  checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
443
- run starts, so capture stderr to get the exact directory — but `<dir>` also accepts the run-dir root
460
+ run starts, so capture stderr to get the exact directory — **unless you passed `--compact` (or `--demo`,
461
+ which implies it), which suppress that line** (it is a raw, un-tildeified host path, exactly what those shareable-output
462
+ modes exist to withhold; `status.json` is still written either way, so `status` still works) — but
463
+ `<dir>` also accepts the run-dir root
444
464
  passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
445
465
  newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
446
466
  rather than hanging forever. (Fuller recipe in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) — repo-only, not in the installed
@@ -508,9 +528,28 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
508
528
  call-count/timing table, the sub-agent dispatch tree, the gate lifecycle, the tool/error rollups, …);
509
529
  bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
510
530
  the current list rather than relying on a fixed enumeration here.
531
+ - **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
532
+ is held to. Cowork picks the lane per session ("Run this task: In the cloud / On your computer") and
533
+ cloud is the default for new sessions; the lanes disagree about what *delivered* means. On `remote`,
534
+ location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
535
+ session end), `present_files` is NOT served, and `user_visible_artifact` /
536
+ `present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
537
+ to check a skill's delivery survives the lane most new sessions get. Orthogonal to `fidelity` — a
538
+ `lane: remote` scenario still runs locally.
511
539
  - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
512
- `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`.
513
- - **`result.json` carries the raw fields** the assertions read: `verdict`, `toolDurations`, `models`, `toolErrors`,
540
+ `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
541
+ plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
542
+ `--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
543
+ tier — instead of aggregating across them (a window spanning >1 generation warns — see Gotcha 6;
544
+ >1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
545
+ individual runs behind each summary with their `skillHash`/`runLabel`, so
546
+ you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
547
+ - **`result.json` carries the raw fields** the assertions read: `verdict`, `lane` (which Cowork delivery
548
+ contract the run was held to — see Gotcha 24), `scratchpadEvidenceComplete` (did a COMPLETE scratchpad
549
+ walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
550
+ `total_cost_usd` for the run — the authoritative single-run spend; NOT the same source as summing
551
+ `modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
552
+ `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations`, `models`, `toolErrors`,
514
553
  `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
515
554
  `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
516
555
  `context` (tools/mcpServers/availableSkills), `tasks`,
@@ -539,7 +578,7 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
539
578
  cost spike from fan-out reads as `trace --view dispatches` (how many, which agent) against that model's
540
579
  per-model usage — the harness doesn't line-item each sub-agent's tokens.
541
580
  - **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
542
- `tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input, with a
581
+ `tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input/scratchpad — the last being the agent's working area outside every user-visible root, with a
543
582
  `trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
544
583
  **Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
545
584
  **absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
@@ -767,11 +806,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
767
806
 
768
807
  22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
769
808
  `manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
770
- assertion keys, unconditionally. The linter is **static**: it never reads your cassettes, so it cannot
771
- know whether yours already carry an `artifacts` manifest and `controlOut` (a current cassette does).
772
- On a healthy fleet every one of those lines is a false alarm. *Fix:* `lint --min-severity WARN` in CI
773
- (≥1.11.0) — the INFO advisories stay one flag away for interactive use. `--strict --min-severity ERROR`
774
- behaves as a plain lint, not a contradiction.
809
+ assertion keys. The linter is **static**: it never reads your cassettes, so it cannot know whether
810
+ yours already carry an `artifacts` manifest and `controlOut` (a current cassette does). On a healthy
811
+ fleet every one of those lines is a false alarm. One exception: `manifest-needs-snapshot` is
812
+ suppressed for `user_visible_artifact` on `lane: remote` — the only manifest-backed key that lane also
813
+ rejects outright (`lane-remote-incompatible-key`, an ERROR), so the INFO would be redundant advice
814
+ about a key the scenario can never even load with. `gate-needs-controlout` has no such exception. *Fix:*
815
+ `lint --min-severity WARN` in CI (≥1.11.0) — the INFO advisories stay one flag away for interactive use.
816
+ `--strict --min-severity ERROR` behaves as a plain lint, not a contradiction.
775
817
  23. **`verify-cassettes`/`replay` report a `discovery-surface` note on cassettes you just recorded fine.**
776
818
  *Why:* the cassette froze its `system/init` tool inventory from before the skills/plugins discovery
777
819
  servers existed at that tier (added 1.10.0). It is a non-gating **note**, never a finding — it cannot
@@ -779,6 +821,18 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
779
821
  `mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
780
822
  re-recording would never produce those tools anyway.
781
823
 
824
+ 24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
825
+ lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
826
+ emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
827
+ Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
828
+ optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
829
+ other — and probing a remote session makes this harness look like it emulates the wrong tool under the
830
+ wrong schema. It doesn't; the lanes genuinely disagree. *Fix:* describe the **outcome** ("deliver the
831
+ file to the user") and let the model pick its surface's tool. The `no_scratchpad_leak` /
832
+ `present_files_called` assertion keys are harness-side names and stay valid either way.
833
+ ([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
834
+ → "File delivery" has the binary-verified detail; repo-only.)
835
+
782
836
  For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
783
837
  files in `references/` (the gotchas above are the full list; the references repeat only the
784
838
  assertion/replay-relevant ones).
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`).
3
+ Self-contained reference. Tracks `cowork-harness 1.14.0` (baseline `desktop-1.24012.9`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.13.2"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.14.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -58,7 +58,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
58
58
  GitHub-hosted runners, no token/Docker/agent:
59
59
 
60
60
  ```yaml
61
- - run: npm i -g "cowork-harness@>=1.13.2"
61
+ - run: npm i -g "cowork-harness@>=1.14.0"
62
62
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
63
63
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
64
64
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -219,10 +219,10 @@ jobs:
219
219
  steps:
220
220
  - uses: actions/checkout@v4
221
221
  - uses: actions/setup-node@v4
222
- with: { node-version: '20' }
222
+ with: { node-version: '24' }
223
223
  - uses: actions/setup-python@v5
224
224
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
225
- - run: npm i -g "cowork-harness@>=1.13.2"
225
+ - run: npm i -g "cowork-harness@>=1.14.0"
226
226
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
227
227
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
228
228
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -251,7 +251,7 @@ jobs:
251
251
  echo "live=true" >> "$GITHUB_OUTPUT"
252
252
  fi
253
253
  - if: steps.guard.outputs.live == 'true'
254
- run: npm i -g "cowork-harness@>=1.13.2"
254
+ run: npm i -g "cowork-harness@>=1.14.0"
255
255
  - if: steps.guard.outputs.live == 'true'
256
256
  run: cowork-harness run scenarios/ --output-format json
257
257
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 1.14.0` (baseline `desktop-1.24012.9`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -30,7 +30,8 @@ turn rows carry `critiqueRole:"task"` / `"reflection"`.
30
30
 
31
31
  Two things a harvester needs: roll-ups are excluded from `stats` aggregation (they carry no verdict —
32
32
  counting them adds a phantom run and drags `passRate` toward 1), so **filter them out of any pass-rate
33
- computed over raw rows**; and a roll-up with `result:"error"` had an unpriced workload, so its totals
33
+ computed over raw rows** — that exclusion also governs `stats --group-by skill-hash`, so a per-generation
34
+ **total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
34
35
  UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
35
36
 
36
37
  ## The report's item shape — no `title`, no `summary`
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`).
3
+ Self-contained reference. Tracks `cowork-harness 1.14.0` (baseline `desktop-1.24012.9`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -211,14 +211,18 @@ up often enough to spell out:
211
211
  (`staleness[]` flags skill/baseline drift as a hint; only a live `run` re-confirms current
212
212
  behavior). See [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § "Still skipped on replay" and [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) § "Which
213
213
  assertions survive replay."
214
- - **Container-only assertions can't verify off the `container` tier.** `no_scratchpad_leak` and
215
- `present_files_called` check the `present_files` delivery path, which is served **only** on
216
- `container` — not `hostloop`/`microvm`. Asserting them off-container hard-fails at runtime (a red
217
- run, not a false green), so you won't be fooled if you write the assertion. The quieter trap is a
218
- scenario that runs at `hostloop`/`microvm`/`protocol` and simply omits these assertions: a green
219
- run there proves **nothing** about scratchpad-leak safety or present_files delivery, because that
220
- tier never exercises the delivery path the assertions would check. Use `fidelity: container` for
221
- present_files/scratchpad-delivery coverage.
214
+ - **`present_files` assertions can't verify off the tiers that serve the tool.** `no_scratchpad_leak` and
215
+ `present_files_called` check the `present_files` delivery path — the desktop-local lane's tool; remote
216
+ Cowork delivers via the agent-native `SendUserFile` instead, so never hardcode a delivery tool name in
217
+ a SKILL.md (SKILL.md Gotcha 24). **The harness** serves `present_files` on `container` **and `hostloop`**
218
+ — not `microvm`/`protocol`. `present_files_called` works at both; `no_scratchpad_leak` stays
219
+ `container`-only, because hostloop's handler passes a validated path through without promoting, so
220
+ there is no scratch→outputs copy to leak. Asserting either key off the tiers that serve it hard-fails at
221
+ runtime (a red run, not a false green), so you won't be fooled if you write the assertion. The quieter
222
+ trap is a scenario that runs at `microvm`/`protocol` and simply omits these assertions: a green run
223
+ there proves **nothing** about scratchpad-leak safety or present_files delivery, because that tier
224
+ never exercises the delivery path the assertions would check. Use `fidelity: container` or `hostloop`
225
+ for present_files-delivery coverage; `no_scratchpad_leak` still needs `container` specifically.
222
226
  - **The harness doesn't observe rendered-artifact interactions, browser downloads, or human
223
227
  clicks.** It runs the agent headless — no webview, no browser, no person clicking "Submit." A
224
228
  class of Cowork bug (a client-side write-back to a relative URL that resolves-but-fails against
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.13.2`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.14.0`
4
4
  (baseline `desktop-1.24012.9`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -37,6 +37,17 @@ fidelity: container # protocol | container | microvm | hostloop
37
37
  execution: local # OPTIONAL — orthogonal to fidelity (a privilege/sandbox tier, all
38
38
  # local): local (default) | cloud-describe (RESERVED — no runner
39
39
  # exists yet; authoring it is a load-time error, not a silent no-op)
40
+ lane: local # OPTIONAL — which Cowork lane's DELIVERY CONTRACT the run is held to:
41
+ # local (default) | remote. On `remote`, location delivers NOTHING (a
42
+ # remote container has no auto-delivering outputs dir and is reclaimed
43
+ # at session end) and `present_files` is NOT served (a local MCP server
44
+ # can't reach a remote session) — so `user_visible_artifact` and
45
+ # present_files_called/no_scratchpad_leak are REJECTED AT SCENARIO-LOAD
46
+ # TIME (before the run starts, before any spend), not left to fail
47
+ # unverifiable/can't-verify at assertion time. Orthogonal to fidelity
48
+ # and execution — a `lane: remote` scenario still runs locally.
49
+ # Delivery semantics only; the remote device bridge is deliberately
50
+ # unmodeled.
40
51
  on_unanswered: fail # policy for unscripted gates: fail | prompt | first | llm — run rejects prompt
41
52
  # ("agent" is retired — no longer a valid value)
42
53
 
@@ -298,8 +309,8 @@ same set live from the schema.
298
309
  | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
299
310
  | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
300
311
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
301
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **`fidelity: container` only** — `present_files` is not served on hostloop/microvm, so a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). **Only `true` is valid** |
302
- | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` only** — `present_files` is not served on hostloop/microvm. **Only `true` is valid** |
312
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
313
+ | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
303
314
  | `question_asked: <regex>` | the agent asked an AskUserQuestion whose text matches |
304
315
  | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
305
316
  | `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min` to also require a gate |
@@ -314,6 +325,7 @@ same set live from the schema.
314
325
  | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
315
326
  | `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
316
327
  | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider) |
328
+ | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
317
329
  | `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
318
330
  | `egress_denied: <host>` | the host was blocked by the egress proxy |
319
331
  | `egress_allowed: <host>` | the host was allowed through |
@@ -358,6 +370,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
358
370
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
359
371
  | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
360
372
  | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
373
+ | `undelivered_deliverables` | warn | The skill produced file(s) OUTSIDE every user-visible root and never delivered them — on a remote Cowork session the workspace is reclaimed at session end; on a local one they stay invisible to the user. Fires on every run without opting in, because the scenarios that most need it are the ones whose author never considered delivery. Silent when the evidence cannot answer the question (no workspace walk, a tier that runs no scratchpad walk, absent delivery telemetry, or a resumed turn) — never a vacuous clean. Opt out: `allow_undelivered_deliverables` |
361
374
  | `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
362
375
 
363
376
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 1.14.0` (baseline `desktop-1.24012.9`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -118,7 +118,8 @@ have:
118
118
  4. **Replay-class note:** `max_turns` / `tool_calls_max` / `dispatch_count_max` re-evaluate
119
119
  token-free on replay; `questions_count_max` needs `controlOut` (Recipe 1's tree applies).
120
120
  `max_cost_usd` / `max_tokens` on replay assert the FROZEN recording's spend — near-zero signal
121
- as a regression gate; if you need a cost gate, put it on the live lane via `stats`.
121
+ as a regression gate; if you need a cost gate, put it on the live lane via `stats` (history) or
122
+ `--max-budget-usd` (a pre-flight refusal before the run spends).
122
123
 
123
124
  ## Recipe 5 — Evaluate your skill's ANSWER QUALITY (semantic regression gate)
124
125
 
@@ -174,7 +175,9 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
174
175
  output. **Reproduce before acting on a finding:** `cowork-harness skill <folder> "<prompt>" --repeat 5 --label gen-1`
175
176
  runs the same skill+prompt N times (2-100) and prints a variance rollup instead of a single pass/fail —
176
177
  `--repeat` works on the `skill` lane, not just `run`. A single green run proves it passed *once*.
177
- Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`. It rejects `--session-id`/
178
+ Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`.
179
+ (`--max-budget-usd` also works WITHOUT `--repeat`, where it pre-flights a single run against that
180
+ scenario's cost history and refuses before spending; the others are batch-only.) It rejects `--session-id`/
178
181
  `--resume` (both pin one run dir) and `--decider-cmd`/`--decider-dir` (a driving agent x N is not a
179
182
  measurement).
180
183
  The full loop (harvest -> reproduce -> fix -> prove freshness -> compare) is written out end-to-end in
@@ -6,6 +6,7 @@
6
6
  "allow_missing_capability",
7
7
  "allow_permissive_auto_allow",
8
8
  "allow_stall",
9
+ "allow_undelivered_deliverables",
9
10
  "artifact_json",
10
11
  "compaction_occurred",
11
12
  "computer_links_resolve",
@@ -80,6 +81,7 @@
80
81
  "execution",
81
82
  "expect_denied",
82
83
  "fidelity",
84
+ "lane",
83
85
  "name",
84
86
  "on_unanswered",
85
87
  "prompt",
@@ -92,6 +94,7 @@
92
94
  "allow_l0_plugin_divergence",
93
95
  "allow_missing_capability",
94
96
  "allow_permissive_auto_allow",
95
- "allow_stall"
97
+ "allow_stall",
98
+ "allow_undelivered_deliverables"
96
99
  ]
97
100
  }
@@ -17,8 +17,10 @@ Two subcommands:
17
17
  lint flags (see references/scenario-schema.md for the why of each):
18
18
  E egress assertion on `fidelity: protocol` (the harness rejects this run)
19
19
  E `transcript_no_host_path` on hostloop/protocol (fails BY DESIGN at those tiers)
20
- E `no_scratchpad_leak`/`present_files_called` off (container-only: served only at fidelity:
21
- protocol/microvm/hostloop container; off-container = cannot-verify)
20
+ E `no_scratchpad_leak` off container (container-only: hostloop serves the tool but never promotes)
21
+ E `present_files_called` on protocol/microvm (served at container+hostloop)
22
+ E `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote`
23
+ (runtime rejects at LOAD time; tier rules suppressed)
22
24
  E `requires_capabilities` on `fidelity: protocol` (probe can't run → hard-fails
23
25
  unless allow_missing_capability)
24
26
  E on_unanswered: agent / invalid value (schema rejects `agent`)
@@ -26,8 +28,7 @@ lint flags (see references/scenario-schema.md for the why of each):
26
28
  E `assertions:` instead of `assert:` (block ignored → every check no-ops)
27
29
  W `transcript_no_host_path` on `fidelity: cowork` (tier resolves per baseline gate —
28
30
  incompatible if it lands hostloop)
29
- W `no_scratchpad_leak`/`present_files_called` on (tier resolves per baseline gate — cannot-verify
30
- `fidelity: cowork` if it resolves off-container)
31
+ W `no_scratchpad_leak` on `fidelity: cowork` (tier resolves per baseline gate)
31
32
  W no content assertion → no-op on a replay gate (every assertion is fs/egress)
32
33
  W mixed-class assert item → fs/egress half dropped on replay
33
34
  W unknown top-level / assertion key (typo or hallucinated schema)
@@ -135,16 +136,26 @@ LIVE_ONLY_KEYS = {
135
136
  "no_lost_write_back",
136
137
  }
137
138
  EGRESS_KEYS = {"egress_denied", "egress_allowed"}
138
- # container-only: served only at fidelity: container (present_files / the scratchpad promotion path
139
- # it depends on). Off-container these report cannot-verify, not a meaningful pass/fail — same tier-fidelity
140
- # class as transcript_no_host_path below, just the opposite direction (container-only vs container-hostile).
141
- CONTAINER_ONLY_KEYS = {"no_scratchpad_leak", "present_files_called"}
139
+ # container-only: promotion/leak semantics exist only at fidelity: container. Production's host-loop
140
+ # branch validates a path and passes it through WITHOUT promoting, so at hostloop there is no
141
+ # scratch->outputs copy that could ever leak -- cannot-verify, not a meaningful pass/fail.
142
+ CONTAINER_ONLY_KEYS = {"no_scratchpad_leak"}
143
+ # container+hostloop: the harness serves present_files at BOTH tiers, so the DELIVERY RECORD this key
144
+ # asserts is meaningful at both. Only protocol (no /sessions/ layout) and microvm (stages into a
145
+ # different tree than the artifact scan walks) deterministically cannot serve it.
146
+ CONTAINER_HOSTLOOP_KEYS = {"present_files_called"}
147
+ # `lane: remote` + any of these is rejected at scenario LOAD time by the runtime (src/run/execute.ts):
148
+ # that lane serves no present_files and delivers nothing by location, so the key can only ever report
149
+ # cannot-verify. Catching it offline is the whole point of the linter -- otherwise the author finds out
150
+ # when a paid run refuses to start.
151
+ LANE_REMOTE_INCOMPATIBLE_KEYS = {"present_files_called", "no_scratchpad_leak", "user_visible_artifact"}
142
152
  # verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
143
153
  VERDICT_MODIFIER_KEYS = {
144
154
  "allow_permissive_auto_allow",
145
155
  "allow_l0_plugin_divergence",
146
156
  "allow_missing_capability",
147
157
  "allow_stall",
158
+ "allow_undelivered_deliverables",
148
159
  }
149
160
 
150
161
  # Every key the replay-class logic knows how to handle. `replay_protocol_fidelity` is valid-but-not-authorable
@@ -192,6 +203,7 @@ _EMBEDDED_TOP_LEVEL_KEYS = {
192
203
  "session",
193
204
  "fidelity",
194
205
  "execution", # execution-location axis, orthogonal to fidelity; "cloud-describe" is reserved (load-time error)
206
+ "lane", # Cowork product-lane axis: which delivery contract the run is held to (local | remote)
195
207
  "on_unanswered",
196
208
  "prompt",
197
209
  "timeout_ms", # wall-clock budget → kill + errorSource:timeout on expiry
@@ -310,6 +322,7 @@ def lint_doc(doc, path, raw_lines):
310
322
  return findings
311
323
 
312
324
  fidelity = (doc.get("fidelity") or "container")
325
+ lane = (doc.get("lane") or "local")
313
326
  items = _assert_items(doc)
314
327
  assert_keys = _all_assert_keys(items)
315
328
  has_expect_denied = bool(doc.get("expect_denied"))
@@ -402,37 +415,72 @@ def lint_doc(doc, path, raw_lines):
402
415
  )
403
416
  )
404
417
 
405
- # E/W: CONTAINER_ONLY_KEYS (no_scratchpad_leak / present_files_called) are served only at
406
- # fidelity: container — off-container they report cannot-verify, not a meaningful check. Mirrors
407
- # the transcript_no_host_path tier check above, just the opposite direction: ERROR on the tiers
408
- # where the runtime deterministically can't serve it, WARN on cowork (baseline-gate-resolution
409
- # dependent — the linter stays offline, so the message names the dependency instead of resolving it).
410
- container_only_present = sorted(assert_keys & CONTAINER_ONLY_KEYS)
411
- if container_only_present:
412
- if fidelity in ("protocol", "microvm", "hostloop"):
413
- findings.append(
414
- Finding(
415
- "ERROR",
416
- "container-only-key-off-container",
417
- f"{container_only_present} on `fidelity: {fidelity}` — container-only (present_files "
418
- "is served only there); off-container it reports cannot-verify. Use fidelity: "
419
- "container for present_files/scratchpad delivery.",
420
- "Use fidelity: container (or drop the assertion for this tier).",
421
- path,
422
- )
418
+ # E: keys the runtime rejects outright on `lane: remote` (load-time throw, before any spend).
419
+ lane_incompatible = sorted(assert_keys & LANE_REMOTE_INCOMPATIBLE_KEYS)
420
+ if lane == "remote" and lane_incompatible:
421
+ findings.append(
422
+ Finding(
423
+ "ERROR",
424
+ "lane-remote-incompatible-key",
425
+ f"{lane_incompatible} on `lane: remote` -- that lane serves no present_files and "
426
+ "delivers nothing by location, so these keys can only report cannot-verify. The "
427
+ "runtime rejects this at scenario LOAD time, before the run starts.",
428
+ "Assert the delivery itself, or set `lane: local` if this scenario models the desktop lane.",
429
+ path,
423
430
  )
424
- elif fidelity == "cowork":
431
+ )
432
+
433
+ # `lane: remote` already rejected these above; tier advice there is unreachable (the lane check
434
+ # fires first, at load, regardless of tier) -- do not tell an author to change a tier that cannot help.
435
+ if lane != "remote":
436
+ # E/W: CONTAINER_ONLY_KEYS (no_scratchpad_leak) -- promotion/leak semantics are container-shaped.
437
+ # NOT a harness coverage gap: the harness serves present_files at hostloop too, but production's
438
+ # host-loop branch validates a path and passes it through WITHOUT promoting, so there is no
439
+ # scratch->outputs copy that could ever leak there.
440
+ container_only_present = sorted(assert_keys & CONTAINER_ONLY_KEYS)
441
+ if container_only_present:
442
+ if fidelity in ("protocol", "microvm", "hostloop"):
443
+ findings.append(
444
+ Finding(
445
+ "ERROR",
446
+ "container-only-key-off-container",
447
+ f"{container_only_present} on `fidelity: {fidelity}` -- promotion/leak semantics "
448
+ "apply only at the container tier (hostloop serves present_files but never "
449
+ "promotes, so there is no scratch->outputs copy to leak); off container it "
450
+ "reports cannot-verify.",
451
+ "Use fidelity: container (or drop the assertion for this tier).",
452
+ path,
453
+ )
454
+ )
455
+ elif fidelity == "cowork":
456
+ findings.append(
457
+ Finding(
458
+ "WARN",
459
+ "container-only-key-off-container",
460
+ f"{container_only_present} on `fidelity: cowork` -- the tier resolves to "
461
+ f"container or hostloop per gate {HOST_LOOP_GATE_ID}; on hostloop this key "
462
+ "reports cannot-verify (no promotion happens there).",
463
+ "Pin fidelity: container if you need this asserted deterministically.",
464
+ path,
465
+ )
466
+ )
467
+
468
+ # E: CONTAINER_HOSTLOOP_KEYS (present_files_called) -- served at container AND hostloop, so only
469
+ # protocol and microvm are flagged. Deliberately NO `cowork` arm: cowork resolves to
470
+ # hostloop|container ONLY (src/run/execute.ts), and both serve the tool -- an advisory there would
471
+ # fire on every applicable scenario with no inapplicable case to distinguish (AGENTS.md, Advisory
472
+ # design: "actionable by construction, or aggregated").
473
+ present_files_present = sorted(assert_keys & CONTAINER_HOSTLOOP_KEYS)
474
+ if present_files_present and fidelity in ("protocol", "microvm"):
425
475
  findings.append(
426
476
  Finding(
427
- "WARN",
428
- "container-only-key-off-container",
429
- f"{container_only_present} on `fidelity: cowork` — container-only (present_files "
430
- "is served only there); off-container it reports cannot-verify. Use fidelity: "
431
- f"container for present_files/scratchpad delivery. (The tier resolves per the "
432
- f"baseline's host-loop gate ({HOST_LOOP_GATE_ID}); if it resolves off-container this "
433
- "assertion cannot verify anything.)",
434
- "Pin fidelity: container if the assertion is load-bearing; keep cowork only if "
435
- "you accept the gate-resolution dependency.",
477
+ "ERROR",
478
+ "present-files-key-off-tier",
479
+ f"{present_files_present} on `fidelity: {fidelity}` -- present_files is served only "
480
+ "at container/hostloop (protocol has no /sessions/ layout for the handler's path "
481
+ "model; microvm stages into a different tree than the artifact scan walks); there it "
482
+ "reports cannot-verify.",
483
+ "Use fidelity: container or hostloop (or drop the assertion for this tier).",
436
484
  path,
437
485
  )
438
486
  )
@@ -523,8 +571,15 @@ def lint_doc(doc, path, raw_lines):
523
571
  )
524
572
  )
525
573
 
526
- # I: manifest-backed keys need an artifacts manifest on replay
574
+ # I: manifest-backed keys need an artifacts manifest on replay. On `lane: remote`, a key that is ALSO
575
+ # in LANE_REMOTE_INCOMPATIBLE_KEYS (user_visible_artifact) already got the ERROR above and is rejected
576
+ # at scenario-LOAD time -- it can never reach a replay to re-record for, so "re-record so it evaluates"
577
+ # is unreachable advice for that key (same rationale as the tier-rule suppression above). Filtered
578
+ # per-key, not the whole block: the other manifest keys (file_exists, artifact_json, ...) are NOT
579
+ # lane-rejected and stay genuinely reachable and worth advising about on `lane: remote`.
527
580
  manifest_present = sorted(assert_keys & MANIFEST_KEYS)
581
+ if lane == "remote":
582
+ manifest_present = [k for k in manifest_present if k not in LANE_REMOTE_INCOMPATIBLE_KEYS]
528
583
  if manifest_present:
529
584
  findings.append(
530
585
  Finding(
package/AGENTS.md CHANGED
@@ -23,7 +23,7 @@ protocol layer or run-loop bookkeeping in the CLI.
23
23
  `npm run test:live` lane. Python fast lane (from `python/`): `pytest -m 'not cowork'`.
24
24
  - CLI binary `cowork-harness`; env vars `COWORK_HARNESS_*` (+ `COWORK_AGENT_BINARY` / `COWORK_AGENT_IMAGE`) —
25
25
  see README's [Reproducibility knobs](./README.md#reproducibility-knobs) for the full env-var list.
26
- Node ≥ 20.
26
+ Node ≥ 22.
27
27
  - `cowork-harness sync` is **local-only** (needs Desktop + `app.asar`; not on CI). The committed
28
28
  `baselines/*.json` are CI's source of truth — never hand-edit release facts into source; they come from
29
29
  `sync` (see `docs/maintenance.md`).