cowork-harness 1.13.2 → 1.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (63) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +87 -21
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +10 -8
  3. package/.claude/skills/cowork-harness/references/critique.md +3 -2
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +13 -9
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +25 -6
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +11 -3
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +4 -1
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +92 -37
  9. package/AGENTS.md +1 -1
  10. package/CHANGELOG.md +410 -0
  11. package/CONTRIBUTING.md +4 -3
  12. package/DESIGN.md +10 -4
  13. package/README.md +32 -25
  14. package/RELEASING.md +14 -7
  15. package/SPEC.md +1 -1
  16. package/dist/assert.js +35 -16
  17. package/dist/cli.js +183 -26
  18. package/dist/critique/command.js +87 -12
  19. package/dist/critique/limitations.js +0 -9
  20. package/dist/dotenv.js +30 -5
  21. package/dist/egress/proxy.js +1 -1
  22. package/dist/egress/sidecar.js +8 -4
  23. package/dist/hostloop/cowork-handler-hostloop.js +140 -0
  24. package/dist/hostloop/cowork-handler.js +23 -15
  25. package/dist/run/analyze-skill.js +2 -2
  26. package/dist/run/artifacts.js +59 -10
  27. package/dist/run/budget.js +119 -0
  28. package/dist/run/cassette.js +180 -14
  29. package/dist/run/chat-result.js +2 -0
  30. package/dist/run/doctor.js +8 -4
  31. package/dist/run/execute.js +38 -4
  32. package/dist/run/renderer.js +27 -1
  33. package/dist/run/repeat-flags.js +5 -2
  34. package/dist/run/run-index.js +208 -25
  35. package/dist/run/run.js +28 -1
  36. package/dist/run/skill-flag-surface.js +15 -2
  37. package/dist/run/verdict.js +118 -1
  38. package/dist/runtime/container.js +21 -12
  39. package/dist/runtime/hostloop.js +44 -7
  40. package/dist/session.js +5 -1
  41. package/dist/types.js +21 -2
  42. package/docker/Dockerfile.proxy +4 -2
  43. package/docker/compose.yml +5 -0
  44. package/docs/README.md +2 -2
  45. package/docs/cassette.md +14 -8
  46. package/docs/critique.md +14 -15
  47. package/docs/debugging.md +8 -6
  48. package/docs/fidelity-gaps.md +117 -1
  49. package/docs/gotchas.md +1 -1
  50. package/docs/invariants.md +1 -0
  51. package/docs/run-status.md +3 -0
  52. package/docs/scenario.md +120 -8
  53. package/docs/stats.md +92 -13
  54. package/examples/README.md +7 -5
  55. package/examples/replays/README.md +1 -1
  56. package/llms.txt +3 -1
  57. package/package.json +3 -3
  58. package/python/README.md +3 -2
  59. package/python/test_scenario_lint.py +176 -34
  60. package/schema/critique-report.json +6 -1
  61. package/schema/run-result.json +15 -3
  62. package/schema/scenario.schema.json +16 -2
  63. package/scripts/bump-version.ts +1 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.13.2
7
- tracks-harness: cowork-harness 1.13.2 (baseline desktop-1.24012.9)
6
+ version: 1.15.0
7
+ tracks-harness: cowork-harness 1.15.0 (baseline desktop-1.24012.9)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.13.2` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.15.0` (baseline
26
26
  > `desktop-1.24012.9`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.13.2**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.13.2" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.13.2"`. **Pin `@>=1.13.2`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.15.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.15.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.15.0"`. **Pin `@>=1.15.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
44
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
45
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -284,12 +284,22 @@ quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `
284
284
  (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
285
285
  baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
286
286
  `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
287
- unverifiable), a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) off the
288
- `container` tier (ERROR on `protocol`/`microvm`/`hostloop`, WARN on `cowork`), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
287
+ unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
288
+ `present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
289
+ to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
290
+ `protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
289
291
  and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
290
292
  (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
291
293
  emits a scenario `lint` would reject.
292
294
 
295
+ **`lint` is the LENIENT check — the loader is the strict one.** An unknown top-level key is a ⚠ WARN in
296
+ `lint` (exit 0) but a **hard error** in the runtime (`Unrecognized key: "<k>"`, exit 2) — so a scenario
297
+ that lints with warnings may still not run. To check whether a scenario actually loads, without spending:
298
+ `cowork-harness record <file.yaml> --dry-run` (exit 2 on a schema error; a directory reports each
299
+ `✗ broken:` file and exits 1). Corollary: **a key from a newer harness fails LOUD on an older CLI**, never
300
+ silently — `lane:` on a pre-1.14.0 CLI does not load rather than defaulting to `lane: local`. That is why a
301
+ new scenario key ships with a version floor.
302
+
293
303
  ## Part II — RUN, RECORD & LOCK
294
304
 
295
305
  You have an authored scenario. This Part runs it, reads the verdict, locks it into a
@@ -368,7 +378,12 @@ cassette — has its own recipe:
368
378
  input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
369
379
  pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
370
380
  produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
371
- kept run predates the current skill). **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
381
+ kept run predates the current skill). **The hazard is general, not critique-specific:** repeated
382
+ `run`/`skill` invocations of one scenario accumulate in the SAME scenario directory regardless of skill
383
+ version, so a plain `stats <scenario>` silently averages pre-fix and post-fix runs together. Compare
384
+ generations with **`stats <scenario> --group-by skill-hash`** (or narrow with `--skill-hash <prefix>` /
385
+ `--label <tag>`); an un-split window spanning more than one generation now warns.
386
+ **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
372
387
  skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
373
388
  critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
374
389
  pairing step there silently groups on an absent key instead of erroring — check the field is present, or
@@ -406,6 +421,16 @@ Recognize these before "fixing" a non-bug:
406
421
  covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
407
422
  the final message before acting — a legitimate question-posing answer that wrote a file never fires.
408
423
  Assert `allow_stall: true` if ending on a question is the intended terminal state.
424
+ - **`undelivered_deliverables`** (`WARN`) — the skill produced file(s) **outside every user-visible root**
425
+ and never delivered them. On a **remote** Cowork session the workspace is reclaimed at session end, so
426
+ they are destroyed; on a **local** one they persist but stay invisible to the user. Either way the user
427
+ does not get them. It fires with no assertion written — `present_files_called` covers the positive case
428
+ only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
429
+ **Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
430
+ scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
431
+ Fix by writing deliverables under `outputs/` or a connected folder, or delivering them explicitly; assert
432
+ **`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
433
+ downloaded inputs) rather than a delivery gap.
409
434
  - **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
410
435
  paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
411
436
  `container`/`microvm`, but only *fires* on an actual scanned leak with no authored
@@ -424,7 +449,7 @@ Recognize these before "fixing" a non-bug:
424
449
  `RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
425
450
  pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
426
451
 
427
- The full 14-code signal table (severity + per-signal opt-out) is in
452
+ The full 17-code signal table (severity + per-signal opt-out) is in
428
453
  [`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
429
454
  the fuller narrative.
430
455
 
@@ -440,7 +465,10 @@ harness writes/updates throughout the run's lifecycle (including a crash-safety
440
465
  error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
441
466
  either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
442
467
  checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
443
- run starts, so capture stderr to get the exact directory — but `<dir>` also accepts the run-dir root
468
+ run starts, so capture stderr to get the exact directory — **unless you passed `--compact` (or `--demo`,
469
+ which implies it), which suppress that line** (it is a raw, un-tildeified host path, exactly what those shareable-output
470
+ modes exist to withhold; `status.json` is still written either way, so `status` still works) — but
471
+ `<dir>` also accepts the run-dir root
444
472
  passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
445
473
  newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
446
474
  rather than hanging forever. (Fuller recipe in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) — repo-only, not in the installed
@@ -508,9 +536,28 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
508
536
  call-count/timing table, the sub-agent dispatch tree, the gate lifecycle, the tool/error rollups, …);
509
537
  bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
510
538
  the current list rather than relying on a fixed enumeration here.
539
+ - **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
540
+ is held to. Cowork picks the lane per session ("Run this task: In the cloud / On your computer") and
541
+ cloud is the default for new sessions; the lanes disagree about what *delivered* means. On `remote`,
542
+ location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
543
+ session end), `present_files` is NOT served, and `user_visible_artifact` /
544
+ `present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
545
+ to check a skill's delivery survives the lane most new sessions get. Orthogonal to `fidelity` — a
546
+ `lane: remote` scenario still runs locally.
511
547
  - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
512
- `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`.
513
- - **`result.json` carries the raw fields** the assertions read: `verdict`, `toolDurations`, `models`, `toolErrors`,
548
+ `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
549
+ plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
550
+ `--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
551
+ tier — instead of aggregating across them (a window spanning >1 generation warns — see Gotcha 6;
552
+ >1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
553
+ individual runs behind each summary with their `skillHash`/`runLabel`, so
554
+ you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
555
+ - **`result.json` carries the raw fields** the assertions read: `verdict`, `lane` (which Cowork delivery
556
+ contract the run was held to — see Gotcha 24), `scratchpadEvidenceComplete` (did a COMPLETE scratchpad
557
+ walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
558
+ `total_cost_usd` for the run — the authoritative single-run spend; NOT the same source as summing
559
+ `modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
560
+ `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations`, `models`, `toolErrors`,
514
561
  `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
515
562
  `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
516
563
  `context` (tools/mcpServers/availableSkills), `tasks`,
@@ -539,7 +586,7 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
539
586
  cost spike from fan-out reads as `trace --view dispatches` (how many, which agent) against that model's
540
587
  per-model usage — the harness doesn't line-item each sub-agent's tokens.
541
588
  - **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
542
- `tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input, with a
589
+ `tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input/scratchpad — the last being the agent's working area outside every user-visible root, with a
543
590
  `trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
544
591
  **Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
545
592
  **absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
@@ -723,10 +770,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
723
770
  isolated** (its own egress sidecar network + proxy, its own per-session run dir), so parallel records don't
724
771
  cross-talk — the concurrency bound exists only for the Docker address pool + API rate limits, not correctness.
725
772
 
726
- 17. **Editing `scenarios/*.yaml` `assert:` does NOT change a plain `replay`.** *Why:* `replay` evaluates the
727
- assertions **frozen in the cassette** by default — it is byte-deterministic and ignores the working tree (so
728
- a committed cassette can't silently re-interpret against an uncommitted YAML). This is *loud* rather than a *silent*
729
- no-op; now plain `replay` prints a `::notice::` when a sibling's `assert:` differs and points you at the fix.
773
+ 17. **Editing `scenarios/*.yaml` does NOT change a plain `replay` — the WHOLE scenario is frozen, not just
774
+ `assert:`.** *Why:* a cassette captures every key (`lane:`, `fidelity:`, `baseline:`, `prompt:`, `skills:` …)
775
+ and `replay` evaluates all of them from that frozen copy — byte-deterministic, ignoring the working tree (so
776
+ a committed cassette can't silently re-interpret against an uncommitted YAML). **Only `assert:`
777
+ (+`expect_denied:`) can be opted back to disk; any other edited key reaches a replay only by re-recording.**
778
+ This is *loud* rather than a *silent* no-op: plain `replay` prints a `::notice::` when a sibling's
779
+ `assert:`/`prompt:` differs, and when the sibling **fails to load** at all (a typo'd or too-new key) — and
780
+ points you at the fix.
730
781
  *Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
731
782
  `--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
732
783
  `prompt`/`answers`/`baseline`/`fidelity`/`skills`/`requires_capabilities` or the skill content (when a
@@ -767,11 +818,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
767
818
 
768
819
  22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
769
820
  `manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
770
- assertion keys, unconditionally. The linter is **static**: it never reads your cassettes, so it cannot
771
- know whether yours already carry an `artifacts` manifest and `controlOut` (a current cassette does).
772
- On a healthy fleet every one of those lines is a false alarm. *Fix:* `lint --min-severity WARN` in CI
773
- (≥1.11.0) — the INFO advisories stay one flag away for interactive use. `--strict --min-severity ERROR`
774
- behaves as a plain lint, not a contradiction.
821
+ assertion keys. The linter is **static**: it never reads your cassettes, so it cannot know whether
822
+ yours already carry an `artifacts` manifest and `controlOut` (a current cassette does). On a healthy
823
+ fleet every one of those lines is a false alarm. One exception: `manifest-needs-snapshot` is
824
+ suppressed for `user_visible_artifact` on `lane: remote` — the only manifest-backed key that lane also
825
+ rejects outright (`lane-remote-incompatible-key`, an ERROR), so the INFO would be redundant advice
826
+ about a key the scenario can never even load with. `gate-needs-controlout` has no such exception. *Fix:*
827
+ `lint --min-severity WARN` in CI (≥1.11.0) — the INFO advisories stay one flag away for interactive use.
828
+ `--strict --min-severity ERROR` behaves as a plain lint, not a contradiction.
775
829
  23. **`verify-cassettes`/`replay` report a `discovery-surface` note on cassettes you just recorded fine.**
776
830
  *Why:* the cassette froze its `system/init` tool inventory from before the skills/plugins discovery
777
831
  servers existed at that tier (added 1.10.0). It is a non-gating **note**, never a finding — it cannot
@@ -779,6 +833,18 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
779
833
  `mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
780
834
  re-recording would never produce those tools anyway.
781
835
 
836
+ 24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
837
+ lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
838
+ emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
839
+ Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
840
+ optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
841
+ other — and probing a remote session makes this harness look like it emulates the wrong tool under the
842
+ wrong schema. It doesn't; the lanes genuinely disagree. *Fix:* describe the **outcome** ("deliver the
843
+ file to the user") and let the model pick its surface's tool. The `no_scratchpad_leak` /
844
+ `present_files_called` assertion keys are harness-side names and stay valid either way.
845
+ ([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
846
+ → "File delivery" has the binary-verified detail; repo-only.)
847
+
782
848
  For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
783
849
  files in `references/` (the gotchas above are the full list; the references repeat only the
784
850
  assertion/replay-relevant ones).
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`).
3
+ Self-contained reference. Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.13.2"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.15.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -58,7 +58,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
58
58
  GitHub-hosted runners, no token/Docker/agent:
59
59
 
60
60
  ```yaml
61
- - run: npm i -g "cowork-harness@>=1.13.2"
61
+ - run: npm i -g "cowork-harness@>=1.15.0"
62
62
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
63
63
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
64
64
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -107,8 +107,10 @@ The split is not just about tokens — it decides **where each lane can run**:
107
107
  one (see [docs/cassette.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md)). This is your **always-on PR gate**. With `--output-format json`, read each
108
108
  `results[].verdict.{pass,signals}` for **per-cassette** pass/fail and the reason (the top-level `ok`
109
109
  collapses the whole batch) — e.g. a `stalled` signal fails a cassette whose assertions all passed.
110
- The PR gate evaluates the assertions **frozen in the cassette**; to re-check against an edited on-disk
111
- `assert:` without a paid re-record, use `replay --assert-from <scenario.yaml>` (or `--reassert`).
110
+ The PR gate evaluates the **whole scenario frozen in the cassette** (`lane:`/`fidelity:`/`baseline:` too,
111
+ not only `assert:`) — a YAML edit cannot move this gate's verdict. To re-check against an edited on-disk
112
+ `assert:` without a paid re-record, use `replay --assert-from <scenario.yaml>` (or `--reassert`); any other
113
+ edited key needs a re-record.
112
114
  - **`run` / `record` (live).** Spawns the real agent in a sandbox: real model tokens + Docker **+ the
113
115
  staged Claude Code agent ELF**, bind-mounted from a local Claude Desktop install or pointed to via
114
116
  `COWORK_AGENT_BINARY`. Nothing is bundled, and **the agent binary is not redistributable** — a clean
@@ -219,10 +221,10 @@ jobs:
219
221
  steps:
220
222
  - uses: actions/checkout@v4
221
223
  - uses: actions/setup-node@v4
222
- with: { node-version: '20' }
224
+ with: { node-version: '24' }
223
225
  - uses: actions/setup-python@v5
224
226
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
225
- - run: npm i -g "cowork-harness@>=1.13.2"
227
+ - run: npm i -g "cowork-harness@>=1.15.0"
226
228
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
227
229
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
228
230
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -251,7 +253,7 @@ jobs:
251
253
  echo "live=true" >> "$GITHUB_OUTPUT"
252
254
  fi
253
255
  - if: steps.guard.outputs.live == 'true'
254
- run: npm i -g "cowork-harness@>=1.13.2"
256
+ run: npm i -g "cowork-harness@>=1.15.0"
255
257
  - if: steps.guard.outputs.live == 'true'
256
258
  run: cowork-harness run scenarios/ --output-format json
257
259
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -30,7 +30,8 @@ turn rows carry `critiqueRole:"task"` / `"reflection"`.
30
30
 
31
31
  Two things a harvester needs: roll-ups are excluded from `stats` aggregation (they carry no verdict —
32
32
  counting them adds a phantom run and drags `passRate` toward 1), so **filter them out of any pass-rate
33
- computed over raw rows**; and a roll-up with `result:"error"` had an unpriced workload, so its totals
33
+ computed over raw rows** — that exclusion also governs `stats --group-by skill-hash`, so a per-generation
34
+ **total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
34
35
  UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
35
36
 
36
37
  ## The report's item shape — no `title`, no `summary`
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`).
3
+ Self-contained reference. Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -211,14 +211,18 @@ up often enough to spell out:
211
211
  (`staleness[]` flags skill/baseline drift as a hint; only a live `run` re-confirms current
212
212
  behavior). See [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § "Still skipped on replay" and [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) § "Which
213
213
  assertions survive replay."
214
- - **Container-only assertions can't verify off the `container` tier.** `no_scratchpad_leak` and
215
- `present_files_called` check the `present_files` delivery path, which is served **only** on
216
- `container` — not `hostloop`/`microvm`. Asserting them off-container hard-fails at runtime (a red
217
- run, not a false green), so you won't be fooled if you write the assertion. The quieter trap is a
218
- scenario that runs at `hostloop`/`microvm`/`protocol` and simply omits these assertions: a green
219
- run there proves **nothing** about scratchpad-leak safety or present_files delivery, because that
220
- tier never exercises the delivery path the assertions would check. Use `fidelity: container` for
221
- present_files/scratchpad-delivery coverage.
214
+ - **`present_files` assertions can't verify off the tiers that serve the tool.** `no_scratchpad_leak` and
215
+ `present_files_called` check the `present_files` delivery path — the desktop-local lane's tool; remote
216
+ Cowork delivers via the agent-native `SendUserFile` instead, so never hardcode a delivery tool name in
217
+ a SKILL.md (SKILL.md Gotcha 24). **The harness** serves `present_files` on `container` **and `hostloop`**
218
+ — not `microvm`/`protocol`. `present_files_called` works at both; `no_scratchpad_leak` stays
219
+ `container`-only, because hostloop's handler passes a validated path through without promoting, so
220
+ there is no scratch→outputs copy to leak. Asserting either key off the tiers that serve it hard-fails at
221
+ runtime (a red run, not a false green), so you won't be fooled if you write the assertion. The quieter
222
+ trap is a scenario that runs at `microvm`/`protocol` and simply omits these assertions: a green run
223
+ there proves **nothing** about scratchpad-leak safety or present_files delivery, because that tier
224
+ never exercises the delivery path the assertions would check. Use `fidelity: container` or `hostloop`
225
+ for present_files-delivery coverage; `no_scratchpad_leak` still needs `container` specifically.
222
226
  - **The harness doesn't observe rendered-artifact interactions, browser downloads, or human
223
227
  clicks.** It runs the agent headless — no webview, no browser, no person clicking "Submit." A
224
228
  class of Cowork bug (a client-side write-back to a relative URL that resolves-but-fails against
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.13.2`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.15.0`
4
4
  (baseline `desktop-1.24012.9`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -37,6 +37,20 @@ fidelity: container # protocol | container | microvm | hostloop
37
37
  execution: local # OPTIONAL — orthogonal to fidelity (a privilege/sandbox tier, all
38
38
  # local): local (default) | cloud-describe (RESERVED — no runner
39
39
  # exists yet; authoring it is a load-time error, not a silent no-op)
40
+ lane: local # OPTIONAL — which Cowork lane's DELIVERY CONTRACT the run is held to:
41
+ # local (default) | remote. On `remote`, location delivers NOTHING (a
42
+ # remote container has no auto-delivering outputs dir and is reclaimed
43
+ # at session end) and `present_files` is NOT served (a local MCP server
44
+ # can't reach a remote session) — so `user_visible_artifact` and
45
+ # present_files_called/no_scratchpad_leak are REJECTED AT SCENARIO-LOAD
46
+ # TIME (before the run starts, before any spend), not left to fail
47
+ # unverifiable/can't-verify at assertion time. Orthogonal to fidelity
48
+ # and execution — a `lane: remote` scenario still runs locally.
49
+ # Delivery semantics only; the remote device bridge is deliberately
50
+ # unmodeled.
51
+ # NEEDS >= 1.14.0: on an older CLI a scenario carrying `lane:` does NOT
52
+ # load (`Unrecognized key: "lane"`, exit 2) — it is NOT reinterpreted as
53
+ # `lane: local`. Adopting the key means raising your floor.
40
54
  on_unanswered: fail # policy for unscripted gates: fail | prompt | first | llm — run rejects prompt
41
55
  # ("agent" is retired — no longer a valid value)
42
56
 
@@ -298,8 +312,8 @@ same set live from the schema.
298
312
  | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
299
313
  | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
300
314
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
301
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **`fidelity: container` only** — `present_files` is not served on hostloop/microvm, so a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). **Only `true` is valid** |
302
- | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` only** — `present_files` is not served on hostloop/microvm. **Only `true` is valid** |
315
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
316
+ | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
303
317
  | `question_asked: <regex>` | the agent asked an AskUserQuestion whose text matches |
304
318
  | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
305
319
  | `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min` to also require a gate |
@@ -314,6 +328,7 @@ same set live from the schema.
314
328
  | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
315
329
  | `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
316
330
  | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider) |
331
+ | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
317
332
  | `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
318
333
  | `egress_denied: <host>` | the host was blocked by the egress proxy |
319
334
  | `egress_allowed: <host>` | the host was allowed through |
@@ -358,6 +373,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
358
373
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
359
374
  | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
360
375
  | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
376
+ | `undelivered_deliverables` | warn | The skill produced file(s) OUTSIDE every user-visible root and never delivered them — on a remote Cowork session the workspace is reclaimed at session end; on a local one they stay invisible to the user. Fires on every run without opting in, because the scenarios that most need it are the ones whose author never considered delivery. Silent when the evidence cannot answer the question (no workspace walk, a tier that runs no scratchpad walk, absent delivery telemetry, or a resumed turn) — never a vacuous clean. Opt out: `allow_undelivered_deliverables` |
361
377
  | `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
362
378
 
363
379
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
@@ -369,9 +385,12 @@ overall run verdict and exit code — `assert result: success` alone won't catch
369
385
  A cassette (`record`/`replay`) has **no filesystem and no network**. `replay` re-evaluates only the
370
386
  **content** assertions. The authoritative list is `ALWAYS_CONTENT_KEYS`/`QUESTION_GATE_KEYS`/`MANIFEST_KEYS` (composed) in `src/run/cassette.ts`.
371
387
 
372
- **Assertion source — frozen by default, on-disk by opt-in.** A plain `replay` evaluates the `assert:` block
373
- **frozen in the cassette** (byte-deterministic, ignores the working tree); editing `scenarios/<name>.yaml` does
374
- not change it — replay only prints a `::notice::` when a sibling's `assert:` differs. `replay --assert-from
388
+ **Scenario source — the WHOLE scenario is frozen; only `assert:` can be opted back to disk.** A cassette
389
+ captures every key (`name`/`prompt`/`session`/`baseline`/`fidelity`/`lane`/`skills`/`answers`/`execution`/
390
+ `requires_capabilities`/`expect_denied`/`assert`), and a plain `replay` evaluates **all** of them from that
391
+ frozen copy (byte-deterministic, ignores the working tree); editing `scenarios/<name>.yaml` does not change
392
+ it — replay only prints a `::notice::` when a sibling's `assert:`/`prompt:` differs, or when it fails to
393
+ load at all. An edited `lane:`/`fidelity:`/`baseline:` reaches a replay ONLY by re-recording. `replay --assert-from
375
394
  <scenario.yaml>` / `--reassert` is the opt-in token-free re-check against the on-disk block; it **hard-fails**
376
395
  on recording-shaping drift (`prompt`/`answers`/`baseline`/`skills`) or skill-content staleness (it implies
377
396
  `--fail-on-skill-drift`). `expect_denied`/filesystem/egress keys are sourced from on-disk but stay live-only —
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -118,7 +118,8 @@ have:
118
118
  4. **Replay-class note:** `max_turns` / `tool_calls_max` / `dispatch_count_max` re-evaluate
119
119
  token-free on replay; `questions_count_max` needs `controlOut` (Recipe 1's tree applies).
120
120
  `max_cost_usd` / `max_tokens` on replay assert the FROZEN recording's spend — near-zero signal
121
- as a regression gate; if you need a cost gate, put it on the live lane via `stats`.
121
+ as a regression gate; if you need a cost gate, put it on the live lane via `stats` (history) or
122
+ `--max-budget-usd` (a pre-flight refusal before the run spends).
122
123
 
123
124
  ## Recipe 5 — Evaluate your skill's ANSWER QUALITY (semantic regression gate)
124
125
 
@@ -174,7 +175,14 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
174
175
  output. **Reproduce before acting on a finding:** `cowork-harness skill <folder> "<prompt>" --repeat 5 --label gen-1`
175
176
  runs the same skill+prompt N times (2-100) and prints a variance rollup instead of a single pass/fail —
176
177
  `--repeat` works on the `skill` lane, not just `run`. A single green run proves it passed *once*.
177
- Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`. It rejects `--session-id`/
178
+ Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`.
179
+ (`--max-budget-usd` also works WITHOUT `--repeat`, where it pre-flights a single run against that
180
+ scenario's cost history and refuses before spending; the others are batch-only.
181
+ **`record` accepts it too** — there it is CUMULATIVE across the batch (a `dir/` or `--rerecord-stale`
182
+ sweep is the least predictable spend in the CLI), refusing up front if the summed history exceeds the
183
+ cap. At `--concurrency 1` a running total also stops the batch once the cap is reached; above that the
184
+ stop is disabled and the tool says so, because with N runs in flight the total is only known after an
185
+ overshoot is already paid for. Unpriced scenarios contribute $0 and are named as a LOWER BOUND.) It rejects `--session-id`/
178
186
  `--resume` (both pin one run dir) and `--decider-cmd`/`--decider-dir` (a driving agent x N is not a
179
187
  measurement).
180
188
  The full loop (harvest -> reproduce -> fix -> prove freshness -> compare) is written out end-to-end in
@@ -6,6 +6,7 @@
6
6
  "allow_missing_capability",
7
7
  "allow_permissive_auto_allow",
8
8
  "allow_stall",
9
+ "allow_undelivered_deliverables",
9
10
  "artifact_json",
10
11
  "compaction_occurred",
11
12
  "computer_links_resolve",
@@ -80,6 +81,7 @@
80
81
  "execution",
81
82
  "expect_denied",
82
83
  "fidelity",
84
+ "lane",
83
85
  "name",
84
86
  "on_unanswered",
85
87
  "prompt",
@@ -92,6 +94,7 @@
92
94
  "allow_l0_plugin_divergence",
93
95
  "allow_missing_capability",
94
96
  "allow_permissive_auto_allow",
95
- "allow_stall"
97
+ "allow_stall",
98
+ "allow_undelivered_deliverables"
96
99
  ]
97
100
  }