cowork-harness 3.9.0 → 3.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +62 -23
- package/.claude/skills/cowork-harness/references/ci-recipe.md +16 -14
- package/.claude/skills/cowork-harness/references/critique.md +10 -5
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +5 -3
- package/.claude/skills/cowork-harness/references/scenario-schema.md +17 -15
- package/.claude/skills/cowork-harness/references/task-recipes.md +3 -2
- package/.claude/skills/cowork-harness/scripts/scenario.py +59 -0
- package/.cowork-redact.json +18 -7
- package/AGENTS.md +1 -1
- package/CHANGELOG.md +370 -0
- package/CONTRIBUTING.md +7 -5
- package/DESIGN.md +1 -1
- package/README.md +16 -5
- package/SPEC.md +26 -2
- package/dist/agent/session.js +7 -0
- package/dist/assert.js +24 -8
- package/dist/baseline.js +6 -0
- package/dist/cli.js +11 -4
- package/dist/critique/command.js +167 -127
- package/dist/critique/limitations.js +1 -1
- package/dist/critique/mount-check.js +52 -0
- package/dist/critique/package-evidence.js +33 -17
- package/dist/decide/decider.js +4 -8
- package/dist/decide/external-channel.js +17 -17
- package/dist/egress/sidecar.js +18 -24
- package/dist/errors.js +38 -5
- package/dist/hostloop/pretooluse-path-hook.js +64 -0
- package/dist/hostloop/process-cwd.js +124 -0
- package/dist/redact.js +10 -2
- package/dist/run/cassette.js +7 -0
- package/dist/run/chat-result.js +1 -0
- package/dist/run/chat.js +3 -1
- package/dist/run/doctor.js +2 -2
- package/dist/run/execute.js +621 -84
- package/dist/run/lint-load.js +233 -0
- package/dist/run/outputs-delete-tier.js +59 -0
- package/dist/run/pre-run-manifest.js +57 -4
- package/dist/run/run.js +53 -6
- package/dist/run/scenario-tool.js +74 -10
- package/dist/run/verdict.js +47 -6
- package/dist/runtime/agent-image.js +20 -1
- package/dist/runtime/argv.js +1 -0
- package/dist/runtime/hostloop-prompt.js +1 -1
- package/dist/runtime/hostloop.js +61 -16
- package/dist/runtime/microvm.js +111 -4
- package/dist/scan.js +69 -8
- package/dist/secrets.js +1 -1
- package/dist/termination.js +143 -0
- package/dist/types.js +2 -2
- package/docs/README.md +4 -4
- package/docs/boundary.md +7 -6
- package/docs/cassette.md +73 -11
- package/docs/ci.md +2 -2
- package/docs/cli.md +105 -25
- package/docs/companion-skill.md +2 -2
- package/docs/critique.md +37 -18
- package/docs/decider-dir.md +3 -1
- package/docs/fidelity-gaps.md +43 -55
- package/docs/gotchas.md +1 -1
- package/docs/invariants.md +5 -5
- package/docs/maintenance.md +3 -0
- package/docs/run-status.md +12 -4
- package/docs/scenario.md +83 -47
- package/docs/session.md +6 -5
- package/docs/subagents.md +14 -11
- package/examples/README.md +2 -1
- package/examples/replays/README.md +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +59 -60
- package/llms.txt +3 -3
- package/package.json +1 -1
- package/python/test_scenario_lint.py +98 -0
- package/schema/run-result.json +35 -3
- package/schema/scenario.schema.json +3 -3
- package/scripts/check-versions.ts +46 -4
- package/scripts/gen-schema.ts +6 -5
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.10.0
|
|
7
|
+
tracks-harness: cowork-harness 3.10.0 (baseline desktop-2.9939.2)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.10.0` (baseline
|
|
29
29
|
> `desktop-2.9939.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.10.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.10.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.10.0"`. **Pin `@^3.10.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -372,7 +372,8 @@ python3 "$S" scaffold --name report-check --skill ./skills/report-gen \
|
|
|
372
372
|
```
|
|
373
373
|
|
|
374
374
|
Then lint every scenario — it encodes the no-silent-false-green invariants. Use the CLI wrapper
|
|
375
|
-
`cowork-harness lint
|
|
375
|
+
`cowork-harness lint`: it runs the bundled `scenario.py lint` **and** the harness's own scenario loader,
|
|
376
|
+
so a file `run`/`record` would refuse fails lint too (running `scenario.py lint` directly skips the loader):
|
|
376
377
|
|
|
377
378
|
```bash
|
|
378
379
|
cowork-harness lint scenarios/*.yaml
|
|
@@ -391,11 +392,17 @@ and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is
|
|
|
391
392
|
(CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
|
|
392
393
|
emits a scenario `lint` would reject.
|
|
393
394
|
|
|
394
|
-
**`lint`
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
395
|
+
**`cowork-harness lint` runs the loader: a file it calls clean is one `run`/`record` will load.** Anything
|
|
396
|
+
the loader refuses — an unknown key, a wrong value type (a scalar `semantic_matches.rubric`), a bad regex,
|
|
397
|
+
a reserved value — is ✗ ERROR `scenario-invalid` (exit 1, with or without `--strict`), and a `baseline:`
|
|
398
|
+
naming no baseline this installed CLI ships is ✗ ERROR `baseline-unknown` (`latest` always resolves). It
|
|
399
|
+
does not check what depends on the machine the run happens on (the session file and its mounts, an
|
|
400
|
+
absolute `baseline:` path, environment variables). A session or matrix YAML in a linted directory is not
|
|
401
|
+
a scenario and is reported as one that does not load — keep those out of the linted set. `python3
|
|
402
|
+
scenario.py lint` run directly stays offline and lenient: there an unknown key is only a ⚠ WARN (exit 0).
|
|
403
|
+
`cowork-harness record <file.yaml> --dry-run` also runs the loader and adds the pre-spend refusals (exit 2
|
|
404
|
+
on a schema error; a directory reports each `✗ broken:` file and exits 1). **Read the exit code, not just
|
|
405
|
+
its sign:** `record <file>` — with or without
|
|
399
406
|
`--dry-run` — answers `2` for "did not load" and `1` for "loaded fine, but this record is refused" (a
|
|
400
407
|
pre-spend policy refusal; `--max-budget-usd` is the one refusal that keeps exit 2). Treating any non-zero
|
|
401
408
|
as "scenario broken" mis-reports every refused-but-valid scenario. Corollary: **the loader** fails LOUD on an unknown key (never silently) —
|
|
@@ -421,8 +428,9 @@ When the scenario declares `answers:`, verify-run **also** checks they still mat
|
|
|
421
428
|
reworded gate or a `choose:` the run never offered fails here in ~1s instead of on a paid re-record). Or skip
|
|
422
429
|
the discovery/encode/record dance entirely and answer gates **live during the recording** with
|
|
423
430
|
`record --decider-dir`/`--decider-llm` (the cassette is flagged non-deterministic but replays deterministically).
|
|
424
|
-
`run` takes no `--dry-run`: to check that a scenario **loads** without spending,
|
|
425
|
-
|
|
431
|
+
`run` takes no `--dry-run`: to check that a scenario **loads** without spending, `cowork-harness lint
|
|
432
|
+
<file.yaml>` runs the real loader (and resolves a named `baseline:`); `cowork-harness record <file.yaml>
|
|
433
|
+
--dry-run` runs the real loader too AND the same scenario-level
|
|
426
434
|
refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
|
|
427
435
|
cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
|
|
428
436
|
guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
|
|
@@ -432,12 +440,14 @@ the verdict kind, and portability can only ever warn — that do NOT affect the
|
|
|
432
440
|
takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
|
|
433
441
|
contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
|
|
434
442
|
real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
|
|
435
|
-
reports every offender and the batch cost estimate. `lint` checks the assertion invariants (
|
|
443
|
+
reports every offender and the batch cost estimate. `lint` checks the assertion invariants AND that each file loads (the same loader, plus a named `baseline:`), but not the pre-spend refusals.
|
|
436
444
|
|
|
437
445
|
**Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
|
|
438
446
|
concluded the free pre-flight was unavailable:
|
|
439
|
-
- **"Does my whole corpus still load?"** →
|
|
440
|
-
|
|
447
|
+
- **"Does my whole corpus still load?"** → `cowork-harness lint scenarios/` answers it (every file the
|
|
448
|
+
loader rejects is an ERROR, and so is a `baseline:` naming no shipped baseline), or the **directory** arm
|
|
449
|
+
(`record scenarios/ --dry-run --quiet`, the CI shape in `references/ci-recipe.md`) when you also want the
|
|
450
|
+
pre-spend refusals. The directory arm reports every offender in one pass, and the
|
|
441
451
|
destination-policy verdict cannot red it: that arm knows no `--out`, so host-inventory and portability
|
|
442
452
|
are advisory `notes[]` at exit 0 while a file that cannot load is `✗ broken:` at exit 1. Limits worth
|
|
443
453
|
knowing: it is **non-recursive** (`readdirSync` — scenarios in subdirectories are never opened), a file
|
|
@@ -593,11 +603,15 @@ Recognize these before "fixing" a non-bug:
|
|
|
593
603
|
**Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
|
|
594
604
|
scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
|
|
595
605
|
**The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
|
|
596
|
-
can see them,
|
|
597
|
-
|
|
598
|
-
|
|
599
|
-
|
|
600
|
-
|
|
606
|
+
can see them, and give the file tools an **absolute** path under the outputs directory the agent's prompt
|
|
607
|
+
names. On the desktop-local host-loop lane (what production runs), against Desktop **2.7032.0 and later**,
|
|
608
|
+
the agent process runs outside the session (`/var/empty`), so a relative `Read`/`Write`/`Edit` — a bare
|
|
609
|
+
filename or `outputs/x.md` alike — is **refused** ("File is in a directory that is denied by your
|
|
610
|
+
permission settings."); only a pathless or relative `Grep`/`Glob` is redirected to outputs. (Before
|
|
611
|
+
2.7032.0 the file tools were rooted at `outputs/`, so a bare filename landed there and `outputs/x.md`
|
|
612
|
+
doubled to `outputs/outputs/x.md`.) At `fidelity: container`/`microvm` (VM-loop, the harness default) the
|
|
613
|
+
base is the session root, so a bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an
|
|
614
|
+
explicit delivery. Addressing
|
|
601
615
|
a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
|
|
602
616
|
decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
|
|
603
617
|
[docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
|
|
@@ -638,11 +652,36 @@ Recognize these before "fixing" a non-bug:
|
|
|
638
652
|
`Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
|
|
639
653
|
failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
|
|
640
654
|
suspiciously empty.
|
|
655
|
+
- **`outputs_delete_unconfirmed`** (`WARN`) — a delete-shaped command near `mnt/outputs` that nothing
|
|
656
|
+
confirms: the per-turn filesystem diff shows no output present at turn start was deleted, and no flagged
|
|
657
|
+
delete has an `outputs/` path as its own operand. The classic case is a Python variable named `rm` in a
|
|
658
|
+
`python3 -c` body that also reads a report from outputs: `rm = json.load(open(".../outputs/r.json"))`.
|
|
659
|
+
A file the turn created and then deleted is invisible to the diff, so real deletes land here too:
|
|
660
|
+
- a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`);
|
|
661
|
+
- a `cd` then a relative path;
|
|
662
|
+
- chained variables (`A=…; B=$A/x; rm "$B"`);
|
|
663
|
+
- a Python path held in a variable set on another line (`p = …` then `os.remove(p)`);
|
|
664
|
+
- wrapper flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`);
|
|
665
|
+
- calls outside the modelled set, such as Node's `fs.promises.rm(…)`.
|
|
666
|
+
|
|
667
|
+
Read the command before dismissing it. A literal-path delete (`rm -f mnt/outputs/x`,
|
|
668
|
+
`os.remove(".../outputs/x")`) still fails `outputs_delete`. So do two non-deletes: quoted text where a
|
|
669
|
+
delete command with an outputs operand follows a separator, subshell or keyword
|
|
670
|
+
(`echo 'note; rm mnt/outputs/x'` — the classifier does not track quotes), and a heredoc that *writes* a
|
|
671
|
+
script instead of running it. A statement over 4 KiB or a command over 16 KiB is judged by the stricter
|
|
672
|
+
original rule, so a huge one-line body with a variable named `rm` fails again, and a command whose variable
|
|
673
|
+
expansion would exceed the scanner's work budget (about a hundred distinct variables in one 10 KB line) is
|
|
674
|
+
not expanded — every mount it names literally counts as deleted in. Waive any of these with
|
|
675
|
+
`allow_outputs_delete`. The warn is raised even when `no_delete_in_outputs` is authored (the assertion
|
|
676
|
+
passes; this warn is how the hit stays visible in text output).
|
|
677
|
+
- **`outputs_diff_unavailable`** (`WARN`) — the outputs filesystem diff could not verify this turn and the
|
|
678
|
+
text scan saw nothing, so a delete by a script file or a non-bash tool would have gone unseen.
|
|
641
679
|
- **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
|
|
642
|
-
`RunResult.scan` is undefined and the host-path
|
|
680
|
+
`RunResult.scan` is undefined and the host-path guard and the outputs-delete **text scan did not run this
|
|
681
|
+
run** (the outputs filesystem diff still did, and a delete it proves still fails). Not a
|
|
643
682
|
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
644
683
|
|
|
645
|
-
The full
|
|
684
|
+
The full 22-code signal table (severity + per-signal opt-out) is in
|
|
646
685
|
[`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
|
|
647
686
|
the fuller narrative.
|
|
648
687
|
|
|
@@ -1066,7 +1105,7 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
1066
1105
|
20. **A `mode: r` connected folder's contents are recorded body-less, not excluded.** `record` captures a
|
|
1067
1106
|
read-only folder's files as path + hash only (`truncated: true`, no `body`) — it's an input the agent
|
|
1068
1107
|
read, not a deliverable it wrote. `file_exists`/`computer_links_resolve` still pass against it on replay
|
|
1069
|
-
(the hash-only entry still materializes a placeholder); `artifact_json`
|
|
1108
|
+
(the hash-only entry still materializes a placeholder); `artifact_json`/`artifact_text` report a clear
|
|
1070
1109
|
evidence-unavailable on every lane (live/verify-run/replay agree — no green-record/red-replay). This is
|
|
1071
1110
|
also why a `mode: r` input never trips the `binary` privacy finding or needs `--allow` — only a
|
|
1072
1111
|
*committed* body is scanned. `scaffold` won't emit `file_exists` for one either (it's not in
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.10.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -77,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
77
77
|
GitHub-hosted runners, no token/Docker/agent:
|
|
78
78
|
|
|
79
79
|
```yaml
|
|
80
|
-
- run: npm i -g "cowork-harness@^3.
|
|
80
|
+
- run: npm i -g "cowork-harness@^3.10.0"
|
|
81
81
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
82
82
|
# no silent false-greens. WITHOUT --strict this
|
|
83
83
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -208,7 +208,7 @@ no live filesystem and no network — and it verifies an artifact's *content* on
|
|
|
208
208
|
carries an `artifacts` manifest (recorded `outputs/` + connected folders; then `file_exists` /
|
|
209
209
|
`user_visible_artifact` / `artifact_json` evaluate on replay). On a manifest-less cassette the
|
|
210
210
|
deliverable is invisible to the gate. Don't let a green replay gate convince you the deliverable is correct.
|
|
211
|
-
Run `cowork-harness lint` (the bundled `scenario.py lint`) in CI to catch a scenario that put a
|
|
211
|
+
Run `cowork-harness lint` (the bundled `scenario.py lint` plus the harness's own loader) in CI to catch a scenario that put a
|
|
212
212
|
filesystem/egress-only check on the replay lane (a silent no-op). Author new scenarios with
|
|
213
213
|
`scenario.py scaffold` so they start from a valid, self-linted skeleton.
|
|
214
214
|
|
|
@@ -238,8 +238,9 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
238
238
|
`.cowork-redact.json` next to your scenarios (or set `COWORK_HARNESS_REDACT_PATTERNS` /
|
|
239
239
|
`_KEYS`); empty by default. The policy is searched in **cwd → the scenario's dir → the cassette's
|
|
240
240
|
dir** (first file found per dir; env vars merge on top). `cowork-harness init-redact` copies the
|
|
241
|
-
packaged reference template (local-path prefixes
|
|
242
|
-
|
|
241
|
+
packaged reference template (local-path prefixes, incl. macOS temp roots and slugged home segments
|
|
242
|
+
like `-Users-<user>-…`, + a generic email regex) into the cwd as a starting point — review and tailor
|
|
243
|
+
it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` or add them. Redaction is **verdict-preserving** — `record` refuses to write if
|
|
243
244
|
redaction would flip an assertion (a manufactured green). `--no-redact` skips it for known-synthetic
|
|
244
245
|
inputs.
|
|
245
246
|
- **Pre-spawn preflight**: `record` warns (`::warning::`, before the paid run starts — once per batch
|
|
@@ -307,14 +308,15 @@ cowork-harness verify-cassettes cassettes/ --skip-privacy # staleness onl
|
|
|
307
308
|
A typical skill repo runs four stages, fastest/cheapest first:
|
|
308
309
|
|
|
309
310
|
1. **Unit** — your skill's own tests (pytest/vitest of its scripts). Not the harness's job.
|
|
310
|
-
2. **Boundary / lint** — `cowork-harness lint scenarios/*.yaml` (no-silent-false-green invariants
|
|
311
|
-
python3 — PyYAML is bundled) + `cowork-harness verify-cassettes cassettes/` (privacy scan + staleness) +
|
|
311
|
+
2. **Boundary / lint** — `cowork-harness lint scenarios/*.yaml` (no-silent-false-green invariants, and
|
|
312
|
+
every scenario the loader would refuse is an ERROR; needs python3 — PyYAML is bundled) + `cowork-harness verify-cassettes cassettes/` (privacy scan + staleness) +
|
|
312
313
|
`cowork-harness boundary-check` where relevant. Token-free, agent-free. **Don't `|| true` the lint
|
|
313
314
|
step** — a missing python3 (exit 127) or a lint error makes `scenario.py` exit non-zero, and swallowing
|
|
314
315
|
that turns the false-green guard itself into a silent no-op.
|
|
315
316
|
|
|
316
|
-
**
|
|
317
|
-
|
|
317
|
+
**Optionally add a load gate next to `lint`** — `lint` already runs the loader (an unknown key, a wrong
|
|
318
|
+
value type or an unknown `baseline:` name is an ERROR), and `record --dry-run` adds the pre-spend
|
|
319
|
+
refusals a real record would apply:
|
|
318
320
|
|
|
319
321
|
```bash
|
|
320
322
|
cowork-harness record scenarios/ --dry-run --quiet # does every scenario LOAD? no tokens, writes nothing
|
|
@@ -324,8 +326,8 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
324
326
|
failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
|
|
325
327
|
offending file *and* the rejected key, one line per file, and the step still exits 1. (Silent is
|
|
326
328
|
literal only when your scenarios name a `fidelity:`; one still in the deprecation window prints one
|
|
327
|
-
defaulted-fidelity notice per scenario.)
|
|
328
|
-
|
|
329
|
+
defaulted-fidelity notice per scenario.) Point `lint` at scenarios only: a session or matrix YAML in
|
|
330
|
+
the linted set is reported as a file that does not load.
|
|
329
331
|
|
|
330
332
|
**If the repo pays for `critique`, gate the evidence corpus here first, for free:**
|
|
331
333
|
|
|
@@ -364,7 +366,7 @@ jobs:
|
|
|
364
366
|
with: { node-version: '24' }
|
|
365
367
|
- uses: actions/setup-python@v5
|
|
366
368
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
367
|
-
- run: npm i -g "cowork-harness@^3.
|
|
369
|
+
- run: npm i -g "cowork-harness@^3.10.0"
|
|
368
370
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
369
371
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
370
372
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -393,7 +395,7 @@ jobs:
|
|
|
393
395
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
394
396
|
fi
|
|
395
397
|
- if: steps.guard.outputs.live == 'true'
|
|
396
|
-
run: npm i -g "cowork-harness@^3.
|
|
398
|
+
run: npm i -g "cowork-harness@^3.10.0"
|
|
397
399
|
- if: steps.guard.outputs.live == 'true'
|
|
398
400
|
run: cowork-harness run scenarios/ --output-format json
|
|
399
401
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -144,10 +144,15 @@ non-git folder is measured raw, as staging copies it). Unlike `lint-skill`'s sta
|
|
|
144
144
|
packager: untracked files and symlinks outside the plugin are excluded by the same filter and containment
|
|
145
145
|
rule a critique applies, bytes are the same UTF-8-decoded measurement, and the one clause no static
|
|
146
146
|
instrument can see (a plugin-root reference read at run time) is stated as the floor rather than guessed
|
|
147
|
-
at.
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
147
|
+
at. A skill that is a git submodule of its plugin (or any `--skill` subdirectory with nothing tracked
|
|
148
|
+
under it) is refused in staging's terms by `--corpus-only` and by a paid critique alike, before any spend;
|
|
149
|
+
the corpus only ever holds files the mount carries.
|
|
150
|
+
|
|
151
|
+
`critique <plugin>/skills/<name>` is the same run as `critique <plugin> --skill <name>`: critique mounts
|
|
152
|
+
the plugin, as Cowork does, and grades `<name>` (same corpus, `skillHash`, `gradedSkill`), with a notice.
|
|
153
|
+
It mounts the skill folder alone — and says why — only when `--skill` cannot reach it (not at exactly
|
|
154
|
+
`skills/<name>`, a submodule, or a case mismatch); that run and its corpus lack the plugin's agents and
|
|
155
|
+
shared references.
|
|
151
156
|
|
|
152
157
|
The report's `evidenceBudget` object says exactly what was shown — read it instead of inferring budgets
|
|
153
158
|
from `dist/` source:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -309,8 +309,10 @@ up often enough to spell out:
|
|
|
309
309
|
|
|
310
310
|
- `COWORK_HARNESS_RUNS_DIR` (or `--run-dir <path>`) — override the default run-output root `~/.cowork-harness/runs` (out of any working tree). flag > env > default.
|
|
311
311
|
- `COWORK_HARNESS_DECIDER_CMD_TIMEOUT_MS` / `COWORK_HARNESS_LLM_TIMEOUT_MS` — decider backstops
|
|
312
|
-
(default 600 s; **fail loud** on timeout).
|
|
313
|
-
|
|
312
|
+
(default 600 s; **fail loud** on timeout). A `--decider-cmd` timeout, or a helper that exits before
|
|
313
|
+
answering, ends the run as an unanswered-gate partial (`result.json` written, exit 2); a timeout is
|
|
314
|
+
recorded as `errorSource: "decider_timeout"`.
|
|
315
|
+
- `COWORK_HARNESS_DECIDER_DIR_POLL_MS` / `_TIMEOUT_MS` — the `--decider-dir` rendezvous (poll defaults: 300 ms for the run-side rendezvous, 500 ms for `gates --follow`). A backstop timeout ends the run as an unanswered-gate partial recorded as `errorSource: "decider_timeout"`.
|
|
314
316
|
- `COWORK_HARNESS_DIALOG_TIMEOUT_MS` — dialog auto-cancel (default 6 s).
|
|
315
317
|
- `COWORK_HARNESS_LLM_MAX_BYTES` — stdout bound on `--decider-llm` (default 8 MiB).
|
|
316
318
|
- `COWORK_HARNESS_LLM_RETRIES` — bounded retries for a transient non-zero `claude -p` exit on the
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.10.0`
|
|
4
4
|
(baseline `desktop-2.9939.2`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -307,7 +307,7 @@ same set live from the schema.
|
|
|
307
307
|
| `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
308
308
|
| `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
|
|
309
309
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
310
|
-
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. **Omitting it does NOT allow deletes** — a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored; use `allow_outputs_delete: true` to accept an intended one. Detection is a post-run bash-command scan
|
|
310
|
+
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. **Omitting it does NOT allow deletes** — a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored; use `allow_outputs_delete: true` to accept an intended one. Detection is a post-run bash-command scan plus a per-turn filesystem diff of `outputs/`, not mount enforcement, so a green means none was *detected*. Fails when the diff proves the delete, when a delete in command/call position has an `outputs/` path as its own operand, or when the diff could not verify; a hit resting only on the detector's inference (e.g. a Python variable named `rm`) with a clean diff passes, and the `outputs_delete_unconfirmed` warn is still raised in the run output — see that code for the classes of real delete that land there |
|
|
311
311
|
| `no_delete_in_mounts: true` | no delete op touched ANY delete-denied mount — `outputs` plus every `rw` connected folder — except those waived by `allow_delete_in`. Production denies `unlink`/`rmdir` on every such mount, so `no_delete_in_outputs` covers only part of the real rule. **only `true` is valid** |
|
|
312
312
|
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
|
|
313
313
|
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
|
|
@@ -315,7 +315,7 @@ same set live from the schema.
|
|
|
315
315
|
| `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); evidence-unavailable on `lane: remote` / `preRunOrigin: remote-unavailable` |
|
|
316
316
|
| `artifact_text: {artifact, contains?, not_contains?, matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read |
|
|
317
317
|
| `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
|
|
318
|
-
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert.
|
|
318
|
+
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. `cowork-harness lint` reports it too (ERROR `scenario-invalid` — the wrapper runs the same loader); the bundled `scenario.py lint` run directly does NOT |
|
|
319
319
|
| `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
|
|
320
320
|
| **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
|
|
321
321
|
| **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
|
|
@@ -358,7 +358,7 @@ same set live from the schema.
|
|
|
358
358
|
| `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
|
|
359
359
|
| `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
|
|
360
360
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
361
|
-
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs
|
|
361
|
+
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, `/private/var/empty` from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
362
362
|
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
363
363
|
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
364
364
|
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
@@ -367,7 +367,7 @@ same set live from the schema.
|
|
|
367
367
|
| `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
|
|
368
368
|
| `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
|
|
369
369
|
| `gate_answer_count_min: <N>` | at least N AskUserQuestion gates fired AND were delivered non-error — presence companion to `gate_answers_delivered`'s vacuous-pass. **`: 0` asserts nothing** and does not satisfy that pairing; `>= 1` is **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
|
|
370
|
-
| `hook_blocked: <regex>` | a PreToolUse hook blocked a tool whose name matches the regex (`RunResult.hookEvents`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream) |
|
|
370
|
+
| `hook_blocked: <regex>` | a PreToolUse hook blocked a tool whose name matches the regex (`RunResult.hookEvents`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream). **Does not see the agent's own refusal**: on `hostloop` against Desktop 2.7032.0+ a relative `Read`/`Write`/`Edit` is denied by the agent's permission rules before any hook runs, so neither this key nor `path_denied` records it — assert it with `tool_result_contains: "denied by your permission settings"` |
|
|
371
371
|
| `no_hook_blocked: true` | no tool was hook-blocked during the run (distinguishes a real tool crash from an intentional hook block) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette. **Only `true` is valid** |
|
|
372
372
|
| `hook_event_fired: <HookEvent>` | a **command hook** for this event (a plugin's `hooks/hooks.json` or manifest hook — `Stop`, `SessionStart`, `PostToolUse`, …) ran: a `hook_response` system frame with that `hook_event` was recorded (`RunResult.contextEvents`). Any outcome counts. The harness passes `--include-hook-events` whenever a staged plugin declares hooks — that is what puts events other than SessionStart/Setup on the stream — so a recording made without it reports "never fired". Content-class, grades on replay. Recorded end-to-end for `Stop` ([stop-hook-probe.scenario.yaml](https://github.com/yaniv-golan/cowork-harness/blob/main/examples/probes/stop-hook-probe.scenario.yaml)); the other names match the same frame but have not each been recorded |
|
|
373
373
|
| `hook_event_blocked: <HookEvent>` | that command hook **blocked** at least once — a `hook_response` frame for the event carried `exit_code: 2`. Fails naming the exit codes seen when it fired without blocking (a frame with no `exit_code` is reported as such, never counted); fails "never fired" otherwise; cannot-verify when the run has no context events. Content-class |
|
|
@@ -379,9 +379,9 @@ same set live from the schema.
|
|
|
379
379
|
| `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
|
|
380
380
|
| `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
|
|
381
381
|
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
382
|
-
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
382
|
+
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Silences `outputs_delete`, `outputs_delete_unconfirmed` and `outputs_diff_unavailable`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
383
383
|
| `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
|
|
384
|
-
| `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
|
|
384
|
+
| `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
|
|
385
385
|
| `egress_denied: <host>` | the host was blocked by the egress proxy |
|
|
386
386
|
| `egress_allowed: <host>` | the host was allowed through |
|
|
387
387
|
| `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
|
|
@@ -405,7 +405,7 @@ dotted path.
|
|
|
405
405
|
|
|
406
406
|
**VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; eleven
|
|
407
407
|
are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
|
|
408
|
-
`"success"`) and
|
|
408
|
+
`"success"`) and eleven are **warn**-severity (informational, never flip pass/fail). All twenty-two signal
|
|
409
409
|
codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
410
410
|
|
|
411
411
|
| Code | Severity | Meaning |
|
|
@@ -415,7 +415,9 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
415
415
|
| `usage_limit` | fail | Usage/quota limit hit (not a skill failure) — retry after the limit resets. Emitted when `RunResult.resultErrorKind === "usage_limit"` |
|
|
416
416
|
| `transport_error` | fail | The connection dropped mid/after-run |
|
|
417
417
|
| `permissive_auto_allow` | fail | A cowork-parity auto-allow real Cowork would block (opt out: `allow_permissive_auto_allow`) |
|
|
418
|
-
| `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs`
|
|
418
|
+
| `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs`, confirmed: the per-turn filesystem diff proves it, a delete in command/call position has an `outputs/` path as its own operand, or the diff could not verify the turn. Authoring `no_delete_in_outputs` moves it into that assertion; `allow_outputs_delete` waives it |
|
|
419
|
+
| `outputs_delete_unconfirmed` | warn | A delete-shaped command near `mnt/outputs` nothing confirms: no output present at turn start was deleted and no flagged delete has an outputs path as its own operand (e.g. a Python variable named `rm` next to an outputs path, quoted prose, a `sed`/`grep` pattern). A *statement* is a fragment split on newline/`;`/`&&`/`\|\|`, quote-blind, so some real deletes also land here — a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`), a `cd` then a relative path, chained variables (`A=…; B=$A/x; rm "$B"`), a Python path held in a variable set on another line (`p = …` then `os.remove(p)`, or `for p in …:` then `p.unlink()`), wrappers with flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`), and calls outside the modelled set such as Node's `fs.promises.rm(…)` — read the command. Raised even when `no_delete_in_outputs` is authored (that assertion passes on it). Kept false positives that still fail `outputs_delete`: quoted text in which a delete command with an outputs operand follows a shell separator, subshell or keyword — the classifier does not track quotes (`echo 'note; rm mnt/outputs/x'`, `echo "a & rm …/outputs/x"`), and a heredoc that *writes* a script rather than running it (`cat <<EOF > clean.sh` with an `rm …/outputs/x` line). Waive: `allow_outputs_delete` |
|
|
420
|
+
| `outputs_diff_unavailable` | warn | The per-turn filesystem diff of `outputs/` could not verify this turn and the text scan flagged nothing — a delete made without a bash command would have gone undetected. With a text hit the turn fails `outputs_delete` instead |
|
|
419
421
|
| `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
|
|
420
422
|
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
421
423
|
| `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
|
|
@@ -425,7 +427,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
425
427
|
| `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
|
|
426
428
|
| `model_fallback` | warn | The agent fell back off the requested model mid-run (SDK `model_fallback` event); a `model_not_found`/`model_blocked` trigger repeats every run until the pin changes |
|
|
427
429
|
| `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
|
|
428
|
-
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete
|
|
430
|
+
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path guard and the outputs-delete text scan did not run this run; the outputs filesystem diff still ran |
|
|
429
431
|
| `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
|
|
430
432
|
| `undelivered_deliverables` | warn | The skill produced file(s) OUTSIDE every user-visible root and never delivered them, so they stay invisible to the user. Fires on every run without opting in, because the scenarios that most need it are the ones whose author never considered delivery. Silent when the evidence cannot answer the question (no workspace walk, a tier that runs no scratchpad walk, absent delivery telemetry, a resumed turn, or a lane where delivery is unobservable — see `delivery_unobservable`) — never a vacuous clean. **`lane: local` only**: on remote, delivery cannot be measured at all, so that lane reports `delivery_unobservable` instead of guessing. Opt out: `allow_undelivered_deliverables` |
|
|
431
433
|
| `delivery_unobservable` | warn | `lane: remote` only — the run produced file(s), but whether any reached the user CANNOT be verified: nothing is delivered by location on that lane and the harness models no remote delivery tool (production uses the agent-native `SendUserFile`). The honest counterpart to `undelivered_deliverables`, which would otherwise fire on every remote run that writes anything — a signal that always fires carries no information. Mutually exclusive with it; quiet when the run produced nothing to deliver. A harness coverage gap, not a skill defect. Opt out: `allow_undelivered_deliverables` |
|
|
@@ -433,7 +435,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
433
435
|
|
|
434
436
|
A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
|
|
435
437
|
overall run verdict and exit code — `assert result: success` alone won't catch it; check
|
|
436
|
-
`result.verdict.signals[].severity` or the run's exit code. Only the
|
|
438
|
+
`result.verdict.signals[].severity` or the run's exit code. Only the eleven **warn** codes are truly benign.
|
|
437
439
|
|
|
438
440
|
## Replay class
|
|
439
441
|
|
|
@@ -580,8 +582,8 @@ debugging a run's behavior. The two are **numbered independently**: a bare "gotc
|
|
|
580
582
|
the stochastic path flags the run `nonDeterministic`. The LLM decider is one mechanism, two
|
|
581
583
|
spellings: `on_unanswered: llm` (YAML) and `--decider-llm` (CLI). The bare `--on-unanswered llm`
|
|
582
584
|
is rejected (use `--decider-llm`). `agent` is **retired** — `on_unanswered: agent` is rejected by
|
|
583
|
-
the schema. (`src/types.ts` — the `on_unanswered` enum; `src/cli.ts
|
|
584
|
-
`--on-unanswered` value check
|
|
585
|
+
the schema. (`src/types.ts` — the `on_unanswered` enum; `src/cli.ts` — the CLI-side
|
|
586
|
+
`--on-unanswered` value check; grep `--on-unanswered llm is not a user flag`.)
|
|
585
587
|
|
|
586
588
|
4. **`--on-unanswered first` is non-deterministic too** — it picks option 1 and is flagged
|
|
587
589
|
`nonDeterministic`; not a deterministic substitute for scripted answers.
|
|
@@ -662,8 +664,8 @@ debugging a run's behavior. The two are **numbered independently**: a bare "gotc
|
|
|
662
664
|
21. **A `mode: r` connected folder's contents are recorded body-less, not excluded.** `record` captures a
|
|
663
665
|
read-only folder's files as `path` + `bytes` + `sha256` only (`truncated: true`, no `body`) — it's an
|
|
664
666
|
input the agent read, not a deliverable it wrote. `file_exists`/`computer_links_resolve` still pass
|
|
665
|
-
against it on replay (the hash-only entry still materializes a 0-byte placeholder); `artifact_json`
|
|
666
|
-
|
|
667
|
+
against it on replay (the hash-only entry still materializes a 0-byte placeholder); `artifact_json`/`artifact_text`
|
|
668
|
+
report a clear evidence-unavailable on every lane (live/verify-run/replay agree). This is also why a
|
|
667
669
|
`mode: r` input never trips the `binary` privacy finding
|
|
668
670
|
or needs `--allow` in `verify-cassettes` — only a *committed* body is scanned. `scaffold` won't emit
|
|
669
671
|
`file_exists` for one either, since it isn't in `RunResult.artifacts`. A `mode: rw`/`rwd` folder's
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -84,7 +84,8 @@ the transcript and would be committed inside the cassette. Set the policy up **b
|
|
|
84
84
|
record — retrofitting means re-recording:
|
|
85
85
|
|
|
86
86
|
1. `cowork-harness init-redact` — copies the reference `.cowork-redact.json` (host-path +
|
|
87
|
-
email patterns) into the current directory. Version-control it.
|
|
87
|
+
email patterns) into the current directory. Version-control it. Re-run with `--force` after an
|
|
88
|
+
upgrade to pick up new rules (save your tailoring first).
|
|
88
89
|
2. **Search set:** record looks for `.cowork-redact.json` in the **cwd, the scenario's directory,
|
|
89
90
|
and the cassette's output directory** — every distinct file found is MERGED (plus
|
|
90
91
|
`COWORK_HARNESS_REDACT_PATTERNS`/`_KEYS` from the env). Repo root (= cwd in CI) is the
|
|
@@ -50,6 +50,15 @@ lint-skill flags (skill bodies + any sibling hooks.json):
|
|
|
50
50
|
live-verified) but has no assertion key, so a scenario can't gate on it
|
|
51
51
|
W `${CLAUDE_PLUGIN_ROOT}` in a VM bash step / host-side hook seeding (host-loop footguns)
|
|
52
52
|
|
|
53
|
+
Through the `cowork-harness lint` CLI wrapper (not when this script is run directly), lint ALSO reports
|
|
54
|
+
every file the harness's own scenario loader rejects -- the check `run`/`record` apply before anything
|
|
55
|
+
runs:
|
|
56
|
+
E `scenario-invalid` the loader refuses the file (schema, value types, unknown keys, bad regex,
|
|
57
|
+
reserved values) -- including a non-scenario YAML in a linted directory
|
|
58
|
+
E `baseline-unknown` `baseline:` names no baseline this installed cowork-harness ships
|
|
59
|
+
Run directly, this script stays offline and never parses with the loader, so prefer the CLI wrapper
|
|
60
|
+
when it is installed.
|
|
61
|
+
|
|
53
62
|
Designed for agents and CI: non-interactive, --help, --json, meaningful exit codes,
|
|
54
63
|
idempotent. `lint` exits 1 on any ERROR (or any finding with --strict); else 0.
|
|
55
64
|
|
|
@@ -1322,6 +1331,55 @@ def _print_findings(findings, n_files, kind="scenario", clean_suffix=" — no si
|
|
|
1322
1331
|
print(f"\n{n_err} error(s), {n_warn} warning(s), {n_info} info across {n_files} file(s).")
|
|
1323
1332
|
|
|
1324
1333
|
|
|
1334
|
+
_EXTRA_FINDINGS_ENV = "COWORK_HARNESS_LINT_EXTRA_FINDINGS"
|
|
1335
|
+
|
|
1336
|
+
|
|
1337
|
+
def _wrapper_loader_findings():
|
|
1338
|
+
"""Findings from the harness's own scenario loader, handed over by the `cowork-harness lint` wrapper.
|
|
1339
|
+
|
|
1340
|
+
The wrapper runs the loader `run`/`record` use before spawning this script and writes what it rejects
|
|
1341
|
+
(`scenario-invalid`, `baseline-unknown`) to a JSON file named by COWORK_HARNESS_LINT_EXTRA_FINDINGS.
|
|
1342
|
+
Merged here so ONE renderer, ONE --min-severity filter and ONE exit rule apply to both sets.
|
|
1343
|
+
|
|
1344
|
+
Honoured only together with COWORK_HARNESS_PROG, which the wrapper always sets: a direct
|
|
1345
|
+
`python3 scenario.py lint` never reads the variable, even if the shell happens to carry it. A blank
|
|
1346
|
+
value is unset. A file that can't be read or doesn't have the expected shape is an ERROR, never a
|
|
1347
|
+
silent drop -- a dropped loader finding would be a false green."""
|
|
1348
|
+
path = (os.environ.get(_EXTRA_FINDINGS_ENV) or "").strip()
|
|
1349
|
+
if not path or not os.environ.get("COWORK_HARNESS_PROG"):
|
|
1350
|
+
return []
|
|
1351
|
+
|
|
1352
|
+
def bad(why):
|
|
1353
|
+
return [
|
|
1354
|
+
Finding(
|
|
1355
|
+
"ERROR",
|
|
1356
|
+
"linter-extra-findings-invalid",
|
|
1357
|
+
f"could not read the scenario-loader findings handed over by cowork-harness: {why}",
|
|
1358
|
+
"This is a harness bug, not a scenario problem -- please report it. "
|
|
1359
|
+
"`cowork-harness record <file> --dry-run` checks that a scenario loads.",
|
|
1360
|
+
"(scenario.py)",
|
|
1361
|
+
)
|
|
1362
|
+
]
|
|
1363
|
+
|
|
1364
|
+
try:
|
|
1365
|
+
entries = json.loads(Path(path).read_text(encoding="utf-8"))
|
|
1366
|
+
except (OSError, ValueError) as e:
|
|
1367
|
+
return bad(str(e))
|
|
1368
|
+
if not isinstance(entries, list):
|
|
1369
|
+
return bad("expected a JSON array")
|
|
1370
|
+
out = []
|
|
1371
|
+
for i, x in enumerate(entries):
|
|
1372
|
+
if (
|
|
1373
|
+
not isinstance(x, dict)
|
|
1374
|
+
or x.get("severity") not in SEV_ORDER
|
|
1375
|
+
or not all(isinstance(x.get(k), str) for k in ("rule", "message", "fix", "file"))
|
|
1376
|
+
or not (x.get("line") is None or (isinstance(x.get("line"), int) and not isinstance(x.get("line"), bool)))
|
|
1377
|
+
):
|
|
1378
|
+
return bad(f"entry {i} is not a finding")
|
|
1379
|
+
out.append(Finding(x["severity"], x["rule"], x["message"], x["fix"], x["file"], x.get("line")))
|
|
1380
|
+
return out
|
|
1381
|
+
|
|
1382
|
+
|
|
1325
1383
|
def cmd_lint(args):
|
|
1326
1384
|
all_findings = []
|
|
1327
1385
|
# Expand directory args to their scenario files — mirrors src/run/inputs.ts `resolveInputs`: a SINGLE
|
|
@@ -1363,6 +1421,7 @@ def cmd_lint(args):
|
|
|
1363
1421
|
)
|
|
1364
1422
|
for f in args.files:
|
|
1365
1423
|
all_findings.extend(lint_file(f))
|
|
1424
|
+
all_findings.extend(_wrapper_loader_findings())
|
|
1366
1425
|
# Filter BEFORE rendering AND before the exit computation — deliberately, so --min-severity narrows
|
|
1367
1426
|
# what the run actually cares about. Filtering at render only would make `--strict --min-severity ERROR`
|
|
1368
1427
|
# print "0 findings" and still exit 1 (because --strict keys off the unfiltered set), which is
|