cowork-harness 3.1.0 → 3.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +55 -22
- package/.claude/skills/cowork-harness/references/ci-recipe.md +8 -6
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +14 -13
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +54 -1
- package/.claude/skills/cowork-harness/scripts/scenario.py +128 -16
- package/CHANGELOG.md +101 -0
- package/README.md +4 -4
- package/RELEASING.md +24 -2
- package/SECURITY.md +2 -0
- package/SPEC.md +20 -0
- package/dist/cli.js +1 -1
- package/dist/errors.js +63 -1
- package/dist/run/cassette.js +117 -57
- package/dist/run/execute.js +19 -3
- package/docs/README.md +12 -4
- package/docs/cassette.md +10 -0
- package/docs/ci.md +9 -3
- package/docs/cli.md +26 -15
- package/docs/companion-skill.md +3 -3
- package/docs/debugging.md +3 -3
- package/docs/decider-dir.md +3 -1
- package/docs/gotchas.md +2 -2
- package/docs/scenario.md +25 -10
- package/docs/session.md +1 -1
- package/docs/stats.md +1 -1
- package/examples/replays/README.md +1 -1
- package/llms.txt +2 -2
- package/package.json +1 -1
- package/python/README.md +7 -4
- package/python/test_scenario_lint.py +131 -0
- package/scripts/gen-schema.ts +53 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.2.0
|
|
7
|
+
tracks-harness: cowork-harness 3.2.0 (baseline desktop-1.40609.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.2.0` (baseline
|
|
29
29
|
> `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.2.0"`. **Pin `@^3.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -57,21 +57,15 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
57
57
|
Everything you do with the harness is one of **three loops**, and the rest of this skill is organized
|
|
58
58
|
into three Parts to match: **author** a scenario (Part I), **run / record / lock** it into a
|
|
59
59
|
reproducible regression (Part II), and **debug** a run that misbehaved or greened when it shouldn't
|
|
60
|
-
(Part III — reachable straight from here in one hop).
|
|
60
|
+
(Part III — reachable straight from here in one hop).
|
|
61
|
+
|
|
62
|
+
Pick the entry point you need. The first three are the everyday path — a quick liveness check, the
|
|
63
|
+
CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that hang off them:
|
|
61
64
|
|
|
62
65
|
- **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>"`. Fastest; no
|
|
63
66
|
scenario file.
|
|
64
67
|
- **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
|
|
65
68
|
This is the CI-grade path and most of this skill.
|
|
66
|
-
- **Regression-test your skill's ANSWER quality** (not just its behavior — does its guidance still lead to
|
|
67
|
-
correct answers after you edit it?) → author `semantic_matches` scenarios and gate on the per-claim
|
|
68
|
-
profile. See **Recipe 5** in `references/task-recipes.md` (validity, N≥3, discrimination — the traps).
|
|
69
|
-
- **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
|
|
70
|
-
<folder> --prompt "<probe>"`. Four model workloads and 10–20 minutes; budget from
|
|
71
|
-
`report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
|
|
72
|
-
**DO**" — routing, artifact location, narration — use `skill` instead**: no evaluator, a fraction of
|
|
73
|
-
the cost, and it answers that question directly. Report and evidence-package shapes:
|
|
74
|
-
`references/critique.md`.
|
|
75
69
|
- **A run failed — or greened and you don't trust it** (the debugging loop) → don't re-run and hope.
|
|
76
70
|
The run already wrote its evidence to a **kept run dir** (`~/.cowork-harness/runs/…`; `--keep` prints
|
|
77
71
|
the path, `trace <run-id>` finds it). **Localize the failure post-hoc** from that evidence:
|
|
@@ -83,6 +77,15 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
83
77
|
**"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
|
|
84
78
|
grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
|
|
85
79
|
surface; see `references/critique.md`.
|
|
80
|
+
- **Regression-test your skill's ANSWER quality** (not just its behavior — does its guidance still lead to
|
|
81
|
+
correct answers after you edit it?) → author `semantic_matches` scenarios and gate on the per-claim
|
|
82
|
+
profile. See **Recipe 5** in `references/task-recipes.md` (validity, N≥3, discrimination — the traps).
|
|
83
|
+
- **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
|
|
84
|
+
<folder> --prompt "<probe>"`. Four model workloads and 10–20 minutes; budget from
|
|
85
|
+
`report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
|
|
86
|
+
**DO**" — routing, artifact location, narration — use `skill` instead**: no evaluator, a fraction of
|
|
87
|
+
the cost, and it answers that question directly. Report and evidence-package shapes:
|
|
88
|
+
`references/critique.md`.
|
|
86
89
|
- **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
|
|
87
90
|
TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
|
|
88
91
|
**"Interactive" splits two ways — don't take the wrong branch.** Want to answer gates yourself *and*
|
|
@@ -380,7 +383,10 @@ emits a scenario `lint` would reject.
|
|
|
380
383
|
`lint` (exit 0) but a **hard error** in the runtime (`Unrecognized key: "<k>"`, exit 2) — so a scenario
|
|
381
384
|
that lints with warnings may still not run. To check whether a scenario actually loads, without spending:
|
|
382
385
|
`cowork-harness record <file.yaml> --dry-run` (exit 2 on a schema error; a directory reports each
|
|
383
|
-
`✗ broken:` file and exits 1).
|
|
386
|
+
`✗ broken:` file and exits 1). **Read the exit code, not just its sign:** `record <file>` — with or without
|
|
387
|
+
`--dry-run` — answers `2` for "did not load" and `1` for "loaded fine, but this record is refused" (a
|
|
388
|
+
pre-spend policy refusal; `--max-budget-usd` is the one refusal that keeps exit 2). Treating any non-zero
|
|
389
|
+
as "scenario broken" mis-reports every refused-but-valid scenario. Corollary: **the loader** fails LOUD on an unknown key (never silently) —
|
|
384
390
|
but **`replay` does not**: a frozen top-level key it doesn't recognize (e.g. `lane:` recorded pre-1.16.0) is
|
|
385
391
|
silently ignored and can flip a lane-sensitive verdict green; only frozen **assertion** keys stay
|
|
386
392
|
hard-rejected there. Full split + the v11 version-regime:
|
|
@@ -416,6 +422,28 @@ contradiction, duplicate cassette target) gate the batch. So a directory dry-run
|
|
|
416
422
|
real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
|
|
417
423
|
reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
|
|
418
424
|
|
|
425
|
+
**Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
|
|
426
|
+
concluded the free pre-flight was unavailable:
|
|
427
|
+
- **"Does my whole corpus still load?"** → the **directory** arm (`record scenarios/ --dry-run --quiet`,
|
|
428
|
+
the CI shape in `references/ci-recipe.md`). It reports every offender in one pass, and the
|
|
429
|
+
destination-policy verdict cannot red it: that arm knows no `--out`, so host-inventory and portability
|
|
430
|
+
are advisory `notes[]` at exit 0 while a file that cannot load is `✗ broken:` at exit 1. Limits worth
|
|
431
|
+
knowing: it is **non-recursive** (`readdirSync` — scenarios in subdirectories are never opened), a file
|
|
432
|
+
with no `prompt:` key reports as `· skipped:` rather than broken (so a renamed or mis-indented
|
|
433
|
+
`prompt:` reads as "not a scenario" and the batch still exits 0), exit 1 covers **refused**
|
|
434
|
+
(path-independent only — prompt policy, assert contradiction, duplicate target) as well as broken, and
|
|
435
|
+
`--quiet` suppresses the advisory notes entirely — it keeps `✗ broken:` / `✗ refused:` / `· skipped:`,
|
|
436
|
+
which is what you want in CI but means the notes are not a thing you will see there.
|
|
437
|
+
- **"Would THIS record be refused?"** → the **single-file** arm **with the flags and `--out` the real
|
|
438
|
+
record will get**. The destination it evaluates is `--out` if given, else `cassettes/<slug>.cassette.json`
|
|
439
|
+
*relative to your cwd* — so previewing from the repo root a record that really runs from a subdirectory
|
|
440
|
+
asks about a path that may not even exist, and a scenario whose cassette IS committed can come back
|
|
441
|
+
refused. Point it at the real destination and the answer is binding.
|
|
442
|
+
|
|
443
|
+
Neither question is answered by passing `--allow-host-inventory-fixture` to get past the refusal: that
|
|
444
|
+
flag is consent for a recording you intend to make, and reaching for it as a load-check habit is how it
|
|
445
|
+
stops meaning anything.
|
|
446
|
+
|
|
419
447
|
**Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
|
|
420
448
|
Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
|
|
421
449
|
default); pass `--out <path>` to put it somewhere tracked, e.g. `examples/replays/<name>.cassette.json`.
|
|
@@ -429,7 +457,7 @@ warns when the cassette would be written outside the scenario's tree, or when `s
|
|
|
429
457
|
outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
|
|
430
458
|
cassette lands). A warning, not a refusal — an out-of-tree throwaway cassette is legitimate; what was
|
|
431
459
|
missing was anything saying so while you could still act. Related: recording at a **host-inheriting** tier
|
|
432
|
-
(`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25).
|
|
460
|
+
(`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25 below).
|
|
433
461
|
The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
|
|
434
462
|
redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
|
|
435
463
|
for a cassette that cannot verify staleness from its own location — recoverable only by passing
|
|
@@ -600,7 +628,7 @@ Recognize these before "fixing" a non-bug:
|
|
|
600
628
|
`RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
|
|
601
629
|
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
602
630
|
|
|
603
|
-
The full
|
|
631
|
+
The full 20-code signal table (severity + per-signal opt-out) is in
|
|
604
632
|
[`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
|
|
605
633
|
the fuller narrative.
|
|
606
634
|
|
|
@@ -750,7 +778,8 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
750
778
|
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
|
|
751
779
|
plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
|
|
752
780
|
`--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
|
|
753
|
-
tier — instead of aggregating across them (a window spanning >1 generation warns
|
|
781
|
+
tier — instead of aggregating across them (a window spanning >1 generation warns, so an un-split A/B
|
|
782
|
+
average announces itself rather than passing as one number;
|
|
754
783
|
>1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
|
|
755
784
|
individual runs behind each summary with their `skillHash`/`runLabel`, so
|
|
756
785
|
you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
|
|
@@ -773,7 +802,7 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
773
802
|
`outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
|
|
774
803
|
`jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
|
|
775
804
|
`{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
|
|
776
|
-
semantics:
|
|
805
|
+
semantics: [`docs/cli.md` → What you get out](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md#what-you-get-out-inspectable-output) (repo-only); [`schema/run-result.json`](https://github.com/yaniv-golan/cowork-harness/blob/main/schema/run-result.json) is the
|
|
777
806
|
machine source.)
|
|
778
807
|
- **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
|
|
779
808
|
**`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
|
|
@@ -830,8 +859,12 @@ The startup banner now reads `type your message (/help for commands)` as a remin
|
|
|
830
859
|
|
|
831
860
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
832
861
|
|
|
833
|
-
Stated as *symptom → why → fix*.
|
|
834
|
-
|
|
862
|
+
Stated as *symptom → why → fix*. This is the **workflow/record/answer-path** view — the broader of the
|
|
863
|
+
two lists, but **not a strict superset**: `references/scenario-schema.md`'s *Authoring gotcha list*
|
|
864
|
+
carries a few assertion-level landmines this one omits (`transcript_no_host_path`'s scan width,
|
|
865
|
+
`egress.extra_allow`'s no-op on the provenanced `web_fetch` path, `replay_protocol_fidelity` not being
|
|
866
|
+
authorable). Reach for this list when debugging a run's behavior, that one while authoring `assert:`.
|
|
867
|
+
**The two lists are numbered independently** — a bare "gotcha N" means the list you are reading.
|
|
835
868
|
|
|
836
869
|
1. **An assertion passed but tested nothing on the PR gate.** *Why:* on a manifest-less cassette
|
|
837
870
|
`replay` skips filesystem/egress keys (`file_exists`, `user_visible_artifact`, `artifact_json`,
|
|
@@ -857,7 +890,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
857
890
|
|
|
858
891
|
3. **A multi-key `assert:` item is an AND.** A single list item with more than one key passes iff
|
|
859
892
|
**every** key passes. *Fix:* one concern per item unless you genuinely mean conjunction (and a
|
|
860
|
-
mixed-class conjunction still loses its filesystem half on replay — see gotcha 1).
|
|
893
|
+
mixed-class conjunction still loses its filesystem half on replay — see gotcha 1 below).
|
|
861
894
|
|
|
862
895
|
4. **`tool_called` doesn't mean "attempted".** Tool counts are authoritative and de-duped: a tool
|
|
863
896
|
that was *requested then denied* does **not** register as called. *Fix:* don't assert `tool_called`
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.2.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.2.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^3.
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.2.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -312,7 +312,9 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
312
312
|
|
|
313
313
|
This is the shape a CI step wants: **silent on success (no output, exit 0), loud and specific on
|
|
314
314
|
failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
|
|
315
|
-
offending file *and* the rejected key, and the step still exits 1.
|
|
315
|
+
offending file *and* the rejected key, one line per file, and the step still exits 1. (Silent is
|
|
316
|
+
literal only when your scenarios name a `fidelity:`; one still in the deprecation window prints one
|
|
317
|
+
defaulted-fidelity notice per scenario.) A scenario that lints with only
|
|
316
318
|
warnings can still be unloadable, so a green `lint` is not evidence the suite runs.
|
|
317
319
|
3. **Scenarios (replay)** — `cowork-harness replay cassettes/` on every PR (the committed `*.cassette.json`).
|
|
318
320
|
Token-free; content + structure + gate delivery.
|
|
@@ -342,7 +344,7 @@ jobs:
|
|
|
342
344
|
with: { node-version: '24' }
|
|
343
345
|
- uses: actions/setup-python@v5
|
|
344
346
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^3.
|
|
347
|
+
- run: npm i -g "cowork-harness@^3.2.0"
|
|
346
348
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
349
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
350
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +373,7 @@ jobs:
|
|
|
371
373
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
374
|
fi
|
|
373
375
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^3.
|
|
376
|
+
run: npm i -g "cowork-harness@^3.2.0"
|
|
375
377
|
- if: steps.guard.outputs.live == 'true'
|
|
376
378
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
379
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.2.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
|
-
# Scenario & session schema, assertion catalog, web_fetch,
|
|
1
|
+
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.2.0`
|
|
4
4
|
(baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -22,7 +22,7 @@ Everything below is the full field + assertion catalog.
|
|
|
22
22
|
- [Replay class — which assertions survive `replay`](#replay-class)
|
|
23
23
|
- [The web_fetch model](#the-web_fetch-model)
|
|
24
24
|
- [Scenario YAML vs the pytest lane](#scenario-yaml-vs-the-pytest-lane)
|
|
25
|
-
- [
|
|
25
|
+
- [Authoring gotcha list](#authoring-gotcha-list)
|
|
26
26
|
|
|
27
27
|
## Scenario YAML
|
|
28
28
|
|
|
@@ -307,7 +307,7 @@ same set live from the schema.
|
|
|
307
307
|
| `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
308
308
|
| `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
|
|
309
309
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
310
|
-
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected
|
|
310
|
+
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. **Omitting it does NOT allow deletes** — a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored; use `allow_outputs_delete: true` to accept an intended one. Detection is a post-run bash-command scan, not mount enforcement, so a green means none was *detected* |
|
|
311
311
|
| `no_delete_in_mounts: true` | no delete op touched ANY delete-denied mount — `outputs` plus every `rw` connected folder — except those waived by `allow_delete_in`. Production denies `unlink`/`rmdir` on every such mount, so `no_delete_in_outputs` covers only part of the real rule. **only `true` is valid** |
|
|
312
312
|
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
|
|
313
313
|
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
|
|
@@ -401,9 +401,9 @@ paraphrases (that re-records red). Structured JSON → assert it in YAML with **
|
|
|
401
401
|
`path` + operator); use the pytest lane (`assert_artifact_json`) only for predicates too complex for a
|
|
402
402
|
dotted path.
|
|
403
403
|
|
|
404
|
-
**VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`;
|
|
404
|
+
**VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; eleven
|
|
405
405
|
are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
|
|
406
|
-
`"success"`) and
|
|
406
|
+
`"success"`) and nine are **warn**-severity (informational, never flip pass/fail). All twenty signal
|
|
407
407
|
codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
408
408
|
|
|
409
409
|
| Code | Severity | Meaning |
|
|
@@ -431,7 +431,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
431
431
|
|
|
432
432
|
A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
|
|
433
433
|
overall run verdict and exit code — `assert result: success` alone won't catch it; check
|
|
434
|
-
`result.verdict.signals[].severity` or the run's exit code. Only the
|
|
434
|
+
`result.verdict.signals[].severity` or the run's exit code. Only the nine **warn** codes are truly benign.
|
|
435
435
|
|
|
436
436
|
## Replay class
|
|
437
437
|
|
|
@@ -546,14 +546,14 @@ real predicate over a skill's **structured JSON output**:
|
|
|
546
546
|
Find an artifact's real field paths by running once with `--keep`, then `cowork-harness inspect <run-dir>`
|
|
547
547
|
(a shallow field preview of each JSON artifact) or by reading the JSON directly.
|
|
548
548
|
|
|
549
|
-
##
|
|
549
|
+
## Authoring gotcha list
|
|
550
550
|
|
|
551
551
|
The "✓ passed ≠ correct" landmines relevant to **scenario/assertion authoring**, as
|
|
552
552
|
*symptom → why → fix*. `file:line` pointers track the version at the top of this file.
|
|
553
|
-
**Scope note:** this is the assertion/replay-focused view; the
|
|
554
|
-
|
|
555
|
-
|
|
556
|
-
|
|
553
|
+
**Scope note:** this is the assertion/replay-focused view; the companion `SKILL.md`'s Gotchas section
|
|
554
|
+
is the broader one (it adds workflow/record/answer-path landmines this reference omits). **Neither list
|
|
555
|
+
is a strict superset of the other** — reach for this one while authoring `assert:`, and SKILL's when
|
|
556
|
+
debugging a run's behavior. The two are **numbered independently**: a bare "gotcha N" means this list.
|
|
557
557
|
|
|
558
558
|
1. **Replay skips filesystem/egress assertions (two shapes) — with a loud warning.** *Full skip:* a pure
|
|
559
559
|
live-only `egress_*`/`no_delete_in_outputs`/`self_heal_ran`/`transcript_no_host_path` item on a
|
|
@@ -591,7 +591,8 @@ reference omits). Neither list is a strict superset of the other — reach for t
|
|
|
591
591
|
one concatenated string → use `[\s\S]`, not `.`. `transcript_matches` is case-insensitive.
|
|
592
592
|
|
|
593
593
|
7. **Multi-key assertion item = AND.** Passes iff every key passes. One concern per item unless
|
|
594
|
-
conjunction is intended (and a mixed-class conjunction loses its filesystem half on replay — gotcha 1
|
|
594
|
+
conjunction is intended (and a mixed-class conjunction loses its filesystem half on replay — gotcha 1
|
|
595
|
+
of THIS list; the two gotcha lists are numbered independently).
|
|
595
596
|
|
|
596
597
|
8. **`tool_called` proves a tool ran, not that it was attempted.** Tool counts are authoritative and
|
|
597
598
|
de-duped: a requested-then-denied tool does NOT register as called; the synthetic
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.2.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -122,5 +122,58 @@
|
|
|
122
122
|
"Stop",
|
|
123
123
|
"SubagentStop",
|
|
124
124
|
"UserPromptSubmit"
|
|
125
|
-
]
|
|
125
|
+
],
|
|
126
|
+
"enums": {
|
|
127
|
+
"answers.decide": [
|
|
128
|
+
"allow",
|
|
129
|
+
"deny"
|
|
130
|
+
],
|
|
131
|
+
"answers.else": [
|
|
132
|
+
"allow",
|
|
133
|
+
"deny"
|
|
134
|
+
],
|
|
135
|
+
"answers.grant": [
|
|
136
|
+
"once",
|
|
137
|
+
"domain"
|
|
138
|
+
],
|
|
139
|
+
"assert.path_denied.agent_scope": [
|
|
140
|
+
"main",
|
|
141
|
+
"subagent",
|
|
142
|
+
"any"
|
|
143
|
+
],
|
|
144
|
+
"assert.path_denied.source": [
|
|
145
|
+
"pretooluse",
|
|
146
|
+
"can_use_tool",
|
|
147
|
+
"permission_denied"
|
|
148
|
+
],
|
|
149
|
+
"assert.question_options.order": [
|
|
150
|
+
"exact",
|
|
151
|
+
"any"
|
|
152
|
+
],
|
|
153
|
+
"assert.result": [
|
|
154
|
+
"success",
|
|
155
|
+
"error"
|
|
156
|
+
],
|
|
157
|
+
"execution": [
|
|
158
|
+
"local",
|
|
159
|
+
"cloud-describe"
|
|
160
|
+
],
|
|
161
|
+
"fidelity": [
|
|
162
|
+
"protocol",
|
|
163
|
+
"container",
|
|
164
|
+
"microvm",
|
|
165
|
+
"hostloop",
|
|
166
|
+
"cowork"
|
|
167
|
+
],
|
|
168
|
+
"lane": [
|
|
169
|
+
"local",
|
|
170
|
+
"remote"
|
|
171
|
+
],
|
|
172
|
+
"on_unanswered": [
|
|
173
|
+
"fail",
|
|
174
|
+
"prompt",
|
|
175
|
+
"llm",
|
|
176
|
+
"first"
|
|
177
|
+
]
|
|
178
|
+
}
|
|
126
179
|
}
|
|
@@ -23,7 +23,10 @@ lint flags (see references/scenario-schema.md for the why of each):
|
|
|
23
23
|
(runtime rejects at LOAD time; tier rules suppressed)
|
|
24
24
|
E `requires_capabilities` on `fidelity: protocol` (probe can't run → hard-fails
|
|
25
25
|
unless allow_missing_capability)
|
|
26
|
-
E
|
|
26
|
+
E `enum-value-invalid` any enum-valued field carries a value the schema rejects (fidelity,
|
|
27
|
+
execution, lane, on_unanswered, answers[].decide/else/grant,
|
|
28
|
+
assert[].result/path_denied.*/question_options.order) — `agent` gets a
|
|
29
|
+
`on_unanswered: agent` → `llm` rename hint
|
|
27
30
|
E authored `replay_protocol_fidelity` assertion (replay-synthesized only)
|
|
28
31
|
E `assertions:` instead of `assert:` (block ignored → every check no-ops)
|
|
29
32
|
E a presence assert + its absence sibling (unsatisfiable; run/skill/record refuse it:
|
|
@@ -302,8 +305,82 @@ REGEX_KEYS = {
|
|
|
302
305
|
"reference_read",
|
|
303
306
|
"no_observed_reference_access",
|
|
304
307
|
}
|
|
305
|
-
|
|
306
|
-
|
|
308
|
+
# The authoritative enum-value map: field id -> allowed values, generated from the Zod schemas into
|
|
309
|
+
# `assertion-keys.json`'s `enums` block by walking BOTH the top-level ScenarioObject fields and the
|
|
310
|
+
# nested answers[]/assert[] item shapes (gen-schema.ts's buildEnumMap). Backs the generic
|
|
311
|
+
# `enum-value-invalid` lint rule below, and replaces two independent hand-copies that used to drift from
|
|
312
|
+
# the schema on their own: VALID_TIERS (now `ENUM_VALUES["fidelity"]`) and the old bespoke
|
|
313
|
+
# on_unanswered check's VALID_ON_UNANSWERED set (now `ENUM_VALUES["on_unanswered"]`).
|
|
314
|
+
_EMBEDDED_ENUMS = {
|
|
315
|
+
"fidelity": ["protocol", "container", "microvm", "hostloop", "cowork"],
|
|
316
|
+
# "cloud-describe" IS a valid schema value here -- the runtime rejects it as RESERVED at load
|
|
317
|
+
# (src/run/execute.ts: no cloud runner exists yet), so it must never be handed back to an author as
|
|
318
|
+
# the FIX for some other invalid `execution:` value. See _enum_fix_values below, which is where that
|
|
319
|
+
# carve-out is applied -- this map itself stays a faithful mirror of what the schema accepts.
|
|
320
|
+
"execution": ["local", "cloud-describe"],
|
|
321
|
+
"lane": ["local", "remote"],
|
|
322
|
+
"on_unanswered": ["fail", "prompt", "llm", "first"],
|
|
323
|
+
"answers.decide": ["allow", "deny"],
|
|
324
|
+
"answers.else": ["allow", "deny"],
|
|
325
|
+
"answers.grant": ["once", "domain"],
|
|
326
|
+
"assert.result": ["success", "error"],
|
|
327
|
+
"assert.path_denied.source": ["pretooluse", "can_use_tool", "permission_denied"],
|
|
328
|
+
"assert.path_denied.agent_scope": ["main", "subagent", "any"],
|
|
329
|
+
"assert.question_options.order": ["exact", "any"],
|
|
330
|
+
}
|
|
331
|
+
|
|
332
|
+
|
|
333
|
+
def _load_enums():
|
|
334
|
+
"""The authoritative enum-value map, from the generated assertion-keys.json sidecar. Falls back to
|
|
335
|
+
the embedded `_EMBEDDED_ENUMS` (kept equal to the generated map) with a loud warning if the file is
|
|
336
|
+
missing or predates the `enums` field."""
|
|
337
|
+
p = Path(__file__).resolve().parent / "assertion-keys.json"
|
|
338
|
+
try:
|
|
339
|
+
enums = json.loads(p.read_text(encoding="utf-8")).get("enums")
|
|
340
|
+
if not enums:
|
|
341
|
+
raise KeyError("enums")
|
|
342
|
+
return enums
|
|
343
|
+
except Exception:
|
|
344
|
+
print(
|
|
345
|
+
f"::warning:: assertion-keys.json missing or has no enums next to scenario.py ({p}) — "
|
|
346
|
+
"using a built-in enum-value map that may be stale (run `npm run schema`).",
|
|
347
|
+
file=sys.stderr,
|
|
348
|
+
)
|
|
349
|
+
return dict(_EMBEDDED_ENUMS)
|
|
350
|
+
|
|
351
|
+
|
|
352
|
+
# field id -> allowed values (generated from the zod schemas; see _load_enums)
|
|
353
|
+
ENUM_VALUES = _load_enums()
|
|
354
|
+
VALID_TIERS = tuple(ENUM_VALUES.get("fidelity", _EMBEDDED_ENUMS["fidelity"]))
|
|
355
|
+
|
|
356
|
+
|
|
357
|
+
def _enum_fix_values(field, allowed):
|
|
358
|
+
"""Allowed values to SHOW an author as the fix for an invalid `field`. Identical to the validation
|
|
359
|
+
set except for `execution`, where `cloud-describe` is schema-valid but runtime-RESERVED (a load-time
|
|
360
|
+
error at record/run/skill — src/run/execute.ts) — offering it as a remedy would just relocate the
|
|
361
|
+
failure from `lint` to `record --dry-run` rather than resolve it."""
|
|
362
|
+
if field == "execution":
|
|
363
|
+
return [v for v in allowed if v != "cloud-describe"]
|
|
364
|
+
return allowed
|
|
365
|
+
|
|
366
|
+
|
|
367
|
+
def _enum_finding(field, value, path):
|
|
368
|
+
"""An `enum-value-invalid` Finding if `value` is not in `field`'s allowed set, else None. Unknown
|
|
369
|
+
`field`s (not in ENUM_VALUES) are silently skipped -- that would be a linter bug, not an authoring
|
|
370
|
+
one, and every field this is called with is a schema enum by construction."""
|
|
371
|
+
allowed = ENUM_VALUES.get(field)
|
|
372
|
+
if not allowed or value in allowed:
|
|
373
|
+
return None
|
|
374
|
+
hint = " (`agent` was renamed to `llm`)" if field == "on_unanswered" and value == "agent" else ""
|
|
375
|
+
# A present-but-null key reads as `fidelity: None` in Python; show the YAML the author wrote.
|
|
376
|
+
shown = "null" if value is None else value
|
|
377
|
+
return Finding(
|
|
378
|
+
"ERROR",
|
|
379
|
+
"enum-value-invalid",
|
|
380
|
+
f"`{field}: {shown}` is not a valid value{hint}.",
|
|
381
|
+
f"Use one of: {' | '.join(_enum_fix_values(field, allowed))}.",
|
|
382
|
+
path,
|
|
383
|
+
)
|
|
307
384
|
|
|
308
385
|
# Gate-id tripwire: the `host-path-assert-cowork` WARN below embeds Cowork's
|
|
309
386
|
# host-loop gate id in offline Python (the linter never reads a baseline). The
|
|
@@ -747,19 +824,54 @@ def lint_doc(doc, path, raw_lines):
|
|
|
747
824
|
)
|
|
748
825
|
)
|
|
749
826
|
|
|
750
|
-
# E:
|
|
751
|
-
|
|
752
|
-
|
|
753
|
-
|
|
754
|
-
|
|
755
|
-
|
|
756
|
-
|
|
757
|
-
|
|
758
|
-
|
|
759
|
-
|
|
760
|
-
|
|
761
|
-
|
|
762
|
-
|
|
827
|
+
# E: enum-value-invalid — every enum-valued scenario field, top-level and nested. Schema-invalid
|
|
828
|
+
# values here are refused at load (`run`/`skill`/`record` all raise zod's `invalid_value`, exit 2)
|
|
829
|
+
# regardless of what `lint` says, so a scenario `lint --strict` calls clean while the loader hard-
|
|
830
|
+
# rejects it is exactly the false-green class this linter exists to catch. Generalizes what used to
|
|
831
|
+
# be a single bespoke check for `on_unanswered` alone — a fifth of the schema's actual enum surface
|
|
832
|
+
# (see ENUM_VALUES above); the `agent` -> `llm` rename hint survives as a field-specific case in
|
|
833
|
+
# _enum_finding.
|
|
834
|
+
# Membership, NOT `.get(...) is not None`: a key present with an empty value (`fidelity:` and
|
|
835
|
+
# nothing after it -- a one-keystroke slip) parses as null, which the loader rejects and which the
|
|
836
|
+
# retired on_unanswered check let through. Absent keys stay absent; they default legitimately.
|
|
837
|
+
for _field in ("fidelity", "execution", "lane", "on_unanswered"):
|
|
838
|
+
if _field in doc:
|
|
839
|
+
_value = doc[_field]
|
|
840
|
+
_f = _enum_finding(_field, _value, path)
|
|
841
|
+
if _f is not None:
|
|
842
|
+
findings.append(_f)
|
|
843
|
+
_answers = doc.get("answers")
|
|
844
|
+
if isinstance(_answers, list):
|
|
845
|
+
for _rule in _answers:
|
|
846
|
+
if not isinstance(_rule, dict):
|
|
847
|
+
continue
|
|
848
|
+
for _key in ("decide", "else", "grant"):
|
|
849
|
+
if _key in _rule:
|
|
850
|
+
_value = _rule[_key]
|
|
851
|
+
_f = _enum_finding(f"answers.{_key}", _value, path)
|
|
852
|
+
if _f is not None:
|
|
853
|
+
findings.append(_f)
|
|
854
|
+
for _item in items:
|
|
855
|
+
if "result" in _item:
|
|
856
|
+
_value = _item["result"]
|
|
857
|
+
_f = _enum_finding("assert.result", _value, path)
|
|
858
|
+
if _f is not None:
|
|
859
|
+
findings.append(_f)
|
|
860
|
+
_path_denied = _item.get("path_denied")
|
|
861
|
+
if isinstance(_path_denied, dict):
|
|
862
|
+
for _key in ("source", "agent_scope"):
|
|
863
|
+
if _key in _path_denied:
|
|
864
|
+
_value = _path_denied[_key]
|
|
865
|
+
_f = _enum_finding(f"assert.path_denied.{_key}", _value, path)
|
|
866
|
+
if _f is not None:
|
|
867
|
+
findings.append(_f)
|
|
868
|
+
_question_options = _item.get("question_options")
|
|
869
|
+
if isinstance(_question_options, dict):
|
|
870
|
+
if "order" in _question_options:
|
|
871
|
+
_value = _question_options["order"]
|
|
872
|
+
_f = _enum_finding("assert.question_options.order", _value, path)
|
|
873
|
+
if _f is not None:
|
|
874
|
+
findings.append(_f)
|
|
763
875
|
|
|
764
876
|
# E: authored replay_protocol_fidelity
|
|
765
877
|
if "replay_protocol_fidelity" in assert_keys:
|