cowork-harness 3.0.1 → 3.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +61 -22
- package/.claude/skills/cowork-harness/references/ci-recipe.md +8 -6
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +15 -13
- package/.claude/skills/cowork-harness/references/task-recipes.md +14 -6
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +54 -1
- package/.claude/skills/cowork-harness/scripts/scenario.py +128 -16
- package/CHANGELOG.md +161 -0
- package/README.md +4 -4
- package/RELEASING.md +24 -2
- package/SECURITY.md +2 -0
- package/SPEC.md +22 -1
- package/dist/cli.js +62 -6
- package/dist/errors.js +63 -1
- package/dist/run/cassette.js +230 -68
- package/dist/run/chat-result.js +2 -0
- package/dist/run/chat.js +6 -0
- package/dist/run/execute.js +42 -5
- package/dist/run/model-provenance.js +100 -0
- package/dist/run/run.js +12 -0
- package/dist/run/verdict.js +26 -0
- package/docs/README.md +12 -4
- package/docs/cassette.md +19 -2
- package/docs/ci.md +9 -3
- package/docs/cli.md +28 -16
- package/docs/companion-skill.md +3 -3
- package/docs/debugging.md +10 -7
- package/docs/decider-dir.md +3 -1
- package/docs/fidelity-gaps.md +55 -0
- package/docs/gotchas.md +2 -2
- package/docs/invariants.md +1 -1
- package/docs/scenario.md +31 -10
- package/docs/session.md +5 -3
- package/docs/stats.md +1 -1
- package/examples/replays/README.md +1 -1
- package/examples/sessions/default.yaml +1 -1
- package/llms.txt +2 -2
- package/package.json +1 -1
- package/python/README.md +7 -4
- package/python/test_scenario_lint.py +131 -0
- package/schema/cassette.v12.json +15 -3
- package/schema/run-result.json +30 -1
- package/scripts/gen-schema.ts +53 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.0
|
|
7
|
-
tracks-harness: cowork-harness 3.0
|
|
6
|
+
version: 3.2.0
|
|
7
|
+
tracks-harness: cowork-harness 3.2.0 (baseline desktop-1.40609.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.2.0` (baseline
|
|
29
29
|
> `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.2.0"`. **Pin `@^3.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -57,21 +57,15 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
57
57
|
Everything you do with the harness is one of **three loops**, and the rest of this skill is organized
|
|
58
58
|
into three Parts to match: **author** a scenario (Part I), **run / record / lock** it into a
|
|
59
59
|
reproducible regression (Part II), and **debug** a run that misbehaved or greened when it shouldn't
|
|
60
|
-
(Part III — reachable straight from here in one hop).
|
|
60
|
+
(Part III — reachable straight from here in one hop).
|
|
61
|
+
|
|
62
|
+
Pick the entry point you need. The first three are the everyday path — a quick liveness check, the
|
|
63
|
+
CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that hang off them:
|
|
61
64
|
|
|
62
65
|
- **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>"`. Fastest; no
|
|
63
66
|
scenario file.
|
|
64
67
|
- **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
|
|
65
68
|
This is the CI-grade path and most of this skill.
|
|
66
|
-
- **Regression-test your skill's ANSWER quality** (not just its behavior — does its guidance still lead to
|
|
67
|
-
correct answers after you edit it?) → author `semantic_matches` scenarios and gate on the per-claim
|
|
68
|
-
profile. See **Recipe 5** in `references/task-recipes.md` (validity, N≥3, discrimination — the traps).
|
|
69
|
-
- **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
|
|
70
|
-
<folder> --prompt "<probe>"`. Four model workloads and 10–20 minutes; budget from
|
|
71
|
-
`report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
|
|
72
|
-
**DO**" — routing, artifact location, narration — use `skill` instead**: no evaluator, a fraction of
|
|
73
|
-
the cost, and it answers that question directly. Report and evidence-package shapes:
|
|
74
|
-
`references/critique.md`.
|
|
75
69
|
- **A run failed — or greened and you don't trust it** (the debugging loop) → don't re-run and hope.
|
|
76
70
|
The run already wrote its evidence to a **kept run dir** (`~/.cowork-harness/runs/…`; `--keep` prints
|
|
77
71
|
the path, `trace <run-id>` finds it). **Localize the failure post-hoc** from that evidence:
|
|
@@ -83,6 +77,15 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
83
77
|
**"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
|
|
84
78
|
grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
|
|
85
79
|
surface; see `references/critique.md`.
|
|
80
|
+
- **Regression-test your skill's ANSWER quality** (not just its behavior — does its guidance still lead to
|
|
81
|
+
correct answers after you edit it?) → author `semantic_matches` scenarios and gate on the per-claim
|
|
82
|
+
profile. See **Recipe 5** in `references/task-recipes.md` (validity, N≥3, discrimination — the traps).
|
|
83
|
+
- **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
|
|
84
|
+
<folder> --prompt "<probe>"`. Four model workloads and 10–20 minutes; budget from
|
|
85
|
+
`report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
|
|
86
|
+
**DO**" — routing, artifact location, narration — use `skill` instead**: no evaluator, a fraction of
|
|
87
|
+
the cost, and it answers that question directly. Report and evidence-package shapes:
|
|
88
|
+
`references/critique.md`.
|
|
86
89
|
- **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
|
|
87
90
|
TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
|
|
88
91
|
**"Interactive" splits two ways — don't take the wrong branch.** Want to answer gates yourself *and*
|
|
@@ -380,7 +383,10 @@ emits a scenario `lint` would reject.
|
|
|
380
383
|
`lint` (exit 0) but a **hard error** in the runtime (`Unrecognized key: "<k>"`, exit 2) — so a scenario
|
|
381
384
|
that lints with warnings may still not run. To check whether a scenario actually loads, without spending:
|
|
382
385
|
`cowork-harness record <file.yaml> --dry-run` (exit 2 on a schema error; a directory reports each
|
|
383
|
-
`✗ broken:` file and exits 1).
|
|
386
|
+
`✗ broken:` file and exits 1). **Read the exit code, not just its sign:** `record <file>` — with or without
|
|
387
|
+
`--dry-run` — answers `2` for "did not load" and `1` for "loaded fine, but this record is refused" (a
|
|
388
|
+
pre-spend policy refusal; `--max-budget-usd` is the one refusal that keeps exit 2). Treating any non-zero
|
|
389
|
+
as "scenario broken" mis-reports every refused-but-valid scenario. Corollary: **the loader** fails LOUD on an unknown key (never silently) —
|
|
384
390
|
but **`replay` does not**: a frozen top-level key it doesn't recognize (e.g. `lane:` recorded pre-1.16.0) is
|
|
385
391
|
silently ignored and can flip a lane-sensitive verdict green; only frozen **assertion** keys stay
|
|
386
392
|
hard-rejected there. Full split + the v11 version-regime:
|
|
@@ -416,6 +422,28 @@ contradiction, duplicate cassette target) gate the batch. So a directory dry-run
|
|
|
416
422
|
real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
|
|
417
423
|
reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
|
|
418
424
|
|
|
425
|
+
**Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
|
|
426
|
+
concluded the free pre-flight was unavailable:
|
|
427
|
+
- **"Does my whole corpus still load?"** → the **directory** arm (`record scenarios/ --dry-run --quiet`,
|
|
428
|
+
the CI shape in `references/ci-recipe.md`). It reports every offender in one pass, and the
|
|
429
|
+
destination-policy verdict cannot red it: that arm knows no `--out`, so host-inventory and portability
|
|
430
|
+
are advisory `notes[]` at exit 0 while a file that cannot load is `✗ broken:` at exit 1. Limits worth
|
|
431
|
+
knowing: it is **non-recursive** (`readdirSync` — scenarios in subdirectories are never opened), a file
|
|
432
|
+
with no `prompt:` key reports as `· skipped:` rather than broken (so a renamed or mis-indented
|
|
433
|
+
`prompt:` reads as "not a scenario" and the batch still exits 0), exit 1 covers **refused**
|
|
434
|
+
(path-independent only — prompt policy, assert contradiction, duplicate target) as well as broken, and
|
|
435
|
+
`--quiet` suppresses the advisory notes entirely — it keeps `✗ broken:` / `✗ refused:` / `· skipped:`,
|
|
436
|
+
which is what you want in CI but means the notes are not a thing you will see there.
|
|
437
|
+
- **"Would THIS record be refused?"** → the **single-file** arm **with the flags and `--out` the real
|
|
438
|
+
record will get**. The destination it evaluates is `--out` if given, else `cassettes/<slug>.cassette.json`
|
|
439
|
+
*relative to your cwd* — so previewing from the repo root a record that really runs from a subdirectory
|
|
440
|
+
asks about a path that may not even exist, and a scenario whose cassette IS committed can come back
|
|
441
|
+
refused. Point it at the real destination and the answer is binding.
|
|
442
|
+
|
|
443
|
+
Neither question is answered by passing `--allow-host-inventory-fixture` to get past the refusal: that
|
|
444
|
+
flag is consent for a recording you intend to make, and reaching for it as a load-check habit is how it
|
|
445
|
+
stops meaning anything.
|
|
446
|
+
|
|
419
447
|
**Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
|
|
420
448
|
Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
|
|
421
449
|
default); pass `--out <path>` to put it somewhere tracked, e.g. `examples/replays/<name>.cassette.json`.
|
|
@@ -429,7 +457,7 @@ warns when the cassette would be written outside the scenario's tree, or when `s
|
|
|
429
457
|
outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
|
|
430
458
|
cassette lands). A warning, not a refusal — an out-of-tree throwaway cassette is legitimate; what was
|
|
431
459
|
missing was anything saying so while you could still act. Related: recording at a **host-inheriting** tier
|
|
432
|
-
(`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25).
|
|
460
|
+
(`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25 below).
|
|
433
461
|
The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
|
|
434
462
|
redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
|
|
435
463
|
for a cassette that cannot verify staleness from its own location — recoverable only by passing
|
|
@@ -568,6 +596,12 @@ Recognize these before "fixing" a non-bug:
|
|
|
568
596
|
would claim more than the evidence supports, and staying silent would read as clean. Mutually exclusive
|
|
569
597
|
with `undelivered_deliverables`, and quiet on a run that produced nothing to deliver. Not a skill defect —
|
|
570
598
|
a harness coverage gap (see the *File delivery* section of fidelity-gaps).
|
|
599
|
+
- **`model_fallback`** (`WARN`) — the agent switched off the requested model mid-run. Read the `trigger`:
|
|
600
|
+
`model_not_found` / `model_blocked` / `permission_denied` are properties of the **pin**, so every run of
|
|
601
|
+
this scenario falls back the same way until you change the pinned id; `overloaded` / `server_error` are
|
|
602
|
+
transient and a re-run may hold. The run's assertions still mean what they say — but they were produced
|
|
603
|
+
by a different model than the scenario names, so treat a green as evidence about the fallback model.
|
|
604
|
+
|
|
571
605
|
- **`mount_delete`** (`WARN`) — a delete touched a **delete-denied mount other than `outputs`**: a `rw`
|
|
572
606
|
connected folder. Production denies `unlink`/`rmdir` on *every* Cowork FUSE mount until per-mount
|
|
573
607
|
approval, not just outputs — a connected folder shows the identical default — so this run diverged from
|
|
@@ -594,7 +628,7 @@ Recognize these before "fixing" a non-bug:
|
|
|
594
628
|
`RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
|
|
595
629
|
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
596
630
|
|
|
597
|
-
The full
|
|
631
|
+
The full 20-code signal table (severity + per-signal opt-out) is in
|
|
598
632
|
[`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
|
|
599
633
|
the fuller narrative.
|
|
600
634
|
|
|
@@ -744,7 +778,8 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
744
778
|
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
|
|
745
779
|
plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
|
|
746
780
|
`--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
|
|
747
|
-
tier — instead of aggregating across them (a window spanning >1 generation warns
|
|
781
|
+
tier — instead of aggregating across them (a window spanning >1 generation warns, so an un-split A/B
|
|
782
|
+
average announces itself rather than passing as one number;
|
|
748
783
|
>1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
|
|
749
784
|
individual runs behind each summary with their `skillHash`/`runLabel`, so
|
|
750
785
|
you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
|
|
@@ -767,7 +802,7 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
767
802
|
`outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
|
|
768
803
|
`jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
|
|
769
804
|
`{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
|
|
770
|
-
semantics:
|
|
805
|
+
semantics: [`docs/cli.md` → What you get out](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md#what-you-get-out-inspectable-output) (repo-only); [`schema/run-result.json`](https://github.com/yaniv-golan/cowork-harness/blob/main/schema/run-result.json) is the
|
|
771
806
|
machine source.)
|
|
772
807
|
- **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
|
|
773
808
|
**`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
|
|
@@ -824,8 +859,12 @@ The startup banner now reads `type your message (/help for commands)` as a remin
|
|
|
824
859
|
|
|
825
860
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
826
861
|
|
|
827
|
-
Stated as *symptom → why → fix*.
|
|
828
|
-
|
|
862
|
+
Stated as *symptom → why → fix*. This is the **workflow/record/answer-path** view — the broader of the
|
|
863
|
+
two lists, but **not a strict superset**: `references/scenario-schema.md`'s *Authoring gotcha list*
|
|
864
|
+
carries a few assertion-level landmines this one omits (`transcript_no_host_path`'s scan width,
|
|
865
|
+
`egress.extra_allow`'s no-op on the provenanced `web_fetch` path, `replay_protocol_fidelity` not being
|
|
866
|
+
authorable). Reach for this list when debugging a run's behavior, that one while authoring `assert:`.
|
|
867
|
+
**The two lists are numbered independently** — a bare "gotcha N" means the list you are reading.
|
|
829
868
|
|
|
830
869
|
1. **An assertion passed but tested nothing on the PR gate.** *Why:* on a manifest-less cassette
|
|
831
870
|
`replay` skips filesystem/egress keys (`file_exists`, `user_visible_artifact`, `artifact_json`,
|
|
@@ -851,7 +890,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
851
890
|
|
|
852
891
|
3. **A multi-key `assert:` item is an AND.** A single list item with more than one key passes iff
|
|
853
892
|
**every** key passes. *Fix:* one concern per item unless you genuinely mean conjunction (and a
|
|
854
|
-
mixed-class conjunction still loses its filesystem half on replay — see gotcha 1).
|
|
893
|
+
mixed-class conjunction still loses its filesystem half on replay — see gotcha 1 below).
|
|
855
894
|
|
|
856
895
|
4. **`tool_called` doesn't mean "attempted".** Tool counts are authoritative and de-duped: a tool
|
|
857
896
|
that was *requested then denied* does **not** register as called. *Fix:* don't assert `tool_called`
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.0
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.2.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.0
|
|
20
|
+
(e.g. `version: "3.2.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^3.0
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.2.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -312,7 +312,9 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
312
312
|
|
|
313
313
|
This is the shape a CI step wants: **silent on success (no output, exit 0), loud and specific on
|
|
314
314
|
failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
|
|
315
|
-
offending file *and* the rejected key, and the step still exits 1.
|
|
315
|
+
offending file *and* the rejected key, one line per file, and the step still exits 1. (Silent is
|
|
316
|
+
literal only when your scenarios name a `fidelity:`; one still in the deprecation window prints one
|
|
317
|
+
defaulted-fidelity notice per scenario.) A scenario that lints with only
|
|
316
318
|
warnings can still be unloadable, so a green `lint` is not evidence the suite runs.
|
|
317
319
|
3. **Scenarios (replay)** — `cowork-harness replay cassettes/` on every PR (the committed `*.cassette.json`).
|
|
318
320
|
Token-free; content + structure + gate delivery.
|
|
@@ -342,7 +344,7 @@ jobs:
|
|
|
342
344
|
with: { node-version: '24' }
|
|
343
345
|
- uses: actions/setup-python@v5
|
|
344
346
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^3.0
|
|
347
|
+
- run: npm i -g "cowork-harness@^3.2.0"
|
|
346
348
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
349
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
350
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +373,7 @@ jobs:
|
|
|
371
373
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
374
|
fi
|
|
373
375
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^3.0
|
|
376
|
+
run: npm i -g "cowork-harness@^3.2.0"
|
|
375
377
|
- if: steps.guard.outputs.live == 'true'
|
|
376
378
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
379
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.0
|
|
3
|
+
Tracks `cowork-harness 3.2.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
|
-
# Scenario & session schema, assertion catalog, web_fetch,
|
|
1
|
+
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.2.0`
|
|
4
4
|
(baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -22,7 +22,7 @@ Everything below is the full field + assertion catalog.
|
|
|
22
22
|
- [Replay class — which assertions survive `replay`](#replay-class)
|
|
23
23
|
- [The web_fetch model](#the-web_fetch-model)
|
|
24
24
|
- [Scenario YAML vs the pytest lane](#scenario-yaml-vs-the-pytest-lane)
|
|
25
|
-
- [
|
|
25
|
+
- [Authoring gotcha list](#authoring-gotcha-list)
|
|
26
26
|
|
|
27
27
|
## Scenario YAML
|
|
28
28
|
|
|
@@ -307,7 +307,7 @@ same set live from the schema.
|
|
|
307
307
|
| `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
308
308
|
| `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
|
|
309
309
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
310
|
-
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected
|
|
310
|
+
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. **Omitting it does NOT allow deletes** — a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored; use `allow_outputs_delete: true` to accept an intended one. Detection is a post-run bash-command scan, not mount enforcement, so a green means none was *detected* |
|
|
311
311
|
| `no_delete_in_mounts: true` | no delete op touched ANY delete-denied mount — `outputs` plus every `rw` connected folder — except those waived by `allow_delete_in`. Production denies `unlink`/`rmdir` on every such mount, so `no_delete_in_outputs` covers only part of the real rule. **only `true` is valid** |
|
|
312
312
|
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
|
|
313
313
|
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
|
|
@@ -401,9 +401,9 @@ paraphrases (that re-records red). Structured JSON → assert it in YAML with **
|
|
|
401
401
|
`path` + operator); use the pytest lane (`assert_artifact_json`) only for predicates too complex for a
|
|
402
402
|
dotted path.
|
|
403
403
|
|
|
404
|
-
**VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`;
|
|
404
|
+
**VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; eleven
|
|
405
405
|
are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
|
|
406
|
-
`"success"`) and
|
|
406
|
+
`"success"`) and nine are **warn**-severity (informational, never flip pass/fail). All twenty signal
|
|
407
407
|
codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
408
408
|
|
|
409
409
|
| Code | Severity | Meaning |
|
|
@@ -421,6 +421,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
421
421
|
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
422
422
|
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
423
423
|
| `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
|
|
424
|
+
| `model_fallback` | warn | The agent fell back off the requested model mid-run (SDK `model_fallback` event); a `model_not_found`/`model_blocked` trigger repeats every run until the pin changes |
|
|
424
425
|
| `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
|
|
425
426
|
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
|
|
426
427
|
| `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
|
|
@@ -430,7 +431,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
430
431
|
|
|
431
432
|
A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
|
|
432
433
|
overall run verdict and exit code — `assert result: success` alone won't catch it; check
|
|
433
|
-
`result.verdict.signals[].severity` or the run's exit code. Only the
|
|
434
|
+
`result.verdict.signals[].severity` or the run's exit code. Only the nine **warn** codes are truly benign.
|
|
434
435
|
|
|
435
436
|
## Replay class
|
|
436
437
|
|
|
@@ -545,14 +546,14 @@ real predicate over a skill's **structured JSON output**:
|
|
|
545
546
|
Find an artifact's real field paths by running once with `--keep`, then `cowork-harness inspect <run-dir>`
|
|
546
547
|
(a shallow field preview of each JSON artifact) or by reading the JSON directly.
|
|
547
548
|
|
|
548
|
-
##
|
|
549
|
+
## Authoring gotcha list
|
|
549
550
|
|
|
550
551
|
The "✓ passed ≠ correct" landmines relevant to **scenario/assertion authoring**, as
|
|
551
552
|
*symptom → why → fix*. `file:line` pointers track the version at the top of this file.
|
|
552
|
-
**Scope note:** this is the assertion/replay-focused view; the
|
|
553
|
-
|
|
554
|
-
|
|
555
|
-
|
|
553
|
+
**Scope note:** this is the assertion/replay-focused view; the companion `SKILL.md`'s Gotchas section
|
|
554
|
+
is the broader one (it adds workflow/record/answer-path landmines this reference omits). **Neither list
|
|
555
|
+
is a strict superset of the other** — reach for this one while authoring `assert:`, and SKILL's when
|
|
556
|
+
debugging a run's behavior. The two are **numbered independently**: a bare "gotcha N" means this list.
|
|
556
557
|
|
|
557
558
|
1. **Replay skips filesystem/egress assertions (two shapes) — with a loud warning.** *Full skip:* a pure
|
|
558
559
|
live-only `egress_*`/`no_delete_in_outputs`/`self_heal_ran`/`transcript_no_host_path` item on a
|
|
@@ -590,7 +591,8 @@ reference omits). Neither list is a strict superset of the other — reach for t
|
|
|
590
591
|
one concatenated string → use `[\s\S]`, not `.`. `transcript_matches` is case-insensitive.
|
|
591
592
|
|
|
592
593
|
7. **Multi-key assertion item = AND.** Passes iff every key passes. One concern per item unless
|
|
593
|
-
conjunction is intended (and a mixed-class conjunction loses its filesystem half on replay — gotcha 1
|
|
594
|
+
conjunction is intended (and a mixed-class conjunction loses its filesystem half on replay — gotcha 1
|
|
595
|
+
of THIS list; the two gotcha lists are numbered independently).
|
|
594
596
|
|
|
595
597
|
8. **`tool_called` proves a tool ran, not that it was attempted.** Tool counts are authoritative and
|
|
596
598
|
de-duped: a requested-then-denied tool does NOT register as called; the synthetic
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.0
|
|
5
|
+
Tracks `cowork-harness 3.2.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -46,6 +46,11 @@ what production would do. Two layers of defense:
|
|
|
46
46
|
statically knowable and only produce a non-failing informational note.
|
|
47
47
|
- **Manual (one-liner):** `grep -h '"effectiveFidelity"' cassettes/*.cassette.json | sort | uniq -c`
|
|
48
48
|
shows the tier distribution of the fleet at a glance.
|
|
49
|
+
- **Do NOT re-record for a `[note]` alone.** A `·`/`[note]` row is informational and the run still exits
|
|
50
|
+
**0** — `session-fingerprint: predates \`model\` coverage` and `prompt-assets: cassette predates …` both
|
|
51
|
+
mean "this cassette was recorded before that field existed, and everything the hash *does* cover still
|
|
52
|
+
matches". Re-recording buys the new coverage and nothing else, so on a fleet of heavy cassettes it is a
|
|
53
|
+
real bill for no verdict change. Re-record when a `✗` says to (baseline moved, skill drift, tier moved).
|
|
49
54
|
|
|
50
55
|
### Cassette anatomy (what you're looking at when you open one)
|
|
51
56
|
|
|
@@ -66,7 +71,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](htt
|
|
|
66
71
|
| `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
|
|
67
72
|
| `scenarioSource` | Relative path to the authored YAML this was recorded from |
|
|
68
73
|
| `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
|
|
69
|
-
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
74
|
+
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (model/folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
70
75
|
| `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
|
|
71
76
|
| `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
|
|
72
77
|
| `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
|
|
@@ -241,10 +246,13 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
241
246
|
unrecoverable if skipped. `skillHash` is content-exact, so an edit mid-batch silently splits the
|
|
242
247
|
dataset into two generations — `stats --group-by skill-hash` separates them afterwards, but a hash
|
|
243
248
|
whose source was never committed names a generation that is unrecoverable, which makes the
|
|
244
|
-
comparison uninterpretable rather than merely noisy. And with no `model:` in the session
|
|
245
|
-
`--model` on the
|
|
246
|
-
before/after can silently straddle two models
|
|
247
|
-
|
|
249
|
+
comparison uninterpretable rather than merely noisy. And with no `model:` in the session and no
|
|
250
|
+
`--model` on the command (every lane takes it), each run uses whatever the staged agent binary
|
|
251
|
+
defaults to, so a before/after can silently straddle two models — the run warns when nothing pinned
|
|
252
|
+
one. Read `result.json` back to confirm: `modelSource` says whether anything pinned the model at all,
|
|
253
|
+
and `modelPinHonored` whether the pin survived (**absent means unverifiable, not "yes"**). `models`
|
|
254
|
+
lists what served the run — ignore any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent
|
|
255
|
+
fabricated locally, not a model).
|
|
248
256
|
3. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
|
|
249
257
|
`result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
|
|
250
258
|
content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
|
|
@@ -122,5 +122,58 @@
|
|
|
122
122
|
"Stop",
|
|
123
123
|
"SubagentStop",
|
|
124
124
|
"UserPromptSubmit"
|
|
125
|
-
]
|
|
125
|
+
],
|
|
126
|
+
"enums": {
|
|
127
|
+
"answers.decide": [
|
|
128
|
+
"allow",
|
|
129
|
+
"deny"
|
|
130
|
+
],
|
|
131
|
+
"answers.else": [
|
|
132
|
+
"allow",
|
|
133
|
+
"deny"
|
|
134
|
+
],
|
|
135
|
+
"answers.grant": [
|
|
136
|
+
"once",
|
|
137
|
+
"domain"
|
|
138
|
+
],
|
|
139
|
+
"assert.path_denied.agent_scope": [
|
|
140
|
+
"main",
|
|
141
|
+
"subagent",
|
|
142
|
+
"any"
|
|
143
|
+
],
|
|
144
|
+
"assert.path_denied.source": [
|
|
145
|
+
"pretooluse",
|
|
146
|
+
"can_use_tool",
|
|
147
|
+
"permission_denied"
|
|
148
|
+
],
|
|
149
|
+
"assert.question_options.order": [
|
|
150
|
+
"exact",
|
|
151
|
+
"any"
|
|
152
|
+
],
|
|
153
|
+
"assert.result": [
|
|
154
|
+
"success",
|
|
155
|
+
"error"
|
|
156
|
+
],
|
|
157
|
+
"execution": [
|
|
158
|
+
"local",
|
|
159
|
+
"cloud-describe"
|
|
160
|
+
],
|
|
161
|
+
"fidelity": [
|
|
162
|
+
"protocol",
|
|
163
|
+
"container",
|
|
164
|
+
"microvm",
|
|
165
|
+
"hostloop",
|
|
166
|
+
"cowork"
|
|
167
|
+
],
|
|
168
|
+
"lane": [
|
|
169
|
+
"local",
|
|
170
|
+
"remote"
|
|
171
|
+
],
|
|
172
|
+
"on_unanswered": [
|
|
173
|
+
"fail",
|
|
174
|
+
"prompt",
|
|
175
|
+
"llm",
|
|
176
|
+
"first"
|
|
177
|
+
]
|
|
178
|
+
}
|
|
126
179
|
}
|