cowork-harness 1.22.0 → 1.24.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +94 -13
- package/.claude/skills/cowork-harness/references/ci-recipe.md +46 -11
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +2 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +34 -11
- package/.claude/skills/cowork-harness/references/task-recipes.md +24 -7
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +57 -10
- package/CHANGELOG.md +516 -0
- package/DESIGN.md +3 -3
- package/README.md +40 -15
- package/SPEC.md +1 -1
- package/baselines/desktop-1.28929.0.json +13 -3
- package/baselines/desktop-1.30096.1.json +577 -0
- package/baselines/desktop-1.32352.0.json +580 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +241 -7
- package/dist/assert.js +259 -2
- package/dist/cli.js +36 -2
- package/dist/critique/command.js +53 -6
- package/dist/run/artifacts.js +11 -0
- package/dist/run/cassette.js +263 -29
- package/dist/run/execute.js +4 -0
- package/dist/run/run.js +14 -1
- package/dist/run/verdict.js +5 -3
- package/dist/runtime/image-capabilities.js +33 -0
- package/dist/sync/cowork-sync.js +667 -42
- package/dist/types.js +47 -1
- package/docs/README.md +9 -2
- package/docs/cassette.md +39 -4
- package/docs/critique.md +25 -5
- package/docs/debugging.md +22 -0
- package/docs/discovery.md +9 -0
- package/docs/fidelity-gaps.md +88 -9
- package/docs/gotchas.md +6 -0
- package/docs/maintenance.md +5 -3
- package/docs/scenario.md +27 -8
- package/docs/session.md +10 -1
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/llms.txt +2 -0
- package/package.json +1 -1
- package/schema/cassette.v10.json +18 -0
- package/schema/cassette.v11.json +18 -0
- package/schema/run-result.json +15 -2
- package/schema/scenario.schema.json +82 -1
- package/scripts/check-versions.ts +198 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.24.0
|
|
7
|
+
tracks-harness: cowork-harness 1.24.0 (baseline desktop-1.32352.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
26
|
-
> `desktop-1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.24.0` (baseline
|
|
26
|
+
> `desktop-1.32352.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.24.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.24.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.24.0"`. **Pin `@>=1.24.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -293,6 +293,8 @@ them by what you're trying to prove:
|
|
|
293
293
|
| the skill didn't error out of a tool | `tool_no_error: <regex>`, `max_tool_errors: <N>` |
|
|
294
294
|
| it didn't waste repeated identical calls | `max_redundant_tool_calls: <N>` |
|
|
295
295
|
| a deliverable reached the user | `user_visible_artifact: <path>` (+ `no_scratchpad_leak: true` if it delivers via `present_files` — **`container` only**) |
|
|
296
|
+
| an internal name/path did **not** leak into a delivered file | `artifact_text: {artifact, not_contains}` — `artifact_json`'s companion for non-JSON bodies; literal path, no glob, so one entry per delivered surface |
|
|
297
|
+
| a named path must **not** exist after the run | `file_absent: <path>` (**live/verify-run only**) — do NOT invert `no_unexpected_files`: that is an allowlist over *newly created* files and needs a pre-run manifest |
|
|
296
298
|
| a to-do workflow finished | `all_tasks_completed: true`, `task_status: {match, status}` |
|
|
297
299
|
| a skill / connector / tool was **offered** | `skill_available`, `connector_available`, `tool_available` (all `<regex>`) |
|
|
298
300
|
| a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
|
|
@@ -301,6 +303,7 @@ them by what you're trying to prove:
|
|
|
301
303
|
| a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
|
|
302
304
|
| no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
|
|
303
305
|
| a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
|
|
306
|
+
| the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
|
|
304
307
|
| a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
|
|
305
308
|
| every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
|
|
306
309
|
| a context compaction happened | `compaction_occurred: true` |
|
|
@@ -393,10 +396,28 @@ the discovery/encode/record dance entirely and answer gates **live during the re
|
|
|
393
396
|
`record --decider-dir`/`--decider-llm` (the cassette is flagged non-deterministic but replays deterministically).
|
|
394
397
|
`run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
|
|
395
398
|
`cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
|
|
396
|
-
refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing)
|
|
397
|
-
cannot green something a paid run would reject. On a directory it reports every offender and the batch
|
|
399
|
+
refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
|
|
400
|
+
cassette-portability pre-flight below**, so it cannot green something a paid run would reject. On a directory it reports every offender and the batch
|
|
398
401
|
cost estimate. `lint` checks the assertion invariants (both above).
|
|
399
402
|
|
|
403
|
+
**Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
|
|
404
|
+
Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
|
|
405
|
+
default); pass `--out <path>` to put it somewhere tracked, e.g. `examples/replays/<name>.cassette.json`.
|
|
406
|
+
That choice is permanent: the cassette rewrites `scenario.session` and `scenarioSource` **relative to
|
|
407
|
+
its own directory** at record time, so moving the file later — a different `--out`, a `git mv`, a copy
|
|
408
|
+
into another repo — leaves those unresolvable and
|
|
409
|
+
`verify-cassettes` reports `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you
|
|
410
|
+
re-record at the new location. **`record` now says so BEFORE it spends:** a pre-flight — at the same
|
|
411
|
+
pre-spend point as the host-inventory refusal, and in `record --dry-run`, so the rehearsal is free —
|
|
412
|
+
warns when the cassette would be written outside the scenario's tree, or when `session:` itself lives
|
|
413
|
+
outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
|
|
414
|
+
cassette lands). A warning, not a refusal — an out-of-tree throwaway cassette is legitimate; what was
|
|
415
|
+
missing was anything saying so while you could still act. Related: recording at a **host-inheriting** tier
|
|
416
|
+
(`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25).
|
|
417
|
+
The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
|
|
418
|
+
redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
|
|
419
|
+
for a permanently unverifiable cassette.
|
|
420
|
+
|
|
400
421
|
**Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
|
|
401
422
|
scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
|
|
402
423
|
(and `verify-run`) read the gates + offered option labels out of that run's `events.jsonl` for free. Iterate
|
|
@@ -548,6 +569,42 @@ The full 17-code signal table (severity + per-signal opt-out) is in
|
|
|
548
569
|
[`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
|
|
549
570
|
the fuller narrative.
|
|
550
571
|
|
|
572
|
+
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
573
|
+
|
|
574
|
+
A single green proves the run passed **once**. Two questions need more than that, and both have a
|
|
575
|
+
discipline that is cheap to follow and expensive to skip.
|
|
576
|
+
|
|
577
|
+
**"Did it pass, or pass once?"** → `--repeat N` (2-100, on `skill` AND `run`) samples the same
|
|
578
|
+
skill+prompt N times and prints a variance rollup instead of a single verdict. `--min-pass-rate` sets
|
|
579
|
+
the batch threshold, `--stop-on-diverge` stops the moment flakiness is proven, `--max-budget-usd` caps
|
|
580
|
+
spend.
|
|
581
|
+
|
|
582
|
+
**"Does the skill actually help?"** → `--ablate-skill` runs the prompt with every skill/plugin
|
|
583
|
+
discovery source removed, so the agent answers from its own priors. **It is ONE arm, not a paired
|
|
584
|
+
experiment**: this invocation is the control. Run the same prompt a second time *without* the flag for
|
|
585
|
+
the treatment arm and compare them yourself. Composed with `--repeat 5` it produces **5 ablated runs
|
|
586
|
+
and 0 treatment runs** — N samples of the control, which is the intended reading and is not an A/B.
|
|
587
|
+
Every ablated run is stamped `ablated: true` in `result.json`; a run that isn't stamped is a real run.
|
|
588
|
+
What the harness gives you here is the run execution and the control arm — designing the comparison
|
|
589
|
+
(scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
|
|
590
|
+
|
|
591
|
+
**Measurement hygiene — four things that silently invalidate a batch:**
|
|
592
|
+
|
|
593
|
+
1. **Pin the model.** With no `model:` in the session (or `--model` on the `skill` lane) the run uses
|
|
594
|
+
whatever the staged agent binary defaults to — not a harness constant, and it can move under a
|
|
595
|
+
baseline bump. Read `result.json`'s `models` back before believing any cross-run comparison — and when
|
|
596
|
+
you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
|
|
597
|
+
fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
|
|
598
|
+
array purely by whether such a turn occurred.
|
|
599
|
+
2. **Commit the skill first.** `fingerprint.skillHash` is content-exact, so an edit mid-batch silently
|
|
600
|
+
splits your dataset into two generations — and a hash whose source was never committed identifies a
|
|
601
|
+
generation that is unrecoverable. `stats --group-by skill-hash` separates them after the fact;
|
|
602
|
+
nothing recovers the source.
|
|
603
|
+
3. **Check which arm you actually ran** before analysing anything: `ablated` and
|
|
604
|
+
`context.availableSkills` in each `result.json`.
|
|
605
|
+
4. **Read `skillsInvoked`.** A rep where the skill never triggered is a measurement of the model, not
|
|
606
|
+
of your skill — discard or re-run it.
|
|
607
|
+
|
|
551
608
|
### Checking whether a background run is alive
|
|
552
609
|
|
|
553
610
|
Never use `ps aux` to check on a `cowork-harness` run you launched in the background — it only sees
|
|
@@ -741,11 +798,13 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
741
798
|
|
|
742
799
|
1. **An assertion passed but tested nothing on the PR gate.** *Why:* on a manifest-less cassette
|
|
743
800
|
`replay` skips filesystem/egress keys (`file_exists`, `user_visible_artifact`, `artifact_json`,
|
|
744
|
-
`egress_*`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`); a
|
|
801
|
+
`artifact_text`, `egress_*`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`); a
|
|
802
|
+
*mixed* item like
|
|
745
803
|
`{result, egress_denied}` greens on `result` while its `egress_denied` half is dropped. (`record`
|
|
746
804
|
snapshots an `artifacts` manifest, which makes
|
|
747
|
-
`file_exists`/`user_visible_artifact`/`artifact_json`/`computer_links_resolve`
|
|
748
|
-
replay-checkable — but the live-only egress keys stay skipped
|
|
805
|
+
`file_exists`/`user_visible_artifact`/`artifact_json`/`artifact_text`/`computer_links_resolve`
|
|
806
|
+
replay-checkable — but the live-only egress keys stay skipped, and `file_absent` is never
|
|
807
|
+
replay-checkable at all: proving absence needs an exhaustive, healthy walk a manifest does not record.) *Fix:* put egress/live-only checks on
|
|
749
808
|
a live gate; keep one concern per `assert:` item; run the linter. The harness warns loudly on skip.
|
|
750
809
|
|
|
751
810
|
2. **A steered gate answer never reached the model.** *Why:* `serializeDecision` must emit
|
|
@@ -754,8 +813,8 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
754
813
|
channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
|
|
755
814
|
reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
|
|
756
815
|
Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
|
|
757
|
-
string — send an array; a scalar is treated as one selection.) `question_asked` / `
|
|
758
|
-
`gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
|
|
816
|
+
string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` /
|
|
817
|
+
`questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
|
|
759
818
|
old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
|
|
760
819
|
on unobserved delivery (absence of evidence is failure, not neutral).
|
|
761
820
|
|
|
@@ -865,7 +924,9 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
865
924
|
14. **A positional `choose` (`first` / index) is order-dependent.** `choose: "2"` survives label drift
|
|
866
925
|
but NOT option *re-ordering* — if the gate presents its options in a different order run-to-run, the
|
|
867
926
|
index lands on a different option (a silent re-record flake). Prefer an exact label when order is
|
|
868
|
-
stable; `lint` flags positional `choose` with an advisory.
|
|
927
|
+
stable; `lint` flags positional `choose` with an advisory. Unstable option order is also what the
|
|
928
|
+
**user** sees — a reordered gate puts a different choice in the default slot — so pin what was shown
|
|
929
|
+
with `question_options`, rather than only hardening the answer rule against it.
|
|
869
930
|
15. **A scripted `choose:` matching no offered option HARD-fails the run — `on_unanswered: first` does NOT
|
|
870
931
|
backstop it.** This is distinct from an *unanswered* gate (no rule matched → falls to `on_unanswered`): a
|
|
871
932
|
rule that DID match the gate but whose `choose:` names a label the gate never offered (the model reworded
|
|
@@ -970,6 +1031,26 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
970
1031
|
scanner's `host-inventory` class on an already-committed cassette. Passing one where the other command
|
|
971
1032
|
wants it fails as an unrecognized flag — they don't interchange. Depth: `references/ci-recipe.md`.
|
|
972
1033
|
|
|
1034
|
+
26. **A `skill`-lane `PASS` does not mean the skill ran, or that the run was the one you wanted.** *Why:*
|
|
1035
|
+
an open-ended `skill` run has no `assert:` block, so its verdict reports only that **no guard fired**
|
|
1036
|
+
(no error, stall, host-path leak, `outputs/` delete, permissive auto-allow or capability gap). On
|
|
1037
|
+
`run` the same word additionally means *your assertions held*; on `skill --repeat N`, `PASS — N/N`
|
|
1038
|
+
means N runs cleared the guards — it says nothing about which model served them, whether the skill
|
|
1039
|
+
was invoked, or whether they were the ablated arm. *Fix:* read the three fields the record already
|
|
1040
|
+
carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all),
|
|
1041
|
+
`models` (which model), `ablated` + `context.availableSkills` (which arm). An answer that reads
|
|
1042
|
+
exactly like skill output is not evidence: the skill's own source is mounted where the model can
|
|
1043
|
+
read it — in production too — so on a self-referential prompt it may read `SKILL.md` and answer
|
|
1044
|
+
directly, with `skillActivity` empty.
|
|
1045
|
+
|
|
1046
|
+
27. **`allow_stall: true` is a scenario assertion, so the `skill` lane cannot use it.** *Why:* the
|
|
1047
|
+
`stalled` guard fires when a run's final message ends in `?` with no productive tool call after the
|
|
1048
|
+
last gate — which includes a complete answer that closes by *offering* a follow-up ("want me to run
|
|
1049
|
+
this through a structured pass?"). The documented opt-out lives in an `assert:` block, and an
|
|
1050
|
+
open-ended `skill` run has none, so the failure message names a remedy that lane can't perform.
|
|
1051
|
+
*Fix:* on `skill`, read the final message before believing `stalled`, or move the check to a
|
|
1052
|
+
`run` scenario where `allow_stall: true` is authorable.
|
|
1053
|
+
|
|
973
1054
|
For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
|
|
974
1055
|
files in `references/` (the gotchas above are the full list; the references repeat only the
|
|
975
1056
|
assertion/replay-relevant ones).
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.24.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -32,7 +32,7 @@ jobs:
|
|
|
32
32
|
- uses: actions/checkout@v4
|
|
33
33
|
- name: Stage the agent binary (official channel, sha256-verified — see https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md)
|
|
34
34
|
run: |
|
|
35
|
-
V=2.1.
|
|
35
|
+
V=2.1.229 # match your scenario's pinned baseline's agentVersion
|
|
36
36
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
37
37
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
38
38
|
# verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
|
|
@@ -58,13 +58,17 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
58
58
|
GitHub-hosted runners, no token/Docker/agent:
|
|
59
59
|
|
|
60
60
|
```yaml
|
|
61
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
61
|
+
- run: npm i -g "cowork-harness@>=1.24.0"
|
|
62
62
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
63
63
|
# no silent false-greens. WITHOUT --strict this
|
|
64
64
|
# step cannot fail on a WARN-class rule (e.g.
|
|
65
65
|
# vacuous-gate-assert) — it would print the
|
|
66
66
|
# finding and still exit 0. --min-severity WARN
|
|
67
|
-
# keeps the advisory INFO class advisory.
|
|
67
|
+
# keeps the advisory INFO class advisory — pair them.
|
|
68
|
+
# Bare `lint --strict` fails on INFO too, which reds
|
|
69
|
+
# on scenarios that are perfectly fine. (`lint-skill
|
|
70
|
+
# --strict` never fails on INFO: same flag name, a
|
|
71
|
+
# different rule. Do not carry one over to the other.)
|
|
68
72
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness — FAILS on a stale recording
|
|
69
73
|
# ALSO fails on a leaked host inventory: recording at
|
|
70
74
|
# protocol/hostloop freezes YOUR machine's MCP servers,
|
|
@@ -114,16 +118,22 @@ The split is not just about tokens — it decides **where each lane can run**:
|
|
|
114
118
|
`transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
|
|
115
119
|
`max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
|
|
116
120
|
`allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
|
|
117
|
-
`allow_stall` (no-op passes); plus the gate keys `question_asked` / `
|
|
118
|
-
`gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
119
|
-
(`file_exists` / `user_visible_artifact` / `artifact_json`) **if** it carries an artifact
|
|
121
|
+
`allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
|
|
122
|
+
`questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
123
|
+
(`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
|
|
124
|
+
manifest. `file_absent` is in neither class — it is live/verify-run only.
|
|
120
125
|
**That list is illustrative, not the authoritative set** — more keys are replay-checkable than fit a
|
|
121
126
|
paragraph, and a hand-typed enumeration is exactly what goes stale. For the current set, ask the CLI:
|
|
122
127
|
|
|
123
128
|
```bash
|
|
124
|
-
cowork-harness assertions --list --output-format json # every key
|
|
129
|
+
cowork-harness assertions --list --output-format json # every key + its one-line semantics
|
|
125
130
|
```
|
|
126
131
|
|
|
132
|
+
It emits `{key, description}` — there is no structured replay-class field to filter on; the
|
|
133
|
+
live-only / manifest / `controlOut` preconditions are stated in each key's `description` prose and,
|
|
134
|
+
in full, in the catalog's replay-class tables in
|
|
135
|
+
[`references/scenario-schema.md`](./scenario-schema.md).
|
|
136
|
+
|
|
127
137
|
**`replay --mutate`** is a distinct, reporting-only diagnostic on this same lane: it perturbs each
|
|
128
138
|
recorded JSON artifact value one at a time, re-runs the assertions against the perturbed cassette,
|
|
129
139
|
and reports which perturbations NOTHING caught — those are the fields your `assert:` block leaves
|
|
@@ -279,7 +289,7 @@ jobs:
|
|
|
279
289
|
with: { node-version: '24' }
|
|
280
290
|
- uses: actions/setup-python@v5
|
|
281
291
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
282
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
292
|
+
- run: npm i -g "cowork-harness@>=1.24.0"
|
|
283
293
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
284
294
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
285
295
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -308,7 +318,7 @@ jobs:
|
|
|
308
318
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
309
319
|
fi
|
|
310
320
|
- if: steps.guard.outputs.live == 'true'
|
|
311
|
-
run: npm i -g "cowork-harness@>=1.
|
|
321
|
+
run: npm i -g "cowork-harness@>=1.24.0"
|
|
312
322
|
- if: steps.guard.outputs.live == 'true'
|
|
313
323
|
run: cowork-harness run scenarios/ --output-format json
|
|
314
324
|
env:
|
|
@@ -335,6 +345,31 @@ scenario = `result === "success" && assertions.every(pass)`. Exit code is non-ze
|
|
|
335
345
|
fails or a run errors, so a plain `cowork-harness run scenarios/` is already CI-ready without parsing
|
|
336
346
|
JSON.
|
|
337
347
|
|
|
348
|
+
**Telling *why* a run failed, without scraping stderr.** Each result carries a `verdict` whose
|
|
349
|
+
`failures[]` is the one place every failure reason is enumerated, in one shape (the same object lands
|
|
350
|
+
in `result.json` and in the stdout envelope, by construction). Each entry is
|
|
351
|
+
`{kind, assertion?: "<key>", message}`, and **`kind` is the discriminator**:
|
|
352
|
+
|
|
353
|
+
- **`assertion`** — one of *your* `assert:` items failed; `assertion` names its key (`file_exists`,
|
|
354
|
+
`semantic_matches`, …).
|
|
355
|
+
- **`guard`** — a signal you didn't author: `stalled`, `permissive_auto_allow`, `missing_capability`, a
|
|
356
|
+
host-path leak, an `outputs/` delete, an infra or transport error, an unanswered gate.
|
|
357
|
+
- **`staleness`** — on `replay --strict` / `--assert-from` / `--reassert`, skill or baseline drift,
|
|
358
|
+
which those modes escalate to a hard failure on purpose (frozen events must not green an edited
|
|
359
|
+
assert against a skill whose current source produces something else).
|
|
360
|
+
- **`cassette-format`** — the cassette is too new for this build to interpret.
|
|
361
|
+
- **`coverage`** — a `verify-run` answer-coverage miss: a gate the run fired that your `answers:` block
|
|
362
|
+
does not cover.
|
|
363
|
+
|
|
364
|
+
So `jq '[.verdict.failures[] | select(.kind=="assertion")]'` answers "did MY assertions pass?" and
|
|
365
|
+
`select(.kind=="staleness")` answers "is the cassette stale?", from the envelope alone. Do not infer
|
|
366
|
+
either from the exit code: every kind lands on exit 1.
|
|
367
|
+
|
|
368
|
+
> **Do not filter on whether `assertion` is present.** That was the only discriminator before `kind`
|
|
369
|
+
> existed and it never worked in both directions: `coverage` entries carry a key too (an internal
|
|
370
|
+
> `answer_coverage` marker), so they read as authored asserts, while `guard`, `staleness` and
|
|
371
|
+
> `cassette-format` all arrive key-less and indistinguishable from one another.
|
|
372
|
+
|
|
338
373
|
**For the commands in this recipe, stdout carries the machine envelope and nothing else.** Without
|
|
339
374
|
`--output-format json`, `run` / `record` / `replay` / `verify-cassettes` / `status` write their whole
|
|
340
375
|
human rendering — warnings, verdict, `status`'s summary line — to **stderr**, and stdout stays empty.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 1.
|
|
3
|
+
Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -210,7 +210,7 @@ up often enough to spell out:
|
|
|
210
210
|
- **A green `replay` proves "same as when recorded," not "correct today."** `replay` never touches
|
|
211
211
|
a filesystem or network — it re-evaluates assertions from the frozen cassette. A fixed set of
|
|
212
212
|
keys is live-only and **skipped outright** on replay (absent from `assertions[]`, not vacuously
|
|
213
|
-
passed): `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
|
|
213
|
+
passed): `file_absent`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
|
|
214
214
|
`egress_allowed`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back`, and `expect_denied`.
|
|
215
215
|
Everything else that *is* evaluated is checked against the **recording**, not fresh behavior — a
|
|
216
216
|
green replay says the skill produced these events when it was recorded, not that it still does
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.24.0`
|
|
4
|
+
(baseline `desktop-1.32352.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -103,6 +103,13 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
|
|
|
103
103
|
Relative paths resolve from the file's own directory, so a scenario + session + referenced files
|
|
104
104
|
form a relocatable bundle. `~` expands to home.
|
|
105
105
|
|
|
106
|
+
> **A recorded cassette is NOT part of that bundle and is NOT relocatable.** It rewrites its own
|
|
107
|
+
> references relative to **its own** directory at record time (`scenario.session` and
|
|
108
|
+
> `scenarioSource`), so moving it afterwards — a different `--out`, a
|
|
109
|
+
> `git mv`, a copy into another repo — leaves them unresolvable and `verify-cassettes` reports
|
|
110
|
+
> `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you re-record at the new location.
|
|
111
|
+
> Decide where a cassette will live *before* you record it.
|
|
112
|
+
|
|
106
113
|
## Session YAML
|
|
107
114
|
|
|
108
115
|
A session (`sessions/*.yaml`) captures everything you'd configure in Cowork **before the first
|
|
@@ -111,7 +118,15 @@ session = your setup.**
|
|
|
111
118
|
|
|
112
119
|
```yaml
|
|
113
120
|
# model & reasoning
|
|
114
|
-
model: claude-opus-4-8 # omit
|
|
121
|
+
model: claude-opus-4-8 # omit ONLY to track the agent default: with no `model:` the harness emits
|
|
122
|
+
# no --model flag at all, so the run uses whatever the STAGED AGENT BINARY
|
|
123
|
+
# defaults to — not a harness constant, and it can move under a baseline
|
|
124
|
+
# bump. PIN IT for anything you will compare (a --repeat batch, a
|
|
125
|
+
# before/after, a with/without); read result.json's `models` back to confirm,
|
|
126
|
+
# IGNORING any `<…>`-wrapped entry (`<synthetic>` = a turn the agent
|
|
127
|
+
# fabricated locally, not a model id).
|
|
128
|
+
# On the ad-hoc `skill` lane there is no session file, so `--model <id>`
|
|
129
|
+
# (or COWORK_HARNESS_MODEL) is the ONLY way to pin it.
|
|
115
130
|
account_name: my-account # OPTIONAL — display name rendered into {{accountName}} / the prompt's
|
|
116
131
|
# "User name:" line; NOT a credential/identity selector (see src/prompt.ts,
|
|
117
132
|
# https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md)
|
|
@@ -282,9 +297,11 @@ same set live from the schema.
|
|
|
282
297
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
283
298
|
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
|
|
284
299
|
| `no_delete_in_mounts: true` | no delete op touched ANY delete-denied mount — `outputs` plus every `rw` connected folder — except those waived by `allow_delete_in`. Production denies `unlink`/`rmdir` on every such mount, so `no_delete_in_outputs` covers only part of the real rule. **only `true` is valid** |
|
|
285
|
-
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning |
|
|
300
|
+
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
|
|
286
301
|
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
|
|
287
302
|
| `self_heal_ran: <bool>` | a plugin-root self-heal script was (not) invoked |
|
|
303
|
+
| `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); evidence-unavailable on `lane: remote` / `preRunOrigin: remote-unavailable` |
|
|
304
|
+
| `artifact_text: {artifact, contains?, not_contains?, matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read |
|
|
288
305
|
| `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
|
|
289
306
|
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
|
|
290
307
|
| `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
|
|
@@ -312,7 +329,7 @@ same set live from the schema.
|
|
|
312
329
|
`skill_triggered`/`no_skill_triggered`, `skill_available`, `connector_available`, `skill_tool_used`,
|
|
313
330
|
`subagent_type` are **regex** (unanchored, case-insensitive). So `tool_called: mcp__workspace__*` (glob) but
|
|
314
331
|
`tool_available: mcp__workspace__.*` (regex) — a `.*` in a `tool_called` glob is a load-time schema error, not silently-matches-nothing.
|
|
315
|
-
| `skill_tool_used: {skill, tool}` | a tool whose name matches `tool` ran inside a skill-activation window whose `skillId` matches `skill` (`RunResult.skillActivity`) — evidence-unavailable if skill-activity telemetry is absent; heuristic for inline skills (a sticky, sequential window matching the agent's `activeSkill` scope, not an exact per-tool boundary) |
|
|
332
|
+
| `skill_tool_used: {skill, tool}` | a tool whose name matches `tool` ran inside a skill-activation window whose `skillId` matches `skill` (`RunResult.skillActivity`) — evidence-unavailable if skill-activity telemetry is absent; heuristic for inline skills (a sticky, sequential window matching the agent's `activeSkill` scope, not an exact per-tool boundary). **Scope:** the window's tool counts **include sub-agent calls** made during it, so this key can't say which agent called (`subagent_tool_used` is the sub-agent-only claim), and it matches tool **names** only — never the path/args, so "did it read *this* file" isn't expressible (per-sub-agent reads are recorded at `subagents[].referencesRead`, readable but not assertable) |
|
|
316
333
|
| `max_cost_usd: <N>` | the run's SDK-reported cost is ≤ N USD — evidence-unavailable if cost telemetry is absent. **Replay asserts the frozen recording's cost, not fresh spend** — a real regression needs a live `run` |
|
|
317
334
|
| `max_tokens: <N>` | `usage.input_tokens + usage.output_tokens` ≤ N (cache tokens excluded) — same replay caveat as `max_cost_usd` |
|
|
318
335
|
| `tool_calls_max: <N>` | total top-level tool calls (sum of `toolCounts`) ≤ N — meaningfully replay-checkable (re-drive recomputes `toolCounts` deterministically) |
|
|
@@ -327,7 +344,8 @@ same set live from the schema.
|
|
|
327
344
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
328
345
|
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
329
346
|
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
330
|
-
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose text matches |
|
|
347
|
+
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
|
|
348
|
+
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered — what the user was actually SHOWN. `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable |
|
|
331
349
|
| `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
|
|
332
350
|
| `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
|
|
333
351
|
| `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
|
|
@@ -340,7 +358,7 @@ same set live from the schema.
|
|
|
340
358
|
| `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
|
|
341
359
|
| `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
|
|
342
360
|
| `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
|
|
343
|
-
| `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider) |
|
|
361
|
+
| `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
|
|
344
362
|
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
345
363
|
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
346
364
|
| `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
|
|
@@ -423,7 +441,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
|
|
|
423
441
|
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
|
|
424
442
|
`allow_stall` are also kept on replay, evaluated as no-op passes.
|
|
425
443
|
|
|
426
|
-
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `questions_count_max`,
|
|
444
|
+
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `questions_count_max`,
|
|
427
445
|
`gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`,
|
|
428
446
|
`path_denied`, `no_path_denied` (the latter three are also `fidelity: hostloop`-only — see the assertion
|
|
429
447
|
table). With `controlOut` present they evaluate; on an old
|
|
@@ -435,13 +453,16 @@ the question keys: a custom hook's block/allow decision is an opaque async reply
|
|
|
435
453
|
built-in Task hook's view and could vacuously pass `no_hook_blocked` even if a custom hook genuinely blocked.
|
|
436
454
|
|
|
437
455
|
**Filesystem — replay-checkable WITH an artifact manifest:** `file_exists`, `user_visible_artifact`,
|
|
438
|
-
`artifact_json`, `computer_links_resolve`
|
|
456
|
+
`artifact_json`, `artifact_text`, `computer_links_resolve` (+ `computer_links_resolve_if_present`) run on
|
|
457
|
+
replay when the cassette carries an `artifacts` snapshot
|
|
439
458
|
(`record` captures `outputs/` + connected folders; `replay` materializes it). `artifact_json` needs the
|
|
440
459
|
small-file JSON `body` inlined; a hash-only entry still satisfies `file_exists`. `computer_links_resolve`
|
|
441
460
|
resolves a `/sessions/…/mnt/…`-shaped link directly against the manifest, and a host-shaped (hostloop) link
|
|
442
461
|
by first normalizing it to a mount-relative path (recorded connected-folder prefixes + the outputs/uploads
|
|
443
462
|
mounts) — replay has no live filesystem to check a host path against directly (that only happens on a live
|
|
444
|
-
`run`/`verify-run`).
|
|
463
|
+
`run`/`verify-run`). `artifact_text` is manifest-class for the same reason `artifact_json` is — a
|
|
464
|
+
body-less, symlinked or over-cap entry fails evidence-unavailable rather than passing.
|
|
465
|
+
Without a manifest (older cassettes) all six are skipped (six need the manifest; two
|
|
445
466
|
more — `no_unexpected_files` and `input_unmodified` — need the pre-run path/hash capture, below); `no_unexpected_files` also
|
|
446
467
|
needs `preRunPaths` (≥0.24 recordings) — without it the key is excluded with a loud warning (live/verify-run
|
|
447
468
|
hard-fails evidence-unavailable instead). `input_unmodified` is the same shape but needs `preRunHashes`
|
|
@@ -451,7 +472,9 @@ warning. A green replay re-confirms
|
|
|
451
472
|
staleness `fingerprint` shows ANY skill/baseline drift, or `replay --fail-on-skill-drift` only on
|
|
452
473
|
skill-source drift; every replay result also reports it class-tagged in `staleness[]` for a JSON gate.
|
|
453
474
|
|
|
454
|
-
**Egress + other filesystem — still skipped on replay (live-only):** `
|
|
475
|
+
**Egress + other filesystem — still skipped on replay (live-only):** `file_absent` (proving a path is
|
|
476
|
+
ABSENT needs an exhaustive, healthy walk; a manifest records no walk health, so "not captured" and "not
|
|
477
|
+
there" are indistinguishable and the key would pass while proving nothing), `no_delete_in_outputs`,
|
|
455
478
|
`self_heal_ran`, `transcript_no_host_path`, `egress_*` / `expect_denied`, `no_mcp_error`, `max_peak_rss_bytes`,
|
|
456
479
|
`no_lost_write_back`. These run only on a live `run`/`record`.
|
|
457
480
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 1.
|
|
5
|
+
Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -17,11 +17,12 @@ a paid live re-record? Walk this tree — the answer is usually no:
|
|
|
17
17
|
`cowork-harness replay <cassette> --assert-from <scenario.yaml>`. Token-free, no re-record.
|
|
18
18
|
If the recording genuinely lacks the telemetry a key needs (very old cassettes), the key fails
|
|
19
19
|
**loud** as `evidence-unavailable` — that is correct behavior, not a bug; only then re-record.
|
|
20
|
-
2. **Gate keys** (`question_asked`, `questions_count_max`, `gate_answers_delivered`) on a cassette
|
|
20
|
+
2. **Gate keys** (`question_asked`, `question_options`, `questions_count_max`, `gate_answers_delivered`) on a cassette
|
|
21
21
|
**with `controlOut`** (any modern recording) → same token-free `--assert-from` path.
|
|
22
22
|
3. **Gate keys** on a **pre-`controlOut`** cassette → one re-record unlocks gate asserts for that
|
|
23
23
|
cassette permanently.
|
|
24
|
-
4. **Filesystem / egress keys** (`file_exists` without an artifact manifest, `
|
|
24
|
+
4. **Filesystem / egress keys** (`file_exists` without an artifact manifest, `file_absent` — live-only
|
|
25
|
+
whatever the cassette carries — `egress_allowed`,
|
|
25
26
|
`egress_denied`, `no_delete_in_outputs`) → live lane **by design**; replay skips them with a
|
|
26
27
|
loud `::warning::`. Keep them in the scenario, run them on the nightly live gate.
|
|
27
28
|
|
|
@@ -69,6 +70,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v11.json`](htt
|
|
|
69
70
|
| `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
|
|
70
71
|
| `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
|
|
71
72
|
| `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
|
|
73
|
+
| `environment.agentImage` | The `agentImage` sub-block records the rootfs image the recording ran against, stamped only for the tiers whose capabilities come from it (`container`, `hostloop`; `microvm` probes the Lima guest instead). `ref` is the resolved image ref, `configId` the local config id (built **and** pulled images, NOT comparable across machines), `registryDigest` the registry manifest digest (pulled only, and the only cross-machine-comparable identity). The image decides `missingCapabilityUse`, which `computeVerdict` fails on, so replaying against a different rootfs can move a verdict with nothing in the cassette having changed — `replay` compares `registryDigest` first and prints an advisory note, never a failure. Re-inspected at cassette-WRITE time, so a rebuild between run and record records the later image. **Absent** on cassettes recorded before this field existed, and never backfilled |
|
|
72
74
|
|
|
73
75
|
## Recipe 3 — Set up redaction BEFORE your first hostloop/protocol record
|
|
74
76
|
|
|
@@ -175,9 +177,16 @@ degrade the advice. It is real work to calibrate; these steps are the traps that
|
|
|
175
177
|
from `RunResult.assertions[].semanticClaims`. Do **not** chase single-run all-pass — set `min_pass` to
|
|
176
178
|
the reliably-hit core for a green verdict, and treat the per-claim rates as the real signal. (N=1
|
|
177
179
|
routinely mislabels a stable 0/3 as "intermittent" and vice-versa.)
|
|
178
|
-
5. **Check discrimination — does the skill actually help?** Run one rep with the skill NOT installed
|
|
179
|
-
|
|
180
|
-
|
|
180
|
+
5. **Check discrimination — does the skill actually help?** Run one rep with the skill NOT installed and
|
|
181
|
+
compare. `--ablate-skill` is the flag for it: it empties every skill/plugin discovery source for
|
|
182
|
+
**that one invocation**, so the agent answers from its own priors, and stamps the result
|
|
183
|
+
`ablated: true`. It is the **control arm only** — run the same prompt again *without* the flag for
|
|
184
|
+
the treatment arm; `--ablate-skill --repeat 5` gives you 5 control runs and 0 treatment runs, not an
|
|
185
|
+
A/B. (Inspecting an organically not-invoked rep works too, but is not a substitute: that rep may
|
|
186
|
+
differ for other reasons.) If the answer still scores high without the skill, that claim is
|
|
187
|
+
answerable from priors and tests the model, not your skill — strengthen it (a skill-specific fact) or
|
|
188
|
+
drop it. Everything past "run both arms" — scrubbing giveaways, shuffling, judging blind, unblinding
|
|
189
|
+
after grading — is yours to build; the harness supplies the runs and the control.
|
|
181
190
|
6. **Gate a change on the profile diff.** Capture the per-claim profile before your edit (the baseline),
|
|
182
191
|
make the edit, re-capture, and compare per claim: a claim that DROPPED (e.g. 3/3 → 0/3) is a regression
|
|
183
192
|
your edit caused; a claim already at 0/3 (a known gap) cannot regress. That turns "did my SKILL.md
|
|
@@ -228,7 +237,15 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
228
237
|
- `cowork-harness inspect <run-dir>` → what the run produced, plus the run's `label` and `skillHash`.
|
|
229
238
|
- In-run alternative: dispatch a checker **sub-agent** (maker/checker) whose result folds into the
|
|
230
239
|
verdict.
|
|
231
|
-
2. **
|
|
240
|
+
2. **Commit the skill before a measurement batch, and pin the model.** Two cheap disciplines that are
|
|
241
|
+
unrecoverable if skipped. `skillHash` is content-exact, so an edit mid-batch silently splits the
|
|
242
|
+
dataset into two generations — `stats --group-by skill-hash` separates them afterwards, but a hash
|
|
243
|
+
whose source was never committed names a generation that is unrecoverable, which makes the
|
|
244
|
+
comparison uninterpretable rather than merely noisy. And with no `model:` in the session (or
|
|
245
|
+
`--model` on the `skill` lane) each run uses whatever the staged agent binary defaults to, so a
|
|
246
|
+
before/after can silently straddle two models; read `result.json`'s `models` back to confirm —
|
|
247
|
+
ignoring any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent fabricated locally, not a model).
|
|
248
|
+
3. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
|
|
232
249
|
`result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
|
|
233
250
|
content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
|
|
234
251
|
no `skillHash` on the `skill` lane at all, so verify the field is present before pairing on it** (a run that mounts nothing
|
|
@@ -10,6 +10,7 @@
|
|
|
10
10
|
"allow_stall",
|
|
11
11
|
"allow_undelivered_deliverables",
|
|
12
12
|
"artifact_json",
|
|
13
|
+
"artifact_text",
|
|
13
14
|
"compaction_occurred",
|
|
14
15
|
"computer_links_resolve",
|
|
15
16
|
"computer_links_resolve_if_present",
|
|
@@ -17,6 +18,7 @@
|
|
|
17
18
|
"dispatch_count_max",
|
|
18
19
|
"egress_allowed",
|
|
19
20
|
"egress_denied",
|
|
21
|
+
"file_absent",
|
|
20
22
|
"file_exists",
|
|
21
23
|
"gate_answer_count_min",
|
|
22
24
|
"gate_answers_delivered",
|
|
@@ -41,6 +43,7 @@
|
|
|
41
43
|
"path_denied",
|
|
42
44
|
"present_files_called",
|
|
43
45
|
"question_asked",
|
|
46
|
+
"question_options",
|
|
44
47
|
"questions_count_max",
|
|
45
48
|
"replay_protocol_fidelity",
|
|
46
49
|
"result",
|