cowork-harness 2.2.0 → 2.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +31 -12
- package/.claude/skills/cowork-harness/references/ci-recipe.md +15 -13
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +14 -12
- package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +8 -1
- package/CHANGELOG.md +236 -0
- package/DESIGN.md +3 -3
- package/README.md +41 -11
- package/baselines/desktop-1.37937.1.json +816 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +27 -0
- package/dist/assert.js +77 -0
- package/dist/cli.js +15 -4
- package/dist/run/cassette.js +255 -77
- package/dist/run/execute.js +59 -0
- package/dist/run/trace-view.js +39 -1
- package/dist/sync/cowork-sync.js +207 -8
- package/dist/types.js +39 -9
- package/docs/README.md +3 -0
- package/docs/cassette.md +36 -11
- package/docs/debugging.md +25 -0
- package/docs/fidelity-gaps.md +5 -1
- package/docs/maintenance.md +15 -1
- package/docs/scenario.md +13 -10
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +41 -46
- package/examples/replays/example-pdf-skill.cassette.json +90 -90
- package/examples/replays/hostloop-computer-links.cassette.json +62 -68
- package/llms.txt +3 -1
- package/package.json +1 -1
- package/schema/scenario.schema.json +28 -9
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 2.
|
|
7
|
-
tracks-harness: cowork-harness 2.
|
|
6
|
+
version: 2.3.0
|
|
7
|
+
tracks-harness: cowork-harness 2.3.0 (baseline desktop-1.37937.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -16,14 +16,17 @@ in the shell; this skill tells you how to author scenarios, pick a fidelity tier
|
|
|
16
16
|
path, place assertions in the right CI lane, and avoid the harness's "✓ passed ≠ actually correct"
|
|
17
17
|
traps.
|
|
18
18
|
|
|
19
|
+
`cowork-harness` is an unofficial, independent project — not affiliated with or endorsed by
|
|
20
|
+
Anthropic. Say so if a user asks what it is.
|
|
21
|
+
|
|
19
22
|
The single most important idea: **a green run is not automatically a correct run.** The harness has
|
|
20
23
|
several ways to no-op a check while still producing a green run (skip an assertion on replay — now
|
|
21
24
|
flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an empty egress
|
|
22
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
26
|
the highest-value part. Read it.
|
|
24
27
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.
|
|
26
|
-
> `desktop-1.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.3.0` (baseline
|
|
29
|
+
> `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
31
|
|
|
29
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
42
|
|
|
40
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.3.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.3.0"`. **Pin `@^2.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
46
|
|
|
44
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -309,6 +312,7 @@ them by what you're trying to prove:
|
|
|
309
312
|
| no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
|
|
310
313
|
| a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
|
|
311
314
|
| the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
|
|
315
|
+
| the user was **told something specific** at a gate | `question_context: {when_question, matches}` — a regex over the question label + option labels + option **descriptions**. Reach for this when the wording may land in an option's `description`, which `question_asked` and `question_options` cannot see |
|
|
312
316
|
| a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
|
|
313
317
|
| every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
|
|
314
318
|
| a context compaction happened | `compaction_occurred: true` |
|
|
@@ -402,8 +406,15 @@ the discovery/encode/record dance entirely and answer gates **live during the re
|
|
|
402
406
|
`run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
|
|
403
407
|
`cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
|
|
404
408
|
refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
|
|
405
|
-
cassette-portability pre-flight below**, so it cannot green something a paid run would reject.
|
|
406
|
-
|
|
409
|
+
cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
|
|
410
|
+
guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
|
|
411
|
+
one a paid run would give. On a **directory** the path-dependent verdicts (host-inventory, cassette
|
|
412
|
+
portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
|
|
413
|
+
the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
|
|
414
|
+
takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
|
|
415
|
+
contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
|
|
416
|
+
real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
|
|
417
|
+
reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
|
|
407
418
|
|
|
408
419
|
**Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
|
|
409
420
|
Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
|
|
@@ -426,7 +437,12 @@ for a cassette that cannot verify staleness from its own location — recoverabl
|
|
|
426
437
|
|
|
427
438
|
**Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
|
|
428
439
|
scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
|
|
429
|
-
(and `verify-run`) read the gates + offered option
|
|
440
|
+
(and `verify-run`) read the gates + every offered option's **label and `description`** out of that run's
|
|
441
|
+
`events.jsonl` for free — a skill routinely puts the sentence the user is actually deciding on in a
|
|
442
|
+
`description`, and `question_context:` is the key that gates on it (`question_options:` compares labels
|
|
443
|
+
only). When a view renders no field you need, read `events.jsonl` directly rather than concluding the text
|
|
444
|
+
was never delivered — the views are a digest, and the record is wider (`jq` recipes in
|
|
445
|
+
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)). Iterate
|
|
430
446
|
your `answers:` against that kept run, then record once. **But the kept run is a snapshot:** if you change the
|
|
431
447
|
skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
|
|
432
448
|
current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
|
|
@@ -821,7 +837,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
821
837
|
channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
|
|
822
838
|
reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
|
|
823
839
|
Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
|
|
824
|
-
string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` /
|
|
840
|
+
string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` / `question_context` /
|
|
825
841
|
`questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
|
|
826
842
|
old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
|
|
827
843
|
on unobserved delivery (absence of evidence is failure, not neutral).
|
|
@@ -1034,11 +1050,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
1034
1050
|
([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
|
|
1035
1051
|
→ "File delivery" has the binary-verified detail; repo-only.)
|
|
1036
1052
|
|
|
1037
|
-
25. **
|
|
1038
|
-
--allow-host-inventory-fixture`
|
|
1053
|
+
25. **Three host-inventory flags — two on `record`, one on `verify-cassettes`.** `record
|
|
1054
|
+
--allow-host-inventory-fixture` proceeds past the PRE-FLIGHT refusal when recording a host-inheriting
|
|
1039
1055
|
(`protocol`/`hostloop`/`cowork`-resolving-to-hostloop) cassette into a repo-visible path — otherwise
|
|
1040
1056
|
`record` refuses before it spends (freezing this machine's MCP servers/agents/account into a committed
|
|
1041
|
-
fixture is the risk).
|
|
1057
|
+
fixture is the risk). It bypasses that check and **nothing else**: the finished recording is still
|
|
1058
|
+
scanned, and a real finding still quarantines it, so you never have to audit the session by hand to
|
|
1059
|
+
pass it. Writing a recording the scan DID flag is the separate `record
|
|
1060
|
+
--allow-host-inventory-findings`. That pre-spend check **warns rather than refuses when the cassette already
|
|
1042
1061
|
exists** — refusing would fire on every `--rerecord-stale` pass — and it reads the tier and the
|
|
1043
1062
|
destination path, never the bytes. So `record` also scans the FINISHED recording, after redaction and
|
|
1044
1063
|
before the write: a `host-inventory`/`machine-inventory` finding on a repo-visible path is
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "2.
|
|
20
|
+
(e.g. `version: "2.3.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.246 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
41
41
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
42
42
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^2.
|
|
70
|
+
- run: npm i -g "cowork-harness@^2.3.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -149,7 +149,7 @@ The split is not just about tokens — it decides **where each lane can run**:
|
|
|
149
149
|
`max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
|
|
150
150
|
`allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
|
|
151
151
|
`allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
|
|
152
|
-
`questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
152
|
+
`question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
153
153
|
(`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
|
|
154
154
|
manifest. `file_absent` is in neither class — it is live/verify-run only.
|
|
155
155
|
**That list is illustrative, not the authoritative set** — more keys are replay-checkable than fit a
|
|
@@ -237,12 +237,14 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
237
237
|
(`hostloop`, `protocol`) has an **empty** redaction policy — that combination commits real host paths
|
|
238
238
|
the `path` scanner then hard-fails at `verify-cassettes` time. The always-on scanner remains the
|
|
239
239
|
universal net (container-tier recordings can trip it too).
|
|
240
|
-
- **Host-inheriting record refused by default — `--allow-host-inventory-fixture`
|
|
241
|
-
`protocol`/`hostloop`/`cowork`-resolving-to-hostloop record into a
|
|
242
|
-
freeze THIS machine's MCP server names, agents, and account metadata
|
|
243
|
-
`record` refuses **before the paid spawn**.
|
|
244
|
-
|
|
245
|
-
|
|
240
|
+
- **Host-inheriting record refused by default — `--allow-host-inventory-fixture` bypasses the
|
|
241
|
+
PRE-FLIGHT, not the scan.** A `protocol`/`hostloop`/`cowork`-resolving-to-hostloop record into a
|
|
242
|
+
repo-visible cassette path would freeze THIS machine's MCP server names, agents, and account metadata
|
|
243
|
+
into a committed fixture, so `record` refuses **before the paid spawn**. `--allow-host-inventory-fixture`
|
|
244
|
+
proceeds past that check — and nothing more: the finished recording is still scanned, and a real finding
|
|
245
|
+
still refuses the write and quarantines it. **You do not need to audit the session by hand to pass this
|
|
246
|
+
flag**; that precondition was never decidable by the operator, and the scan is the real gate. Writing a
|
|
247
|
+
recording the scan DID flag is the separate `--allow-host-inventory-findings`.
|
|
246
248
|
|
|
247
249
|
Two details that matter for a re-record loop. The pre-spend check **warns rather than refuses when the
|
|
248
250
|
cassette already exists**, deliberately: refusing there would fire on every `--rerecord-stale` pass and
|
|
@@ -340,7 +342,7 @@ jobs:
|
|
|
340
342
|
with: { node-version: '24' }
|
|
341
343
|
- uses: actions/setup-python@v5
|
|
342
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
343
|
-
- run: npm i -g "cowork-harness@^2.
|
|
345
|
+
- run: npm i -g "cowork-harness@^2.3.0"
|
|
344
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
345
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
346
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -369,7 +371,7 @@ jobs:
|
|
|
369
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
370
372
|
fi
|
|
371
373
|
- if: steps.guard.outputs.live == 'true'
|
|
372
|
-
run: npm i -g "cowork-harness@^2.
|
|
374
|
+
run: npm i -g "cowork-harness@^2.3.0"
|
|
373
375
|
- if: steps.guard.outputs.live == 'true'
|
|
374
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
375
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 2.
|
|
3
|
+
Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.3.0`
|
|
4
|
+
(baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -289,10 +289,10 @@ same set live from the schema.
|
|
|
289
289
|
| Assertion | Passes when |
|
|
290
290
|
|---|---|
|
|
291
291
|
| `result: success \| error` | the run ended with that status |
|
|
292
|
-
| `transcript_contains: <str>` | the assistant transcript includes the literal string |
|
|
293
|
-
| `transcript_not_contains: <str>` | it does not |
|
|
294
|
-
| `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose |
|
|
295
|
-
| `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) |
|
|
292
|
+
| `transcript_contains: <str>` | the assistant transcript includes the literal string **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
293
|
+
| `transcript_not_contains: <str>` | it does not **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
294
|
+
| `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
295
|
+
| `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
296
296
|
| `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
|
|
297
297
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
298
298
|
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
|
|
@@ -344,8 +344,9 @@ same set live from the schema.
|
|
|
344
344
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
345
345
|
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
346
346
|
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
347
|
-
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
|
|
348
|
-
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered —
|
|
347
|
+
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
348
|
+
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
349
|
+
| `question_context: {when_question?, matches}` | a regex over **everything the gate showed the user** — question label + every option label + every option **description**. Use it when the sentence you need to prove reached the founder may land in any of those fields: `question_asked` sees only the question text, `question_options` compares only labels, so a phrase delivered in an option's `description` is invisible to both. `when_question` narrows; omitting it searches every gate (NOT ambiguous here — this key asks whether the text was shown, not which gate offered which set). Ask-time payload only, never a `tool_result` (a producer that also writes the phrase to its own gate-state file would otherwise false-green it). Zero gates FAILS ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
349
350
|
| `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
|
|
350
351
|
| `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
|
|
351
352
|
| `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
|
|
@@ -369,8 +370,8 @@ same set live from the schema.
|
|
|
369
370
|
| `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
|
|
370
371
|
| `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
|
|
371
372
|
| `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
|
|
372
|
-
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) |
|
|
373
|
-
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** |
|
|
373
|
+
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
374
|
+
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
374
375
|
|
|
375
376
|
`expect_denied: [host, …]` adds one `egress_denied` per host. Run `cowork-harness assertions --list` for this
|
|
376
377
|
table from the live schema. Example: `artifact_json: { artifact: outputs/cap.json, path: me.run_id, equals: "r1" }`.
|
|
@@ -441,7 +442,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
|
|
|
441
442
|
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
|
|
442
443
|
`allow_stall` are also kept on replay, evaluated as no-op passes.
|
|
443
444
|
|
|
444
|
-
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `questions_count_max`,
|
|
445
|
+
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
|
|
445
446
|
`gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`,
|
|
446
447
|
`path_denied`, `no_path_denied` (the latter three are also `fidelity: hostloop`-only — see the assertion
|
|
447
448
|
table). With `controlOut` present they evaluate; on an old
|
|
@@ -545,7 +546,8 @@ reference omits). Neither list is a strict superset of the other — reach for t
|
|
|
545
546
|
`egress_denied` is dropped. Both now warn loudly. → put egress/live-only checks on a live gate; one
|
|
546
547
|
concern per item; run the linter. (`LIVE_ONLY_KEYS`/`MANIFEST_KEYS` in `src/run/cassette.ts`.)
|
|
547
548
|
|
|
548
|
-
2. **Gate keys need a `controlOut` cassette.** `question_asked`, `
|
|
549
|
+
2. **Gate keys need a `controlOut` cassette.** `question_asked`, `question_options`, `question_context`,
|
|
550
|
+
`questions_count_max`,
|
|
549
551
|
`gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked` only evaluate on
|
|
550
552
|
replay with `controlOut`; on an old cassette they warn and are excluded (not passed).
|
|
551
553
|
`gate_answers_delivered` **fails on
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 2.
|
|
5
|
+
Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -17,7 +17,7 @@ a paid live re-record? Walk this tree — the answer is usually no:
|
|
|
17
17
|
`cowork-harness replay <cassette> --assert-from <scenario.yaml>`. Token-free, no re-record.
|
|
18
18
|
If the recording genuinely lacks the telemetry a key needs (very old cassettes), the key fails
|
|
19
19
|
**loud** as `evidence-unavailable` — that is correct behavior, not a bug; only then re-record.
|
|
20
|
-
2. **Gate keys** (`question_asked`, `question_options`, `questions_count_max`, `gate_answers_delivered`) on a cassette
|
|
20
|
+
2. **Gate keys** (`question_asked`, `question_options`, `question_context`, `questions_count_max`, `gate_answers_delivered`) on a cassette
|
|
21
21
|
**with `controlOut`** (any modern recording) → same token-free `--assert-from` path.
|
|
22
22
|
3. **Gate keys** on a **pre-`controlOut`** cassette → one re-record unlocks gate asserts for that
|
|
23
23
|
cassette permanently.
|
|
@@ -114,6 +114,7 @@ CONTENT_KEYS = {
|
|
|
114
114
|
GATE_KEYS = {
|
|
115
115
|
"question_asked",
|
|
116
116
|
"question_options",
|
|
117
|
+
"question_context",
|
|
117
118
|
"questions_count_max",
|
|
118
119
|
"gate_answers_delivered",
|
|
119
120
|
"gate_answer_count_min",
|
|
@@ -284,6 +285,7 @@ def _load_top_level_keys():
|
|
|
284
285
|
# every valid top-level scenario key (generated from the zod ScenarioObject schema; see _load_top_level_keys)
|
|
285
286
|
TOP_LEVEL_KEYS = _load_top_level_keys()
|
|
286
287
|
REGEX_KEYS = {
|
|
288
|
+
"matches",
|
|
287
289
|
"transcript_matches",
|
|
288
290
|
"transcript_not_matches",
|
|
289
291
|
"when_question",
|
|
@@ -754,7 +756,11 @@ def lint_doc(doc, path, raw_lines):
|
|
|
754
756
|
# YAML 1.1 but are rejected by the loader as strings — see _numeric for the same dialect gap.)
|
|
755
757
|
delivered_values = _assert_values(items, "gate_answers_delivered")
|
|
756
758
|
if any(v is not False for v in delivered_values):
|
|
757
|
-
has_companion =
|
|
759
|
+
has_companion = (
|
|
760
|
+
"question_asked" in assert_keys
|
|
761
|
+
or "question_options" in assert_keys
|
|
762
|
+
or "question_context" in assert_keys
|
|
763
|
+
) or any(
|
|
758
764
|
(n := _numeric(v)) is not None and n >= 1 for v in _assert_values(items, "gate_answer_count_min")
|
|
759
765
|
)
|
|
760
766
|
if not has_companion:
|
|
@@ -837,6 +843,7 @@ def lint_doc(doc, path, raw_lines):
|
|
|
837
843
|
),
|
|
838
844
|
("question_asked", "question_asked" in assert_keys),
|
|
839
845
|
("question_options", "question_options" in assert_keys),
|
|
846
|
+
("question_context", "question_context" in assert_keys),
|
|
840
847
|
("gate_answers_delivered: false", any(v is False for v in _assert_values(items, "gate_answers_delivered"))),
|
|
841
848
|
],
|
|
842
849
|
"a delivered gate records at least one question, so requiring a gate to be present contradicts requiring zero questions",
|
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,242 @@ All notable changes to this project are documented here. The format is based on
|
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
|
|
5
5
|
[Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
|
|
6
6
|
|
|
7
|
+
## [Unreleased]
|
|
8
|
+
|
|
9
|
+
## [2.3.0] — 2026-08-26
|
|
10
|
+
|
|
11
|
+
### Parity
|
|
12
|
+
|
|
13
|
+
- **Baseline `desktop-1.37937.1` (agent `2.1.246`).** `sync` had been refusing to write on three
|
|
14
|
+
`spawn.env` deltas; all three are now classified.
|
|
15
|
+
|
|
16
|
+
**Two new pinned keys.** The Cowork spawn sets `CLAUDE_CODE_PROMPT_CACHE_TTL="1h"` and
|
|
17
|
+
`CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL="5m"` unconditionally — no gate, no session or deployment
|
|
18
|
+
branch — so every first-party session receives them. They are **additive**:
|
|
19
|
+
`ENABLE_PROMPT_CACHING_1H="1"` is still set alongside. They read zero times in agent `2.1.241` and six
|
|
20
|
+
times each in `2.1.246`, so the contract went live one agent release after Desktop began setting it.
|
|
21
|
+
|
|
22
|
+
**`MCP_TOOL_TIMEOUT` is now classified per SITE, not per key.** Its first-party construction is
|
|
23
|
+
unchanged (still resolving to `180000`), but the third-party-only branch gained a second, settings-
|
|
24
|
+
conditional construction of the same key whose value expression the const resolver cannot reach.
|
|
25
|
+
Allowlisting the key — the obvious fix — would have been a silent contract loss: the allowlist is
|
|
26
|
+
checked *before* the pin list, so the key would have vanished from the generated env entirely, and it
|
|
27
|
+
is not a `REQUIRED_SPAWN_KEYS` member, so nothing would have hard-failed. Instead the 3p-only branch is
|
|
28
|
+
located by content and its inner keys are classified by name without resolving their values, which is
|
|
29
|
+
what the branch already meant. A brand-new key there still hard-fails.
|
|
30
|
+
|
|
31
|
+
**`CLAUDE_PREVIEW_CLASSIFIER_FLOOR` is now inert on the shipping agent** — recorded, not changed. Agent
|
|
32
|
+
`2.1.246` renamed the flag it reads to `CLAUDE_CHROME_CLASSIFIER_FLOOR` (and its consumer field
|
|
33
|
+
`previewClassifierFloorEnabled` → `chromeClassifierFloorEnabled`) while Desktop still sets only the old
|
|
34
|
+
name, so the classifier floor now falls through to its GrowthBook default. The key stays pinned: the
|
|
35
|
+
baseline records what the spawn constructs, and a Desktop-side rename must surface as a diff line
|
|
36
|
+
rather than as silence. Nothing in the harness reads it behaviourally. Together with the cache-TTL keys
|
|
37
|
+
above this is the same lesson pointing both ways — Desktop and the agent version the spawn env
|
|
38
|
+
independently, so "Desktop sets X" and "the agent reads X" are separately-dated claims.
|
|
39
|
+
|
|
40
|
+
Also found in the same pass and recorded in [docs/fidelity-gaps.md](./docs/fidelity-gaps.md), with no
|
|
41
|
+
harness surface: the deliberately-unmodeled remote device-tool family gained `device_fs` and
|
|
42
|
+
`device_request_delete_permission` plus a folder-access announce mode — the Desktop half of the six
|
|
43
|
+
`cowork_*` risk categories that had appeared in agent `2.1.237`'s auto-mode rubric.
|
|
44
|
+
|
|
45
|
+
Also verified unchanged against the retained previous asar: the Cowork system prompt (the retained raw
|
|
46
|
+
files for 1.32885.1 through 1.37937.1 are byte-identical on disk), both sub-agent appends, `tools[]`,
|
|
47
|
+
the `canUseTool` chain, the mount modes, the egress contract, and the VM rootfs image.
|
|
48
|
+
|
|
49
|
+
- **All three committed example cassettes re-recorded against this baseline** — `example-pdf-skill`
|
|
50
|
+
(`container`), `example-multiselect-gate` (`protocol`) and `hostloop-computer-links` (`hostloop`).
|
|
51
|
+
The container one shows no behavioural change (transcript wording only). `verify-cassettes` is clean,
|
|
52
|
+
and the recorded MCP inventory is the Cowork lane's own servers (`cowork`, `plugins`, `skills`,
|
|
53
|
+
`workspace`) — no host inventory reached the fixtures, checked independently of the built-in scan
|
|
54
|
+
because the 2026-08-04 leak hid from a `mcp__` grep and surfaced only through NAME fields.
|
|
55
|
+
|
|
56
|
+
- **Live end-to-end pass re-run against this baseline.** `npm run test:live` — **4 suites, 19 assertions,
|
|
57
|
+
19 green, 0 skipped** — across `protocol`, `container` and `hostloop`, on agent `2.1.246`. Every
|
|
58
|
+
`describe` in that lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated
|
|
59
|
+
case reports as skipped rather than passing vacuously; zero skips means the whole population executed.
|
|
60
|
+
`DESIGN.md`'s scope note is re-stamped accordingly and now records that no baseline is unverified for
|
|
61
|
+
want of a live run. Not claimed: the `boundary-check` sandbox proof and the example-scenario suite were
|
|
62
|
+
part of the previous stamp and were not run here.
|
|
63
|
+
|
|
64
|
+
### Added
|
|
65
|
+
|
|
66
|
+
- **`question_context: {when_question?, matches}` — assert what a gate actually put in front of the user.**
|
|
67
|
+
A regex tested against a gate's founder-visible payload FIELD BY FIELD — the question label, every option
|
|
68
|
+
**label**, and every option **description**, each as a separate string. Matching per field rather than over
|
|
69
|
+
one joined blob is deliberate and load-bearing: a pattern cannot straddle two fields, so
|
|
70
|
+
`invoicing[\s\S]*Audit logging` will not match by stitching one option's description to the next option's
|
|
71
|
+
label — a "sentence" nobody was shown. The neighbouring transcript keys' docs teach `[\s\S]` for spanning
|
|
72
|
+
turns, so that is the habit an author brings here. `matches` is required NON-EMPTY: an empty pattern
|
|
73
|
+
compiles to `//i` and would green any run that fired a gate at all. `question_asked` matches question text only and `question_options`
|
|
74
|
+
compares labels only, so a sentence the model delivered inside an option's `description` was invisible to
|
|
75
|
+
every assertion key — a false-negative generator for any skill that puts context there, which the tool's
|
|
76
|
+
own schema invites. Measured on a consumer's paid run: a producer-authored sentence arrived verbatim in
|
|
77
|
+
the question, reworded inside it, and relocated into the proceed option's `description` across three runs
|
|
78
|
+
of one scenario; the third redded a lane on a run where the founder had in fact been told.
|
|
79
|
+
|
|
80
|
+
Evidence is the **ask-time** `AskUserQuestion` payload, never a `tool_result` — a skill's producer
|
|
81
|
+
typically also writes the same sentence into its own gate-state file, so `tool_result_matches` on that
|
|
82
|
+
phrase grades true whether or not the model ever surfaced it. Unlike `question_options`, omitting
|
|
83
|
+
`when_question` on a multi-gate run is **not** ambiguous: this key asks whether the text was shown at all.
|
|
84
|
+
Zero gates recorded fails; unreadable gate evidence fails evidence-unavailable, never vacuously.
|
|
85
|
+
|
|
86
|
+
### Changed
|
|
87
|
+
|
|
88
|
+
- **`trace --view questions` now renders each gate's offered options — labels AND `description`s** — under
|
|
89
|
+
an `offered:` block, sub-question-labelled on a bundled gate. It previously printed the question label
|
|
90
|
+
alone, so the option `description` a skill routinely puts the deciding sentence in was reachable only by
|
|
91
|
+
hand-reading `events.jsonl`; a reader who found nothing in the view could reasonably conclude the text was
|
|
92
|
+
never delivered. The payload was always recorded — this was purely a rendering gap, and the same run dir
|
|
93
|
+
answers the question either way. The row (`--output-format json`) gains `subQuestions[]` carrying the
|
|
94
|
+
untruncated ask-time options; the text view caps each description at 240 chars, so nothing is lost, only
|
|
95
|
+
wrapped. This pairs with `question_context`: the view is how you *find* the text, that key is how you
|
|
96
|
+
*gate* on it.
|
|
97
|
+
|
|
98
|
+
- **The `tool_use` blindness of `transcript_contains`/`_not_contains`/`_matches`/`_not_matches` and
|
|
99
|
+
`computer_links_resolve`/`_if_present` is now documented and guarded.** `semantic_matches` has carried a
|
|
100
|
+
⚠️ spelling out that its corpus excludes every `tool_use` — "a rubric claim about whether a tool was
|
|
101
|
+
called is unassertable" — while the six keys with the identical blindness said only "the assistant
|
|
102
|
+
transcript". A consumer wrote a `transcript_matches` against text living in a gate question; it could not
|
|
103
|
+
have matched at any phrasing, and the recording fail-closed after the spend. The caveat is now a property
|
|
104
|
+
of an enumerable set (`TOOL_USE_BLIND_KEYS`) enforced across every surface that documents a key — the
|
|
105
|
+
docs tables, the zod `.describe()` behind `assertions --list` and the generated JSON schema, and the
|
|
106
|
+
packaged skill reference — so a newly-added blind key cannot ship without it.
|
|
107
|
+
- **`question_asked`/`question_options`/`question_context` now warn that they match model-authored text.**
|
|
108
|
+
Gate question text and option labels are composed by the model and reworded run to run. The `choose:`
|
|
109
|
+
side already documented this (stable leading anchor, 1-based index); the assert side documented it
|
|
110
|
+
nowhere. Guarded by `MODEL_AUTHORED_TEXT_KEYS`.
|
|
111
|
+
- **A bad regex in a NESTED assertion field is now caught at load, not after the paid spawn.** The pre-compile
|
|
112
|
+
pass reached only top-level string keys, so every regex one level down — `artifact_text.matches`,
|
|
113
|
+
`artifact_text.not_matches`, `path_denied.path_matches`, `skill_tool_used.skill`/`.tool`,
|
|
114
|
+
`subagent_dispatch_healthy.type`, `subagent_output_contains.match`, `task_status.match`,
|
|
115
|
+
`question_options.when_question`, and the new `question_context.*` — was first compiled inside the
|
|
116
|
+
evaluator. All eleven are now validated at load, and `test/nested-regex-leaves.test.ts` reads `assert.ts`
|
|
117
|
+
and fails if the evaluator compiles a nested leaf the load-time table does not carry, so the gap cannot
|
|
118
|
+
reopen silently. `lint`'s double-quoted-regex warning also now covers nested `matches:` leaves.
|
|
119
|
+
- **`diff` with a single positional now names the missing operand** instead of printing bare usage. There is
|
|
120
|
+
deliberately no one-argument form: `diff` is polymorphic over baselines, run dirs and cassettes, so it
|
|
121
|
+
would need type dispatch plus a defined source for "the committed version".
|
|
122
|
+
- **The two cost keys are cross-referenced.** `RunResult.cost.usd` is one invocation's SDK
|
|
123
|
+
`total_cost_usd`; the critique report's `costUsd.totalUsd` aggregates the task turn, the reflection turn
|
|
124
|
+
and both evaluator passes. Reading the wrong one returns `undefined` rather than erroring, which reads as
|
|
125
|
+
"no cost recorded". Kept as two shapes on purpose — collapsing them would destroy the per-phase split.
|
|
126
|
+
|
|
127
|
+
### Fixed
|
|
128
|
+
- **Two spawn-contract sentinels were weaker than their names implied; both now bite.**
|
|
129
|
+
|
|
130
|
+
`checkSpawnContractFacts` pinned the `allowedTools[]` built-in head and the built-in→`mcp__` boundary
|
|
131
|
+
but nothing between the boundary and the closing bracket — so `mcp__plugins__search_connectors` was
|
|
132
|
+
added to the array and both checks stayed green. The `mcp__` membership is now pinned as a set, and the
|
|
133
|
+
flag names the added and removed entries instead of reporting that something "moved". (That tool is
|
|
134
|
+
declared only on the third-party deployment, so the first-party inventory the harness serves is
|
|
135
|
+
unaffected — but the addition should not have been invisible.)
|
|
136
|
+
|
|
137
|
+
`checkMountModeFacts` asserted each read-only mount with a single `regex.test`, while `uploads`,
|
|
138
|
+
`.claude/skills` and `.projects/<uuid>` are each built at **two** sites — the VM-loop mount-set builder
|
|
139
|
+
and host-loop `computeBashMounts`. Either site satisfied the check, so a one-lane `ro`→`rw` flip — a
|
|
140
|
+
containment change on exactly one execution tier — passed green. It now compares the site count to the
|
|
141
|
+
count carrying `mode:"ro"` and names the lane count in the flag. Its fifth fact — the delete-deny
|
|
142
|
+
resolver `?"rwd":"rw"` — had the same shape and the same two-lane reality (it went 1 site to 2 in this
|
|
143
|
+
release) and now guards a floor on that count: a lane losing the resolver flags, a lane gaining one
|
|
144
|
+
does not.
|
|
145
|
+
|
|
146
|
+
- **All eight asar sentinels now carry a committed mutation case.** They are green on the previous
|
|
147
|
+
release's asar too, so a green proves nothing on its own unless the checker is known to bite; five of
|
|
148
|
+
them (`checkCodeTripwires`, `checkWebFetchFacts`, `checkEgressContractFacts`, `checkSyspromptMapFacts`,
|
|
149
|
+
`checkNormalizationSanity`) had no such case. Each new one changes an **inner** character of its
|
|
150
|
+
anchor, because a suffix rename still satisfies a substring regex — a mutation that cannot fail is not
|
|
151
|
+
evidence — and asserts that the mutation actually applied.
|
|
152
|
+
|
|
153
|
+
- **`record --dry-run <dir/>` no longer announces a WARN as a refusal.** The batch arm labelled every
|
|
154
|
+
advisory note `⚠ would-refuse (advisory)`, but `cassettePortabilityPreflight` can only ever return
|
|
155
|
+
`ok`/`warn` — it has no refuse path at all — and `hostInventoryPreflight` returns `warn` whenever the
|
|
156
|
+
target cassette already exists, which is every re-record corpus sitting at the default path. So the
|
|
157
|
+
preview told operators the real `record` would refuse runs it would in fact accept, on the one arm whose
|
|
158
|
+
design principle is that a guess must not gate. The label now follows the verdict kind
|
|
159
|
+
(`⚠ would-warn (advisory)`), the "ADVISORY, not this run's verdict" footer is unchanged, and exit codes
|
|
160
|
+
were never affected either way.
|
|
161
|
+
|
|
162
|
+
- **The 3p-branch rule refuses to blank W1.** W1 is the window every modeled first-party key is derived
|
|
163
|
+
from, so a branch marker appearing there hard-fails instead of blanking; deleting real pinned keys from
|
|
164
|
+
the derived env with nothing failing is the worse of the two outcomes.
|
|
165
|
+
|
|
166
|
+
- **`record --dry-run` now runs every pre-spend refusal the real record runs, and one refusal moved from
|
|
167
|
+
after the paid run to before it.** The rehearsal re-implemented the checks by hand, so it drifted:
|
|
168
|
+
`hostInventoryPreflight` shipped 2026-08-04 and the commit three days later titled *"make `--dry-run`
|
|
169
|
+
refuse what the real record refuses"* swept in the two checks returning `string | undefined` and missed
|
|
170
|
+
the one returning a `{kind}` verdict — for 19 days an operator could not discover that refusal without
|
|
171
|
+
spending. Separately, the slug-collision refusal (*"refusing to overwrite … it belongs to scenario X"*)
|
|
172
|
+
sat **after** `executeScenario`: you paid for the run and were then refused, though it is a pure function
|
|
173
|
+
of path + `--force` + the existing cassette's name. Both now run in one shared pre-spend block.
|
|
174
|
+
|
|
175
|
+
Refusals are also uniform now. `promptPolicyRejection` threw while `hostInventoryPreflight` called
|
|
176
|
+
`fail()`, which `process.exit`s — and the dir-batch loop catches a throw per item but cannot catch an
|
|
177
|
+
exit, so a host-inventory refusal mid-batch abandoned concurrent runs already paid for.
|
|
178
|
+
|
|
179
|
+
Under `--output-format json` a directory batch carries those advisory verdicts in a `notes[]` array, kept
|
|
180
|
+
separate from `refusals[]` so automation cannot read a guess as a binding verdict.
|
|
181
|
+
|
|
182
|
+
**On a directory, the path-dependent verdicts are ADVISORY (`⚠ would-refuse`) and do not affect the exit
|
|
183
|
+
code**, because a directory target takes no `--out`: the preview would have to guess the destination, and
|
|
184
|
+
a guess must not gate. Measured on a real consumer, gating on that guess would have refused 26 of 27
|
|
185
|
+
scenarios the real record accepts. Dry-run a single scenario file with the real flags for a binding
|
|
186
|
+
answer. `--quiet` suppresses the notes; it never suppresses a refusal.
|
|
187
|
+
|
|
188
|
+
Also new in the batch preview: the duplicate-cassette-path refusal the real batch already had.
|
|
189
|
+
|
|
190
|
+
### Upgrade impact
|
|
191
|
+
|
|
192
|
+
- **`--allow-host-inventory-fixture` no longer waives a MEASURED host-inventory finding.** It was one flag
|
|
193
|
+
doing two jobs: bypassing the pre-flight refusal (a precondition the operator cannot check — "use only
|
|
194
|
+
when the session has no personal MCP servers or plugins") *and* downgrading the write-time scan's refusal
|
|
195
|
+
to a warning. So an operator who passed it to get past the undecidable precondition also switched off the
|
|
196
|
+
scan that would have caught a real leak. It is now the pre-flight bypass only; the finished recording is
|
|
197
|
+
still scanned, and a finding still refuses the write and quarantines the recording. Writing a flagged
|
|
198
|
+
recording needs the new, narrower `--allow-host-inventory-findings`. **A batch recorder that passes the
|
|
199
|
+
old flag on a new-fixture path will now abort on a genuine finding where it previously warned and wrote.**
|
|
200
|
+
|
|
201
|
+
A `--dry-run` that reports what *would* be captured was considered and declined: the inventory does not
|
|
202
|
+
exist until the agent has run, so a preview would either re-implement the scanner against a hypothetical
|
|
203
|
+
(a second oracle free to disagree with the real one) or require the spend it was meant to avoid. Reasoning
|
|
204
|
+
is in [docs/cassette.md](./docs/cassette.md).
|
|
205
|
+
|
|
206
|
+
This does not break a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)):
|
|
207
|
+
no command or flag is removed and no exit-code meaning changes — one flag's consent narrows, and the
|
|
208
|
+
capability it shed is reachable through the new one. So it ships in a minor.
|
|
209
|
+
|
|
210
|
+
### Documentation
|
|
211
|
+
|
|
212
|
+
Five gaps found by asking a consumer which harness properties actually changed an outcome on a real
|
|
213
|
+
working day, then checking whether the docs said so. Four of the five were documented only as features,
|
|
214
|
+
never as the failure they prevent — the sentence a reader needs to recognise their own situation.
|
|
215
|
+
|
|
216
|
+
- **The blocking gate is now stated as a blocker.** `AskUserQuestion` *blocks*: it is a question to a
|
|
217
|
+
human and `claude -p` has no human, so a gated skill under a plain CLI run stalls or never reaches the
|
|
218
|
+
code behind the gate. That made half the skill untestable, not merely awkward to test — the README had
|
|
219
|
+
only "untestable headless unless something answers it", buried as the last of two afterthought bullets.
|
|
220
|
+
- **"Assert on the run, not the output"** — a new README section naming the class of claim the harness
|
|
221
|
+
exists for (`subagent_tool_absent`, `dispatch_count_max`, `no_delete_in_outputs`, `subagent_file_write`)
|
|
222
|
+
and why no output diff can reach it: a correct run and one that quietly handed a restricted sub-agent
|
|
223
|
+
shell access produce byte-identical files. `subagent_tool_absent` did not appear in the README at all.
|
|
224
|
+
- **The raw-`events.jsonl` escape hatch is documented**, with verified `jq` recipes, in
|
|
225
|
+
[docs/debugging.md](./docs/debugging.md). Every documented route into the event log went through
|
|
226
|
+
`trace`, whose views are a digest — so the run dir's *evidence* surface is wider than any view's
|
|
227
|
+
*observation* surface, and wider still than the assertion catalog. Concluding "it never happened" from
|
|
228
|
+
a view that doesn't render the field is a false negative, and the doc now says so and shows the read.
|
|
229
|
+
- **Sub-agent delivery has a route.** The tier-qualified outputs contract — the reason a hand-off path
|
|
230
|
+
that works on one loop lands in sandbox scratch on the other — was correctly documented in
|
|
231
|
+
[docs/subagents.md](./docs/subagents.md) but filed under "read on demand", reachable only by someone
|
|
232
|
+
who already knew the answer. It now has a "Common tasks" row keyed to the symptom.
|
|
233
|
+
- **Replay's cost claim carries a number.** "Zero spend" is the price; the wall-clock is what makes an
|
|
234
|
+
always-on per-PR gate obviously affordable, and no doc stated it (well under a second per cassette).
|
|
235
|
+
|
|
236
|
+
- **The project's provenance is stated on every surface the framing travels to.** `README.md` (which
|
|
237
|
+
renders on the npm page), `llms.txt`, both `.claude-plugin/marketplace.json` descriptions and `SKILL.md`
|
|
238
|
+
now say the same thing: an independent project, not affiliated with, endorsed by, or supported by
|
|
239
|
+
Anthropic; it bundles no Anthropic code; it is not Cowork. `SKILL.md` additionally tells the agent to say
|
|
240
|
+
so when a user asks what it is. Nothing enforces the four staying in step, so an edit can still drop it
|
|
241
|
+
from one of them unnoticed.
|
|
242
|
+
|
|
7
243
|
## [2.2.0] — 2026-08-25
|
|
8
244
|
|
|
9
245
|
### Upgrade impact
|