cowork-harness 1.0.6 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +14 -8
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  3. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +35 -2
  4. package/.claude/skills/cowork-harness/references/scenario-schema.md +9 -6
  5. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -0
  6. package/.claude/skills/cowork-harness/scripts/scenario.py +48 -0
  7. package/CHANGELOG.md +203 -0
  8. package/README.md +15 -12
  9. package/SPEC.md +42 -4
  10. package/dist/assert.js +214 -3
  11. package/dist/baseline.js +36 -8
  12. package/dist/cli.js +46 -9
  13. package/dist/decide/decider.js +9 -1
  14. package/dist/hostloop/canusetool-gate.js +86 -2
  15. package/dist/hostloop/provenance.js +7 -2
  16. package/dist/hostloop/workspace-handler.js +4 -2
  17. package/dist/run/analyze-artifact-runtime.js +547 -0
  18. package/dist/run/analyze-artifact.js +1190 -0
  19. package/dist/run/analyze-skill.js +99 -21
  20. package/dist/run/artifacts.js +11 -4
  21. package/dist/run/cassette.js +37 -6
  22. package/dist/run/chat-result.js +4 -2
  23. package/dist/run/doctor.js +98 -18
  24. package/dist/run/execute.js +52 -10
  25. package/dist/run/latest-run.js +43 -0
  26. package/dist/run/pre-run-manifest.js +3 -2
  27. package/dist/run/renderer.js +4 -0
  28. package/dist/run/status-target.js +28 -0
  29. package/dist/run/trace-view.js +1 -1
  30. package/dist/runtime/hostloop.js +23 -4
  31. package/dist/runtime/microvm.js +44 -7
  32. package/dist/scan.js +17 -14
  33. package/dist/sync/cowork-sync.js +10 -1
  34. package/dist/types.js +15 -1
  35. package/docs/cassette.md +18 -9
  36. package/docs/fidelity-gaps.md +47 -0
  37. package/docs/maintenance.md +2 -0
  38. package/docs/run-status.md +5 -0
  39. package/docs/scenario.md +12 -5
  40. package/docs/subagents.md +68 -0
  41. package/examples/replays/README.md +1 -1
  42. package/package.json +4 -1
  43. package/python/test_scenario_lint.py +88 -0
  44. package/schema/doctor.json +81 -0
  45. package/schema/scenario.schema.json +16 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.0.6
7
- tracks-harness: cowork-harness 1.0.6 (baseline desktop-1.21459.0)
6
+ version: 1.2.0
7
+ tracks-harness: cowork-harness 1.2.0 (baseline desktop-1.21459.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0.6` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.2.0` (baseline
26
26
  > `desktop-1.21459.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0.6**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.0.6" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.0.6"`. **Pin `@>=1.0.6`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.2.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.2.0"`. **Pin `@>=1.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
- What the ≥ 1.0.6 floor gates, by release:
44
+ What the ≥ 1.2.0 floor gates, by release:
45
45
 
46
46
  - **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
47
47
  - **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
@@ -52,6 +52,8 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
52
52
  - **0.32.0:** `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict`.
53
53
  - **0.33.0:** the `redacted` marker on display-omitted reasoning — `subagents[].reasoning` and the top-level `thinking[]` now carry `{text:"", redacted:true}` when the model returns a signed-but-empty thought (so "reasoned, text omitted" is distinct from "no thought"), plus the fenced `debug.thinking_display` escape hatch.
54
54
  - **1.0.0:** first stable release — the SPEC §12 compatibility contract takes effect (covered CLI/schema/env/Action surfaces are now stable; breaking changes need a major bump). No new author-facing command; the floor simply tracks the 1.0 release.
55
+ - **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
56
+ - **1.2.0:** three new assertion keys — `no_lost_write_back: true` (the write-back detector above wired as a per-scenario gate over the run's authored files, live-only) and the regex siblings `tool_result_matches`/`tool_result_not_matches` (case-insensitive per-result, for an error-signature *family* a literal substring can't express). **`microvm` outputs are now observable** — its session tree is snapshotted from the VM into the run dir, so `file_exists`/`artifact_json`/`user_visible_artifact`/`no_unexpected_files`/`input_unmodified`/`no_lost_write_back`/`semantic_matches` all work there, no longer `container`/`hostloop`-only. `status <dir>` also resolves the newest session under a `--run-dir` root; the completion footer prints a `→ result: …/result.json` pointer; the `on_unanswered=fail` error also points to `on_unanswered: llm`. `analyze-skill` hardening: a phantom `<script>` prose block no longer sinks a real verdict, a delete/remove flow claiming success classifies as lost (error) not just suspect, and write-back detection widened (optional-call `?.` spellings, member-spelled/aliased `fetch`/`sendBeacon`, axios instance/config forms).
55
57
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
56
58
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
57
59
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
@@ -220,7 +222,8 @@ them by what you're trying to prove:
220
222
  | a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
221
223
  | a tool ran **inside** a skill's scope | `skill_tool_used: {skill, tool}` |
222
224
  | a sub-agent did the work | `subagent_output_contains: {contains}`, `subagent_dispatched: <regex>`, `dispatch_count_max: <N>` |
223
- | a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run; not microvm) |
225
+ | a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
226
+ | no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
224
227
  | a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
225
228
  | a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
226
229
  | every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
@@ -277,7 +280,8 @@ quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `
277
280
  (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
278
281
  baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
279
282
  `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
280
- unverifiable), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
283
+ unverifiable), a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) off the
284
+ `container` tier (ERROR on `protocol`/`microvm`/`hostloop`, WARN on `cowork`), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
281
285
  and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
282
286
  (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
283
287
  emits a scenario `lint` would reject.
@@ -372,7 +376,9 @@ harness writes/updates throughout the run's lifecycle (including a crash-safety
372
376
  error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
373
377
  either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
374
378
  checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
375
- run starts, so capture stderr to get the exact directory. `--follow` fails loud on a timeout/staleness
379
+ run starts, so capture stderr to get the exact directory — but `<dir>` also accepts the run-dir root
380
+ passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
381
+ newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
376
382
  rather than hanging forever. (Fuller recipe in `docs/run-status.md` — repo-only, not in the installed
377
383
  payload; `cowork-harness status --help` has the flags.)
378
384
 
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.0.6` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.2.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.0.6"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.2.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.0.6"
60
+ - run: npm i -g "cowork-harness@>=1.2.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.0.6"
200
+ - run: npm i -g "cowork-harness@>=1.2.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.0.6"
229
+ run: npm i -g "cowork-harness@>=1.2.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.0.6` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.2.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -130,7 +130,10 @@ hands `fn` exactly this dict.
130
130
  ### Determinism contract
131
131
 
132
132
  - `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
133
- `--answer`/`choose` to add. Correct, but flaky for skills whose gates appear stochastically.
133
+ `--answer`/`choose` to add, and also suggests `on_unanswered: llm` (in the scenario YAML) as a
134
+ secondary escape valve for a gate whose wording drifts run-to-run — a regex chases a moving target,
135
+ at the cost of non-determinism (one model call per gate). Correct, but flaky for skills whose gates
136
+ appear stochastically.
134
137
  - `first` — picks option 1 and warns loudly. **Flagged `nonDeterministic`** — not a deterministic
135
138
  substitute for scripted answers. (For a web_fetch approval gate it abstains → fail-closed.)
136
139
  - `prompt` — asks at the TTY (`skill` only).
@@ -171,6 +174,36 @@ writes a `result.json` (marked `partial: true`) with the artifacts the agent pro
171
174
  the work isn't discarded — then exits 2. Inspect it with `cowork-harness inspect <run-dir>`. `verify-run`
172
175
  and `scaffold` refuse to treat a partial run's half-finished output as a passing result.
173
176
 
177
+ ## What a green does NOT prove
178
+
179
+ A passing run is evidence for exactly what it checked, not a blanket certificate. Three gaps come
180
+ up often enough to spell out:
181
+
182
+ - **A green `replay` proves "same as when recorded," not "correct today."** `replay` never touches
183
+ a filesystem or network — it re-evaluates assertions from the frozen cassette. A fixed set of
184
+ keys is live-only and **skipped outright** on replay (absent from `assertions[]`, not vacuously
185
+ passed): `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
186
+ `egress_allowed`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back`, and `expect_denied`.
187
+ Everything else that *is* evaluated is checked against the **recording**, not fresh behavior — a
188
+ green replay says the skill produced these events when it was recorded, not that it still does
189
+ (`staleness[]` flags skill/baseline drift as a hint; only a live `run` re-confirms current
190
+ behavior). See `docs/cassette.md` § "Still skipped on replay" and `docs/scenario.md` § "Which
191
+ assertions survive replay."
192
+ - **Container-only assertions can't verify off the `container` tier.** `no_scratchpad_leak` and
193
+ `present_files_called` check the `present_files` delivery path, which is served **only** on
194
+ `container` — not `hostloop`/`microvm`. Asserting them off-container hard-fails at runtime (a red
195
+ run, not a false green), so you won't be fooled if you write the assertion. The quieter trap is a
196
+ scenario that runs at `hostloop`/`microvm`/`protocol` and simply omits these assertions: a green
197
+ run there proves **nothing** about scratchpad-leak safety or present_files delivery, because that
198
+ tier never exercises the delivery path the assertions would check. Use `fidelity: container` for
199
+ present_files/scratchpad-delivery coverage.
200
+ - **The harness doesn't observe rendered-artifact interactions, browser downloads, or human
201
+ clicks.** It runs the agent headless — no webview, no browser, no person clicking "Submit." A
202
+ class of Cowork bug (a client-side write-back to a relative URL that resolves-but-fails against
203
+ Cowork's own origin, or a broken blob-download fallback) is invisible to any live run, however
204
+ faithfully sandboxed, because it only manifests in a rendered DOM a human is driving. See
205
+ `docs/fidelity-gaps.md` § "Browser↔webview↔human-interaction boundary."
206
+
174
207
  ## Relevant environment variables
175
208
 
176
209
  - `COWORK_HARNESS_RUNS_DIR` (or `--run-dir <path>`) — override the default run-output root `~/.cowork-harness/runs` (out of any working tree). flag > env > default.
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0.6`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.2.0`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
@@ -251,13 +251,16 @@ same set live from the schema.
251
251
  | `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
252
252
  | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
253
253
  | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
254
- | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest/microvm case, which fails for a different reason; **microvm cannot capture** (use container/hostloop); replay needs `cassette.preRunPaths` (≥0.24 container/hostloop recordings) — cassettes without it **exclude** the key with a loud warning |
255
- | `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail; **microvm cannot capture**; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
254
+ | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning |
255
+ | `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
256
256
  | `self_heal_ran: <bool>` | a plugin-root self-heal script was (not) invoked |
257
+ | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
257
258
  | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
258
259
  | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
259
260
  | `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
260
261
  | `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
262
+ | `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
263
+ | `tool_result_not_matches: <regex>` | the regex sibling of `tool_result_not_contains` — same fails-loud-on-absent-evidence semantics |
261
264
  | `subagent_tool_used: <glob>` | a sub-agent used a tool matching this glob (same `*`/`?`, anchored, case-sensitive semantics as `tool_called`, including the empty/regex-ish rejection) |
262
265
  | `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
263
266
  | `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
@@ -312,7 +315,7 @@ same set live from the schema.
312
315
  | `egress_allowed: <host>` | the host was allowed through |
313
316
  | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
314
317
  | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
315
- | `semantic_matches: {rubric: [...], min_pass?, judge_model?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose (authored-file evidence is unavailable on the **microvm** tier — no pre-run manifest to diff against, so no files are captured; `container`/`hostloop` do capture them). Beyond the microvm case, when the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
318
+ | `semantic_matches: {rubric: [...], min_pass?, judge_model?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
316
319
  | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
317
320
  | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) |
318
321
  | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** |
@@ -407,8 +410,8 @@ staleness `fingerprint` shows ANY skill/baseline drift, or `replay --fail-on-ski
407
410
  skill-source drift; every replay result also reports it class-tagged in `staleness[]` for a JSON gate.
408
411
 
409
412
  **Egress + other filesystem — still skipped on replay (live-only):** `no_delete_in_outputs`,
410
- `self_heal_ran`, `transcript_no_host_path`, `egress_*` / `expect_denied`, `no_mcp_error`, `max_peak_rss_bytes`.
411
- These run only on a live `run`/`record`.
413
+ `self_heal_ran`, `transcript_no_host_path`, `egress_*` / `expect_denied`, `no_mcp_error`, `max_peak_rss_bytes`,
414
+ `no_lost_write_back`. These run only on a live `run`/`record`.
412
415
 
413
416
  **Mixed assertions on replay:** before evaluating, `replay` strips each assertion to its replay-checkable
414
417
  keys and drops any left empty. So `{result, egress_denied}` evaluates on replay as `{result}` alone — its
@@ -27,6 +27,7 @@
27
27
  "max_turns",
28
28
  "no_delete_in_outputs",
29
29
  "no_hook_blocked",
30
+ "no_lost_write_back",
30
31
  "no_mcp_error",
31
32
  "no_path_denied",
32
33
  "no_scratchpad_leak",
@@ -60,7 +61,9 @@
60
61
  "tool_no_error_if_called",
61
62
  "tool_not_called",
62
63
  "tool_result_contains",
64
+ "tool_result_matches",
63
65
  "tool_result_not_contains",
66
+ "tool_result_not_matches",
64
67
  "transcript_contains",
65
68
  "transcript_matches",
66
69
  "transcript_no_host_path",
@@ -17,6 +17,8 @@ Two subcommands:
17
17
  lint flags (see references/scenario-schema.md for the why of each):
18
18
  E egress assertion on `fidelity: protocol` (the harness rejects this run)
19
19
  E `transcript_no_host_path` on hostloop/protocol (fails BY DESIGN at those tiers)
20
+ E `no_scratchpad_leak`/`present_files_called` off (container-only: served only at fidelity:
21
+ protocol/microvm/hostloop container; off-container = cannot-verify)
20
22
  E `requires_capabilities` on `fidelity: protocol` (probe can't run → hard-fails
21
23
  unless allow_missing_capability)
22
24
  E on_unanswered: agent / invalid value (schema rejects `agent`)
@@ -24,6 +26,8 @@ lint flags (see references/scenario-schema.md for the why of each):
24
26
  E `assertions:` instead of `assert:` (block ignored → every check no-ops)
25
27
  W `transcript_no_host_path` on `fidelity: cowork` (tier resolves per baseline gate —
26
28
  incompatible if it lands hostloop)
29
+ W `no_scratchpad_leak`/`present_files_called` on (tier resolves per baseline gate — cannot-verify
30
+ `fidelity: cowork` if it resolves off-container)
27
31
  W no content assertion → no-op on a replay gate (every assertion is fs/egress)
28
32
  W mixed-class assert item → fs/egress half dropped on replay
29
33
  W unknown top-level / assertion key (typo or hallucinated schema)
@@ -58,6 +62,8 @@ CONTENT_KEYS = {
58
62
  "transcript_not_matches",
59
63
  "tool_result_contains",
60
64
  "tool_result_not_contains",
65
+ "tool_result_matches",
66
+ "tool_result_not_matches",
61
67
  "tool_called",
62
68
  "tool_not_called",
63
69
  "subagent_tool_used",
@@ -125,8 +131,13 @@ LIVE_ONLY_KEYS = {
125
131
  "no_mcp_error",
126
132
  "max_peak_rss_bytes",
127
133
  "semantic_matches",
134
+ "no_lost_write_back",
128
135
  }
129
136
  EGRESS_KEYS = {"egress_denied", "egress_allowed"}
137
+ # container-only: served only at fidelity: container (present_files / the scratchpad promotion path
138
+ # it depends on). Off-container these report cannot-verify, not a meaningful pass/fail — same tier-fidelity
139
+ # class as transcript_no_host_path below, just the opposite direction (container-only vs container-hostile).
140
+ CONTAINER_ONLY_KEYS = {"no_scratchpad_leak", "present_files_called"}
130
141
  # verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
131
142
  VERDICT_MODIFIER_KEYS = {
132
143
  "allow_permissive_auto_allow",
@@ -221,6 +232,8 @@ REGEX_KEYS = {
221
232
  "subagent_dispatched",
222
233
  "question_asked",
223
234
  "hook_blocked",
235
+ "tool_result_matches",
236
+ "tool_result_not_matches",
224
237
  }
225
238
  VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
226
239
  VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
@@ -388,6 +401,41 @@ def lint_doc(doc, path, raw_lines):
388
401
  )
389
402
  )
390
403
 
404
+ # E/W: CONTAINER_ONLY_KEYS (no_scratchpad_leak / present_files_called) are served only at
405
+ # fidelity: container — off-container they report cannot-verify, not a meaningful check. Mirrors
406
+ # the transcript_no_host_path tier check above, just the opposite direction: ERROR on the tiers
407
+ # where the runtime deterministically can't serve it, WARN on cowork (baseline-gate-resolution
408
+ # dependent — the linter stays offline, so the message names the dependency instead of resolving it).
409
+ container_only_present = sorted(assert_keys & CONTAINER_ONLY_KEYS)
410
+ if container_only_present:
411
+ if fidelity in ("protocol", "microvm", "hostloop"):
412
+ findings.append(
413
+ Finding(
414
+ "ERROR",
415
+ "container-only-key-off-container",
416
+ f"{container_only_present} on `fidelity: {fidelity}` — container-only (present_files "
417
+ "is served only there); off-container it reports cannot-verify. Use fidelity: "
418
+ "container for present_files/scratchpad delivery.",
419
+ "Use fidelity: container (or drop the assertion for this tier).",
420
+ path,
421
+ )
422
+ )
423
+ elif fidelity == "cowork":
424
+ findings.append(
425
+ Finding(
426
+ "WARN",
427
+ "container-only-key-off-container",
428
+ f"{container_only_present} on `fidelity: cowork` — container-only (present_files "
429
+ "is served only there); off-container it reports cannot-verify. Use fidelity: "
430
+ f"container for present_files/scratchpad delivery. (The tier resolves per the "
431
+ f"baseline's host-loop gate ({HOST_LOOP_GATE_ID}); if it resolves off-container this "
432
+ "assertion cannot verify anything.)",
433
+ "Pin fidelity: container if the assertion is load-bearing; keep cowork only if "
434
+ "you accept the gate-resolution dependency.",
435
+ path,
436
+ )
437
+ )
438
+
391
439
  # E: requires_capabilities on protocol — the capability probe cannot run at protocol tier
392
440
  # (clause b of the requires_capabilities contract), so the run HARD-FAILS unless an assert item
393
441
  # opts out via allow_missing_capability: true. Offline-detectable fails-by-design, same class as
package/CHANGELOG.md CHANGED
@@ -6,6 +6,209 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.2.0] — 2026-07-18
10
+
11
+ ### Added
12
+
13
+ - **`no_lost_write_back: true` scenario assertion** — gate a scenario on "the agent didn't emit an
14
+ interactive artifact whose Submit is lost under Cowork". It runs the shipped static Tier A analyzer
15
+ (`analyze-artifact`, deterministic, no headless DOM) over the files the run authored — diffed against the
16
+ pre-run manifest — so a lost relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back becomes a
17
+ per-scenario verdict, not just an out-of-band `analyze-skill` scan. A lost write-back on an **added**
18
+ agent-authored source (`outputs/` or the scratchpad) fails; a **pre-existing** file the skill only modified
19
+ on a read-write connected mount is advisory (not the skill's to own); `-suspect` findings surface but pass.
20
+ Honest evidence-unavailable semantics: could-not-verify (fail-closed, never a silent clean) on a `--resume`
21
+ scratchpad walk or an authored candidate that couldn't be analyzed. Runs on every live sandbox tier
22
+ including **microvm** (its outputs are snapshotted from the VM into the run dir — see the #52 entry below).
23
+ **Live/verify-run only** — skipped-loud on replay (a cassette embedding the key never
24
+ hard-fails its replay); `verify-run` recomputes the authored set from the kept work dir. Only `true` is
25
+ valid (omit to skip).
26
+ - **`tool_result_matches` / `tool_result_not_matches` scenario assertion keys** — the case-insensitive regex
27
+ siblings of `tool_result_contains`/`tool_result_not_contains`, evaluated per captured tool result (subject
28
+ to the same 10 KB per-result assertText cap). Useful for catching an error-signature *family* (e.g. a
29
+ script's non-zero exit swallowed by its wrapper, but the message still printed) that a literal substring
30
+ match can't express. Same evidence-unavailable wording as the `_contains` pair: a bad regex fails the
31
+ assertion with a compile error, and a no-match against a display-truncated result is reported as
32
+ could-not-verify rather than a silent pass/fail.
33
+
34
+ - **A folder-grant refusal for `request_cowork_directory`, ported from Desktop 1.22209.0.** Cowork now
35
+ refuses (pre-prompt) a mid-session folder grant that targets a security-sensitive home-directory path —
36
+ `.ssh`, `.aws`, `.gnupg`, `.kube`, `.docker`, `.claude`, `.config/{gcloud,gh,powershell}`, the darwin
37
+ `Library/{Keychains,LaunchAgents,LaunchDaemons,Application Support,Cookies}` paths, or a protected shell
38
+ dotfile (`.zshrc`, `.netrc`, etc.) — either directly, as a descendant, or as an ancestor whose grant would
39
+ incidentally expose one (e.g. requesting the home directory itself). The harness's `hostloop` canUseTool
40
+ gate ports this byte-faithfully, denying with Desktop's own message. **Currently dead code in a stock
41
+ run**: no built-in workspace/cowork server registers `request_cowork_directory` yet (a pre-existing,
42
+ separately tracked gap), so this only fires for a scenario that supplies its own `mcp_config` registering
43
+ a matching tool name. Ported ahead of that gap closing so the refusal semantics are ready the moment it
44
+ does. Two GrowthBook feature-gate ids Desktop 1.22209.0 introduced for a related "auto mode always-allow"
45
+ tool-approval feature are pinned as drift sentinels (`sync`'s `PINNED_GATES`) without being behaviorally
46
+ modeled — this harness has no persistent per-tool permission concept to model them against.
47
+
48
+ ### Changed
49
+
50
+ - **`cowork-harness status <run-dir>` now resolves the newest session under a run-dir root.** Previously
51
+ `status` only worked against the exact per-session out-dir printed at run start; pointing it at the root
52
+ passed to `run --run-dir` failed with "no status.json". It now scans up to two levels under a directory
53
+ lacking its own `status.json` for the newest session that has one, and reads that instead.
54
+ - **`doctor`'s two staged-agent checks are now titled distinctly** — "Staged agent binary (VM/container
55
+ ELF)" vs. "Staged native agent binary (hostloop)" — so a failure on either is attributable to the right
56
+ one instead of reading as an ambiguous duplicate.
57
+ - **The run-completion footer now prints a `→ result: <run-dir>/result.json` pointer**, on both success and
58
+ failure, since the run directory is always kept on disk. Suppressed on the replay lane, which never
59
+ writes a `result.json`.
60
+ - **The `on_unanswered=fail` unscripted-gate error now also mentions `on_unanswered: llm` as a secondary
61
+ escape valve.** Previously it suggested only `--answer "<regex>=<choice>"`, which is the right primary
62
+ fix but the wrong tool for a gate whose wording drifts run-to-run — a regex chases a moving target. The
63
+ added line explicitly says "in the scenario YAML" (`--on-unanswered llm` is rejected on the CLI in favor
64
+ of `--decider-llm`) and notes the tradeoff (non-deterministic, one model call per gate) so it doesn't
65
+ read as unconditionally preferable to fixing the script.
66
+
67
+ ### Fixed
68
+
69
+ - **Silent false-green on a missing workspace root (`#52`).** When the workspace root (`outDir/work/session/mnt`)
70
+ couldn't be walked — the canonical case is a **microvm** run, whose outputs stage into the VM work tree, not
71
+ into the run dir — `RunResult.workspaceFiles`/`artifacts` persisted as `[]`, **indistinguishable from a run
72
+ that genuinely wrote nothing.** A consumer reading `result.json` (e.g. skill-creator-plus) saw "zero
73
+ artifacts, clean." They now persist as **`undefined` (unavailable)** — the same convention replay already
74
+ uses for "no live filesystem to scan" — with a loud `::warning::` naming the reason. Applied across the
75
+ success, partial-salvage, and chat lanes; the walk's `complete`/root-absent health (F18) is now *consumed*
76
+ at the call site rather than discarded. A genuinely-empty run (root present, no files) still correctly
77
+ reports `[]`; only an unobservable root flips to unavailable.
78
+
79
+ - **microvm outputs are now observable — root-cause fix for `#52`.** A microvm run's outputs already live on
80
+ host disk (`VM_WORK_HOST` is mounted writable into the VM at `/sessions`), just at a different path than the
81
+ run dir the post-run pipeline walks. The run now **snapshots the session-root tree from the VM mount into
82
+ `outDir/work/session`** (mirroring how `hostloop` snapshots connected folders) — rm-before-copy,
83
+ symlinks copied verbatim (`dereference: false`), fail-loud if the tree is unexpectedly absent — plus captures
84
+ the pre-run manifest on this tier. Result: **`workspaceFiles`/`artifacts`, `file_exists`, `artifact_json`,
85
+ `user_visible_artifact`, `no_lost_write_back`, `no_unexpected_files`, and `input_unmodified` now work on
86
+ `microvm`** instead of being evidence-unavailable — verified live to be identical to `container`. It's also
87
+ fidelity-positive: real Cowork's VM outputs are host-observable too. (`no_scratchpad_leak` /
88
+ `present_files_called` stay `container`-only — they key off the `present_files` tool, not workspace
89
+ observability.) Doc/message sweep: the "use container/hostloop" carve-outs for these keys are removed.
90
+
91
+ - **`sync`'s non-macOS guard no longer blames Claude Desktop for a limitation that's actually this
92
+ harness's own.** The error previously read "sync requires macOS (the Cowork Desktop app is macOS-only)"
93
+ — false, Desktop ships a Windows build too; only this harness's `sync` tooling doesn't support non-macOS
94
+ install layouts yet. The message now says so.
95
+ - **The web_fetch provenance-miss denial is synced to Desktop 1.22209.0's wording.** The message now notes
96
+ that a URL surfaced in a WebSearch result also counts as provenance, and tells a subagent that can't ask
97
+ the user to continue without the page and report the blocked URL rather than stall. No logic change —
98
+ Desktop's underlying provenance rules are unchanged between releases, only the wording moved.
99
+ - **The hostloop path-gate no longer emits a spurious "cwd mismatch" warning when the run-dir is reached
100
+ through a symlink** (e.g. macOS `/tmp` → `/private/tmp`). The diagnostic now compares the wire and
101
+ spawner cwds after best-effort realpath canonicalization, so the same directory reached via two spellings
102
+ no longer false-alarms; the path-gate's actual allow/deny decision was already realpath-rooted and is
103
+ unaffected.
104
+ - **`verify-cassettes` no longer flags `claude.com` as a `domain` PII finding on every MCP-session
105
+ cassette.** The scanner's capability-manifest exclusion (`isCapabilityManifest()`) recognized only two
106
+ structural forms (the `system/init` event and the `initialize` registry `control_response`); it missed
107
+ the MCP `initialize` handshake itself — both Claude Code's own `control_request` (`clientInfo.websiteUrl`)
108
+ and the configured MCP server's `control_response` (`serverInfo`) — which fell through to the full
109
+ `domain`/`currency` net on every recording that talks to an MCP server. The only previous workaround,
110
+ `--allow-domain 'claude\.com'` in CI, was a class-scoped (not location-scoped) allow that would have
111
+ silently cleared a genuine `claude.com`-hosted leak anywhere else in the same cassette. Both handshake
112
+ forms are now recognized (shape-matched on the response side, since `serverInfo` is server-authored, not
113
+ Claude Code's own fixed string), and the CI gate no longer needs any `--allow-domain`/`--allow-email`
114
+ flag to pass on the committed example cassettes.
115
+ - **`hostloop`/`cowork` runs and `doctor --tier cowork` no longer hard-block when the pinned VM/container
116
+ ELF was pruned by a Desktop update but a patch-newer sibling is staged.** At that tier the ELF is
117
+ bind-mounted into the bash sidecar for parity and is not run by any harness-spawned process (the native
118
+ binary is the agent), so a same-major.minor patch bump is now auto-tolerated (loud note, advisory sha) —
119
+ matching the native binary's existing policy. The sha-pinned strictness is unchanged for
120
+ `container`/`microvm`, where the ELF is the executed agent.
121
+ - **`doctor --tier cowork` now mirrors the resolved loop** (`decideLoopFromBaseline`) for both agent
122
+ binaries, so it neither false-greens nor false-not-readies. When `cowork` resolves to host-loop it
123
+ tolerates the ELF patch bump and requires the native binary (the executed agent there); when it resolves
124
+ to VM-loop it keeps the ELF strict like `container` **and** stops requiring the native binary that a
125
+ VM-loop run doesn't use. Previously the tier's checks were unconditional, disagreeing with the actual
126
+ run on a VM-loop-resolving baseline.
127
+
128
+ - **`analyze-skill`: a `<script>…</script>` pair inside a docstring or comment no longer aborts a whole
129
+ file to could-not-verify.** The lexical block extractor could pull English prose out of a docstring or
130
+ comment as a phantom "script block"; its parse failure short-circuited the entire file to a
131
+ could-not-verify (exit 3), discarding the verdict already computed for the file's real, parseable
132
+ write-back block. The per-block analysis now accumulates: a block that fails to parse with no
133
+ `fetch`/XHR-`open`/`sendBeacon`/`axios`/`.post()` write-back hint is discounted as prose, while one that
134
+ carries a hint — or any block large enough to trip the analysis cap — is still reported as a
135
+ could-not-verify surfaced alongside any finding. A candidate whose every isolated `<script>` block is
136
+ unparseable stays a could-not-verify (fail-closed), never a silent clean pass. Several follow-on gaps in
137
+ that accumulation are also closed. First: whenever at least one `<script>` block was discounted as
138
+ prose, a parseable sibling block (or an already-flagged `<form method=post>`) no longer vouches for
139
+ write-back surface OUTSIDE every extracted block (top-level `.js`/`.ts` code, an inline `on*=` handler,
140
+ or surrounding template markup) — any write-back hint left in that un-analyzed remainder is now its own
141
+ could-not-verify, reported alongside any finding rather than silently passed; a source with no
142
+ discounted block, or an inline-handler write-back with nothing else in play, is unaffected. Second: the
143
+ write-back hint check (and the earlier candidacy check) now also recognizes optional-call spellings —
144
+ `fetch?.(`, `xhr?.open?.(`, `$.post?.(`, `navigator?.sendBeacon?.(` — so a source whose only write-back
145
+ uses `?.` is neither missed as a candidate nor discounted as prose inside an unparseable block; that
146
+ optional-call matching is also linearized (no more quadratic backtracking on a long non-matching
147
+ whitespace run). Third: a member-spelled write-back inside a block that DOES parse —
148
+ `window.fetch(...)`, `globalThis.fetch(...)`, `self.fetch(...)`, or the same spelling inside a same-file
149
+ fetch-wrapper's own body — is now classified the same as a bare `fetch(...)` call instead of going
150
+ unrecognized and falling through as clean; a bare `sendBeacon(...)` identifier call (e.g. a locally
151
+ bound alias of `navigator.sendBeacon`) is now recognized the same way as the member-spelled
152
+ `navigator.sendBeacon(...)`. Fourth: a `.post(...)`/`.put(...)`/`.patch(...)` call on a receiver outside
153
+ the known `axios`/`$`/`jQuery` set — the common miss being an axios instance,
154
+ `const api = axios.create(); api.post("/api/save", data)` — targeting a relative URL is no longer
155
+ invisible; it is now reported as an advisory finding (never escalated to an error, since the receiver
156
+ isn't provably a write-back client and could be unrelated code; `.delete(...)` is deliberately excluded
157
+ from this, since it's common on non-HTTP collection types). Fifth: the axios/`$`/`jQuery` verb set
158
+ recognized on the whitelisted identifier itself is widened from `.post(...)` alone to
159
+ `.post(...)`/`.put(...)`/`.patch(...)`/`.delete(...)`/`.postForm(...)`/`.putForm(...)`/`.patchForm(...)`
160
+ (the last three are axios v1's multipart form-data verb aliases; `.delete(...)` is INCLUDED here,
161
+ unlike the any-receiver advisory case above, since the literal `axios`/`$`/`jQuery` identifier has no
162
+ ambiguity about what it means); a bare config-object call — `axios({method:"POST", url:"/api/save",
163
+ ...})` — and `axios.request({...})` are both recognized, whether the config argument is an inline
164
+ object literal or a hoisted identifier (`const cfg = {...}; axios(cfg)`). This is not exhaustive: axios's
165
+ alternate `$.ajax({...})`-style config-key vocabulary, and any computed-member, whitespace-separated, or
166
+ aliased/re-exported spelling, remain a documented, lexically/structurally invisible accepted class (see
167
+ the relevant doc comments in `analyze-artifact.ts`) — as does a `formaction`/`formmethod` override on a
168
+ submit button that redirects an otherwise-remote `<form>` back to a relative, in-scope URL.
169
+ - **`analyze-skill`: a delete/remove flow that claims success now classifies as a lost write-back (error),
170
+ not just a suspect (advisory).** The success-claim vocabulary that distinguishes a lost write-back (an
171
+ unconditional "it worked" toast) from a merely-suspect one gained delete-flow words (`deleted`,
172
+ `removed`, …) alongside the existing `saved`/`submitted`/`persisted`/`completed`/`success`. A relative
173
+ `DELETE` write-back whose only success signal is a "Deleted!"/"Removed!" toast — no `resp.ok`/status
174
+ check — is now flagged at error severity like its save-flow equivalent, since under Cowork it resolves
175
+ non-ok against Cowork's own origin and the false confirmation is identical.
176
+
177
+ ### Documentation
178
+
179
+ - **Documented an `analyze-skill --runtime` recipe for agent-generated artifacts.**
180
+ `analyze-skill --runtime <run-dir>/work/session/mnt/outputs` confirms interactive-artifact write-backs in
181
+ HTML the agent *generates during a run* — content the source-only static scan can't see until a run has
182
+ happened. Notes the tier-specific output paths and the microvm/replay caveats.
183
+
184
+ ## [1.1.0] — 2026-07-16
185
+
186
+ Minor: `analyze-skill` gains interactive-artifact write-back detection (static + an optional `--runtime`
187
+ headless-DOM confirmer), `lint` gains a container-only-key tier check, and the `doctor` JSON envelope is
188
+ frozen as a covered SPEC §12 surface. All additive.
189
+
190
+ ### Added
191
+
192
+ - **`analyze-skill` now detects interactive-artifact write-backs lost under Cowork.** Alongside the
193
+ existing `/sessions` path scan, it statically analyzes `.html/.htm/.js/.mjs/.ts/.jsx/.tsx/.py` sources
194
+ under the target for a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back that silently
195
+ fails under Cowork (the artifact is served from Cowork's own origin, so a relative write-back resolves
196
+ non-ok and a page that doesn't check `resp.ok` shows a false "Saved"). Findings: `artifact-write-back-lost`
197
+ (error — gates under `--strict`), `artifact-write-back-suspect` (advisory), and a separate top-level
198
+ `analysisFailures` **could-not-verify** channel (a candidate that couldn't be parsed/analyzed) that
199
+ always exits `3`, `--strict`-independent. A guard that isn't statically provable-truthy is `suspect`,
200
+ never silently clean; the blanket `analyze-skill: ignore` marker does not silence artifact rules.
201
+ Each `SkillFinding` now carries a `severity` (`error|advisory`); `--strict` gates on any error finding.
202
+ - **`analyze-skill --runtime`** — an optional headless-DOM confirmation that drives a materialized `.html`
203
+ artifact in jsdom (stubbed network + synthetic user actions, run twice) to *observe* whether a relative
204
+ write-back fires and is lost. Enrichment only (never changes the exit code); trusted-source scope. `jsdom`
205
+ is an optional, dynamically-imported dependency — absent it reports "run `npm i jsdom` to enable".
206
+ - **`schema/doctor.json`** — the `doctor --output-format json` envelope is now a covered SPEC §12 surface
207
+ (`oneOf` the completed-probe shape and the shared error envelope for every category). `doctor`'s normal
208
+ JSON output is standardized through the shared envelope frame.
209
+ - **`lint` flags container-only assertion keys off-container** — `no_scratchpad_leak`/`present_files_called`
210
+ on `fidelity: protocol|microvm|hostloop` is an ERROR, on `fidelity: cowork` a WARN, clean on `container`.
211
+
9
212
  ## [1.0.6] — 2026-07-15
10
213
 
11
214
  Patch: platform baseline synced to Claude Desktop `1.21459.0`. The spawn contract and rendered system