cowork-harness 1.0.6 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +14 -8
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +35 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +9 -6
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +48 -0
- package/CHANGELOG.md +203 -0
- package/README.md +15 -12
- package/SPEC.md +42 -4
- package/dist/assert.js +214 -3
- package/dist/baseline.js +36 -8
- package/dist/cli.js +46 -9
- package/dist/decide/decider.js +9 -1
- package/dist/hostloop/canusetool-gate.js +86 -2
- package/dist/hostloop/provenance.js +7 -2
- package/dist/hostloop/workspace-handler.js +4 -2
- package/dist/run/analyze-artifact-runtime.js +547 -0
- package/dist/run/analyze-artifact.js +1190 -0
- package/dist/run/analyze-skill.js +99 -21
- package/dist/run/artifacts.js +11 -4
- package/dist/run/cassette.js +37 -6
- package/dist/run/chat-result.js +4 -2
- package/dist/run/doctor.js +98 -18
- package/dist/run/execute.js +52 -10
- package/dist/run/latest-run.js +43 -0
- package/dist/run/pre-run-manifest.js +3 -2
- package/dist/run/renderer.js +4 -0
- package/dist/run/status-target.js +28 -0
- package/dist/run/trace-view.js +1 -1
- package/dist/runtime/hostloop.js +23 -4
- package/dist/runtime/microvm.js +44 -7
- package/dist/scan.js +17 -14
- package/dist/sync/cowork-sync.js +10 -1
- package/dist/types.js +15 -1
- package/docs/cassette.md +18 -9
- package/docs/fidelity-gaps.md +47 -0
- package/docs/maintenance.md +2 -0
- package/docs/run-status.md +5 -0
- package/docs/scenario.md +12 -5
- package/docs/subagents.md +68 -0
- package/examples/replays/README.md +1 -1
- package/package.json +4 -1
- package/python/test_scenario_lint.py +88 -0
- package/schema/doctor.json +81 -0
- package/schema/scenario.schema.json +16 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.0
|
|
7
|
-
tracks-harness: cowork-harness 1.0
|
|
6
|
+
version: 1.2.0
|
|
7
|
+
tracks-harness: cowork-harness 1.2.0 (baseline desktop-1.21459.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.0
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.2.0` (baseline
|
|
26
26
|
> `desktop-1.21459.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.0
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.2.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.2.0"`. **Pin `@>=1.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
|
-
What the ≥ 1.0
|
|
44
|
+
What the ≥ 1.2.0 floor gates, by release:
|
|
45
45
|
|
|
46
46
|
- **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
|
|
47
47
|
- **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
|
|
@@ -52,6 +52,8 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
52
52
|
- **0.32.0:** `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict`.
|
|
53
53
|
- **0.33.0:** the `redacted` marker on display-omitted reasoning — `subagents[].reasoning` and the top-level `thinking[]` now carry `{text:"", redacted:true}` when the model returns a signed-but-empty thought (so "reasoned, text omitted" is distinct from "no thought"), plus the fenced `debug.thinking_display` escape hatch.
|
|
54
54
|
- **1.0.0:** first stable release — the SPEC §12 compatibility contract takes effect (covered CLI/schema/env/Action surfaces are now stable; breaking changes need a major bump). No new author-facing command; the floor simply tracks the 1.0 release.
|
|
55
|
+
- **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
|
|
56
|
+
- **1.2.0:** three new assertion keys — `no_lost_write_back: true` (the write-back detector above wired as a per-scenario gate over the run's authored files, live-only) and the regex siblings `tool_result_matches`/`tool_result_not_matches` (case-insensitive per-result, for an error-signature *family* a literal substring can't express). **`microvm` outputs are now observable** — its session tree is snapshotted from the VM into the run dir, so `file_exists`/`artifact_json`/`user_visible_artifact`/`no_unexpected_files`/`input_unmodified`/`no_lost_write_back`/`semantic_matches` all work there, no longer `container`/`hostloop`-only. `status <dir>` also resolves the newest session under a `--run-dir` root; the completion footer prints a `→ result: …/result.json` pointer; the `on_unanswered=fail` error also points to `on_unanswered: llm`. `analyze-skill` hardening: a phantom `<script>` prose block no longer sinks a real verdict, a delete/remove flow claiming success classifies as lost (error) not just suspect, and write-back detection widened (optional-call `?.` spellings, member-spelled/aliased `fetch`/`sendBeacon`, axios instance/config forms).
|
|
55
57
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
56
58
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
57
59
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
@@ -220,7 +222,8 @@ them by what you're trying to prove:
|
|
|
220
222
|
| a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
|
|
221
223
|
| a tool ran **inside** a skill's scope | `skill_tool_used: {skill, tool}` |
|
|
222
224
|
| a sub-agent did the work | `subagent_output_contains: {contains}`, `subagent_dispatched: <regex>`, `dispatch_count_max: <N>` |
|
|
223
|
-
| a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run
|
|
225
|
+
| a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
|
|
226
|
+
| no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
|
|
224
227
|
| a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
|
|
225
228
|
| a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
|
|
226
229
|
| every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
|
|
@@ -277,7 +280,8 @@ quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `
|
|
|
277
280
|
(ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
|
|
278
281
|
baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
|
|
279
282
|
`allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
|
|
280
|
-
unverifiable), a
|
|
283
|
+
unverifiable), a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) off the
|
|
284
|
+
`container` tier (ERROR on `protocol`/`microvm`/`hostloop`, WARN on `cowork`), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
|
|
281
285
|
and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
|
|
282
286
|
(CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
|
|
283
287
|
emits a scenario `lint` would reject.
|
|
@@ -372,7 +376,9 @@ harness writes/updates throughout the run's lifecycle (including a crash-safety
|
|
|
372
376
|
error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
|
|
373
377
|
either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
|
|
374
378
|
checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
|
|
375
|
-
run starts, so capture stderr to get the exact directory
|
|
379
|
+
run starts, so capture stderr to get the exact directory — but `<dir>` also accepts the run-dir root
|
|
380
|
+
passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
|
|
381
|
+
newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
|
|
376
382
|
rather than hanging forever. (Fuller recipe in `docs/run-status.md` — repo-only, not in the installed
|
|
377
383
|
payload; `cowork-harness status --help` has the flags.)
|
|
378
384
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.0
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.2.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.0
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.2.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.0
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.2.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.0
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.2.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.0
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.2.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.0
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.2.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -130,7 +130,10 @@ hands `fn` exactly this dict.
|
|
|
130
130
|
### Determinism contract
|
|
131
131
|
|
|
132
132
|
- `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
|
|
133
|
-
`--answer`/`choose` to add
|
|
133
|
+
`--answer`/`choose` to add, and also suggests `on_unanswered: llm` (in the scenario YAML) as a
|
|
134
|
+
secondary escape valve for a gate whose wording drifts run-to-run — a regex chases a moving target,
|
|
135
|
+
at the cost of non-determinism (one model call per gate). Correct, but flaky for skills whose gates
|
|
136
|
+
appear stochastically.
|
|
134
137
|
- `first` — picks option 1 and warns loudly. **Flagged `nonDeterministic`** — not a deterministic
|
|
135
138
|
substitute for scripted answers. (For a web_fetch approval gate it abstains → fail-closed.)
|
|
136
139
|
- `prompt` — asks at the TTY (`skill` only).
|
|
@@ -171,6 +174,36 @@ writes a `result.json` (marked `partial: true`) with the artifacts the agent pro
|
|
|
171
174
|
the work isn't discarded — then exits 2. Inspect it with `cowork-harness inspect <run-dir>`. `verify-run`
|
|
172
175
|
and `scaffold` refuse to treat a partial run's half-finished output as a passing result.
|
|
173
176
|
|
|
177
|
+
## What a green does NOT prove
|
|
178
|
+
|
|
179
|
+
A passing run is evidence for exactly what it checked, not a blanket certificate. Three gaps come
|
|
180
|
+
up often enough to spell out:
|
|
181
|
+
|
|
182
|
+
- **A green `replay` proves "same as when recorded," not "correct today."** `replay` never touches
|
|
183
|
+
a filesystem or network — it re-evaluates assertions from the frozen cassette. A fixed set of
|
|
184
|
+
keys is live-only and **skipped outright** on replay (absent from `assertions[]`, not vacuously
|
|
185
|
+
passed): `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
|
|
186
|
+
`egress_allowed`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back`, and `expect_denied`.
|
|
187
|
+
Everything else that *is* evaluated is checked against the **recording**, not fresh behavior — a
|
|
188
|
+
green replay says the skill produced these events when it was recorded, not that it still does
|
|
189
|
+
(`staleness[]` flags skill/baseline drift as a hint; only a live `run` re-confirms current
|
|
190
|
+
behavior). See `docs/cassette.md` § "Still skipped on replay" and `docs/scenario.md` § "Which
|
|
191
|
+
assertions survive replay."
|
|
192
|
+
- **Container-only assertions can't verify off the `container` tier.** `no_scratchpad_leak` and
|
|
193
|
+
`present_files_called` check the `present_files` delivery path, which is served **only** on
|
|
194
|
+
`container` — not `hostloop`/`microvm`. Asserting them off-container hard-fails at runtime (a red
|
|
195
|
+
run, not a false green), so you won't be fooled if you write the assertion. The quieter trap is a
|
|
196
|
+
scenario that runs at `hostloop`/`microvm`/`protocol` and simply omits these assertions: a green
|
|
197
|
+
run there proves **nothing** about scratchpad-leak safety or present_files delivery, because that
|
|
198
|
+
tier never exercises the delivery path the assertions would check. Use `fidelity: container` for
|
|
199
|
+
present_files/scratchpad-delivery coverage.
|
|
200
|
+
- **The harness doesn't observe rendered-artifact interactions, browser downloads, or human
|
|
201
|
+
clicks.** It runs the agent headless — no webview, no browser, no person clicking "Submit." A
|
|
202
|
+
class of Cowork bug (a client-side write-back to a relative URL that resolves-but-fails against
|
|
203
|
+
Cowork's own origin, or a broken blob-download fallback) is invisible to any live run, however
|
|
204
|
+
faithfully sandboxed, because it only manifests in a rendered DOM a human is driving. See
|
|
205
|
+
`docs/fidelity-gaps.md` § "Browser↔webview↔human-interaction boundary."
|
|
206
|
+
|
|
174
207
|
## Relevant environment variables
|
|
175
208
|
|
|
176
209
|
- `COWORK_HARNESS_RUNS_DIR` (or `--run-dir <path>`) — override the default run-output root `~/.cowork-harness/runs` (out of any working tree). flag > env > default.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.0
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.2.0`
|
|
4
4
|
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
@@ -251,13 +251,16 @@ same set live from the schema.
|
|
|
251
251
|
| `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
|
|
252
252
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
253
253
|
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
|
|
254
|
-
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest
|
|
255
|
-
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail;
|
|
254
|
+
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning |
|
|
255
|
+
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
|
|
256
256
|
| `self_heal_ran: <bool>` | a plugin-root self-heal script was (not) invoked |
|
|
257
|
+
| `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
|
|
257
258
|
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
|
|
258
259
|
| `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
|
|
259
260
|
| `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
|
|
260
261
|
| `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
|
|
262
|
+
| `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
|
|
263
|
+
| `tool_result_not_matches: <regex>` | the regex sibling of `tool_result_not_contains` — same fails-loud-on-absent-evidence semantics |
|
|
261
264
|
| `subagent_tool_used: <glob>` | a sub-agent used a tool matching this glob (same `*`/`?`, anchored, case-sensitive semantics as `tool_called`, including the empty/regex-ish rejection) |
|
|
262
265
|
| `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
|
|
263
266
|
| `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
|
|
@@ -312,7 +315,7 @@ same set live from the schema.
|
|
|
312
315
|
| `egress_allowed: <host>` | the host was allowed through |
|
|
313
316
|
| `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
|
|
314
317
|
| `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
|
|
315
|
-
| `semantic_matches: {rubric: [...], min_pass?, judge_model?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose (authored-file evidence is
|
|
318
|
+
| `semantic_matches: {rubric: [...], min_pass?, judge_model?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
|
|
316
319
|
| `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
|
|
317
320
|
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) |
|
|
318
321
|
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** |
|
|
@@ -407,8 +410,8 @@ staleness `fingerprint` shows ANY skill/baseline drift, or `replay --fail-on-ski
|
|
|
407
410
|
skill-source drift; every replay result also reports it class-tagged in `staleness[]` for a JSON gate.
|
|
408
411
|
|
|
409
412
|
**Egress + other filesystem — still skipped on replay (live-only):** `no_delete_in_outputs`,
|
|
410
|
-
`self_heal_ran`, `transcript_no_host_path`, `egress_*` / `expect_denied`, `no_mcp_error`, `max_peak_rss_bytes
|
|
411
|
-
These run only on a live `run`/`record`.
|
|
413
|
+
`self_heal_ran`, `transcript_no_host_path`, `egress_*` / `expect_denied`, `no_mcp_error`, `max_peak_rss_bytes`,
|
|
414
|
+
`no_lost_write_back`. These run only on a live `run`/`record`.
|
|
412
415
|
|
|
413
416
|
**Mixed assertions on replay:** before evaluating, `replay` strips each assertion to its replay-checkable
|
|
414
417
|
keys and drops any left empty. So `{result, egress_denied}` evaluates on replay as `{result}` alone — its
|
|
@@ -27,6 +27,7 @@
|
|
|
27
27
|
"max_turns",
|
|
28
28
|
"no_delete_in_outputs",
|
|
29
29
|
"no_hook_blocked",
|
|
30
|
+
"no_lost_write_back",
|
|
30
31
|
"no_mcp_error",
|
|
31
32
|
"no_path_denied",
|
|
32
33
|
"no_scratchpad_leak",
|
|
@@ -60,7 +61,9 @@
|
|
|
60
61
|
"tool_no_error_if_called",
|
|
61
62
|
"tool_not_called",
|
|
62
63
|
"tool_result_contains",
|
|
64
|
+
"tool_result_matches",
|
|
63
65
|
"tool_result_not_contains",
|
|
66
|
+
"tool_result_not_matches",
|
|
64
67
|
"transcript_contains",
|
|
65
68
|
"transcript_matches",
|
|
66
69
|
"transcript_no_host_path",
|
|
@@ -17,6 +17,8 @@ Two subcommands:
|
|
|
17
17
|
lint flags (see references/scenario-schema.md for the why of each):
|
|
18
18
|
E egress assertion on `fidelity: protocol` (the harness rejects this run)
|
|
19
19
|
E `transcript_no_host_path` on hostloop/protocol (fails BY DESIGN at those tiers)
|
|
20
|
+
E `no_scratchpad_leak`/`present_files_called` off (container-only: served only at fidelity:
|
|
21
|
+
protocol/microvm/hostloop container; off-container = cannot-verify)
|
|
20
22
|
E `requires_capabilities` on `fidelity: protocol` (probe can't run → hard-fails
|
|
21
23
|
unless allow_missing_capability)
|
|
22
24
|
E on_unanswered: agent / invalid value (schema rejects `agent`)
|
|
@@ -24,6 +26,8 @@ lint flags (see references/scenario-schema.md for the why of each):
|
|
|
24
26
|
E `assertions:` instead of `assert:` (block ignored → every check no-ops)
|
|
25
27
|
W `transcript_no_host_path` on `fidelity: cowork` (tier resolves per baseline gate —
|
|
26
28
|
incompatible if it lands hostloop)
|
|
29
|
+
W `no_scratchpad_leak`/`present_files_called` on (tier resolves per baseline gate — cannot-verify
|
|
30
|
+
`fidelity: cowork` if it resolves off-container)
|
|
27
31
|
W no content assertion → no-op on a replay gate (every assertion is fs/egress)
|
|
28
32
|
W mixed-class assert item → fs/egress half dropped on replay
|
|
29
33
|
W unknown top-level / assertion key (typo or hallucinated schema)
|
|
@@ -58,6 +62,8 @@ CONTENT_KEYS = {
|
|
|
58
62
|
"transcript_not_matches",
|
|
59
63
|
"tool_result_contains",
|
|
60
64
|
"tool_result_not_contains",
|
|
65
|
+
"tool_result_matches",
|
|
66
|
+
"tool_result_not_matches",
|
|
61
67
|
"tool_called",
|
|
62
68
|
"tool_not_called",
|
|
63
69
|
"subagent_tool_used",
|
|
@@ -125,8 +131,13 @@ LIVE_ONLY_KEYS = {
|
|
|
125
131
|
"no_mcp_error",
|
|
126
132
|
"max_peak_rss_bytes",
|
|
127
133
|
"semantic_matches",
|
|
134
|
+
"no_lost_write_back",
|
|
128
135
|
}
|
|
129
136
|
EGRESS_KEYS = {"egress_denied", "egress_allowed"}
|
|
137
|
+
# container-only: served only at fidelity: container (present_files / the scratchpad promotion path
|
|
138
|
+
# it depends on). Off-container these report cannot-verify, not a meaningful pass/fail — same tier-fidelity
|
|
139
|
+
# class as transcript_no_host_path below, just the opposite direction (container-only vs container-hostile).
|
|
140
|
+
CONTAINER_ONLY_KEYS = {"no_scratchpad_leak", "present_files_called"}
|
|
130
141
|
# verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
|
|
131
142
|
VERDICT_MODIFIER_KEYS = {
|
|
132
143
|
"allow_permissive_auto_allow",
|
|
@@ -221,6 +232,8 @@ REGEX_KEYS = {
|
|
|
221
232
|
"subagent_dispatched",
|
|
222
233
|
"question_asked",
|
|
223
234
|
"hook_blocked",
|
|
235
|
+
"tool_result_matches",
|
|
236
|
+
"tool_result_not_matches",
|
|
224
237
|
}
|
|
225
238
|
VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
|
|
226
239
|
VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
|
|
@@ -388,6 +401,41 @@ def lint_doc(doc, path, raw_lines):
|
|
|
388
401
|
)
|
|
389
402
|
)
|
|
390
403
|
|
|
404
|
+
# E/W: CONTAINER_ONLY_KEYS (no_scratchpad_leak / present_files_called) are served only at
|
|
405
|
+
# fidelity: container — off-container they report cannot-verify, not a meaningful check. Mirrors
|
|
406
|
+
# the transcript_no_host_path tier check above, just the opposite direction: ERROR on the tiers
|
|
407
|
+
# where the runtime deterministically can't serve it, WARN on cowork (baseline-gate-resolution
|
|
408
|
+
# dependent — the linter stays offline, so the message names the dependency instead of resolving it).
|
|
409
|
+
container_only_present = sorted(assert_keys & CONTAINER_ONLY_KEYS)
|
|
410
|
+
if container_only_present:
|
|
411
|
+
if fidelity in ("protocol", "microvm", "hostloop"):
|
|
412
|
+
findings.append(
|
|
413
|
+
Finding(
|
|
414
|
+
"ERROR",
|
|
415
|
+
"container-only-key-off-container",
|
|
416
|
+
f"{container_only_present} on `fidelity: {fidelity}` — container-only (present_files "
|
|
417
|
+
"is served only there); off-container it reports cannot-verify. Use fidelity: "
|
|
418
|
+
"container for present_files/scratchpad delivery.",
|
|
419
|
+
"Use fidelity: container (or drop the assertion for this tier).",
|
|
420
|
+
path,
|
|
421
|
+
)
|
|
422
|
+
)
|
|
423
|
+
elif fidelity == "cowork":
|
|
424
|
+
findings.append(
|
|
425
|
+
Finding(
|
|
426
|
+
"WARN",
|
|
427
|
+
"container-only-key-off-container",
|
|
428
|
+
f"{container_only_present} on `fidelity: cowork` — container-only (present_files "
|
|
429
|
+
"is served only there); off-container it reports cannot-verify. Use fidelity: "
|
|
430
|
+
f"container for present_files/scratchpad delivery. (The tier resolves per the "
|
|
431
|
+
f"baseline's host-loop gate ({HOST_LOOP_GATE_ID}); if it resolves off-container this "
|
|
432
|
+
"assertion cannot verify anything.)",
|
|
433
|
+
"Pin fidelity: container if the assertion is load-bearing; keep cowork only if "
|
|
434
|
+
"you accept the gate-resolution dependency.",
|
|
435
|
+
path,
|
|
436
|
+
)
|
|
437
|
+
)
|
|
438
|
+
|
|
391
439
|
# E: requires_capabilities on protocol — the capability probe cannot run at protocol tier
|
|
392
440
|
# (clause b of the requires_capabilities contract), so the run HARD-FAILS unless an assert item
|
|
393
441
|
# opts out via allow_missing_capability: true. Offline-detectable fails-by-design, same class as
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,209 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.2.0] — 2026-07-18
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`no_lost_write_back: true` scenario assertion** — gate a scenario on "the agent didn't emit an
|
|
14
|
+
interactive artifact whose Submit is lost under Cowork". It runs the shipped static Tier A analyzer
|
|
15
|
+
(`analyze-artifact`, deterministic, no headless DOM) over the files the run authored — diffed against the
|
|
16
|
+
pre-run manifest — so a lost relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back becomes a
|
|
17
|
+
per-scenario verdict, not just an out-of-band `analyze-skill` scan. A lost write-back on an **added**
|
|
18
|
+
agent-authored source (`outputs/` or the scratchpad) fails; a **pre-existing** file the skill only modified
|
|
19
|
+
on a read-write connected mount is advisory (not the skill's to own); `-suspect` findings surface but pass.
|
|
20
|
+
Honest evidence-unavailable semantics: could-not-verify (fail-closed, never a silent clean) on a `--resume`
|
|
21
|
+
scratchpad walk or an authored candidate that couldn't be analyzed. Runs on every live sandbox tier
|
|
22
|
+
including **microvm** (its outputs are snapshotted from the VM into the run dir — see the #52 entry below).
|
|
23
|
+
**Live/verify-run only** — skipped-loud on replay (a cassette embedding the key never
|
|
24
|
+
hard-fails its replay); `verify-run` recomputes the authored set from the kept work dir. Only `true` is
|
|
25
|
+
valid (omit to skip).
|
|
26
|
+
- **`tool_result_matches` / `tool_result_not_matches` scenario assertion keys** — the case-insensitive regex
|
|
27
|
+
siblings of `tool_result_contains`/`tool_result_not_contains`, evaluated per captured tool result (subject
|
|
28
|
+
to the same 10 KB per-result assertText cap). Useful for catching an error-signature *family* (e.g. a
|
|
29
|
+
script's non-zero exit swallowed by its wrapper, but the message still printed) that a literal substring
|
|
30
|
+
match can't express. Same evidence-unavailable wording as the `_contains` pair: a bad regex fails the
|
|
31
|
+
assertion with a compile error, and a no-match against a display-truncated result is reported as
|
|
32
|
+
could-not-verify rather than a silent pass/fail.
|
|
33
|
+
|
|
34
|
+
- **A folder-grant refusal for `request_cowork_directory`, ported from Desktop 1.22209.0.** Cowork now
|
|
35
|
+
refuses (pre-prompt) a mid-session folder grant that targets a security-sensitive home-directory path —
|
|
36
|
+
`.ssh`, `.aws`, `.gnupg`, `.kube`, `.docker`, `.claude`, `.config/{gcloud,gh,powershell}`, the darwin
|
|
37
|
+
`Library/{Keychains,LaunchAgents,LaunchDaemons,Application Support,Cookies}` paths, or a protected shell
|
|
38
|
+
dotfile (`.zshrc`, `.netrc`, etc.) — either directly, as a descendant, or as an ancestor whose grant would
|
|
39
|
+
incidentally expose one (e.g. requesting the home directory itself). The harness's `hostloop` canUseTool
|
|
40
|
+
gate ports this byte-faithfully, denying with Desktop's own message. **Currently dead code in a stock
|
|
41
|
+
run**: no built-in workspace/cowork server registers `request_cowork_directory` yet (a pre-existing,
|
|
42
|
+
separately tracked gap), so this only fires for a scenario that supplies its own `mcp_config` registering
|
|
43
|
+
a matching tool name. Ported ahead of that gap closing so the refusal semantics are ready the moment it
|
|
44
|
+
does. Two GrowthBook feature-gate ids Desktop 1.22209.0 introduced for a related "auto mode always-allow"
|
|
45
|
+
tool-approval feature are pinned as drift sentinels (`sync`'s `PINNED_GATES`) without being behaviorally
|
|
46
|
+
modeled — this harness has no persistent per-tool permission concept to model them against.
|
|
47
|
+
|
|
48
|
+
### Changed
|
|
49
|
+
|
|
50
|
+
- **`cowork-harness status <run-dir>` now resolves the newest session under a run-dir root.** Previously
|
|
51
|
+
`status` only worked against the exact per-session out-dir printed at run start; pointing it at the root
|
|
52
|
+
passed to `run --run-dir` failed with "no status.json". It now scans up to two levels under a directory
|
|
53
|
+
lacking its own `status.json` for the newest session that has one, and reads that instead.
|
|
54
|
+
- **`doctor`'s two staged-agent checks are now titled distinctly** — "Staged agent binary (VM/container
|
|
55
|
+
ELF)" vs. "Staged native agent binary (hostloop)" — so a failure on either is attributable to the right
|
|
56
|
+
one instead of reading as an ambiguous duplicate.
|
|
57
|
+
- **The run-completion footer now prints a `→ result: <run-dir>/result.json` pointer**, on both success and
|
|
58
|
+
failure, since the run directory is always kept on disk. Suppressed on the replay lane, which never
|
|
59
|
+
writes a `result.json`.
|
|
60
|
+
- **The `on_unanswered=fail` unscripted-gate error now also mentions `on_unanswered: llm` as a secondary
|
|
61
|
+
escape valve.** Previously it suggested only `--answer "<regex>=<choice>"`, which is the right primary
|
|
62
|
+
fix but the wrong tool for a gate whose wording drifts run-to-run — a regex chases a moving target. The
|
|
63
|
+
added line explicitly says "in the scenario YAML" (`--on-unanswered llm` is rejected on the CLI in favor
|
|
64
|
+
of `--decider-llm`) and notes the tradeoff (non-deterministic, one model call per gate) so it doesn't
|
|
65
|
+
read as unconditionally preferable to fixing the script.
|
|
66
|
+
|
|
67
|
+
### Fixed
|
|
68
|
+
|
|
69
|
+
- **Silent false-green on a missing workspace root (`#52`).** When the workspace root (`outDir/work/session/mnt`)
|
|
70
|
+
couldn't be walked — the canonical case is a **microvm** run, whose outputs stage into the VM work tree, not
|
|
71
|
+
into the run dir — `RunResult.workspaceFiles`/`artifacts` persisted as `[]`, **indistinguishable from a run
|
|
72
|
+
that genuinely wrote nothing.** A consumer reading `result.json` (e.g. skill-creator-plus) saw "zero
|
|
73
|
+
artifacts, clean." They now persist as **`undefined` (unavailable)** — the same convention replay already
|
|
74
|
+
uses for "no live filesystem to scan" — with a loud `::warning::` naming the reason. Applied across the
|
|
75
|
+
success, partial-salvage, and chat lanes; the walk's `complete`/root-absent health (F18) is now *consumed*
|
|
76
|
+
at the call site rather than discarded. A genuinely-empty run (root present, no files) still correctly
|
|
77
|
+
reports `[]`; only an unobservable root flips to unavailable.
|
|
78
|
+
|
|
79
|
+
- **microvm outputs are now observable — root-cause fix for `#52`.** A microvm run's outputs already live on
|
|
80
|
+
host disk (`VM_WORK_HOST` is mounted writable into the VM at `/sessions`), just at a different path than the
|
|
81
|
+
run dir the post-run pipeline walks. The run now **snapshots the session-root tree from the VM mount into
|
|
82
|
+
`outDir/work/session`** (mirroring how `hostloop` snapshots connected folders) — rm-before-copy,
|
|
83
|
+
symlinks copied verbatim (`dereference: false`), fail-loud if the tree is unexpectedly absent — plus captures
|
|
84
|
+
the pre-run manifest on this tier. Result: **`workspaceFiles`/`artifacts`, `file_exists`, `artifact_json`,
|
|
85
|
+
`user_visible_artifact`, `no_lost_write_back`, `no_unexpected_files`, and `input_unmodified` now work on
|
|
86
|
+
`microvm`** instead of being evidence-unavailable — verified live to be identical to `container`. It's also
|
|
87
|
+
fidelity-positive: real Cowork's VM outputs are host-observable too. (`no_scratchpad_leak` /
|
|
88
|
+
`present_files_called` stay `container`-only — they key off the `present_files` tool, not workspace
|
|
89
|
+
observability.) Doc/message sweep: the "use container/hostloop" carve-outs for these keys are removed.
|
|
90
|
+
|
|
91
|
+
- **`sync`'s non-macOS guard no longer blames Claude Desktop for a limitation that's actually this
|
|
92
|
+
harness's own.** The error previously read "sync requires macOS (the Cowork Desktop app is macOS-only)"
|
|
93
|
+
— false, Desktop ships a Windows build too; only this harness's `sync` tooling doesn't support non-macOS
|
|
94
|
+
install layouts yet. The message now says so.
|
|
95
|
+
- **The web_fetch provenance-miss denial is synced to Desktop 1.22209.0's wording.** The message now notes
|
|
96
|
+
that a URL surfaced in a WebSearch result also counts as provenance, and tells a subagent that can't ask
|
|
97
|
+
the user to continue without the page and report the blocked URL rather than stall. No logic change —
|
|
98
|
+
Desktop's underlying provenance rules are unchanged between releases, only the wording moved.
|
|
99
|
+
- **The hostloop path-gate no longer emits a spurious "cwd mismatch" warning when the run-dir is reached
|
|
100
|
+
through a symlink** (e.g. macOS `/tmp` → `/private/tmp`). The diagnostic now compares the wire and
|
|
101
|
+
spawner cwds after best-effort realpath canonicalization, so the same directory reached via two spellings
|
|
102
|
+
no longer false-alarms; the path-gate's actual allow/deny decision was already realpath-rooted and is
|
|
103
|
+
unaffected.
|
|
104
|
+
- **`verify-cassettes` no longer flags `claude.com` as a `domain` PII finding on every MCP-session
|
|
105
|
+
cassette.** The scanner's capability-manifest exclusion (`isCapabilityManifest()`) recognized only two
|
|
106
|
+
structural forms (the `system/init` event and the `initialize` registry `control_response`); it missed
|
|
107
|
+
the MCP `initialize` handshake itself — both Claude Code's own `control_request` (`clientInfo.websiteUrl`)
|
|
108
|
+
and the configured MCP server's `control_response` (`serverInfo`) — which fell through to the full
|
|
109
|
+
`domain`/`currency` net on every recording that talks to an MCP server. The only previous workaround,
|
|
110
|
+
`--allow-domain 'claude\.com'` in CI, was a class-scoped (not location-scoped) allow that would have
|
|
111
|
+
silently cleared a genuine `claude.com`-hosted leak anywhere else in the same cassette. Both handshake
|
|
112
|
+
forms are now recognized (shape-matched on the response side, since `serverInfo` is server-authored, not
|
|
113
|
+
Claude Code's own fixed string), and the CI gate no longer needs any `--allow-domain`/`--allow-email`
|
|
114
|
+
flag to pass on the committed example cassettes.
|
|
115
|
+
- **`hostloop`/`cowork` runs and `doctor --tier cowork` no longer hard-block when the pinned VM/container
|
|
116
|
+
ELF was pruned by a Desktop update but a patch-newer sibling is staged.** At that tier the ELF is
|
|
117
|
+
bind-mounted into the bash sidecar for parity and is not run by any harness-spawned process (the native
|
|
118
|
+
binary is the agent), so a same-major.minor patch bump is now auto-tolerated (loud note, advisory sha) —
|
|
119
|
+
matching the native binary's existing policy. The sha-pinned strictness is unchanged for
|
|
120
|
+
`container`/`microvm`, where the ELF is the executed agent.
|
|
121
|
+
- **`doctor --tier cowork` now mirrors the resolved loop** (`decideLoopFromBaseline`) for both agent
|
|
122
|
+
binaries, so it neither false-greens nor false-not-readies. When `cowork` resolves to host-loop it
|
|
123
|
+
tolerates the ELF patch bump and requires the native binary (the executed agent there); when it resolves
|
|
124
|
+
to VM-loop it keeps the ELF strict like `container` **and** stops requiring the native binary that a
|
|
125
|
+
VM-loop run doesn't use. Previously the tier's checks were unconditional, disagreeing with the actual
|
|
126
|
+
run on a VM-loop-resolving baseline.
|
|
127
|
+
|
|
128
|
+
- **`analyze-skill`: a `<script>…</script>` pair inside a docstring or comment no longer aborts a whole
|
|
129
|
+
file to could-not-verify.** The lexical block extractor could pull English prose out of a docstring or
|
|
130
|
+
comment as a phantom "script block"; its parse failure short-circuited the entire file to a
|
|
131
|
+
could-not-verify (exit 3), discarding the verdict already computed for the file's real, parseable
|
|
132
|
+
write-back block. The per-block analysis now accumulates: a block that fails to parse with no
|
|
133
|
+
`fetch`/XHR-`open`/`sendBeacon`/`axios`/`.post()` write-back hint is discounted as prose, while one that
|
|
134
|
+
carries a hint — or any block large enough to trip the analysis cap — is still reported as a
|
|
135
|
+
could-not-verify surfaced alongside any finding. A candidate whose every isolated `<script>` block is
|
|
136
|
+
unparseable stays a could-not-verify (fail-closed), never a silent clean pass. Several follow-on gaps in
|
|
137
|
+
that accumulation are also closed. First: whenever at least one `<script>` block was discounted as
|
|
138
|
+
prose, a parseable sibling block (or an already-flagged `<form method=post>`) no longer vouches for
|
|
139
|
+
write-back surface OUTSIDE every extracted block (top-level `.js`/`.ts` code, an inline `on*=` handler,
|
|
140
|
+
or surrounding template markup) — any write-back hint left in that un-analyzed remainder is now its own
|
|
141
|
+
could-not-verify, reported alongside any finding rather than silently passed; a source with no
|
|
142
|
+
discounted block, or an inline-handler write-back with nothing else in play, is unaffected. Second: the
|
|
143
|
+
write-back hint check (and the earlier candidacy check) now also recognizes optional-call spellings —
|
|
144
|
+
`fetch?.(`, `xhr?.open?.(`, `$.post?.(`, `navigator?.sendBeacon?.(` — so a source whose only write-back
|
|
145
|
+
uses `?.` is neither missed as a candidate nor discounted as prose inside an unparseable block; that
|
|
146
|
+
optional-call matching is also linearized (no more quadratic backtracking on a long non-matching
|
|
147
|
+
whitespace run). Third: a member-spelled write-back inside a block that DOES parse —
|
|
148
|
+
`window.fetch(...)`, `globalThis.fetch(...)`, `self.fetch(...)`, or the same spelling inside a same-file
|
|
149
|
+
fetch-wrapper's own body — is now classified the same as a bare `fetch(...)` call instead of going
|
|
150
|
+
unrecognized and falling through as clean; a bare `sendBeacon(...)` identifier call (e.g. a locally
|
|
151
|
+
bound alias of `navigator.sendBeacon`) is now recognized the same way as the member-spelled
|
|
152
|
+
`navigator.sendBeacon(...)`. Fourth: a `.post(...)`/`.put(...)`/`.patch(...)` call on a receiver outside
|
|
153
|
+
the known `axios`/`$`/`jQuery` set — the common miss being an axios instance,
|
|
154
|
+
`const api = axios.create(); api.post("/api/save", data)` — targeting a relative URL is no longer
|
|
155
|
+
invisible; it is now reported as an advisory finding (never escalated to an error, since the receiver
|
|
156
|
+
isn't provably a write-back client and could be unrelated code; `.delete(...)` is deliberately excluded
|
|
157
|
+
from this, since it's common on non-HTTP collection types). Fifth: the axios/`$`/`jQuery` verb set
|
|
158
|
+
recognized on the whitelisted identifier itself is widened from `.post(...)` alone to
|
|
159
|
+
`.post(...)`/`.put(...)`/`.patch(...)`/`.delete(...)`/`.postForm(...)`/`.putForm(...)`/`.patchForm(...)`
|
|
160
|
+
(the last three are axios v1's multipart form-data verb aliases; `.delete(...)` is INCLUDED here,
|
|
161
|
+
unlike the any-receiver advisory case above, since the literal `axios`/`$`/`jQuery` identifier has no
|
|
162
|
+
ambiguity about what it means); a bare config-object call — `axios({method:"POST", url:"/api/save",
|
|
163
|
+
...})` — and `axios.request({...})` are both recognized, whether the config argument is an inline
|
|
164
|
+
object literal or a hoisted identifier (`const cfg = {...}; axios(cfg)`). This is not exhaustive: axios's
|
|
165
|
+
alternate `$.ajax({...})`-style config-key vocabulary, and any computed-member, whitespace-separated, or
|
|
166
|
+
aliased/re-exported spelling, remain a documented, lexically/structurally invisible accepted class (see
|
|
167
|
+
the relevant doc comments in `analyze-artifact.ts`) — as does a `formaction`/`formmethod` override on a
|
|
168
|
+
submit button that redirects an otherwise-remote `<form>` back to a relative, in-scope URL.
|
|
169
|
+
- **`analyze-skill`: a delete/remove flow that claims success now classifies as a lost write-back (error),
|
|
170
|
+
not just a suspect (advisory).** The success-claim vocabulary that distinguishes a lost write-back (an
|
|
171
|
+
unconditional "it worked" toast) from a merely-suspect one gained delete-flow words (`deleted`,
|
|
172
|
+
`removed`, …) alongside the existing `saved`/`submitted`/`persisted`/`completed`/`success`. A relative
|
|
173
|
+
`DELETE` write-back whose only success signal is a "Deleted!"/"Removed!" toast — no `resp.ok`/status
|
|
174
|
+
check — is now flagged at error severity like its save-flow equivalent, since under Cowork it resolves
|
|
175
|
+
non-ok against Cowork's own origin and the false confirmation is identical.
|
|
176
|
+
|
|
177
|
+
### Documentation
|
|
178
|
+
|
|
179
|
+
- **Documented an `analyze-skill --runtime` recipe for agent-generated artifacts.**
|
|
180
|
+
`analyze-skill --runtime <run-dir>/work/session/mnt/outputs` confirms interactive-artifact write-backs in
|
|
181
|
+
HTML the agent *generates during a run* — content the source-only static scan can't see until a run has
|
|
182
|
+
happened. Notes the tier-specific output paths and the microvm/replay caveats.
|
|
183
|
+
|
|
184
|
+
## [1.1.0] — 2026-07-16
|
|
185
|
+
|
|
186
|
+
Minor: `analyze-skill` gains interactive-artifact write-back detection (static + an optional `--runtime`
|
|
187
|
+
headless-DOM confirmer), `lint` gains a container-only-key tier check, and the `doctor` JSON envelope is
|
|
188
|
+
frozen as a covered SPEC §12 surface. All additive.
|
|
189
|
+
|
|
190
|
+
### Added
|
|
191
|
+
|
|
192
|
+
- **`analyze-skill` now detects interactive-artifact write-backs lost under Cowork.** Alongside the
|
|
193
|
+
existing `/sessions` path scan, it statically analyzes `.html/.htm/.js/.mjs/.ts/.jsx/.tsx/.py` sources
|
|
194
|
+
under the target for a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back that silently
|
|
195
|
+
fails under Cowork (the artifact is served from Cowork's own origin, so a relative write-back resolves
|
|
196
|
+
non-ok and a page that doesn't check `resp.ok` shows a false "Saved"). Findings: `artifact-write-back-lost`
|
|
197
|
+
(error — gates under `--strict`), `artifact-write-back-suspect` (advisory), and a separate top-level
|
|
198
|
+
`analysisFailures` **could-not-verify** channel (a candidate that couldn't be parsed/analyzed) that
|
|
199
|
+
always exits `3`, `--strict`-independent. A guard that isn't statically provable-truthy is `suspect`,
|
|
200
|
+
never silently clean; the blanket `analyze-skill: ignore` marker does not silence artifact rules.
|
|
201
|
+
Each `SkillFinding` now carries a `severity` (`error|advisory`); `--strict` gates on any error finding.
|
|
202
|
+
- **`analyze-skill --runtime`** — an optional headless-DOM confirmation that drives a materialized `.html`
|
|
203
|
+
artifact in jsdom (stubbed network + synthetic user actions, run twice) to *observe* whether a relative
|
|
204
|
+
write-back fires and is lost. Enrichment only (never changes the exit code); trusted-source scope. `jsdom`
|
|
205
|
+
is an optional, dynamically-imported dependency — absent it reports "run `npm i jsdom` to enable".
|
|
206
|
+
- **`schema/doctor.json`** — the `doctor --output-format json` envelope is now a covered SPEC §12 surface
|
|
207
|
+
(`oneOf` the completed-probe shape and the shared error envelope for every category). `doctor`'s normal
|
|
208
|
+
JSON output is standardized through the shared envelope frame.
|
|
209
|
+
- **`lint` flags container-only assertion keys off-container** — `no_scratchpad_leak`/`present_files_called`
|
|
210
|
+
on `fidelity: protocol|microvm|hostloop` is an ERROR, on `fidelity: cowork` a WARN, clean on `container`.
|
|
211
|
+
|
|
9
212
|
## [1.0.6] — 2026-07-15
|
|
10
213
|
|
|
11
214
|
Patch: platform baseline synced to Claude Desktop `1.21459.0`. The spawn contract and rendered system
|