cowork-harness 1.2.0 → 1.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +68 -13
- package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +8 -5
- package/.claude/skills/cowork-harness/references/scenario-schema.md +7 -6
- package/.claude/skills/cowork-harness/references/task-recipes.md +27 -0
- package/CHANGELOG.md +131 -0
- package/README.md +19 -17
- package/SPEC.md +33 -5
- package/baselines/desktop-1.22209.3.json +392 -0
- package/dist/agent/session.js +77 -8
- package/dist/assert.js +125 -37
- package/dist/cli.js +95 -37
- package/dist/decide/decider.js +52 -6
- package/dist/decide/external-channel.js +19 -6
- package/dist/hostloop/webfetch-dedup.js +50 -0
- package/dist/hostloop/workspace-handler.js +42 -4
- package/dist/loop-decision.js +30 -0
- package/dist/run/analyze-artifact-runtime.js +309 -65
- package/dist/run/analyze-artifact.js +581 -206
- package/dist/run/analyze-skill.js +182 -37
- package/dist/run/artifacts.js +95 -19
- package/dist/run/cassette.js +37 -1
- package/dist/run/chat-result.js +3 -0
- package/dist/run/chat.js +10 -1
- package/dist/run/diff.js +2 -0
- package/dist/run/execute.js +31 -8
- package/dist/run/inspect-view.js +19 -0
- package/dist/run/pre-run-manifest.js +39 -8
- package/dist/run/renderer.js +11 -0
- package/dist/run/run-index.js +16 -5
- package/dist/run/run-status.js +1 -0
- package/dist/run/run.js +49 -7
- package/dist/run/trace-view.js +27 -7
- package/dist/run/verdict.js +19 -0
- package/dist/runtime/hostloop.js +1 -0
- package/dist/types.js +2 -3
- package/docs/README.md +5 -1
- package/docs/cassette.md +9 -2
- package/docs/debugging.md +34 -0
- package/docs/decisions/README.md +7 -0
- package/docs/gotchas.md +9 -0
- package/docs/maintenance.md +1 -1
- package/docs/scenario.md +55 -4
- package/docs/stats.md +5 -1
- package/docs/subagents.md +7 -3
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/llms.txt +1 -0
- package/package.json +1 -1
- package/python/README.md +13 -4
- package/schema/run-result.json +19 -1
- package/schema/scenario.schema.json +2 -2
- package/scripts/check-baseline-staleness.ts +147 -0
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: cowork-harness
|
|
3
|
-
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
3
|
+
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.4.0
|
|
7
|
+
tracks-harness: cowork-harness 1.4.0 (baseline desktop-1.22209.3)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
26
|
-
> `desktop-1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.4.0` (baseline
|
|
26
|
+
> `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.4.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.4.0"`. **Pin `@>=1.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
|
-
What the ≥ 1.
|
|
44
|
+
What the ≥ 1.4.0 floor gates, by release:
|
|
45
45
|
|
|
46
46
|
- **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
|
|
47
47
|
- **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
|
|
@@ -54,6 +54,8 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
54
54
|
- **1.0.0:** first stable release — the SPEC §12 compatibility contract takes effect (covered CLI/schema/env/Action surfaces are now stable; breaking changes need a major bump). No new author-facing command; the floor simply tracks the 1.0 release.
|
|
55
55
|
- **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
|
|
56
56
|
- **1.2.0:** three new assertion keys — `no_lost_write_back: true` (the write-back detector above wired as a per-scenario gate over the run's authored files, live-only) and the regex siblings `tool_result_matches`/`tool_result_not_matches` (case-insensitive per-result, for an error-signature *family* a literal substring can't express). **`microvm` outputs are now observable** — its session tree is snapshotted from the VM into the run dir, so `file_exists`/`artifact_json`/`user_visible_artifact`/`no_unexpected_files`/`input_unmodified`/`no_lost_write_back`/`semantic_matches` all work there, no longer `container`/`hostloop`-only. `status <dir>` also resolves the newest session under a `--run-dir` root; the completion footer prints a `→ result: …/result.json` pointer; the `on_unanswered=fail` error also points to `on_unanswered: llm`. `analyze-skill` hardening: a phantom `<script>` prose block no longer sinks a real verdict, a delete/remove flow claiming success classifies as lost (error) not just suspect, and write-back detection widened (optional-call `?.` spellings, member-spelled/aliased `fetch`/`sendBeacon`, axios instance/config forms).
|
|
57
|
+
- **1.3.0:** run-identity for the iterate-across-fixes loop — `skill`/`run` take `--label <tag>` (a generation tag surfaced in `result.json` `runLabel`, the run-index row, `inspect`, and `status.json`) alongside an auto-recorded `skillCommit`, layered on the **authoritative** content-exact `fingerprint.skillHash` a harvest step should pair critiques by (`inspect`/index surface a short `skillHash` prefix). `trace --full-results` captures the full input+result of **every** tool call — successful ones too, not just errors — so an external grader can ground a self-critique finding against the call it cites. `verify-run` now **warns** on skillHash drift for answer-less scenarios (was silent unless the scenario declared scripted `answers`). Plus `skill --allow-missing-capability` (the open-ended-run opt-out for a capability FALSE-NEGATIVE on the lean `core` image), a new **warn-severity `ended_with_question`** verdict signal (the agent's final answer contains a question and the run wrote no `outputs/` deliverable — the lenient sibling of the strict `stalled`), and LLM-decider `OTHER:` free-text answers now marked `[via Other free-text]` in gate provenance.
|
|
58
|
+
- **1.4.0:** platform baseline synced to Desktop **1.22209.3** (agent 2.1.215; no prompt/spawn/egress drift), and **`coworkWebFetchDedup` is now enacted** on the hostloop `web_fetch` path — a repeat fetch of the same normalized URL within the TTL (15 min, cap 100) makes **no network request** and returns a marker to re-use the earlier result, with no egress event. Baseline-gated (Desktop ≥ 1.22209.3), keyed under the request + terminal `destination_url`, never caching errors/empty; a hit is assertable via `tool_result_contains: "Already fetched"`.
|
|
57
59
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
58
60
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
59
61
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
@@ -311,8 +313,10 @@ your `answers:` against that kept run, then record once. **But the kept run is a
|
|
|
311
313
|
skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
|
|
312
314
|
current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
|
|
313
315
|
deliberately. (Same fail-closed family: corrupt gate evidence — unparseable `events.jsonl` lines, or fewer
|
|
314
|
-
gates than `trace.json` recorded questions —
|
|
315
|
-
|
|
316
|
+
gates than `trace.json` recorded questions — a structurally invalid `result.json`, a `command:"replay"`
|
|
317
|
+
result (a replay is a re-check of a recorded cassette, not run evidence — verify the original live run dir),
|
|
318
|
+
and a `mode:"chat"` result (chat carries no assertions or verdict by contract) also refuse rather than
|
|
319
|
+
certify.) (A token-free probe of "which gates fire" isn't possible — gates are model-decided per run.)
|
|
316
320
|
|
|
317
321
|
Run artifacts are written to `~/.cowork-harness/runs/…` by default — **outside any working tree**, so a run
|
|
318
322
|
launched from a repo root never drops sensitive skill inputs/outputs into it. Pass `--run-dir <path>` (or set
|
|
@@ -354,6 +358,15 @@ cassette — has its own recipe:
|
|
|
354
358
|
unless the scenario asserts `allow_missing_capability: true`, which downgrades it to a notice and
|
|
355
359
|
proceeds. Rebuild with `--build-arg COWORK_FULL_PARITY=1` and point `COWORK_AGENT_IMAGE` at it for those
|
|
356
360
|
skills.
|
|
361
|
+
6. **Iterate across fixes — verify before you trust, and don't cross-pair generations.** A green run is
|
|
362
|
+
not a correct run, and a skill's self-reported finding is not real until its cited evidence is found in
|
|
363
|
+
the run's own output. Ground each finding against `result.json` (`finalMessage` = the skill's own
|
|
364
|
+
answer/critique; `toolResults` = tool outputs) and the tool-call stream via
|
|
365
|
+
`cowork-harness trace <run-dir> --output-format json` — add `--full-results` so a successful call's full
|
|
366
|
+
input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
|
|
367
|
+
pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
|
|
368
|
+
produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
|
|
369
|
+
kept run predates the current skill). See `docs/debugging.md` (repo-only) for the full loop.
|
|
357
370
|
|
|
358
371
|
#### Interpreting verdict signals
|
|
359
372
|
|
|
@@ -364,6 +377,42 @@ The run verdict may include `WARN`-severity signals in addition to pass/fail. On
|
|
|
364
377
|
the run can still green. If you see it, fix the asset path — a green with a missing asset is
|
|
365
378
|
not a valid pass.
|
|
366
379
|
|
|
380
|
+
**False negatives — signals that are tier/image artifacts, not skill defects.** Some fail-severity
|
|
381
|
+
signals read like a skill gap but are really a property of the reduced test image or the fidelity tier.
|
|
382
|
+
Recognize these before "fixing" a non-bug:
|
|
383
|
+
|
|
384
|
+
- **`missing_capability`** — the lean `core` agent image is a deliberate partial mirror of real Cowork's
|
|
385
|
+
rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
|
|
386
|
+
`markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
|
|
387
|
+
(`magick`) can trip this even though real Cowork **ships** those. The message says so ("likely a FALSE
|
|
388
|
+
NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
|
|
389
|
+
`COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
|
|
390
|
+
`allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
|
|
391
|
+
or a declared `requires_capabilities` the tier can't provide, both lanes — an unknown family name
|
|
392
|
+
hard-fails rather than silently passing.) **On an open-ended `skill` run** (no `assert:` block to carry
|
|
393
|
+
the modifier), pass **`--allow-missing-capability`** — the CLI equivalent of the assertion.
|
|
394
|
+
- **`ended_with_question`** (`WARN`, live lane) — a heuristic: the agent's final answer contains a
|
|
395
|
+
question and the run wrote **no deliverable to `outputs/`** — it may have ended on a request for input
|
|
396
|
+
instead of finishing. Warn-only; the fix is scripting/steering the answer (`answer:` / `--answer` / a
|
|
397
|
+
decider, or `--decider-llm --intent`), not editing the skill's prose. The strict, fail-severity sibling
|
|
398
|
+
`stalled` already catches a *trailing*-`?` final turn that did no tool work after the last gate; this
|
|
399
|
+
covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
|
|
400
|
+
the final message before acting — a legitimate question-posing answer that wrote a file never fires.
|
|
401
|
+
Assert `allow_stall: true` if ending on a question is the intended terminal state.
|
|
402
|
+
- **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
|
|
403
|
+
paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
|
|
404
|
+
`container`/`microvm`, but only *fires* on an actual scanned leak with no authored
|
|
405
|
+
`transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
|
|
406
|
+
run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
|
|
407
|
+
it's valid.
|
|
408
|
+
- **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
|
|
409
|
+
`RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
|
|
410
|
+
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
411
|
+
|
|
412
|
+
The full 14-code signal table (severity + per-signal opt-out) is in
|
|
413
|
+
[`references/scenario-schema.md`](./references/scenario-schema.md); `docs/scenario.md` (repo-only) carries
|
|
414
|
+
the fuller narrative.
|
|
415
|
+
|
|
367
416
|
### Checking whether a background run is alive
|
|
368
417
|
|
|
369
418
|
Never use `ps aux` to check on a `cowork-harness` run you launched in the background — it only sees
|
|
@@ -472,7 +521,11 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
472
521
|
- **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
|
|
473
522
|
`tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input, with a
|
|
474
523
|
`trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
|
|
475
|
-
**Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field.
|
|
524
|
+
**Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
|
|
525
|
+
**absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
|
|
526
|
+
collection) is evidence **UNAVAILABLE**, not an empty run — `trace --view files` reports a loud
|
|
527
|
+
UNAVAILABLE marker (`workspaceFilesRecorded: false` in JSON, and no phantom "removed" diff rows) and
|
|
528
|
+
`inspect` prints `artifacts: UNAVAILABLE` (`artifactsRecorded: false`) instead of `artifacts (0):`.
|
|
476
529
|
|
|
477
530
|
### Debugging with `chat`
|
|
478
531
|
|
|
@@ -615,9 +668,11 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
615
668
|
- `on_unanswered` governs **unanswered** `AskUserQuestion` gates; the `stalled` signal covers
|
|
616
669
|
stalling *after* one is answered — two different failure modes.
|
|
617
670
|
- **Free-text aside:** a "type-it-in-notes" option has **no scripted deterministic answer** today
|
|
618
|
-
(the `OTHER:` directive works only on the LLM-decider path, not scripted `choose
|
|
619
|
-
|
|
620
|
-
|
|
671
|
+
(the `OTHER:` directive works only on the LLM-decider path, not scripted `choose:`, and only on
|
|
672
|
+
**single-select** gates — a **multi-select** gate is index-only, so `OTHER:` fails loud there; on an
|
|
673
|
+
options-bearing single-select gate a bare out-of-set LLM answer also fails loud (exit 2) — see the
|
|
674
|
+
LLM-decider free-text note in `references/fidelity-and-answers.md`). An LLM decision answered via
|
|
675
|
+
`OTHER:` is marked `[via Other free-text]` in its `gateProvenance` rationale.
|
|
621
676
|
14. **A positional `choose` (`first` / index) is order-dependent.** `choose: "2"` survives label drift
|
|
622
677
|
but NOT option *re-ordering* — if the gate presents its options in a different order run-to-run, the
|
|
623
678
|
index lands on a different option (a silent re-record flake). Prefer an exact label when order is
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.4.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.4.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -32,7 +32,7 @@ jobs:
|
|
|
32
32
|
- uses: actions/checkout@v4
|
|
33
33
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
34
34
|
run: |
|
|
35
|
-
V=2.1.
|
|
35
|
+
V=2.1.215 # match your scenario's pinned baseline's agentVersion
|
|
36
36
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
37
37
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
38
38
|
# verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.4.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.4.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.4.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.4.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -164,10 +164,13 @@ Caveat: `decide` only builds a **single-select** sample (set choices with `--opt
|
|
|
164
164
|
multiSelect flag), so its printed request shows `options[].label` but never `multiSelect:true` — to
|
|
165
165
|
exercise the array reply path, run a real multiSelect gate or unit-test the helper directly.
|
|
166
166
|
|
|
167
|
-
LLM-decider free-text goes via `OTHER: <value>` on an **options-bearing** gate; a bare
|
|
168
|
-
answer (no matching label, no `OTHER:`) fails loud (`UnansweredError` → exit 2) — it never
|
|
169
|
-
guesses an option.
|
|
170
|
-
|
|
167
|
+
LLM-decider free-text goes via `OTHER: <value>` on an **options-bearing single-select** gate; a bare
|
|
168
|
+
out-of-set answer (no matching label, no `OTHER:`) fails loud (`UnansweredError` → exit 2) — it never
|
|
169
|
+
stalls or guesses an option. A **multi-select** gate is **index-only**: it accepts comma-separated option
|
|
170
|
+
numbers, and `OTHER:` fails loud there (no free-text escape on that path). Open-ended (no-option) gates
|
|
171
|
+
need no `OTHER:` prefix: free text is delivered verbatim. A decision answered via any free-text path is
|
|
172
|
+
marked `[via Other free-text]` in its `gateProvenance` rationale, so a `result.json` consumer can tell it
|
|
173
|
+
from an offered-option pick. (Scripted scenarios use the separate `answer:` escape hatch.)
|
|
171
174
|
|
|
172
175
|
A gate that fails loud (`on_unanswered: fail`, the default) still **salvages a PARTIAL run**: the harness
|
|
173
176
|
writes a `result.json` (marked `partial: true`) with the artifacts the agent produced before the whiff, so
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.4.0`
|
|
4
4
|
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
@@ -252,7 +252,7 @@ same set live from the schema.
|
|
|
252
252
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
253
253
|
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
|
|
254
254
|
| `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning |
|
|
255
|
-
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
|
|
255
|
+
| `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
|
|
256
256
|
| `self_heal_ran: <bool>` | a plugin-root self-heal script was (not) invoked |
|
|
257
257
|
| `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
|
|
258
258
|
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
|
|
@@ -334,7 +334,7 @@ dotted path.
|
|
|
334
334
|
|
|
335
335
|
**VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; most
|
|
336
336
|
are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
|
|
337
|
-
`"success"`) and only
|
|
337
|
+
`"success"`) and only four are **warn**-severity (informational, never flip pass/fail). Current signal
|
|
338
338
|
codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
339
339
|
|
|
340
340
|
| Code | Severity | Meaning |
|
|
@@ -347,16 +347,17 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
347
347
|
| `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
|
|
348
348
|
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
349
349
|
| `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
|
|
350
|
-
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`) |
|
|
350
|
+
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
351
351
|
| `infra_error` | fail | A VM/egress sidecar crashed mid-run — not author-suppressible |
|
|
352
|
-
| `stalled` | fail | The run ended on an unanswered question (opt out: `allow_stall`) |
|
|
352
|
+
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
353
353
|
| `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
|
|
354
354
|
| `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
|
|
355
355
|
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
|
|
356
|
+
| `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
|
|
356
357
|
|
|
357
358
|
A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
|
|
358
359
|
overall run verdict and exit code — `assert result: success` alone won't catch it; check
|
|
359
|
-
`result.verdict.signals[].severity` or the run's exit code. Only the
|
|
360
|
+
`result.verdict.signals[].severity` or the run's exit code. Only the four **warn** codes are truly benign.
|
|
360
361
|
|
|
361
362
|
## Replay class
|
|
362
363
|
|
|
@@ -163,3 +163,30 @@ degrade the advice. It is real work to calibrate; these steps are the traps that
|
|
|
163
163
|
**Lane note:** `semantic_matches` is **live-only** (the judge is a live model call), so these scenarios
|
|
164
164
|
run on the `run` lane, never token-free `replay` — the linter's "all assertions live-only" warning is
|
|
165
165
|
expected and correct here.
|
|
166
|
+
|
|
167
|
+
## Recipe 6 — Iterate a skill across fixes (ground findings, don't cross-pair generations)
|
|
168
|
+
|
|
169
|
+
Hardening a skill is a loop: run → read what it did → fix → run again. Two disciplines keep it honest.
|
|
170
|
+
|
|
171
|
+
1. **Verify before you trust.** A green run is not a correct run, and a skill's self-reported finding (a
|
|
172
|
+
self-critique appendix, "I extracted X") is not real until its cited evidence is found in the run's own
|
|
173
|
+
output. The harness emits the substrate; the grader is yours (it lives outside the harness):
|
|
174
|
+
- `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
|
|
175
|
+
- `cowork-harness trace <run-dir> --output-format json` → the tool-call stream. Add `--full-results` so
|
|
176
|
+
a **successful** call's full input + result are captured (the default view slices them to ~100/120
|
|
177
|
+
chars) — this is what lets your grader confirm "the skill claims it read X and derived Y" against the
|
|
178
|
+
actual call.
|
|
179
|
+
- `cowork-harness inspect <run-dir>` → what the run produced, plus the run's `label` and `skillHash`.
|
|
180
|
+
- In-run alternative: dispatch a checker **sub-agent** (maker/checker) whose result folds into the
|
|
181
|
+
verdict.
|
|
182
|
+
2. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
|
|
183
|
+
`result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
|
|
184
|
+
content-exact, on every live run, changes on any tracked edit. **Group/pair on it** (`inspect` and the
|
|
185
|
+
run-index row surface a short prefix). Add `--label <tag>` for a human-readable generation name
|
|
186
|
+
(skillHash is the correctness key; the label is ergonomics). `cowork-harness verify-run <run-dir>
|
|
187
|
+
<scenario.yaml>` is the native staleness guard: it **warns** when a kept run predates the current
|
|
188
|
+
skill, and with scripted `answers` **hard-fails** rather than vouch for a stale gate snapshot.
|
|
189
|
+
|
|
190
|
+
**Lane note:** the exploratory driver is `skill <dir> --decider-llm --intent "<what this run tests>"`,
|
|
191
|
+
which is flagged non-deterministic (a green here is exploration, not a scripted pass) — pin the
|
|
192
|
+
load-bearing gates with `--answer` once you know which fire.
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,137 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.4.0] — 2026-07-19
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`coworkWebFetchDedup` enacted (hostloop `web_fetch`).** Real Cowork keeps a per-session negative-work
|
|
14
|
+
cache: a repeat `web_fetch` of the same normalized URL within a TTL (default 15 min; cap 100; FIFO
|
|
15
|
+
eviction; a hit does not refresh recency) makes **no network request** and returns a marker telling the
|
|
16
|
+
model to re-use the earlier result. The harness now reproduces this on the host-API (`coworkWebFetchViaApi`)
|
|
17
|
+
path — **baseline-gated** (only when the resolved baseline's `coworkWebFetchDedup` gate is on, i.e. Desktop
|
|
18
|
+
≥ 1.22209.3), keyed under both the request URL and the terminal `destination_url`, never caching errors /
|
|
19
|
+
empty / non-2xx responses, and emitting **no egress event** on a hit (matching production's zero-network
|
|
20
|
+
dedup). A hit is observable via the marker text (`tool_result_contains: "Already fetched"`).
|
|
21
|
+
|
|
22
|
+
### Changed
|
|
23
|
+
|
|
24
|
+
- **Platform baseline synced to Desktop 1.22209.3** (agent `2.1.215`). No prompt / spawn-env / egress-allowlist
|
|
25
|
+
drift vs 1.21459.0; the sync captured the new `coworkWebFetchDedup` runtime config (enacted above) plus a
|
|
26
|
+
few new (off) GrowthBook gates. The skill/README/reference version floors and agent-binary pins track the
|
|
27
|
+
new baseline.
|
|
28
|
+
|
|
29
|
+
## [1.3.0] — 2026-07-19
|
|
30
|
+
|
|
31
|
+
### Added
|
|
32
|
+
|
|
33
|
+
- **`skill --allow-missing-capability`** — the open-ended-run equivalent of a scenario asserting
|
|
34
|
+
`allow_missing_capability: true`. An open-ended `skill` run has no `assert:` block to carry the opt-out,
|
|
35
|
+
so a self-flagged capability FALSE-NEGATIVE on the lean `core` image would hard-fail with no escape
|
|
36
|
+
hatch; the flag merges the modifier onto the synthesized success assertion, suppressing both the
|
|
37
|
+
post-run `missing_capability` fail and the pre-run capability abort.
|
|
38
|
+
- **`ended_with_question` verdict signal (WARN)** — a heuristic run-level classifier: the agent's final
|
|
39
|
+
answer contains a question and the run wrote no deliverable to `outputs/` — a likely conversational
|
|
40
|
+
dead-end that still exited `result:"success"`. The lenient warn-severity sibling of the strict, fail
|
|
41
|
+
`stalled` (which already catches a trailing-`?` final turn with no post-gate tool work); this covers the
|
|
42
|
+
residual — a mid-message `?`, or tool work after the last gate that still ended asking. Never flips a
|
|
43
|
+
verdict; `verify-run` over historical results may surface it on matching old runs (benign).
|
|
44
|
+
- **`--decider-llm` "Other" answers are marked in gate provenance** — a decision answered via the
|
|
45
|
+
`OTHER:` free-text path (or a no-option free-text gate) now carries `[via Other free-text]` in its
|
|
46
|
+
`rationale`, so a `result.json` consumer can distinguish it from an offered-option pick. (Multi-select
|
|
47
|
+
gates remain index-only — `OTHER:` is rejected there, now pinned by a test.)
|
|
48
|
+
|
|
49
|
+
- **Run-identity metadata for the iterate-across-fixes loop.** `skill`/`run` accept `--label <tag>`, a
|
|
50
|
+
human-readable generation tag surfaced in `result.json` (`runLabel`), the run-index row, `inspect`, and
|
|
51
|
+
`status.json`. Each live run also records `skillCommit` — best-effort git `HEAD` of the session's skill
|
|
52
|
+
source dirs (commit provenance; `null` when the dirs span >1 repo or aren't a git work tree). These are
|
|
53
|
+
ergonomics on top of the **authoritative** content-exact version key `fingerprint.skillHash` (already
|
|
54
|
+
recorded on every run): a harvest step should group/pair a critique against a matching `skillHash`.
|
|
55
|
+
`inspect` and the run-index row now surface a short `skillHash` prefix so a pairing check needs no
|
|
56
|
+
`result.json` open. (Chat runs carry no `skillHash` and take no `--label`.)
|
|
57
|
+
- **`trace --full-results`** captures the FULL input + result of every tool call — successful ones too,
|
|
58
|
+
not just errors (`resultTextFull`/`detailFull`, 4 KB cap) — so an external grader can ground a
|
|
59
|
+
self-critique finding against the call it cites. The default view keeps its 100/120-char slices, so
|
|
60
|
+
existing JSON consumers are unaffected.
|
|
61
|
+
- **`verify-run` now warns on skill drift for answer-less scenarios.** Previously the skillHash-drift
|
|
62
|
+
check ran only when a scenario declared scripted `answers` (a hard fail). An answer-less `verify-run`
|
|
63
|
+
now emits a `::warning::` ("the kept run predates the current skill … findings describe an older skill
|
|
64
|
+
version") instead of staying silent — a WARN, not a fail, since re-asserting a new `assert:` block
|
|
65
|
+
against a frozen run dir is legitimate.
|
|
66
|
+
|
|
67
|
+
### Fixed
|
|
68
|
+
|
|
69
|
+
- **Filesystem-evidence assertions no longer pass on incomplete evidence.** `input_unmodified` now fails
|
|
70
|
+
loud when its glob matches **no** pre-run path (a typo or renamed mount was a silent vacuous pass), and
|
|
71
|
+
fails evidence-unavailable when a manifest path escapes the workspace root or a matched file exceeds the
|
|
72
|
+
hash cap. The pre-run baseline records its provenance, so `no_unexpected_files` / `input_unmodified` fail
|
|
73
|
+
evidence-unavailable on an unreadable connected-folder baseline instead of diffing a partial tree
|
|
74
|
+
(`RunResult` gains `preRunOrigin`). `no_lost_write_back` no longer silently misses a modified-but-unreadable
|
|
75
|
+
or over-cap file, an authored file under an unreadable subtree, or a scratchpad deliverable behind a
|
|
76
|
+
symlink/hardlink — each now surfaces as could-not-verify.
|
|
77
|
+
- **Run-dir consumers no longer read absent evidence as empty, or a replay re-check as run evidence.**
|
|
78
|
+
`verify-run` now refuses (exit `2`, "can't verify ⇒ not green") a run dir whose `result.json` was produced
|
|
79
|
+
by `replay` (`command:"replay"` — a re-check of a recorded cassette, not run evidence) or by `chat`
|
|
80
|
+
(`mode:"chat"` — no assertions or verdict by contract); the refusal keys on `command`/`mode`, never on
|
|
81
|
+
`workspaceFiles`, so a live run merely lacking optional evidence fields still verifies. `stats --reindex`
|
|
82
|
+
skips a stray `command:"replay"` `result.json` instead of stripping the label and relabeling it `"run"`,
|
|
83
|
+
and its report line separates `skipped — replay re-check, not evidence` from `skipped — missing/corrupt
|
|
84
|
+
result.json`. `trace --view files` reports workspace-file evidence **UNAVAILABLE**
|
|
85
|
+
(`workspaceFilesRecorded: false` in JSON) when `workspaceFiles` is absent — distinct from a run that
|
|
86
|
+
genuinely wrote nothing — and no longer emits phantom "removed" diff rows against a persisted
|
|
87
|
+
`preRunHashes` in that case. `inspect` likewise prints `artifacts: UNAVAILABLE`
|
|
88
|
+
(`artifactsRecorded: false`) instead of `artifacts (0):` when `result.artifacts` is absent.
|
|
89
|
+
- **Static artifact analysis (`analyze-skill` / `no_lost_write_back` Tier A)** closes several false-green and
|
|
90
|
+
false-positive holes: recognizes `axios`/jQuery and library write-backs, computed member calls
|
|
91
|
+
(`xhr["open"]`), unquoted and submitter-overridden `<form>` actions, and ES-module sources; classifies URLs
|
|
92
|
+
with the WHATWG parser so `localhost.evil.com`, protocol-relative `//host`, and `mailto:`/`data:` are no
|
|
93
|
+
longer misread as local; scope-aware constant folding and member-mutation tracking stop live code being
|
|
94
|
+
proven dead; an unresolved URL or method (including spread request options) yields could-not-verify instead
|
|
95
|
+
of a silent clean; and every write-back in a file is reported, not just the first. The advertised
|
|
96
|
+
`.ts/.tsx/.jsx` source extensions are dropped (no compatible parser) — treated as out-of-scope rather than
|
|
97
|
+
parse-noise.
|
|
98
|
+
- **Runtime artifact confirmation (Tier B)** now confirms edit-fired autosaves and load-time write-backs (not
|
|
99
|
+
only explicit commits), handles `fetch(Request)`, models browser-faithful XHR response/listener semantics,
|
|
100
|
+
skips disabled/hidden controls and honors a submitter's `formaction`, analyzes observed writes before
|
|
101
|
+
downgrading on an external script, matches loopback hosts exactly, and no longer swallows unrelated harness
|
|
102
|
+
exceptions.
|
|
103
|
+
- **`analyze-skill` orchestration** fails could-not-verify (exit `3`) when a `references`/`agents`/`commands`/
|
|
104
|
+
`skills` subtree is unreadable (was a silent clean); an ignore marker inside a fenced or block-quoted
|
|
105
|
+
example no longer suppresses real findings; a mixed invocation fails on a positional that resolves to no
|
|
106
|
+
scannable source; JSON coverage lists clean artifact sources; and runtime mode honors the 3 MB read cap.
|
|
107
|
+
- **Protocol / decision handling fails closed on drift.** Unknown control-request subtypes and duplicate
|
|
108
|
+
outstanding request IDs are rejected as protocol errors; malformed user/tool-result blocks increment a new
|
|
109
|
+
`evidenceErrors.protocolMalformed` counter; the ABSTAIN permission/dialog fallback reconciles answer
|
|
110
|
+
delivery like a normal decision; `present_files` leak classification normalizes `..` paths; gate answers
|
|
111
|
+
require a request id; and an `allow_if` expression no longer fails to compile when a permission input key is
|
|
112
|
+
a reserved word.
|
|
113
|
+
- **CI / release gates.** The baseline-staleness check rejects non-finite, future-dated, and corrupt
|
|
114
|
+
timestamps (extracted to a unit-tested `scripts/check-baseline-staleness.ts`); manual container-image
|
|
115
|
+
publishes require a green CI run unless an explicit break-glass input is set; and the composite Action
|
|
116
|
+
passes inputs via the environment (closing a shell-injection surface) and accepts a JSON-array `extra-args`
|
|
117
|
+
that preserves quoting and spaces.
|
|
118
|
+
- **`analyze-skill` top-level `--help`** now documents the `--runtime` flag and exit code `3`
|
|
119
|
+
(could-not-verify); previously these appeared only in the per-command `analyze-skill --help`.
|
|
120
|
+
- Corrected a phantom assertion key in the `record --margins` documentation
|
|
121
|
+
(`max_tool_calls` → the real `tool_calls_max`).
|
|
122
|
+
|
|
123
|
+
### Documentation
|
|
124
|
+
|
|
125
|
+
- Documented previously-undocumented CLI flags: `record --force` (narrowly overrides the
|
|
126
|
+
different-scenario slug-collision overwrite refusal) and `record --decider-model`; `probe-dispatch`'s
|
|
127
|
+
inherited `--decider-cmd` / `--decider-dir` / `--on-unanswered` / `--ablate-skill`; and
|
|
128
|
+
`skill --timeout` / `--answer-policy`.
|
|
129
|
+
- Clarified the scenario-schema descriptions for `user_visible_artifact` (the assertion value is
|
|
130
|
+
workRoot-relative — e.g. `outputs/x.md`, not `mnt/`-prefixed) and `gate_answers_delivered`
|
|
131
|
+
(documented the `: false` confirmed-non-delivery inverse). Regenerated `schema/scenario.schema.json`.
|
|
132
|
+
- Qualified the stable JSON-envelope contract: SPEC §11 now maps commands to envelope families by
|
|
133
|
+
mechanism (`jsonEnvelope` / `jsonPayloadEnvelope` / dedicated), and README defers to it; clarified
|
|
134
|
+
that `replay` exit `2` is a whole-cassette operational failure, distinct from an in-cassette
|
|
135
|
+
malformation (which fails as an exit-`1` assertion).
|
|
136
|
+
- Added a `docs/decisions/` ADR index, `python/README.md` cross-links into the main doc spine, and
|
|
137
|
+
`RELEASING.md` to `llms.txt`; signposted the specialized `docs/*.md` guides and added `lint-skill` /
|
|
138
|
+
`analyze-skill` rows to the docs index.
|
|
139
|
+
|
|
9
140
|
## [1.2.0] — 2026-07-18
|
|
10
141
|
|
|
11
142
|
### Added
|