cowork-harness 3.2.1 → 3.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +13 -7
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -3
- package/.claude/skills/cowork-harness/references/task-recipes.md +5 -1
- package/CHANGELOG.md +337 -0
- package/DESIGN.md +2 -2
- package/README.md +4 -4
- package/SPEC.md +1 -1
- package/baselines/desktop-1.40609.1.json +5 -4
- package/baselines/desktop-1.44121.1.json +918 -0
- package/baselines/desktop-1.46388.3.json +942 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +7 -0
- package/baselines/prompts/desktop-1.46388.3/subagent-append-hl.md +37 -0
- package/dist/assert.js +271 -26
- package/dist/cli.js +30 -12
- package/dist/prompt/subagent-manifest.js +86 -0
- package/dist/prompt.js +22 -2
- package/dist/run/artifacts.js +80 -9
- package/dist/run/cassette.js +8 -0
- package/dist/run/chat.js +14 -0
- package/dist/run/execute.js +55 -12
- package/dist/session.js +4 -0
- package/dist/sync/cowork-sync.js +382 -30
- package/dist/types.js +35 -2
- package/docs/cassette.md +5 -1
- package/docs/ci.md +9 -3
- package/docs/cli.md +15 -4
- package/docs/companion-skill.md +3 -3
- package/docs/fidelity-gaps.md +19 -10
- package/docs/invariants.md +1 -1
- package/docs/maintenance.md +43 -15
- package/docs/scenario.md +4 -3
- package/docs/subagents.md +43 -0
- package/examples/data/manifest-probe/ledger.csv +5 -0
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +3 -3
- package/examples/replays/example-pdf-skill.cassette.json +97 -114
- package/examples/replays/hostloop-computer-links.cassette.json +3 -3
- package/examples/scenarios/subagent-manifest-probe.yaml +75 -0
- package/examples/sessions/subagent-manifest-probe.yaml +17 -0
- package/package.json +1 -1
- package/schema/cassette.v10.json +1 -1
- package/schema/cassette.v11.json +1 -1
- package/schema/cassette.v12.json +1 -1
- package/schema/run-result.json +1 -1
- package/schema/scenario.schema.json +9 -0
- package/scripts/check-versions.ts +52 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.4.0
|
|
7
|
+
tracks-harness: cowork-harness 3.4.0 (baseline desktop-1.46388.3)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
29
|
-
> `desktop-1.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.4.0` (baseline
|
|
29
|
+
> `desktop-1.46388.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.4.0"`. **Pin `@^3.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.4.0` (baseline `desktop-1.46388.3`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.4.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,12 +36,18 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.260 # match your scenario's pinned baseline's agentVersion
|
|
40
|
+
# The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
|
|
41
|
+
# served only from .../claude-code-releases/rc/<commit>/ — the stable path 404s for those, and
|
|
42
|
+
# 2.1.255 is one. Take B from your pinned baseline's agentBinary.releaseBaseUrl; baselines
|
|
43
|
+
# written before that field existed were stable-staged, so their base is the plain
|
|
44
|
+
# https://downloads.claude.ai/claude-code-releases.
|
|
45
|
+
B=https://downloads.claude.ai/claude-code-releases
|
|
40
46
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
41
47
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
42
48
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
43
49
|
EXPECTED=<paste agentBinary.sha256 for $V>
|
|
44
|
-
curl -fSL "
|
|
50
|
+
curl -fSL "$B/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
45
51
|
echo "$EXPECTED $RUNNER_TEMP/claude-$V" | sha256sum -c -
|
|
46
52
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
47
53
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
@@ -67,7 +73,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
73
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
74
|
|
|
69
75
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^3.
|
|
76
|
+
- run: npm i -g "cowork-harness@^3.4.0"
|
|
71
77
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
78
|
# no silent false-greens. WITHOUT --strict this
|
|
73
79
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -344,7 +350,7 @@ jobs:
|
|
|
344
350
|
with: { node-version: '24' }
|
|
345
351
|
- uses: actions/setup-python@v5
|
|
346
352
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
347
|
-
- run: npm i -g "cowork-harness@^3.
|
|
353
|
+
- run: npm i -g "cowork-harness@^3.4.0"
|
|
348
354
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
349
355
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
350
356
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -373,7 +379,7 @@ jobs:
|
|
|
373
379
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
374
380
|
fi
|
|
375
381
|
- if: steps.guard.outputs.live == 'true'
|
|
376
|
-
run: npm i -g "cowork-harness@^3.
|
|
382
|
+
run: npm i -g "cowork-harness@^3.4.0"
|
|
377
383
|
- if: steps.guard.outputs.live == 'true'
|
|
378
384
|
run: cowork-harness run scenarios/ --output-format json
|
|
379
385
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.4.0` (baseline `desktop-1.46388.3`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.4.0`
|
|
4
|
+
(baseline `desktop-1.46388.3`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -384,7 +384,7 @@ same set live from the schema.
|
|
|
384
384
|
| `egress_allowed: <host>` | the host was allowed through |
|
|
385
385
|
| `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
|
|
386
386
|
| `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
|
|
387
|
-
| `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
|
|
387
|
+
| `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, evidence_files?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) **`evidence_files: [globs]` scopes which authored files are graded** — reach for it the moment a run authors more than a couple of files. The capture budget (64 KiB total by default) is spent prefix-major then alphabetically, so a pipeline that stages intermediates (`outputs/_work/*.json`) exhausts it before reaching its own deliverable and the verdict is refused evidence-unavailable over files no rubric mentions. Scoping also makes the capture spend the budget on the named files FIRST and exempts them from the per-file cap. Paths are `<root>/<rel>` (`outputs/report.md`, never a bare `report.md`; session-root writes are `scratchpad/<rel>`); globs are `*`/`?`/`**`, not regex. A glob matching nothing FAILS and the message lists every authored path — read it rather than guessing. Still too big? Raise `$COWORK_HARNESS_AUTHORED_TOTAL_BYTES`. The typed reason is on `RunResult.assertions[].semanticEvidence` — check `.reason` instead of parsing the message |
|
|
388
388
|
| `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
|
|
389
389
|
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
390
390
|
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.4.0` (baseline `desktop-1.46388.3`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -169,6 +169,10 @@ degrade the advice. It is real work to calibrate; these steps are the traps that
|
|
|
169
169
|
rate. Assert tool use with the structural keys instead (`tool_called`, `present_files_called`,
|
|
170
170
|
`subagent_dispatched`, `hook_blocked`), and reserve `semantic_matches` for what the agent *said* or
|
|
171
171
|
*wrote*. For a fan-out skill whose real work happens in sub-agents, add `include_subagent_text: true`.
|
|
172
|
+
If the run authors more than a couple of files, add `evidence_files: ["outputs/report.md"]` naming the
|
|
173
|
+
deliverable — otherwise the 64 KiB capture budget is spent alphabetically and an unrelated intermediate
|
|
174
|
+
dropped at the cap refuses the whole verdict. Paths are `<root>/<rel>`; a glob that matches nothing
|
|
175
|
+
fails and prints the paths the run actually authored.
|
|
172
176
|
|
|
173
177
|
2. **Write DISCRIMINATING claims, and verify each against ground truth — not memory.** A claim that
|
|
174
178
|
contradicts how the tool actually behaves can *never* pass (the correct skill will contradict it), and
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,343 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.4.0] — 2026-09-05
|
|
10
|
+
|
|
11
|
+
### Upgrade notes
|
|
12
|
+
|
|
13
|
+
- **`promptAssetsHash` now also covers prompt text the harness GENERATES**, not only the committed
|
|
14
|
+
prompt-asset files — specifically the host-loop sub-agent folder manifest and the trailing skills
|
|
15
|
+
sentence added in this release. If you hold a cassette recorded with 3.3.0 against a baseline whose
|
|
16
|
+
`appVersion` still matches your live one, it will report `prompt-assets` staleness; re-record it. A
|
|
17
|
+
cassette on an older baseline already reports `baseline` staleness and is unaffected by this change.
|
|
18
|
+
- **A host-loop sub-agent is now told which folders exist and how to address them.** If you assert on a
|
|
19
|
+
sub-agent's prose, note it now has per-session path information it did not have before.
|
|
20
|
+
- **Sub-agent prompt assets are no longer the whole append.** If you maintain your own baseline,
|
|
21
|
+
`spawn.subagentAppendHostLoop` now points at the overridable SECTION only; the manifest and trailing
|
|
22
|
+
sentence are generated. Do not restore them into the asset — they would render twice.
|
|
23
|
+
|
|
24
|
+
### Verification
|
|
25
|
+
|
|
26
|
+
Verified **locally**, against the Desktop install this baseline was synced from — Desktop `1.46388.3`,
|
|
27
|
+
agent `2.1.260`, macOS arm64, agent image `cowork-agent-base:2` — on 2026-09-05:
|
|
28
|
+
|
|
29
|
+
| suite | result |
|
|
30
|
+
|---|---|
|
|
31
|
+
| `boundary-check` | **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress) |
|
|
32
|
+
| `npm run test:live` | 4 files, **19 passed, 0 skipped** |
|
|
33
|
+
| `run examples/scenarios/` | **7/7 success**, 39 assertions, 0 failed |
|
|
34
|
+
| e2e self-tests | **8/8 success** — askuserquestion, multiselect, multiselect-deciderdir, l1-container, l1-egress, present-files, semantic-evidence-files, canary-hostloop |
|
|
35
|
+
|
|
36
|
+
Tiers exercised: **`protocol`, `container`, `hostloop`**. **`microvm` was NOT exercised** (it needs a
|
|
37
|
+
real VM; CI excludes it for the same reason), so `smoke-l2-microvm` did not run.
|
|
38
|
+
|
|
39
|
+
`subagent-manifest-probe` is new here and is the first live coverage the sub-agent append has ever
|
|
40
|
+
had. It proves the composed append is *actionable* — the sub-agent shelled against
|
|
41
|
+
`/sessions/<id>/mnt/<name>/`, summed the same ledger through its file tools to the same total, and a
|
|
42
|
+
bare relative write landed at the host cwd — not that the paraphrase matches Desktop's wording, which
|
|
43
|
+
is what the `manifest`/`suffix` fingerprint axes are for.
|
|
44
|
+
|
|
45
|
+
### Parity — Desktop 1.46388.3 (agent 2.1.260)
|
|
46
|
+
|
|
47
|
+
New baseline `desktop-1.46388.3`. `sync` refused with 7 unknown deltas; all are classified and it now
|
|
48
|
+
runs clean.
|
|
49
|
+
|
|
50
|
+
#### Fixed — the sub-agent append sentinel could not see part of the append
|
|
51
|
+
|
|
52
|
+
Desktop 1.46388.3 restructured the per-session sub-agent append into three parts: the overridable
|
|
53
|
+
`## Cowork environment` section, a host-loop-only folder manifest, and one trailing sentence appended to
|
|
54
|
+
**both** branches. The sentinel fingerprinted only the two ternary branch texts, so the trailing sentence
|
|
55
|
+
changed the rendered append on both branches while the VM fingerprint did not move — and the pointer
|
|
56
|
+
coupling guard only demands a repoint on an axis whose fingerprint moved. A counterfactual confirms it:
|
|
57
|
+
Desktop 1.44121.1 with only that sentence added raises **no flags and syncs green**.
|
|
58
|
+
|
|
59
|
+
Two fingerprint axes now cover the composed parts (`manifest`, `suffix`). Both become mandatory once
|
|
60
|
+
either is recorded, and only the newest entry is compared, so entries predating the restructure are not
|
|
61
|
+
retro-failed.
|
|
62
|
+
|
|
63
|
+
Related, and separately load-bearing: the host-loop branch extractor was selecting the **wrong template**.
|
|
64
|
+
It anchored on a phrase that left the branch at this release and survives in unrelated prose in the same
|
|
65
|
+
module, so the drift it reported named text that is not the branch. It is now anchored on the ternary's
|
|
66
|
+
structure rather than on wording. The module selector keyed on the same dead phrase and was one chunk
|
|
67
|
+
move away from dropping the entire branch sentinel behind a single generic flag.
|
|
68
|
+
|
|
69
|
+
#### Added — the host-loop folder manifest is modeled
|
|
70
|
+
|
|
71
|
+
A host-loop sub-agent is now told the same thing production tells it: which folders exist on the user's
|
|
72
|
+
computer, each with its host path and its shell path, or a different sentence when there is nothing to
|
|
73
|
+
list. It is generated from the run's real mounts rather than read from a static asset — the same reason
|
|
74
|
+
the main-loop shell section became a generator at Desktop 1.14271.0 — and it uses **canonical** folder
|
|
75
|
+
paths, matching what production renders. Folders dropped by `COWORK_HARNESS_SOFT_MISSING=1` are listed as
|
|
76
|
+
unreachable rather than silently disappearing, which is production's behaviour for a folder that fails to
|
|
77
|
+
mount.
|
|
78
|
+
|
|
79
|
+
The text this generator produces is mixed into `promptAssetsHash`, so editing it stales recorded cassettes
|
|
80
|
+
exactly as editing a committed prompt asset does.
|
|
81
|
+
|
|
82
|
+
#### Added — `CLAUDE_CODE_QUESTION_EXTENDED`
|
|
83
|
+
|
|
84
|
+
A new gate-conditional key in the Cowork spawn window, pinned along with its gate. The gate reads
|
|
85
|
+
served-and-off today, so the key stays out of `spawn.env` and an upstream flip will surface as a diff. It
|
|
86
|
+
drives an AskUserQuestion behaviour the harness models, and Desktop and the agent shipped it in the same
|
|
87
|
+
release.
|
|
88
|
+
|
|
89
|
+
#### Cassettes
|
|
90
|
+
|
|
91
|
+
`example-pdf-skill` was **re-recorded** at `container`: it embeds the VM sub-agent append in its
|
|
92
|
+
recorded frames, and that text changed, so re-stamping would have left a committed fixture
|
|
93
|
+
contradicting what the harness now sends. Tool calls moved 5 → 4 against the cassette it replaced.
|
|
94
|
+
|
|
95
|
+
`example-multiselect-gate` (protocol) and `hostloop-computer-links` (hostloop) were **re-stamped**
|
|
96
|
+
rather than re-recorded. No sub-agent was dispatched in either recording, protocol receives no
|
|
97
|
+
sub-agent append at all by design, and neither cassette embeds the changed text — so the prompt change
|
|
98
|
+
provably never entered either recording, while re-recording at those tiers would bake the recording
|
|
99
|
+
machine's own MCP servers and skills into a public fixture.
|
|
100
|
+
|
|
101
|
+
#### Fixed — a mutation test made ambiguous by the new build
|
|
102
|
+
|
|
103
|
+
A second `=31999` literal appeared in an unrelated chunk, and the two were in different chunks at
|
|
104
|
+
1.44121.1, so no chunk marker separates them durably. The S4 mutation now changes every occurrence and
|
|
105
|
+
asserts the guard's **resolved value** rather than merely that it flagged, so a mutation that never
|
|
106
|
+
reached the guarded site cannot read as a pass.
|
|
107
|
+
|
|
108
|
+
## [3.3.0] — 2026-09-02
|
|
109
|
+
|
|
110
|
+
### Verification
|
|
111
|
+
|
|
112
|
+
Verified **locally**, against the Desktop install this baseline was synced from — Desktop `1.44121.1`,
|
|
113
|
+
agent `2.1.258`, macOS arm64, agent image `cowork-agent-base:2` — on 2026-09-02:
|
|
114
|
+
|
|
115
|
+
| suite | result |
|
|
116
|
+
|---|---|
|
|
117
|
+
| `boundary-check` | pass (3/3: allowlist-permits, loopback-not-proxied, hostloop-bash-egress) |
|
|
118
|
+
| `npm run test:live` | 4 files, **19 passed, 0 skipped** |
|
|
119
|
+
| `run examples/scenarios/` | **6/6 success** |
|
|
120
|
+
| 8 of 9 e2e scenarios | **8/8 success** — askuserquestion, multiselect, multiselect-deciderdir, l1-container, l1-egress, present-files, **semantic-evidence-files**, canary-hostloop. `smoke-l2-microvm` is the ninth and was not run. |
|
|
121
|
+
|
|
122
|
+
Tiers exercised: **`protocol`, `container`, `hostloop`**. **`microvm` was NOT exercised.**
|
|
123
|
+
|
|
124
|
+
**The `evidence_files` scoped path was exercised live**, by `smoke-semantic-evidence-files` at container
|
|
125
|
+
tier, which is new in this release and exists because the feature had no live or e2e coverage at all: it
|
|
126
|
+
passes scoped (`semanticEvidence.reason "graded"`, `paths ["outputs/report.md"]`) and refuses with the
|
|
127
|
+
`evidence_files` line removed (`reason "evidence_incomplete"`, naming all seven authored files), so the
|
|
128
|
+
scenario is falsifiable rather than decorative. Both readings are quoted in the scenario's own header
|
|
129
|
+
with their run ids. The **truncation refusal**, the **unscoped-starvation** case and the
|
|
130
|
+
**`no_pre_run_manifest` refusal** remain covered by unit tests only — the last of those because arming the
|
|
131
|
+
manifest for `semantic_matches` is precisely what stops a scenario reaching it; its live trigger is a
|
|
132
|
+
`--resume` turn.
|
|
133
|
+
|
|
134
|
+
Nothing was skipped and nothing was gated out. An earlier pass of this same suite reported 18 passed and
|
|
135
|
+
1 skipped — `live-outputs-delete`'s "touches outputs without deleting" case, whose agent issued no Bash
|
|
136
|
+
call, so its guard observed nothing. That case ran and passed here. It is model variance in the fixture,
|
|
137
|
+
and it is recorded because the suite reports such a case as a SKIP rather than a pass, by design.
|
|
138
|
+
|
|
139
|
+
**CI does not run any of this.** There is no `ANTHROPIC_API_KEY` repository secret, so `ci.yml`'s live
|
|
140
|
+
scenario stage soft-skips in seconds on both the pull request and the merge commit. A green CI run for
|
|
141
|
+
this release therefore covers build, unit tests, the agent-image recipe and the boundary check — **not
|
|
142
|
+
live inference**. The evidence above is one machine and one account, which is not the same guarantee.
|
|
143
|
+
|
|
144
|
+
### Fixed
|
|
145
|
+
|
|
146
|
+
- **A `semantic_matches` assert could grade a document containing NO authored files and report it as
|
|
147
|
+
complete.** The authored set is derived by diffing the work tree against a pre-run manifest, and the
|
|
148
|
+
manifest is only captured when a scenario asserts one of `no_unexpected_files`, `input_unmodified`,
|
|
149
|
+
`no_delete_in_outputs`, `no_delete_in_mounts` or `no_lost_write_back` — or the run is a `record`.
|
|
150
|
+
`semantic_matches` was not on that list. A scenario whose asserts include `semantic_matches` but none of
|
|
151
|
+
those five, and which is not a recording, therefore never armed the baseline: the capture returned zero
|
|
152
|
+
files, and the judge was handed the final message and transcript alone — with the result stamped as
|
|
153
|
+
complete authored evidence. Found on a live run that authored 7 files under `outputs/` and whose assert
|
|
154
|
+
reported none.
|
|
155
|
+
|
|
156
|
+
**This is a FALSE GREEN, not a new regression, and it is not new in 3.3.0.** The arming predicate has
|
|
157
|
+
carried the same five keys since authored-file evidence first reached the judge in `0.29.0`, so every
|
|
158
|
+
release from `0.29.0` through `3.2.1` graded that shape against no authored evidence at all. It failed
|
|
159
|
+
silently, which is why it outlived a 404ing recipe that failed loudly.
|
|
160
|
+
|
|
161
|
+
**Re-run your affected scenarios.** If a scenario asserts `semantic_matches` without any of the five keys
|
|
162
|
+
above, a previous green graded no authored files — whatever it was measuring, it was not the run's
|
|
163
|
+
output. Re-running is the only way to learn what the verdict should have been.
|
|
164
|
+
|
|
165
|
+
Both halves are fixed: `semantic_matches` now arms the manifest, and an absent manifest is recorded in the
|
|
166
|
+
capture's health so the assert fails `evidence unavailable` with
|
|
167
|
+
`semanticEvidence.reason: "no_pre_run_manifest"` instead of grading an empty set. The distinction matters
|
|
168
|
+
— every other reason describes evidence that exists and could not be fully shown; this one describes
|
|
169
|
+
evidence that was never derivable.
|
|
170
|
+
|
|
171
|
+
**This changes the verdict of existing scenarios**, in the same way as the truncation fix above: a
|
|
172
|
+
`semantic_matches` that was passing on final-message-and-transcript evidence alone will now either grade
|
|
173
|
+
against the authored files it should always have seen, or refuse if no baseline can be captured (a
|
|
174
|
+
`--resume` turn, where the baseline belongs to the first turn — re-run live without `--resume`). A rubric
|
|
175
|
+
written against the weaker document may need revisiting, and that is the point: it was never being graded
|
|
176
|
+
against what the skill produced.
|
|
177
|
+
|
|
178
|
+
- **A deliverable larger than the 16 KiB per-file capture cap was graded as a PREFIX, and could return a
|
|
179
|
+
pass on a claim its tail disproved.** With one 87 KB `outputs/report.md` at stock settings the capture
|
|
180
|
+
keeps the first 16 KiB and flags it `truncated` — but nothing was omitted, nothing was unreadable, and the
|
|
181
|
+
composed document sits far under the aggregate cap, so the assert graded it and stamped `graded`. A
|
|
182
|
+
negative rubric ("the report contains no unmitigated risk line") passed over a prefix that never included
|
|
183
|
+
the line. The only signal was a ` (truncated)` suffix on the file's heading; the evidence-health note,
|
|
184
|
+
which tells the judge not to read an absence as a negative, never mentioned truncation at all.
|
|
185
|
+
|
|
186
|
+
Truncation is now treated exactly like omission: any file the judge would grade that was kept only as a
|
|
187
|
+
prefix fails `evidence unavailable`, and the health note lists truncated paths in the same "do not infer
|
|
188
|
+
absence from the cut" register as dropped ones. **This changes the verdict of existing scenarios** — a run
|
|
189
|
+
whose deliverable exceeded 16 KiB was previously graded on its first 16 KiB and now refuses. The fix is to
|
|
190
|
+
scope the assert (`semantic_matches.evidence_files: ["outputs/report.md"]`), which exempts the file from
|
|
191
|
+
the per-file cap; raising `COWORK_HARNESS_AUTHORED_TOTAL_BYTES` alone does **not** lift that cap, and the
|
|
192
|
+
failure message says so rather than sending you to the wrong knob.
|
|
193
|
+
|
|
194
|
+
The feature made this shape more reachable, which is why it is fixed here: a scope on one assert exempts
|
|
195
|
+
its file from the per-file cap and consumes the shared budget, so an unscoped sibling's per-file allowance
|
|
196
|
+
is `min(16 KiB, whatever is left)` and can fall to almost nothing. The multi-assert warning now says so.
|
|
197
|
+
|
|
198
|
+
- **Raising the capture budget could have moved incompleteness from a loud refusal into a silent cut.**
|
|
199
|
+
Authored evidence sits near the tail of the judged document, so an overflow of the 262144-char aggregate
|
|
200
|
+
cap can eat it. Any `semantic_matches` — scoped or not — now refuses `authored_evidence_truncated` rather
|
|
201
|
+
than grading a document whose evidence was cut, and the message names the section that overflowed (a
|
|
202
|
+
document overflowing on `include_subagent_text` is not fixed by narrowing a file scope). The check
|
|
203
|
+
measures the offset at which the authored region ends, so a trim that reaches only the trailing
|
|
204
|
+
evidence-health note is an advisory warning, not a refusal.
|
|
205
|
+
|
|
206
|
+
**This changes the verdict of some existing scenarios, and you should expect it.** Affected: any live
|
|
207
|
+
`semantic_matches` whose composed judge document already exceeded 262144 chars far enough to cut into the
|
|
208
|
+
authored-file sections — in practice a run with a large deliverable, a long transcript, or
|
|
209
|
+
`include_subagent_text: true` with several substantial sub-agents. At stock settings a single large
|
|
210
|
+
deliverable does **not** reach the cap (32768 + 131072 + 65536 = 229376 chars with every section maxed),
|
|
211
|
+
so the two triggers that do not require raising the budget are **many** authored files — the per-file
|
|
212
|
+
`## Authored file: <path>` headings are counted, so several hundred small files overflow at the same total
|
|
213
|
+
— and `include_subagent_text: true`, whose per-dispatch 16 KiB is multiplied by an unbounded dispatch
|
|
214
|
+
count. Those runs were being graded against a document whose evidence had been silently truncated, and
|
|
215
|
+
could return a **pass** on a claim the judge never had the text to verify. They now fail `evidence unavailable` with
|
|
216
|
+
`semanticEvidence.reason: "authored_evidence_truncated"` and a message naming the section that overflowed.
|
|
217
|
+
|
|
218
|
+
This is a fix to a false PASS rather than a break in a covered surface — but it is **not** free: the check
|
|
219
|
+
is a property of the composed document, not of the rubric, and the harness cannot know which sections a
|
|
220
|
+
given rubric actually depended on. A rubric that was satisfiable from the agent's final message alone (the
|
|
221
|
+
first section, which is never cut) was being graded correctly and now turns red too. That is a deliberate
|
|
222
|
+
trade — a conservative refusal you can see and act on, over a silent grade you cannot — but if a newly-red
|
|
223
|
+
assert was genuinely answerable from the final message, this is why. What to do when you hit it, in order — (1) scope the judge with
|
|
224
|
+
`semantic_matches.evidence_files: ["outputs/<your deliverable>"]`, which shrinks the document to the files
|
|
225
|
+
the rubric is about and is the right answer in almost every case; (2) if the graded files themselves do not
|
|
226
|
+
fit, raise `COWORK_HARNESS_AUTHORED_TOTAL_BYTES` so the capture keeps them whole — but note the composed
|
|
227
|
+
document is still capped at 262144 chars, so past that point only scoping helps; (3) if the overflow is
|
|
228
|
+
sub-agent text, set `include_subagent_text: false`. The failure message names which of these applies.
|
|
229
|
+
|
|
230
|
+
|
|
231
|
+
- **`agentBinary.manifestChecksumMatch` was cross-checking the wrong release channel, and had been for a
|
|
232
|
+
month.** `sync` hard-coded the *stable* versioned manifest path. Desktop also stages release
|
|
233
|
+
**candidates**, served only from `…/claude-code-releases/rc/<commit>/`, so for agent `2.1.255` the
|
|
234
|
+
stable path 404s — and because the helper swallows every fetch failure to stay offline-capable, the
|
|
235
|
+
baseline recorded `"unknown"` for a build whose published checksum matches the staged ELF exactly.
|
|
236
|
+
`sync` now reads the channel out of the asar's own SDK descriptor and queries that, so
|
|
237
|
+
`baselines/desktop-1.40609.1.json` records `manifestChecksumMatch: true`.
|
|
238
|
+
|
|
239
|
+
This was not new with `2.1.255`. Of the 25 Desktop builds on record **3 are RC-staged**, and the two
|
|
240
|
+
earlier ones (`1.24012.9`/`.11`, agent `2.1.219`) recorded `true` only because that version had *also*
|
|
241
|
+
been promoted to stable — the wrong-channel query happened to resolve. Nothing in the output
|
|
242
|
+
distinguished that lucky pass from a real one, which is the defect this closes.
|
|
243
|
+
|
|
244
|
+
- **`sync` now says *why* a checksum cross-check did not run.** An HTTP status from the channel is a
|
|
245
|
+
`WARNING` (it does not serve this version's manifest); a transport failure is a `NOTE` (the rig has no
|
|
246
|
+
egress, which says nothing about the release). Previously both wrote `"unknown"` silently — and a
|
|
247
|
+
sandbox-blocked fetch was once mistaken for a finding about the release because of it. The recorded
|
|
248
|
+
field stays two-valued (`boolean | "unknown"`); the distinction is in the operator-facing output.
|
|
249
|
+
|
|
250
|
+
- **The documented ELF-download recipes `curl`ed a URL that 404s.** `docs/ci.md`, the companion skill's
|
|
251
|
+
`references/ci-recipe.md` and `docs/maintenance.md`'s recovery runbook all assumed the stable path
|
|
252
|
+
while pinning `V=2.1.255`, which is RC-staged — so the copy-paste CI step failed. They now take the
|
|
253
|
+
base URL from the baseline. `scripts/check-versions.ts` pins it: the invariant that kept `V=` honest
|
|
254
|
+
had no idea the URL had stopped working, and it is what propagated the broken pin into all three.
|
|
255
|
+
|
|
256
|
+
**If you copied the recipe from `3.2.1` or earlier, re-copy it.** The `curl` now reads its host from a
|
|
257
|
+
new `B=` line alongside the existing `V=` pin, because the stable path is not always where Desktop
|
|
258
|
+
staged the agent from. An older copy has no `B=` line and hard-codes the stable URL, so it keeps
|
|
259
|
+
working only for as long as the agent you pin happens to be a stable-channel build.
|
|
260
|
+
|
|
261
|
+
### Added
|
|
262
|
+
|
|
263
|
+
- **`agentBinary.releaseBaseUrl` in the platform baseline** — the release channel Desktop staged the
|
|
264
|
+
agent from. It names the source `manifestChecksumMatch` agreed with (without it the boolean is an
|
|
265
|
+
unattributed claim), makes the ELF-recovery runbook work for an RC-staged version, and surfaces a
|
|
266
|
+
stable↔RC flip as a `sync --diff` line. Recomputed from the local asar every sync, so it needs no
|
|
267
|
+
network; absent on baselines written earlier, all of which were stable-staged or later promoted.
|
|
268
|
+
|
|
269
|
+
- **`baselines/desktop-1.40609.1.json` re-synced** against the live Desktop 1.40609.1 install to pick up
|
|
270
|
+
the corrected row. (It is no longer the newest baseline — Desktop self-updated later in the same cycle and
|
|
271
|
+
`desktop-1.44121.1` ships alongside it; see **Parity** below. Both moved, which is why this release
|
|
272
|
+
touches two baseline files.) `provenance.fcache` moved with it — that payload is server-refreshed on Desktop's
|
|
273
|
+
own schedule and drifts between syncs; it is not a change this fix caused. `spawn`, `network`, all 28
|
|
274
|
+
gate rows and every fingerprint are unchanged.
|
|
275
|
+
|
|
276
|
+
- **`semantic_matches.evidence_files` — scope which authored files the judge grades.** A run that authors
|
|
277
|
+
many files could not pass a `semantic_matches` assert at all: the authored-file capture spends a fixed
|
|
278
|
+
64 KiB budget prefix-major then alphabetically, so a pipeline staging intermediates under
|
|
279
|
+
`outputs/_work/` exhausted it before reaching its own deliverable, and any omission refused the verdict —
|
|
280
|
+
over files no rubric mentioned. Naming the deliverable now (a) sends only those files to the judge,
|
|
281
|
+
(b) spends the capture budget on them **first**, (c) exempts them from the per-file cap, and (d) narrows
|
|
282
|
+
the evidence-unavailable refusal to in-scope omissions. A scenario with no `evidence_files` anywhere keeps
|
|
283
|
+
its previous grading behaviour; note that scoping ONE assert changes the shared capture for its unscoped
|
|
284
|
+
siblings (they share a single authored-file capture, and a scope exempts its files from the per-file cap),
|
|
285
|
+
which now warns.
|
|
286
|
+
|
|
287
|
+
Guards that come with it, because scoping is a new way to manufacture a vacuous green: globs matching
|
|
288
|
+
**nothing** fail evidence-unavailable (and the message lists every path the run authored — paths are
|
|
289
|
+
`<root>/<rel>`, which nothing else in the CLI surfaces); an empty list is a load-time error; an in-scope
|
|
290
|
+
file that is *truncated* rather than dropped also refuses, since a partial deliverable grades as a partial
|
|
291
|
+
document; and a scoped **pass** records the files it graded in its `evidence`.
|
|
292
|
+
|
|
293
|
+
The judged document is memoized per assert, and its cache key now includes the scope — keying it on
|
|
294
|
+
`include_subagent_text` alone was sufficient before scopes existed and is not any more.
|
|
295
|
+
|
|
296
|
+
- **`RunResult.assertions[].semanticEvidence`** — the typed reason a `semantic_matches` assert refused
|
|
297
|
+
(`scope_matched_nothing` | `in_scope_omitted` | `in_scope_truncated` | `evidence_incomplete` |
|
|
298
|
+
`authored_evidence_truncated`) or what it graded (`graded`, recorded on a substantive fail too — the bug
|
|
299
|
+
this guards against is a false ABSENCE, so a red is only actionable next to what the judge was shown),
|
|
300
|
+
with the paths. Five causes with five different fixes previously shared one prose message; a consumer had
|
|
301
|
+
to regex English to tell them apart.
|
|
302
|
+
|
|
303
|
+
- **`COWORK_HARNESS_AUTHORED_TOTAL_BYTES`** — raise the total authored-file evidence budget (default
|
|
304
|
+
65536) when a deliverable legitimately exceeds it. A malformed value throws rather than silently
|
|
305
|
+
defaulting: a quietly-defaulted evidence budget resurfaces later as an unexplained refusal. Raising it
|
|
306
|
+
enlarges the document sent to an external judge, so cost, latency and disclosure scale with it.
|
|
307
|
+
|
|
308
|
+
### Parity
|
|
309
|
+
|
|
310
|
+
- **Baseline `desktop-1.44121.1` (agent `2.1.258`).** Desktop self-updated during this release cycle, so
|
|
311
|
+
3.3.0 carries the sync as well as the fix above. Measured unchanged against `desktop-1.40609.1`:
|
|
312
|
+
`spawn` and `network` **byte-identical**, all 28 recorded gate rows identical in value, `guest`,
|
|
313
|
+
`mountLayout` and `settings` identical. The three committed example cassettes are **re-stamped, not
|
|
314
|
+
re-recorded** — `promptAssetsHash` resolves to the same `491afe2862dc67ea` under both baselines.
|
|
315
|
+
|
|
316
|
+
- **The gate-id set moved even though every pinned gate's value held, and one removal matters.**
|
|
317
|
+
`provenance.asarGateIds` goes **291 → 328 (+43, −6)**. Among the six ids Desktop dropped is
|
|
318
|
+
**`3246569822`, the `canSaveSkill` gate** — 3 occurrences in the 1.40609.1 asar, **0** in 1.44121.1. As
|
|
319
|
+
of Desktop 1.44121.1 the skill-saving capability is no longer gate-guarded for a standard session; it
|
|
320
|
+
rests on the `skillsEnabled` conjunct alone. The baseline still carries the `canSaveSkill:3246569822`
|
|
321
|
+
row — the server still serves the flag, and that row is fcache provenance, not an asar reading — but it
|
|
322
|
+
now carries a `note` recording that the id is absent from the asar, so it cannot be mistaken for
|
|
323
|
+
evidence that a gate still guards the feature. `provenance.spawnEnvSpreadCount` also moves 32 → 33, the
|
|
324
|
+
new nested conditional spread carrying the third-party key below.
|
|
325
|
+
|
|
326
|
+
- **Agent `2.1.255` → `2.1.258`.** The agent's `CLAUDE_*` env-flag export table moves **588 → 590**
|
|
327
|
+
(+3, −1): `CLAUDE_CODE_ARTIFACT_MULTI_FILE`, `CLAUDE_CODE_ARTIFACT_TOOLSET` and
|
|
328
|
+
`CLAUDE_CODE_MODEL_CATALOG_URL` arrive, `CLAUDE_CODE_PRINT_ENGINE_LOOP` goes. **None is set by the
|
|
329
|
+
Cowork spawn**, and no flag the spawn does set changed from or to zero consumers, so no harness change
|
|
330
|
+
follows from the bump.
|
|
331
|
+
|
|
332
|
+
- **This is the first sync to exercise `agentBinary.releaseBaseUrl`, and it moved in both directions
|
|
333
|
+
within a day.** `1.40609.1` staged a release candidate; `1.44121.1` is back on the stable channel, so
|
|
334
|
+
the recorded base flips from `…/claude-code-releases/rc/aa8f2d98…` to `…/claude-code-releases` and
|
|
335
|
+
`manifestChecksumMatch` reads `true` from the stable manifest. The guard added with the field caught a
|
|
336
|
+
real error while doing it: the documented `B=` pin still named the RC channel, which would have shipped
|
|
337
|
+
a `curl` that 404s for anyone on the new agent.
|
|
338
|
+
|
|
339
|
+
- **A new third-party spawn key, `CLAUDE_STREAM_IDLE_TIMEOUT_MS`, is classified as out-of-scope for the
|
|
340
|
+
modeled session.** Desktop constructs it only inside the `accountType === "3p"` branch and only when a
|
|
341
|
+
gateway provider supplies a stream-idle timeout, so a first-party session never receives it and it is
|
|
342
|
+
absent from the baseline's `spawn.env` (still 24 keys). `sync` refused to write until it was
|
|
343
|
+
classified, which is the refusal working as intended.
|
|
344
|
+
|
|
345
|
+
|
|
9
346
|
## [3.2.1] — 2026-09-02
|
|
10
347
|
|
|
11
348
|
### Parity
|
package/DESIGN.md
CHANGED
|
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
47
47
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
48
48
|
|
|
49
49
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
50
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
50
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.260**, per `baselines/desktop-1.46388.3.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
|
|
51
51
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
52
52
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
53
53
|
|
|
@@ -176,7 +176,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
176
176
|
|
|
177
177
|
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
178
178
|
|
|
179
|
-
> **Scope of that claim, stated plainly.** `2026-
|
|
179
|
+
> **Scope of that claim, stated plainly.** `2026-09-05 / desktop-1.46388.3` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.260** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, and covered **four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 passed, 0 skipped**; `run examples/scenarios/` **7/7 success, 39 assertions, 0 failed** (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe); and the e2e self-tests **8/8 success** (smoke-askuserquestion, smoke-multiselect, smoke-multiselect-deciderdir, smoke-l1-container, smoke-l1-egress, smoke-present-files, smoke-semantic-evidence-files, canary-hostloop). **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported none. Tiers exercised: **`protocol`, `container`, `hostloop`**. **`microvm` was NOT exercised** — it needs a real VM and CI excludes it for the same reason, so the ninth e2e scenario (`smoke-l2-microvm`) did not run. **Scope-out, so this is not read as more than it is.** (a) `subagent-manifest-probe` is new in this release and is the first live coverage the sub-agent append has ever had; it proves the composed append is ACTIONABLE — a sub-agent reaches the right files through both tool families, and a bare relative write lands at the host cwd — **not** that the harness's paraphrase matches Desktop's wording, which is the `manifest`/`suffix` fingerprint axes' job. (b) A live pass verifies observed behaviour, not the whole spawn contract by construction. (c) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: `example-pdf-skill` was re-recorded against this baseline and the other two committed cassettes were re-stamped (see the CHANGELOG for why each), so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
180
180
|
|
|
181
181
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
182
182
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^3.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^3.4.0"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.4.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.4.0"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Not sure a harness is what you need? The next two sections are the argument.
|
|
@@ -353,6 +353,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
353
353
|
## Status
|
|
354
354
|
|
|
355
355
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
356
|
-
**`desktop-1.
|
|
356
|
+
**`desktop-1.46388.3`**. Release-by-release verification notes (what was re-verified against
|
|
357
357
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
358
358
|
this section would otherwise duplicate lives in the sections above.
|
package/SPEC.md
CHANGED
|
@@ -630,7 +630,7 @@ and a fail observed) always fails that batch, regardless of the numeric rate.
|
|
|
630
630
|
"nonDeterministic?": bool, // true if any decision came from a non-deterministic source → not reproducible
|
|
631
631
|
"gateProvenance?": { "total": number, "bySource": {…}, "gates": [{ "question","answeredBy","answer","model?" }] }, // how each AskUserQuestion gate was answered; informational (never fails the verdict); live/partial lane only (absent on replay)
|
|
632
632
|
"permissiveAutoAllow?": ["string"], // tools auto-allowed by cowork parity that real Cowork BLOCKS → green is NOT faithful
|
|
633
|
-
"staleness?": [{ "class": "baseline|skill|shared-root|format|unverifiable-baseline|unverifiable-skill|resolved-tier|unverifiable-tier|prompt-assets|unverifiable-prompt-assets", "message" }], // replay only; cassette-staleness findings, surfaced for a JSON gate. Drift classes are non-failing by default (a stale but passing replay stays ok:true); `unverifiable-skill` FAILS by default since 2.0.0. `--strict` fails on every class, `--fail-on-skill-drift` adds skill/shared-root. `resolved-tier` = a `fidelity: cowork` cassette's recorded effectiveFidelity no longer matches the tier the scenario's baseline (pinned `baseline:` or `latest`) resolves to today — the recording exercises the wrong tier; `unverifiable-tier` = the tier check couldn't run for a baseline-dependent (`fidelity: cowork`) cassette (no recorded effectiveFidelity, or its pinned baseline failed to load). Tier resolution is baseline-only (the CLAUDE_FORCE_HOST_LOOP env override is suppressed) so verify results can't differ across machines. `prompt-assets` = the baseline's committed prompt-asset files (spawn.promptTemplate/subagentAppend/subagentAppendHostLoop) changed since record under the SAME appVersion — warn by default, `--strict` fails, re-record; `unverifiable-prompt-assets` = a recorded `fingerprint.promptAssetsHash` exists but the live baseline's prompt assets can't be hashed (a moved/dangling pointer) — can't verify ⇒ not green.
|
|
633
|
+
"staleness?": [{ "class": "baseline|skill|shared-root|format|unverifiable-baseline|unverifiable-skill|resolved-tier|unverifiable-tier|prompt-assets|unverifiable-prompt-assets", "message" }], // replay only; cassette-staleness findings, surfaced for a JSON gate. Drift classes are non-failing by default (a stale but passing replay stays ok:true); `unverifiable-skill` FAILS by default since 2.0.0. `--strict` fails on every class, `--fail-on-skill-drift` adds skill/shared-root. `resolved-tier` = a `fidelity: cowork` cassette's recorded effectiveFidelity no longer matches the tier the scenario's baseline (pinned `baseline:` or `latest`) resolves to today — the recording exercises the wrong tier; `unverifiable-tier` = the tier check couldn't run for a baseline-dependent (`fidelity: cowork`) cassette (no recorded effectiveFidelity, or its pinned baseline failed to load). Tier resolution is baseline-only (the CLAUDE_FORCE_HOST_LOOP env override is suppressed) so verify results can't differ across machines. `prompt-assets` = the baseline's committed prompt-asset files (spawn.promptTemplate/subagentAppend/subagentAppendHostLoop), or the sub-agent prompt text the harness generates rather than reads from an asset (the folder manifest + trailing sentence, Desktop >=1.46388.3), changed since record under the SAME appVersion — warn by default, `--strict` fails, re-record; `unverifiable-prompt-assets` = a recorded `fingerprint.promptAssetsHash` exists but the live baseline's prompt assets can't be hashed (a moved/dangling pointer) — can't verify ⇒ not green.
|
|
634
634
|
"skippedAssertions?": { "full": number, "partial": number }, // replay only; count of live-only assertions NOT evaluated (full = whole assertion skipped; partial = content half ran, fs/egress half dropped). The skipped ones are absent from `assertions[]`.
|
|
635
635
|
"toolResults?": [{ "toolUseId?","isError","text","assertText?" }], // tool-result text at assertion-fidelity cap (10 KB); backs tool_result_contains/tool_result_not_contains and their regex siblings tool_result_matches/tool_result_not_matches
|
|
636
636
|
"skillsInvoked?": ["string"], // Wave 1: skill/plugin ids invoked via the Skill tool_use event, call order, duplicates kept. Backs skill_triggered/no_skill_triggered.
|