cowork-harness 3.4.1 → 3.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +29 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +54 -21
- package/CHANGELOG.md +246 -0
- package/DESIGN.md +3 -3
- package/README.md +4 -4
- package/baselines/desktop-1.46388.3.json +2 -2
- package/baselines/desktop-1.46388.4.json +952 -0
- package/baselines/desktop-2.2553.1.json +1018 -0
- package/dist/agent/session.js +88 -10
- package/dist/baseline.js +5 -0
- package/dist/decide/llm-transport.js +31 -10
- package/dist/run/hook-events.js +20 -8
- package/dist/runtime/argv.js +27 -1
- package/dist/runtime/host-env.js +25 -5
- package/dist/runtime/hostloop.js +1 -1
- package/dist/scan.js +12 -0
- package/dist/session.js +34 -4
- package/dist/sync/cowork-sync.js +266 -12
- package/dist/types.js +23 -5
- package/docs/ci.md +2 -2
- package/docs/cli.md +3 -3
- package/docs/companion-skill.md +3 -3
- package/docs/fidelity-gaps.md +147 -12
- package/docs/invariants.md +2 -0
- package/docs/maintenance.md +25 -1
- package/docs/session.md +11 -5
- package/docs/subagents.md +13 -6
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +65 -43
- package/examples/replays/example-pdf-skill.cassette.json +95 -94
- package/examples/replays/hostloop-computer-links.cassette.json +75 -63
- package/package.json +2 -1
- package/scripts/check-claims.ts +103 -0
- package/scripts/check-versions.ts +1 -1
- package/scripts/gen-schema.ts +6 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.6.0
|
|
7
|
+
tracks-harness: cowork-harness 3.6.0 (baseline desktop-2.2553.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
29
|
-
> `desktop-
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.6.0` (baseline
|
|
29
|
+
> `desktop-2.2553.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.6.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.6.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.6.0"`. **Pin `@^3.6.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.6.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.275 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
|
|
41
41
|
# served only from .../claude-code-releases/rc/<commit>/ — the stable path 404s for those, and
|
|
42
42
|
# 2.1.255 is one. Take B from your pinned baseline's agentBinary.releaseBaseUrl; baselines
|
|
@@ -73,7 +73,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
73
73
|
GitHub-hosted runners, no token/Docker/agent:
|
|
74
74
|
|
|
75
75
|
```yaml
|
|
76
|
-
- run: npm i -g "cowork-harness@^3.
|
|
76
|
+
- run: npm i -g "cowork-harness@^3.6.0"
|
|
77
77
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
78
78
|
# no silent false-greens. WITHOUT --strict this
|
|
79
79
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -350,7 +350,7 @@ jobs:
|
|
|
350
350
|
with: { node-version: '24' }
|
|
351
351
|
- uses: actions/setup-python@v5
|
|
352
352
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
353
|
-
- run: npm i -g "cowork-harness@^3.
|
|
353
|
+
- run: npm i -g "cowork-harness@^3.6.0"
|
|
354
354
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
355
355
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
356
356
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -379,7 +379,7 @@ jobs:
|
|
|
379
379
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
380
380
|
fi
|
|
381
381
|
- if: steps.guard.outputs.live == 'true'
|
|
382
|
-
run: npm i -g "cowork-harness@^3.
|
|
382
|
+
run: npm i -g "cowork-harness@^3.6.0"
|
|
383
383
|
- if: steps.guard.outputs.live == 'true'
|
|
384
384
|
run: cowork-harness run scenarios/ --output-format json
|
|
385
385
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
4
|
-
(baseline `desktop-
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.6.0`
|
|
4
|
+
(baseline `desktop-2.2553.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -113,14 +113,43 @@
|
|
|
113
113
|
"PreToolUse"
|
|
114
114
|
],
|
|
115
115
|
"knownHookEvents": [
|
|
116
|
+
"ConfigChange",
|
|
117
|
+
"CwdChanged",
|
|
118
|
+
"DirectoryAdded",
|
|
119
|
+
"Elicitation",
|
|
120
|
+
"ElicitationResult",
|
|
121
|
+
"FileChanged",
|
|
122
|
+
"InstructionsLoaded",
|
|
123
|
+
"MessageDisplay",
|
|
116
124
|
"Notification",
|
|
125
|
+
"PermissionDenied",
|
|
126
|
+
"PermissionRequest",
|
|
127
|
+
"PostCompact",
|
|
128
|
+
"PostModelSwitch",
|
|
129
|
+
"PostToolBatch",
|
|
117
130
|
"PostToolUse",
|
|
131
|
+
"PostToolUseFailure",
|
|
118
132
|
"PreCompact",
|
|
133
|
+
"PreModelSwitch",
|
|
119
134
|
"PreToolUse",
|
|
120
135
|
"SessionEnd",
|
|
121
136
|
"SessionStart",
|
|
137
|
+
"Setup",
|
|
122
138
|
"Stop",
|
|
139
|
+
"StopFailure",
|
|
140
|
+
"SubagentStart",
|
|
123
141
|
"SubagentStop",
|
|
142
|
+
"TaskCompleted",
|
|
143
|
+
"TaskCreated",
|
|
144
|
+
"TeammateIdle",
|
|
145
|
+
"UserPromptExpansion",
|
|
146
|
+
"UserPromptSubmit",
|
|
147
|
+
"WorktreeCreate",
|
|
148
|
+
"WorktreeRemove"
|
|
149
|
+
],
|
|
150
|
+
"liveVerifiedHookEvents": [
|
|
151
|
+
"PostToolUse",
|
|
152
|
+
"SessionStart",
|
|
124
153
|
"UserPromptSubmit"
|
|
125
154
|
],
|
|
126
155
|
"enums": {
|
|
@@ -215,26 +215,49 @@ ASSERT_KEYS = _load_assert_keys()
|
|
|
215
215
|
# reason the key lists are: a hand-copied served-set silently stops warning about the event it was later
|
|
216
216
|
# extended to cover — the exact drift this check exists to prevent.
|
|
217
217
|
_FALLBACK_SERVED_HOOK_EVENTS = {"PreToolUse"}
|
|
218
|
+
# Re-sourced 2026-09-06 from the agent's OWN hooks-config validator array (ELF 2.1.260), not from a grep
|
|
219
|
+
# for event-name constants. The previous 9-name set reported the other 24 -- PostCompact and
|
|
220
|
+
# MessageDisplay among them -- identically to a misspelling, at ERROR severity below.
|
|
218
221
|
_FALLBACK_KNOWN_HOOK_EVENTS = {
|
|
219
|
-
"PreToolUse", "PostToolUse", "
|
|
220
|
-
"
|
|
222
|
+
"PreToolUse", "PostToolUse", "PostToolUseFailure", "PostToolBatch",
|
|
223
|
+
"Notification", "UserPromptSubmit", "UserPromptExpansion", "SessionStart",
|
|
224
|
+
"SessionEnd", "Stop", "StopFailure", "SubagentStart", "SubagentStop",
|
|
225
|
+
"PreCompact", "PostCompact", "PreModelSwitch", "PostModelSwitch",
|
|
226
|
+
"PermissionRequest", "PermissionDenied", "Setup", "TeammateIdle",
|
|
227
|
+
"TaskCreated", "TaskCompleted", "Elicitation", "ElicitationResult",
|
|
228
|
+
"ConfigChange", "WorktreeCreate", "WorktreeRemove", "InstructionsLoaded",
|
|
229
|
+
"CwdChanged", "FileChanged", "DirectoryAdded", "MessageDisplay",
|
|
221
230
|
}
|
|
231
|
+
# The subset a plugin hook has been OBSERVED to fire for here (live-verified 2026-08-01, container +
|
|
232
|
+
# hostloop). Kept apart from the known set because the message wording depends on which claim we can
|
|
233
|
+
# make: accepted-by-the-validator is not reached-by-a-run.
|
|
234
|
+
_FALLBACK_LIVE_VERIFIED_HOOK_EVENTS = {"SessionStart", "UserPromptSubmit", "PostToolUse"}
|
|
222
235
|
|
|
223
236
|
|
|
224
237
|
def _load_hook_events():
|
|
225
|
-
"""(served, known) hook-event sets, from the generated assertion-keys.json sidecar."""
|
|
238
|
+
"""(served, known, live-verified) hook-event sets, from the generated assertion-keys.json sidecar."""
|
|
226
239
|
p = Path(__file__).resolve().parent / "assertion-keys.json"
|
|
227
240
|
try:
|
|
228
241
|
d = json.loads(p.read_text(encoding="utf-8"))
|
|
229
242
|
served, known = set(d["servedHookEvents"]), set(d["knownHookEvents"])
|
|
243
|
+
# Tolerate a sidecar generated before liveVerifiedHookEvents existed: fall back rather than
|
|
244
|
+
# dropping to the embedded KNOWN set wholesale, which would undo the 33-name fix. Membership
|
|
245
|
+
# test, NOT truthiness: an explicitly EMPTY list means the firing receipt was withdrawn, and
|
|
246
|
+
# `or` would silently restore the 3-name claim -- widening a receipt, which is the one direction
|
|
247
|
+
# this constant exists to prevent.
|
|
248
|
+
live = set(d["liveVerifiedHookEvents"]) if "liveVerifiedHookEvents" in d else set(_FALLBACK_LIVE_VERIFIED_HOOK_EVENTS)
|
|
230
249
|
if served and known:
|
|
231
|
-
return served, known
|
|
250
|
+
return served, known, live
|
|
232
251
|
except Exception:
|
|
233
252
|
pass
|
|
234
|
-
return
|
|
253
|
+
return (
|
|
254
|
+
set(_FALLBACK_SERVED_HOOK_EVENTS),
|
|
255
|
+
set(_FALLBACK_KNOWN_HOOK_EVENTS),
|
|
256
|
+
set(_FALLBACK_LIVE_VERIFIED_HOOK_EVENTS),
|
|
257
|
+
)
|
|
235
258
|
|
|
236
259
|
|
|
237
|
-
SERVED_HOOK_EVENTS, KNOWN_HOOK_EVENTS = _load_hook_events()
|
|
260
|
+
SERVED_HOOK_EVENTS, KNOWN_HOOK_EVENTS, LIVE_VERIFIED_HOOK_EVENTS = _load_hook_events()
|
|
238
261
|
|
|
239
262
|
# Self-check: every valid assertion key must be classified, else the replay-class lint logic mishandles it.
|
|
240
263
|
# Surfaced loudly at load AND as a lint ERROR in cmd_lint (so --strict / exit codes flow). Never sys.exit here.
|
|
@@ -1489,15 +1512,17 @@ def _lint_hook_events(path):
|
|
|
1489
1512
|
"""Flag hook events a plugin DECLARES that this harness does not SERVE.
|
|
1490
1513
|
|
|
1491
1514
|
Why this exists: the harness installs `PreToolUse` only, while real Cowork installs three event types
|
|
1492
|
-
and the agent binary
|
|
1515
|
+
and the agent binary accepts thirty-three. A plugin declaring `UserPromptSubmit` therefore mounted, ran,
|
|
1493
1516
|
and produced no comment of any kind — the surface was discoverable only by grepping the harness's own
|
|
1494
1517
|
compiled output, which is exactly what one consumer had to do.
|
|
1495
1518
|
|
|
1496
|
-
|
|
1497
|
-
|
|
1498
|
-
|
|
1499
|
-
|
|
1500
|
-
|
|
1519
|
+
Severity is three-way and each level means something different. A name the agent's validator ACCEPTS
|
|
1520
|
+
but this harness does not serve is INFO, not a skill defect: the agent loads a plugin's hooks.json
|
|
1521
|
+
through its own `--plugin-dir` channel, which the harness neither serves nor blocks. Of those, only
|
|
1522
|
+
the events in LIVE_VERIFIED_HOOK_EVENTS have been observed firing here, so the confident wording is
|
|
1523
|
+
scoped to them and every other accepted name says so. A name the validator REJECTS is ERROR — that
|
|
1524
|
+
one genuinely never runs on any surface. Claiming "this will not fire" for an accepted event would
|
|
1525
|
+
assert more than is known; claiming nothing
|
|
1501
1526
|
leaves the consumer to reverse-engineer it. So say precisely what is known.
|
|
1502
1527
|
"""
|
|
1503
1528
|
findings = []
|
|
@@ -1533,17 +1558,25 @@ def _lint_hook_events(path):
|
|
|
1533
1558
|
continue
|
|
1534
1559
|
line_no = next((i for i, ln in enumerate(lines, 1) if f'"{name}"' in ln), 1)
|
|
1535
1560
|
if name in KNOWN_HOOK_EVENTS:
|
|
1561
|
+
fires = (
|
|
1562
|
+
"fires here — a plugin's own `hooks/hooks.json` is loaded and executed by the agent "
|
|
1563
|
+
"binary (live-verified at both `container` and `hostloop`)"
|
|
1564
|
+
if name in LIVE_VERIFIED_HOOK_EVENTS
|
|
1565
|
+
else "is a hook event the agent accepts, and it loads a plugin's own `hooks/hooks.json` "
|
|
1566
|
+
"itself — though whether a harness run ever reaches this event's trigger has not "
|
|
1567
|
+
"been verified here"
|
|
1568
|
+
)
|
|
1536
1569
|
findings.append(Finding(
|
|
1537
1570
|
"INFO", "hook-event-not-served",
|
|
1538
|
-
f"`{name}` fires
|
|
1539
|
-
f"agent binary (live-verified at both `container` and `hostloop`) — but cowork-harness "
|
|
1571
|
+
f"`{name}` {fires} — but cowork-harness "
|
|
1540
1572
|
f"itself installs only {', '.join(sorted(SERVED_HOOK_EVENTS))} on `initialize`. Two "
|
|
1541
1573
|
f"consequences: there is no assertion key for this event, so a scenario cannot GATE on it; "
|
|
1542
|
-
f"and
|
|
1543
|
-
f"anything driven by
|
|
1544
|
-
"
|
|
1545
|
-
"
|
|
1546
|
-
"
|
|
1574
|
+
f"and if real Cowork installs a `{name}` hook of its own, the harness does not reproduce it, "
|
|
1575
|
+
f"so anything driven by that is absent here. (Cowork installs hooks of its own for "
|
|
1576
|
+
f"PreToolUse, PostToolUse and UserPromptSubmit only.)",
|
|
1577
|
+
"The harness does not block your hook — this is about assertability, not breakage. To gate "
|
|
1578
|
+
"on its effect, assert the OBSERVABLE result instead (a file it writes, a tool it blocks), "
|
|
1579
|
+
"not the hook itself.",
|
|
1547
1580
|
path, line_no,
|
|
1548
1581
|
))
|
|
1549
1582
|
elif name.lower() in {e.lower() for e in KNOWN_HOOK_EVENTS}:
|
|
@@ -1558,8 +1591,8 @@ def _lint_hook_events(path):
|
|
|
1558
1591
|
else:
|
|
1559
1592
|
findings.append(Finding(
|
|
1560
1593
|
"ERROR", "hook-event-unknown",
|
|
1561
|
-
f"`{name}` is not a
|
|
1562
|
-
f"hook would never run on any surface.",
|
|
1594
|
+
f"`{name}` is not a hook event the agent recognizes — an unrecognized event name is "
|
|
1595
|
+
f"ignored, so this hook would never run on any surface.",
|
|
1563
1596
|
f"Check spelling and capitalization. Valid events: {', '.join(sorted(KNOWN_HOOK_EVENTS))}.",
|
|
1564
1597
|
path, line_no,
|
|
1565
1598
|
))
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,252 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.6.0] — 2026-09-18
|
|
10
|
+
|
|
11
|
+
### Upgrade notes
|
|
12
|
+
|
|
13
|
+
- **If you run against Claude Desktop 2.2553.1 (agent 2.1.275), upgrade — `critique` and `--decider-llm`
|
|
14
|
+
are broken on 3.5.0 there.** That agent makes an auxiliary Haiku call in `-p` mode, and 3.5.0's LLM
|
|
15
|
+
transport hard-fails on the two-model envelope it produces (`critique` exits 2 with no error text).
|
|
16
|
+
3.6.0 identifies the primary model instead of counting keys. Nothing changes on older agents.
|
|
17
|
+
- **`latest` now resolves to `desktop-2.2553.1`.** A cassette you recorded against `1.46388.4` with
|
|
18
|
+
`baseline: latest` reports `baseline` staleness on replay (warn by default; `--strict` fails). Re-record
|
|
19
|
+
it, or pin the scenario to `desktop-1.46388.4` if you are not ready to move.
|
|
20
|
+
- **The spawned agent's env gains `CLAUDE_CODE_DESKTOP_APP_VERSION`** on baselines from 2.2553.1 on. It
|
|
21
|
+
is the value the agent uses for the `anthropic-client-version` request header; older baselines are
|
|
22
|
+
unaffected. If you snapshot the spawn env, expect the new key.
|
|
23
|
+
- **Every `-p` call and every run on agent 2.1.275 now carries a ~$0.001 Haiku entry** in `modelUsage`
|
|
24
|
+
and in the run result's cost. Cost comparisons across the 2.1.260→2.1.275 bump will show it — it is
|
|
25
|
+
the agent's spend, not the harness's.
|
|
26
|
+
|
|
27
|
+
### Added
|
|
28
|
+
- **Parity: baseline `desktop-2.2553.1` (agent 2.1.275)** — the first `2.x` Claude Desktop. `sync` refused
|
|
29
|
+
to write with **10 unknown deltas**; all are resolved and the baseline is clean. Two of the ten turned
|
|
30
|
+
out to be defects in this repo's own extractor rather than changes in Desktop:
|
|
31
|
+
- **The S6c Artifact-gate flag was a false alarm.** The frame-artifacts predicate is byte-identical; it
|
|
32
|
+
merely stopped being its own statement (it now shares a declaration with the Artifact host-grant
|
|
33
|
+
binding). The sentinel's value capture ran past the top-level comma and swallowed the sibling binding,
|
|
34
|
+
so an anchored whole-expression match rejected an unchanged predicate — and the message it printed
|
|
35
|
+
("cached-arm/HIPAA/trailing-term change") was simply wrong. The value is now sliced brace/paren/quote
|
|
36
|
+
aware to the first top-level `,` or `;`. Both directions are pinned by tests: the sibling-binding shape
|
|
37
|
+
stays clean, and a real widening hidden before the comma still fires.
|
|
38
|
+
- **The path-gate tool set and path keys were refactored, not removed.** Desktop replaced two array
|
|
39
|
+
literals with one tool→path-key map plus `Object.keys` / `[...new Set(Object.values(…))]` derivations,
|
|
40
|
+
which accounted for 4 of the 10 deltas at once. Same five tools, same two keys — no contract change.
|
|
41
|
+
The extractor now accepts either form. The map is matched **exactly and in order**, because
|
|
42
|
+
`Object.keys` order reaches the sub-agent prompt through `.join(", ")` while the manifest fingerprint
|
|
43
|
+
hashes generator source — a reordered map would otherwise change what the model reads with every check
|
|
44
|
+
still green. Keys and values must resolve to the *same* map, and an ambiguous binding flags rather than
|
|
45
|
+
taking the first match: the bundle carries two more same-shaped maps that add Bash/NotebookEdit/MultiEdit.
|
|
46
|
+
|
|
47
|
+
- **`npm run check:claims` — a staleness report for this repo's "binary-verified" claims.** It lists every
|
|
48
|
+
version-stamped claim in `src/`, `scripts/` and `docs/` that is behind the currently pinned agent and
|
|
49
|
+
`app.asar` versions. First run: **42 of 49 claims behind the pin**, the oldest about 34 baselines back.
|
|
50
|
+
It **exits 0 by design and is not a gate** — a stale stamp is not a wrong claim, and hard-failing would
|
|
51
|
+
force a version bump every sync that anyone could satisfy by editing the digit without re-reading the
|
|
52
|
+
binary. Maintainers get it as step 3b of the parity-sync ritual; contributors need not run it.
|
|
53
|
+
- **`test/subagent-model-precedence-elf.test.ts`** — pins the sub-agent model resolution order against the
|
|
54
|
+
agent binary itself, so the precedence corrected in 3.5.0 cannot silently drift back. It reads the
|
|
55
|
+
binary's own telemetry labels rather than a minified symbol (the gate accessor beside that code renamed
|
|
56
|
+
between two Desktop patch builds). Its second assertion is **not** gated on a staged binary: it fails if
|
|
57
|
+
`docs/session.md`, `docs/subagents.md` or `src/session.ts` reintroduces the reversed order, which is the
|
|
58
|
+
half CI can enforce and the way that claim went wrong in three places at once.
|
|
59
|
+
|
|
60
|
+
**Why these two:** the repo carries ~49 version-stamped claims about the agent binary and,
|
|
61
|
+
before 3.5.0, exactly one was re-derived from the binary by a test. Both claims spot-checked during that
|
|
62
|
+
release turned out wrong — the hook-event list and the model precedence. Two for two is not a sample that
|
|
63
|
+
justifies leaving the rest unexamined, but it also does not justify pretending a report verifies them: it
|
|
64
|
+
shows the population and its age, and a human decides what to re-read.
|
|
65
|
+
|
|
66
|
+
### Fixed
|
|
67
|
+
|
|
68
|
+
- **The LLM decider transport no longer hard-fails on agent 2.1.275's two-model envelope.** The agent that
|
|
69
|
+
ships with Desktop 2.2553.1 makes an auxiliary Haiku call in `-p` mode, so `claude -p --output-format json`
|
|
70
|
+
now reports two `modelUsage` keys where 2.1.260 reported one — measured with the same prompt and flags
|
|
71
|
+
against both native binaries. The transport asserted exactly one key, which turned **every** critique
|
|
72
|
+
evaluator pass and every `--decider-llm` gate into an instrument failure on the new agent (`critique`
|
|
73
|
+
exited 2 with no error text; found by the live lane, which had this test gated off in the previous pass).
|
|
74
|
+
The primary model is now *identified* as the key that resolves the requested `--model` (exact id, or the
|
|
75
|
+
id carrying a floating alias like `sonnet` as a dash-separated segment) rather than *assumed* from the
|
|
76
|
+
count. Zero or several keys resolving the request still fails closed — that ambiguity is the contract
|
|
77
|
+
break the check exists to catch. The whole usage map is still passed through, so the auxiliary call's
|
|
78
|
+
cost is not lost.
|
|
79
|
+
- **The spawned agent now sends the client-identity headers production sends.** Desktop 2.2553.1 sets
|
|
80
|
+
`CLAUDE_CODE_DESKTOP_APP_VERSION` unconditionally on first-party sessions, and the agent reads it on the
|
|
81
|
+
`local-agent` entrypoint — which the harness pins — as the fallback source of the `anthropic-client-version`
|
|
82
|
+
header (its companion `anthropic-client-platform` is the hard-coded literal `desktop_app`) whenever
|
|
83
|
+
`ANTHROPIC_CUSTOM_HEADERS` carries none, which is the harness's case. The key is host-derived (an Electron `app.getVersion()` call), so it is allowlisted in the sync
|
|
84
|
+
**and** injected from the baseline's `appVersion` in both the container and native spawn envs — **version-gated** to baselines from 2.2553.1 on, since injecting it on an older baseline would hand the agent a key that baseline's Desktop never set (not symmetric with `CLAUDE_CODE_HOST_PLATFORM`, which every asar on record sets). Allowlisting
|
|
85
|
+
it alone would have been silent: the allowlist is consulted before the pin list and the key is not
|
|
86
|
+
required, so the harness would simply have stopped sending those headers with nothing failing.
|
|
87
|
+
- **`CLAUDE_ARTIFACT_HOST_GRANT` is guarded, not merely allowlisted.** The new key is allowlisted on the
|
|
88
|
+
grounds that a default session never receives it — a claim that rests entirely on its guard. A new
|
|
89
|
+
sentinel requires it to stay gated on the *same* predicate as the Artifact tool spread, and fires if it
|
|
90
|
+
is ever constructed unconditionally or re-keyed. (The existing frame-artifacts assertion could not reach
|
|
91
|
+
it: the two spreads have different shapes.)
|
|
92
|
+
- **`CLAUDE_CODE_DISABLE_CRON` gained a second disjunct** (a managed-settings scheduled-tasks switch). The
|
|
93
|
+
pinned value is unchanged at `"1"`, and it is *earned* rather than assumed — the spawn window still passes
|
|
94
|
+
`disableCron:!0`, which short-circuits. The resolver and its anchor admit the new shape; the disjunct is
|
|
95
|
+
not inert in general, only under that short-circuit.
|
|
96
|
+
|
|
97
|
+
### Changed
|
|
98
|
+
|
|
99
|
+
|
|
100
|
+
- `CLAUDE_CODE_MODEL_CATALOG` (new, third-party-only branch) is allowlisted, matching the standing rule for
|
|
101
|
+
third-party-only keys.
|
|
102
|
+
- **`design` added to the host-inventory scan's known-built-in skill roster.** It surfaced as a finding on
|
|
103
|
+
the first fresh `container` recording after this sync, on a cassette whose scenario declares no skills.
|
|
104
|
+
It qualifies under the roster's existing three criteria: the recording was sealed (`container`, so
|
|
105
|
+
`HOME=/tmp` and no host `~/.claude`), `"design"` is a bare literal in both the staged agent ELF and the
|
|
106
|
+
host CLI, and five personal skill names from the same machine are absent from that binary. It is not new
|
|
107
|
+
to this agent — the `design-consent` / `design-revoke` slash commands were already in the previously
|
|
108
|
+
shipped cassette; what changed is that the feature now also registers in `skills[]`, an axis the scan
|
|
109
|
+
treats more strictly.
|
|
110
|
+
- **All three committed cassettes in `examples/replays/` are re-recorded against `desktop-2.2553.1`**, each reporting no behavioural change versus the recording it replaced. The `protocol` fixture was recorded on the hermetic managed config dir (`ANTHROPIC_API_KEY` path) and the `container` one in a sealed container; `verify-cassettes` reports zero host-inventory findings on all three.
|
|
111
|
+
- **A full live pass was run against `desktop-2.2553.1` / agent 2.1.275**, all four suites and all four tiers: `boundary-check` 6/6; e2e self-tests 9/9 including `smoke-l2-microvm` in a real VM and `smoke-multiselect-deciderdir` through the `--decider-llm` path; `npm run test:live` 19 tests, 18 passed, 1 failed, **0 skipped** (the previous pass had one skip — the hostloop `critique` case — which this pass exercised for the first time and which found the transport defect fixed above); `run examples/scenarios/` 7/7. The one live red is a pre-existing `live-matrix` case on old baselines where the model sometimes answers as text instead of calling `AskUserQuestion` — model variance, re-run and flipped, logged for hardening.
|
|
112
|
+
- **`test/model-provenance.test.ts`'s pre-coverage-note test now builds its own fixture.** Every committed cassette now carries `model` coverage, so no shipped fixture emits the note the test reads. Rather than asserting the note's shape only when one happens to be present — a test that could not fail — it rewrites a real cassette's session fingerprint to the pre-`model` hash in a temp tree that preserves the relative session layout.
|
|
113
|
+
|
|
114
|
+
|
|
115
|
+
## [3.5.0] — 2026-09-06
|
|
116
|
+
|
|
117
|
+
**Live verification for this release** (macOS arm64, agent **2.1.260**, agent image `cowork-agent-base:2`,
|
|
118
|
+
Desktop 1.46388.4):
|
|
119
|
+
|
|
120
|
+
| Suite | Result |
|
|
121
|
+
|---|---|
|
|
122
|
+
| `boundary-check` | **6/6** — host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress |
|
|
123
|
+
| `npm run test:live` | **4 files, 18 passed, 1 skipped** — the skip is `live-outputs-delete`'s silent-guard case, which reports SKIPPED when the agent issues no Bash call (documented as expected-rare, not a failure) |
|
|
124
|
+
| `run examples/scenarios/` | **7/7 success** across `container`, `hostloop` and `protocol` |
|
|
125
|
+
| e2e self-tests | **8/8 success** — askuserquestion, multiselect, multiselect-deciderdir, l1-container, l1-egress, present-files, semantic-evidence-files, canary-hostloop |
|
|
126
|
+
|
|
127
|
+
**Two things stated rather than glossed.** (a) The first `test:live` run showed one red — a sub-agent
|
|
128
|
+
WebSearch that the model simply did not perform — and the first scenario batch showed one red on
|
|
129
|
+
`example-pdf-skill`. Both were **model variance**: each passed on re-run with nothing changed, which is
|
|
130
|
+
the disposition the tests themselves prescribe for this class. They are recorded because a suite that
|
|
131
|
+
only ever reports its green run is not evidence of anything. (b) `smoke-l2-microvm` was **not run** this
|
|
132
|
+
release. It is the ninth e2e scenario and the tier CI can never cover (Apple-VZ is macOS-arm64 only); the
|
|
133
|
+
most recent microvm pass remains the one recorded under 3.4.0.
|
|
134
|
+
|
|
135
|
+
### Added
|
|
136
|
+
|
|
137
|
+
- **`liveVerifiedHookEvents` in the generated `assertion-keys.json` sidecar.** A new key alongside
|
|
138
|
+
`servedHookEvents` and `knownHookEvents`, naming the subset of hook events a plugin's own hook has
|
|
139
|
+
actually been **observed** to fire for in a harness run (three, live-verified at `container` and
|
|
140
|
+
`hostloop`). It exists so the linter can distinguish "the agent's validator accepts this name" from "a
|
|
141
|
+
run reaches this trigger" — two claims a single list would have conflated. Additive: a consumer reading
|
|
142
|
+
only the two existing keys is unaffected, and an older `scenario.py` ignores it.
|
|
143
|
+
- **`test/hook-events-elf-parity.test.ts`** — re-extracts the hook-event list from the staged agent binary
|
|
144
|
+
at test time instead of comparing two hand-maintained files. It **skips where no Desktop is staged**, CI
|
|
145
|
+
included, and says so in its own header: a green CI is not evidence for this invariant.
|
|
146
|
+
- **Two rows in [`docs/invariants.md`](./docs/invariants.md)**: the hook-event invariant below, and
|
|
147
|
+
`provenance.spawnEnvKeys` + `spawnEnvSpreadCount` as the spawn-env drift alarm. The latter documents
|
|
148
|
+
machinery that has existed since Desktop 1.24012.1 but appeared **zero times** outside `src/` — which is
|
|
149
|
+
why a later design exercise set about re-inventing it. [`docs/maintenance.md`](./docs/maintenance.md)
|
|
150
|
+
now tells a maintainer to read those two values in a `sync` diff first, and why a key-set delta beats a
|
|
151
|
+
count: it names the key.
|
|
152
|
+
|
|
153
|
+
### Fixed
|
|
154
|
+
|
|
155
|
+
- **`lint-skill` and the mount-time hook warning reported 24 valid hook events as misspellings.**
|
|
156
|
+
`KNOWN_HOOK_EVENTS` held **9** names, assembled by grepping the agent binary for event-name constants;
|
|
157
|
+
the agent's own hooks-config validator accepts **33**. A plugin declaring `PostCompact` or
|
|
158
|
+
`MessageDisplay` — both accepted by the agent, both of which run — got "not a recognized hook event …
|
|
159
|
+
Check spelling/capitalization", at **ERROR** severity in `lint-skill` and as a `::warning::` on the run
|
|
160
|
+
path. **The message text changed on both paths**, so a CI job grepping for the old strings needs
|
|
161
|
+
updating: the ERROR now reads "is not a hook event the agent recognizes", and an accepted-but-unserved
|
|
162
|
+
event reports at INFO / `::notice::` with the confident "it WILL fire" wording reserved for the three
|
|
163
|
+
events actually observed firing here. The list is now sourced from the validator array itself, and
|
|
164
|
+
`test/hook-events-elf-parity.test.ts` re-reads it from the staged agent binary rather than from a
|
|
165
|
+
committed fixture, because a fixture-vs-const test compares two hand-maintained files and only moves
|
|
166
|
+
when someone re-extracts by hand — the step that had failed. That test **skips where no Desktop is
|
|
167
|
+
staged**, CI included, so a green CI is not evidence for this invariant.
|
|
168
|
+
- **The sub-agent model precedence was documented backwards in three places.** `docs/session.md`,
|
|
169
|
+
`docs/subagents.md` and `src/session.ts` stated `env > dispatch param > frontmatter > inherit`. The
|
|
170
|
+
binary resolves **dispatch param > frontmatter > env > inherit** (verified in agent 2.1.260; promoting
|
|
171
|
+
the env override to the top is what `CLAUDE_CODE_SUBAGENT_MODEL_FORCE` does, and the Cowork spawn sets
|
|
172
|
+
neither that flag nor `CLAUDE_CODE_COORDINATOR_FORCE_WORKER_INHERIT_MODEL`). So
|
|
173
|
+
`agent_env.subagent_model` does **not** outrank a sub-agent's own `model:` frontmatter, contrary to
|
|
174
|
+
what the docs promised.
|
|
175
|
+
- **Two model-forcing env vars leaked asymmetrically into `hostloop` and `protocol`.**
|
|
176
|
+
`SCRUBBED_AGENT_ENV_KEYS` matches exact keys, so `CLAUDE_CODE_SUBAGENT_MODEL` never covered
|
|
177
|
+
`CLAUDE_CODE_SUBAGENT_MODEL_FORCE` or `CLAUDE_CODE_COORDINATOR_FORCE_WORKER_INHERIT_MODEL`. An
|
|
178
|
+
operator with either exported got different sub-agent model resolution on the two inheriting tiers than
|
|
179
|
+
on `container`/`microvm`, which is the asymmetry that constant exists to prevent. Both are now scrubbed.
|
|
180
|
+
**Operator-visible consequence, and the reason it is filed as a fix rather than a cleanup:** unlike the
|
|
181
|
+
three keys already on that list, these two have **no `agent_env` knob**, so exporting one in your shell
|
|
182
|
+
now has it silently deleted with nothing authored to put it back. `docs/session.md` says so. If you
|
|
183
|
+
need either, set it inside the run rather than in the environment the harness inherits. Real Cowork
|
|
184
|
+
sets neither. Companion tests drive the real `hostloop` and `protocol` env builders with each key set —
|
|
185
|
+
not a hand-built object — and pin the exact-key *mechanism*, plus the deliberate decision to leave
|
|
186
|
+
`CLAUDE_CODE_COORDINATOR_MODE` (which enables one of them, and leaks the same way) unscrubbed while no
|
|
187
|
+
coordinator surface is modelled.
|
|
188
|
+
|
|
189
|
+
### Documentation
|
|
190
|
+
|
|
191
|
+
- **Corrected what this repo says about Cowork's force-ask PreToolUse hook.** It gates **nine** tools,
|
|
192
|
+
not four (the four named ones plus `create`/`update`/`delete_scheduled_task` and
|
|
193
|
+
`start`/`stop_watching`), and its decision is **not** unconditional: two gate-conditioned early returns
|
|
194
|
+
defer to the auto-mode permission classifier, so in a real auto-mode session 7 of the 9 raise no
|
|
195
|
+
prompt. The scheduled-task branch ships in Desktop 1.22209.0 — before 1.24012.9, the first baseline to
|
|
196
|
+
record this hook — so the note in `spawn.hooks` was already wrong for 5 of the 9 tools when it was
|
|
197
|
+
first written; the second branch lands three baselines later at 1.26832.0, making it wrong for all
|
|
198
|
+
nine from there on. Inaccurate in all fourteen baselines carrying it, either way.
|
|
199
|
+
**Nothing about the harness changes**: auto mode is structurally unreachable here, so for every mode a
|
|
200
|
+
scenario can express production still answers `ask`, and serving that hook unconditionally would remain
|
|
201
|
+
faithful. Corrected in `desktop-1.46388.3` forward; the older baselines keep their wording.
|
|
202
|
+
- **`spawn.hooks` is documented as hand-pinned documentation, not as a drift tripwire.** It cannot be
|
|
203
|
+
one — `sync` spreads `spawn` forward from the base baseline, so the field carries through untouched and
|
|
204
|
+
nothing re-derives it from the app bundle. The claim was disproved by its own subject, above.
|
|
205
|
+
- **The desktop-local lane boundary is stated as a missing transport flag rather than an entrypoint
|
|
206
|
+
string.** `--sdk-url` is absent from the entire app bundle and present 41× in the agent, and the
|
|
207
|
+
`ccr-session` host resolver throws without it — so features routed through that host (including
|
|
208
|
+
`cowork_memory_context`) are *structurally* unreachable locally, not merely disabled. An entrypoint
|
|
209
|
+
test can be relaxed in one release; a flag Desktop never passes cannot be worked around agent-side.
|
|
210
|
+
- **New fidelity note: the silent-turn reminder.** Whether the agent narrates between tool calls is
|
|
211
|
+
decided by a server-delivered model capability, and that narration lands in the corpus
|
|
212
|
+
`semantic_matches` grades — so an assertion resting on the presence or wording of inter-tool text rests
|
|
213
|
+
on something an account-level capability can change. Assert the observable result instead.
|
|
214
|
+
- **New fidelity note: auto-memory is delivered through four env keys the harness never sets.** When a
|
|
215
|
+
session has an auto-memory directory, Desktop ships it to the agent as `CLAUDE_COWORK_MEMORY_PATH_OVERRIDE`,
|
|
216
|
+
`CLAUDE_COWORK_MEMORY_INDEX_CONTENT`, `CLAUDE_COWORK_MEMORY_EXTRA_GUIDELINES` and — separately gated —
|
|
217
|
+
`CLAUDE_COWORK_MEMORY_GUIDELINES`. The last two carry **prompt text**: the agent reads them into its
|
|
218
|
+
memory prompt, so this is model-visible content rather than configuration. With no memory directory the
|
|
219
|
+
same branch sends `CLAUDE_CODE_DISABLE_AUTO_MEMORY:"1"` instead. `docs/fidelity-gaps.md` now also
|
|
220
|
+
tabulates the **three distinct gates** involved, which are easy to conflate: the one keyed on the memory
|
|
221
|
+
*directory* is not the one governing the guidelines env key, and neither is the one that gates only the
|
|
222
|
+
resolver's third branch — so a session can get a memory directory with that third gate off.
|
|
223
|
+
- **Corrected an overstated divergence in how the agent's credential reaches it.** A comment in
|
|
224
|
+
`src/runtime/host-env.ts` said real Cowork passes only `CLAUDE_CODE_OAUTH_TOKEN`, without scoping the
|
|
225
|
+
claim. Desktop does swap that env var for a file descriptor — staging the token into a `0600` temp file,
|
|
226
|
+
opening it, and unlinking it — but that wrapper is reached from the **host-loop branch only**. At
|
|
227
|
+
`container` and `microvm`, production passes the plain token exactly as this harness does. The comment
|
|
228
|
+
now records the host-loop-only scope and why the divergence is deliberately not emulated: the env var
|
|
229
|
+
remains a first-class credential source for the agent, Desktop's own helper falls back to it on I/O
|
|
230
|
+
failure, and no boundary is crossed at host-loop that was not already open.
|
|
231
|
+
- **Release-channel facts added to the recovery runbook.** `/stable` is a rollout pointer, not "latest";
|
|
232
|
+
not every published version is served (2.1.255 is 404 on both channels while its neighbours are 200),
|
|
233
|
+
so **a 404 is not evidence of a wrong channel** — use a positive control on a neighbouring version; and
|
|
234
|
+
`manifest.zst.json` is served *beside* `manifest.json`, additive rather than a migration.
|
|
235
|
+
|
|
236
|
+
### Changed
|
|
237
|
+
|
|
238
|
+
- **Parity sync to Desktop 1.46388.4** (agent **2.1.260**, unmoved from 1.46388.3). Baseline
|
|
239
|
+
`desktop-1.46388.4` written with zero unknown deltas; the whole app-bundle delta is three build chunks.
|
|
240
|
+
Two gates newly pinned and both recorded `force`/ON: `builtinToolsApprovableByAutoMode:4202409342` (the
|
|
241
|
+
unpinned sibling of `scheduledTaskToolsApprovableByAutoMode`, and one of the two gates that release
|
|
242
|
+
tools from the force-ask hook) and `cuCanUseToolEnabled:2486083521`, which had moved `off` → `ON` while
|
|
243
|
+
unpinned. Gates deliberately left unpinned now carry their reasoning in `cowork-sync.ts` rather than
|
|
244
|
+
being silently absent. Committed cassettes re-stamped rather than re-recorded, on a field-level diff of the two baselines:
|
|
245
|
+
the entire delta is `appVersion`, `capturedAt`, the `$comment` date, `provenance.asarFingerprint`,
|
|
246
|
+
`provenance.fcache.embeddedTimestamp` and the two new gate rows. Nothing a replay depends on —
|
|
247
|
+
`spawn.env`, `spawn.hooks`, `permissionMode`, `agentBinary.sha256`, the egress allowlist, the mount
|
|
248
|
+
layout — moved at all.
|
|
249
|
+
|
|
250
|
+
### Security
|
|
251
|
+
|
|
252
|
+
- **Bumped the transitive `fast-uri` 3.1.5 → 3.1.7**, clearing four Dependabot high-severity advisories.
|
|
253
|
+
Lockfile only — no direct dependency changed and no runtime behaviour is affected.
|
|
254
|
+
|
|
9
255
|
## [3.4.1] — 2026-09-05
|
|
10
256
|
|
|
11
257
|
### Documentation
|
package/DESIGN.md
CHANGED
|
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
47
47
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
48
48
|
|
|
49
49
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
50
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
50
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.275**, per `baselines/desktop-2.2553.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
|
|
51
51
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
52
52
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
53
53
|
|
|
@@ -174,9 +174,9 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
174
174
|
|
|
175
175
|
> **Machine-readable form:** the five shapes below are schema'd as `schema/protocol.v1.json`, with a golden vector pack at `fixtures/protocol/v1/` — see [docs/protocol.md](./docs/protocol.md) for scope, versioning, and how to conformance-test against them.
|
|
176
176
|
|
|
177
|
-
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS.
|
|
177
|
+
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. **This heading deliberately carries no version or baseline figures**: it used to restate the agent version and baseline of the last live pass, and those went stale independently of the note below it — at one point naming three different agent versions across two adjacent sentences. The note directly below is the single authority for which baseline and agent were actually exercised, and what the pass did and did not cover. Read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
178
178
|
|
|
179
|
-
> **Scope of that claim
|
|
179
|
+
> **Scope of that claim.** `2026-09-18 / desktop-2.2553.1` is the baseline carrying the latest live pass, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.275** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 tests — 18 passed, 1 failed, 0 skipped** (the fail is `live-matrix`'s 2-cell baseline-axis case on the OLD baselines 1.17377.2/1.18286.0, where `claude-opus-4-8` sometimes answers the prompt's question as text instead of calling `AskUserQuestion`; it passed 2 of 4 attempts on the 1.18286.0 cell and never wavered on the other — model variance on a test that predates this sync, logged for hardening); and `run examples/scenarios/` **7/7 success** on its first run (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe), then **6/7** on an accidental second run where `subagent-manifest-probe`'s sub-agent chose an absolute host path outside the connected folders and the hostloop path gate correctly denied it — the gate working as production does, on a model choice. and the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files) — `smoke-multiselect-deciderdir` runs the `--decider-llm` path, so it is live confirmation that the transport fix holds there as well as in `critique`. **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — and the previous pass HAD one skip, the hostloop `critique` case, which this pass exercised for the first time and which found a real defect (the LLM decider transport hard-failed on agent 2.1.275's two-model `-p` envelope; fixed in this release). Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`** (`smoke-l2-microvm` passed in a real VM with its own kernel, 151.9s). **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise — both reds above were re-run and both flipped. (c) `run examples/scenarios/` was accidentally run twice; the first run is the one counted, and the second's single red is described above for what it is. Separately and not a live matter: all three committed cassettes in `examples/replays/` were re-recorded against `desktop-2.2553.1` on 2026-09-18 (each reporting "no behavioural change — transcript wording only" versus the recording it replaced); the `protocol` fixture on the hermetic managed config dir and the `container` one in a sealed container; `verify-cassettes` reports `privacyScanned:true` and zero host-inventory findings on all three. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
180
180
|
|
|
181
181
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
182
182
|
|