cowork-harness 2.5.0 → 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +7 -7
- package/.claude/skills/cowork-harness/references/ci-recipe.md +17 -17
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +6 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +16 -5
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -2
- package/.claude/skills/cowork-harness/scripts/scenario.py +3 -2
- package/CHANGELOG.md +119 -0
- package/DESIGN.md +2 -2
- package/README.md +16 -16
- package/SPEC.md +2 -2
- package/baselines/desktop-1.40609.0.json +878 -0
- package/baselines/provisioning/rootfs-provisioning.json +32 -40
- package/dist/cli.js +8 -1
- package/dist/run/cassette.js +4 -3
- package/dist/run/chat-result.js +1 -1
- package/dist/run/chat.js +15 -1
- package/dist/run/execute.js +15 -6
- package/dist/run/hook-events.js +51 -0
- package/dist/run/skill-flag-surface.js +11 -0
- package/dist/run/verdict.js +17 -8
- package/dist/runtime/argv.js +16 -1
- package/dist/runtime/lima.js +37 -3
- package/dist/runtime/protocol.js +85 -13
- package/dist/scan.js +1 -0
- package/dist/sync/cowork-sync.js +39 -2
- package/dist/types.js +11 -3
- package/docs/cassette.md +2 -2
- package/docs/discovery.md +1 -1
- package/docs/fidelity-gaps.md +78 -6
- package/docs/maintenance.md +1 -1
- package/docs/scenario.md +6 -3
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/examples/sessions/l0-plugin-delivery.yaml +6 -0
- package/package.json +1 -1
- package/python/test_scenario_lint.py +1 -1
- package/schema/run-result.json +2 -2
- package/schema/scenario.schema.json +6 -2
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version:
|
|
7
|
-
tracks-harness: cowork-harness
|
|
6
|
+
version: 3.0.0
|
|
7
|
+
tracks-harness: cowork-harness 3.0.0 (baseline desktop-1.40609.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness
|
|
29
|
-
> `desktop-1.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.0` (baseline
|
|
29
|
+
> `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,13 +42,13 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.0"`. **Pin `@^3.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
49
49
|
— upgrade rather than work around it, since this skill's `file:line` pointers and flag names track the floor.
|
|
50
50
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
51
|
-
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
51
|
+
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
|
|
52
52
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
53
53
|
- **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
|
|
54
54
|
|
|
@@ -679,7 +679,7 @@ than a stuck `"running"`.)
|
|
|
679
679
|
### Place assertions in the right CI lane
|
|
680
680
|
|
|
681
681
|
CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
|
|
682
|
-
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@
|
|
682
|
+
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
|
|
683
683
|
PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
|
|
684
684
|
the four-stage pipeline.
|
|
685
685
|
|
|
@@ -1,23 +1,23 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
7
7
|
|
|
8
8
|
```yaml
|
|
9
|
-
- uses: yaniv-golan/cowork-harness@
|
|
9
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
10
10
|
with:
|
|
11
11
|
command: replay
|
|
12
12
|
path: cassettes/
|
|
13
|
-
version: "^
|
|
13
|
+
version: "^3" # hold the major; see below
|
|
14
14
|
```
|
|
15
15
|
|
|
16
|
-
**These recipes pin `version: "^
|
|
16
|
+
**These recipes pin `version: "^3"`.** The Action's `version` input *defaults* to `latest`, which means a
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "
|
|
20
|
+
(e.g. `version: "3.0.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.247 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
41
41
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
42
42
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -47,11 +47,11 @@ jobs:
|
|
|
47
47
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
48
48
|
# Background on the provenance chain: the "Agent-binary provenance" section of
|
|
49
49
|
# https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
|
|
50
|
-
- uses: yaniv-golan/cowork-harness@
|
|
50
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
51
51
|
with:
|
|
52
52
|
command: run
|
|
53
53
|
path: scenarios/
|
|
54
|
-
version: "^
|
|
54
|
+
version: "^3"
|
|
55
55
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
56
56
|
```
|
|
57
57
|
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.0.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -119,18 +119,18 @@ Action has no input for, and it creates a coupling nothing checks:
|
|
|
119
119
|
you have a reason:
|
|
120
120
|
|
|
121
121
|
```yaml
|
|
122
|
-
- uses: yaniv-golan/cowork-harness@
|
|
122
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
123
123
|
with:
|
|
124
124
|
command: lint
|
|
125
125
|
path: scenarios/
|
|
126
|
-
version: "^
|
|
127
|
-
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any
|
|
126
|
+
version: "^3" # holds the major
|
|
127
|
+
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 3.x satisfies that
|
|
128
128
|
```
|
|
129
129
|
|
|
130
130
|
**If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
|
|
131
131
|
floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
|
|
132
132
|
recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
|
|
133
|
-
major instead — `version: "^
|
|
133
|
+
major instead — `version: "^3"`, which is what the steps above use — keeping the floor's intent while
|
|
134
134
|
stopping at the major boundary. An exact
|
|
135
135
|
pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
|
|
136
136
|
moment a recipe adopts a newer flag.
|
|
@@ -147,7 +147,7 @@ The split is not just about tokens — it decides **where each lane can run**:
|
|
|
147
147
|
Docker, no agent binary** — runs on a stock GitHub Actions runner. Evaluates **content** assertions —
|
|
148
148
|
`transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
|
|
149
149
|
`max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
|
|
150
|
-
`allow_permissive_auto_allow` / `allow_missing_capability` / `
|
|
150
|
+
`allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
|
|
151
151
|
`allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
|
|
152
152
|
`question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
153
153
|
(`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
|
|
@@ -322,7 +322,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
322
322
|
|
|
323
323
|
## GitHub Actions sketch
|
|
324
324
|
|
|
325
|
-
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@
|
|
325
|
+
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v3` does
|
|
326
326
|
in one step (see the top of this doc) — reach for this form when you need independent per-command
|
|
327
327
|
gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
|
|
328
328
|
equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
|
|
@@ -342,7 +342,7 @@ jobs:
|
|
|
342
342
|
with: { node-version: '24' }
|
|
343
343
|
- uses: actions/setup-python@v5
|
|
344
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^
|
|
345
|
+
- run: npm i -g "cowork-harness@^3.0.0"
|
|
346
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +371,7 @@ jobs:
|
|
|
371
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
372
|
fi
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^
|
|
374
|
+
run: npm i -g "cowork-harness@^3.0.0"
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness
|
|
3
|
+
Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -20,6 +20,11 @@ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.379
|
|
|
20
20
|
- A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
|
|
21
21
|
true` — with no container around the native file tools, that combination gives the agent genuine,
|
|
22
22
|
software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
|
|
23
|
+
- A `protocol` scenario staging a plugin that declares runnable hooks needs `allow_host_hooks: true`
|
|
24
|
+
(`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
|
|
25
|
+
hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
|
|
26
|
+
Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
|
|
27
|
+
does not fall back to the default.
|
|
23
28
|
- **Set the tier in the scenario's `fidelity:` field — not a flag.** `--fidelity` is accepted only by
|
|
24
29
|
`skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
|
|
25
30
|
positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.0`
|
|
4
|
+
(baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -98,6 +98,17 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
|
|
|
98
98
|
# around hostloop's native file tools, that combination gives the
|
|
99
99
|
# agent genuine, software-checked-only host filesystem access.
|
|
100
100
|
# Read-only folders and folder-less runs need no opt-in.
|
|
101
|
+
|
|
102
|
+
allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
|
|
103
|
+
# declares runnable hooks (`<plugin>/hooks/hooks.json`): L0 passes
|
|
104
|
+
# --plugin-dir, so the CLI executes those hooks as NATIVE HOST
|
|
105
|
+
# processes under your account, with no container sandbox. A plugin
|
|
106
|
+
# that declares no hooks needs no opt-in, and a misplaced root-level
|
|
107
|
+
# `hooks.json` cannot execute so it does not trigger the gate.
|
|
108
|
+
# Use `--fidelity container` to run them sandboxed instead.
|
|
109
|
+
# NEEDS cowork-harness >= 3.0.0. The loader is a strict object, so an
|
|
110
|
+
# OLDER CLI does not default it — it hard-errors
|
|
111
|
+
# `Unrecognized key: "allow_host_hooks"` and exits 2.
|
|
101
112
|
```
|
|
102
113
|
|
|
103
114
|
Relative paths resolve from the file's own directory, so a scenario + session + referenced files
|
|
@@ -362,7 +373,7 @@ same set live from the schema.
|
|
|
362
373
|
| `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
|
|
363
374
|
| `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
|
|
364
375
|
| `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
|
|
365
|
-
| `
|
|
376
|
+
| `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
|
|
366
377
|
| `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
|
|
367
378
|
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
368
379
|
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
@@ -404,7 +415,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
404
415
|
| `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
|
|
405
416
|
| `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
|
|
406
417
|
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
407
|
-
| `
|
|
418
|
+
| `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
|
|
408
419
|
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
409
420
|
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
410
421
|
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
@@ -443,7 +454,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
|
|
|
443
454
|
`skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
|
|
444
455
|
`max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
|
|
445
456
|
(`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
|
|
446
|
-
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `
|
|
457
|
+
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
|
|
447
458
|
`allow_stall` are also kept on replay, evaluated as no-op passes.
|
|
448
459
|
|
|
449
460
|
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness
|
|
5
|
+
Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"keys": [
|
|
4
4
|
"all_tasks_completed",
|
|
5
5
|
"allow_delete_in",
|
|
6
|
-
"
|
|
6
|
+
"allow_l0_host_config_contamination",
|
|
7
7
|
"allow_missing_capability",
|
|
8
8
|
"allow_outputs_delete",
|
|
9
9
|
"allow_permissive_auto_allow",
|
|
@@ -83,6 +83,7 @@
|
|
|
83
83
|
"vm_path_denied"
|
|
84
84
|
],
|
|
85
85
|
"topLevelKeys": [
|
|
86
|
+
"allow_host_hooks",
|
|
86
87
|
"allow_host_writes",
|
|
87
88
|
"answers",
|
|
88
89
|
"assert",
|
|
@@ -101,7 +102,7 @@
|
|
|
101
102
|
],
|
|
102
103
|
"verdictModifierKeys": [
|
|
103
104
|
"allow_delete_in",
|
|
104
|
-
"
|
|
105
|
+
"allow_l0_host_config_contamination",
|
|
105
106
|
"allow_missing_capability",
|
|
106
107
|
"allow_outputs_delete",
|
|
107
108
|
"allow_permissive_auto_allow",
|
|
@@ -173,7 +173,7 @@ LANE_REMOTE_INCOMPATIBLE_KEYS = {"present_files_called", "no_scratchpad_leak", "
|
|
|
173
173
|
# verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
|
|
174
174
|
VERDICT_MODIFIER_KEYS = {
|
|
175
175
|
"allow_permissive_auto_allow",
|
|
176
|
-
"
|
|
176
|
+
"allow_l0_host_config_contamination",
|
|
177
177
|
"allow_missing_capability",
|
|
178
178
|
"allow_stall",
|
|
179
179
|
"allow_undelivered_deliverables",
|
|
@@ -262,7 +262,8 @@ _EMBEDDED_TOP_LEVEL_KEYS = {
|
|
|
262
262
|
"assert",
|
|
263
263
|
"skills", # opt-in skill-staleness hash scope
|
|
264
264
|
"requires_capabilities", # Fix 4b: scenario-level required-capability declaration (pre-flight gate)
|
|
265
|
-
"allow_host_writes",
|
|
265
|
+
"allow_host_writes",
|
|
266
|
+
"allow_host_hooks", # protocol consent: a staged plugin's hooks run as NATIVE HOST processes # hostloop native-split: consent for a writable connected folder (pre-run gate)
|
|
266
267
|
}
|
|
267
268
|
|
|
268
269
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,125 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.0.0] — 2026-08-29
|
|
10
|
+
|
|
11
|
+
### Breaking
|
|
12
|
+
|
|
13
|
+
- **`l0_plugin_divergence` is renamed `l0_host_config_contamination`**, and its modifier
|
|
14
|
+
`allow_l0_plugin_divergence` is renamed `allow_l0_host_config_contamination`. The `RunResult` field
|
|
15
|
+
`l0PluginDivergence` becomes `l0HostConfigContamination`. Verdict-signal codes and the scenario schema
|
|
16
|
+
are covered surfaces ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)), so this is
|
|
17
|
+
a MAJOR bump. The old name described plugin *delivery* diverging at L0; delivery now works, and what the
|
|
18
|
+
signal actually reports is that the run read the operator's real config dir. A scenario asserting the old
|
|
19
|
+
key must rename it; a consumer keying on the old code must too.
|
|
20
|
+
|
|
21
|
+
- **The signal's firing conditions changed with it.** It fired when a session declared plugin dirs; it now
|
|
22
|
+
fires when `protocol` reads the operator's real config dir — tested on the dir the agent will actually
|
|
23
|
+
read, so a pinned `plugins.config_dir` is caught even with the managed branch nominally active. A
|
|
24
|
+
`skills.local`-only protocol run that previously passed can now fail, and a sealed protocol run with
|
|
25
|
+
plugins that previously failed now passes.
|
|
26
|
+
|
|
27
|
+
|
|
28
|
+
### Added
|
|
29
|
+
|
|
30
|
+
- **`allow_host_hooks` (scenario key) and `--allow-host-hooks` (`chat` / `skill`).** Consent to
|
|
31
|
+
running a staged plugin's hooks as native host processes at `protocol`. Top-level like
|
|
32
|
+
`allow_host_writes` rather than a verdict modifier — it gates the SPAWN, it does not suppress a signal.
|
|
33
|
+
|
|
34
|
+
**Version floor: `allow_host_hooks` needs cowork-harness ≥ 3.0.0.** The scenario loader is a `z.strictObject`, so an older CLI does NOT fall back to the default — it hard-errors `Unrecognized key: "allow_host_hooks"` and exits 2 (verified). Adopting the key is a floor bump for every consumer of that scenario.
|
|
35
|
+
|
|
36
|
+
- **A committed live probe for L0 plugin delivery** — `examples/probes/l0-plugin-delivery.scenario.yaml`
|
|
37
|
+
plus its fixture and session. It asserts the plugin under test reaches the agent's init inventory at
|
|
38
|
+
`protocol`, which no unit test can see: the argv builder was correct in isolation and the defect was a
|
|
39
|
+
missing call site. It pins the sealed config-dir branch deliberately — off it, the operator's own
|
|
40
|
+
installed plugins are in the inventory and a name collision would satisfy the assertion with or without
|
|
41
|
+
`--plugin-dir`, measuring the machine instead of the argv. The fixture also ships an agent, deliberately
|
|
42
|
+
unasserted: there is no `agent_available` assertion key (tool / skill / connector have one, agents do
|
|
43
|
+
not), and naming that gap is better than inventing a key to hide it.
|
|
44
|
+
|
|
45
|
+
### Changed
|
|
46
|
+
|
|
47
|
+
- **`protocol` (L0) now passes `--plugin-dir`, so a declared plugin or skill dir is actually delivered.**
|
|
48
|
+
It previously passed no `--plugin-dir` and `local_plugins` never reached the generated `settings.json`,
|
|
49
|
+
so **the positional argument was silently inert**: `cowork-harness chat <dir> --fidelity protocol` (and
|
|
50
|
+
the `skill` / `probe-dispatch` equivalents) measured whatever the operator had installed rather than the
|
|
51
|
+
tree they passed. Live-verified: a bare skill dir and a plugin root both register at L0 now, with their
|
|
52
|
+
skills and their declared agents. Expect scenarios that previously passed *vacuously* — asserting against
|
|
53
|
+
an agent that never had the plugin — to start failing honestly.
|
|
54
|
+
|
|
55
|
+
- **A plugin's hooks and MCP servers now run at `protocol`, and hooks require consent.** Loading a plugin
|
|
56
|
+
means the CLI executes its `<plugin>/hooks/hooks.json` as **native host processes** — the operator's
|
|
57
|
+
account and environment, no container sandbox — and opens its declared MCP servers. `protocol` therefore
|
|
58
|
+
refuses to spawn when a staged plugin declares runnable hooks unless the scenario sets
|
|
59
|
+
`allow_host_hooks: true` (or `--allow-host-hooks` for `chat`/`skill`); a plugin without hooks
|
|
60
|
+
needs no opt-in, and a misplaced root-level `hooks.json` (which cannot execute) does not trigger it. A
|
|
61
|
+
per-run disclosure is printed even when consent was given. Mirrors the existing `allow_host_writes` gate.
|
|
62
|
+
|
|
63
|
+
- **`protocol` takes the managed config dir when any credential is in the environment.**
|
|
64
|
+
`CLAUDE_CODE_OAUTH_TOKEN` and `ANTHROPIC_AUTH_TOKEN` now select it alongside `ANTHROPIC_API_KEY`, and the
|
|
65
|
+
token is injected into the agent's env (a managed config dir with no credential yields "Not logged in").
|
|
66
|
+
This severs host plugin/skill/MCP discovery — verified: a token-only L0 run delivered the plugin under
|
|
67
|
+
test and **zero** host plugins. `COWORK_MANAGED_CONFIG=0` suppresses only the token-derived branch,
|
|
68
|
+
never the `ANTHROPIC_API_KEY` CI path; an unrecognized value is now rejected rather than silently
|
|
69
|
+
selecting the managed branch (`COWORK_MANAGED_CONFIG=false` previously did).
|
|
70
|
+
|
|
71
|
+
- **`l0_plugin_divergence` now reports contamination rather than the `--plugin-dir` layout.** Delivery is
|
|
72
|
+
fixed, so the signal fires when L0 runs against the operator's **real** config dir, and it tests the dir
|
|
73
|
+
the agent will actually read — a pinned `plugins.config_dir` reaches host discovery with the managed
|
|
74
|
+
branch nominally active. It stays a hard fail: nothing else catches this (the host-inventory scan runs
|
|
75
|
+
only at cassette-record time, and `host_path_leak`'s default-fail is skipped at this tier).
|
|
76
|
+
|
|
77
|
+
- **`workflow-authoring` added to the built-in skill roster** (`KNOWN_BUILTIN_SKILLS`), measured against a
|
|
78
|
+
sealed `protocol` run rather than a contaminated one.
|
|
79
|
+
|
|
80
|
+
- **Platform baseline `desktop-1.40609.0` (agent `2.1.247`).** The Cowork system prompt, both sub-agent
|
|
81
|
+
appends, the egress allowlist, `spawn.env`, `tools[]`/`allowedTools` and `mountLayout` are all re-derived
|
|
82
|
+
unchanged. The one `sync` delta was a pure refactor: W2's `CLAUDE_CODE_ENTRYPOINT` deployment ternary is
|
|
83
|
+
now a hoisted helper, and W1 hard-sets `local-agent` after the spread either way, so the derived value
|
|
84
|
+
does not move. The VM rootfs is re-captured — node `v22.22.3` → `v22.23.2`, pip 144 → 136 packages
|
|
85
|
+
(`xlrd` added; the nine removals are Ubuntu system packages, no analysis library dropped).
|
|
86
|
+
|
|
87
|
+
### Fixed
|
|
88
|
+
|
|
89
|
+
- **`microvm` resolves its agent binary through the shared resolver instead of deriving the path itself.**
|
|
90
|
+
It read `agentBinary.stagedPath` raw and handed it to the guest mount, so a pin that Claude Desktop had
|
|
91
|
+
pruned surfaced as `env: 'claude': No such file or directory` and exit 127 — the least informative
|
|
92
|
+
message possible for a condition the other three tiers name precisely. The quieter half mattered more:
|
|
93
|
+
the raw path also skipped `verifiedElf`, leaving the one tier that actually **executes** the ELF in a VM
|
|
94
|
+
as the only one not verifying it against the baseline pin, while `container` hard-fails on the same
|
|
95
|
+
mismatch. Routing it through `resolveAgentBinary` restores all three safeguards at once — existence
|
|
96
|
+
check, sha verification and the pruned-binary fallback — and removes a fourth derivation of a rule
|
|
97
|
+
`baseline.ts` already documented `microvm` as following. Resolution happens **before** the
|
|
98
|
+
already-Running reuse short-circuit, since a VM created while the binary was present keeps a mount at
|
|
99
|
+
the pruned path — the originally reported state. `vm status` / `vm prune` / `doctor` keep working when
|
|
100
|
+
the binary is missing (that is when an operator reaches for them) and degrade to the pinned path rather
|
|
101
|
+
than throwing.
|
|
102
|
+
|
|
103
|
+
- **`sync` resolves a spawn-env value expression that is a hoisted one-line helper** (`uH(n.type)`), where
|
|
104
|
+
it previously refused the baseline. The branch is narrow by construction: the argument must be a `.type`
|
|
105
|
+
member and the callee body must be exactly the literal deployment ternary over its own parameter, resolved
|
|
106
|
+
in the window's own chunk — a two-character minified name hopped against the joined bundle lands on an
|
|
107
|
+
unrelated helper. Anything else stays unresolvable rather than guessed.
|
|
108
|
+
|
|
109
|
+
- **`provenance.asarGateIds` now includes the gate-defaults map's bare-numeric keys.** The scan matched
|
|
110
|
+
quoted literals only, so any gate that is keyed in that map and never read through a quoted id was
|
|
111
|
+
missing — `provenance.asarGateIds` 252 → 291, and the 1.40609.0 delta corrects from +27/−4 to +30/−4.
|
|
112
|
+
This is not the bare-*number* scan the extractor deliberately rejects: that matches any numeric and adds
|
|
113
|
+
1687 ids over the same bundle, while the defaults-map entry shape adds 39.
|
|
114
|
+
|
|
115
|
+
### Documentation
|
|
116
|
+
|
|
117
|
+
- **A plugin's declared MCP servers are a documented fidelity gap.** Production replaces them with
|
|
118
|
+
zero-tool SDK stubs named `plugin:<plugin>:<server>` — remote (`url` + http/sse) unconditionally, local
|
|
119
|
+
and `.mcpb` under an MCP policy — while the harness stages plugins with `--plugin-dir` and lets the CLI
|
|
120
|
+
open the real ones. A plugin under test therefore sees a tool surface production would not give it, in
|
|
121
|
+
both directions and silently. Not modeled: the remote rule is mechanically reproducible, but the
|
|
122
|
+
local/`.mcpb` rule is conditioned on Desktop policy state the harness has no source for.
|
|
123
|
+
|
|
124
|
+
- **The auto-mode permission rubric gap is tier-independent.** `docs/fidelity-gaps.md` scopes it to both
|
|
125
|
+
loops rather than VM-loop only, and records that the rubric's Filesystem section is tier-dependent
|
|
126
|
+
model-visible text — a different kind of divergence from a permission verdict.
|
|
127
|
+
|
|
9
128
|
## [2.5.0] — 2026-08-28
|
|
10
129
|
|
|
11
130
|
### Added
|
package/DESIGN.md
CHANGED
|
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
47
47
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
48
48
|
|
|
49
49
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
50
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
50
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.247**, per `baselines/desktop-1.40609.0.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
|
|
51
51
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
52
52
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
53
53
|
|
|
@@ -176,7 +176,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
176
176
|
|
|
177
177
|
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
178
178
|
|
|
179
|
-
> **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is
|
|
179
|
+
> **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is no longer the newest committed baseline: **one** baseline has shipped since (`1.40609.0`), **one** of which moved the agent ELF, most recently to **2.1.247** — so the newest baseline is **not** live-verified, and this paragraph describes the 1.37937.1 pass only. The pass ran on agent `2.1.246` (the staged VM ELF and the native `.app` were both at that version) and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe, resume continuity on the native binary, critique at the unpinned tier, uploads readability, sub-agent WebSearch capture, and the discovery-server declaration check). It was one invocation of `npm run test:live` — **4 suites, 19 assertions, 19 green / 0 skipped**. **Nothing was gated out**, which is the part worth stating: every `describe` in this lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported zero skips, and the 19 that ran are exactly the 19 the suite enumerates. **Scope-out, so this is not read as more than it is.** (a) The assertion population is 19 here against the 24 recorded for the 1.32885.1 pass; cases have been retired and consolidated since (the `live-outputs-delete` whole-line-`#`-comment case was retired in 1.25.0 after its pinned command stopped being executed by the model), so the two counts are not comparable and the drop is not coverage lost in this pass. (b) The `boundary-check` sandbox proof and the example-scenario suite were part of the 1.32885.1 stamp and were **NOT** run here — this paragraph claims `npm run test:live` only. (c) A live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: the three committed example cassettes (`example-pdf-skill` at `container`, `example-multiselect-gate` at `protocol`, `hostloop-computer-links` at `hostloop`) were re-recorded against this baseline in the same change, so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
180
180
|
|
|
181
181
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
182
182
|
|