cowork-harness 1.6.0 → 1.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +49 -18
- package/.claude/skills/cowork-harness/references/ci-recipe.md +10 -9
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +2 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +12 -10
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +250 -0
- package/CONTRIBUTING.md +3 -0
- package/README.md +39 -24
- package/SPEC.md +14 -4
- package/baselines/desktop-1.24012.0.json +405 -0
- package/baselines/desktop-1.24012.1.json +485 -0
- package/dist/agent/session.js +9 -3
- package/dist/agent/timeline.js +76 -3
- package/dist/assert.js +4 -2
- package/dist/baseline.js +20 -0
- package/dist/cli.js +194 -33
- package/dist/critique/command.js +61 -8
- package/dist/critique/evidence.js +65 -41
- package/dist/critique/limitations.js +118 -0
- package/dist/critique/package-evidence.js +29 -18
- package/dist/dotenv.js +38 -3
- package/dist/errors.js +12 -0
- package/dist/hostloop/workspace-handler.js +35 -1
- package/dist/io.js +15 -7
- package/dist/run/cassette.js +4 -0
- package/dist/run/chat-result.js +6 -2
- package/dist/run/chat.js +39 -6
- package/dist/run/diff.js +26 -5
- package/dist/run/doctor.js +2 -1
- package/dist/run/execute.js +158 -51
- package/dist/run/inspect-view.js +11 -4
- package/dist/run/latest-run.js +33 -7
- package/dist/run/migrate-run-dir.js +817 -0
- package/dist/run/renderer.js +11 -3
- package/dist/run/run-index.js +225 -45
- package/dist/run/run.js +13 -4
- package/dist/run/runs-gc.js +45 -8
- package/dist/run/scaffold.js +28 -4
- package/dist/run/skill-files.js +18 -1
- package/dist/run/trace-view.js +72 -12
- package/dist/run/turn-events.js +34 -0
- package/dist/run/turn-layout.js +189 -0
- package/dist/run/verdict.js +17 -2
- package/dist/runtime/host-env.js +8 -2
- package/dist/runtime/hostloop.js +58 -17
- package/dist/runtime/image-capabilities.js +22 -8
- package/dist/runtime/resource-sampler.js +20 -6
- package/dist/sync/cowork-sync.js +137 -17
- package/dist/types.js +17 -2
- package/docs/cassette.md +5 -5
- package/docs/chat.md +4 -3
- package/docs/critique.md +54 -7
- package/docs/debugging.md +38 -3
- package/docs/fidelity-gaps.md +51 -11
- package/docs/gotchas.md +7 -1
- package/docs/maintenance.md +2 -2
- package/docs/run-status.md +4 -3
- package/docs/scenario.md +21 -5
- package/docs/session.md +1 -1
- package/docs/stats.md +29 -7
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/llms.txt +1 -1
- package/package.json +1 -1
- package/python/cowork_harness.py +35 -11
- package/schema/run-result.json +10 -4
- package/schema/scenario.schema.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.7.0
|
|
7
|
+
tracks-harness: cowork-harness 1.7.0 (baseline desktop-1.24012.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
26
|
-
> `desktop-1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.7.0` (baseline
|
|
26
|
+
> `desktop-1.24012.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,9 +39,30 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
43
|
-
|
|
44
|
-
What the ≥ 1.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.7.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.7.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.7.0"`. **Pin `@>=1.7.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
|
+
|
|
44
|
+
What the ≥ 1.7.0 floor gates, by release:
|
|
45
|
+
|
|
46
|
+
- **1.7.0 (per-turn run-directory layout, single shape):** every run dir — `run`/`skill`/`chat`,
|
|
47
|
+
single-turn or multi-turn (any `--session-id` + `--resume`, and **every `critique`**: task turn +
|
|
48
|
+
reflection turn) — writes each turn's `result.json` / `run.jsonl` / `trace.json` / `resources.jsonl`
|
|
49
|
+
into **`turns/<N>/`**, written once and never renamed. **There is no root compat copy of anything —
|
|
50
|
+
`<run-dir>/result.json` does not exist.** **When a user asks about a run, point them at
|
|
51
|
+
`turns/1/result.json`** (the only completion on a single-turn dir) — or on a `critique` dir, better yet
|
|
52
|
+
at `result.graded.json` / `trace.graded.json`, the role-stable aliases `critique` writes (the graded
|
|
53
|
+
task turn, never the reflection one). `events.jsonl` / `timeline.jsonl` stay cumulative at the run-dir
|
|
54
|
+
root, always. A dir written before this layout existed is refused LOUD, by name, naming the shape found
|
|
55
|
+
— never silently misread as if it were `turns/1/`. **Convert it in place with
|
|
56
|
+
`cowork-harness migrate-run-dir` (dry-run by default); do NOT tell the user to re-run or re-record,
|
|
57
|
+
which throws away history the migrator recovers.** Its `events.jsonl` still fully supports `trace`,
|
|
58
|
+
which is why every refusal points there. `verify-run` REFUSES a multi-turn dir rather than certifying
|
|
59
|
+
the wrong turn, and `trace` shows the latest turn with a `::notice::` when earlier turns exist.
|
|
60
|
+
|
|
61
|
+
- **1.7.0 (limitation provenance):** every `critique` limitation in `critique --help` and
|
|
62
|
+
docs/critique.md is tagged with WHY it exists — `[structural]` (permanent), `[unverified]` (unproven,
|
|
63
|
+
**not** known-impossible), `[deliberate]`, `[not-built]`. **Read the tag before telling a user to
|
|
64
|
+
design around a limitation.** In particular `critique`'s container-tier pin is `[unverified]`, not
|
|
65
|
+
permanent: do not advise building a second hostloop test lane on the assumption it can never lift.
|
|
45
66
|
|
|
46
67
|
- **1.6.0 (`critique`'s skill-flag parity):** `critique` accepts most `skill` flags under the same names,
|
|
47
68
|
which is what makes a document-analysis skill (cap table, deck, financial model, transcript)
|
|
@@ -81,7 +102,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
81
102
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
82
103
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
83
104
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
84
|
-
- **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) —
|
|
105
|
+
- **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
|
|
85
106
|
|
|
86
107
|
## Orient — the three loops
|
|
87
108
|
|
|
@@ -113,7 +134,7 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
113
134
|
> `node_modules/cowork-harness/docs/<name>.md` before assuming a pointer dangles. A **plugin**
|
|
114
135
|
> install loads a trimmed source-only cache where those pointers genuinely do dangle.
|
|
115
136
|
|
|
116
|
-
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
|
|
137
|
+
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · migrate-run-dir · lint ·
|
|
117
138
|
lint-skill · analyze-skill · probe-dispatch ·
|
|
118
139
|
verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
119
140
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
@@ -155,7 +176,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
|
|
|
155
176
|
> is your **working tree**, so an uncommitted edit to an already-tracked file *is* tested — you needn't
|
|
156
177
|
> commit to iterate. Only brand-new (untracked) files must be `git add`-ed to appear. Commit before you
|
|
157
178
|
> record the **locking cassette**, though: real Cowork ships the *committed* tree, so a green on
|
|
158
|
-
> uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder
|
|
179
|
+
> uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder mounts *empty* and the agent reports "the skill isn't
|
|
159
180
|
> installed" then did the work itself — a green-looking run where the skill never loaded. That now
|
|
160
181
|
> **hard-fails** (`BoundaryError`, exit 3) naming the dir, and a partially-tracked folder emits a loud
|
|
161
182
|
> `::notice:: [stage]` listing the excluded files. Fix: `git add` the skill, or `COWORK_HARNESS_GITSET=0`
|
|
@@ -372,7 +393,7 @@ cassette — has its own recipe:
|
|
|
372
393
|
"<q>=Israeli company"` binds whichever option starts with `Israeli company`. It is uniqueness-guarded and
|
|
373
394
|
**fails loud** if the anchor ever matches two options (the documented trade: drift-tolerance, not strict
|
|
374
395
|
CI reproducibility — for that, pin a full exact label or a free-text `answer:`).
|
|
375
|
-
3. **Budget ~1 re-run per file.** If a gate whiffs, the run
|
|
396
|
+
3. **Budget ~1 re-run per file.** If a gate whiffs, the run does not vanish — it exits non-zero but
|
|
376
397
|
**salvages a PARTIAL run** (the extraction the agent already did is written to disk). So the cost of a
|
|
377
398
|
missed gate is one re-run with a better `--intent` or a scripted answer, not a lost paid run.
|
|
378
399
|
4. **Inspect the outputs to judge correctness.** `cowork-harness inspect <run-dir>` shows what the run
|
|
@@ -435,6 +456,14 @@ Recognize these before "fixing" a non-bug:
|
|
|
435
456
|
`transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
|
|
436
457
|
run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
|
|
437
458
|
it's valid.
|
|
459
|
+
- **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
|
|
460
|
+
infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
|
|
461
|
+
rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
|
|
462
|
+
the fail-severity `infra_error`, where a **supervising process** died and contaminated everything. Note
|
|
463
|
+
a model-requested `timeout_ms` expiry is *not* this: it returns the command's partial output with
|
|
464
|
+
`Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
|
|
465
|
+
failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
|
|
466
|
+
suspiciously empty.
|
|
438
467
|
- **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
|
|
439
468
|
`RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
|
|
440
469
|
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
@@ -638,12 +667,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
638
667
|
`mcp__workspace__bash` alias, so a "sub-agent used no shell" check must glob **both** `Bash` and
|
|
639
668
|
`mcp__workspace__*` to hold across every tier.
|
|
640
669
|
|
|
641
|
-
6. **`dispatch_count_max` is
|
|
642
|
-
assertion: passing means "happened to dispatch ≤N this
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
`
|
|
646
|
-
only
|
|
670
|
+
6. **`dispatch_count_max` is your author-chosen budget UNDER Cowork's production cap, not a
|
|
671
|
+
reproduction of it.** It's a post-hoc count assertion: passing means "happened to dispatch ≤N this
|
|
672
|
+
run." Cowork DOES cap `Task` fan-out **agent-side** (`taskRegistry`: concurrent **20** /
|
|
673
|
+
per-session **200**, landed 2.1.212/2.1.217) — SEPARATE from the scheduled-task session limiter
|
|
674
|
+
(gate `1648655587`'s `{perTask:1, global:3}`, a different mechanism; binary-verified, `SPEC.md` §10
|
|
675
|
+
— repo-only). The harness **inherits** the production cap by spawning the real agent binary, so a
|
|
676
|
+
`dispatch_count_max` pass means "your tighter budget held," not "near a real limit"; use it to catch
|
|
677
|
+
a fan-out you don't want.
|
|
647
678
|
|
|
648
679
|
7. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
|
|
649
680
|
assertions need a sandboxed tier (`container`+). Good: this one fails loud by design.
|
|
@@ -727,7 +758,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
727
758
|
|
|
728
759
|
17. **Editing `scenarios/*.yaml` `assert:` does NOT change a plain `replay`.** *Why:* `replay` evaluates the
|
|
729
760
|
assertions **frozen in the cassette** by default — it is byte-deterministic and ignores the working tree (so
|
|
730
|
-
a committed cassette can't silently re-interpret against an uncommitted YAML). This
|
|
761
|
+
a committed cassette can't silently re-interpret against an uncommitted YAML). This is *loud* rather than a *silent*
|
|
731
762
|
no-op; now plain `replay` prints a `::notice::` when a sibling's `assert:` differs and points you at the fix.
|
|
732
763
|
*Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
|
|
733
764
|
`--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.7.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.7.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -32,7 +32,7 @@ jobs:
|
|
|
32
32
|
- uses: actions/checkout@v4
|
|
33
33
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
34
34
|
run: |
|
|
35
|
-
V=2.1.
|
|
35
|
+
V=2.1.217 # match your scenario's pinned baseline's agentVersion
|
|
36
36
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
37
37
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
38
38
|
# verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.7.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.7.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.7.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -263,8 +263,9 @@ JSON, not the human-readable text (which is explicitly NOT stable).
|
|
|
263
263
|
|
|
264
264
|
A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
|
|
265
265
|
set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
|
|
266
|
-
artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl
|
|
267
|
-
`
|
|
266
|
+
artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl` and
|
|
267
|
+
`egress.log` at the root, plus each turn's `run.jsonl` / `trace.json` / `result.json` under `turns/<N>/`
|
|
268
|
+
(a single-turn run has just `turns/1/`; there is no root compat copy of any of these). Digest one with `cowork-harness trace <run-id | dir>`.
|
|
268
269
|
Secrets are scrubbed from every persisted log by value.
|
|
269
270
|
|
|
270
271
|
## Don't assume a fixed assertion count across lanes
|
|
@@ -293,7 +294,7 @@ does **not** imply the recording is still valid. Each replay result carries `sta
|
|
|
293
294
|
(A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
|
|
294
295
|
check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
|
|
295
296
|
`·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
|
|
296
|
-
gate (`ok:false`) — but it
|
|
297
|
+
gate (`ok:false`) — but it is class-AWARE on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
|
|
297
298
|
`format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
|
|
298
299
|
while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
|
|
299
300
|
fail it either way.)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.7.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -122,7 +122,7 @@ hands `fn` exactly this dict.
|
|
|
122
122
|
|
|
123
123
|
- On `skill`: `fail | prompt | first`.
|
|
124
124
|
- On `run`: `fail | first` (`prompt` rejected — it would break determinism).
|
|
125
|
-
- On `critique
|
|
125
|
+
- On `critique`: `fail | first` (`prompt` rejected — there is no TTY inside the spawned turn).
|
|
126
126
|
- **`llm` is NOT an `--on-unanswered` value.** The bare flag `--on-unanswered llm` is rejected; use
|
|
127
127
|
`--decider-llm` (CLI) or `on_unanswered: llm` (YAML).
|
|
128
128
|
- `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.7.0`
|
|
4
4
|
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
@@ -269,7 +269,7 @@ same set live from the schema.
|
|
|
269
269
|
| `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
|
|
270
270
|
| `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
|
|
271
271
|
| `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
|
|
272
|
-
| `dispatch_count_max: <N>` | at most N sub-agents dispatched —
|
|
272
|
+
| `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 |
|
|
273
273
|
| `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked via the `Skill` tool — evidence-unavailable (not a normal fail) if the agent's init tools have no `Skill` tool |
|
|
274
274
|
| `no_skill_triggered: <regex>` | no invoked skill id matched — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent or the `Skill` tool is unobservable |
|
|
275
275
|
| `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
|
|
@@ -348,16 +348,17 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
348
348
|
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
349
349
|
| `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
|
|
350
350
|
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
351
|
-
| `infra_error` | fail | A VM/egress sidecar
|
|
351
|
+
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
352
352
|
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
353
353
|
| `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
|
|
354
354
|
| `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
|
|
355
355
|
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
|
|
356
356
|
| `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
|
|
357
|
+
| `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
|
|
357
358
|
|
|
358
359
|
A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
|
|
359
360
|
overall run verdict and exit code — `assert result: success` alone won't catch it; check
|
|
360
|
-
`result.verdict.signals[].severity` or the run's exit code. Only the
|
|
361
|
+
`result.verdict.signals[].severity` or the run's exit code. Only the five **warn** codes are truly benign.
|
|
361
362
|
|
|
362
363
|
## Replay class
|
|
363
364
|
|
|
@@ -417,7 +418,7 @@ skill-source drift; every replay result also reports it class-tagged in `stalene
|
|
|
417
418
|
**Mixed assertions on replay:** before evaluating, `replay` strips each assertion to its replay-checkable
|
|
418
419
|
keys and drops any left empty. So `{result, egress_denied}` evaluates on replay as `{result}` alone — its
|
|
419
420
|
`egress_denied` half is removed (not AND-ed against an unreadable value); with a manifest, `file_exists`/
|
|
420
|
-
`artifact_json` are
|
|
421
|
+
`artifact_json` are not stripped. The harness is **loud in two classes**: a *full skip* (`::warning::`
|
|
421
422
|
with the count of pure live-only assertions not evaluated) and a *partial skip* (`::warning::` when a mixed
|
|
422
423
|
assertion's live-only half was dropped).
|
|
423
424
|
Two CI consequences: skipped assertions are **absent** from `results[].assertions[]` (not
|
|
@@ -529,11 +530,12 @@ reference omits). Neither list is a strict superset of the other — reach for t
|
|
|
529
530
|
11. **`subagent_declared_but_unused` fires on declared-but-didn't-use-THAT-tool**, even if the
|
|
530
531
|
sub-agent used other tools.
|
|
531
532
|
|
|
532
|
-
12. **`dispatch_count_max` is
|
|
533
|
-
and asserts on it; passing means "dispatched ≤N this
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
533
|
+
12. **`dispatch_count_max` is your author-chosen budget under Cowork's production cap, not a
|
|
534
|
+
reproduction of it.** It records the count and asserts on it; passing means "dispatched ≤N this
|
|
535
|
+
run," not "the harness capped it." Cowork DOES cap `Task` fan-out agent-side (`taskRegistry`:
|
|
536
|
+
concurrent 20 / per-session 200, landed 2.1.212/2.1.217), which the harness inherits by spawning the
|
|
537
|
+
real agent binary — SEPARATE from gate `1648655587`'s `{perTask:1, global:3}` scheduled/cron-task
|
|
538
|
+
session limiter, a different mechanism (binary-verified; SPEC §10).
|
|
537
539
|
|
|
538
540
|
13. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
|
|
539
541
|
assertions need `container`+. Fails loud by design.
|
|
@@ -182,7 +182,7 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
182
182
|
(blinded evaluator + mechanical citation checking). **Critiquing a document-analysis skill?** The probe
|
|
183
183
|
attaches nothing on its own — pass `--upload <path>` (repeatable) or `--folder <dir>` exactly as you
|
|
184
184
|
would to `skill`, or the graded run has no file and "there was no file attached" is the correct finding,
|
|
185
|
-
not a skill defect. Source flags reach both spawned turns automatically
|
|
185
|
+
not a skill defect. Source flags reach both spawned turns automatically — see SKILL.md's
|
|
186
186
|
floor list. See docs/critique.md for the full flag table, cost and limits. If you
|
|
187
187
|
prefer to build your own grader, the substrate is still here:
|
|
188
188
|
- `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,256 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.7.0] — 2026-07-22
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`migrate-run-dir` — convert pre-layout run dirs to the per-turn `turns/<N>/` layout, in place.**
|
|
14
|
+
A run dir written before the per-turn layout keeps `result.json` / `run.jsonl` / `trace.json` /
|
|
15
|
+
`resources.jsonl` at its root. Once the legacy read layer is removed, those dirs become unreadable to
|
|
16
|
+
`verify-run` / `diff` / `inspect` / `stats`; this command converts them so the history survives the
|
|
17
|
+
change instead of having to be re-run.
|
|
18
|
+
|
|
19
|
+
**Dry-run by default** — `--write` applies, and `--scenario <name>` scopes the run to a single
|
|
20
|
+
scenario so a rollout can be staged: migrate one, verify it, then do the rest. It renames rather than copies, so file mtimes (the recency
|
|
21
|
+
signal `stats` and `status --latest-for` rank by) survive untouched, and it restores directory mtimes
|
|
22
|
+
afterwards. An interrupted run records a journal outside the run dir and is finished by re-running the
|
|
23
|
+
command. Anything it cannot resolve unambiguously — a root artifact that is neither a duplicate nor
|
|
24
|
+
placeable, telemetry whose turn boundary cannot be dated or that spans more than two turns or whose
|
|
25
|
+
samples would land in a turn no transcript or result evidences, a dir with
|
|
26
|
+
no transcript at all — is **refused and named**. The one inference it makes is positional: an EMPTY file
|
|
27
|
+
has no content to attribute, so it follows its position to an EVIDENCED turn (its own, by name or by
|
|
28
|
+
rootArtifactTurn when a root transcript or result exists to move there) — never one it would mint. The
|
|
29
|
+
same never-mint rule holds on the content path: a stray `resources.turn-N.jsonl` cannot manufacture
|
|
30
|
+
`turns/N/` whether it is empty or carries samples, and a fully-archived dir's trailing telemetry cannot
|
|
31
|
+
manufacture the next turn out of arithmetic alone. Exit `1` when anything was refused, so a CI caller sees unfinished work.
|
|
32
|
+
After a `--write` that migrated or recovered anything, it prints a reminder to rebuild the index
|
|
33
|
+
(`cowork-harness stats --reindex`), since the index keys a row's timestamp off `result.json`'s mtime and
|
|
34
|
+
the files have moved.
|
|
35
|
+
|
|
36
|
+
- **`prune` skips scenarios with a migration in flight.** Between an interrupted migration and its
|
|
37
|
+
recovery a run dir's mtime reflects the migration, not the run — and `prune` ranks keep-slots by that
|
|
38
|
+
mtime, so it could evict a newer run in favour of a half-migrated older one. It now defers those
|
|
39
|
+
scenarios and says so.
|
|
40
|
+
|
|
41
|
+
- **`critique` surfaces the GRADED turn's `outcome` and `skillHash` in its own report**
|
|
42
|
+
(`gradedOutcome` / `gradedSkillHash` in JSON, and in the text header), and writes the graded result
|
|
43
|
+
under the stable name **`result.graded.json`**. `critique` runs two turns into one run directory, so
|
|
44
|
+
after the resume `result.json` is the *reflection* turn's and the graded turn is archived as
|
|
45
|
+
`result.turn-1.json` — the correct file to read was the *lower* number, the opposite of every other
|
|
46
|
+
multi-run convention. A harvester reading `result.json` silently ingested the reflection turn's numbers:
|
|
47
|
+
valid-looking, wrong, and unsignalled. Reported by a consumer building exactly that harvester; a
|
|
48
|
+
documentation-only fix would have helped only readers who already knew to look.
|
|
49
|
+
|
|
50
|
+
- **`exec_infra_error` verdict signal (`WARN`)** — a container `exec` that failed for infrastructure
|
|
51
|
+
reasons, as distinct from the fail-severity `infra_error` (a supervising process died). One failed
|
|
52
|
+
command no longer contaminates a whole run's evidence.
|
|
53
|
+
- **`RunResult.infraErrors[].source` is now an enum** — `hostloop-sidecar` / `hostloop-exec` /
|
|
54
|
+
`egress-sidecar`. The origin is what drives severity, and it is carried through the frozen cassette so
|
|
55
|
+
replay reaches the same verdict as the live run.
|
|
56
|
+
- **Capability use-scan health** — an unreadable or partially unparseable `events.jsonl` is now reported
|
|
57
|
+
as a degraded scan instead of being indistinguishable from a complete scan that found nothing.
|
|
58
|
+
|
|
59
|
+
- **Every `critique` limitation is now tagged with WHY it exists**, not just what it is — `structural`
|
|
60
|
+
(permanent, architect around it), `unverified` (unproven, **not** known-impossible), `deliberate` (a
|
|
61
|
+
design choice), `not-built` (simply absent). The distinction a reader needs is rarely "what can't it
|
|
62
|
+
do" but "should I design around this forever, or wait for it?" **Container-tier-only is `unverified`**:
|
|
63
|
+
the resume-continuity proof was run against the container tier's Linux ELF, and hostloop runs a
|
|
64
|
+
different (native) agent binary, so the proof does not transfer — nothing suggests hostloop would fail,
|
|
65
|
+
nobody has run it. **Lifting the pin needs BOTH** a live resume-continuity proof at hostloop against its
|
|
66
|
+
native binary AND the follow-on work that proof unblocks (unpinning three hard-coded container sites,
|
|
67
|
+
stamping the tier on the session manifest so a cross-tier resume fails loud, and plumbing host-write
|
|
68
|
+
consent) — evidence alone is not sufficient. A consumer read that pin as permanent and built a second
|
|
69
|
+
test lane around it. The tags appear in `critique --help` and [docs/critique.md](./docs/critique.md),
|
|
70
|
+
generated from one source.
|
|
71
|
+
- **`critique --help`'s KNOWN LIMITATIONS block is generated** from that source, and CI asserts the
|
|
72
|
+
shipped binary's output, the docs bullets, and their tags all agree.
|
|
73
|
+
|
|
74
|
+
### Changed
|
|
75
|
+
|
|
76
|
+
- ⚠️ **BREAKING: per-turn run-directory layout, single shape — the root-level `result.json` compatibility
|
|
77
|
+
copy is REMOVED.** A run directory that holds several turns (any `--resume`, and every `critique`) writes
|
|
78
|
+
each turn's `result.json`, `run.jsonl`, `trace.json` and `resources.jsonl` into **`turns/<N>/`**, once,
|
|
79
|
+
under its final name — nothing is renamed or overwritten as later turns arrive. `chat` now goes through
|
|
80
|
+
the same layout too (always `turns/1/` — a `chat` session mints a fresh dir per invocation and never
|
|
81
|
+
resumes). **`<outDir>/result.json` no longer exists — there is no root compat copy of any per-turn
|
|
82
|
+
artifact, on ANY run dir.** Read `turns/<N>/result.json` directly (`turns/1/` for a single-turn run), or
|
|
83
|
+
— for `critique` — the unchanged `result.graded.json` / `trace.graded.json` role aliases. Cumulative
|
|
84
|
+
streams (`events.jsonl`, `timeline.jsonl`) and session state are unchanged, so `critique`'s byte-offset
|
|
85
|
+
turn-isolation proof and cassette capture are unaffected.
|
|
86
|
+
|
|
87
|
+
**Two prior shapes are now REFUSED, loudly, by name, instead of being silently misread:**
|
|
88
|
+
- a **pre-layout** run dir (written before `turns/<N>/` existed: root `result.json`/`run.jsonl`, or a
|
|
89
|
+
name-mangled `result.turn-<N>.json` archive, no `turns/`);
|
|
90
|
+
- a **mixed** run dir (a pre-layout dir resumed under CURRENT code before this release — `turns/` present
|
|
91
|
+
*and* a stray root/archived file).
|
|
92
|
+
|
|
93
|
+
`verify-run`, `inspect`, `scaffold`, `diff`, `status --latest-for`, and a resumed `--session-id` all
|
|
94
|
+
refuse these with a message naming the shape found and pointing at `trace <dir>` — which still works
|
|
95
|
+
fully, since every one of its views derives from `events.jsonl`, which never moves. `stats --reindex`
|
|
96
|
+
counts them as skipped and names the remedy rather than dropping them from the index quietly.
|
|
97
|
+
|
|
98
|
+
**Migration: `cowork-harness migrate-run-dir`** converts a pre-layout dir in place (dry-run by default),
|
|
99
|
+
preserving the file timestamps `stats` and `status --latest-for` rank by. `diff` and
|
|
100
|
+
`status --latest-for` are called out because their pre-refusal behaviour was the dangerous kind: `diff`
|
|
101
|
+
reported two genuinely different runs as `identical` and exited 0, and `status --latest-for` could
|
|
102
|
+
select a *different* run than the newest and report its verdict — a CI script reading `.verdict.pass`
|
|
103
|
+
got a green light for a red run.
|
|
104
|
+
|
|
105
|
+
Previously the latest turn lived at the root while earlier ones were name-mangled archives, so a file's
|
|
106
|
+
name depended on whether a later turn ever happened; that shape produced a wrong-turn read, a destroyed
|
|
107
|
+
trace, and a dropped index row — this release's read-side (`turnArtifactPath` / `listTurns` in
|
|
108
|
+
`turn-layout.ts`, with the old `readTurnResult` deleted for having zero production callers) no longer has
|
|
109
|
+
a legacy-resolving branch at all, so that class of bug is now unrepresentable rather than merely fixed. The Python SDK's `_latest_run_jsonl` likewise now raises loudly
|
|
110
|
+
on a pre-layout dir instead of silently falling back to a root `run.jsonl` that (for any current-layout
|
|
111
|
+
dir) is a path to nowhere.
|
|
112
|
+
|
|
113
|
+
- **Platform baseline synced to Desktop 1.24012.1** (`baselines/desktop-1.24012.1.json`, now what
|
|
114
|
+
`baseline: latest` resolves to). The staged **agent binary is `2.1.217`** (native app + VM ELF, new
|
|
115
|
+
sha256 for each). The baseline moved in two steps this release — an earlier sync to **1.24012.0** (agent
|
|
116
|
+
`2.1.215`), then to 1.24012.1 — with **no prompt, spawn-env, or egress-allowlist drift across either**:
|
|
117
|
+
`spawn.env` is byte-identical to 1.22209.3, the same 15-domain allowlist and `gvisor` mode carry over
|
|
118
|
+
(the effort map is the one spawn field that changed — see the sonnet-5 delta below), and the
|
|
119
|
+
`deriveSpawnEnv` / `checkSpawnContractFacts` oracles stay green against the live asar. The
|
|
120
|
+
substantive deltas all came from the 1.24012.0 step and carry forward unchanged: `claude-sonnet-5` joins
|
|
121
|
+
the per-model effort map (`low|medium|high|xhigh|max`, recommended `medium`, modes `auto`); the
|
|
122
|
+
`coworkRuntimeConfig` gate drops its `pluginsFullSyncStalenessMs` key (never modeled here, inert); and
|
|
123
|
+
the dormant `autoModeOverridesAlwaysAllow` sentinel fired — see below. 1.24012.1 itself adds only the
|
|
124
|
+
agent bump: the `2.1.217` binary can emit the VCS SDK events `code_change_published` /
|
|
125
|
+
`vcs_state_changed` (SDK floor `2.1.216`), which the harness surfaces as a `system_event` (its existing
|
|
126
|
+
graceful degradation of an unmodeled system event, unchanged), and the binary's native skill-discovery
|
|
127
|
+
enable predicate widened to three branches — still inert here, since real Cowork's model-visible surface
|
|
128
|
+
is the Desktop SDK-MCP discovery servers, not the native tools. The example cassettes'
|
|
129
|
+
`fingerprint.baseline` tracks the new baseline.
|
|
130
|
+
- **The `autoModeOverridesAlwaysAllow` gate (`4200321681`) flipped absent → on** (`source: force`) and was
|
|
131
|
+
revisited as its pin intended. It stays **unmodeled, deliberately**: binary-verified in 1.24012.0, both
|
|
132
|
+
call sites only override an *already-existing* always-allow decision — the session rule cache
|
|
133
|
+
(`approvedToolNames`) and scheduled-task auto-approval — each further gated on `permissionMode` and
|
|
134
|
+
`isDestructiveConnectorTool`. The harness persists neither, so it already prompts wherever the gate makes
|
|
135
|
+
Cowork prompt; enabling it moves real Cowork *toward* harness behavior rather than away. Revisit only if
|
|
136
|
+
the harness gains a persistent per-tool approval cache.
|
|
137
|
+
- **The staged agent (`2.1.217`) enforces sub-agent fan-out caps, so `dispatch_count_max` is now framed as
|
|
138
|
+
a budget UNDER Cowork's cap, not a reproduction of it.** Because the harness spawns the real binary, a
|
|
139
|
+
run that fans out past the agent's caps now errors from the binary itself: a **concurrent** cap
|
|
140
|
+
(`CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS`, default 20; error `subagent_concurrency_cap`, new in 2.1.217)
|
|
141
|
+
and a **per-session** cap (`CLAUDE_CODE_MAX_SUBAGENTS_PER_SESSION`, default 200; error
|
|
142
|
+
`subagent_count_cap`, present since ≤2.1.215). The scenario schema's `dispatch_count_max` description,
|
|
143
|
+
the `assert` over-budget message, and SPEC §10 no longer claim "Cowork imposes no in-conversation
|
|
144
|
+
Task-dispatch cap" — that claim was stale. The harness does not reproduce the caps; it inherits them by
|
|
145
|
+
running the binary.
|
|
146
|
+
- **A host-loop `exec` infrastructure failure now WARNS instead of failing the run.** ⚠️ **Upgrade note:**
|
|
147
|
+
a run that previously exited `1` because one `docker exec` failed will now exit `0`. A dead sidecar
|
|
148
|
+
still hard-fails. Known residual, documented in `docs/scenario.md`: if *every* exec failed the agent ran
|
|
149
|
+
nothing and the run still only warns — inspect `result.infraErrors` when a run looks suspiciously empty.
|
|
150
|
+
- **A model-requested bash `timeout_ms` expiry is no longer classified as an infrastructure failure.**
|
|
151
|
+
The model now receives its command's own partial output with `Command timed out after <duration>`
|
|
152
|
+
merged into stderr — matching real Cowork, verified against the staged agent binary — instead of an
|
|
153
|
+
opaque `[infrastructure error: see run log for details]`.
|
|
154
|
+
- **The agent's spawn env now always carries a normalized IANA `TZ`** — matching Desktop, which injects
|
|
155
|
+
`Intl.DateTimeFormat().resolvedOptions().timeZone` unconditionally. Previously `TZ` was forwarded only
|
|
156
|
+
when the host shell exported it, and forwarded raw, so a host with no `TZ` set — or a legacy/non-IANA
|
|
157
|
+
export (`US/Eastern`, `EST5EDT`) — diverged from real Cowork's date/"today" rendering inside the agent.
|
|
158
|
+
- **The `tool_available` assertion now names its evidence limit.** It evaluates against the run's
|
|
159
|
+
*eagerly-loaded* tool set (the SDK init manifest in `result.json`); a factory-deferred tool — e.g. the
|
|
160
|
+
skill-discovery MCP tools, loaded on demand via a `ToolSearch` round-trip — can be genuinely available in
|
|
161
|
+
the run yet miss here, a false negative. The assertion still fails on a miss; the failure message now
|
|
162
|
+
states that eagerly-loaded scope rather than implying provable unavailability.
|
|
163
|
+
- **An explicitly requested `--dotenv` file now fails loud.** ⚠️ **Upgrade note:** an unreadable file, or
|
|
164
|
+
a path that is a directory, previously fell through to lower-precedence `.env` sources *while still
|
|
165
|
+
printing a success line* — so a typo'd path silently ran against the wrong credentials. It is now a
|
|
166
|
+
usage error. Automatic `.env` discovery is unchanged (still best-effort).
|
|
167
|
+
- **`diff` no longer reports `identical` when only one side has an artifact manifest.** ⚠️ **Upgrade
|
|
168
|
+
note:** such a comparison previously exited `0`; it now exits `1`, because unavailable evidence is not
|
|
169
|
+
evidence of equality. Both-sides-missing still does not veto identity. The `--output-format json`
|
|
170
|
+
envelope gained an `artifactsAvailability` key.
|
|
171
|
+
- **`stats --reindex` merges rows by per-completion identity** (`outDir` + a new `turn` field) rather than
|
|
172
|
+
by `outDir` alone, and reports rejected symlinked run directories.
|
|
173
|
+
|
|
174
|
+
### Fixed
|
|
175
|
+
|
|
176
|
+
- **`verify-run` now REFUSES a multi-turn run directory** instead of certifying the wrong turn. Root
|
|
177
|
+
`result.json` is the latest turn; on a `critique` directory that is the *reflection* turn while the
|
|
178
|
+
scenario describes the *graded* one. Previously the cumulative gate scan false-FAILED on the other
|
|
179
|
+
turn's gates — wrong, but loud. The refusal names `result.graded.json` / `turns/1/result.json` so the
|
|
180
|
+
caller can still reach the graded turn.
|
|
181
|
+
- **`trace` no longer mixes turn scopes.** After timeline reads became turn-scoped, `--view
|
|
182
|
+
tool-durations` showed the latest turn while the tools/questions/dispatches views still showed every
|
|
183
|
+
turn — two views of one run directory describing different scopes. All views are now the latest turn,
|
|
184
|
+
and a `::notice::` reports when earlier turns exist rather than hiding them. Its cache-read footer and
|
|
185
|
+
gate-provenance (`answeredBy`) views also now say when a result is not turn-addressable (a pre-layout
|
|
186
|
+
dir) instead of silently omitting the cache-read ratio / labels.
|
|
187
|
+
- **`prune` no longer demotes an unmigrated pre-layout run dir to the junk tier.** Its real-run predicate
|
|
188
|
+
keyed on `hasTurnDirs || events.jsonl`, reasoning only about what current writers produce — but `prune`
|
|
189
|
+
ranks *history*, including the legacy dirs `migrate-run-dir` exists to preserve, so a legacy dir with no
|
|
190
|
+
`events.jsonl` could be evicted ahead of an empty scaffold. It now also counts a `legacy` / `mixed` shape
|
|
191
|
+
as a real run. (Distinct from the in-flight-migration deferral above — this is about which dirs count as
|
|
192
|
+
real at all.)
|
|
193
|
+
|
|
194
|
+
- **A resumed turn was judged on the PRIOR turn's evidence — three wrong-verdict paths.** `events.jsonl`
|
|
195
|
+
is append-only across turns with no per-turn marker, and three whole-file scanners decide a run's
|
|
196
|
+
outcome: `scanEvents` (outputs-delete / host-path-leak → fail signals, and an authored
|
|
197
|
+
`no_delete_in_outputs`), `findUngatedPathToolCalls` (→ a run-level `error` at hostloop), and
|
|
198
|
+
`detectCapabilityUse` (→ `missing_capability`, a fail signal, which fires on the default lean image).
|
|
199
|
+
So on any `--resume` — and every `critique` reflection turn — turn 1's delete, gated tool call, or
|
|
200
|
+
capability use FAILED turn 2. A turn-start marker now scopes all three to the current turn.
|
|
201
|
+
`resources.jsonl` had the same shape (turn 1's peak RSS judged against turn 2's `max_peak_rss_bytes`)
|
|
202
|
+
and is archived per turn. Single-turn runs write no marker, so their `events.jsonl` is byte-identical
|
|
203
|
+
and no cassette is affected. Missing marker ⇒ whole-file scan, i.e. fail-closed.
|
|
204
|
+
|
|
205
|
+
- **A resumed turn's telemetry included the PRIOR turn's events, and could produce a false PASS.**
|
|
206
|
+
`timeline.jsonl` is append-mode with a fresh header per turn, but `readTimeline` returned every line
|
|
207
|
+
after the first as an event — so on any `--resume` (and every `critique` reflection turn) the current
|
|
208
|
+
turn's `toolDurations`/`skillActivity`/`subagents` folded in the previous turn's tool calls. Because
|
|
209
|
+
the **`skill_tool_used` assertion** evaluates against that same `skillActivity`, a turn-1 skill window
|
|
210
|
+
could satisfy a turn-2 assertion. The reader now returns only the current turn's segment. The file
|
|
211
|
+
stays one append-only stream, so `critique`'s byte-offset turn-isolation proof is unaffected.
|
|
212
|
+
|
|
213
|
+
- **A resumed turn destroyed the prior turn's `trace.json`.** Because it is rebuilt and overwritten on
|
|
214
|
+
every completion, the earlier turn's trace was deleted rather than preserved, so a `critique` lost the
|
|
215
|
+
graded turn's trace entirely. Each turn now owns its own `turns/<N>/trace.json`, written once and never
|
|
216
|
+
overwritten, and `critique` additionally writes **`trace.graded.json`** beside `result.graded.json`.
|
|
217
|
+
- **`stats --reindex` dropped every non-latest turn when rebuilding from the runs tree.** It read only the
|
|
218
|
+
root `result.json` per run directory, so a resumed session's earlier turns vanished — and on a
|
|
219
|
+
`critique` directory the root file is the *reflection* turn, so it was the **graded** rows that were
|
|
220
|
+
lost. Every turn under `turns/<N>/` is now indexed as its own completion; `result.graded.json` — a
|
|
221
|
+
root-level copy of the graded turn — is deliberately not matched, so it cannot double-count. A dir that
|
|
222
|
+
has not been migrated is counted as `skippedLegacy` and reported with the remedy, never dropped
|
|
223
|
+
silently.
|
|
224
|
+
|
|
225
|
+
- **An ambient `GIT_DIR` silently computed the wrong skill file set.** Git hooks export `GIT_DIR` (and
|
|
226
|
+
`GIT_INDEX_FILE`) into every child process, and with `GIT_DIR` set but no `GIT_WORK_TREE` git stops
|
|
227
|
+
inferring the work tree from `cwd` and treats `cwd` as the repo root. `gitTrackedSet`'s
|
|
228
|
+
`rev-parse --show-toplevel` probe therefore still succeeded — so the not-a-repo raw-walk fallback never
|
|
229
|
+
fired — while `git ls-files -- .` returned the **entire repo index as root-relative paths** instead of
|
|
230
|
+
the directory-relative ones. Measured on this repo: 2 tracked files became 625 wrong ones. That set
|
|
231
|
+
feeds both `skillHash` and the mount-copy filter, so any run invoked from a git hook (or from CI that
|
|
232
|
+
exports `GIT_DIR`) got a wrong hash and a mount filter pointed at paths that do not exist under the
|
|
233
|
+
skill dir. The visible symptom was the repo's own pre-commit hook reporting committed example cassettes
|
|
234
|
+
as `[stale] skill files changed since record` on every parity sync. `skillCommit` had the same defect:
|
|
235
|
+
`git -C <dir>` is overridden by an ambient `GIT_DIR`, so every skill dir resolved to that foreign repo's
|
|
236
|
+
HEAD — recording a foreign commit as the skill's provenance and masking dirs that are genuinely in
|
|
237
|
+
different repos. Both call sites now spawn git with `GIT_DIR` / `GIT_WORK_TREE` / `GIT_INDEX_FILE`
|
|
238
|
+
stripped, via one shared helper so they cannot drift. `run-index`'s `gitInfo` and `doctor`'s worktree
|
|
239
|
+
probe deliberately keep inheriting — they are asking about the *ambient* repo.
|
|
240
|
+
- **`stats --reindex` destroyed multi-turn history.** Every `--resume` turn — and `critique`'s task +
|
|
241
|
+
reflection pair — writes to one `outDir`, so keying by directory collapsed N completions into one,
|
|
242
|
+
silently changing run counts, pass rates and costs.
|
|
243
|
+
- **Host-loop sidecar failures never reached the verdict.** They were appended straight to `events.jsonl`,
|
|
244
|
+
which no live drive re-reads, so a dying sidecar left the run green; a signal-only termination (OOM,
|
|
245
|
+
`SIGKILL`) was recorded nowhere at all.
|
|
246
|
+
- **`result.json` was written non-atomically** at all three producers, so an interrupted write could leave
|
|
247
|
+
the canonical record truncated.
|
|
248
|
+
- **Corrupt index rows were blind-cast**, letting one malformed row crash `stats` or fabricate a
|
|
249
|
+
pass/cost value; `reindex` also followed symlinks out of the runs root.
|
|
250
|
+
- **`scaffold` turned unavailable artifact evidence into "no artifacts"**, permanently encoding a false
|
|
251
|
+
"this run produced nothing" claim into a generated scenario.
|
|
252
|
+
- **`critique` treated a vanished turn-1 evidence file as genuinely empty evidence** rather than an
|
|
253
|
+
integrity failure. A stream that was legitimately zero bytes at capture is still reported clean.
|
|
254
|
+
- **`critique`'s exit-code table omitted `1`.** Exit `1` is reachable on operator interrupt
|
|
255
|
+
(SIGINT/SIGTERM); a sweep wrapper treating it as impossible misreads a cancelled run as a crash.
|
|
256
|
+
- Documented that after critique's resume, `result.json` is the **reflection** turn's result — the graded
|
|
257
|
+
turn is archived as `result.turn-1.json`. Reading the wrong one yields a valid-looking wrong number.
|
|
258
|
+
|
|
9
259
|
## [1.6.0] — 2026-07-20
|
|
10
260
|
|
|
11
261
|
### Added
|
package/CONTRIBUTING.md
CHANGED
|
@@ -44,6 +44,9 @@ src/
|
|
|
44
44
|
decide/ Decider — answer policy (scripted / LLM / external-channel deciders)
|
|
45
45
|
run/ Run — turn loop + RunRecord (run.ts); executeScenario (execute.ts); cassette replay
|
|
46
46
|
runtime/ protocol (L0) / container (L1) / microvm (L2) / hostloop / lima
|
|
47
|
+
hostloop/ Cowork host-loop handlers — can_use_tool gate, web_fetch dedup, workspace/path hooks
|
|
48
|
+
critique/ critique — run a skill, then grade its self-report against a frozen run record
|
|
49
|
+
staging/ agent-binary resolution + mount naming for the sandboxed live tiers
|
|
47
50
|
egress/ default-deny allowlist proxy
|
|
48
51
|
boundary.ts sandbox self-test probes
|
|
49
52
|
assert.ts synchronous assertion evaluator
|