cowork-harness 1.6.0 → 1.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +49 -18
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +10 -9
  3. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +2 -2
  4. package/.claude/skills/cowork-harness/references/scenario-schema.md +12 -10
  5. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  6. package/CHANGELOG.md +250 -0
  7. package/CONTRIBUTING.md +3 -0
  8. package/README.md +39 -24
  9. package/SPEC.md +14 -4
  10. package/baselines/desktop-1.24012.0.json +405 -0
  11. package/baselines/desktop-1.24012.1.json +485 -0
  12. package/dist/agent/session.js +9 -3
  13. package/dist/agent/timeline.js +76 -3
  14. package/dist/assert.js +4 -2
  15. package/dist/baseline.js +20 -0
  16. package/dist/cli.js +194 -33
  17. package/dist/critique/command.js +61 -8
  18. package/dist/critique/evidence.js +65 -41
  19. package/dist/critique/limitations.js +118 -0
  20. package/dist/critique/package-evidence.js +29 -18
  21. package/dist/dotenv.js +38 -3
  22. package/dist/errors.js +12 -0
  23. package/dist/hostloop/workspace-handler.js +35 -1
  24. package/dist/io.js +15 -7
  25. package/dist/run/cassette.js +4 -0
  26. package/dist/run/chat-result.js +6 -2
  27. package/dist/run/chat.js +39 -6
  28. package/dist/run/diff.js +26 -5
  29. package/dist/run/doctor.js +2 -1
  30. package/dist/run/execute.js +158 -51
  31. package/dist/run/inspect-view.js +11 -4
  32. package/dist/run/latest-run.js +33 -7
  33. package/dist/run/migrate-run-dir.js +817 -0
  34. package/dist/run/renderer.js +11 -3
  35. package/dist/run/run-index.js +225 -45
  36. package/dist/run/run.js +13 -4
  37. package/dist/run/runs-gc.js +45 -8
  38. package/dist/run/scaffold.js +28 -4
  39. package/dist/run/skill-files.js +18 -1
  40. package/dist/run/trace-view.js +72 -12
  41. package/dist/run/turn-events.js +34 -0
  42. package/dist/run/turn-layout.js +189 -0
  43. package/dist/run/verdict.js +17 -2
  44. package/dist/runtime/host-env.js +8 -2
  45. package/dist/runtime/hostloop.js +58 -17
  46. package/dist/runtime/image-capabilities.js +22 -8
  47. package/dist/runtime/resource-sampler.js +20 -6
  48. package/dist/sync/cowork-sync.js +137 -17
  49. package/dist/types.js +17 -2
  50. package/docs/cassette.md +5 -5
  51. package/docs/chat.md +4 -3
  52. package/docs/critique.md +54 -7
  53. package/docs/debugging.md +38 -3
  54. package/docs/fidelity-gaps.md +51 -11
  55. package/docs/gotchas.md +7 -1
  56. package/docs/maintenance.md +2 -2
  57. package/docs/run-status.md +4 -3
  58. package/docs/scenario.md +21 -5
  59. package/docs/session.md +1 -1
  60. package/docs/stats.md +29 -7
  61. package/examples/replays/README.md +1 -1
  62. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  63. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  64. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  65. package/llms.txt +1 -1
  66. package/package.json +1 -1
  67. package/python/cowork_harness.py +35 -11
  68. package/schema/run-result.json +10 -4
  69. package/schema/scenario.schema.json +1 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.6.0
7
- tracks-harness: cowork-harness 1.6.0 (baseline desktop-1.22209.3)
6
+ version: 1.7.0
7
+ tracks-harness: cowork-harness 1.7.0 (baseline desktop-1.24012.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.6.0` (baseline
26
- > `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.7.0` (baseline
26
+ > `desktop-1.24012.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,9 +39,30 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.6.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.6.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.6.0"`. **Pin `@>=1.6.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
-
44
- What the ≥ 1.6.0 floor gates, by release:
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.7.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.7.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.7.0"`. **Pin `@>=1.7.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
+
44
+ What the ≥ 1.7.0 floor gates, by release:
45
+
46
+ - **1.7.0 (per-turn run-directory layout, single shape):** every run dir — `run`/`skill`/`chat`,
47
+ single-turn or multi-turn (any `--session-id` + `--resume`, and **every `critique`**: task turn +
48
+ reflection turn) — writes each turn's `result.json` / `run.jsonl` / `trace.json` / `resources.jsonl`
49
+ into **`turns/<N>/`**, written once and never renamed. **There is no root compat copy of anything —
50
+ `<run-dir>/result.json` does not exist.** **When a user asks about a run, point them at
51
+ `turns/1/result.json`** (the only completion on a single-turn dir) — or on a `critique` dir, better yet
52
+ at `result.graded.json` / `trace.graded.json`, the role-stable aliases `critique` writes (the graded
53
+ task turn, never the reflection one). `events.jsonl` / `timeline.jsonl` stay cumulative at the run-dir
54
+ root, always. A dir written before this layout existed is refused LOUD, by name, naming the shape found
55
+ — never silently misread as if it were `turns/1/`. **Convert it in place with
56
+ `cowork-harness migrate-run-dir` (dry-run by default); do NOT tell the user to re-run or re-record,
57
+ which throws away history the migrator recovers.** Its `events.jsonl` still fully supports `trace`,
58
+ which is why every refusal points there. `verify-run` REFUSES a multi-turn dir rather than certifying
59
+ the wrong turn, and `trace` shows the latest turn with a `::notice::` when earlier turns exist.
60
+
61
+ - **1.7.0 (limitation provenance):** every `critique` limitation in `critique --help` and
62
+ docs/critique.md is tagged with WHY it exists — `[structural]` (permanent), `[unverified]` (unproven,
63
+ **not** known-impossible), `[deliberate]`, `[not-built]`. **Read the tag before telling a user to
64
+ design around a limitation.** In particular `critique`'s container-tier pin is `[unverified]`, not
65
+ permanent: do not advise building a second hostloop test lane on the assumption it can never lift.
45
66
 
46
67
  - **1.6.0 (`critique`'s skill-flag parity):** `critique` accepts most `skill` flags under the same names,
47
68
  which is what makes a document-analysis skill (cap table, deck, financial model, transcript)
@@ -81,7 +102,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
81
102
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
82
103
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
83
104
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
84
- - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — UNRELEASED, see the floor list; `--run-dir` stays global-only everywhere.
105
+ - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
85
106
 
86
107
  ## Orient — the three loops
87
108
 
@@ -113,7 +134,7 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
113
134
  > `node_modules/cowork-harness/docs/<name>.md` before assuming a pointer dangles. A **plugin**
114
135
  > install loads a trimmed source-only cache where those pointers genuinely do dangle.
115
136
 
116
- Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
137
+ Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · migrate-run-dir · lint ·
117
138
  lint-skill · analyze-skill · probe-dispatch ·
118
139
  verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
119
140
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
@@ -155,7 +176,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
155
176
  > is your **working tree**, so an uncommitted edit to an already-tracked file *is* tested — you needn't
156
177
  > commit to iterate. Only brand-new (untracked) files must be `git add`-ed to appear. Commit before you
157
178
  > record the **locking cassette**, though: real Cowork ships the *committed* tree, so a green on
158
- > uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder used to mount *empty* and the agent reported "the skill isn't
179
+ > uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder mounts *empty* and the agent reports "the skill isn't
159
180
  > installed" then did the work itself — a green-looking run where the skill never loaded. That now
160
181
  > **hard-fails** (`BoundaryError`, exit 3) naming the dir, and a partially-tracked folder emits a loud
161
182
  > `::notice:: [stage]` listing the excluded files. Fix: `git add` the skill, or `COWORK_HARNESS_GITSET=0`
@@ -372,7 +393,7 @@ cassette — has its own recipe:
372
393
  "<q>=Israeli company"` binds whichever option starts with `Israeli company`. It is uniqueness-guarded and
373
394
  **fails loud** if the anchor ever matches two options (the documented trade: drift-tolerance, not strict
374
395
  CI reproducibility — for that, pin a full exact label or a free-text `answer:`).
375
- 3. **Budget ~1 re-run per file.** If a gate whiffs, the run no longer vanishes — it exits non-zero but
396
+ 3. **Budget ~1 re-run per file.** If a gate whiffs, the run does not vanish — it exits non-zero but
376
397
  **salvages a PARTIAL run** (the extraction the agent already did is written to disk). So the cost of a
377
398
  missed gate is one re-run with a better `--intent` or a scripted answer, not a lost paid run.
378
399
  4. **Inspect the outputs to judge correctness.** `cowork-harness inspect <run-dir>` shows what the run
@@ -435,6 +456,14 @@ Recognize these before "fixing" a non-bug:
435
456
  `transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
436
457
  run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
437
458
  it's valid.
459
+ - **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
460
+ infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
461
+ rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
462
+ the fail-severity `infra_error`, where a **supervising process** died and contaminated everything. Note
463
+ a model-requested `timeout_ms` expiry is *not* this: it returns the command's partial output with
464
+ `Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
465
+ failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
466
+ suspiciously empty.
438
467
  - **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
439
468
  `RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
440
469
  pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
@@ -638,12 +667,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
638
667
  `mcp__workspace__bash` alias, so a "sub-agent used no shell" check must glob **both** `Bash` and
639
668
  `mcp__workspace__*` to hold across every tier.
640
669
 
641
- 6. **`dispatch_count_max` is an author-chosen budget, not a production cap.** It's a post-hoc count
642
- assertion: passing means "happened to dispatch ≤N this run," nothing more. Cowork imposes **no**
643
- in-conversation `Task`-dispatch cap — gate `1648655587`'s `{perTask:1, global:3}` governs the
644
- separate scheduled/cron-task session scheduler, not the `Task` tool (binary-verified; details in
645
- `SPEC.md` §10 — repo-only). So there is no "skip-on-cap" for the harness to reproduce; use this key
646
- only to catch a fan-out you don't want.
670
+ 6. **`dispatch_count_max` is your author-chosen budget UNDER Cowork's production cap, not a
671
+ reproduction of it.** It's a post-hoc count assertion: passing means "happened to dispatch ≤N this
672
+ run." Cowork DOES cap `Task` fan-out **agent-side** (`taskRegistry`: concurrent **20** /
673
+ per-session **200**, landed 2.1.212/2.1.217) — SEPARATE from the scheduled-task session limiter
674
+ (gate `1648655587`'s `{perTask:1, global:3}`, a different mechanism; binary-verified, `SPEC.md` §10
675
+ — repo-only). The harness **inherits** the production cap by spawning the real agent binary, so a
676
+ `dispatch_count_max` pass means "your tighter budget held," not "near a real limit"; use it to catch
677
+ a fan-out you don't want.
647
678
 
648
679
  7. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
649
680
  assertions need a sandboxed tier (`container`+). Good: this one fails loud by design.
@@ -727,7 +758,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
727
758
 
728
759
  17. **Editing `scenarios/*.yaml` `assert:` does NOT change a plain `replay`.** *Why:* `replay` evaluates the
729
760
  assertions **frozen in the cassette** by default — it is byte-deterministic and ignores the working tree (so
730
- a committed cassette can't silently re-interpret against an uncommitted YAML). This used to be a *silent*
761
+ a committed cassette can't silently re-interpret against an uncommitted YAML). This is *loud* rather than a *silent*
731
762
  no-op; now plain `replay` prints a `::notice::` when a sibling's `assert:` differs and points you at the fix.
732
763
  *Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
733
764
  `--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.6.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.7.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.6.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.7.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -32,7 +32,7 @@ jobs:
32
32
  - uses: actions/checkout@v4
33
33
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
34
34
  run: |
35
- V=2.1.215 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.217 # match your scenario's pinned baseline's agentVersion
36
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
37
37
  chmod +x "$RUNNER_TEMP/claude-$V"
38
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.6.0"
60
+ - run: npm i -g "cowork-harness@>=1.7.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.6.0"
200
+ - run: npm i -g "cowork-harness@>=1.7.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.6.0"
229
+ run: npm i -g "cowork-harness@>=1.7.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -263,8 +263,9 @@ JSON, not the human-readable text (which is explicitly NOT stable).
263
263
 
264
264
  A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
265
265
  set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
266
- artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl`, `run.jsonl`,
267
- `trace.json`, `egress.log`, `result.json`. Digest one with `cowork-harness trace <run-id | dir>`.
266
+ artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl` and
267
+ `egress.log` at the root, plus each turn's `run.jsonl` / `trace.json` / `result.json` under `turns/<N>/`
268
+ (a single-turn run has just `turns/1/`; there is no root compat copy of any of these). Digest one with `cowork-harness trace <run-id | dir>`.
268
269
  Secrets are scrubbed from every persisted log by value.
269
270
 
270
271
  ## Don't assume a fixed assertion count across lanes
@@ -293,7 +294,7 @@ does **not** imply the recording is still valid. Each replay result carries `sta
293
294
  (A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
294
295
  check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
295
296
  `·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
296
- gate (`ok:false`) — but it's no longer class-blind on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
297
+ gate (`ok:false`) — but it is class-AWARE on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
297
298
  `format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
298
299
  while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
299
300
  fail it either way.)
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.6.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.7.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -122,7 +122,7 @@ hands `fn` exactly this dict.
122
122
 
123
123
  - On `skill`: `fail | prompt | first`.
124
124
  - On `run`: `fail | first` (`prompt` rejected — it would break determinism).
125
- - On `critique` (UNRELEASED): `fail | first` (`prompt` rejected — there is no TTY inside the spawned turn).
125
+ - On `critique`: `fail | first` (`prompt` rejected — there is no TTY inside the spawned turn).
126
126
  - **`llm` is NOT an `--on-unanswered` value.** The bare flag `--on-unanswered llm` is rejected; use
127
127
  `--decider-llm` (CLI) or `on_unanswered: llm` (YAML).
128
128
  - `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.6.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.7.0`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
@@ -269,7 +269,7 @@ same set live from the schema.
269
269
  | `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
270
270
  | `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
271
271
  | `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
272
- | `dispatch_count_max: <N>` | at most N sub-agents dispatched — an author-chosen budget (Cowork imposes no in-conversation Task-dispatch cap; records only, enforces nothing — see gotcha 12) |
272
+ | `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 |
273
273
  | `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked via the `Skill` tool — evidence-unavailable (not a normal fail) if the agent's init tools have no `Skill` tool |
274
274
  | `no_skill_triggered: <regex>` | no invoked skill id matched — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent or the `Skill` tool is unobservable |
275
275
  | `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
@@ -348,16 +348,17 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
348
348
  | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
349
349
  | `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
350
350
  | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
351
- | `infra_error` | fail | A VM/egress sidecar crashed mid-run — not author-suppressible |
351
+ | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
352
352
  | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
353
353
  | `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
354
354
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
355
355
  | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
356
356
  | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
357
+ | `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
357
358
 
358
359
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
359
360
  overall run verdict and exit code — `assert result: success` alone won't catch it; check
360
- `result.verdict.signals[].severity` or the run's exit code. Only the four **warn** codes are truly benign.
361
+ `result.verdict.signals[].severity` or the run's exit code. Only the five **warn** codes are truly benign.
361
362
 
362
363
  ## Replay class
363
364
 
@@ -417,7 +418,7 @@ skill-source drift; every replay result also reports it class-tagged in `stalene
417
418
  **Mixed assertions on replay:** before evaluating, `replay` strips each assertion to its replay-checkable
418
419
  keys and drops any left empty. So `{result, egress_denied}` evaluates on replay as `{result}` alone — its
419
420
  `egress_denied` half is removed (not AND-ed against an unreadable value); with a manifest, `file_exists`/
420
- `artifact_json` are no longer stripped. The harness is **loud in two classes**: a *full skip* (`::warning::`
421
+ `artifact_json` are not stripped. The harness is **loud in two classes**: a *full skip* (`::warning::`
421
422
  with the count of pure live-only assertions not evaluated) and a *partial skip* (`::warning::` when a mixed
422
423
  assertion's live-only half was dropped).
423
424
  Two CI consequences: skipped assertions are **absent** from `results[].assertions[]` (not
@@ -529,11 +530,12 @@ reference omits). Neither list is a strict superset of the other — reach for t
529
530
  11. **`subagent_declared_but_unused` fires on declared-but-didn't-use-THAT-tool**, even if the
530
531
  sub-agent used other tools.
531
532
 
532
- 12. **`dispatch_count_max` is an author-chosen budget, not a production cap.** It records the count
533
- and asserts on it; passing means "dispatched ≤N this run," not "the harness capped it." Cowork
534
- imposes no in-conversation Task-dispatch cap to reproduce — gate `1648655587`'s
535
- `{perTask:1, global:3}` is the scheduled/cron-task session limiter, a different mechanism
536
- (binary-verified; SPEC §10).
533
+ 12. **`dispatch_count_max` is your author-chosen budget under Cowork's production cap, not a
534
+ reproduction of it.** It records the count and asserts on it; passing means "dispatched ≤N this
535
+ run," not "the harness capped it." Cowork DOES cap `Task` fan-out agent-side (`taskRegistry`:
536
+ concurrent 20 / per-session 200, landed 2.1.212/2.1.217), which the harness inherits by spawning the
537
+ real agent binary — SEPARATE from gate `1648655587`'s `{perTask:1, global:3}` scheduled/cron-task
538
+ session limiter, a different mechanism (binary-verified; SPEC §10).
537
539
 
538
540
  13. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
539
541
  assertions need `container`+. Fails loud by design.
@@ -182,7 +182,7 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
182
182
  (blinded evaluator + mechanical citation checking). **Critiquing a document-analysis skill?** The probe
183
183
  attaches nothing on its own — pass `--upload <path>` (repeatable) or `--folder <dir>` exactly as you
184
184
  would to `skill`, or the graded run has no file and "there was no file attached" is the correct finding,
185
- not a skill defect. Source flags reach both spawned turns automatically. UNRELEASED — see SKILL.md's
185
+ not a skill defect. Source flags reach both spawned turns automatically — see SKILL.md's
186
186
  floor list. See docs/critique.md for the full flag table, cost and limits. If you
187
187
  prefer to build your own grader, the substrate is still here:
188
188
  - `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
package/CHANGELOG.md CHANGED
@@ -6,6 +6,256 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.7.0] — 2026-07-22
10
+
11
+ ### Added
12
+
13
+ - **`migrate-run-dir` — convert pre-layout run dirs to the per-turn `turns/<N>/` layout, in place.**
14
+ A run dir written before the per-turn layout keeps `result.json` / `run.jsonl` / `trace.json` /
15
+ `resources.jsonl` at its root. Once the legacy read layer is removed, those dirs become unreadable to
16
+ `verify-run` / `diff` / `inspect` / `stats`; this command converts them so the history survives the
17
+ change instead of having to be re-run.
18
+
19
+ **Dry-run by default** — `--write` applies, and `--scenario <name>` scopes the run to a single
20
+ scenario so a rollout can be staged: migrate one, verify it, then do the rest. It renames rather than copies, so file mtimes (the recency
21
+ signal `stats` and `status --latest-for` rank by) survive untouched, and it restores directory mtimes
22
+ afterwards. An interrupted run records a journal outside the run dir and is finished by re-running the
23
+ command. Anything it cannot resolve unambiguously — a root artifact that is neither a duplicate nor
24
+ placeable, telemetry whose turn boundary cannot be dated or that spans more than two turns or whose
25
+ samples would land in a turn no transcript or result evidences, a dir with
26
+ no transcript at all — is **refused and named**. The one inference it makes is positional: an EMPTY file
27
+ has no content to attribute, so it follows its position to an EVIDENCED turn (its own, by name or by
28
+ rootArtifactTurn when a root transcript or result exists to move there) — never one it would mint. The
29
+ same never-mint rule holds on the content path: a stray `resources.turn-N.jsonl` cannot manufacture
30
+ `turns/N/` whether it is empty or carries samples, and a fully-archived dir's trailing telemetry cannot
31
+ manufacture the next turn out of arithmetic alone. Exit `1` when anything was refused, so a CI caller sees unfinished work.
32
+ After a `--write` that migrated or recovered anything, it prints a reminder to rebuild the index
33
+ (`cowork-harness stats --reindex`), since the index keys a row's timestamp off `result.json`'s mtime and
34
+ the files have moved.
35
+
36
+ - **`prune` skips scenarios with a migration in flight.** Between an interrupted migration and its
37
+ recovery a run dir's mtime reflects the migration, not the run — and `prune` ranks keep-slots by that
38
+ mtime, so it could evict a newer run in favour of a half-migrated older one. It now defers those
39
+ scenarios and says so.
40
+
41
+ - **`critique` surfaces the GRADED turn's `outcome` and `skillHash` in its own report**
42
+ (`gradedOutcome` / `gradedSkillHash` in JSON, and in the text header), and writes the graded result
43
+ under the stable name **`result.graded.json`**. `critique` runs two turns into one run directory, so
44
+ after the resume `result.json` is the *reflection* turn's and the graded turn is archived as
45
+ `result.turn-1.json` — the correct file to read was the *lower* number, the opposite of every other
46
+ multi-run convention. A harvester reading `result.json` silently ingested the reflection turn's numbers:
47
+ valid-looking, wrong, and unsignalled. Reported by a consumer building exactly that harvester; a
48
+ documentation-only fix would have helped only readers who already knew to look.
49
+
50
+ - **`exec_infra_error` verdict signal (`WARN`)** — a container `exec` that failed for infrastructure
51
+ reasons, as distinct from the fail-severity `infra_error` (a supervising process died). One failed
52
+ command no longer contaminates a whole run's evidence.
53
+ - **`RunResult.infraErrors[].source` is now an enum** — `hostloop-sidecar` / `hostloop-exec` /
54
+ `egress-sidecar`. The origin is what drives severity, and it is carried through the frozen cassette so
55
+ replay reaches the same verdict as the live run.
56
+ - **Capability use-scan health** — an unreadable or partially unparseable `events.jsonl` is now reported
57
+ as a degraded scan instead of being indistinguishable from a complete scan that found nothing.
58
+
59
+ - **Every `critique` limitation is now tagged with WHY it exists**, not just what it is — `structural`
60
+ (permanent, architect around it), `unverified` (unproven, **not** known-impossible), `deliberate` (a
61
+ design choice), `not-built` (simply absent). The distinction a reader needs is rarely "what can't it
62
+ do" but "should I design around this forever, or wait for it?" **Container-tier-only is `unverified`**:
63
+ the resume-continuity proof was run against the container tier's Linux ELF, and hostloop runs a
64
+ different (native) agent binary, so the proof does not transfer — nothing suggests hostloop would fail,
65
+ nobody has run it. **Lifting the pin needs BOTH** a live resume-continuity proof at hostloop against its
66
+ native binary AND the follow-on work that proof unblocks (unpinning three hard-coded container sites,
67
+ stamping the tier on the session manifest so a cross-tier resume fails loud, and plumbing host-write
68
+ consent) — evidence alone is not sufficient. A consumer read that pin as permanent and built a second
69
+ test lane around it. The tags appear in `critique --help` and [docs/critique.md](./docs/critique.md),
70
+ generated from one source.
71
+ - **`critique --help`'s KNOWN LIMITATIONS block is generated** from that source, and CI asserts the
72
+ shipped binary's output, the docs bullets, and their tags all agree.
73
+
74
+ ### Changed
75
+
76
+ - ⚠️ **BREAKING: per-turn run-directory layout, single shape — the root-level `result.json` compatibility
77
+ copy is REMOVED.** A run directory that holds several turns (any `--resume`, and every `critique`) writes
78
+ each turn's `result.json`, `run.jsonl`, `trace.json` and `resources.jsonl` into **`turns/<N>/`**, once,
79
+ under its final name — nothing is renamed or overwritten as later turns arrive. `chat` now goes through
80
+ the same layout too (always `turns/1/` — a `chat` session mints a fresh dir per invocation and never
81
+ resumes). **`<outDir>/result.json` no longer exists — there is no root compat copy of any per-turn
82
+ artifact, on ANY run dir.** Read `turns/<N>/result.json` directly (`turns/1/` for a single-turn run), or
83
+ — for `critique` — the unchanged `result.graded.json` / `trace.graded.json` role aliases. Cumulative
84
+ streams (`events.jsonl`, `timeline.jsonl`) and session state are unchanged, so `critique`'s byte-offset
85
+ turn-isolation proof and cassette capture are unaffected.
86
+
87
+ **Two prior shapes are now REFUSED, loudly, by name, instead of being silently misread:**
88
+ - a **pre-layout** run dir (written before `turns/<N>/` existed: root `result.json`/`run.jsonl`, or a
89
+ name-mangled `result.turn-<N>.json` archive, no `turns/`);
90
+ - a **mixed** run dir (a pre-layout dir resumed under CURRENT code before this release — `turns/` present
91
+ *and* a stray root/archived file).
92
+
93
+ `verify-run`, `inspect`, `scaffold`, `diff`, `status --latest-for`, and a resumed `--session-id` all
94
+ refuse these with a message naming the shape found and pointing at `trace <dir>` — which still works
95
+ fully, since every one of its views derives from `events.jsonl`, which never moves. `stats --reindex`
96
+ counts them as skipped and names the remedy rather than dropping them from the index quietly.
97
+
98
+ **Migration: `cowork-harness migrate-run-dir`** converts a pre-layout dir in place (dry-run by default),
99
+ preserving the file timestamps `stats` and `status --latest-for` rank by. `diff` and
100
+ `status --latest-for` are called out because their pre-refusal behaviour was the dangerous kind: `diff`
101
+ reported two genuinely different runs as `identical` and exited 0, and `status --latest-for` could
102
+ select a *different* run than the newest and report its verdict — a CI script reading `.verdict.pass`
103
+ got a green light for a red run.
104
+
105
+ Previously the latest turn lived at the root while earlier ones were name-mangled archives, so a file's
106
+ name depended on whether a later turn ever happened; that shape produced a wrong-turn read, a destroyed
107
+ trace, and a dropped index row — this release's read-side (`turnArtifactPath` / `listTurns` in
108
+ `turn-layout.ts`, with the old `readTurnResult` deleted for having zero production callers) no longer has
109
+ a legacy-resolving branch at all, so that class of bug is now unrepresentable rather than merely fixed. The Python SDK's `_latest_run_jsonl` likewise now raises loudly
110
+ on a pre-layout dir instead of silently falling back to a root `run.jsonl` that (for any current-layout
111
+ dir) is a path to nowhere.
112
+
113
+ - **Platform baseline synced to Desktop 1.24012.1** (`baselines/desktop-1.24012.1.json`, now what
114
+ `baseline: latest` resolves to). The staged **agent binary is `2.1.217`** (native app + VM ELF, new
115
+ sha256 for each). The baseline moved in two steps this release — an earlier sync to **1.24012.0** (agent
116
+ `2.1.215`), then to 1.24012.1 — with **no prompt, spawn-env, or egress-allowlist drift across either**:
117
+ `spawn.env` is byte-identical to 1.22209.3, the same 15-domain allowlist and `gvisor` mode carry over
118
+ (the effort map is the one spawn field that changed — see the sonnet-5 delta below), and the
119
+ `deriveSpawnEnv` / `checkSpawnContractFacts` oracles stay green against the live asar. The
120
+ substantive deltas all came from the 1.24012.0 step and carry forward unchanged: `claude-sonnet-5` joins
121
+ the per-model effort map (`low|medium|high|xhigh|max`, recommended `medium`, modes `auto`); the
122
+ `coworkRuntimeConfig` gate drops its `pluginsFullSyncStalenessMs` key (never modeled here, inert); and
123
+ the dormant `autoModeOverridesAlwaysAllow` sentinel fired — see below. 1.24012.1 itself adds only the
124
+ agent bump: the `2.1.217` binary can emit the VCS SDK events `code_change_published` /
125
+ `vcs_state_changed` (SDK floor `2.1.216`), which the harness surfaces as a `system_event` (its existing
126
+ graceful degradation of an unmodeled system event, unchanged), and the binary's native skill-discovery
127
+ enable predicate widened to three branches — still inert here, since real Cowork's model-visible surface
128
+ is the Desktop SDK-MCP discovery servers, not the native tools. The example cassettes'
129
+ `fingerprint.baseline` tracks the new baseline.
130
+ - **The `autoModeOverridesAlwaysAllow` gate (`4200321681`) flipped absent → on** (`source: force`) and was
131
+ revisited as its pin intended. It stays **unmodeled, deliberately**: binary-verified in 1.24012.0, both
132
+ call sites only override an *already-existing* always-allow decision — the session rule cache
133
+ (`approvedToolNames`) and scheduled-task auto-approval — each further gated on `permissionMode` and
134
+ `isDestructiveConnectorTool`. The harness persists neither, so it already prompts wherever the gate makes
135
+ Cowork prompt; enabling it moves real Cowork *toward* harness behavior rather than away. Revisit only if
136
+ the harness gains a persistent per-tool approval cache.
137
+ - **The staged agent (`2.1.217`) enforces sub-agent fan-out caps, so `dispatch_count_max` is now framed as
138
+ a budget UNDER Cowork's cap, not a reproduction of it.** Because the harness spawns the real binary, a
139
+ run that fans out past the agent's caps now errors from the binary itself: a **concurrent** cap
140
+ (`CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS`, default 20; error `subagent_concurrency_cap`, new in 2.1.217)
141
+ and a **per-session** cap (`CLAUDE_CODE_MAX_SUBAGENTS_PER_SESSION`, default 200; error
142
+ `subagent_count_cap`, present since ≤2.1.215). The scenario schema's `dispatch_count_max` description,
143
+ the `assert` over-budget message, and SPEC §10 no longer claim "Cowork imposes no in-conversation
144
+ Task-dispatch cap" — that claim was stale. The harness does not reproduce the caps; it inherits them by
145
+ running the binary.
146
+ - **A host-loop `exec` infrastructure failure now WARNS instead of failing the run.** ⚠️ **Upgrade note:**
147
+ a run that previously exited `1` because one `docker exec` failed will now exit `0`. A dead sidecar
148
+ still hard-fails. Known residual, documented in `docs/scenario.md`: if *every* exec failed the agent ran
149
+ nothing and the run still only warns — inspect `result.infraErrors` when a run looks suspiciously empty.
150
+ - **A model-requested bash `timeout_ms` expiry is no longer classified as an infrastructure failure.**
151
+ The model now receives its command's own partial output with `Command timed out after <duration>`
152
+ merged into stderr — matching real Cowork, verified against the staged agent binary — instead of an
153
+ opaque `[infrastructure error: see run log for details]`.
154
+ - **The agent's spawn env now always carries a normalized IANA `TZ`** — matching Desktop, which injects
155
+ `Intl.DateTimeFormat().resolvedOptions().timeZone` unconditionally. Previously `TZ` was forwarded only
156
+ when the host shell exported it, and forwarded raw, so a host with no `TZ` set — or a legacy/non-IANA
157
+ export (`US/Eastern`, `EST5EDT`) — diverged from real Cowork's date/"today" rendering inside the agent.
158
+ - **The `tool_available` assertion now names its evidence limit.** It evaluates against the run's
159
+ *eagerly-loaded* tool set (the SDK init manifest in `result.json`); a factory-deferred tool — e.g. the
160
+ skill-discovery MCP tools, loaded on demand via a `ToolSearch` round-trip — can be genuinely available in
161
+ the run yet miss here, a false negative. The assertion still fails on a miss; the failure message now
162
+ states that eagerly-loaded scope rather than implying provable unavailability.
163
+ - **An explicitly requested `--dotenv` file now fails loud.** ⚠️ **Upgrade note:** an unreadable file, or
164
+ a path that is a directory, previously fell through to lower-precedence `.env` sources *while still
165
+ printing a success line* — so a typo'd path silently ran against the wrong credentials. It is now a
166
+ usage error. Automatic `.env` discovery is unchanged (still best-effort).
167
+ - **`diff` no longer reports `identical` when only one side has an artifact manifest.** ⚠️ **Upgrade
168
+ note:** such a comparison previously exited `0`; it now exits `1`, because unavailable evidence is not
169
+ evidence of equality. Both-sides-missing still does not veto identity. The `--output-format json`
170
+ envelope gained an `artifactsAvailability` key.
171
+ - **`stats --reindex` merges rows by per-completion identity** (`outDir` + a new `turn` field) rather than
172
+ by `outDir` alone, and reports rejected symlinked run directories.
173
+
174
+ ### Fixed
175
+
176
+ - **`verify-run` now REFUSES a multi-turn run directory** instead of certifying the wrong turn. Root
177
+ `result.json` is the latest turn; on a `critique` directory that is the *reflection* turn while the
178
+ scenario describes the *graded* one. Previously the cumulative gate scan false-FAILED on the other
179
+ turn's gates — wrong, but loud. The refusal names `result.graded.json` / `turns/1/result.json` so the
180
+ caller can still reach the graded turn.
181
+ - **`trace` no longer mixes turn scopes.** After timeline reads became turn-scoped, `--view
182
+ tool-durations` showed the latest turn while the tools/questions/dispatches views still showed every
183
+ turn — two views of one run directory describing different scopes. All views are now the latest turn,
184
+ and a `::notice::` reports when earlier turns exist rather than hiding them. Its cache-read footer and
185
+ gate-provenance (`answeredBy`) views also now say when a result is not turn-addressable (a pre-layout
186
+ dir) instead of silently omitting the cache-read ratio / labels.
187
+ - **`prune` no longer demotes an unmigrated pre-layout run dir to the junk tier.** Its real-run predicate
188
+ keyed on `hasTurnDirs || events.jsonl`, reasoning only about what current writers produce — but `prune`
189
+ ranks *history*, including the legacy dirs `migrate-run-dir` exists to preserve, so a legacy dir with no
190
+ `events.jsonl` could be evicted ahead of an empty scaffold. It now also counts a `legacy` / `mixed` shape
191
+ as a real run. (Distinct from the in-flight-migration deferral above — this is about which dirs count as
192
+ real at all.)
193
+
194
+ - **A resumed turn was judged on the PRIOR turn's evidence — three wrong-verdict paths.** `events.jsonl`
195
+ is append-only across turns with no per-turn marker, and three whole-file scanners decide a run's
196
+ outcome: `scanEvents` (outputs-delete / host-path-leak → fail signals, and an authored
197
+ `no_delete_in_outputs`), `findUngatedPathToolCalls` (→ a run-level `error` at hostloop), and
198
+ `detectCapabilityUse` (→ `missing_capability`, a fail signal, which fires on the default lean image).
199
+ So on any `--resume` — and every `critique` reflection turn — turn 1's delete, gated tool call, or
200
+ capability use FAILED turn 2. A turn-start marker now scopes all three to the current turn.
201
+ `resources.jsonl` had the same shape (turn 1's peak RSS judged against turn 2's `max_peak_rss_bytes`)
202
+ and is archived per turn. Single-turn runs write no marker, so their `events.jsonl` is byte-identical
203
+ and no cassette is affected. Missing marker ⇒ whole-file scan, i.e. fail-closed.
204
+
205
+ - **A resumed turn's telemetry included the PRIOR turn's events, and could produce a false PASS.**
206
+ `timeline.jsonl` is append-mode with a fresh header per turn, but `readTimeline` returned every line
207
+ after the first as an event — so on any `--resume` (and every `critique` reflection turn) the current
208
+ turn's `toolDurations`/`skillActivity`/`subagents` folded in the previous turn's tool calls. Because
209
+ the **`skill_tool_used` assertion** evaluates against that same `skillActivity`, a turn-1 skill window
210
+ could satisfy a turn-2 assertion. The reader now returns only the current turn's segment. The file
211
+ stays one append-only stream, so `critique`'s byte-offset turn-isolation proof is unaffected.
212
+
213
+ - **A resumed turn destroyed the prior turn's `trace.json`.** Because it is rebuilt and overwritten on
214
+ every completion, the earlier turn's trace was deleted rather than preserved, so a `critique` lost the
215
+ graded turn's trace entirely. Each turn now owns its own `turns/<N>/trace.json`, written once and never
216
+ overwritten, and `critique` additionally writes **`trace.graded.json`** beside `result.graded.json`.
217
+ - **`stats --reindex` dropped every non-latest turn when rebuilding from the runs tree.** It read only the
218
+ root `result.json` per run directory, so a resumed session's earlier turns vanished — and on a
219
+ `critique` directory the root file is the *reflection* turn, so it was the **graded** rows that were
220
+ lost. Every turn under `turns/<N>/` is now indexed as its own completion; `result.graded.json` — a
221
+ root-level copy of the graded turn — is deliberately not matched, so it cannot double-count. A dir that
222
+ has not been migrated is counted as `skippedLegacy` and reported with the remedy, never dropped
223
+ silently.
224
+
225
+ - **An ambient `GIT_DIR` silently computed the wrong skill file set.** Git hooks export `GIT_DIR` (and
226
+ `GIT_INDEX_FILE`) into every child process, and with `GIT_DIR` set but no `GIT_WORK_TREE` git stops
227
+ inferring the work tree from `cwd` and treats `cwd` as the repo root. `gitTrackedSet`'s
228
+ `rev-parse --show-toplevel` probe therefore still succeeded — so the not-a-repo raw-walk fallback never
229
+ fired — while `git ls-files -- .` returned the **entire repo index as root-relative paths** instead of
230
+ the directory-relative ones. Measured on this repo: 2 tracked files became 625 wrong ones. That set
231
+ feeds both `skillHash` and the mount-copy filter, so any run invoked from a git hook (or from CI that
232
+ exports `GIT_DIR`) got a wrong hash and a mount filter pointed at paths that do not exist under the
233
+ skill dir. The visible symptom was the repo's own pre-commit hook reporting committed example cassettes
234
+ as `[stale] skill files changed since record` on every parity sync. `skillCommit` had the same defect:
235
+ `git -C <dir>` is overridden by an ambient `GIT_DIR`, so every skill dir resolved to that foreign repo's
236
+ HEAD — recording a foreign commit as the skill's provenance and masking dirs that are genuinely in
237
+ different repos. Both call sites now spawn git with `GIT_DIR` / `GIT_WORK_TREE` / `GIT_INDEX_FILE`
238
+ stripped, via one shared helper so they cannot drift. `run-index`'s `gitInfo` and `doctor`'s worktree
239
+ probe deliberately keep inheriting — they are asking about the *ambient* repo.
240
+ - **`stats --reindex` destroyed multi-turn history.** Every `--resume` turn — and `critique`'s task +
241
+ reflection pair — writes to one `outDir`, so keying by directory collapsed N completions into one,
242
+ silently changing run counts, pass rates and costs.
243
+ - **Host-loop sidecar failures never reached the verdict.** They were appended straight to `events.jsonl`,
244
+ which no live drive re-reads, so a dying sidecar left the run green; a signal-only termination (OOM,
245
+ `SIGKILL`) was recorded nowhere at all.
246
+ - **`result.json` was written non-atomically** at all three producers, so an interrupted write could leave
247
+ the canonical record truncated.
248
+ - **Corrupt index rows were blind-cast**, letting one malformed row crash `stats` or fabricate a
249
+ pass/cost value; `reindex` also followed symlinks out of the runs root.
250
+ - **`scaffold` turned unavailable artifact evidence into "no artifacts"**, permanently encoding a false
251
+ "this run produced nothing" claim into a generated scenario.
252
+ - **`critique` treated a vanished turn-1 evidence file as genuinely empty evidence** rather than an
253
+ integrity failure. A stream that was legitimately zero bytes at capture is still reported clean.
254
+ - **`critique`'s exit-code table omitted `1`.** Exit `1` is reachable on operator interrupt
255
+ (SIGINT/SIGTERM); a sweep wrapper treating it as impossible misreads a cancelled run as a crash.
256
+ - Documented that after critique's resume, `result.json` is the **reflection** turn's result — the graded
257
+ turn is archived as `result.turn-1.json`. Reading the wrong one yields a valid-looking wrong number.
258
+
9
259
  ## [1.6.0] — 2026-07-20
10
260
 
11
261
  ### Added
package/CONTRIBUTING.md CHANGED
@@ -44,6 +44,9 @@ src/
44
44
  decide/ Decider — answer policy (scripted / LLM / external-channel deciders)
45
45
  run/ Run — turn loop + RunRecord (run.ts); executeScenario (execute.ts); cassette replay
46
46
  runtime/ protocol (L0) / container (L1) / microvm (L2) / hostloop / lima
47
+ hostloop/ Cowork host-loop handlers — can_use_tool gate, web_fetch dedup, workspace/path hooks
48
+ critique/ critique — run a skill, then grade its self-report against a frozen run record
49
+ staging/ agent-binary resolution + mount naming for the sandboxed live tiers
47
50
  egress/ default-deny allowlist proxy
48
51
  boundary.ts sandbox self-test probes
49
52
  assert.ts synchronous assertion evaluator