cowork-harness 1.6.0 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +80 -21
- package/.claude/skills/cowork-harness/references/ci-recipe.md +10 -9
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +18 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +13 -11
- package/.claude/skills/cowork-harness/references/task-recipes.md +3 -3
- package/CHANGELOG.md +420 -0
- package/CONTRIBUTING.md +10 -6
- package/DESIGN.md +1 -1
- package/README.md +50 -29
- package/RELEASING.md +7 -5
- package/SPEC.md +14 -4
- package/baselines/desktop-1.24012.0.json +405 -0
- package/baselines/desktop-1.24012.1.json +485 -0
- package/dist/agent/session.js +9 -3
- package/dist/agent/timeline.js +76 -3
- package/dist/assert.js +4 -2
- package/dist/baseline.js +20 -0
- package/dist/cli.js +216 -36
- package/dist/critique/command.js +450 -43
- package/dist/critique/evaluator.js +111 -43
- package/dist/critique/evidence.js +80 -43
- package/dist/critique/limitations.js +134 -0
- package/dist/critique/package-evidence.js +149 -30
- package/dist/decide/llm-transport.js +3 -1
- package/dist/dotenv.js +38 -3
- package/dist/errors.js +12 -0
- package/dist/hostloop/canusetool-gate.js +4 -4
- package/dist/hostloop/safety.js +6 -5
- package/dist/hostloop/workspace-handler.js +35 -1
- package/dist/io.js +15 -7
- package/dist/run/analyze-artifact-runtime.js +3 -4
- package/dist/run/analyze-artifact.js +1 -1
- package/dist/run/artifacts.js +12 -0
- package/dist/run/cassette.js +4 -0
- package/dist/run/chat-result.js +14 -7
- package/dist/run/chat.js +39 -6
- package/dist/run/diff.js +26 -5
- package/dist/run/doctor.js +67 -1
- package/dist/run/execute.js +206 -67
- package/dist/run/inspect-view.js +11 -4
- package/dist/run/latest-run.js +33 -7
- package/dist/run/migrate-run-dir.js +817 -0
- package/dist/run/renderer.js +11 -3
- package/dist/run/run-index.js +225 -45
- package/dist/run/run.js +13 -4
- package/dist/run/runs-gc.js +45 -8
- package/dist/run/scaffold.js +28 -4
- package/dist/run/skill-files.js +18 -1
- package/dist/run/skill-flag-surface.js +13 -2
- package/dist/run/subagent-reasoning.js +58 -3
- package/dist/run/trace-view.js +110 -12
- package/dist/run/turn-events.js +34 -0
- package/dist/run/turn-layout.js +189 -0
- package/dist/run/verdict.js +17 -2
- package/dist/runtime/host-env.js +8 -2
- package/dist/runtime/hostloop-prompt.js +23 -4
- package/dist/runtime/hostloop.js +67 -20
- package/dist/runtime/image-capabilities.js +22 -8
- package/dist/runtime/resource-sampler.js +20 -6
- package/dist/sync/cowork-sync.js +137 -17
- package/dist/types.js +17 -2
- package/docker/Dockerfile.agent +9 -0
- package/docs/README.md +8 -3
- package/docs/boundary.md +1 -1
- package/docs/cassette.md +5 -5
- package/docs/chat.md +5 -4
- package/docs/critique.md +168 -10
- package/docs/debugging.md +47 -3
- package/docs/fidelity-gaps.md +51 -11
- package/docs/gotchas.md +9 -1
- package/docs/maintenance.md +2 -2
- package/docs/run-status.md +4 -3
- package/docs/scenario.md +21 -5
- package/docs/session.md +1 -1
- package/docs/stats.md +29 -7
- package/docs/subagents.md +1 -1
- package/examples/README.md +1 -1
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/llms.txt +2 -2
- package/package.json +1 -1
- package/python/README.md +7 -2
- package/python/cowork_harness.py +35 -11
- package/schema/run-result.json +37 -4
- package/schema/scenario.schema.json +1 -1
- package/scripts/bump-version.ts +5 -3
- package/scripts/check-versions.ts +45 -11
- package/scripts/release-preflight.ts +0 -2
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.8.0
|
|
7
|
+
tracks-harness: cowork-harness 1.8.0 (baseline desktop-1.24012.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
26
|
-
> `desktop-1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.8.0` (baseline
|
|
26
|
+
> `desktop-1.24012.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,9 +39,52 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
43
|
-
|
|
44
|
-
What the ≥ 1.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.8.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.8.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.8.0"`. **Pin `@>=1.8.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
|
+
|
|
44
|
+
What the ≥ 1.8.0 floor gates, by release:
|
|
45
|
+
|
|
46
|
+
- **1.8.0 (critique hardening + hostloop unpin):** `critique --skill <name>` (multi-skill-plugin
|
|
47
|
+
grading — a plugin root with no `--skill` is refused pre-spend; the evidence package gains the
|
|
48
|
+
invoked skill's `agents/<name>.md` + bounded `references/*.md` CONTENT), run-dir artifacts
|
|
49
|
+
(`critique-report.json` always, `critique-evidence-package.txt` when the evaluator ran,
|
|
50
|
+
`critique-salvage.json` on exit 2) + `--out <path>`, per-critique `costUsd` across all four
|
|
51
|
+
workloads, `findingFingerprint` per item (cross-INPUT clustering; skillHash stays the cross-FIX
|
|
52
|
+
key), gate-answer `--answer` echo in the report, per-item-tolerant evaluator parse
|
|
53
|
+
(`droppedEvaluatorItems`), `skillMdTruncated` (a readable-but-oversized SKILL.md is flagged
|
|
54
|
+
"graded a cut copy" — distinct from missing/unreadable), `subagents[].webSearches` +
|
|
55
|
+
`trace --view subagent-research` (live/record lane), and the hostloop uploads-bullet fix (Read of
|
|
56
|
+
an uploaded file at the advertised path works — no copy-into-outputs workaround needed). Also:
|
|
57
|
+
**`critique --fidelity hostloop`** (the container-tier pin is lifted — resume-continuity proven
|
|
58
|
+
live on the native binary; `microvm`/`protocol`/`cowork` stay refused, each with a tagged reason),
|
|
59
|
+
**`skill --allow-host-writes`** (hostloop writable-folder consent; `critique` forwards it to both
|
|
60
|
+
turns), `verdictProvenance` stamped on every critique report (advisory self-run, not an
|
|
61
|
+
attestation), evidence caps 64KB SKILL.md / 32KB transcript / 144KB package (a flagship-sized
|
|
62
|
+
SKILL.md grades untruncated; ~2–2.5× evaluator cost on large skills), pinned sessions stamp their
|
|
63
|
+
fidelity and a cross-tier `--resume` fails loud pre-spawn, and `doctor`'s advisory
|
|
64
|
+
`image-freshness` check (pulled agent image behind the published GHCR one → warn + re-pull
|
|
65
|
+
remedy).
|
|
66
|
+
|
|
67
|
+
- **1.7.0 (per-turn run-directory layout, single shape):** every run dir — `run`/`skill`/`chat`,
|
|
68
|
+
single-turn or multi-turn (any `--session-id` + `--resume`, and **every `critique`**: task turn +
|
|
69
|
+
reflection turn) — writes each turn's `result.json` / `run.jsonl` / `trace.json` / `resources.jsonl`
|
|
70
|
+
into **`turns/<N>/`**, written once and never renamed. **There is no root compat copy of anything —
|
|
71
|
+
`<run-dir>/result.json` does not exist.** **When a user asks about a run, point them at
|
|
72
|
+
`turns/1/result.json`** (the only completion on a single-turn dir) — or on a `critique` dir, better yet
|
|
73
|
+
at `result.graded.json` / `trace.graded.json`, the role-stable aliases `critique` writes (the graded
|
|
74
|
+
task turn, never the reflection one). `events.jsonl` / `timeline.jsonl` stay cumulative at the run-dir
|
|
75
|
+
root, always. A dir written before this layout existed is refused LOUD, by name, naming the shape found
|
|
76
|
+
— never silently misread as if it were `turns/1/`. **Convert it in place with
|
|
77
|
+
`cowork-harness migrate-run-dir` (dry-run by default); do NOT tell the user to re-run or re-record,
|
|
78
|
+
which throws away history the migrator recovers.** Its `events.jsonl` still fully supports `trace`,
|
|
79
|
+
which is why every refusal points there. `verify-run` REFUSES a multi-turn dir rather than certifying
|
|
80
|
+
the wrong turn, and `trace` shows the latest turn with a `::notice::` when earlier turns exist.
|
|
81
|
+
|
|
82
|
+
- **1.7.0 (limitation provenance):** every `critique` limitation in `critique --help` and
|
|
83
|
+
docs/critique.md is tagged with WHY it exists — `[structural]` (permanent), `[unverified]` (unproven,
|
|
84
|
+
**not** known-impossible), `[deliberate]`, `[not-built]`. **Read the tag before telling a user to
|
|
85
|
+
design around a limitation.** Worked example of why the tag matters: `critique`'s container-tier pin
|
|
86
|
+
was tagged `[unverified]`, not permanent — and 1.8.0 lifted it (hostloop allowed, see above) exactly
|
|
87
|
+
as the tag predicted a proof could.
|
|
45
88
|
|
|
46
89
|
- **1.6.0 (`critique`'s skill-flag parity):** `critique` accepts most `skill` flags under the same names,
|
|
47
90
|
which is what makes a document-analysis skill (cap table, deck, financial model, transcript)
|
|
@@ -81,7 +124,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
81
124
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
82
125
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
83
126
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
84
|
-
- **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) —
|
|
127
|
+
- **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
|
|
85
128
|
|
|
86
129
|
## Orient — the three loops
|
|
87
130
|
|
|
@@ -113,7 +156,7 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
113
156
|
> `node_modules/cowork-harness/docs/<name>.md` before assuming a pointer dangles. A **plugin**
|
|
114
157
|
> install loads a trimmed source-only cache where those pointers genuinely do dangle.
|
|
115
158
|
|
|
116
|
-
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
|
|
159
|
+
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · migrate-run-dir · lint ·
|
|
117
160
|
lint-skill · analyze-skill · probe-dispatch ·
|
|
118
161
|
verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
119
162
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
@@ -155,7 +198,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
|
|
|
155
198
|
> is your **working tree**, so an uncommitted edit to an already-tracked file *is* tested — you needn't
|
|
156
199
|
> commit to iterate. Only brand-new (untracked) files must be `git add`-ed to appear. Commit before you
|
|
157
200
|
> record the **locking cassette**, though: real Cowork ships the *committed* tree, so a green on
|
|
158
|
-
> uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder
|
|
201
|
+
> uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder mounts *empty* and the agent reports "the skill isn't
|
|
159
202
|
> installed" then did the work itself — a green-looking run where the skill never loaded. That now
|
|
160
203
|
> **hard-fails** (`BoundaryError`, exit 3) naming the dir, and a partially-tracked folder emits a loud
|
|
161
204
|
> `::notice:: [stage]` listing the excluded files. Fix: `git add` the skill, or `COWORK_HARNESS_GITSET=0`
|
|
@@ -372,7 +415,7 @@ cassette — has its own recipe:
|
|
|
372
415
|
"<q>=Israeli company"` binds whichever option starts with `Israeli company`. It is uniqueness-guarded and
|
|
373
416
|
**fails loud** if the anchor ever matches two options (the documented trade: drift-tolerance, not strict
|
|
374
417
|
CI reproducibility — for that, pin a full exact label or a free-text `answer:`).
|
|
375
|
-
3. **Budget ~1 re-run per file.** If a gate whiffs, the run
|
|
418
|
+
3. **Budget ~1 re-run per file.** If a gate whiffs, the run does not vanish — it exits non-zero but
|
|
376
419
|
**salvages a PARTIAL run** (the extraction the agent already did is written to disk). So the cost of a
|
|
377
420
|
missed gate is one re-run with a better `--intent` or a scripted answer, not a lost paid run.
|
|
378
421
|
4. **Inspect the outputs to judge correctness.** `cowork-harness inspect <run-dir>` shows what the run
|
|
@@ -435,6 +478,14 @@ Recognize these before "fixing" a non-bug:
|
|
|
435
478
|
`transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
|
|
436
479
|
run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
|
|
437
480
|
it's valid.
|
|
481
|
+
- **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
|
|
482
|
+
infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
|
|
483
|
+
rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
|
|
484
|
+
the fail-severity `infra_error`, where a **supervising process** died and contaminated everything. Note
|
|
485
|
+
a model-requested `timeout_ms` expiry is *not* this: it returns the command's partial output with
|
|
486
|
+
`Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
|
|
487
|
+
failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
|
|
488
|
+
suspiciously empty.
|
|
438
489
|
- **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
|
|
439
490
|
`RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
|
|
440
491
|
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
@@ -638,19 +689,27 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
638
689
|
`mcp__workspace__bash` alias, so a "sub-agent used no shell" check must glob **both** `Bash` and
|
|
639
690
|
`mcp__workspace__*` to hold across every tier.
|
|
640
691
|
|
|
641
|
-
6. **`dispatch_count_max` is
|
|
642
|
-
assertion: passing means "happened to dispatch ≤N this
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
`
|
|
646
|
-
only
|
|
692
|
+
6. **`dispatch_count_max` is your author-chosen budget UNDER Cowork's production cap, not a
|
|
693
|
+
reproduction of it.** It's a post-hoc count assertion: passing means "happened to dispatch ≤N this
|
|
694
|
+
run." Cowork DOES cap `Task` fan-out **agent-side** (`taskRegistry`: concurrent **20** /
|
|
695
|
+
per-session **200**, landed 2.1.212/2.1.217) — SEPARATE from the scheduled-task session limiter
|
|
696
|
+
(gate `1648655587`'s `{perTask:1, global:3}`, a different mechanism; binary-verified, `SPEC.md` §10
|
|
697
|
+
— repo-only). The harness **inherits** the production cap by spawning the real agent binary, so a
|
|
698
|
+
`dispatch_count_max` pass means "your tighter budget held," not "near a real limit"; use it to catch
|
|
699
|
+
a fan-out you don't want.
|
|
647
700
|
|
|
648
701
|
7. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
|
|
649
702
|
assertions need a sandboxed tier (`container`+). Good: this one fails loud by design.
|
|
650
703
|
|
|
651
|
-
8. **Read-only mounts are enforced; delete-deny is
|
|
652
|
-
(a write fails in-guest). But `rw` vs `rwd`
|
|
653
|
-
|
|
704
|
+
8. **Read-only mounts are enforced; delete-deny is a HARNESS gap — production DOES enforce it.**
|
|
705
|
+
`mode:r` mounts get a real `:ro` bind (a write fails in-guest). But `rw` vs `rwd`
|
|
706
|
+
(write-but-no-delete on `outputs/` / connected folders) is *not* mount-enforced **in the harness** —
|
|
707
|
+
`rm` succeeds and is only caught post-hoc by `no_delete_in_outputs`. **Real Cowork enforces it live:**
|
|
708
|
+
`rm` under `outputs/` fails `Operation not permitted`, and a skill must request approval via
|
|
709
|
+
`allow_cowork_file_delete` to delete. Two consequences: a skill should not stage disposable scratch
|
|
710
|
+
under `outputs/` (in production, cleanup there costs an approval prompt), and a skill's
|
|
711
|
+
"catch-EPERM-then-request-approval" branch cannot be exercised at any harness tier (the `rm` just
|
|
712
|
+
succeeds here). Do not read this gotcha as "delete-deny may not be real in production" — it is real.
|
|
654
713
|
|
|
655
714
|
9. **Keep `.env` out of any mounted folder** — it is copied into the sandbox and the token could
|
|
656
715
|
leak. Put it at a working-dir or install root (token resolution: env > `--dotenv` > `./.env` >
|
|
@@ -727,7 +786,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
727
786
|
|
|
728
787
|
17. **Editing `scenarios/*.yaml` `assert:` does NOT change a plain `replay`.** *Why:* `replay` evaluates the
|
|
729
788
|
assertions **frozen in the cassette** by default — it is byte-deterministic and ignores the working tree (so
|
|
730
|
-
a committed cassette can't silently re-interpret against an uncommitted YAML). This
|
|
789
|
+
a committed cassette can't silently re-interpret against an uncommitted YAML). This is *loud* rather than a *silent*
|
|
731
790
|
no-op; now plain `replay` prints a `::notice::` when a sibling's `assert:` differs and points you at the fix.
|
|
732
791
|
*Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
|
|
733
792
|
`--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.8.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -32,7 +32,7 @@ jobs:
|
|
|
32
32
|
- uses: actions/checkout@v4
|
|
33
33
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
34
34
|
run: |
|
|
35
|
-
V=2.1.
|
|
35
|
+
V=2.1.217 # match your scenario's pinned baseline's agentVersion
|
|
36
36
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
37
37
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
38
38
|
# verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.8.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.8.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.8.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -263,8 +263,9 @@ JSON, not the human-readable text (which is explicitly NOT stable).
|
|
|
263
263
|
|
|
264
264
|
A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
|
|
265
265
|
set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
|
|
266
|
-
artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl
|
|
267
|
-
`
|
|
266
|
+
artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl` and
|
|
267
|
+
`egress.log` at the root, plus each turn's `run.jsonl` / `trace.json` / `result.json` under `turns/<N>/`
|
|
268
|
+
(a single-turn run has just `turns/1/`; there is no root compat copy of any of these). Digest one with `cowork-harness trace <run-id | dir>`.
|
|
268
269
|
Secrets are scrubbed from every persisted log by value.
|
|
269
270
|
|
|
270
271
|
## Don't assume a fixed assertion count across lanes
|
|
@@ -293,7 +294,7 @@ does **not** imply the recording is still valid. Each replay result carries `sta
|
|
|
293
294
|
(A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
|
|
294
295
|
check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
|
|
295
296
|
`·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
|
|
296
|
-
gate (`ok:false`) — but it
|
|
297
|
+
gate (`ok:false`) — but it is class-AWARE on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
|
|
297
298
|
`format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
|
|
298
299
|
while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
|
|
299
300
|
fail it either way.)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -24,6 +24,22 @@ Self-contained reference. Tracks `cowork-harness 1.6.0` (baseline `desktop-1.201
|
|
|
24
24
|
`skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
|
|
25
25
|
positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
|
|
26
26
|
|
|
27
|
+
### hostloop paths: what Read sees vs what bash sees (uploads included)
|
|
28
|
+
|
|
29
|
+
At `hostloop` the native file tools (Read/Write/Edit/Glob/Grep) run **on the host** and are gated by a
|
|
30
|
+
path-containment hook (production's own model), while `mcp__workspace__bash` runs **in the VM sidecar**
|
|
31
|
+
and sees `/sessions/<id>/...` paths. Consequences an agent (or a debugging author) must know:
|
|
32
|
+
|
|
33
|
+
- A `/sessions/...` path handed to `Read` fails ("is a VM path") — that split is **production-faithful**,
|
|
34
|
+
not a harness bug. Translate via the prompt's path-mapping table.
|
|
35
|
+
- **Uploads ARE Read-able at hostloop**: the staged uploads dir on the host is in the containment
|
|
36
|
+
allowlist (read-only — writes/edits there are blocked, matching production). If a Read of an uploads
|
|
37
|
+
path fails "outside this session's connected folders", the path *form* is wrong (e.g. a relative
|
|
38
|
+
`mnt/uploads/...`, which resolves against the outputs cwd) — the file itself is reachable at the
|
|
39
|
+
staged host uploads dir; it does not need to be copied into `outputs/` first.
|
|
40
|
+
- In real Cowork an upload is a **hardlink** in the session's uploads dir; the harness **copies** it
|
|
41
|
+
instead. Read-reachability is the same; only the advertised parent dir can differ.
|
|
42
|
+
|
|
27
43
|
### `microvm` prerequisites & lifecycle
|
|
28
44
|
|
|
29
45
|
Requires macOS on arm64 and Lima (`brew install lima`; binary expected at
|
|
@@ -122,7 +138,7 @@ hands `fn` exactly this dict.
|
|
|
122
138
|
|
|
123
139
|
- On `skill`: `fail | prompt | first`.
|
|
124
140
|
- On `run`: `fail | first` (`prompt` rejected — it would break determinism).
|
|
125
|
-
- On `critique
|
|
141
|
+
- On `critique`: `fail | first` (`prompt` rejected — there is no TTY inside the spawned turn).
|
|
126
142
|
- **`llm` is NOT an `--on-unanswered` value.** The bare flag `--on-unanswered llm` is rejected; use
|
|
127
143
|
`--decider-llm` (CLI) or `on_unanswered: llm` (YAML).
|
|
128
144
|
- `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.8.0`
|
|
4
|
+
(baseline `desktop-1.24012.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -269,7 +269,7 @@ same set live from the schema.
|
|
|
269
269
|
| `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
|
|
270
270
|
| `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
|
|
271
271
|
| `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
|
|
272
|
-
| `dispatch_count_max: <N>` | at most N sub-agents dispatched —
|
|
272
|
+
| `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 |
|
|
273
273
|
| `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked via the `Skill` tool — evidence-unavailable (not a normal fail) if the agent's init tools have no `Skill` tool |
|
|
274
274
|
| `no_skill_triggered: <regex>` | no invoked skill id matched — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent or the `Skill` tool is unobservable |
|
|
275
275
|
| `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
|
|
@@ -348,16 +348,17 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
348
348
|
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
349
349
|
| `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
|
|
350
350
|
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
351
|
-
| `infra_error` | fail | A VM/egress sidecar
|
|
351
|
+
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
352
352
|
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
353
353
|
| `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
|
|
354
354
|
| `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
|
|
355
355
|
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
|
|
356
356
|
| `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
|
|
357
|
+
| `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
|
|
357
358
|
|
|
358
359
|
A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
|
|
359
360
|
overall run verdict and exit code — `assert result: success` alone won't catch it; check
|
|
360
|
-
`result.verdict.signals[].severity` or the run's exit code. Only the
|
|
361
|
+
`result.verdict.signals[].severity` or the run's exit code. Only the five **warn** codes are truly benign.
|
|
361
362
|
|
|
362
363
|
## Replay class
|
|
363
364
|
|
|
@@ -417,7 +418,7 @@ skill-source drift; every replay result also reports it class-tagged in `stalene
|
|
|
417
418
|
**Mixed assertions on replay:** before evaluating, `replay` strips each assertion to its replay-checkable
|
|
418
419
|
keys and drops any left empty. So `{result, egress_denied}` evaluates on replay as `{result}` alone — its
|
|
419
420
|
`egress_denied` half is removed (not AND-ed against an unreadable value); with a manifest, `file_exists`/
|
|
420
|
-
`artifact_json` are
|
|
421
|
+
`artifact_json` are not stripped. The harness is **loud in two classes**: a *full skip* (`::warning::`
|
|
421
422
|
with the count of pure live-only assertions not evaluated) and a *partial skip* (`::warning::` when a mixed
|
|
422
423
|
assertion's live-only half was dropped).
|
|
423
424
|
Two CI consequences: skipped assertions are **absent** from `results[].assertions[]` (not
|
|
@@ -529,11 +530,12 @@ reference omits). Neither list is a strict superset of the other — reach for t
|
|
|
529
530
|
11. **`subagent_declared_but_unused` fires on declared-but-didn't-use-THAT-tool**, even if the
|
|
530
531
|
sub-agent used other tools.
|
|
531
532
|
|
|
532
|
-
12. **`dispatch_count_max` is
|
|
533
|
-
and asserts on it; passing means "dispatched ≤N this
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
533
|
+
12. **`dispatch_count_max` is your author-chosen budget under Cowork's production cap, not a
|
|
534
|
+
reproduction of it.** It records the count and asserts on it; passing means "dispatched ≤N this
|
|
535
|
+
run," not "the harness capped it." Cowork DOES cap `Task` fan-out agent-side (`taskRegistry`:
|
|
536
|
+
concurrent 20 / per-session 200, landed 2.1.212/2.1.217), which the harness inherits by spawning the
|
|
537
|
+
real agent binary — SEPARATE from gate `1648655587`'s `{perTask:1, global:3}` scheduled/cron-task
|
|
538
|
+
session limiter, a different mechanism (binary-verified; SPEC §10).
|
|
537
539
|
|
|
538
540
|
13. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
|
|
539
541
|
assertions need `container`+. Fails loud by design.
|
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
# Task recipes — end-to-end paths for the jobs consumers actually do
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
|
-
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
|
|
4
|
+
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
+
Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -182,7 +182,7 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
182
182
|
(blinded evaluator + mechanical citation checking). **Critiquing a document-analysis skill?** The probe
|
|
183
183
|
attaches nothing on its own — pass `--upload <path>` (repeatable) or `--folder <dir>` exactly as you
|
|
184
184
|
would to `skill`, or the graded run has no file and "there was no file attached" is the correct finding,
|
|
185
|
-
not a skill defect. Source flags reach both spawned turns automatically
|
|
185
|
+
not a skill defect. Source flags reach both spawned turns automatically — see SKILL.md's
|
|
186
186
|
floor list. See docs/critique.md for the full flag table, cost and limits. If you
|
|
187
187
|
prefer to build your own grader, the substrate is still here:
|
|
188
188
|
- `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
|