cowork-harness 4.1.0 → 4.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +14 -10
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +4 -4
- package/.claude/skills/cowork-harness/references/assertions-guide.md +2 -2
- package/.claude/skills/cowork-harness/references/authoring.md +65 -14
- package/.claude/skills/cowork-harness/references/ci-recipe.md +23 -11
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/debugging.md +9 -3
- package/.claude/skills/cowork-harness/references/eval.md +79 -0
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/gotchas.md +19 -5
- package/.claude/skills/cowork-harness/references/measurement.md +10 -2
- package/.claude/skills/cowork-harness/references/run-record-replay.md +14 -7
- package/.claude/skills/cowork-harness/references/scenario-schema.md +16 -4
- package/.claude/skills/cowork-harness/references/task-recipes.md +32 -13
- package/.claude/skills/cowork-harness/scripts/scenario.py +56 -1
- package/CHANGELOG.md +381 -0
- package/DESIGN.md +2 -2
- package/README.md +10 -5
- package/SPEC.md +25 -13
- package/baselines/desktop-2.16120.0.json +1148 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +10 -1
- package/dist/agent/session.js +26 -0
- package/dist/assert.js +102 -13
- package/dist/baseline.js +39 -1
- package/dist/cli.js +156 -373
- package/dist/critique/command.js +19 -9
- package/dist/critique/skill-invocation.js +61 -1
- package/dist/decide/decider.js +25 -3
- package/dist/decide/semantic-judge.js +170 -37
- package/dist/decide/usage.js +52 -0
- package/dist/errors.js +41 -0
- package/dist/eval/classify.js +308 -0
- package/dist/eval/command.js +591 -0
- package/dist/eval/invocation.js +44 -0
- package/dist/eval/job-runner.js +50 -0
- package/dist/eval/manifest.js +11 -0
- package/dist/eval/pins.js +45 -0
- package/dist/eval/report.js +479 -0
- package/dist/eval/runs.js +127 -0
- package/dist/eval/schedule.js +31 -0
- package/dist/eval/snapshot.js +328 -0
- package/dist/eval/stats.js +240 -0
- package/dist/eval/usage.js +62 -0
- package/dist/hillclimb/schema-check.js +657 -0
- package/dist/run/api-retries.js +31 -0
- package/dist/run/artifacts.js +5 -4
- package/dist/run/authored-capture-opts.js +23 -0
- package/dist/run/cassette.js +211 -40
- package/dist/run/chat-result.js +9 -1
- package/dist/run/chat.js +125 -63
- package/dist/run/command-globals.js +15 -2
- package/dist/run/doctor.js +62 -45
- package/dist/run/execute.js +280 -80
- package/dist/run/lint-load.js +5 -2
- package/dist/run/model-provenance.js +50 -3
- package/dist/run/provenance.js +29 -11
- package/dist/run/renderer.js +15 -1
- package/dist/run/run-index.js +6 -0
- package/dist/run/run.js +8 -0
- package/dist/run/runs-gc.js +26 -1
- package/dist/run/verify-context.js +403 -0
- package/dist/runtime/agent-tree.js +480 -0
- package/dist/runtime/hostloop.js +6 -5
- package/dist/runtime/protocol.js +6 -1
- package/dist/scan.js +20 -0
- package/dist/session.js +5 -0
- package/dist/sync/cowork-sync.js +266 -5
- package/dist/termination.js +76 -10
- package/dist/types.js +2 -2
- package/docs/README.md +2 -1
- package/docs/boundary.md +7 -0
- package/docs/cassette.md +13 -6
- package/docs/chat.md +7 -1
- package/docs/ci.md +25 -1
- package/docs/cli.md +25 -14
- package/docs/companion-skill.md +2 -2
- package/docs/critique.md +2 -1
- package/docs/debugging.md +12 -4
- package/docs/eval.md +244 -0
- package/docs/fidelity-gaps.md +78 -9
- package/docs/run-status.md +7 -3
- package/docs/scenario.md +15 -12
- package/docs/session.md +1 -1
- package/docs/stats.md +15 -3
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/llms.txt +3 -2
- package/package.json +1 -1
- package/python/test_scenario_lint.py +71 -0
- package/schema/run-result.json +77 -1
- package/schema/scenario.schema.json +2 -2
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: cowork-harness
|
|
3
|
-
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
3
|
+
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did the change make its answers worse? (`eval`: paired, interleaved A/B of two plugin versions, pinned models). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 4.
|
|
7
|
-
tracks-harness: cowork-harness 4.
|
|
6
|
+
version: 4.2.0
|
|
7
|
+
tracks-harness: cowork-harness 4.2.0 (baseline desktop-2.16120.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -26,8 +26,8 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
|
|
|
26
26
|
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
27
|
Read them.
|
|
28
28
|
|
|
29
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.
|
|
30
|
-
> `desktop-2.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.2.0` (baseline
|
|
30
|
+
> `desktop-2.16120.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
31
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
32
32
|
|
|
33
33
|
## Preflight — make sure the harness can actually run
|
|
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
43
43
|
|
|
44
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
45
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
46
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.2.0"`. **Pin `@^4.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
47
47
|
|
|
48
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
49
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -76,8 +76,10 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
|
|
|
76
76
|
grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
|
|
77
77
|
surface; see `references/critique.md`.
|
|
78
78
|
- **Regression-test your skill's ANSWER quality** (not just its behavior — does its guidance still lead to
|
|
79
|
-
correct answers after you edit it?) → author `semantic_matches` scenarios
|
|
80
|
-
|
|
79
|
+
correct answers after you edit it?) → author `semantic_matches` scenarios, then compare the version
|
|
80
|
+
before your edit with the one after using `cowork-harness eval` (EXPERIMENTAL, live: 10 runs per
|
|
81
|
+
scenario at the defaults). See **Recipe 5** step 6 in `references/task-recipes.md` (validity,
|
|
82
|
+
discrimination — the traps) and [`references/eval.md`](references/eval.md).
|
|
81
83
|
- **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
|
|
82
84
|
<folder> --prompt "<probe>"`. Up to four model workloads (zero with `--corpus-only`; pass 2 is skipped with no self-report) and 10–20 minutes; budget from
|
|
83
85
|
`report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
|
|
@@ -97,7 +99,7 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
|
|
|
97
99
|
|
|
98
100
|
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · migrate-run-dir · lint ·
|
|
99
101
|
lint-skill · analyze-skill · probe-dispatch ·
|
|
100
|
-
verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
102
|
+
verify-run · trace · inspect · diff · critique · eval · eval report · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
101
103
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
102
104
|
|
|
103
105
|
## Invariants — how a green run lies
|
|
@@ -107,7 +109,8 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
|
|
|
107
109
|
|
|
108
110
|
1. **`result: success` is not "the task completed".** It means the agent didn't error. Assert the
|
|
109
111
|
deliverable (`file_exists` / `artifact_json` / `transcript_matches`). A `skill`-lane `PASS` only means
|
|
110
|
-
no guard fired: read `skillsInvoked
|
|
112
|
+
no guard fired: read `skillsInvoked` (plus `slashInvokedSkills` — a `/<skill>` prompt runs the skill
|
|
113
|
+
with no `Skill` call), `models` and `ablated` before concluding anything from it.
|
|
111
114
|
2. **`replay` skips live-only keys.** Filesystem and egress keys are skipped on replay (loudly), so a
|
|
112
115
|
mixed item like `{result, egress_denied}` greens on its content half. Keep one concern per `assert:`
|
|
113
116
|
item, put live-only checks on a live gate, and run `cowork-harness lint`.
|
|
@@ -154,4 +157,5 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
|
|
|
154
157
|
| [`references/fidelity-and-answers.md`](references/fidelity-and-answers.md) | tier semantics, answer paths, the determinism contract |
|
|
155
158
|
| [`references/ci-recipe.md`](references/ci-recipe.md) | the GitHub Action, replay-vs-live lanes, the four-stage pipeline |
|
|
156
159
|
| [`references/critique.md`](references/critique.md) | `critique` report and evidence-package shapes |
|
|
160
|
+
| [`references/eval.md`](references/eval.md) | `eval`: paired before/after of two plugin versions — labels, refusals, exit codes, files |
|
|
157
161
|
| `scripts/scenario.py` | `scaffold`, `lint`, `lint-skill`, `resolve-agent-types <plugin-dir>` (validates a pinned `subagent_type` against `plugin.json` + `agents/*.md`) |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertion catalog
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Every `assert:` key with its semantics, and the
|
|
4
4
|
verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
|
|
5
5
|
the scenario and session YAML fields are there too.
|
|
6
6
|
|
|
@@ -55,8 +55,8 @@ same set live from the schema.
|
|
|
55
55
|
| `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
|
|
56
56
|
| `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut. **Covers what the run dispatches** (`Agent`/`Task`, including `Agent(subagent_type:"fork")`), **not a `context: fork` skill** invoked through the `Skill` tool: that skill's own answer is never a dispatch — it comes back as the `Skill` tool result, which agent 2.1.284 builds as `Skill "<name>" completed (forked execution).`, a `Result:` line, then the answer. Assert on it with `tool_result_matches` anchored on that prefix, e.g. `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` (`[\s\S]*` because `.` stops at a newline; the match is case-insensitive, has no multiline flag so `^` is the start of the result, and sees the first 10 KB of each result). This covers a **foreground** fork only: a backgrounded fork's result is the line `Skill "<name>" launched (forked execution, running in the background).`, which carries no answer |
|
|
57
57
|
| `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 in `scenario-schema.md` |
|
|
58
|
-
| `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked via the `Skill` tool — evidence-unavailable (not a normal fail) if the agent's init tools have no `Skill` tool |
|
|
59
|
-
| `no_skill_triggered: <regex>` | no invoked skill id matched — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent
|
|
58
|
+
| `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked — via the `Skill` tool, or by a prompt starting `/<skill> …` / `/<plugin>:<skill> …` (the agent expands that itself with no `Skill` call; recorded as `slashInvokedSkills`) — evidence-unavailable (not a normal fail) if neither matched and the agent's init tools have no `Skill` tool, or the leading `/name` can't be resolved (ambiguous bare name, no skill inventory) |
|
|
59
|
+
| `no_skill_triggered: <regex>` | no invoked skill id matched, counting a slash-command invocation as well as a `Skill` call — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent, the `Skill` tool is unobservable, or the prompt's leading `/name` can't be resolved |
|
|
60
60
|
| `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
|
|
61
61
|
| `connector_available: <regex>` | an MCP server/connector's name matched the regex (available, not necessarily used) — evidence-unavailable if `RunResult.context.mcpServers` is absent |
|
|
62
62
|
| `tool_available: <regex>` | a tool in the init manifest matched the regex (available, not necessarily called — see `tool_called` for invocation) — evidence-unavailable if `RunResult.context.tools` is absent. The `mcp__skills__*`/`mcp__plugins__*` discovery tools are modeled (as `alwaysLoad`) on `container`/`hostloop`/`cowork` — a miss there is a real absence; `microvm`/`protocol` declare no such server, so a miss on those two tiers means "not modeled at this tier", not "provably unavailable" |
|
|
@@ -107,7 +107,7 @@ same set live from the schema.
|
|
|
107
107
|
| `egress_allowed: <host>` | the host was allowed through |
|
|
108
108
|
| `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
|
|
109
109
|
| `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
|
|
110
|
-
| `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, evidence_files?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). **A `context: fork` skill's own answer is not graded either**, even with `include_subagent_text: true`: it is not a dispatch (so it has no `subagents[]` entry to fold in), and it reaches the main agent as the `Skill` tool result, which the judged document excludes. The judge sees it only if the main agent restates it; to check the fork's answer directly, use `tool_result_matches` (see `subagent_output_contains`). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) **`evidence_files: [globs]` scopes which authored files are graded** — reach for it the moment a run authors more than a couple of files. The capture budget (64 KiB total by default) is spent prefix-major then alphabetically, so a pipeline that stages intermediates (`outputs/_work/*.json`) exhausts it before reaching its own deliverable and the verdict is refused evidence-unavailable over files no rubric mentions. Scoping also makes the capture spend the budget on the named files FIRST and exempts them from the per-file cap. Paths are `<root>/<rel>` (`outputs/report.md`, never a bare `report.md`; session-root writes are `scratchpad/<rel>`); globs are `*`/`?`/`**`, not regex. A glob matching nothing FAILS and the message lists every authored path — read it rather than guessing. Still too big? Raise `$COWORK_HARNESS_AUTHORED_TOTAL_BYTES`. The typed reason is on `RunResult.assertions[].semanticEvidence` — check `.reason` instead of parsing the message |
|
|
110
|
+
| `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, evidence_files?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). **A `context: fork` skill's own answer is not graded either**, even with `include_subagent_text: true`: it is not a dispatch (so it has no `subagents[]` entry to fold in), and it reaches the main agent as the `Skill` tool result, which the judged document excludes. The judge sees it only if the main agent restates it; to check the fork's answer directly, use `tool_result_matches` (see `subagent_output_contains`). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass, rationale?}]`, so a consumer can diff the per-claim profile across runs; `rationale` is the judge's one-sentence reason, printed under each failed claim in the failure footer: untrusted model text that can quote the judged document, whose content never affects `pass`, absent when the judge gave none (a reply whose shape is broken, such as unparseable JSON, a malformed `{"results": …}` group beside a valid grade, or a partial restatement that contradicts it, is retried once and then marked `judgeInvalid`), and comparable only between runs that share `judgePromptHash`); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible — or let `eval` compare two skill versions per claim, with the judge pinned). Each graded assert records the judge's provenance: `RunResult.assertions[].judgeModel` (the resolved model), `judgeCostUsd` (judge spend over both attempts, reported beside `cost.usd` and never inside it; absent when unpriced) and `judgePromptHash` (the grading-prompt template — compare only runs that share it). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) **`evidence_files: [globs]` scopes which authored files are graded** — reach for it the moment a run authors more than a couple of files. The capture budget (64 KiB total by default) is spent prefix-major then alphabetically, so a pipeline that stages intermediates (`outputs/_work/*.json`) exhausts it before reaching its own deliverable and the verdict is refused evidence-unavailable over files no rubric mentions. Scoping also makes the capture spend the budget on the named files FIRST and exempts them from the per-file cap. Paths are `<root>/<rel>` (`outputs/report.md`, never a bare `report.md`; session-root writes are `scratchpad/<rel>`); globs are `*`/`?`/`**`, not regex. A glob matching nothing FAILS and the message lists every authored path — read it rather than guessing. Still too big? Raise `$COWORK_HARNESS_AUTHORED_TOTAL_BYTES`. The typed reason is on `RunResult.assertions[].semanticEvidence` — check `.reason` instead of parsing the message |
|
|
111
111
|
| `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
|
|
112
112
|
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
113
113
|
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertions guide
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
|
|
4
4
|
|
|
5
5
|
### Assertions: two orthogonal axes
|
|
6
6
|
|
|
@@ -43,7 +43,7 @@ them by what you're trying to prove:
|
|
|
43
43
|
| a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
|
|
44
44
|
| a tool ran **inside** a skill's scope | `skill_tool_used: {skill, tool}` |
|
|
45
45
|
| a sub-agent did the work | `subagent_output_contains: {contains}`, `subagent_dispatched: <regex>`, `dispatch_count_max: <N>` |
|
|
46
|
-
| a `context: fork` skill answered correctly | `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` — its answer is the `Skill` tool result, not a sub-agent output, so `subagent_output_contains` and `semantic_matches` never see it (foreground fork only — a backgrounded fork's result carries no answer) |
|
|
46
|
+
| a `context: fork` skill answered correctly | `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` — its answer is the `Skill` tool result, not a sub-agent output, so `subagent_output_contains` and `semantic_matches` never see it (foreground fork only — a backgrounded fork's result carries no answer). Only when the MODEL invokes the skill: a `/<skill> …` prompt runs the fork with no `Skill` call and no such result — use `skill_triggered` + `transcript_matches` there |
|
|
47
47
|
| a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
|
|
48
48
|
| no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
|
|
49
49
|
| a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Authoring a scenario
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
|
|
4
4
|
|
|
5
5
|
## Part I — AUTHOR a scenario
|
|
6
6
|
|
|
@@ -220,25 +220,76 @@ so a file `run`/`record` would refuse fails lint too (running `scenario.py lint`
|
|
|
220
220
|
cowork-harness lint scenarios/*.yaml
|
|
221
221
|
```
|
|
222
222
|
|
|
223
|
-
`lint`
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
`
|
|
228
|
-
|
|
229
|
-
`
|
|
230
|
-
|
|
231
|
-
`
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
223
|
+
`lint` exits non-zero on any ERROR (CI-friendly); `--strict` also fails on WARN. Every rule it reports:
|
|
224
|
+
|
|
225
|
+
| Rule | Severity | Fires on |
|
|
226
|
+
|---|---|---|
|
|
227
|
+
| `assert-contradiction` | ERROR | assert items no single run can satisfy together |
|
|
228
|
+
| `assertions-key` | ERROR | `assertions:` instead of `assert:` — none of the checks would run |
|
|
229
|
+
| `authored-replay-fidelity` | ERROR | an authored `replay_protocol_fidelity` (only the replay lane synthesizes it) |
|
|
230
|
+
| `capabilities-on-protocol` | ERROR | non-empty `requires_capabilities` on `protocol` without `allow_missing_capability` — the probe cannot run there, so the run fails as unverifiable |
|
|
231
|
+
| `cassette-evidence-skipped` | INFO | with `--cassette-dir`: a cassette (or the directory) could not be read, so it cannot quiet replay-evidence advice |
|
|
232
|
+
| `container-only-key-off-container` | ERROR | `no_scratchpad_leak` off `container` — hostloop's `present_files` never promotes, so there is nothing to leak (WARN on `cowork`, whose tier resolves per the baseline gate) |
|
|
233
|
+
| `egress-on-protocol` | ERROR | an egress assertion (`egress_*` / `expect_denied`) on `protocol`, which enforces no egress |
|
|
234
|
+
| `enum-value-invalid` | ERROR | a field value outside its allowed set |
|
|
235
|
+
| `fidelity-missing` | ERROR | no `fidelity:` (required since 4.0.0) |
|
|
236
|
+
| `file-absent-contradiction` | ERROR | one path under both `file_exists` and `file_absent` |
|
|
237
|
+
| `gate-needs-controlout` | INFO | gate assertions, which evaluate on replay only when the cassette has `controlOut` |
|
|
238
|
+
| `host-path-assert-cowork` | WARN | `transcript_no_host_path` on `cowork` — it fails by design if the tier resolves to hostloop |
|
|
239
|
+
| `host-path-assert-tier` | ERROR | `transcript_no_host_path` on `hostloop` / `protocol`, where it fails by design |
|
|
240
|
+
| `lane-remote-incompatible-key` | ERROR | `present_files_called` / `no_scratchpad_leak` / `user_visible_artifact` on `lane: remote` (the runtime rejects them at load, so the tier rules are suppressed there) |
|
|
241
|
+
| `linter-extra-findings-invalid` | ERROR | the loader findings `cowork-harness lint` hands the linter could not be read |
|
|
242
|
+
| `linter-unclassified-key` | ERROR | a valid assertion key this linter cannot classify (the linter is out of date) |
|
|
243
|
+
| `manifest-needs-snapshot` | INFO | manifest-backed keys, which evaluate on replay only when the cassette carries an `artifacts` manifest |
|
|
244
|
+
| `mixed-assert-item` | WARN | one assert item mixing replay-checkable and live-only keys (replay drops the live-only half) |
|
|
245
|
+
| `no-scenarios` | ERROR | a linted directory with no `*.yaml` / `*.yml` |
|
|
246
|
+
| `not-found` | ERROR | a named file that does not exist |
|
|
247
|
+
| `parse` | ERROR | a file that is not YAML, or not a mapping |
|
|
248
|
+
| `positional-choose-order` | INFO | an answer rule with a positional `choose` (first / index), which option re-ordering can move |
|
|
249
|
+
| `present-files-key-off-tier` | ERROR | `present_files_called` on `protocol` / `microvm` (served only at `container` / `hostloop`) |
|
|
250
|
+
| `prompt-slash-not-leading` | WARN | a `prompt:` that names `/<skill>` without starting with it, so it is never expanded |
|
|
251
|
+
| `reference-access-contradiction` | ERROR | one reference under both `reference_read` and `no_observed_reference_access` |
|
|
252
|
+
| `regex-double-quoted` | WARN | a double-quoted regex with an unescaped backslash (YAML strips it) |
|
|
253
|
+
| `replay-noop` | WARN | every assertion is live-only or a verdict modifier, so a replay gate verifies nothing |
|
|
254
|
+
| `slash-prompt-forked-result-anchor` | WARN | a `prompt:` starting with `/<skill>` plus a `tool_result_*` anchored on `forked execution` — a slash-invoked skill makes no `Skill` call, so that tool result never exists; assert `skill_triggered` instead |
|
|
255
|
+
| `tool-called-always-passes` | INFO | `tool_called` with `count: {min: 0}` and no `max` — it asserts nothing |
|
|
256
|
+
| `tool-input-regex-redactable` | WARN | a `tool_not_called` input literal the redaction policy rewrites in the committed cassette (or a policy pattern it cannot check offline) |
|
|
257
|
+
| `tool-input-shell-tier` | INFO | the object form with `tool: Bash` and a `command` on `hostloop` / `cowork`, where shell runs as `mcp__workspace__bash` — list both |
|
|
258
|
+
| `tool-not-called-tier-vacuous` | WARN | `tool_not_called` / `subagent_tool_absent` naming a tool the tier never serves |
|
|
259
|
+
| `transcript-command-shaped` | WARN | a `transcript_*` value shaped like a shell command — those keys read prose only, never a tool call |
|
|
260
|
+
| `unknown-assert-key` | WARN | an assertion key not in the catalog (the loader rejects it) |
|
|
261
|
+
| `unknown-top-key` | WARN | a scenario key not in the schema |
|
|
262
|
+
| `vacuous-gate-assert` | WARN | `gate_answers_delivered` with no presence companion (zero gates passes it), or inert beside `questions_count_max: 0` |
|
|
263
|
+
| `scenario-invalid` | ERROR | the harness's scenario loader refuses the file (via `cowork-harness lint` only — see below) |
|
|
264
|
+
| `baseline-unknown` | ERROR | `baseline:` names no baseline this CLI ships (via `cowork-harness lint` only) |
|
|
265
|
+
| `lint-loader-internal` | ERROR | the wrapper could not run its loader check on a file — a harness bug; it never falls back to a lint that skipped the loader |
|
|
266
|
+
|
|
267
|
+
`scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never emits a scenario `lint`
|
|
268
|
+
would reject.
|
|
269
|
+
|
|
270
|
+
**Lint the skill itself: `cowork-harness lint-skill <skill-dir>`.** It checks the skill, not a scenario:
|
|
271
|
+
Cowork host-loop footguns (`${CLAUDE_PLUGIN_ROOT}` in a VM bash step, hook events, a misplaced
|
|
272
|
+
`hooks.json`, an unresolvable `subagent_type`), the evidence corpus a `critique` can package
|
|
273
|
+
(`references/critique.md`), and two size caps. `skill-body-over-reattach-cap` (WARN) fires when the
|
|
274
|
+
`SKILL.md` body, frontmatter excluded, passes 19,000 B — after a compaction the agent re-attaches only the
|
|
275
|
+
start of an invoked skill — and `skill-body-near-reattach-cap` (INFO) from 80% of that;
|
|
276
|
+
`skill-reference-over-read-cap` (WARN) fires on a `references/**.md` over 60,000 B, past which a
|
|
277
|
+
whole-file Read returns a partial view. `--strict` fails on WARN, never on INFO. To accept a reviewed
|
|
278
|
+
judgement-call finding, pass `--ignore-rule <rule>[=<glob>]` (repeatable; the glob matches the finding's
|
|
279
|
+
file) or fence the text in `SKILL.md` with `<!-- lint-skill: ignore-start <rule>[,<rule>…]: <reason> -->`
|
|
280
|
+
… `<!-- lint-skill: ignore-end -->` (outside any code fence). A suppressed finding is still printed; it
|
|
281
|
+
stops gating. A provable rule (an ERROR, a misplaced `hooks.json`, a missing pinned agent) cannot be
|
|
282
|
+
suppressed: naming it, or an unknown rule, in `--ignore-rule` is a usage error (exit 2); in a marker it is
|
|
283
|
+
WARN `lint-skill-ignore-invalid`, as is any other malformed marker. An unclosed marker is WARN
|
|
284
|
+
`lint-skill-ignore-unclosed`, and one that suppresses nothing is INFO `lint-skill-ignore-unused`.
|
|
235
285
|
|
|
236
286
|
**`cowork-harness lint` runs the loader: a file it calls clean is one `run`/`record` will load.** Anything
|
|
237
287
|
the loader refuses — an unknown key, a wrong value type (a scalar `semantic_matches.rubric`), a bad regex,
|
|
238
288
|
a reserved value — is ✗ ERROR `scenario-invalid` (exit 1, with or without `--strict`), and a `baseline:`
|
|
239
289
|
naming no baseline this installed CLI ships is ✗ ERROR `baseline-unknown` (`latest` always resolves). It
|
|
240
290
|
does not check what depends on the machine the run happens on (the session file and its mounts, an
|
|
241
|
-
absolute `baseline:` path
|
|
291
|
+
absolute `baseline:` path that does not exist here — one that exists is checked (4.1.1 and later) —
|
|
292
|
+
environment variables). A session or matrix YAML in a linted directory is not
|
|
242
293
|
a scenario and is reported as one that does not load — keep those out of the linted set. `python3
|
|
243
294
|
scenario.py lint` run directly stays offline and lenient: there an unknown key is only a ⚠ WARN (exit 0).
|
|
244
295
|
`cowork-harness record <file.yaml> --dry-run` also runs the loader and adds the pre-spend refusals (exit 2
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "4.
|
|
20
|
+
(e.g. `version: "4.2.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
82
82
|
GitHub-hosted runners, no token/Docker/agent:
|
|
83
83
|
|
|
84
84
|
```yaml
|
|
85
|
-
- run: npm i -g "cowork-harness@^4.
|
|
85
|
+
- run: npm i -g "cowork-harness@^4.2.0"
|
|
86
86
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
87
87
|
# no silent false-greens. WITHOUT --strict this
|
|
88
88
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -241,7 +241,9 @@ cowork-harness replay cassettes/ # replay every *.cassette.jso
|
|
|
241
241
|
Re-record whenever the protocol or your scenario's expected content changes. An old cassette without
|
|
242
242
|
`controlOut` excludes the gate keys (with a loud warning) — re-record to enable them. `record` **refuses
|
|
243
243
|
to freeze a failing live run** into a cassette (pass `--allow-failing` to override) — a committed red
|
|
244
|
-
cassette is a latent false-signal.
|
|
244
|
+
cassette is a latent false-signal. An `--allow-failing` recording of a red run exits 0 with `ok: true`:
|
|
245
|
+
`record`'s `ok` means a cassette was written, and the run's own verdict is `results[0].verdict.pass`
|
|
246
|
+
(`items[].verdict` on `record <dir/>`).
|
|
245
247
|
|
|
246
248
|
## Privacy: cassettes are committed fixtures → record only against SYNTHETIC inputs
|
|
247
249
|
|
|
@@ -256,7 +258,9 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
256
258
|
dir** (first file found per dir; env vars merge on top). `cowork-harness init-redact` copies the
|
|
257
259
|
packaged reference template (local-path prefixes, incl. macOS temp roots and slugged home segments
|
|
258
260
|
like `-Users-<user>-…`, + a generic email regex) into the cwd as a starting point — review and tailor
|
|
259
|
-
it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force`
|
|
261
|
+
it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` (without it
|
|
262
|
+
`init-redact` refuses to overwrite an existing policy; with it your tailoring is replaced — save and
|
|
263
|
+
re-apply it) or add them by hand. Redaction is **verdict-preserving** — `record` refuses to write if
|
|
260
264
|
redaction would flip an assertion (a manufactured green). `--no-redact` skips it for known-synthetic
|
|
261
265
|
inputs.
|
|
262
266
|
- **Pre-spawn preflight**: `record` warns (`::warning::`, before the paid run starts — once per batch
|
|
@@ -342,8 +346,8 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
342
346
|
This is the shape a CI step wants: **silent on success (no output, exit 0), loud and specific on
|
|
343
347
|
failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
|
|
344
348
|
offending file *and* the rejected key, one line per file, and the step still exits 1. It exits 0,
|
|
345
|
-
though, on an input the real record would refuse (a
|
|
346
|
-
tier-vacuous `tool_not_called`): that prints a `⚠ input error:` line and lands in `inputErrors[]`. To
|
|
349
|
+
though, on an input the real record would refuse (a `session:` file that cannot be read — 4.1.1 and
|
|
350
|
+
later — a missing path, an unknown baseline name, a tier-vacuous `tool_not_called`): that prints a `⚠ input error:` line and lands in `inputErrors[]`. To
|
|
347
351
|
gate on those too, use the JSON form (4.1.0 and later):
|
|
348
352
|
|
|
349
353
|
```bash
|
|
@@ -391,7 +395,7 @@ jobs:
|
|
|
391
395
|
with: { node-version: '24' }
|
|
392
396
|
- uses: actions/setup-python@v5
|
|
393
397
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
394
|
-
- run: npm i -g "cowork-harness@^4.
|
|
398
|
+
- run: npm i -g "cowork-harness@^4.2.0"
|
|
395
399
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
396
400
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
397
401
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -420,7 +424,7 @@ jobs:
|
|
|
420
424
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
421
425
|
fi
|
|
422
426
|
- if: steps.guard.outputs.live == 'true'
|
|
423
|
-
run: npm i -g "cowork-harness@^4.
|
|
427
|
+
run: npm i -g "cowork-harness@^4.2.0"
|
|
424
428
|
- if: steps.guard.outputs.live == 'true'
|
|
425
429
|
run: cowork-harness run scenarios/ --output-format json
|
|
426
430
|
env:
|
|
@@ -443,7 +447,8 @@ sandbox).
|
|
|
443
447
|
|
|
444
448
|
`--output-format json` emits a machine envelope on stdout (human output goes to stderr):
|
|
445
449
|
`{tool, version, command, ok, results[], error}` — one `RunResult` per scenario. **Overall pass for a
|
|
446
|
-
scenario is `verdict.pass`** (envelope-wide: `ok`
|
|
450
|
+
scenario is `verdict.pass`** (envelope-wide: `ok`; on `record`, `ok` only means the command exited 0 —
|
|
451
|
+
see below), and it is strictly stronger than
|
|
447
452
|
`result === "success" && assertions.every(pass)`: the verdict also carries ~20 signal codes that fail a run
|
|
448
453
|
with no failing assertion at all — `stalled`, `outputs_delete`, `mount_delete`, `host_path_leak`,
|
|
449
454
|
`undelivered_deliverables`, `missing_capability`, `permissive_auto_allow`, `ended_with_question`,
|
|
@@ -477,6 +482,11 @@ Keep the `.results[]?` hop and the `?` operators. `.results[0]` silently ignores
|
|
|
477
482
|
the first when you pass a directory, and a bare `.verdict` does not exist at the envelope root at all —
|
|
478
483
|
both read as "no failures" against a run that failed.
|
|
479
484
|
|
|
485
|
+
**`record` is shaped differently.** Its `ok` is the exit code's verdict (a cassette was written), not the
|
|
486
|
+
run's: `record --allow-failing` on a red run is `ok: true`. The verdict is in `results[0].verdict` on
|
|
487
|
+
`record <file>` and in `items[].verdict` on `record <dir/>` and `--rerecord-stale`, which carry no
|
|
488
|
+
`results[]` — so `.results[]?` reads nothing there. Use `[.items[]? | .verdict.failures[]?]` on a batch.
|
|
489
|
+
|
|
480
490
|
> **Do not filter on whether `assertion` is present.** That was the only discriminator before `kind`
|
|
481
491
|
> existed and it never worked in both directions: `coverage` entries carry a key too (an internal
|
|
482
492
|
> `answer_coverage` marker), so they read as authored asserts, while `guard`, `staleness` and
|
|
@@ -486,7 +496,9 @@ both read as "no failures" against a run that failed.
|
|
|
486
496
|
`--output-format json`, `run` / `record` / `replay` / `verify-cassettes` / `status` write their whole
|
|
487
497
|
human rendering — warnings, verdict, `status`'s summary line — to **stderr**, and stdout stays empty.
|
|
488
498
|
A wrapper that captures only stdout gets an empty log and, if it greps that for a state, a silent false
|
|
489
|
-
negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout
|
|
499
|
+
negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout —
|
|
500
|
+
`COWORK_HARNESS_OUTPUT_FORMAT=json` makes JSON the default for every command that takes `--output-format`
|
|
501
|
+
(an explicit flag still wins).
|
|
490
502
|
(Commands whose whole job is to print a value — `--version`, `assertions --list`, `scaffold`, `gates`,
|
|
491
503
|
`skill --dry-run` — write it to stdout by design. Under `--output-format json` the value rides inside the
|
|
492
504
|
envelope: `scaffold`'s YAML is `.scenario`, `skill --dry-run`'s preview is the envelope's own fields;
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Debugging a run
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
|
|
4
4
|
|
|
5
5
|
## Part III — Debug
|
|
6
6
|
|
|
@@ -37,7 +37,11 @@ already does.
|
|
|
37
37
|
**microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
|
|
38
38
|
stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
|
|
39
39
|
`cowork-harness vm status` — a `provisioning` other than `ready` confirms it — and if a run does not
|
|
40
|
-
recover it on its own, `cowork-harness vm delete` and retry.
|
|
40
|
+
recover it on its own, `cowork-harness vm delete` and retry. A run on such a VM can also have cached an
|
|
41
|
+
empty toolchain for it in the capability probe's cache: `vm delete` drops that VM's entry (`vm prune`
|
|
42
|
+
forgets only the orphaned VMs it deletes, never the current one); otherwise delete `capability-cache.json`
|
|
43
|
+
from the runs root (`~/.cowork-harness/runs/` unless
|
|
44
|
+
`COWORK_HARNESS_RUNS_DIR` is set) so the next run probes again.
|
|
41
45
|
|
|
42
46
|
**Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
|
|
43
47
|
**deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
|
|
@@ -80,7 +84,9 @@ decide which assertions from *Assertions: two orthogonal axes* in `assertions-gu
|
|
|
80
84
|
walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
|
|
81
85
|
`total_cost_usd` for the run — the authoritative single-run spend of the agent session, which leaves out the `semantic_matches` judge and the LLM decider calls; NOT the same source as summing
|
|
82
86
|
`modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
|
|
83
|
-
`usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`,
|
|
87
|
+
`usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`,
|
|
88
|
+
`toolCalls` (every tool call in stream order: `name`, top-level `input` fields each capped at 10 KB, and
|
|
89
|
+
`origin` `main`/`subagent`/`unknown` — what the object form of `tool_called`/`tool_not_called` reads), `toolErrors`,
|
|
84
90
|
`redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
|
|
85
91
|
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
|
|
86
92
|
`context` (tools/mcpServers/availableSkills), `tasks`,
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
# `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
|
|
2
|
+
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). The full guide is
|
|
4
|
+
[docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
|
|
5
|
+
need while running it.
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
cowork-harness eval <scenario.yaml | dir/> --arm before=git:HEAD:plugins/my-skill --arm after=./plugins/my-skill \
|
|
9
|
+
--model <concrete id> --judge-model <concrete id> [--holdout <scenario.yaml>]... [--fail-on possible|confirmed]
|
|
10
|
+
cowork-harness eval report <eval-dir> # rebuild the report from the eval dir, no spend
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
## What it does
|
|
14
|
+
|
|
15
|
+
- Two arms: a plugin directory, or `git:<ref>:<path>` (read from the commit, `<path>` relative to the repo
|
|
16
|
+
root). The FIRST arm is the baseline; a drop is the second arm passing less often.
|
|
17
|
+
- Only the session's single `plugins.local_plugins` entry is substituted. Each arm is copied once, before
|
|
18
|
+
the first run, and every rep mounts the copy.
|
|
19
|
+
- The plugin is mounted at the `local_plugins` path (`mnt/.local-plugins/marketplaces/<marketplace>/<plugin>`),
|
|
20
|
+
not the `remote_plugins` path a UI-installed plugin has (`mnt/.remote-plugins/plugin_<id>`). A skill that
|
|
21
|
+
locates its own files at runtime sees a different path under each, so the comparison holds for the
|
|
22
|
+
`local_plugins` layout only.
|
|
23
|
+
- scenarios × 2 × `--reps` live runs (10 per scenario at the default `--reps 5`), interleaved, plus one judge
|
|
24
|
+
call per `semantic_matches` assert per run. The start-up line prints the job count. No budget flag.
|
|
25
|
+
- In a scenario directory, YAML with no `prompt:` (a session file) is skipped.
|
|
26
|
+
- Run an A/A first (`--allow-identical-arms`, the same source twice) to see your scenarios' noise.
|
|
27
|
+
|
|
28
|
+
## Reading the labels
|
|
29
|
+
|
|
30
|
+
| Label | Meaning |
|
|
31
|
+
|---|---|
|
|
32
|
+
| `confirmed drop/rise` | significant after the correction (bh q = 0.10, or holm) |
|
|
33
|
+
| `possible drop/rise` | p ≤ `--alpha` (0.05), not confirmed |
|
|
34
|
+
| `no detectable change` | with the smallest change this n could have detected (MDD) |
|
|
35
|
+
| `underpowered` | no outcome at these sizes could reach `--alpha` — NOT "no change" |
|
|
36
|
+
| `insufficient` | too few valid reps in an arm (4 of 5 needed by default) |
|
|
37
|
+
|
|
38
|
+
A drop is a signal to investigate, not proof: open the run dirs the report links for that row. In order,
|
|
39
|
+
first match wins: an infrastructure failure is excluded and reported — including a rep where no model
|
|
40
|
+
answered (the agent's `Not logged in` / `Authentication required` reply, rule `auth`; a usage or
|
|
41
|
+
spend limit as its final message, even on a nonzero exit after spend, rule `usage_limit`; or only
|
|
42
|
+
`<synthetic>` models at $0, rule `no_model_answered`). An agent-caused failure (timeout, max turns,
|
|
43
|
+
unanswered question, crash) then fails every row of its rep, even with its pin unknown. Only after that
|
|
44
|
+
are a pin the agent did not honour (false, or unknown on a rep that completed) and a snapshot that changed
|
|
45
|
+
excluded and reported. A loud UNCLASSIFIED count means a termination the classifier does not know — read
|
|
46
|
+
those runs. Per scenario: if EVERY rep of both arms errored, or every rep of one arm is infrastructure,
|
|
47
|
+
that scenario compared nothing — its rows are `insufficient` and the eval exits 1. If one arm's every rep
|
|
48
|
+
is the agent's own failure and the other arm ran, the reps are scored (a real drop) — exit 0 unless
|
|
49
|
+
`--fail-on`. Either way the header
|
|
50
|
+
names the arm, scenario, dominant error and a matching hint (`Every rep of arm <label> in <scenario> errored — …`).
|
|
51
|
+
|
|
52
|
+
## Refused before any run (exit 2)
|
|
53
|
+
|
|
54
|
+
- a model or judge that is an alias (`opus`, `best`) rather than a concrete id;
|
|
55
|
+
- an eval dir inside any git work tree (the snapshots would mount empty);
|
|
56
|
+
- identical arms (unless `--allow-identical-arms`);
|
|
57
|
+
- an arm that contains the eval's own scenario or session files (by location, copy or symlink), an
|
|
58
|
+
`evals.json`, or a symlink resolving outside it;
|
|
59
|
+
- a scenario input a run would refuse (a missing path, a `tool_not_called` the tier can never violate);
|
|
60
|
+
- `--fail-on confirmed` when no row could reach `confirmed` at this `--reps`;
|
|
61
|
+
- no usable agent credential for a scenario's tier — the same check as `doctor --tier <tier>`'s `token` row,
|
|
62
|
+
with its fix. A Keychain login or a `.credentials.json` in the config dir, without an env/.env token,
|
|
63
|
+
passes only at `protocol`;
|
|
64
|
+
- a session whose plugin is declared only under `plugins.remote_plugins` (the session must declare exactly
|
|
65
|
+
one `local_plugins` entry). Workaround: eval a copy of the session that declares the same directory under
|
|
66
|
+
`local_plugins`, and check the `remote_plugins` path handling with an ordinary `run`.
|
|
67
|
+
|
|
68
|
+
## Exit codes
|
|
69
|
+
|
|
70
|
+
`0` completed — no drop fails the eval unless you pass `--fail-on`. `1` a drop at the `--fail-on` level,
|
|
71
|
+
every row `insufficient`, a scenario that compared nothing, or the judge model differed across reps (an A/A run under `--fail-on possible`
|
|
72
|
+
can exit 1 on noise). `2` usage or a refusal. `3` an arm snapshot could not be copied or staged.
|
|
73
|
+
|
|
74
|
+
## Files
|
|
75
|
+
|
|
76
|
+
`<eval-dir>/` (default `~/.cowork-harness/evals/<eval-id>/`): `manifest.json`, `arms/`, `runs.jsonl`,
|
|
77
|
+
`report.json` (every rep with its bucket), `report.md`. The runs are ordinary run dirs labelled
|
|
78
|
+
`eval:<eval-id>:<arm>`; a bare `prune` keeps 5 per scenario and says which evals it trimmed — `eval report`
|
|
79
|
+
still works, but the evidence links then dangle.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Gotchas
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). The full "✓ passed ≠ correct" landmine catalog.
|
|
4
4
|
|
|
5
5
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
6
6
|
|
|
@@ -175,9 +175,10 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
175
175
|
`--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
|
|
176
176
|
`prompt`/`answers`/`baseline`/`fidelity`/`lane`/`skills`/`requires_capabilities` or the skill content (when a
|
|
177
177
|
fingerprint exists) drifted from the recording (re-record then), and `expect_denied`/filesystem/egress keys
|
|
178
|
-
are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the
|
|
179
|
-
|
|
180
|
-
|
|
178
|
+
are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the session's
|
|
179
|
+
`model:` IS in the cassette's `sessionFingerprint` (a `--model`/env model is not; `environment.model`
|
|
180
|
+
records what ran). `verify-cassettes` reports a changed session as staleness (exit 1); `replay` never
|
|
181
|
+
checks it — plain, `--strict` or `--assert-from` — so re-record if the session changed. `verify-run` reads
|
|
181
182
|
on-disk `assert:` against a kept *run dir*; `replay --assert-from` is the equivalent for a *cassette*.
|
|
182
183
|
|
|
183
184
|
18. **`questions_count_max` counts sub-questions, not gates.** One `AskUserQuestion` tool call can
|
|
@@ -234,6 +235,17 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
234
235
|
fail your gate. *Fix:* nothing, unless the scenario asserts `tool_available` on
|
|
235
236
|
`mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
|
|
236
237
|
re-recording would never produce those tools anyway.
|
|
238
|
+
A sibling **`agent-version:` note** means the agent version the cassette's own `system/init` event
|
|
239
|
+
reports differs from the one the baseline its `fingerprint.baseline` names pins for that tier: the
|
|
240
|
+
`agentVersion` at `container`/`microvm`, the native agent in `agentBinary.nativeStagedPath` at
|
|
241
|
+
`hostloop`. It never appears at `protocol`, which runs the unpinned `claude` on your `PATH`. *Why:* one
|
|
242
|
+
of a fingerprint re-stamped by hand across an agent bump, a recording made under
|
|
243
|
+
`COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1`, an explicit binary override (`COWORK_AGENT_BINARY`, or
|
|
244
|
+
`COWORK_HOST_AGENT_BINARY` at `hostloop`), or at `hostloop` the default patch-bump substitution of the
|
|
245
|
+
native agent; the note lists that tier's causes and does not pick one. It is non-gating too:
|
|
246
|
+
`verify-cassettes` puts it in the result's `notes[]`, and `replay` prints one
|
|
247
|
+
`::notice:: [replay] <file> — … [agent-version]` line per cassette on stderr (also under
|
|
248
|
+
`--output-format json`). *Fix:* re-record against the pinned agent.
|
|
237
249
|
|
|
238
250
|
24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
|
|
239
251
|
lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
|
|
@@ -270,7 +282,9 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
270
282
|
`run` the same word additionally means *your assertions held*; on `skill --repeat N`, `PASS — N/N`
|
|
271
283
|
means N runs cleared the guards — it says nothing about which model served them, whether the skill
|
|
272
284
|
was invoked, or whether they were the ablated arm. *Fix:* read the three fields the record already
|
|
273
|
-
carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all
|
|
285
|
+
carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all;
|
|
286
|
+
a `/<skill> …` prompt runs the skill with NO `Skill` call, so read `slashInvokedSkills` too — and
|
|
287
|
+
`models` is then just `["<synthetic>"]`, with the real model only in `modelUsage`),
|
|
274
288
|
`models` (which model), `ablated` + `context.availableSkills` (which arm). An answer that reads
|
|
275
289
|
exactly like skill output is not evidence: the skill's own source is mounted where the model can
|
|
276
290
|
read it — in production too — so on a self-referential prompt it may read `SKILL.md` and answer
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Measurement
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
|
|
4
4
|
|
|
5
5
|
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
6
6
|
|
|
@@ -25,6 +25,11 @@ Every ablated run is stamped `ablated: true` in `result.json` and carries `ablat
|
|
|
25
25
|
What the harness gives you here is the run execution and the control arm — designing the comparison
|
|
26
26
|
(scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
|
|
27
27
|
|
|
28
|
+
**"Did my edit change it?"** → `cowork-harness eval` (EXPERIMENTAL): the version before your edit and the
|
|
29
|
+
one after, interleaved, with the agent and judge models pinned, compared per assertion and per rubric
|
|
30
|
+
claim with an exact test. It is a regression signal to investigate, not proof — see
|
|
31
|
+
[`eval.md`](eval.md) and Recipe 5 step 6 in [`task-recipes.md`](task-recipes.md).
|
|
32
|
+
|
|
28
33
|
### Tool timing — what `toolDurations` measures
|
|
29
34
|
|
|
30
35
|
`result.json`'s `toolDurations` and `trace <run> --view tool-durations` report, per tool, the **wall gap
|
|
@@ -40,7 +45,10 @@ Compare timings between runs of the same tier and model only.
|
|
|
40
45
|
|
|
41
46
|
1. **Pin the model in the session.** A run that resolves no model is refused, but one pinned only by
|
|
42
47
|
`COWORK_HARNESS_MODEL` takes its model from the machine, so two shells can run two models. Set
|
|
43
|
-
`model:` in the session (or pass the same `--model` on every `skill` run).
|
|
48
|
+
`model:` in the session (or pass the same `--model` on every `skill` run). Adding `model:` to a session
|
|
49
|
+
that already has cassettes re-stales them (`verify-cassettes` exits 1 — the model is in the session
|
|
50
|
+
fingerprint); to pin without re-recording now, use `--model` or `COWORK_HARNESS_MODEL` and move the
|
|
51
|
+
pin into the session at the next re-record. Read `result.json`'s `models` back before believing any cross-run comparison — and when
|
|
44
52
|
you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
|
|
45
53
|
fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
|
|
46
54
|
array purely by whether such a turn occurred.
|