cowork-harness 4.0.0 → 4.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +11 -14
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +3 -3
- package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
- package/.claude/skills/cowork-harness/references/authoring.md +64 -14
- package/.claude/skills/cowork-harness/references/ci-recipe.md +41 -11
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/debugging.md +16 -3
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +5 -3
- package/.claude/skills/cowork-harness/references/gotchas.md +16 -11
- package/.claude/skills/cowork-harness/references/measurement.md +5 -2
- package/.claude/skills/cowork-harness/references/run-record-replay.md +18 -6
- package/.claude/skills/cowork-harness/references/scenario-schema.md +19 -6
- package/.claude/skills/cowork-harness/references/task-recipes.md +9 -4
- package/.claude/skills/cowork-harness/scripts/scenario.py +638 -52
- package/.cowork-redact.json +3 -3
- package/CHANGELOG.md +234 -3
- package/README.md +3 -3
- package/RELEASING.md +4 -0
- package/SPEC.md +32 -11
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +90 -0
- package/dist/baseline.js +32 -1
- package/dist/cli.js +49 -14
- package/dist/errors.js +41 -0
- package/dist/redactable-literal.js +2 -0
- package/dist/run/cassette.js +111 -35
- package/dist/run/chat.js +6 -2
- package/dist/run/doctor.js +49 -11
- package/dist/run/execute.js +255 -68
- package/dist/run/host-path-tokens.js +73 -0
- package/dist/run/input-host-paths.js +138 -0
- package/dist/run/lint-load.js +12 -5
- package/dist/run/scenario-tool.js +8 -3
- package/dist/run/verdict.js +5 -0
- package/dist/runtime/container.js +4 -0
- package/dist/runtime/image-capabilities.js +21 -2
- package/dist/runtime/lima.js +122 -3
- package/dist/runtime/microvm.js +6 -0
- package/dist/scan.js +50 -2
- package/dist/session.js +5 -0
- package/dist/types.js +1 -1
- package/docs/cassette.md +7 -1
- package/docs/ci.md +13 -1
- package/docs/cli.md +19 -14
- package/docs/companion-skill.md +2 -2
- package/docs/debugging.md +5 -0
- package/docs/gotchas.md +5 -4
- package/docs/plugin-root.md +69 -29
- package/docs/scenario.md +40 -9
- package/docs/session.md +1 -1
- package/docs/subagents.md +3 -3
- package/examples/replays/README.md +1 -1
- package/examples/scenarios/csv-fx-normalize.yaml +1 -1
- package/examples/scenarios/csv-metrics.yaml +1 -1
- package/examples/skills/csv-fx-normalize/skills/csv-fx-normalize/SKILL.md +8 -1
- package/examples/skills/csv-metrics/skills/csv-metrics/SKILL.md +8 -1
- package/package.json +2 -2
- package/python/test_scenario_lint.py +244 -0
- package/schema/run-result.json +10 -0
- package/schema/scenario.schema.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 4.
|
|
7
|
-
tracks-harness: cowork-harness 4.
|
|
6
|
+
version: 4.1.1
|
|
7
|
+
tracks-harness: cowork-harness 4.1.1 (baseline desktop-2.9939.4)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
|
|
|
26
26
|
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
27
|
Read them.
|
|
28
28
|
|
|
29
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.1.1` (baseline
|
|
30
30
|
> `desktop-2.9939.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
31
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
32
32
|
|
|
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
43
43
|
|
|
44
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
45
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
46
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.1.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.1.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.1.1"`. **Pin `@^4.1.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
47
47
|
|
|
48
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
49
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -67,14 +67,11 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
|
|
|
67
67
|
no scenario file.
|
|
68
68
|
- **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
|
|
69
69
|
This is the CI-grade path and most of this skill.
|
|
70
|
-
- **A run failed — or greened and you don't trust it** (the debugging loop) →
|
|
71
|
-
|
|
72
|
-
the
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
|
|
76
|
-
output* sections in [`references/debugging.md`](references/debugging.md) are the detail (the fuller human-facing map lives in
|
|
77
|
-
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only, not shipped with the installed skill).
|
|
70
|
+
- **A run failed — or greened and you don't trust it** (the debugging loop) → **Read
|
|
71
|
+
[`references/debugging.md`](references/debugging.md) before touching the run.** Its *Triage* table first
|
|
72
|
+
splits "the skill misbehaved" from "a green you don't trust" — they need different tools — then names the
|
|
73
|
+
tool for each, all token-free over the **kept run dir** (`~/.cowork-harness/runs/…`; `--keep` prints the
|
|
74
|
+
path, `trace <run-id>` finds it). Don't re-run and hope.
|
|
78
75
|
**"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
|
|
79
76
|
grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
|
|
80
77
|
surface; see `references/critique.md`.
|
|
@@ -136,8 +133,8 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
|
|
|
136
133
|
(a cassette cannot be moved) → `replay` on the PR gate.
|
|
137
134
|
- **Fix answers without paying:** `--keep` one run → `trace <run-dir> --view questions` → edit
|
|
138
135
|
`answers:` → `verify-run <run-dir> <scenario.yaml>` → record once.
|
|
139
|
-
- **Debug:**
|
|
140
|
-
`verify-run`, `diff
|
|
136
|
+
- **Debug:** Read `references/debugging.md` first — its triage table picks the tool (`inspect`,
|
|
137
|
+
`trace --view …`, `verify-run`, `diff`, `replay --explain`).
|
|
141
138
|
- **Measure:** `--repeat N` for flakiness. `--ablate-skill` runs the control arm only; run the treatment
|
|
142
139
|
arm yourself, with the model pinned and a recoverable source frozen.
|
|
143
140
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertion catalog
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Every `assert:` key with its semantics, and the
|
|
4
4
|
verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
|
|
5
5
|
the scenario and session YAML fields are there too.
|
|
6
6
|
|
|
@@ -102,7 +102,7 @@ same set live from the schema.
|
|
|
102
102
|
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
103
103
|
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Silences `outputs_delete`, `outputs_delete_unconfirmed` and `outputs_diff_unavailable`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
104
104
|
| `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
|
|
105
|
-
| `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
|
|
105
|
+
| `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text (a path that came verbatim from the scenario's input files or prompt is exempt; `scan.hostPathsFromInputs` counts such paths) — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
|
|
106
106
|
| `egress_denied: <host>` | the host was blocked by the egress proxy |
|
|
107
107
|
| `egress_allowed: <host>` | the host was allowed through |
|
|
108
108
|
| `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
|
|
@@ -140,7 +140,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
140
140
|
| `outputs_delete_unconfirmed` | warn | A delete-shaped command near `mnt/outputs` nothing confirms: no output present at turn start was deleted and no flagged delete has an outputs path as its own operand (e.g. a Python variable named `rm` next to an outputs path, quoted prose, a `sed`/`grep` pattern). A *statement* is a fragment split on newline/`;`/`&&`/`\|\|`, quote-blind, so some real deletes also land here — a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`), a `cd` then a relative path, chained variables (`A=…; B=$A/x; rm "$B"`), a Python path held in a variable set on another line (`p = …` then `os.remove(p)`, or `for p in …:` then `p.unlink()`), wrappers with flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`), and calls outside the modelled set such as Node's `fs.promises.rm(…)` — read the command. Raised even when `no_delete_in_outputs` is authored (that assertion passes on it). Kept false positives that still fail `outputs_delete`: quoted text in which a delete command with an outputs operand follows a shell separator, subshell or keyword — the classifier does not track quotes (`echo 'note; rm mnt/outputs/x'`, `echo "a & rm …/outputs/x"`), and a heredoc that *writes* a script rather than running it (`cat <<EOF > clean.sh` with an `rm …/outputs/x` line). Waive: `allow_outputs_delete` |
|
|
141
141
|
| `outputs_diff_unavailable` | warn | The per-turn filesystem diff of `outputs/` could not verify this turn and the text scan flagged nothing — a delete made without a bash command would have gone undetected. With a text hit the turn fails `outputs_delete` instead |
|
|
142
142
|
| `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
|
|
143
|
-
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
143
|
+
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`). A path copied verbatim from the scenario's own uploads, connected folders or prompt is not a leak |
|
|
144
144
|
| `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
|
|
145
145
|
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
146
146
|
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertions guide
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
|
|
4
4
|
|
|
5
5
|
### Assertions: two orthogonal axes
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Authoring a scenario
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
|
|
4
4
|
|
|
5
5
|
## Part I — AUTHOR a scenario
|
|
6
6
|
|
|
@@ -220,25 +220,75 @@ so a file `run`/`record` would refuse fails lint too (running `scenario.py lint`
|
|
|
220
220
|
cowork-harness lint scenarios/*.yaml
|
|
221
221
|
```
|
|
222
222
|
|
|
223
|
-
`lint`
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
`
|
|
228
|
-
|
|
229
|
-
`
|
|
230
|
-
|
|
231
|
-
`
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
223
|
+
`lint` exits non-zero on any ERROR (CI-friendly); `--strict` also fails on WARN. Every rule it reports:
|
|
224
|
+
|
|
225
|
+
| Rule | Severity | Fires on |
|
|
226
|
+
|---|---|---|
|
|
227
|
+
| `assert-contradiction` | ERROR | assert items no single run can satisfy together |
|
|
228
|
+
| `assertions-key` | ERROR | `assertions:` instead of `assert:` — none of the checks would run |
|
|
229
|
+
| `authored-replay-fidelity` | ERROR | an authored `replay_protocol_fidelity` (only the replay lane synthesizes it) |
|
|
230
|
+
| `capabilities-on-protocol` | ERROR | non-empty `requires_capabilities` on `protocol` without `allow_missing_capability` — the probe cannot run there, so the run fails as unverifiable |
|
|
231
|
+
| `cassette-evidence-skipped` | INFO | with `--cassette-dir`: a cassette (or the directory) could not be read, so it cannot quiet replay-evidence advice |
|
|
232
|
+
| `container-only-key-off-container` | ERROR | `no_scratchpad_leak` off `container` — hostloop's `present_files` never promotes, so there is nothing to leak (WARN on `cowork`, whose tier resolves per the baseline gate) |
|
|
233
|
+
| `egress-on-protocol` | ERROR | an egress assertion (`egress_*` / `expect_denied`) on `protocol`, which enforces no egress |
|
|
234
|
+
| `enum-value-invalid` | ERROR | a field value outside its allowed set |
|
|
235
|
+
| `fidelity-missing` | ERROR | no `fidelity:` (required since 4.0.0) |
|
|
236
|
+
| `file-absent-contradiction` | ERROR | one path under both `file_exists` and `file_absent` |
|
|
237
|
+
| `gate-needs-controlout` | INFO | gate assertions, which evaluate on replay only when the cassette has `controlOut` |
|
|
238
|
+
| `host-path-assert-cowork` | WARN | `transcript_no_host_path` on `cowork` — it fails by design if the tier resolves to hostloop |
|
|
239
|
+
| `host-path-assert-tier` | ERROR | `transcript_no_host_path` on `hostloop` / `protocol`, where it fails by design |
|
|
240
|
+
| `lane-remote-incompatible-key` | ERROR | `present_files_called` / `no_scratchpad_leak` / `user_visible_artifact` on `lane: remote` (the runtime rejects them at load, so the tier rules are suppressed there) |
|
|
241
|
+
| `linter-extra-findings-invalid` | ERROR | the loader findings `cowork-harness lint` hands the linter could not be read |
|
|
242
|
+
| `linter-unclassified-key` | ERROR | a valid assertion key this linter cannot classify (the linter is out of date) |
|
|
243
|
+
| `manifest-needs-snapshot` | INFO | manifest-backed keys, which evaluate on replay only when the cassette carries an `artifacts` manifest |
|
|
244
|
+
| `mixed-assert-item` | WARN | one assert item mixing replay-checkable and live-only keys (replay drops the live-only half) |
|
|
245
|
+
| `no-scenarios` | ERROR | a linted directory with no `*.yaml` / `*.yml` |
|
|
246
|
+
| `not-found` | ERROR | a named file that does not exist |
|
|
247
|
+
| `parse` | ERROR | a file that is not YAML, or not a mapping |
|
|
248
|
+
| `positional-choose-order` | INFO | an answer rule with a positional `choose` (first / index), which option re-ordering can move |
|
|
249
|
+
| `present-files-key-off-tier` | ERROR | `present_files_called` on `protocol` / `microvm` (served only at `container` / `hostloop`) |
|
|
250
|
+
| `prompt-slash-not-leading` | WARN | a `prompt:` that names `/<skill>` without starting with it, so it is never expanded |
|
|
251
|
+
| `reference-access-contradiction` | ERROR | one reference under both `reference_read` and `no_observed_reference_access` |
|
|
252
|
+
| `regex-double-quoted` | WARN | a double-quoted regex with an unescaped backslash (YAML strips it) |
|
|
253
|
+
| `replay-noop` | WARN | every assertion is live-only or a verdict modifier, so a replay gate verifies nothing |
|
|
254
|
+
| `tool-called-always-passes` | INFO | `tool_called` with `count: {min: 0}` and no `max` — it asserts nothing |
|
|
255
|
+
| `tool-input-regex-redactable` | WARN | a `tool_not_called` input literal the redaction policy rewrites in the committed cassette (or a policy pattern it cannot check offline) |
|
|
256
|
+
| `tool-input-shell-tier` | INFO | the object form with `tool: Bash` and a `command` on `hostloop` / `cowork`, where shell runs as `mcp__workspace__bash` — list both |
|
|
257
|
+
| `tool-not-called-tier-vacuous` | WARN | `tool_not_called` / `subagent_tool_absent` naming a tool the tier never serves |
|
|
258
|
+
| `transcript-command-shaped` | WARN | a `transcript_*` value shaped like a shell command — those keys read prose only, never a tool call |
|
|
259
|
+
| `unknown-assert-key` | WARN | an assertion key not in the catalog (the loader rejects it) |
|
|
260
|
+
| `unknown-top-key` | WARN | a scenario key not in the schema |
|
|
261
|
+
| `vacuous-gate-assert` | WARN | `gate_answers_delivered` with no presence companion (zero gates passes it), or inert beside `questions_count_max: 0` |
|
|
262
|
+
| `scenario-invalid` | ERROR | the harness's scenario loader refuses the file (via `cowork-harness lint` only — see below) |
|
|
263
|
+
| `baseline-unknown` | ERROR | `baseline:` names no baseline this CLI ships (via `cowork-harness lint` only) |
|
|
264
|
+
| `lint-loader-internal` | ERROR | the wrapper could not run its loader check on a file — a harness bug; it never falls back to a lint that skipped the loader |
|
|
265
|
+
|
|
266
|
+
`scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never emits a scenario `lint`
|
|
267
|
+
would reject.
|
|
268
|
+
|
|
269
|
+
**Lint the skill itself: `cowork-harness lint-skill <skill-dir>`.** It checks the skill, not a scenario:
|
|
270
|
+
Cowork host-loop footguns (`${CLAUDE_PLUGIN_ROOT}` in a VM bash step, hook events, a misplaced
|
|
271
|
+
`hooks.json`, an unresolvable `subagent_type`), the evidence corpus a `critique` can package
|
|
272
|
+
(`references/critique.md`), and two size caps. `skill-body-over-reattach-cap` (WARN) fires when the
|
|
273
|
+
`SKILL.md` body, frontmatter excluded, passes 19,000 B — after a compaction the agent re-attaches only the
|
|
274
|
+
start of an invoked skill — and `skill-body-near-reattach-cap` (INFO) from 80% of that;
|
|
275
|
+
`skill-reference-over-read-cap` (WARN) fires on a `references/**.md` over 60,000 B, past which a
|
|
276
|
+
whole-file Read returns a partial view. `--strict` fails on WARN, never on INFO. To accept a reviewed
|
|
277
|
+
judgement-call finding, pass `--ignore-rule <rule>[=<glob>]` (repeatable; the glob matches the finding's
|
|
278
|
+
file) or fence the text in `SKILL.md` with `<!-- lint-skill: ignore-start <rule>[,<rule>…]: <reason> -->`
|
|
279
|
+
… `<!-- lint-skill: ignore-end -->` (outside any code fence). A suppressed finding is still printed; it
|
|
280
|
+
stops gating. A provable rule (an ERROR, a misplaced `hooks.json`, a missing pinned agent) cannot be
|
|
281
|
+
suppressed: naming it, or an unknown rule, in `--ignore-rule` is a usage error (exit 2); in a marker it is
|
|
282
|
+
WARN `lint-skill-ignore-invalid`, as is any other malformed marker. An unclosed marker is WARN
|
|
283
|
+
`lint-skill-ignore-unclosed`, and one that suppresses nothing is INFO `lint-skill-ignore-unused`.
|
|
235
284
|
|
|
236
285
|
**`cowork-harness lint` runs the loader: a file it calls clean is one `run`/`record` will load.** Anything
|
|
237
286
|
the loader refuses — an unknown key, a wrong value type (a scalar `semantic_matches.rubric`), a bad regex,
|
|
238
287
|
a reserved value — is ✗ ERROR `scenario-invalid` (exit 1, with or without `--strict`), and a `baseline:`
|
|
239
288
|
naming no baseline this installed CLI ships is ✗ ERROR `baseline-unknown` (`latest` always resolves). It
|
|
240
289
|
does not check what depends on the machine the run happens on (the session file and its mounts, an
|
|
241
|
-
absolute `baseline:` path
|
|
290
|
+
absolute `baseline:` path that does not exist here — one that exists is checked (4.1.1 and later) —
|
|
291
|
+
environment variables). A session or matrix YAML in a linted directory is not
|
|
242
292
|
a scenario and is reported as one that does not load — keep those out of the linted set. `python3
|
|
243
293
|
scenario.py lint` run directly stays offline and lenient: there an unknown key is only a ⚠ WARN (exit 0).
|
|
244
294
|
`cowork-harness record <file.yaml> --dry-run` also runs the loader and adds the pre-spend refusals (exit 2
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "4.
|
|
20
|
+
(e.g. `version: "4.1.1"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
82
82
|
GitHub-hosted runners, no token/Docker/agent:
|
|
83
83
|
|
|
84
84
|
```yaml
|
|
85
|
-
- run: npm i -g "cowork-harness@^4.
|
|
85
|
+
- run: npm i -g "cowork-harness@^4.1.1"
|
|
86
86
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
87
87
|
# no silent false-greens. WITHOUT --strict this
|
|
88
88
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -93,7 +93,16 @@ GitHub-hosted runners, no token/Docker/agent:
|
|
|
93
93
|
# 3.x CLI bare `--strict` also fails on INFO, so keep
|
|
94
94
|
# the pair explicit. `--min-severity INFO` gates on
|
|
95
95
|
# INFO as well. (`lint-skill --strict` never fails
|
|
96
|
-
# on INFO and has no floor to widen
|
|
96
|
+
# on INFO and has no floor to widen; to accept a
|
|
97
|
+
# reviewed lint-skill WARN, keep --strict and add
|
|
98
|
+
# `--ignore-rule <rule>[=<skill-dir>/<path>]` or an
|
|
99
|
+
# ignore-start/ignore-end marker in the SKILL.md.)
|
|
100
|
+
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity INFO --cassette-dir cassettes/
|
|
101
|
+
# committed *.cassette.json evidence suppresses only
|
|
102
|
+
# replay advisories proven by every matching cassette;
|
|
103
|
+
# skipped or malformed cassettes stay visible.
|
|
104
|
+
# Staleness does not affect this existence check;
|
|
105
|
+
# verify-cassettes checks freshness and drift.
|
|
97
106
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness — FAILS on a stale recording
|
|
98
107
|
# ALSO fails on a leaked host inventory: recording at
|
|
99
108
|
# protocol/hostloop freezes YOUR machine's MCP servers,
|
|
@@ -232,7 +241,9 @@ cowork-harness replay cassettes/ # replay every *.cassette.jso
|
|
|
232
241
|
Re-record whenever the protocol or your scenario's expected content changes. An old cassette without
|
|
233
242
|
`controlOut` excludes the gate keys (with a loud warning) — re-record to enable them. `record` **refuses
|
|
234
243
|
to freeze a failing live run** into a cassette (pass `--allow-failing` to override) — a committed red
|
|
235
|
-
cassette is a latent false-signal.
|
|
244
|
+
cassette is a latent false-signal. An `--allow-failing` recording of a red run exits 0 with `ok: true`:
|
|
245
|
+
`record`'s `ok` means a cassette was written, and the run's own verdict is `results[0].verdict.pass`
|
|
246
|
+
(`items[].verdict` on `record <dir/>`).
|
|
236
247
|
|
|
237
248
|
## Privacy: cassettes are committed fixtures → record only against SYNTHETIC inputs
|
|
238
249
|
|
|
@@ -247,7 +258,9 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
247
258
|
dir** (first file found per dir; env vars merge on top). `cowork-harness init-redact` copies the
|
|
248
259
|
packaged reference template (local-path prefixes, incl. macOS temp roots and slugged home segments
|
|
249
260
|
like `-Users-<user>-…`, + a generic email regex) into the cwd as a starting point — review and tailor
|
|
250
|
-
it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force`
|
|
261
|
+
it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` (without it
|
|
262
|
+
`init-redact` refuses to overwrite an existing policy; with it your tailoring is replaced — save and
|
|
263
|
+
re-apply it) or add them by hand. Redaction is **verdict-preserving** — `record` refuses to write if
|
|
251
264
|
redaction would flip an assertion (a manufactured green). `--no-redact` skips it for known-synthetic
|
|
252
265
|
inputs.
|
|
253
266
|
- **Pre-spawn preflight**: `record` warns (`::warning::`, before the paid run starts — once per batch
|
|
@@ -332,7 +345,16 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
332
345
|
|
|
333
346
|
This is the shape a CI step wants: **silent on success (no output, exit 0), loud and specific on
|
|
334
347
|
failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
|
|
335
|
-
offending file *and* the rejected key, one line per file, and the step still exits 1.
|
|
348
|
+
offending file *and* the rejected key, one line per file, and the step still exits 1. It exits 0,
|
|
349
|
+
though, on an input the real record would refuse (a `session:` file that cannot be read — 4.1.1 and
|
|
350
|
+
later — a missing path, an unknown baseline name, a tier-vacuous `tool_not_called`): that prints a `⚠ input error:` line and lands in `inputErrors[]`. To
|
|
351
|
+
gate on those too, use the JSON form (4.1.0 and later):
|
|
352
|
+
|
|
353
|
+
```bash
|
|
354
|
+
cowork-harness record scenarios/ --dry-run --output-format json | jq -e '.ok and (.inputErrors == [])'
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
Point `lint`
|
|
336
358
|
at scenarios only: a session or matrix YAML in the linted set is reported as a file that does not
|
|
337
359
|
load.
|
|
338
360
|
|
|
@@ -373,7 +395,7 @@ jobs:
|
|
|
373
395
|
with: { node-version: '24' }
|
|
374
396
|
- uses: actions/setup-python@v5
|
|
375
397
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
376
|
-
- run: npm i -g "cowork-harness@^4.
|
|
398
|
+
- run: npm i -g "cowork-harness@^4.1.1"
|
|
377
399
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
378
400
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
379
401
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -402,7 +424,7 @@ jobs:
|
|
|
402
424
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
403
425
|
fi
|
|
404
426
|
- if: steps.guard.outputs.live == 'true'
|
|
405
|
-
run: npm i -g "cowork-harness@^4.
|
|
427
|
+
run: npm i -g "cowork-harness@^4.1.1"
|
|
406
428
|
- if: steps.guard.outputs.live == 'true'
|
|
407
429
|
run: cowork-harness run scenarios/ --output-format json
|
|
408
430
|
env:
|
|
@@ -425,7 +447,8 @@ sandbox).
|
|
|
425
447
|
|
|
426
448
|
`--output-format json` emits a machine envelope on stdout (human output goes to stderr):
|
|
427
449
|
`{tool, version, command, ok, results[], error}` — one `RunResult` per scenario. **Overall pass for a
|
|
428
|
-
scenario is `verdict.pass`** (envelope-wide: `ok`
|
|
450
|
+
scenario is `verdict.pass`** (envelope-wide: `ok`; on `record`, `ok` only means the command exited 0 —
|
|
451
|
+
see below), and it is strictly stronger than
|
|
429
452
|
`result === "success" && assertions.every(pass)`: the verdict also carries ~20 signal codes that fail a run
|
|
430
453
|
with no failing assertion at all — `stalled`, `outputs_delete`, `mount_delete`, `host_path_leak`,
|
|
431
454
|
`undelivered_deliverables`, `missing_capability`, `permissive_auto_allow`, `ended_with_question`,
|
|
@@ -459,6 +482,11 @@ Keep the `.results[]?` hop and the `?` operators. `.results[0]` silently ignores
|
|
|
459
482
|
the first when you pass a directory, and a bare `.verdict` does not exist at the envelope root at all —
|
|
460
483
|
both read as "no failures" against a run that failed.
|
|
461
484
|
|
|
485
|
+
**`record` is shaped differently.** Its `ok` is the exit code's verdict (a cassette was written), not the
|
|
486
|
+
run's: `record --allow-failing` on a red run is `ok: true`. The verdict is in `results[0].verdict` on
|
|
487
|
+
`record <file>` and in `items[].verdict` on `record <dir/>` and `--rerecord-stale`, which carry no
|
|
488
|
+
`results[]` — so `.results[]?` reads nothing there. Use `[.items[]? | .verdict.failures[]?]` on a batch.
|
|
489
|
+
|
|
462
490
|
> **Do not filter on whether `assertion` is present.** That was the only discriminator before `kind`
|
|
463
491
|
> existed and it never worked in both directions: `coverage` entries carry a key too (an internal
|
|
464
492
|
> `answer_coverage` marker), so they read as authored asserts, while `guard`, `staleness` and
|
|
@@ -468,7 +496,9 @@ both read as "no failures" against a run that failed.
|
|
|
468
496
|
`--output-format json`, `run` / `record` / `replay` / `verify-cassettes` / `status` write their whole
|
|
469
497
|
human rendering — warnings, verdict, `status`'s summary line — to **stderr**, and stdout stays empty.
|
|
470
498
|
A wrapper that captures only stdout gets an empty log and, if it greps that for a state, a silent false
|
|
471
|
-
negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout
|
|
499
|
+
negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout —
|
|
500
|
+
`COWORK_HARNESS_OUTPUT_FORMAT=json` makes JSON the default for every command that takes `--output-format`
|
|
501
|
+
(an explicit flag still wins).
|
|
472
502
|
(Commands whose whole job is to print a value — `--version`, `assertions --list`, `scaffold`, `gates`,
|
|
473
503
|
`skill --dry-run` — write it to stdout by design. Under `--output-format json` the value rides inside the
|
|
474
504
|
envelope: `scaffold`'s YAML is `.scenario`, `skill --dry-run`'s preview is the envelope's own fields;
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,13 +1,15 @@
|
|
|
1
1
|
# Debugging a run
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
|
|
4
4
|
|
|
5
5
|
## Part III — Debug
|
|
6
6
|
|
|
7
7
|
A run misbehaved, or greened when you don't trust it. Debugging is a first-class loop, not an
|
|
8
8
|
afterthought: the run already wrote its evidence, so you **localize the failure post-hoc** rather than
|
|
9
9
|
re-run and hope. Start at the triage below, then use the observability output and, when you need to
|
|
10
|
-
reproduce interactively, `chat`.
|
|
10
|
+
reproduce interactively, `chat`. (The fuller human-facing map is
|
|
11
|
+
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only,
|
|
12
|
+
not shipped with the installed skill.)
|
|
11
13
|
|
|
12
14
|
> **"Evidence" below means the run's own record** — events, trace, transcript; what `trace` / `inspect` /
|
|
13
15
|
> `diff` / `verify-run` / `replay --explain` read. `critique`'s **evaluator** grades against a separate,
|
|
@@ -32,6 +34,15 @@ agent stderr) — read those before re-running; a re-record rarely tells you mor
|
|
|
32
34
|
already does.
|
|
33
35
|
<!-- END triage-canonical -->
|
|
34
36
|
|
|
37
|
+
**microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
|
|
38
|
+
stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
|
|
39
|
+
`cowork-harness vm status` — a `provisioning` other than `ready` confirms it — and if a run does not
|
|
40
|
+
recover it on its own, `cowork-harness vm delete` and retry. A run on such a VM can also have cached an
|
|
41
|
+
empty toolchain for it in the capability probe's cache: `vm delete` drops that VM's entry (`vm prune`
|
|
42
|
+
forgets only the orphaned VMs it deletes, never the current one); otherwise delete `capability-cache.json`
|
|
43
|
+
from the runs root (`~/.cowork-harness/runs/` unless
|
|
44
|
+
`COWORK_HARNESS_RUNS_DIR` is set) so the next run probes again.
|
|
45
|
+
|
|
35
46
|
**Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
|
|
36
47
|
**deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
|
|
37
48
|
so a "bug" you see here that real Cowork also has isn't yours to fix. The tier semantics are in
|
|
@@ -73,7 +84,9 @@ decide which assertions from *Assertions: two orthogonal axes* in `assertions-gu
|
|
|
73
84
|
walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
|
|
74
85
|
`total_cost_usd` for the run — the authoritative single-run spend of the agent session, which leaves out the `semantic_matches` judge and the LLM decider calls; NOT the same source as summing
|
|
75
86
|
`modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
|
|
76
|
-
`usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`,
|
|
87
|
+
`usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`,
|
|
88
|
+
`toolCalls` (every tool call in stream order: `name`, top-level `input` fields each capped at 10 KB, and
|
|
89
|
+
`origin` `main`/`subagent`/`unknown` — what the object form of `tool_called`/`tool_not_called` reads), `toolErrors`,
|
|
77
90
|
`redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
|
|
78
91
|
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
|
|
79
92
|
`context` (tools/mcpServers/availableSkills), `tasks`,
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -73,7 +73,9 @@ The instance is `cowork-vm-<config-hash>` — a config or agent-version change y
|
|
|
73
73
|
stale VM is never silently reused (the old one is orphaned until `vm prune`). Pin a fixed name with
|
|
74
74
|
`COWORK_LIMA_INSTANCE`. The optional argument to every `vm` subcommand is a **baseline**
|
|
75
75
|
(`desktop-<version>`, default `latest`), never the `cowork-vm-<hash>` VM name — passing a VM name is a
|
|
76
|
-
usage error that names the baseline(s) deriving it.
|
|
76
|
+
usage error that names the baseline(s) deriving it. A Running VM is used only once its provisioning has
|
|
77
|
+
finished (a run waits up to `COWORK_VM_PROVISION_TIMEOUT_S`, default 900 s); `microvm <instance> never
|
|
78
|
+
finished provisioning (…)` means it will not — `cowork-harness vm delete` and retry.
|
|
77
79
|
|
|
78
80
|
## Answer paths (resolving gates: AskUserQuestion + tool-permission)
|
|
79
81
|
|
|
@@ -325,7 +327,7 @@ up often enough to spell out:
|
|
|
325
327
|
- `COWORK_HARNESS_SCRUB_KEYS` / `COWORK_HARNESS_SCRUB_VALUES` — extra env names / literal values to
|
|
326
328
|
redact from logs (beyond the auth tokens + `ANTHROPIC_CUSTOM_HEADERS`).
|
|
327
329
|
- `COWORK_HARNESS_SOFT_MISSING` — downgrade a missing mount source from hard-error to warn-and-skip.
|
|
328
|
-
- `COWORK_VM_GATEWAY` / `COWORK_VM_PROXY_PORT` / `COWORK_LIMA_INSTANCE` — L2 (microVM) knobs.
|
|
330
|
+
- `COWORK_VM_GATEWAY` / `COWORK_VM_PROXY_PORT` / `COWORK_LIMA_INSTANCE` / `COWORK_VM_PROVISION_TIMEOUT_S` — L2 (microVM) knobs.
|
|
329
331
|
- `COWORK_LOCKDOWN` — default `on`: gates sandbox hardening on every isolated tier — **aborts loudly**
|
|
330
332
|
if the L2 guest firewall fails to apply (no silent unprotected run), and also gates the
|
|
331
333
|
`container`/`hostloop` Docker hardening. Set `=off` to opt out and run without isolation deliberately.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Gotchas
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). The full "✓ passed ≠ correct" landmine catalog.
|
|
4
4
|
|
|
5
5
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
6
6
|
|
|
@@ -175,9 +175,10 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
175
175
|
`--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
|
|
176
176
|
`prompt`/`answers`/`baseline`/`fidelity`/`lane`/`skills`/`requires_capabilities` or the skill content (when a
|
|
177
177
|
fingerprint exists) drifted from the recording (re-record then), and `expect_denied`/filesystem/egress keys
|
|
178
|
-
are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the
|
|
179
|
-
|
|
180
|
-
|
|
178
|
+
are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the session's
|
|
179
|
+
`model:` IS in the cassette's `sessionFingerprint` (a `--model`/env model is not; `environment.model`
|
|
180
|
+
records what ran). `verify-cassettes` reports a changed session as staleness (exit 1); `replay` never
|
|
181
|
+
checks it — plain, `--strict` or `--assert-from` — so re-record if the session changed. `verify-run` reads
|
|
181
182
|
on-disk `assert:` against a kept *run dir*; `replay --assert-from` is the equivalent for a *cassette*.
|
|
182
183
|
|
|
183
184
|
18. **`questions_count_max` counts sub-questions, not gates.** One `AskUserQuestion` tool call can
|
|
@@ -214,13 +215,17 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
214
215
|
|
|
215
216
|
22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
|
|
216
217
|
`manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
|
|
217
|
-
assertion keys. The linter is **static
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
218
|
+
assertion keys. The linter is **static** until you opt in to cassette evidence. With committed
|
|
219
|
+
recordings, pass `lint --cassette-dir <dir>` (or one cassette file): it scans the same `*.cassette.json`
|
|
220
|
+
shape as replay and verify-cassettes, resolves each exact `scenarioSource` relative to its cassette,
|
|
221
|
+
and suppresses an advisory only when **all** matching cassettes carry the evidence the replay lane
|
|
222
|
+
needs. A malformed, wrong-shaped, unsupported-version, or provenance-less cassette is reported as INFO
|
|
223
|
+
and keeps the advisory visible — even beside a healthy sibling. The two dedicated diff assertions also
|
|
224
|
+
require their respective `preRunPaths` / `preRunHashes` baselines, and a non-empty `controlOut` /
|
|
225
|
+
artifact manifest is required where replay needs it. For a strict CI gate that keeps actionable INFO
|
|
226
|
+
rules visible while suppressing only proven replay noise, use `lint --strict --min-severity INFO
|
|
227
|
+
--cassette-dir <dir>`. Without the opt-in path, use `lint --min-severity WARN` in CI (≥1.11.0) to hide
|
|
228
|
+
the INFO class.
|
|
224
229
|
From 4.0.0 WARN is `--strict`'s default floor, so bare `lint --strict` hides and passes INFO; add
|
|
225
230
|
`--min-severity INFO` to fail on it. `--strict --min-severity ERROR` behaves as a plain lint, not a
|
|
226
231
|
contradiction.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Measurement
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
|
|
4
4
|
|
|
5
5
|
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
6
6
|
|
|
@@ -40,7 +40,10 @@ Compare timings between runs of the same tier and model only.
|
|
|
40
40
|
|
|
41
41
|
1. **Pin the model in the session.** A run that resolves no model is refused, but one pinned only by
|
|
42
42
|
`COWORK_HARNESS_MODEL` takes its model from the machine, so two shells can run two models. Set
|
|
43
|
-
`model:` in the session (or pass the same `--model` on every `skill` run).
|
|
43
|
+
`model:` in the session (or pass the same `--model` on every `skill` run). Adding `model:` to a session
|
|
44
|
+
that already has cassettes re-stales them (`verify-cassettes` exits 1 — the model is in the session
|
|
45
|
+
fingerprint); to pin without re-recording now, use `--model` or `COWORK_HARNESS_MODEL` and move the
|
|
46
|
+
pin into the session at the next re-record. Read `result.json`'s `models` back before believing any cross-run comparison — and when
|
|
44
47
|
you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
|
|
45
48
|
fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
|
|
46
49
|
array purely by whether such a turn occurred.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Run, record and lock
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.1` (baseline `desktop-2.9939.4`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
|
|
4
4
|
|
|
5
5
|
## Part II — RUN, RECORD & LOCK
|
|
6
6
|
|
|
@@ -25,12 +25,17 @@ the discovery/encode/record dance entirely and answer gates **live during the re
|
|
|
25
25
|
refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
|
|
26
26
|
cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
|
|
27
27
|
guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
|
|
28
|
-
one a paid run would give.
|
|
28
|
+
one a paid run would give. Two inputs it reports rather than refuses, at exit 0 under `inputErrors[]` with a
|
|
29
|
+
`⚠ input error:` line: a `session:` file that cannot be read (missing, a directory, not valid YAML — 4.1.1
|
|
30
|
+
and later) and a `tool_not_called` the tier can never violate — `run` and the real `record` refuse both. On a **directory** the path-dependent verdicts (host-inventory, cassette
|
|
29
31
|
portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
|
|
30
32
|
the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
|
|
31
33
|
takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
|
|
32
|
-
contradiction, duplicate cassette target) gate the batch.
|
|
33
|
-
|
|
34
|
+
contradiction, duplicate cassette target) gate the batch. An input the real record would refuse — a
|
|
35
|
+
`session:` file that cannot be read (4.1.1 and later), a missing path, an unknown baseline name, a `tool_not_called` the tier can never violate — is listed under
|
|
36
|
+
`inputErrors[]` with a `⚠ input error:` line, also at exit 0. So a directory dry-run CAN exit 0 on a scenario the
|
|
37
|
+
real `record` would refuse; re-run that one file with its real flags for a binding answer, or gate on
|
|
38
|
+
`.ok and (.inputErrors == [])` in the JSON payload (4.1.0 and later). A directory also
|
|
34
39
|
reports every offender and the batch cost estimate. `lint` checks the assertion invariants AND that each file loads (the same loader, plus a named `baseline:`), but not the pre-spend refusals.
|
|
35
40
|
|
|
36
41
|
**Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
|
|
@@ -65,7 +70,10 @@ finished: a failing verdict without `--allow-failing`, an assert on an artifact
|
|
|
65
70
|
quarantined inventory finding, or any other error before the cassette is written. Only the after-the-run
|
|
66
71
|
refusals report the run: under `--output-format json` they carry it in `results[0]` (verdict and
|
|
67
72
|
cost) with `error.category: "runtime"`; every pre-spend refusal, and a run that throws before returning a
|
|
68
|
-
result (an unanswered gate), has `results: []`.
|
|
73
|
+
result (an unanswered gate), has `results: []`. The envelope's `ok` is the exit code (`ok` ⇔ exit 0, a
|
|
74
|
+
cassette was written), not the verdict: an `--allow-failing` recording of a red run is `ok: true` with
|
|
75
|
+
`results[0].verdict.pass: false`. `record <dir/>` reports per scenario under `items[]` (each with `status`,
|
|
76
|
+
and `verdict`/`result` once its run finished), not `results[]`.
|
|
69
77
|
|
|
70
78
|
**Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
|
|
71
79
|
Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
|
|
@@ -249,7 +257,11 @@ Recognize these before "fixing" a non-bug:
|
|
|
249
257
|
`container`/`microvm`, but only *fires* on an actual scanned leak with no authored
|
|
250
258
|
`transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
|
|
251
259
|
run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
|
|
252
|
-
it's valid.
|
|
260
|
+
it's valid. A host path that came verbatim from the scenario's own uploads, connected folders or prompt
|
|
261
|
+
is not a leak (whole-token match; a token cut short where the path goes on, or one naming a location the
|
|
262
|
+
harness created for the run, is never exempt); `scan.inputHostPathTokens` counts the host-path tokens the
|
|
263
|
+
inputs carried, `scan.hostPathsFromInputs` the ones exempted, and a non-zero exemption prints a
|
|
264
|
+
`::notice::` so a clean scan that relied on it is never silent.
|
|
253
265
|
- **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
|
|
254
266
|
infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
|
|
255
267
|
rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
|