cowork-harness 4.0.0 → 4.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +11 -14
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +3 -3
- package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
- package/.claude/skills/cowork-harness/references/authoring.md +1 -1
- package/.claude/skills/cowork-harness/references/ci-recipe.md +25 -7
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/debugging.md +9 -2
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +5 -3
- package/.claude/skills/cowork-harness/references/gotchas.md +12 -8
- package/.claude/skills/cowork-harness/references/measurement.md +1 -1
- package/.claude/skills/cowork-harness/references/run-record-replay.md +9 -4
- package/.claude/skills/cowork-harness/references/scenario-schema.md +5 -4
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/scenario.py +638 -52
- package/.cowork-redact.json +3 -3
- package/CHANGELOG.md +163 -3
- package/README.md +3 -3
- package/RELEASING.md +4 -0
- package/SPEC.md +25 -8
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +90 -0
- package/dist/cli.js +37 -7
- package/dist/redactable-literal.js +2 -0
- package/dist/run/cassette.js +56 -14
- package/dist/run/chat.js +6 -2
- package/dist/run/doctor.js +49 -11
- package/dist/run/execute.js +180 -60
- package/dist/run/host-path-tokens.js +73 -0
- package/dist/run/input-host-paths.js +138 -0
- package/dist/run/lint-load.js +7 -3
- package/dist/run/scenario-tool.js +8 -3
- package/dist/run/verdict.js +5 -0
- package/dist/runtime/container.js +4 -0
- package/dist/runtime/image-capabilities.js +21 -2
- package/dist/runtime/lima.js +122 -3
- package/dist/runtime/microvm.js +6 -0
- package/dist/scan.js +30 -2
- package/dist/types.js +1 -1
- package/docs/cassette.md +7 -1
- package/docs/ci.md +13 -1
- package/docs/cli.md +18 -13
- package/docs/companion-skill.md +2 -2
- package/docs/debugging.md +5 -0
- package/docs/gotchas.md +5 -4
- package/docs/plugin-root.md +69 -29
- package/docs/scenario.md +31 -3
- package/docs/session.md +1 -1
- package/docs/subagents.md +3 -3
- package/examples/replays/README.md +1 -1
- package/examples/scenarios/csv-fx-normalize.yaml +1 -1
- package/examples/scenarios/csv-metrics.yaml +1 -1
- package/examples/skills/csv-fx-normalize/skills/csv-fx-normalize/SKILL.md +8 -1
- package/examples/skills/csv-metrics/skills/csv-metrics/SKILL.md +8 -1
- package/package.json +2 -2
- package/python/test_scenario_lint.py +244 -0
- package/schema/run-result.json +10 -0
- package/schema/scenario.schema.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 4.
|
|
7
|
-
tracks-harness: cowork-harness 4.
|
|
6
|
+
version: 4.1.0
|
|
7
|
+
tracks-harness: cowork-harness 4.1.0 (baseline desktop-2.9939.4)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
|
|
|
26
26
|
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
27
|
Read them.
|
|
28
28
|
|
|
29
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.1.0` (baseline
|
|
30
30
|
> `desktop-2.9939.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
31
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
32
32
|
|
|
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
43
43
|
|
|
44
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
45
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
46
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.1.0"`. **Pin `@^4.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
47
47
|
|
|
48
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
49
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -67,14 +67,11 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
|
|
|
67
67
|
no scenario file.
|
|
68
68
|
- **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
|
|
69
69
|
This is the CI-grade path and most of this skill.
|
|
70
|
-
- **A run failed — or greened and you don't trust it** (the debugging loop) →
|
|
71
|
-
|
|
72
|
-
the
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
|
|
76
|
-
output* sections in [`references/debugging.md`](references/debugging.md) are the detail (the fuller human-facing map lives in
|
|
77
|
-
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only, not shipped with the installed skill).
|
|
70
|
+
- **A run failed — or greened and you don't trust it** (the debugging loop) → **Read
|
|
71
|
+
[`references/debugging.md`](references/debugging.md) before touching the run.** Its *Triage* table first
|
|
72
|
+
splits "the skill misbehaved" from "a green you don't trust" — they need different tools — then names the
|
|
73
|
+
tool for each, all token-free over the **kept run dir** (`~/.cowork-harness/runs/…`; `--keep` prints the
|
|
74
|
+
path, `trace <run-id>` finds it). Don't re-run and hope.
|
|
78
75
|
**"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
|
|
79
76
|
grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
|
|
80
77
|
surface; see `references/critique.md`.
|
|
@@ -136,8 +133,8 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
|
|
|
136
133
|
(a cassette cannot be moved) → `replay` on the PR gate.
|
|
137
134
|
- **Fix answers without paying:** `--keep` one run → `trace <run-dir> --view questions` → edit
|
|
138
135
|
`answers:` → `verify-run <run-dir> <scenario.yaml>` → record once.
|
|
139
|
-
- **Debug:**
|
|
140
|
-
`verify-run`, `diff
|
|
136
|
+
- **Debug:** Read `references/debugging.md` first — its triage table picks the tool (`inspect`,
|
|
137
|
+
`trace --view …`, `verify-run`, `diff`, `replay --explain`).
|
|
141
138
|
- **Measure:** `--repeat N` for flakiness. `--ablate-skill` runs the control arm only; run the treatment
|
|
142
139
|
arm yourself, with the model pinned and a recoverable source frozen.
|
|
143
140
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertion catalog
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Every `assert:` key with its semantics, and the
|
|
4
4
|
verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
|
|
5
5
|
the scenario and session YAML fields are there too.
|
|
6
6
|
|
|
@@ -102,7 +102,7 @@ same set live from the schema.
|
|
|
102
102
|
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
103
103
|
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Silences `outputs_delete`, `outputs_delete_unconfirmed` and `outputs_diff_unavailable`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
104
104
|
| `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
|
|
105
|
-
| `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
|
|
105
|
+
| `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text (a path that came verbatim from the scenario's input files or prompt is exempt; `scan.hostPathsFromInputs` counts such paths) — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
|
|
106
106
|
| `egress_denied: <host>` | the host was blocked by the egress proxy |
|
|
107
107
|
| `egress_allowed: <host>` | the host was allowed through |
|
|
108
108
|
| `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
|
|
@@ -140,7 +140,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
140
140
|
| `outputs_delete_unconfirmed` | warn | A delete-shaped command near `mnt/outputs` nothing confirms: no output present at turn start was deleted and no flagged delete has an outputs path as its own operand (e.g. a Python variable named `rm` next to an outputs path, quoted prose, a `sed`/`grep` pattern). A *statement* is a fragment split on newline/`;`/`&&`/`\|\|`, quote-blind, so some real deletes also land here — a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`), a `cd` then a relative path, chained variables (`A=…; B=$A/x; rm "$B"`), a Python path held in a variable set on another line (`p = …` then `os.remove(p)`, or `for p in …:` then `p.unlink()`), wrappers with flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`), and calls outside the modelled set such as Node's `fs.promises.rm(…)` — read the command. Raised even when `no_delete_in_outputs` is authored (that assertion passes on it). Kept false positives that still fail `outputs_delete`: quoted text in which a delete command with an outputs operand follows a shell separator, subshell or keyword — the classifier does not track quotes (`echo 'note; rm mnt/outputs/x'`, `echo "a & rm …/outputs/x"`), and a heredoc that *writes* a script rather than running it (`cat <<EOF > clean.sh` with an `rm …/outputs/x` line). Waive: `allow_outputs_delete` |
|
|
141
141
|
| `outputs_diff_unavailable` | warn | The per-turn filesystem diff of `outputs/` could not verify this turn and the text scan flagged nothing — a delete made without a bash command would have gone undetected. With a text hit the turn fails `outputs_delete` instead |
|
|
142
142
|
| `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
|
|
143
|
-
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
143
|
+
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`). A path copied verbatim from the scenario's own uploads, connected folders or prompt is not a leak |
|
|
144
144
|
| `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
|
|
145
145
|
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
146
146
|
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertions guide
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
|
|
4
4
|
|
|
5
5
|
### Assertions: two orthogonal axes
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Authoring a scenario
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
|
|
4
4
|
|
|
5
5
|
## Part I — AUTHOR a scenario
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "4.
|
|
20
|
+
(e.g. `version: "4.1.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
82
82
|
GitHub-hosted runners, no token/Docker/agent:
|
|
83
83
|
|
|
84
84
|
```yaml
|
|
85
|
-
- run: npm i -g "cowork-harness@^4.
|
|
85
|
+
- run: npm i -g "cowork-harness@^4.1.0"
|
|
86
86
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
87
87
|
# no silent false-greens. WITHOUT --strict this
|
|
88
88
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -93,7 +93,16 @@ GitHub-hosted runners, no token/Docker/agent:
|
|
|
93
93
|
# 3.x CLI bare `--strict` also fails on INFO, so keep
|
|
94
94
|
# the pair explicit. `--min-severity INFO` gates on
|
|
95
95
|
# INFO as well. (`lint-skill --strict` never fails
|
|
96
|
-
# on INFO and has no floor to widen
|
|
96
|
+
# on INFO and has no floor to widen; to accept a
|
|
97
|
+
# reviewed lint-skill WARN, keep --strict and add
|
|
98
|
+
# `--ignore-rule <rule>[=<skill-dir>/<path>]` or an
|
|
99
|
+
# ignore-start/ignore-end marker in the SKILL.md.)
|
|
100
|
+
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity INFO --cassette-dir cassettes/
|
|
101
|
+
# committed *.cassette.json evidence suppresses only
|
|
102
|
+
# replay advisories proven by every matching cassette;
|
|
103
|
+
# skipped or malformed cassettes stay visible.
|
|
104
|
+
# Staleness does not affect this existence check;
|
|
105
|
+
# verify-cassettes checks freshness and drift.
|
|
97
106
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness — FAILS on a stale recording
|
|
98
107
|
# ALSO fails on a leaked host inventory: recording at
|
|
99
108
|
# protocol/hostloop freezes YOUR machine's MCP servers,
|
|
@@ -332,7 +341,16 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
332
341
|
|
|
333
342
|
This is the shape a CI step wants: **silent on success (no output, exit 0), loud and specific on
|
|
334
343
|
failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
|
|
335
|
-
offending file *and* the rejected key, one line per file, and the step still exits 1.
|
|
344
|
+
offending file *and* the rejected key, one line per file, and the step still exits 1. It exits 0,
|
|
345
|
+
though, on an input the real record would refuse (a missing path, an unknown baseline name, a
|
|
346
|
+
tier-vacuous `tool_not_called`): that prints a `⚠ input error:` line and lands in `inputErrors[]`. To
|
|
347
|
+
gate on those too, use the JSON form (4.1.0 and later):
|
|
348
|
+
|
|
349
|
+
```bash
|
|
350
|
+
cowork-harness record scenarios/ --dry-run --output-format json | jq -e '.ok and (.inputErrors == [])'
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
Point `lint`
|
|
336
354
|
at scenarios only: a session or matrix YAML in the linted set is reported as a file that does not
|
|
337
355
|
load.
|
|
338
356
|
|
|
@@ -373,7 +391,7 @@ jobs:
|
|
|
373
391
|
with: { node-version: '24' }
|
|
374
392
|
- uses: actions/setup-python@v5
|
|
375
393
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
376
|
-
- run: npm i -g "cowork-harness@^4.
|
|
394
|
+
- run: npm i -g "cowork-harness@^4.1.0"
|
|
377
395
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
378
396
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
379
397
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -402,7 +420,7 @@ jobs:
|
|
|
402
420
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
403
421
|
fi
|
|
404
422
|
- if: steps.guard.outputs.live == 'true'
|
|
405
|
-
run: npm i -g "cowork-harness@^4.
|
|
423
|
+
run: npm i -g "cowork-harness@^4.1.0"
|
|
406
424
|
- if: steps.guard.outputs.live == 'true'
|
|
407
425
|
run: cowork-harness run scenarios/ --output-format json
|
|
408
426
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,13 +1,15 @@
|
|
|
1
1
|
# Debugging a run
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
|
|
4
4
|
|
|
5
5
|
## Part III — Debug
|
|
6
6
|
|
|
7
7
|
A run misbehaved, or greened when you don't trust it. Debugging is a first-class loop, not an
|
|
8
8
|
afterthought: the run already wrote its evidence, so you **localize the failure post-hoc** rather than
|
|
9
9
|
re-run and hope. Start at the triage below, then use the observability output and, when you need to
|
|
10
|
-
reproduce interactively, `chat`.
|
|
10
|
+
reproduce interactively, `chat`. (The fuller human-facing map is
|
|
11
|
+
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only,
|
|
12
|
+
not shipped with the installed skill.)
|
|
11
13
|
|
|
12
14
|
> **"Evidence" below means the run's own record** — events, trace, transcript; what `trace` / `inspect` /
|
|
13
15
|
> `diff` / `verify-run` / `replay --explain` read. `critique`'s **evaluator** grades against a separate,
|
|
@@ -32,6 +34,11 @@ agent stderr) — read those before re-running; a re-record rarely tells you mor
|
|
|
32
34
|
already does.
|
|
33
35
|
<!-- END triage-canonical -->
|
|
34
36
|
|
|
37
|
+
**microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
|
|
38
|
+
stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
|
|
39
|
+
`cowork-harness vm status` — a `provisioning` other than `ready` confirms it — and if a run does not
|
|
40
|
+
recover it on its own, `cowork-harness vm delete` and retry.
|
|
41
|
+
|
|
35
42
|
**Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
|
|
36
43
|
**deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
|
|
37
44
|
so a "bug" you see here that real Cowork also has isn't yours to fix. The tier semantics are in
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -73,7 +73,9 @@ The instance is `cowork-vm-<config-hash>` — a config or agent-version change y
|
|
|
73
73
|
stale VM is never silently reused (the old one is orphaned until `vm prune`). Pin a fixed name with
|
|
74
74
|
`COWORK_LIMA_INSTANCE`. The optional argument to every `vm` subcommand is a **baseline**
|
|
75
75
|
(`desktop-<version>`, default `latest`), never the `cowork-vm-<hash>` VM name — passing a VM name is a
|
|
76
|
-
usage error that names the baseline(s) deriving it.
|
|
76
|
+
usage error that names the baseline(s) deriving it. A Running VM is used only once its provisioning has
|
|
77
|
+
finished (a run waits up to `COWORK_VM_PROVISION_TIMEOUT_S`, default 900 s); `microvm <instance> never
|
|
78
|
+
finished provisioning (…)` means it will not — `cowork-harness vm delete` and retry.
|
|
77
79
|
|
|
78
80
|
## Answer paths (resolving gates: AskUserQuestion + tool-permission)
|
|
79
81
|
|
|
@@ -325,7 +327,7 @@ up often enough to spell out:
|
|
|
325
327
|
- `COWORK_HARNESS_SCRUB_KEYS` / `COWORK_HARNESS_SCRUB_VALUES` — extra env names / literal values to
|
|
326
328
|
redact from logs (beyond the auth tokens + `ANTHROPIC_CUSTOM_HEADERS`).
|
|
327
329
|
- `COWORK_HARNESS_SOFT_MISSING` — downgrade a missing mount source from hard-error to warn-and-skip.
|
|
328
|
-
- `COWORK_VM_GATEWAY` / `COWORK_VM_PROXY_PORT` / `COWORK_LIMA_INSTANCE` — L2 (microVM) knobs.
|
|
330
|
+
- `COWORK_VM_GATEWAY` / `COWORK_VM_PROXY_PORT` / `COWORK_LIMA_INSTANCE` / `COWORK_VM_PROVISION_TIMEOUT_S` — L2 (microVM) knobs.
|
|
329
331
|
- `COWORK_LOCKDOWN` — default `on`: gates sandbox hardening on every isolated tier — **aborts loudly**
|
|
330
332
|
if the L2 guest firewall fails to apply (no silent unprotected run), and also gates the
|
|
331
333
|
`container`/`hostloop` Docker hardening. Set `=off` to opt out and run without isolation deliberately.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Gotchas
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). The full "✓ passed ≠ correct" landmine catalog.
|
|
4
4
|
|
|
5
5
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
6
6
|
|
|
@@ -214,13 +214,17 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
214
214
|
|
|
215
215
|
22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
|
|
216
216
|
`manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
|
|
217
|
-
assertion keys. The linter is **static
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
217
|
+
assertion keys. The linter is **static** until you opt in to cassette evidence. With committed
|
|
218
|
+
recordings, pass `lint --cassette-dir <dir>` (or one cassette file): it scans the same `*.cassette.json`
|
|
219
|
+
shape as replay and verify-cassettes, resolves each exact `scenarioSource` relative to its cassette,
|
|
220
|
+
and suppresses an advisory only when **all** matching cassettes carry the evidence the replay lane
|
|
221
|
+
needs. A malformed, wrong-shaped, unsupported-version, or provenance-less cassette is reported as INFO
|
|
222
|
+
and keeps the advisory visible — even beside a healthy sibling. The two dedicated diff assertions also
|
|
223
|
+
require their respective `preRunPaths` / `preRunHashes` baselines, and a non-empty `controlOut` /
|
|
224
|
+
artifact manifest is required where replay needs it. For a strict CI gate that keeps actionable INFO
|
|
225
|
+
rules visible while suppressing only proven replay noise, use `lint --strict --min-severity INFO
|
|
226
|
+
--cassette-dir <dir>`. Without the opt-in path, use `lint --min-severity WARN` in CI (≥1.11.0) to hide
|
|
227
|
+
the INFO class.
|
|
224
228
|
From 4.0.0 WARN is `--strict`'s default floor, so bare `lint --strict` hides and passes INFO; add
|
|
225
229
|
`--min-severity INFO` to fail on it. `--strict --min-severity ERROR` behaves as a plain lint, not a
|
|
226
230
|
contradiction.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Measurement
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
|
|
4
4
|
|
|
5
5
|
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Run, record and lock
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
|
|
4
4
|
|
|
5
5
|
## Part II — RUN, RECORD & LOCK
|
|
6
6
|
|
|
@@ -29,8 +29,11 @@ one a paid run would give. On a **directory** the path-dependent verdicts (host-
|
|
|
29
29
|
portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
|
|
30
30
|
the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
|
|
31
31
|
takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
|
|
32
|
-
contradiction, duplicate cassette target) gate the batch.
|
|
33
|
-
|
|
32
|
+
contradiction, duplicate cassette target) gate the batch. An input the real record would refuse — a missing
|
|
33
|
+
path, an unknown baseline name, a `tool_not_called` the tier can never violate — is listed under
|
|
34
|
+
`inputErrors[]` with a `⚠ input error:` line, also at exit 0. So a directory dry-run CAN exit 0 on a scenario the
|
|
35
|
+
real `record` would refuse; re-run that one file with its real flags for a binding answer, or gate on
|
|
36
|
+
`.ok and (.inputErrors == [])` in the JSON payload (4.1.0 and later). A directory also
|
|
34
37
|
reports every offender and the batch cost estimate. `lint` checks the assertion invariants AND that each file loads (the same loader, plus a named `baseline:`), but not the pre-spend refusals.
|
|
35
38
|
|
|
36
39
|
**Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
|
|
@@ -249,7 +252,9 @@ Recognize these before "fixing" a non-bug:
|
|
|
249
252
|
`container`/`microvm`, but only *fires* on an actual scanned leak with no authored
|
|
250
253
|
`transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
|
|
251
254
|
run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
|
|
252
|
-
it's valid.
|
|
255
|
+
it's valid. A host path that came verbatim from the scenario's own uploads, connected folders or prompt
|
|
256
|
+
is not a leak (whole-token match; a token cut short where the path goes on, or one naming a location the
|
|
257
|
+
harness created for the run, is never exempt); `scan.hostPathsFromInputs` counts the paths exempted.
|
|
253
258
|
- **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
|
|
254
259
|
infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
|
|
255
260
|
rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, replay class, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.1.0`
|
|
4
4
|
(baseline `desktop-2.9939.4`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -73,7 +73,7 @@ answers: # scripted answers (see below)
|
|
|
73
73
|
- when_question: "Which output format"
|
|
74
74
|
choose: "Markdown"
|
|
75
75
|
- when_tool: Bash
|
|
76
|
-
allow_if:
|
|
76
|
+
allow_if: '!/\brm\b/.test(command)' # a word match; `includes('rm')` also denies "normalize"
|
|
77
77
|
else: deny
|
|
78
78
|
- when_tool: Write
|
|
79
79
|
decide: allow
|
|
@@ -238,7 +238,8 @@ with no session file, the CLI flags `--folder <dir>` and `--upload <file>` are t
|
|
|
238
238
|
`plugins.remote_plugins` instead — Cowork serves an installed plugin from `.remote-plugins/plugin_<id>`
|
|
239
239
|
(named by id, not plugin name), while `local_plugins` mounts two levels deeper
|
|
240
240
|
(`.local-plugins/marketplaces/local-desktop-app-uploads/<plugin>`). The choice matters to a skill that
|
|
241
|
-
locates its own files from the shell (
|
|
241
|
+
locates its own files from the shell (at host-loop, Cowork's default, the braced `${CLAUDE_PLUGIN_ROOT}` is
|
|
242
|
+
replaced with a HOST path and a bare `$CLAUDE_PLUGIN_ROOT` is empty in the VM shell).
|
|
242
243
|
Search for the skill's own `SKILL.md`, not for a directory named after the plugin — that finds nothing
|
|
243
244
|
under `.remote-plugins/plugin_<id>` — and set no `-maxdepth` that stops short of the deeper local layout:
|
|
244
245
|
`find /sessions/*/mnt/.local-plugins /sessions/*/mnt/.remote-plugins -path '*/skills/<skill-name>/SKILL.md' 2>/dev/null | head -1`.
|
|
@@ -281,7 +282,7 @@ label validation by intent (mutually exclusive with `choose`):
|
|
|
281
282
|
- when_tool: Write
|
|
282
283
|
decide: allow # allow | deny
|
|
283
284
|
- when_tool: Bash
|
|
284
|
-
allow_if:
|
|
285
|
+
allow_if: '!/\brm\b/.test(command) && !command.includes("curl")' # JS predicate over tool input
|
|
285
286
|
else: deny # decision when predicate is false (default deny)
|
|
286
287
|
- when_tool: "webfetch:example.com" # a web_fetch approval (provenance-miss gate)
|
|
287
288
|
decide: allow
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 4.
|
|
5
|
+
Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|