cowork-harness 1.13.1 → 1.13.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +4 -4
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +7 -5
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/scenario.py +23 -8
- package/CHANGELOG.md +20 -0
- package/README.md +6 -6
- package/docs/critique.md +5 -2
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.13.
|
|
7
|
-
tracks-harness: cowork-harness 1.13.
|
|
6
|
+
version: 1.13.2
|
|
7
|
+
tracks-harness: cowork-harness 1.13.2 (baseline desktop-1.24012.9)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.13.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.13.2` (baseline
|
|
26
26
|
> `desktop-1.24012.9`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.13.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.13.2**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.13.2" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.13.2"`. **Pin `@>=1.13.2`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.13.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.13.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.13.2"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -58,7 +58,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
58
58
|
GitHub-hosted runners, no token/Docker/agent:
|
|
59
59
|
|
|
60
60
|
```yaml
|
|
61
|
-
- run: npm i -g "cowork-harness@>=1.13.
|
|
61
|
+
- run: npm i -g "cowork-harness@>=1.13.2"
|
|
62
62
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
63
63
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
64
64
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -222,7 +222,7 @@ jobs:
|
|
|
222
222
|
with: { node-version: '20' }
|
|
223
223
|
- uses: actions/setup-python@v5
|
|
224
224
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
225
|
-
- run: npm i -g "cowork-harness@>=1.13.
|
|
225
|
+
- run: npm i -g "cowork-harness@>=1.13.2"
|
|
226
226
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
227
227
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
228
228
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -251,7 +251,7 @@ jobs:
|
|
|
251
251
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
252
252
|
fi
|
|
253
253
|
- if: steps.guard.outputs.live == 'true'
|
|
254
|
-
run: npm i -g "cowork-harness@>=1.13.
|
|
254
|
+
run: npm i -g "cowork-harness@>=1.13.2"
|
|
255
255
|
- if: steps.guard.outputs.live == 'true'
|
|
256
256
|
run: cowork-harness run scenarios/ --output-format json
|
|
257
257
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 1.13.
|
|
3
|
+
Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -52,10 +52,12 @@ separately at **128 KiB**, cut **head+tail with an elided middle**, so a run's s
|
|
|
52
52
|
both survive a cut instead of just one end.
|
|
53
53
|
|
|
54
54
|
You do not need a paid run to find out where you stand: **`cowork-harness lint-skill <skill-dir>` sizes
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
55
|
+
your corpus against the same ceiling**, reporting `skill-corpus-near-evidence-ceiling` (INFO) from 80%
|
|
56
|
+
and `skill-corpus-over-evidence-ceiling` (WARN, so it fails `--strict`) past it. It counts the three
|
|
57
|
+
classes the ceiling governs — `SKILL.md`, every file under `references/` (**any extension**: the packager
|
|
58
|
+
applies no extension filter, so JSON schemas and rule packs count toward your total), and a plugin
|
|
59
|
+
skill's `agents/<name>.md`. It does not apply staging's git-tracked filter, so an untracked reference
|
|
60
|
+
inflates the figure; `corpusCuts` below stays the authority.
|
|
59
61
|
|
|
60
62
|
The report's `evidenceBudget` object says exactly what was shown — read it instead of inferring budgets
|
|
61
63
|
from `dist/` source:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.13.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.13.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.13.2`
|
|
4
4
|
(baseline `desktop-1.24012.9`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 1.13.
|
|
5
|
+
Tracks `cowork-harness 1.13.2` (baseline `desktop-1.24012.9`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -1205,19 +1205,33 @@ _EVIDENCE_CORPUS_NOTICE_RATIO = 0.8
|
|
|
1205
1205
|
|
|
1206
1206
|
|
|
1207
1207
|
def _lint_skill_corpus_size(md_path):
|
|
1208
|
-
"""Total skill-authored bytes
|
|
1208
|
+
"""Total skill-authored bytes against the critique evidence ceiling.
|
|
1209
1209
|
|
|
1210
|
-
|
|
1211
|
-
|
|
1212
|
-
|
|
1213
|
-
|
|
1214
|
-
the
|
|
1210
|
+
Counts the SAME THREE CLASSES the ceiling governs: SKILL.md, every file under references/ (any
|
|
1211
|
+
extension -- the packager applies no extension filter, so JSON schemas and rule packs count), and,
|
|
1212
|
+
for a skill inside a multi-skill plugin, the invoked skill's <root>/agents/<name>.md.
|
|
1213
|
+
|
|
1214
|
+
Omitting the agents md is not a rounding error: a plugin whose SKILL.md + references sit in the INFO
|
|
1215
|
+
band while the agents md carries the corpus past the ceiling reported INFO and PASSED --strict on
|
|
1216
|
+
content the packager would cut. A proximity check that greens a corpus destined to be cut is worse
|
|
1217
|
+
than no check.
|
|
1218
|
+
|
|
1219
|
+
Still approximate in ONE direction only, and it now over- rather than under-counts: the packager
|
|
1220
|
+
applies staging's git-tracked filter, so an untracked reference inflates this figure. That errs
|
|
1221
|
+
toward warning early. The report's corpusCuts stays the authority."""
|
|
1215
1222
|
skill_dir = Path(md_path).parent
|
|
1216
1223
|
total = 0
|
|
1217
1224
|
files = [Path(md_path)]
|
|
1218
1225
|
refs = skill_dir / "references"
|
|
1219
1226
|
if refs.is_dir():
|
|
1220
1227
|
files.extend(p for p in sorted(refs.rglob("*")) if p.is_file())
|
|
1228
|
+
# Multi-skill plugin layout: skillDir is <root>/skills/<name> and the invoked skill's sub-agent
|
|
1229
|
+
# system prompt is <root>/agents/<name>.md -- the same resolution the critique command performs.
|
|
1230
|
+
# A standalone skill (no `skills/` parent) has no agents md and is unaffected.
|
|
1231
|
+
if skill_dir.parent.name == "skills":
|
|
1232
|
+
agents_md = skill_dir.parent.parent / "agents" / f"{skill_dir.name}.md"
|
|
1233
|
+
if agents_md.is_file():
|
|
1234
|
+
files.append(agents_md)
|
|
1221
1235
|
for p in files:
|
|
1222
1236
|
try:
|
|
1223
1237
|
total += p.stat().st_size
|
|
@@ -1231,8 +1245,9 @@ def _lint_skill_corpus_size(md_path):
|
|
|
1231
1245
|
"skill-corpus-over-evidence-ceiling",
|
|
1232
1246
|
f"skill content is {total:,} B ({pct:.0f}% of the {_EVIDENCE_CORPUS_CEILING:,} B critique "
|
|
1233
1247
|
f"evidence ceiling) — a critique will cut it before grading.",
|
|
1234
|
-
"Split or trim the largest references/ files. This
|
|
1235
|
-
"
|
|
1248
|
+
"Split or trim the largest references/ files. This counts SKILL.md + references/** + "
|
|
1249
|
+
"agents/<skill>.md, the same three classes the ceiling governs; it does not apply "
|
|
1250
|
+
"staging's git-tracked filter, so an untracked reference inflates it. The critique "
|
|
1236
1251
|
"report's corpusCuts names exactly which files lose bytes.",
|
|
1237
1252
|
str(skill_dir),
|
|
1238
1253
|
)
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,26 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.13.2] — 2026-07-28
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- **`lint-skill`'s corpus check sized a smaller set than the ceiling it warns about, and could pass
|
|
14
|
+
`--strict` on a corpus a critique would cut.** The evidence ceiling governs `SKILL.md` + `references/**`
|
|
15
|
+
+ `agents/<skill>.md` **combined**; the check summed only the first two. A multi-skill plugin whose
|
|
16
|
+
SKILL.md and references sat in the INFO band while its `agents/<name>.md` carried the total past the
|
|
17
|
+
ceiling reported INFO and **exited 0 under `--strict`** — a green gate on content that was already
|
|
18
|
+
destined to be cut. The check now counts all three classes, resolving `<root>/agents/<name>.md` the same
|
|
19
|
+
way `critique --skill <name>` does; a standalone skill (no `skills/` parent) is unaffected.
|
|
20
|
+
|
|
21
|
+
Two clarifications that follow, because both were understated: every file under `references/` counts
|
|
22
|
+
regardless of extension — the packager applies **no** extension filter, so JSON schemas and rule packs
|
|
23
|
+
are part of your corpus — and the remaining approximation now errs in one direction only. The check
|
|
24
|
+
skips staging's git-tracked filter, so an untracked reference inflates the figure and warns early
|
|
25
|
+
rather than late. `corpusCuts` in the report remains the authority.
|
|
26
|
+
|
|
27
|
+
Reported by a consumer, who caught it by measuring their own plugin against the new check on upgrade.
|
|
28
|
+
|
|
9
29
|
## [1.13.1] — 2026-07-28
|
|
10
30
|
|
|
11
31
|
### Fixed
|
package/README.md
CHANGED
|
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
91
91
|
|
|
92
92
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
93
93
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
94
|
-
> From a global install (`npm i -g "cowork-harness@>=1.13.
|
|
94
|
+
> From a global install (`npm i -g "cowork-harness@>=1.13.2"`), point at the package root instead:
|
|
95
95
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
96
96
|
> (or copy the cassette into your own project and pass that path).
|
|
97
97
|
|
|
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
101
101
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
102
102
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
103
103
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
104
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.13.
|
|
104
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.13.2"`.
|
|
105
105
|
|
|
106
106
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
107
107
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
126
126
|
claude plugin install cowork-harness@cowork-harness
|
|
127
127
|
```
|
|
128
128
|
|
|
129
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.13.
|
|
129
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.13.2"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
130
130
|
|
|
131
131
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
132
132
|
|
|
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
|
|
|
147
147
|
To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
|
|
148
148
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
149
149
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
150
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.13.
|
|
150
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.13.2"` — see
|
|
151
151
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
152
152
|
|
|
153
153
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -411,7 +411,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
411
411
|
| `scaffold <run-id>` | Turn a kept run into a starter scenario YAML (gates→answers, artifacts→`file_exists`) | authoring a scenario from a real run instead of guessing |
|
|
412
412
|
| `python3 …/scenario.py scaffold --name <name> --skill <dir>` | Generate a starter scenario skeleton from scratch (the `…` is `.claude/skills/cowork-harness/scripts/`) | starting a new scenario when you have no prior run |
|
|
413
413
|
| `lint <scenario.yaml \| dir/>…` | Check scenarios for silent false-greens — assertions placed on the wrong CI lane, mixed content/live keys, missing `controlOut`-required keys (files or a directory of `*.yaml`/`*.yml`; bundled `scenario.py`; needs python3 — PyYAML is bundled); `--json` emits findings as machine-readable JSON instead of the text report; `--min-severity ERROR|WARN|INFO` drops findings below a floor **before** both the report and the exit computation (so `--strict --min-severity ERROR` behaves as a plain lint), which is the way to mute the unconditional INFO advisories in CI | before committing a new scenario or after changing assertions |
|
|
414
|
-
| `lint-skill <SKILL.md \| skill-dir/>…` | Lint a skill body (and any sibling `hooks.json`) for two Cowork host-loop footguns — a `${CLAUDE_PLUGIN_ROOT}` path used in an in-VM bash context, and a hook command that exports an env var or writes into `/tmp` for the in-VM agent — plus static resolution of any pinned `subagent_type` against the enclosing plugin's `agents/*.md` (bundled `scenario.py`; needs python3); the two footguns are WARN-only, and of the three `subagent_type` outcomes, an in-plugin-prefixed agent missing from the enclosing plugin's fully-enumerated `agents/*.md` (`subagent-type-not-found-in-plugin`) is a **provable typo and is WARN too**, while a cross-plugin (`subagent-type-unresolvable`) or unknown-bare (`subagent-type-unknown`) value stays INFO (no built-in agent-type registry to disprove it against); it also sizes the skill's own content
|
|
414
|
+
| `lint-skill <SKILL.md \| skill-dir/>…` | Lint a skill body (and any sibling `hooks.json`) for two Cowork host-loop footguns — a `${CLAUDE_PLUGIN_ROOT}` path used in an in-VM bash context, and a hook command that exports an env var or writes into `/tmp` for the in-VM agent — plus static resolution of any pinned `subagent_type` against the enclosing plugin's `agents/*.md` (bundled `scenario.py`; needs python3); the two footguns are WARN-only, and of the three `subagent_type` outcomes, an in-plugin-prefixed agent missing from the enclosing plugin's fully-enumerated `agents/*.md` (`subagent-type-not-found-in-plugin`) is a **provable typo and is WARN too**, while a cross-plugin (`subagent-type-unresolvable`) or unknown-bare (`subagent-type-unknown`) value stays INFO (no built-in agent-type registry to disprove it against); it also sizes the skill's own content — `SKILL.md` + every file under `references/` (any extension) + a plugin skill's `agents/<name>.md`, the same three classes the ceiling governs — against the 512 KiB `critique` evidence ceiling — `skill-corpus-near-evidence-ceiling` at ≥ 80% is INFO, and `skill-corpus-over-evidence-ceiling` past it is **WARN**, so a corpus a critique would cut fails `--strict` — pass `--strict` (the CI-recommended invocation; plain `lint-skill` is advisory-only) to fail on any WARN, or `--json` for machine-readable findings | authoring or reviewing a skill before a paid Cowork host-loop run exposes the footgun, or before a pinned `subagent_type` typo breaks a `Task` dispatch |
|
|
415
415
|
| `python3 …/scenario.py resolve-agent-types <plugin-dir> [--json]` | Token-free: print a plugin's valid `<plugin>:<agent>` subagent types, resolved from `.claude-plugin/plugin.json` + `agents/*.md` frontmatter (filename-stem fallback when an agent file has no `name:`); the `…` is `.claude/skills/cowork-harness/scripts/` | "does `founder-skills:deck-review` resolve within this plugin?" without a live dispatch |
|
|
416
416
|
| `analyze-skill <SKILL.md \| skill-dir/ \| glob>…` | Token-free ADVISORY scan: flags a `/sessions/...` path handed to a file tool or used as a dispatch/sub-agent output path — that path class is DENIED on host-loop. Accepts multiple positionals (files, dirs, or a simple `*`/`**` glob), walked recursively across a plugin's `SKILL.md`/`agents/`/`references/`/`commands/`. Findings print but exit 0 by default; `--strict` (the CI-recommended invocation) fails on any unsuppressed finding. Three per-file ignore markers (`ignore-next-line`, `ignore-start`/`ignore-end`, file-wide `ignore`) and a `--output-format json` payload are supported. **It also flags interactive-artifact write-backs lost under Cowork**: it statically analyzes `.html`/`.js`/`.ts`/`.py` sources under the target for a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` write-back that silently fails when the artifact is served from Cowork's own origin (`artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; a candidate that can't be parsed is a could-not-verify exit `3`). `--runtime` adds an optional headless-DOM confirmation (needs `jsdom`) that *observes* the lost write-back. Full reference — ignore-marker syntax, directory-walk rules, JSON shape, and the artifact detector: [docs/subagents.md](./docs/subagents.md#static-path-fidelity-check-analyze-skill) | catching the exact "skill hands `/sessions/...` to a file tool" defect — or an artifact whose Submit silently fails under Cowork — statically, before paying for a live run to discover it |
|
|
417
417
|
| `probe-dispatch <skill-dir> "<prompt>"` | Cheap single-dispatch mechanics probe: a THIN wrapper over `skill` (fidelity `container`/`microvm`/`hostloop`, default `hostloop`) that scopes a prompt to trigger ONE `Task` dispatch, then prints just that dispatch's `{resolvedAgentType, pathDenials, delivered}` — no new data model, a pure projection of the same `RunResult` `skill` already produces; `--expect-write <suffix>` narrows `delivered`; "one dispatch" is prompt-scoped, not enforced (`--output-format json` for machine consumption); also inherits the common decider/answer flags — `--decider-cmd`, `--decider-dir`, `--on-unanswered`, `--ablate-skill` | "did THIS dispatch resolve to the type I expect, avoid a path denial, and actually deliver its write?" without hand-writing a scenario or reading `trace --view dispatches` |
|
|
@@ -694,7 +694,7 @@ jobs:
|
|
|
694
694
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
695
695
|
```
|
|
696
696
|
|
|
697
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.13.
|
|
697
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.13.2` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
|
|
698
698
|
|
|
699
699
|
The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **seven-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
700
700
|
|
package/docs/critique.md
CHANGED
|
@@ -245,8 +245,11 @@ actually breached), `corpusExcluded` (files present on the host but never delive
|
|
|
245
245
|
staging — untracked, with git-mode on), and `trimRecord` (any section the overall belt-and-suspenders cap
|
|
246
246
|
shaved). `cowork-harness lint-skill <skill-dir>` answers the same proximity question **without a paid
|
|
247
247
|
run** — `skill-corpus-near-evidence-ceiling` (INFO) from 80%, `skill-corpus-over-evidence-ceiling` (WARN,
|
|
248
|
-
so it fails `--strict`) past it
|
|
249
|
-
`
|
|
248
|
+
so it fails `--strict`) past it. It counts the same three classes the ceiling governs: `SKILL.md`, every
|
|
249
|
+
file under `references/` (**any extension** — the packager applies no extension filter, so JSON schemas
|
|
250
|
+
and rule packs count), and a plugin skill's `agents/<name>.md`. It does not apply staging's git-tracked
|
|
251
|
+
filter, so an untracked reference inflates the figure — it errs toward warning early, and `corpusCuts`
|
|
252
|
+
stays the authority.
|
|
250
253
|
On a normal skill this is one reassuring line; the other fields only grow teeth on a genuinely
|
|
251
254
|
oversized skill or an untracked-file mistake.
|
|
252
255
|
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.13.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.13.2"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "1.13.
|
|
3
|
+
"version": "1.13.2",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|