cowork-harness 1.24.0 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +17 -9
- package/.claude/skills/cowork-harness/references/ci-recipe.md +63 -18
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -3
- package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
- package/CHANGELOG.md +544 -0
- package/CONTRIBUTING.md +12 -0
- package/DESIGN.md +3 -3
- package/README.md +12 -12
- package/SPEC.md +47 -20
- package/baselines/desktop-1.32885.1.json +521 -0
- package/baselines/desktop-1.34493.1.json +791 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +68 -0
- package/baselines/prompts/desktop-1.32885.1/subagent-append-hl.md +25 -0
- package/dist/cli.js +62 -16
- package/dist/run/cassette.js +939 -116
- package/dist/run/envelope.js +7 -1
- package/dist/run/jcs.js +81 -0
- package/dist/run/matrix.js +4 -2
- package/dist/run/provenance.js +37 -0
- package/dist/run/renderer.js +17 -0
- package/dist/run/repeat.js +36 -0
- package/dist/run/skill-hash.js +182 -45
- package/dist/scan.js +36 -2
- package/dist/sync/baseline-diff.js +9 -0
- package/dist/sync/cowork-sync.js +194 -27
- package/dist/types.js +4 -1
- package/docs/cassette.md +172 -32
- package/docs/debugging.md +15 -1
- package/docs/fidelity-gaps.md +22 -3
- package/docs/invariants.md +2 -1
- package/docs/maintenance.md +69 -8
- package/docs/protocol.md +14 -4
- package/docs/scenario.md +6 -4
- package/docs/session.md +3 -2
- package/examples/replays/README.md +15 -4
- package/examples/replays/example-multiselect-gate.cassette.json +54 -79
- package/examples/replays/example-pdf-skill.cassette.json +104 -98
- package/examples/replays/hostloop-computer-links.cassette.json +103 -68
- package/package.json +1 -1
- package/schema/cassette.v10.json +1 -1
- package/schema/cassette.v11.json +1 -1
- package/schema/cassette.v12.json +318 -0
- package/schema/cassette.v9.json +21 -19
- package/schema/run-result.json +69 -285
- package/schema/verify-cassettes.json +6 -2
- package/scripts/check-versions.ts +4 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version:
|
|
7
|
-
tracks-harness: cowork-harness
|
|
6
|
+
version: 2.0.0
|
|
7
|
+
tracks-harness: cowork-harness 2.0.0 (baseline desktop-1.34493.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness
|
|
26
|
-
> `desktop-1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.0.0` (baseline
|
|
26
|
+
> `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=2.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=2.0.0"`. **Pin `@>=2.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -407,7 +407,7 @@ That choice is permanent: the cassette rewrites `scenario.session` and `scenario
|
|
|
407
407
|
its own directory** at record time, so moving the file later — a different `--out`, a `git mv`, a copy
|
|
408
408
|
into another repo — leaves those unresolvable and
|
|
409
409
|
`verify-cassettes` reports `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you
|
|
410
|
-
re-record at the new location. **`record` now says so BEFORE it spends:** a pre-flight — at the same
|
|
410
|
+
re-record at the new location — or point `replay`/`verify-cassettes` at the session with `--session <file>`, which resolves it without a re-record. Since 2.0.0 a bare `replay` FAILS on this class rather than warning. **`record` now says so BEFORE it spends:** a pre-flight — at the same
|
|
411
411
|
pre-spend point as the host-inventory refusal, and in `record --dry-run`, so the rehearsal is free —
|
|
412
412
|
warns when the cassette would be written outside the scenario's tree, or when `session:` itself lives
|
|
413
413
|
outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
|
|
@@ -584,7 +584,9 @@ discovery source removed, so the agent answers from its own priors. **It is ONE
|
|
|
584
584
|
experiment**: this invocation is the control. Run the same prompt a second time *without* the flag for
|
|
585
585
|
the treatment arm and compare them yourself. Composed with `--repeat 5` it produces **5 ablated runs
|
|
586
586
|
and 0 treatment runs** — N samples of the control, which is the intended reading and is not an A/B.
|
|
587
|
-
|
|
587
|
+
The rollup says so on its verdict line: `repeat "<skill>": PASS [ABLATED — control arm] — 5/5 passed`.
|
|
588
|
+
Every ablated run is stamped `ablated: true` in `result.json` and carries `ablated=true` on its
|
|
589
|
+
`[provenance]` footer line; a run that isn't stamped is a real run.
|
|
588
590
|
What the harness gives you here is the run execution and the control arm — designing the comparison
|
|
589
591
|
(scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
|
|
590
592
|
|
|
@@ -1026,8 +1028,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
1026
1028
|
25. **Two distinct host-inventory consent flags — a record-time one and a verify-time one.** `record
|
|
1027
1029
|
--allow-host-inventory-fixture` is the boolean consent to proceed recording a host-inheriting
|
|
1028
1030
|
(`protocol`/`hostloop`/`cowork`-resolving-to-hostloop) cassette into a repo-visible path — otherwise
|
|
1029
|
-
`record` refuses (freezing this machine's MCP servers/agents/account into a committed
|
|
1030
|
-
risk).
|
|
1031
|
+
`record` refuses before it spends (freezing this machine's MCP servers/agents/account into a committed
|
|
1032
|
+
fixture is the risk). That pre-spend check **warns rather than refuses when the cassette already
|
|
1033
|
+
exists** — refusing would fire on every `--rerecord-stale` pass — and it reads the tier and the
|
|
1034
|
+
destination path, never the bytes. So `record` also scans the FINISHED recording, after redaction and
|
|
1035
|
+
before the write: a `host-inventory`/`machine-inventory` finding on a repo-visible path is
|
|
1036
|
+
**quarantined** to `<runs-root>/quarantine/` with a `.findings.txt` naming what leaked, and the command
|
|
1037
|
+
fails without writing the path you asked for (the recording is not discarded — you paid for it).
|
|
1038
|
+
`verify-cassettes --allow-host-inventory <regex>` is unrelated: a per-finding suppressor for the
|
|
1031
1039
|
scanner's `host-inventory` class on an already-committed cassette. Passing one where the other command
|
|
1032
1040
|
wants it fails as an unrecognized flag — they don't interchange. Depth: `references/ci-recipe.md`.
|
|
1033
1041
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "
|
|
16
|
+
release; pin an exact version (e.g. `version: "2.0.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -32,7 +32,7 @@ jobs:
|
|
|
32
32
|
- uses: actions/checkout@v4
|
|
33
33
|
- name: Stage the agent binary (official channel, sha256-verified — see https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md)
|
|
34
34
|
run: |
|
|
35
|
-
V=2.1.
|
|
35
|
+
V=2.1.237 # match your scenario's pinned baseline's agentVersion
|
|
36
36
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
37
37
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
38
38
|
# verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
|
|
@@ -58,7 +58,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
58
58
|
GitHub-hosted runners, no token/Docker/agent:
|
|
59
59
|
|
|
60
60
|
```yaml
|
|
61
|
-
- run: npm i -g "cowork-harness@>=
|
|
61
|
+
- run: npm i -g "cowork-harness@>=2.0.0"
|
|
62
62
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
63
63
|
# no silent false-greens. WITHOUT --strict this
|
|
64
64
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -77,7 +77,20 @@ GitHub-hosted runners, no token/Docker/agent:
|
|
|
77
77
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
78
78
|
```
|
|
79
79
|
|
|
80
|
-
|
|
80
|
+
If a cassette has MOVED (a `git mv`, a repo reorg, a copy between projects), staleness becomes
|
|
81
|
+
unverifiable — `verify-cassettes` exits 3 and says the skill dirs are not resolvable. Recover with
|
|
82
|
+
`--session <file>` on either command rather than re-recording:
|
|
83
|
+
|
|
84
|
+
```yaml
|
|
85
|
+
- run: cowork-harness verify-cassettes cassettes/moved.cassette.json --session sessions/default.yaml
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
It takes a session (not skill dirs) so `staleness.hash_ignore` survives, refuses a directory target,
|
|
89
|
+
and echoes the dirs it resolved. Full contract:
|
|
90
|
+
[docs/cassette.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md).
|
|
91
|
+
|
|
92
|
+
Both lines matter: for the CONTENT-drift classes (`skill`, `shared-root`) `replay` alone **warns** and
|
|
93
|
+
exits 0, so dropping the
|
|
81
94
|
`verify-cassettes` step means a skill edit silently stops being tested. (One command instead of two:
|
|
82
95
|
`replay --fail-on-skill-drift`.)
|
|
83
96
|
|
|
@@ -210,9 +223,23 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
210
223
|
- **Host-inheriting record refused by default — `--allow-host-inventory-fixture` is the consent.** A
|
|
211
224
|
`protocol`/`hostloop`/`cowork`-resolving-to-hostloop record into a repo-visible cassette path would
|
|
212
225
|
freeze THIS machine's MCP server names, agents, and account metadata into a committed fixture, so
|
|
213
|
-
`record` refuses
|
|
226
|
+
`record` refuses **before the paid spawn**. Pass `--allow-host-inventory-fixture` only when the
|
|
214
227
|
recording session genuinely has no personal MCP servers or plugins to leak — it is a per-record
|
|
215
228
|
boolean consent, not a pattern.
|
|
229
|
+
|
|
230
|
+
Two details that matter for a re-record loop. The pre-spend check **warns rather than refuses when the
|
|
231
|
+
cassette already exists**, deliberately: refusing there would fire on every `--rerecord-stale` pass and
|
|
232
|
+
make the escape flag reflexive. And it is a **prediction** — it reads the tier and the destination path,
|
|
233
|
+
never the resulting bytes, so it can be wrong in both directions.
|
|
234
|
+
|
|
235
|
+
So `record` also checks the **evidence**. After redaction and before the write, the finished cassette is
|
|
236
|
+
scanned; a `host-inventory` or `machine-inventory` finding on a repo-visible path is **quarantined** —
|
|
237
|
+
written to `<runs-root>/quarantine/` (honouring `--run-dir`/`COWORK_HARNESS_RUNS_DIR`) with a
|
|
238
|
+
`.findings.txt` sibling naming what leaked, and the command fails without writing the path you asked for.
|
|
239
|
+
The recording is not discarded; you paid for it. Only the machine-identity classes trigger this —
|
|
240
|
+
`email`/`currency`/`domain`/`path` are frequently legitimate scenario content, and a gate that fires on
|
|
241
|
+
those just teaches you to pass the escape flag. Outside a git repo it warns instead: nothing there
|
|
242
|
+
publishes the file by accident.
|
|
216
243
|
- **Always-on scan gate** — `verify-cassettes` flags email / currency / bare-domain / local-path /
|
|
217
244
|
machine-inventory matches it finds in the committed cassettes and **exits non-zero**, so "no leak" is
|
|
218
245
|
a gate, not discipline. Non-zero is not one thing, though: exit `1` means verification RAN and found a
|
|
@@ -222,6 +249,13 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
222
249
|
$? -ne 0 ]` tripwire treats both the same — if you need to tell "the gate caught something" apart from
|
|
223
250
|
"the gate couldn't run", branch on the exit code (or parse `--output-format json`'s per-file
|
|
224
251
|
`findings`/`staleness` vs `unverifiable`/`version`/`error` buckets).
|
|
252
|
+
|
|
253
|
+
**Do not read `error` as "this file was never scanned".** The privacy scan needs a readable *transcript*
|
|
254
|
+
(an `events` array of strings), not a *valid* cassette — so a file that fails shape validation is still
|
|
255
|
+
scanned, and reports its findings **and** its `error`. Each result carries **`privacyScanned`**, which
|
|
256
|
+
answers that question directly. A gate that must not treat "could not verify" as "verified clean" should
|
|
257
|
+
key on `privacyScanned === false`, where `findings: []` is an absence of evidence rather than evidence of
|
|
258
|
+
absence. `--skip-privacy` also reports `false`, for the same reason.
|
|
225
259
|
Suppress synthetic / public reference names (NVCA, Cooley GO, …) with `--allow <regex>`. (Multi-word
|
|
226
260
|
proper names are NOT a default class — too noisy to gate on; add a pattern via config if your corpus
|
|
227
261
|
needs it.)
|
|
@@ -289,7 +323,7 @@ jobs:
|
|
|
289
323
|
with: { node-version: '24' }
|
|
290
324
|
- uses: actions/setup-python@v5
|
|
291
325
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
292
|
-
- run: npm i -g "cowork-harness@>=
|
|
326
|
+
- run: npm i -g "cowork-harness@>=2.0.0"
|
|
293
327
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
294
328
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
295
329
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -318,7 +352,7 @@ jobs:
|
|
|
318
352
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
319
353
|
fi
|
|
320
354
|
- if: steps.guard.outputs.live == 'true'
|
|
321
|
-
run: npm i -g "cowork-harness@>=
|
|
355
|
+
run: npm i -g "cowork-harness@>=2.0.0"
|
|
322
356
|
- if: steps.guard.outputs.live == 'true'
|
|
323
357
|
run: cowork-harness run scenarios/ --output-format json
|
|
324
358
|
env:
|
|
@@ -357,13 +391,19 @@ in `result.json` and in the stdout envelope, by construction). Each entry is
|
|
|
357
391
|
- **`staleness`** — on `replay --strict` / `--assert-from` / `--reassert`, skill or baseline drift,
|
|
358
392
|
which those modes escalate to a hard failure on purpose (frozen events must not green an edited
|
|
359
393
|
assert against a skill whose current source produces something else).
|
|
360
|
-
- **`cassette-format`** — the cassette
|
|
394
|
+
- **`cassette-format`** — the cassette itself cannot be interpreted: too new a version for this build,
|
|
395
|
+
OR corrupt (duplicate/malformed control frames, a truncated recording). Before 1.25.0 the corrupt
|
|
396
|
+
cases were mis-reported as `assertion`, i.e. as if you had written them.
|
|
361
397
|
- **`coverage`** — a `verify-run` answer-coverage miss: a gate the run fired that your `answers:` block
|
|
362
398
|
does not cover.
|
|
363
399
|
|
|
364
|
-
So `jq '[.verdict.failures[] | select(.kind=="assertion")]'` answers "did MY
|
|
365
|
-
`select(.kind=="staleness")` answers "is the cassette stale?", from
|
|
366
|
-
either from the exit code: every kind lands on exit 1.
|
|
400
|
+
So `jq '[.results[]? | .verdict.failures[]? | select(.kind=="assertion")] | length'` answers "did MY
|
|
401
|
+
assertions pass?" and swapping in `select(.kind=="staleness")` answers "is the cassette stale?", from
|
|
402
|
+
the envelope alone. Do not infer either from the exit code: every kind lands on exit 1.
|
|
403
|
+
|
|
404
|
+
Keep the `.results[]?` hop and the `?` operators. `.results[0]` silently ignores every scenario after
|
|
405
|
+
the first when you pass a directory, and a bare `.verdict` does not exist at the envelope root at all —
|
|
406
|
+
both read as "no failures" against a run that failed.
|
|
367
407
|
|
|
368
408
|
> **Do not filter on whether `assertion` is present.** That was the only discriminator before `kind`
|
|
369
409
|
> existed and it never worked in both directions: `coverage` entries carry a key too (an internal
|
|
@@ -400,19 +440,21 @@ evaluation), not present-and-passing. A CI script that counts assertions will se
|
|
|
400
440
|
on replay vs live — compare by assertion identity / pass-fail, not by total count. The count of
|
|
401
441
|
skipped live-only assertions is reported on each replay result as `skippedAssertions: {full, partial}`.
|
|
402
442
|
|
|
403
|
-
## Staleness does NOT fail a replay
|
|
443
|
+
## Staleness mostly does NOT fail a replay — read it from the JSON
|
|
404
444
|
|
|
405
|
-
A plain `replay` **warns** on a
|
|
406
|
-
does **not** imply the recording is still valid.
|
|
445
|
+
A plain `replay` **warns** on a DRIFTED cassette (skill/baseline drift) but stays `ok:true` — a green
|
|
446
|
+
replay does **not** imply the recording is still valid. **Since 2.0.0 there is one exception:**
|
|
447
|
+
`unverifiable-skill` — staleness that could not be checked at all, most often a cassette that moved —
|
|
448
|
+
FAILS a bare `replay`. Recover with `--session <file>` rather than re-recording. Each replay result carries `staleness[]`, an array of
|
|
407
449
|
`{class, message}`, so a token-free gate can act on it without `ok` being the whole story:
|
|
408
450
|
|
|
409
451
|
| `class` | meaning | concern |
|
|
410
452
|
|---|---|---|
|
|
411
453
|
| `baseline` | platform baseline moved since record | low (format-compatible) |
|
|
412
454
|
| `skill` / `shared-root` | the skill source the assertions validate drifted | **high** (assertions may validate dead code) |
|
|
413
|
-
| `format` |
|
|
455
|
+
| `format` | the git/raw file-set mode or agent-scope differs from the recording | re-record under the same setting (waivable) |
|
|
414
456
|
| `unverifiable-baseline` | the latest baseline couldn't be loaded | couldn't verify (env, not skill) |
|
|
415
|
-
| `unverifiable-skill` | skill dirs unresolvable
|
|
457
|
+
| `unverifiable-skill` | skill dirs unresolvable, **or** the cassette predates the hash-format epoch (v12) so its digest is not comparable | couldn't verify the skill. **Fails a bare `replay`.** For the epoch case try `rehash` first — it migrates without a re-record where it can prove the content unchanged (`rehash <file> --session <s.yaml>` if the cassette moved) |
|
|
416
458
|
| `resolved-tier` | a `fidelity: cowork` cassette's recorded `effectiveFidelity` no longer matches what the baseline resolves to today (the host-loop gate flipped) | **high** (the recording exercises the wrong tier) |
|
|
417
459
|
| `unverifiable-tier` | tier check couldn't run for a `fidelity: cowork` cassette (no recorded `effectiveFidelity`, or its pinned baseline failed to load) | couldn't verify the tier — re-record |
|
|
418
460
|
|
|
@@ -427,8 +469,11 @@ fail it either way.)
|
|
|
427
469
|
To gate in CI, pick the severity you want:
|
|
428
470
|
|
|
429
471
|
- `replay --strict` — fail (exit 1) on **any** staleness class.
|
|
430
|
-
- `replay --fail-on-skill-drift` — fail
|
|
472
|
+
- `replay --fail-on-skill-drift` — fail on the skill-source DRIFT classes (`skill` / `shared-root`); `unverifiable-skill` needs no flag, it fails the default verdict since 2.0.0;
|
|
431
473
|
baseline / format / `unverifiable-baseline` stay non-failing warnings.
|
|
474
|
+
Note `--allow-failing` waives this gate wholesale, including the copy `--assert-from` turns on for you:
|
|
475
|
+
`replay --assert-from … --write --allow-failing` will persist an assert block validated against a
|
|
476
|
+
recording whose skill sources have since moved. Re-record when the drift is real.
|
|
432
477
|
- Or read `results[].staleness[].class` yourself and decide.
|
|
433
478
|
|
|
434
479
|
Both flags realize the gate as failing assertions, so the verdict / `ok` / exit code stay consistent with the
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness
|
|
3
|
+
Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.0.0`
|
|
4
|
+
(baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -107,7 +107,7 @@ form a relocatable bundle. `~` expands to home.
|
|
|
107
107
|
> references relative to **its own** directory at record time (`scenario.session` and
|
|
108
108
|
> `scenarioSource`), so moving it afterwards — a different `--out`, a
|
|
109
109
|
> `git mv`, a copy into another repo — leaves them unresolvable and `verify-cassettes` reports
|
|
110
|
-
> `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you re-record at the new location
|
|
110
|
+
> `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you re-record at the new location, or pass `--session <file>`.
|
|
111
111
|
> Decide where a cassette will live *before* you record it.
|
|
112
112
|
|
|
113
113
|
## Session YAML
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness
|
|
5
|
+
Tracks `cowork-harness 2.0.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -182,7 +182,7 @@ degrade the advice. It is real work to calibrate; these steps are the traps that
|
|
|
182
182
|
**that one invocation**, so the agent answers from its own priors, and stamps the result
|
|
183
183
|
`ablated: true`. It is the **control arm only** — run the same prompt again *without* the flag for
|
|
184
184
|
the treatment arm; `--ablate-skill --repeat 5` gives you 5 control runs and 0 treatment runs, not an
|
|
185
|
-
A/B. (Inspecting an organically not-invoked rep works too, but is not a substitute: that rep may
|
|
185
|
+
A/B — and the rollup labels it `PASS [ABLATED — control arm]` so you cannot bank it as one. (Inspecting an organically not-invoked rep works too, but is not a substitute: that rep may
|
|
186
186
|
differ for other reasons.) If the answer still scores high without the skill, that claim is
|
|
187
187
|
answerable from priors and tests the model, not your skill — strengthen it (a skill-specific fact) or
|
|
188
188
|
drop it. Everything past "run both arms" — scrubbing giveaways, shuffling, judging blind, unblinding
|