cowork-harness 1.13.2 → 1.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +87 -21
- package/.claude/skills/cowork-harness/references/ci-recipe.md +10 -8
- package/.claude/skills/cowork-harness/references/critique.md +3 -2
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +13 -9
- package/.claude/skills/cowork-harness/references/scenario-schema.md +25 -6
- package/.claude/skills/cowork-harness/references/task-recipes.md +11 -3
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +4 -1
- package/.claude/skills/cowork-harness/scripts/scenario.py +92 -37
- package/AGENTS.md +1 -1
- package/CHANGELOG.md +410 -0
- package/CONTRIBUTING.md +4 -3
- package/DESIGN.md +10 -4
- package/README.md +32 -25
- package/RELEASING.md +14 -7
- package/SPEC.md +1 -1
- package/dist/assert.js +35 -16
- package/dist/cli.js +183 -26
- package/dist/critique/command.js +87 -12
- package/dist/critique/limitations.js +0 -9
- package/dist/dotenv.js +30 -5
- package/dist/egress/proxy.js +1 -1
- package/dist/egress/sidecar.js +8 -4
- package/dist/hostloop/cowork-handler-hostloop.js +140 -0
- package/dist/hostloop/cowork-handler.js +23 -15
- package/dist/run/analyze-skill.js +2 -2
- package/dist/run/artifacts.js +59 -10
- package/dist/run/budget.js +119 -0
- package/dist/run/cassette.js +180 -14
- package/dist/run/chat-result.js +2 -0
- package/dist/run/doctor.js +8 -4
- package/dist/run/execute.js +38 -4
- package/dist/run/renderer.js +27 -1
- package/dist/run/repeat-flags.js +5 -2
- package/dist/run/run-index.js +208 -25
- package/dist/run/run.js +28 -1
- package/dist/run/skill-flag-surface.js +15 -2
- package/dist/run/verdict.js +118 -1
- package/dist/runtime/container.js +21 -12
- package/dist/runtime/hostloop.js +44 -7
- package/dist/session.js +5 -1
- package/dist/types.js +21 -2
- package/docker/Dockerfile.proxy +4 -2
- package/docker/compose.yml +5 -0
- package/docs/README.md +2 -2
- package/docs/cassette.md +14 -8
- package/docs/critique.md +14 -15
- package/docs/debugging.md +8 -6
- package/docs/fidelity-gaps.md +117 -1
- package/docs/gotchas.md +1 -1
- package/docs/invariants.md +1 -0
- package/docs/run-status.md +3 -0
- package/docs/scenario.md +120 -8
- package/docs/stats.md +92 -13
- package/examples/README.md +7 -5
- package/examples/replays/README.md +1 -1
- package/llms.txt +3 -1
- package/package.json +3 -3
- package/python/README.md +3 -2
- package/python/test_scenario_lint.py +176 -34
- package/schema/critique-report.json +6 -1
- package/schema/run-result.json +15 -3
- package/schema/scenario.schema.json +16 -2
- package/scripts/bump-version.ts +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.15.0
|
|
7
|
+
tracks-harness: cowork-harness 1.15.0 (baseline desktop-1.24012.9)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.15.0` (baseline
|
|
26
26
|
> `desktop-1.24012.9`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.15.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.15.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.15.0"`. **Pin `@>=1.15.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -284,12 +284,22 @@ quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `
|
|
|
284
284
|
(ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
|
|
285
285
|
baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
|
|
286
286
|
`allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
|
|
287
|
-
unverifiable),
|
|
288
|
-
`
|
|
287
|
+
unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
|
|
288
|
+
`present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
|
|
289
|
+
to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
|
|
290
|
+
`protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
|
|
289
291
|
and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
|
|
290
292
|
(CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
|
|
291
293
|
emits a scenario `lint` would reject.
|
|
292
294
|
|
|
295
|
+
**`lint` is the LENIENT check — the loader is the strict one.** An unknown top-level key is a ⚠ WARN in
|
|
296
|
+
`lint` (exit 0) but a **hard error** in the runtime (`Unrecognized key: "<k>"`, exit 2) — so a scenario
|
|
297
|
+
that lints with warnings may still not run. To check whether a scenario actually loads, without spending:
|
|
298
|
+
`cowork-harness record <file.yaml> --dry-run` (exit 2 on a schema error; a directory reports each
|
|
299
|
+
`✗ broken:` file and exits 1). Corollary: **a key from a newer harness fails LOUD on an older CLI**, never
|
|
300
|
+
silently — `lane:` on a pre-1.14.0 CLI does not load rather than defaulting to `lane: local`. That is why a
|
|
301
|
+
new scenario key ships with a version floor.
|
|
302
|
+
|
|
293
303
|
## Part II — RUN, RECORD & LOCK
|
|
294
304
|
|
|
295
305
|
You have an authored scenario. This Part runs it, reads the verdict, locks it into a
|
|
@@ -368,7 +378,12 @@ cassette — has its own recipe:
|
|
|
368
378
|
input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
|
|
369
379
|
pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
|
|
370
380
|
produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
|
|
371
|
-
kept run predates the current skill). **
|
|
381
|
+
kept run predates the current skill). **The hazard is general, not critique-specific:** repeated
|
|
382
|
+
`run`/`skill` invocations of one scenario accumulate in the SAME scenario directory regardless of skill
|
|
383
|
+
version, so a plain `stats <scenario>` silently averages pre-fix and post-fix runs together. Compare
|
|
384
|
+
generations with **`stats <scenario> --group-by skill-hash`** (or narrow with `--skill-hash <prefix>` /
|
|
385
|
+
`--label <tag>`); an un-split window spanning more than one generation now warns.
|
|
386
|
+
**Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
|
|
372
387
|
skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
|
|
373
388
|
critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
|
|
374
389
|
pairing step there silently groups on an absent key instead of erroring — check the field is present, or
|
|
@@ -406,6 +421,16 @@ Recognize these before "fixing" a non-bug:
|
|
|
406
421
|
covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
|
|
407
422
|
the final message before acting — a legitimate question-posing answer that wrote a file never fires.
|
|
408
423
|
Assert `allow_stall: true` if ending on a question is the intended terminal state.
|
|
424
|
+
- **`undelivered_deliverables`** (`WARN`) — the skill produced file(s) **outside every user-visible root**
|
|
425
|
+
and never delivered them. On a **remote** Cowork session the workspace is reclaimed at session end, so
|
|
426
|
+
they are destroyed; on a **local** one they persist but stay invisible to the user. Either way the user
|
|
427
|
+
does not get them. It fires with no assertion written — `present_files_called` covers the positive case
|
|
428
|
+
only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
|
|
429
|
+
**Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
|
|
430
|
+
scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
|
|
431
|
+
Fix by writing deliverables under `outputs/` or a connected folder, or delivering them explicitly; assert
|
|
432
|
+
**`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
|
|
433
|
+
downloaded inputs) rather than a delivery gap.
|
|
409
434
|
- **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
|
|
410
435
|
paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
|
|
411
436
|
`container`/`microvm`, but only *fires* on an actual scanned leak with no authored
|
|
@@ -424,7 +449,7 @@ Recognize these before "fixing" a non-bug:
|
|
|
424
449
|
`RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
|
|
425
450
|
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
426
451
|
|
|
427
|
-
The full
|
|
452
|
+
The full 17-code signal table (severity + per-signal opt-out) is in
|
|
428
453
|
[`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
|
|
429
454
|
the fuller narrative.
|
|
430
455
|
|
|
@@ -440,7 +465,10 @@ harness writes/updates throughout the run's lifecycle (including a crash-safety
|
|
|
440
465
|
error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
|
|
441
466
|
either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
|
|
442
467
|
checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
|
|
443
|
-
run starts, so capture stderr to get the exact directory —
|
|
468
|
+
run starts, so capture stderr to get the exact directory — **unless you passed `--compact` (or `--demo`,
|
|
469
|
+
which implies it), which suppress that line** (it is a raw, un-tildeified host path, exactly what those shareable-output
|
|
470
|
+
modes exist to withhold; `status.json` is still written either way, so `status` still works) — but
|
|
471
|
+
`<dir>` also accepts the run-dir root
|
|
444
472
|
passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
|
|
445
473
|
newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
|
|
446
474
|
rather than hanging forever. (Fuller recipe in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) — repo-only, not in the installed
|
|
@@ -508,9 +536,28 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
508
536
|
call-count/timing table, the sub-agent dispatch tree, the gate lifecycle, the tool/error rollups, …);
|
|
509
537
|
bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
|
|
510
538
|
the current list rather than relying on a fixed enumeration here.
|
|
539
|
+
- **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
|
|
540
|
+
is held to. Cowork picks the lane per session ("Run this task: In the cloud / On your computer") and
|
|
541
|
+
cloud is the default for new sessions; the lanes disagree about what *delivered* means. On `remote`,
|
|
542
|
+
location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
|
|
543
|
+
session end), `present_files` is NOT served, and `user_visible_artifact` /
|
|
544
|
+
`present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
|
|
545
|
+
to check a skill's delivery survives the lane most new sessions get. Orthogonal to `fidelity` — a
|
|
546
|
+
`lane: remote` scenario still runs locally.
|
|
511
547
|
- **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
|
|
512
|
-
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`.
|
|
513
|
-
-
|
|
548
|
+
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
|
|
549
|
+
plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
|
|
550
|
+
`--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
|
|
551
|
+
tier — instead of aggregating across them (a window spanning >1 generation warns — see Gotcha 6;
|
|
552
|
+
>1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
|
|
553
|
+
individual runs behind each summary with their `skillHash`/`runLabel`, so
|
|
554
|
+
you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
|
|
555
|
+
- **`result.json` carries the raw fields** the assertions read: `verdict`, `lane` (which Cowork delivery
|
|
556
|
+
contract the run was held to — see Gotcha 24), `scratchpadEvidenceComplete` (did a COMPLETE scratchpad
|
|
557
|
+
walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
|
|
558
|
+
`total_cost_usd` for the run — the authoritative single-run spend; NOT the same source as summing
|
|
559
|
+
`modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
|
|
560
|
+
`usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations`, `models`, `toolErrors`,
|
|
514
561
|
`redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
|
|
515
562
|
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
|
|
516
563
|
`context` (tools/mcpServers/availableSkills), `tasks`,
|
|
@@ -539,7 +586,7 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
539
586
|
cost spike from fan-out reads as `trace --view dispatches` (how many, which agent) against that model's
|
|
540
587
|
per-model usage — the harness doesn't line-item each sub-agent's tokens.
|
|
541
588
|
- **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
|
|
542
|
-
`tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input, with a
|
|
589
|
+
`tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input/scratchpad — the last being the agent's working area outside every user-visible root, with a
|
|
543
590
|
`trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
|
|
544
591
|
**Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
|
|
545
592
|
**absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
|
|
@@ -723,10 +770,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
723
770
|
isolated** (its own egress sidecar network + proxy, its own per-session run dir), so parallel records don't
|
|
724
771
|
cross-talk — the concurrency bound exists only for the Docker address pool + API rate limits, not correctness.
|
|
725
772
|
|
|
726
|
-
17. **Editing `scenarios/*.yaml`
|
|
727
|
-
|
|
728
|
-
|
|
729
|
-
|
|
773
|
+
17. **Editing `scenarios/*.yaml` does NOT change a plain `replay` — the WHOLE scenario is frozen, not just
|
|
774
|
+
`assert:`.** *Why:* a cassette captures every key (`lane:`, `fidelity:`, `baseline:`, `prompt:`, `skills:` …)
|
|
775
|
+
and `replay` evaluates all of them from that frozen copy — byte-deterministic, ignoring the working tree (so
|
|
776
|
+
a committed cassette can't silently re-interpret against an uncommitted YAML). **Only `assert:`
|
|
777
|
+
(+`expect_denied:`) can be opted back to disk; any other edited key reaches a replay only by re-recording.**
|
|
778
|
+
This is *loud* rather than a *silent* no-op: plain `replay` prints a `::notice::` when a sibling's
|
|
779
|
+
`assert:`/`prompt:` differs, and when the sibling **fails to load** at all (a typo'd or too-new key) — and
|
|
780
|
+
points you at the fix.
|
|
730
781
|
*Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
|
|
731
782
|
`--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
|
|
732
783
|
`prompt`/`answers`/`baseline`/`fidelity`/`skills`/`requires_capabilities` or the skill content (when a
|
|
@@ -767,11 +818,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
767
818
|
|
|
768
819
|
22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
|
|
769
820
|
`manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
|
|
770
|
-
assertion keys
|
|
771
|
-
|
|
772
|
-
|
|
773
|
-
|
|
774
|
-
|
|
821
|
+
assertion keys. The linter is **static**: it never reads your cassettes, so it cannot know whether
|
|
822
|
+
yours already carry an `artifacts` manifest and `controlOut` (a current cassette does). On a healthy
|
|
823
|
+
fleet every one of those lines is a false alarm. One exception: `manifest-needs-snapshot` is
|
|
824
|
+
suppressed for `user_visible_artifact` on `lane: remote` — the only manifest-backed key that lane also
|
|
825
|
+
rejects outright (`lane-remote-incompatible-key`, an ERROR), so the INFO would be redundant advice
|
|
826
|
+
about a key the scenario can never even load with. `gate-needs-controlout` has no such exception. *Fix:*
|
|
827
|
+
`lint --min-severity WARN` in CI (≥1.11.0) — the INFO advisories stay one flag away for interactive use.
|
|
828
|
+
`--strict --min-severity ERROR` behaves as a plain lint, not a contradiction.
|
|
775
829
|
23. **`verify-cassettes`/`replay` report a `discovery-surface` note on cassettes you just recorded fine.**
|
|
776
830
|
*Why:* the cassette froze its `system/init` tool inventory from before the skills/plugins discovery
|
|
777
831
|
servers existed at that tier (added 1.10.0). It is a non-gating **note**, never a finding — it cannot
|
|
@@ -779,6 +833,18 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
779
833
|
`mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
|
|
780
834
|
re-recording would never produce those tools anyway.
|
|
781
835
|
|
|
836
|
+
24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
|
|
837
|
+
lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
|
|
838
|
+
emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
|
|
839
|
+
Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
|
|
840
|
+
optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
|
|
841
|
+
other — and probing a remote session makes this harness look like it emulates the wrong tool under the
|
|
842
|
+
wrong schema. It doesn't; the lanes genuinely disagree. *Fix:* describe the **outcome** ("deliver the
|
|
843
|
+
file to the user") and let the model pick its surface's tool. The `no_scratchpad_leak` /
|
|
844
|
+
`present_files_called` assertion keys are harness-side names and stay valid either way.
|
|
845
|
+
([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
|
|
846
|
+
→ "File delivery" has the binary-verified detail; repo-only.)
|
|
847
|
+
|
|
782
848
|
For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
|
|
783
849
|
files in `references/` (the gotchas above are the full list; the references repeat only the
|
|
784
850
|
assertion/replay-relevant ones).
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.15.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -58,7 +58,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
58
58
|
GitHub-hosted runners, no token/Docker/agent:
|
|
59
59
|
|
|
60
60
|
```yaml
|
|
61
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
61
|
+
- run: npm i -g "cowork-harness@>=1.15.0"
|
|
62
62
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
63
63
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
64
64
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -107,8 +107,10 @@ The split is not just about tokens — it decides **where each lane can run**:
|
|
|
107
107
|
one (see [docs/cassette.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md)). This is your **always-on PR gate**. With `--output-format json`, read each
|
|
108
108
|
`results[].verdict.{pass,signals}` for **per-cassette** pass/fail and the reason (the top-level `ok`
|
|
109
109
|
collapses the whole batch) — e.g. a `stalled` signal fails a cassette whose assertions all passed.
|
|
110
|
-
The PR gate evaluates the
|
|
111
|
-
`assert:`
|
|
110
|
+
The PR gate evaluates the **whole scenario frozen in the cassette** (`lane:`/`fidelity:`/`baseline:` too,
|
|
111
|
+
not only `assert:`) — a YAML edit cannot move this gate's verdict. To re-check against an edited on-disk
|
|
112
|
+
`assert:` without a paid re-record, use `replay --assert-from <scenario.yaml>` (or `--reassert`); any other
|
|
113
|
+
edited key needs a re-record.
|
|
112
114
|
- **`run` / `record` (live).** Spawns the real agent in a sandbox: real model tokens + Docker **+ the
|
|
113
115
|
staged Claude Code agent ELF**, bind-mounted from a local Claude Desktop install or pointed to via
|
|
114
116
|
`COWORK_AGENT_BINARY`. Nothing is bundled, and **the agent binary is not redistributable** — a clean
|
|
@@ -219,10 +221,10 @@ jobs:
|
|
|
219
221
|
steps:
|
|
220
222
|
- uses: actions/checkout@v4
|
|
221
223
|
- uses: actions/setup-node@v4
|
|
222
|
-
with: { node-version: '
|
|
224
|
+
with: { node-version: '24' }
|
|
223
225
|
- uses: actions/setup-python@v5
|
|
224
226
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
225
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
227
|
+
- run: npm i -g "cowork-harness@>=1.15.0"
|
|
226
228
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
227
229
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
228
230
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -251,7 +253,7 @@ jobs:
|
|
|
251
253
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
252
254
|
fi
|
|
253
255
|
- if: steps.guard.outputs.live == 'true'
|
|
254
|
-
run: npm i -g "cowork-harness@>=1.
|
|
256
|
+
run: npm i -g "cowork-harness@>=1.15.0"
|
|
255
257
|
- if: steps.guard.outputs.live == 'true'
|
|
256
258
|
run: cowork-harness run scenarios/ --output-format json
|
|
257
259
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 1.
|
|
3
|
+
Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -30,7 +30,8 @@ turn rows carry `critiqueRole:"task"` / `"reflection"`.
|
|
|
30
30
|
|
|
31
31
|
Two things a harvester needs: roll-ups are excluded from `stats` aggregation (they carry no verdict —
|
|
32
32
|
counting them adds a phantom run and drags `passRate` toward 1), so **filter them out of any pass-rate
|
|
33
|
-
computed over raw rows
|
|
33
|
+
computed over raw rows** — that exclusion also governs `stats --group-by skill-hash`, so a per-generation
|
|
34
|
+
**total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
|
|
34
35
|
UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
|
|
35
36
|
|
|
36
37
|
## The report's item shape — no `title`, no `summary`
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -211,14 +211,18 @@ up often enough to spell out:
|
|
|
211
211
|
(`staleness[]` flags skill/baseline drift as a hint; only a live `run` re-confirms current
|
|
212
212
|
behavior). See [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § "Still skipped on replay" and [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) § "Which
|
|
213
213
|
assertions survive replay."
|
|
214
|
-
-
|
|
215
|
-
`present_files_called` check the `present_files` delivery path
|
|
216
|
-
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
214
|
+
- **`present_files` assertions can't verify off the tiers that serve the tool.** `no_scratchpad_leak` and
|
|
215
|
+
`present_files_called` check the `present_files` delivery path — the desktop-local lane's tool; remote
|
|
216
|
+
Cowork delivers via the agent-native `SendUserFile` instead, so never hardcode a delivery tool name in
|
|
217
|
+
a SKILL.md (SKILL.md Gotcha 24). **The harness** serves `present_files` on `container` **and `hostloop`**
|
|
218
|
+
— not `microvm`/`protocol`. `present_files_called` works at both; `no_scratchpad_leak` stays
|
|
219
|
+
`container`-only, because hostloop's handler passes a validated path through without promoting, so
|
|
220
|
+
there is no scratch→outputs copy to leak. Asserting either key off the tiers that serve it hard-fails at
|
|
221
|
+
runtime (a red run, not a false green), so you won't be fooled if you write the assertion. The quieter
|
|
222
|
+
trap is a scenario that runs at `microvm`/`protocol` and simply omits these assertions: a green run
|
|
223
|
+
there proves **nothing** about scratchpad-leak safety or present_files delivery, because that tier
|
|
224
|
+
never exercises the delivery path the assertions would check. Use `fidelity: container` or `hostloop`
|
|
225
|
+
for present_files-delivery coverage; `no_scratchpad_leak` still needs `container` specifically.
|
|
222
226
|
- **The harness doesn't observe rendered-artifact interactions, browser downloads, or human
|
|
223
227
|
clicks.** It runs the agent headless — no webview, no browser, no person clicking "Submit." A
|
|
224
228
|
class of Cowork bug (a client-side write-back to a relative URL that resolves-but-fails against
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.15.0`
|
|
4
4
|
(baseline `desktop-1.24012.9`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -37,6 +37,20 @@ fidelity: container # protocol | container | microvm | hostloop
|
|
|
37
37
|
execution: local # OPTIONAL — orthogonal to fidelity (a privilege/sandbox tier, all
|
|
38
38
|
# local): local (default) | cloud-describe (RESERVED — no runner
|
|
39
39
|
# exists yet; authoring it is a load-time error, not a silent no-op)
|
|
40
|
+
lane: local # OPTIONAL — which Cowork lane's DELIVERY CONTRACT the run is held to:
|
|
41
|
+
# local (default) | remote. On `remote`, location delivers NOTHING (a
|
|
42
|
+
# remote container has no auto-delivering outputs dir and is reclaimed
|
|
43
|
+
# at session end) and `present_files` is NOT served (a local MCP server
|
|
44
|
+
# can't reach a remote session) — so `user_visible_artifact` and
|
|
45
|
+
# present_files_called/no_scratchpad_leak are REJECTED AT SCENARIO-LOAD
|
|
46
|
+
# TIME (before the run starts, before any spend), not left to fail
|
|
47
|
+
# unverifiable/can't-verify at assertion time. Orthogonal to fidelity
|
|
48
|
+
# and execution — a `lane: remote` scenario still runs locally.
|
|
49
|
+
# Delivery semantics only; the remote device bridge is deliberately
|
|
50
|
+
# unmodeled.
|
|
51
|
+
# NEEDS >= 1.14.0: on an older CLI a scenario carrying `lane:` does NOT
|
|
52
|
+
# load (`Unrecognized key: "lane"`, exit 2) — it is NOT reinterpreted as
|
|
53
|
+
# `lane: local`. Adopting the key means raising your floor.
|
|
40
54
|
on_unanswered: fail # policy for unscripted gates: fail | prompt | first | llm — run rejects prompt
|
|
41
55
|
# ("agent" is retired — no longer a valid value)
|
|
42
56
|
|
|
@@ -298,8 +312,8 @@ same set live from the schema.
|
|
|
298
312
|
| `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
|
|
299
313
|
| `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
|
|
300
314
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
301
|
-
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature).
|
|
302
|
-
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container`
|
|
315
|
+
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
316
|
+
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
303
317
|
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose text matches |
|
|
304
318
|
| `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
|
|
305
319
|
| `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min` to also require a gate |
|
|
@@ -314,6 +328,7 @@ same set live from the schema.
|
|
|
314
328
|
| `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
|
|
315
329
|
| `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
|
|
316
330
|
| `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider) |
|
|
331
|
+
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
317
332
|
| `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
|
|
318
333
|
| `egress_denied: <host>` | the host was blocked by the egress proxy |
|
|
319
334
|
| `egress_allowed: <host>` | the host was allowed through |
|
|
@@ -358,6 +373,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
358
373
|
| `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
|
|
359
374
|
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
|
|
360
375
|
| `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
|
|
376
|
+
| `undelivered_deliverables` | warn | The skill produced file(s) OUTSIDE every user-visible root and never delivered them — on a remote Cowork session the workspace is reclaimed at session end; on a local one they stay invisible to the user. Fires on every run without opting in, because the scenarios that most need it are the ones whose author never considered delivery. Silent when the evidence cannot answer the question (no workspace walk, a tier that runs no scratchpad walk, absent delivery telemetry, or a resumed turn) — never a vacuous clean. Opt out: `allow_undelivered_deliverables` |
|
|
361
377
|
| `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
|
|
362
378
|
|
|
363
379
|
A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
|
|
@@ -369,9 +385,12 @@ overall run verdict and exit code — `assert result: success` alone won't catch
|
|
|
369
385
|
A cassette (`record`/`replay`) has **no filesystem and no network**. `replay` re-evaluates only the
|
|
370
386
|
**content** assertions. The authoritative list is `ALWAYS_CONTENT_KEYS`/`QUESTION_GATE_KEYS`/`MANIFEST_KEYS` (composed) in `src/run/cassette.ts`.
|
|
371
387
|
|
|
372
|
-
**
|
|
373
|
-
|
|
374
|
-
|
|
388
|
+
**Scenario source — the WHOLE scenario is frozen; only `assert:` can be opted back to disk.** A cassette
|
|
389
|
+
captures every key (`name`/`prompt`/`session`/`baseline`/`fidelity`/`lane`/`skills`/`answers`/`execution`/
|
|
390
|
+
`requires_capabilities`/`expect_denied`/`assert`), and a plain `replay` evaluates **all** of them from that
|
|
391
|
+
frozen copy (byte-deterministic, ignores the working tree); editing `scenarios/<name>.yaml` does not change
|
|
392
|
+
it — replay only prints a `::notice::` when a sibling's `assert:`/`prompt:` differs, or when it fails to
|
|
393
|
+
load at all. An edited `lane:`/`fidelity:`/`baseline:` reaches a replay ONLY by re-recording. `replay --assert-from
|
|
375
394
|
<scenario.yaml>` / `--reassert` is the opt-in token-free re-check against the on-disk block; it **hard-fails**
|
|
376
395
|
on recording-shaping drift (`prompt`/`answers`/`baseline`/`skills`) or skill-content staleness (it implies
|
|
377
396
|
`--fail-on-skill-drift`). `expect_denied`/filesystem/egress keys are sourced from on-disk but stay live-only —
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 1.
|
|
5
|
+
Tracks `cowork-harness 1.15.0` (baseline `desktop-1.24012.9`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -118,7 +118,8 @@ have:
|
|
|
118
118
|
4. **Replay-class note:** `max_turns` / `tool_calls_max` / `dispatch_count_max` re-evaluate
|
|
119
119
|
token-free on replay; `questions_count_max` needs `controlOut` (Recipe 1's tree applies).
|
|
120
120
|
`max_cost_usd` / `max_tokens` on replay assert the FROZEN recording's spend — near-zero signal
|
|
121
|
-
as a regression gate; if you need a cost gate, put it on the live lane via `stats
|
|
121
|
+
as a regression gate; if you need a cost gate, put it on the live lane via `stats` (history) or
|
|
122
|
+
`--max-budget-usd` (a pre-flight refusal before the run spends).
|
|
122
123
|
|
|
123
124
|
## Recipe 5 — Evaluate your skill's ANSWER QUALITY (semantic regression gate)
|
|
124
125
|
|
|
@@ -174,7 +175,14 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
174
175
|
output. **Reproduce before acting on a finding:** `cowork-harness skill <folder> "<prompt>" --repeat 5 --label gen-1`
|
|
175
176
|
runs the same skill+prompt N times (2-100) and prints a variance rollup instead of a single pass/fail —
|
|
176
177
|
`--repeat` works on the `skill` lane, not just `run`. A single green run proves it passed *once*.
|
|
177
|
-
Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`.
|
|
178
|
+
Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`.
|
|
179
|
+
(`--max-budget-usd` also works WITHOUT `--repeat`, where it pre-flights a single run against that
|
|
180
|
+
scenario's cost history and refuses before spending; the others are batch-only.
|
|
181
|
+
**`record` accepts it too** — there it is CUMULATIVE across the batch (a `dir/` or `--rerecord-stale`
|
|
182
|
+
sweep is the least predictable spend in the CLI), refusing up front if the summed history exceeds the
|
|
183
|
+
cap. At `--concurrency 1` a running total also stops the batch once the cap is reached; above that the
|
|
184
|
+
stop is disabled and the tool says so, because with N runs in flight the total is only known after an
|
|
185
|
+
overshoot is already paid for. Unpriced scenarios contribute $0 and are named as a LOWER BOUND.) It rejects `--session-id`/
|
|
178
186
|
`--resume` (both pin one run dir) and `--decider-cmd`/`--decider-dir` (a driving agent x N is not a
|
|
179
187
|
measurement).
|
|
180
188
|
The full loop (harvest -> reproduce -> fix -> prove freshness -> compare) is written out end-to-end in
|
|
@@ -6,6 +6,7 @@
|
|
|
6
6
|
"allow_missing_capability",
|
|
7
7
|
"allow_permissive_auto_allow",
|
|
8
8
|
"allow_stall",
|
|
9
|
+
"allow_undelivered_deliverables",
|
|
9
10
|
"artifact_json",
|
|
10
11
|
"compaction_occurred",
|
|
11
12
|
"computer_links_resolve",
|
|
@@ -80,6 +81,7 @@
|
|
|
80
81
|
"execution",
|
|
81
82
|
"expect_denied",
|
|
82
83
|
"fidelity",
|
|
84
|
+
"lane",
|
|
83
85
|
"name",
|
|
84
86
|
"on_unanswered",
|
|
85
87
|
"prompt",
|
|
@@ -92,6 +94,7 @@
|
|
|
92
94
|
"allow_l0_plugin_divergence",
|
|
93
95
|
"allow_missing_capability",
|
|
94
96
|
"allow_permissive_auto_allow",
|
|
95
|
-
"allow_stall"
|
|
97
|
+
"allow_stall",
|
|
98
|
+
"allow_undelivered_deliverables"
|
|
96
99
|
]
|
|
97
100
|
}
|