cowork-harness 3.0.1 → 3.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.0.1
7
- tracks-harness: cowork-harness 3.0.1 (baseline desktop-1.40609.0)
6
+ version: 3.1.0
7
+ tracks-harness: cowork-harness 3.1.0 (baseline desktop-1.40609.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.1` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.1.0` (baseline
29
29
  > `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.1"`. **Pin `@^3.0.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.1.0"`. **Pin `@^3.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -568,6 +568,12 @@ Recognize these before "fixing" a non-bug:
568
568
  would claim more than the evidence supports, and staying silent would read as clean. Mutually exclusive
569
569
  with `undelivered_deliverables`, and quiet on a run that produced nothing to deliver. Not a skill defect —
570
570
  a harness coverage gap (see the *File delivery* section of fidelity-gaps).
571
+ - **`model_fallback`** (`WARN`) — the agent switched off the requested model mid-run. Read the `trigger`:
572
+ `model_not_found` / `model_blocked` / `permission_denied` are properties of the **pin**, so every run of
573
+ this scenario falls back the same way until you change the pinned id; `overloaded` / `server_error` are
574
+ transient and a re-run may hold. The run's assertions still mean what they say — but they were produced
575
+ by a different model than the scenario names, so treat a green as evidence about the fallback model.
576
+
571
577
  - **`mount_delete`** (`WARN`) — a delete touched a **delete-denied mount other than `outputs`**: a `rw`
572
578
  connected folder. Production denies `unlink`/`rmdir` on *every* Cowork FUSE mount until per-mount
573
579
  approval, not just outputs — a connected folder shows the identical default — so this run diverged from
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.0.1"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.1.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^3.0.1"
70
+ - run: npm i -g "cowork-harness@^3.1.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^3.0.1"
345
+ - run: npm i -g "cowork-harness@^3.1.0"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^3.0.1"
374
+ run: npm i -g "cowork-harness@^3.1.0"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.1`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.1.0`
4
4
  (baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -421,6 +421,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
421
421
  | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
422
422
  | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
423
423
  | `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
424
+ | `model_fallback` | warn | The agent fell back off the requested model mid-run (SDK `model_fallback` event); a `model_not_found`/`model_blocked` trigger repeats every run until the pin changes |
424
425
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
425
426
  | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
426
427
  | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -46,6 +46,11 @@ what production would do. Two layers of defense:
46
46
  statically knowable and only produce a non-failing informational note.
47
47
  - **Manual (one-liner):** `grep -h '"effectiveFidelity"' cassettes/*.cassette.json | sort | uniq -c`
48
48
  shows the tier distribution of the fleet at a glance.
49
+ - **Do NOT re-record for a `[note]` alone.** A `·`/`[note]` row is informational and the run still exits
50
+ **0** — `session-fingerprint: predates \`model\` coverage` and `prompt-assets: cassette predates …` both
51
+ mean "this cassette was recorded before that field existed, and everything the hash *does* cover still
52
+ matches". Re-recording buys the new coverage and nothing else, so on a fleet of heavy cassettes it is a
53
+ real bill for no verdict change. Re-record when a `✗` says to (baseline moved, skill drift, tier moved).
49
54
 
50
55
  ### Cassette anatomy (what you're looking at when you open one)
51
56
 
@@ -66,7 +71,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](htt
66
71
  | `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
67
72
  | `scenarioSource` | Relative path to the authored YAML this was recorded from |
68
73
  | `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
69
- | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
74
+ | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (model/folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
70
75
  | `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
71
76
  | `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
72
77
  | `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
@@ -241,10 +246,13 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
241
246
  unrecoverable if skipped. `skillHash` is content-exact, so an edit mid-batch silently splits the
242
247
  dataset into two generations — `stats --group-by skill-hash` separates them afterwards, but a hash
243
248
  whose source was never committed names a generation that is unrecoverable, which makes the
244
- comparison uninterpretable rather than merely noisy. And with no `model:` in the session (or
245
- `--model` on the `skill` lane) each run uses whatever the staged agent binary defaults to, so a
246
- before/after can silently straddle two models; read `result.json`'s `models` back to confirm —
247
- ignoring any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent fabricated locally, not a model).
249
+ comparison uninterpretable rather than merely noisy. And with no `model:` in the session and no
250
+ `--model` on the command (every lane takes it), each run uses whatever the staged agent binary
251
+ defaults to, so a before/after can silently straddle two models — the run warns when nothing pinned
252
+ one. Read `result.json` back to confirm: `modelSource` says whether anything pinned the model at all,
253
+ and `modelPinHonored` whether the pin survived (**absent means unverifiable, not "yes"**). `models`
254
+ lists what served the run — ignore any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent
255
+ fabricated locally, not a model).
248
256
  3. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
249
257
  `result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
250
258
  content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
package/CHANGELOG.md CHANGED
@@ -6,6 +6,66 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.1.0] — 2026-08-31
10
+
11
+ ### Fixed
12
+
13
+ - **The pre-coverage cassette note was a paragraph repeated per file.** It fires on every cassette
14
+ recorded before `model` joined the session shape — for a consumer with a dozen, a dozen identical
15
+ paragraphs burying the findings that are genuinely per-file. Cut to one line (370 chars to 157), matching
16
+ the prompt-assets note beside it; the reasoning stays in `docs/fidelity-gaps.md`. Notes are deliberately
17
+ **not** aggregated in `verify-cassettes` — it is a per-file audit, and collapsing them would destroy the
18
+ attribution — so length was the part to fix. The note has always been non-failing: a pre-coverage
19
+ cassette stays clean, and `replay --strict` exits 0.
20
+ - **The unpinned-model warning told `run` users the flag lived on other lanes.** It named the `skill`,
21
+ `probe-dispatch` and `chat` lanes for a flag `run` and `record` now accept — caught by running it, not
22
+ by reading it.
23
+
24
+ - **`run --model ""` started a spending run instead of failing.** SPEC §CB-2 requires an empty or
25
+ whitespace model value to be a usage error, never a silently-propagated empty string. `record` inherits
26
+ that from the shared flag parser; `run` hand-rolls its own loop (the session-less lanes take `--model`
27
+ out of the leftovers themselves, so the shared parser cannot consume it) and did not enforce it.
28
+
29
+ ### Added
30
+
31
+ - **`--model <id>` on `run` and `record`.** The `skill`, `probe-dispatch` and `chat` lanes already had it;
32
+ the two lanes that produce cassettes and CI verdicts did not, so the only way to pin them was editing the
33
+ session file. Precedence is explicit: an `--model` flag or a `--matrix` `models:` axis (a per-invocation
34
+ act) outranks the session's `model:`, while `COWORK_HARNESS_MODEL` only fills a gap — a machine-scoped
35
+ variable must never silently outrank a model the scenario declares, or the run's model becomes a
36
+ property of the shell it was launched from.
37
+ - **Model provenance on every run, and a warning when nothing pins the model.** `--model` reaches the
38
+ agent only when a session declares `model:`, so an unset session inherits whatever the local CLI would
39
+ pick — and because the agent selects part of its system prompt by model capability, that moves the
40
+ *instructions* it is given, not just answer quality. `--effort` is not symmetric: it is always emitted
41
+ from the baseline's `spawn.effortDefault`, so an unset session pinned the effort and left the model it
42
+ applies to floating. Runs now warn when no model resolves (`chat` in its own words), naming the key and
43
+ the flag. Omitting it is deprecated and becomes an error in the next major — the same path
44
+ `Scenario.fidelity` took, and deliberately not a hard fail today: the repo chose deprecation for a
45
+ default with a larger blast radius, and a harder gate for a lesser field would be incoherent with that.
46
+ - **A family pin is checked as family membership, not equality.** `--model opus` is resolved by the agent
47
+ to a concrete id (`claude-opus-5`), never echoed back, and which member an account resolves it to is not
48
+ something the harness can know — so comparing the two as strings would report a pinned-and-honored run as
49
+ `modelPinHonored: false`. `opus`/`sonnet`/`haiku`/`fable` now compare by family and still fail on a
50
+ wrong-family model; `best` and `opusplan` name no comparable model and stay unverifiable.
51
+ - **`RunResult.modelFallbacks`, `modelPinHonored` and `modelSource`.** Fallbacks come from the agent's own
52
+ `system`/`model_fallback` event rather than a diff of `models[]` — the event names the trigger, which is
53
+ what separates a retired or blocked pin (every run falls back the same way until the id changes) from a
54
+ transient overload. A new `model_fallback` warn signal surfaces it in the verdict; it cannot fire on a
55
+ healthy run, so it adds no volume to a currently-green one. `modelPinHonored` is three-state on purpose:
56
+ `true`, `false`, and **absent for unverifiable** — nothing pinned, or no live model evidence. The
57
+ natural two-state shape reports "we could not tell" as "the pin held", which is a false green, so the
58
+ no-evidence cases are asserted explicitly rather than left to a truthiness check.
59
+ - **A cassette records the model it was recorded with**, in its `environment` block: the id the agent
60
+ reported (so it survives a mid-run fallback) plus the source. A cassette recorded with nothing pinned
61
+ says `unresolved`, which is the difference between "recorded on opus-5 by choice" and "on whatever that
62
+ laptop happened to pick".
63
+ - **The session fingerprint now covers the pinned model.** Editing `model:` after recording previously
64
+ moved nothing and `verify-cassettes` said nothing, even though the edit changes what the agent was told.
65
+ A cassette recorded before this coverage reports **unverifiable** rather than clean — its hash cannot
66
+ tell "the field was never covered" from "the pin changed since", and claiming otherwise in the remedy
67
+ would reintroduce the false green the coverage exists to close.
68
+
9
69
  ## [3.0.1] — 2026-08-30
10
70
 
11
71
  ### Changed
package/README.md CHANGED
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
49
49
 
50
50
  | I want to… | Start here | Needs |
51
51
  |---|---|---|
52
- | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.0.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
- | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.0.1"` |
52
+ | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.1.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
+ | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.1.0"` |
54
54
  | **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
55
55
 
56
56
  Not sure a harness is what you need? The next two sections are the argument.
package/SPEC.md CHANGED
@@ -78,7 +78,8 @@ claude -p --verbose
78
78
  (--max-thinking-tokens 31999 | --thinking disabled) # session.extended_thinking on|off (default on)
79
79
  # debug.max_thinking_tokens → --max-thinking-tokens <N> (fenced, non-Cowork)
80
80
  [--append-system-prompt <rendered cowork sections>]
81
- [--model <session.model>]
81
+ [--model <resolved model>] # --model flag / matrix axis > session.model > COWORK_HARNESS_MODEL;
82
+ # omitted when none pin one (the agent picks its own default — warns)
82
83
  [--mcp-config <configGuest>/mcp.json] # if session.mcp.config set — HONORED in plain cowork mode (§6)
83
84
  (--plugin-dir <mntRoot>/<p>)… # one per pluginDirs entry
84
85
  --tools <baseline.spawn.tools…> # variadic, LAST
package/dist/cli.js CHANGED
@@ -90,6 +90,7 @@ const HELP = `cowork-harness <command> (v${"$VERSION"})
90
90
 
91
91
  ── Automated scenarios ────────────────────────────────────────────────────────
92
92
  run <scenario.yaml | dir/> run one scenario or every *.yaml in a dir (CI-ready exit code)
93
+ [--model <id>] pin the model (overrides the session's 'model:'; unset warns)
93
94
  [--on-unanswered fail|first] ('prompt' rejected — breaks determinism)
94
95
  [--decider-cmd '<helper>'] answer live questions via a spawned helper
95
96
  [--decider-dir <dir>] answer live questions in-band; then use 'gates'/'answer' to stream/respond
@@ -97,7 +98,7 @@ const HELP = `cowork-harness <command> (v${"$VERSION"})
97
98
  (run 'run --help' for the full flag reference)
98
99
 
99
100
  ── Cassette lifecycle ─────────────────────────────────────────────────────────
100
- record <scenario.yaml> run + save a control-protocol cassette
101
+ record <scenario.yaml> run + save a control-protocol cassette [--model <id>]
101
102
  [--out <file>] cassette path (default: cassettes/<scenario-name>.cassette.json)
102
103
  [--max-artifact-bytes <n>] inline-body cap (default 65536 / $COWORK_HARNESS_MAX_ARTIFACT_BYTES)
103
104
  [--concurrency <N>] record a dir/ batch (or --rerecord-stale) N at a time (default 1; runs are
@@ -325,6 +326,12 @@ const RUN_HELP = `cowork-harness run <scenario.yaml | dir/>
325
326
  ('run' takes no --dry-run: 'record <file.yaml> --dry-run' checks that a scenario LOADS without
326
327
  spending; 'lint <file.yaml>' checks the assertion invariants. Both are token-free.)
327
328
 
329
+ Model:
330
+ --model <id> pin the model, overriding the session's 'model:' for this run
331
+ (env COWORK_HARNESS_MODEL sets a default; a --matrix 'models:' axis
332
+ wins over both). A run that resolves no model warns: omitting it is
333
+ deprecated and becomes an error in the next major.
334
+
328
335
  Input policy:
329
336
  --on-unanswered fail|first policy for an unscripted question (default: fail — deterministic).
330
337
  'prompt' is rejected (it would break reproducibility).
@@ -1358,7 +1365,48 @@ async function cmdRun(rawArgs) {
1358
1365
  // swallow a real arg) instead of the loud reject below, so muscle memory from `skill` doesn't error.
1359
1366
  // Note it so the no-effect is visible.
1360
1367
  const keepRequested = rawRest.includes("--keep");
1361
- const args = rawRest.filter((a) => a !== "--keep");
1368
+ // `--model <id>` — the one-off pin, matching the `skill`/`probe-dispatch`/`chat` lanes (and defaulting
1369
+ // from COWORK_HARNESS_MODEL the same way they do). Parsed HERE rather than in `takeCommonFlags`: those
1370
+ // lanes take `--model` out of the leftovers themselves, so consuming it in the shared parser would
1371
+ // strip it before their own loops ever see it. Extracted before the `--keep` filter so both survive.
1372
+ // The EXPLICIT flag and the env DEFAULT are kept apart deliberately. `run`/`record` resolve a session
1373
+ // FILE, and a machine-scoped env var must never silently outrank a model that file declares — that is
1374
+ // the very defect this whole change exists to close, and collapsing the two here reintroduced it
1375
+ // through the back door (a stray line in a repo `.env` would repoint every clone's runs). The flag is
1376
+ // an explicit per-invocation act and does outrank the file; the env var only fills a gap.
1377
+ let modelFlag;
1378
+ const withoutModel = [];
1379
+ for (let i = 0; i < rawRest.length; i++) {
1380
+ const tok = rawRest[i];
1381
+ // Accept `--model=<id>` too: `record` gets it free from the shared parser, and a form that works on
1382
+ // one lane and is reported as an unexpected argument on the other is just a trap.
1383
+ if (tok.startsWith("--model=")) {
1384
+ // SPEC §CB-2: an empty/whitespace value is a hard usage error, never a silently-propagated empty
1385
+ // model string — the same rule `flagValue()` and chat's own parser enforce.
1386
+ const v = tok.slice("--model=".length);
1387
+ if (v.trim() === "")
1388
+ fail("run", "usage", "--model requires a model id", undefined, isJsonOutput(rawArgs));
1389
+ modelFlag = v;
1390
+ continue;
1391
+ }
1392
+ if (tok === "--model") {
1393
+ const v = rawRest[i + 1];
1394
+ // `takeCommonFlags` has already consumed every flag it knows, so a `-`-prefixed or missing next
1395
+ // token means the value is absent. The `.yaml` check catches the case the `-` guard CANNOT see:
1396
+ // in `run --model --verbose x.yaml`, `--verbose` is consumed upstream, so this loop sees
1397
+ // `["--model","x.yaml"]` and would swallow the scenario path as a model id — then report "no
1398
+ // scenario path", blaming the one argument that was correct. No model id ends in .yaml/.yml.
1399
+ const looksLikeScenario = v !== undefined && /\.ya?ml$/i.test(v);
1400
+ // `v.trim() === ""` is SPEC §CB-2 (an empty model string must never propagate).
1401
+ if (v === undefined || v.trim() === "" || v.startsWith("-") || looksLikeScenario)
1402
+ fail("run", "usage", `--model requires a model id${v === undefined ? "" : ` (found \`${v}\`${looksLikeScenario ? ", which looks like the scenario path" : ""})`}`, undefined, isJsonOutput(rawArgs));
1403
+ modelFlag = v;
1404
+ i++;
1405
+ continue;
1406
+ }
1407
+ withoutModel.push(tok);
1408
+ }
1409
+ const args = withoutModel.filter((a) => a !== "--keep");
1362
1410
  if (keepRequested)
1363
1411
  log("note: `run` always keeps runs (under the runs root); --keep is a no-op here.");
1364
1412
  // A leftover global-only flag is NOT a positional. `takeCommonFlags` has already consumed every real
@@ -1460,7 +1508,9 @@ async function cmdRun(rawArgs) {
1460
1508
  const cellScenario = cell.axes.baseline !== undefined ? { ...scenario, baseline: cell.axes.baseline } : scenario;
1461
1509
  try {
1462
1510
  const session = applySessionOverrides(baseSession, {
1463
- model: cell.axes.model,
1511
+ // The cell axis is the more specific choice; `--model` is the fallback for a matrix whose
1512
+ // cells vary some OTHER axis and still want one pinned model across all of them.
1513
+ model: cell.axes.model ?? modelFlag,
1464
1514
  skillDirSubstitution: cell.axes.skillDir !== undefined
1465
1515
  ? [baseSession.plugins.local_plugins[0], resolvedSkillDirs.get(cell.axes.skillDir)]
1466
1516
  : undefined,
@@ -1584,7 +1634,10 @@ async function cmdRun(rawArgs) {
1584
1634
  policy,
1585
1635
  externalChannel,
1586
1636
  o,
1587
- extra: deciderModel ? { llmModel: deciderModel } : undefined,
1637
+ extra: {
1638
+ ...(deciderModel ? { llmModel: deciderModel } : {}),
1639
+ ...(modelFlag !== undefined ? { modelOverride: modelFlag } : {}),
1640
+ },
1588
1641
  }));
1589
1642
  continue;
1590
1643
  }
@@ -1605,7 +1658,10 @@ async function cmdRun(rawArgs) {
1605
1658
  policy,
1606
1659
  externalChannel,
1607
1660
  o,
1608
- extra: deciderModel ? { llmModel: deciderModel } : undefined,
1661
+ extra: {
1662
+ ...(deciderModel ? { llmModel: deciderModel } : {}),
1663
+ ...(modelFlag !== undefined ? { modelOverride: modelFlag } : {}),
1664
+ },
1609
1665
  rethrowUnanswered: true,
1610
1666
  }),
1611
1667
  onResult: (r) => results.push(r),