cowork-harness 4.4.1 → 4.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +8 -7
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +2 -2
- package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
- package/.claude/skills/cowork-harness/references/authoring.md +10 -10
- package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/debugging.md +6 -8
- package/.claude/skills/cowork-harness/references/eval.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +7 -3
- package/.claude/skills/cowork-harness/references/gotchas.md +3 -3
- package/.claude/skills/cowork-harness/references/hillclimb-recipe.md +1 -1
- package/.claude/skills/cowork-harness/references/hillclimb.md +1 -1
- package/.claude/skills/cowork-harness/references/measurement.md +1 -1
- package/.claude/skills/cowork-harness/references/run-record-replay.md +3 -3
- package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
- package/.claude/skills/cowork-harness/references/semantic-judging.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +123 -1
- package/.claude/skills/cowork-harness/scripts/scenario.py +1 -1
- package/CHANGELOG.md +122 -0
- package/DESIGN.md +6 -6
- package/README.md +10 -14
- package/baselines/desktop-2.19675.1.json +1458 -0
- package/baselines/desktop-2.26454.0.json +1492 -0
- package/dist/cli.js +10 -2
- package/dist/decide/llm-transport.js +32 -9
- package/dist/hostloop/workspace-handler.js +3 -1
- package/dist/run/chat.js +3 -0
- package/dist/run/doctor.js +52 -3
- package/dist/run/execute.js +11 -4
- package/dist/run/lane-notice.js +72 -0
- package/dist/run/verdict.js +1 -1
- package/dist/runtime/hostloop.js +2 -2
- package/dist/session.js +15 -2
- package/dist/sync/baseline-diff.js +62 -2
- package/dist/sync/cowork-sync.js +203 -6
- package/dist/sync/remote-devices.js +430 -0
- package/dist/types.js +22 -4
- package/docs/ci.md +2 -2
- package/docs/cli.md +6 -5
- package/docs/companion-skill.md +2 -2
- package/docs/fidelity-gaps.md +73 -32
- package/docs/invariants.md +1 -0
- package/docs/maintenance.md +32 -8
- package/docs/plugin-root.md +4 -0
- package/docs/scenario.md +137 -16
- package/docs/session.md +5 -2
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/package.json +1 -1
- package/schema/run-result.json +1 -1
- package/schema/scenario.schema.json +3 -3
- package/scripts/check-versions.ts +27 -2
- package/scripts/release-preflight.ts +5 -3
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did it make answers worse? (`eval`: paired A/B, pinned models) — or improving one round by round (`hillclimb`, the `/claude-api hillclimb` runner). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval / hillclimb commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 4.
|
|
7
|
-
tracks-harness: cowork-harness 4.
|
|
6
|
+
version: 4.5.0
|
|
7
|
+
tracks-harness: cowork-harness 4.5.0 (baseline desktop-2.26454.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -26,8 +26,8 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
|
|
|
26
26
|
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
27
|
Read them.
|
|
28
28
|
|
|
29
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.
|
|
30
|
-
> `desktop-2.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.5.0` (baseline
|
|
30
|
+
> `desktop-2.26454.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
31
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
32
32
|
|
|
33
33
|
## Preflight — make sure the harness can actually run
|
|
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
43
43
|
|
|
44
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
45
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
46
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.5.0"`. **Pin `@^4.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
47
47
|
|
|
48
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
49
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -130,7 +130,8 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
|
|
|
130
130
|
recording the cassette that locks it.
|
|
131
131
|
7. **The tier decides what exists.** `protocol` has no sandbox and no egress, tool names differ per tier
|
|
132
132
|
(`container` serves `mcp__workspace__web_fetch`, not `WebFetch`), and every tier models Cowork's
|
|
133
|
-
|
|
133
|
+
local lane only; new Pro and Max tasks do not use it from 2026-10-06. Behaviour-shaped results transfer;
|
|
134
|
+
path, mount, delivery and egress results do not.
|
|
134
135
|
8. **A WARN signal never blocks a green.** Read the verdict signals after every run
|
|
135
136
|
(`prompt_asset_missing`, `undelivered_deliverables`, `model_fallback`, …).
|
|
136
137
|
|
|
@@ -156,7 +157,7 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
|
|
|
156
157
|
| [`references/measurement.md`](references/measurement.md) | `--repeat`, `--ablate-skill`, measurement hygiene |
|
|
157
158
|
| [`references/debugging.md`](references/debugging.md) | triage, `result.json` fields and `trace` views, `chat` |
|
|
158
159
|
| [`references/gotchas.md`](references/gotchas.md) | the full "✓ passed ≠ correct" landmine catalog |
|
|
159
|
-
| [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality |
|
|
160
|
+
| [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality, and goals with no flag (force a compaction, ablate a section, a form reply, resume in a new conversation, hook JSON decisions, schema checks, unattended runs) |
|
|
160
161
|
| [`references/assertion-catalog.md`](references/assertion-catalog.md) | every `assert:` key's semantics, the verdict-signal table |
|
|
161
162
|
| [`references/semantic-judging.md`](references/semantic-judging.md) | `semantic_matches` in full: what the judge reads, fork results, refusal reasons, provenance |
|
|
162
163
|
| [`references/scenario-schema.md`](references/scenario-schema.md) | every YAML field, which keys survive `replay`, the `web_fetch` model |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertion catalog
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Every `assert:` key with its semantics, and the
|
|
4
4
|
verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
|
|
5
5
|
the scenario and session YAML fields are there too.
|
|
6
6
|
|
|
@@ -79,7 +79,7 @@ same set live from the schema.
|
|
|
79
79
|
| `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
|
|
80
80
|
| `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
|
|
81
81
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
82
|
-
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, an empty macOS system dir from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the
|
|
82
|
+
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, an empty macOS system dir from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
83
83
|
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
84
84
|
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
85
85
|
| `question_options: {when_question?, equals?: [..], contains?: [..], order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertions guide
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
|
|
4
4
|
|
|
5
5
|
### Assertions: two orthogonal axes
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Authoring a scenario
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
|
|
4
4
|
|
|
5
5
|
## Part I — AUTHOR a scenario
|
|
6
6
|
|
|
@@ -11,8 +11,10 @@ provenance, and the scaffold/lint tools that keep the YAML honest.
|
|
|
11
11
|
### Two files: session vs scenario
|
|
12
12
|
|
|
13
13
|
- **`sessions/*.yaml`** — pre-prompt setup: `model`, mounts (`folders`), and discovery
|
|
14
|
-
(marketplaces / plugins / skills / mcp). One session is reused by many scenarios. A scenario
|
|
15
|
-
|
|
14
|
+
(marketplaces / plugins / skills / mcp). One session is reused by many scenarios. A scenario's
|
|
15
|
+
`session:` is a **path** to such a file, never a nested block: `session: { plugins: … }` fails to load. A
|
|
16
|
+
scenario that omits `session:` gets an all-defaults session with no plugin declared, so a scenario that tests a
|
|
17
|
+
plugin needs a session file.
|
|
16
18
|
- **`scenarios/*.yaml`** — the test: `prompt`, scripted `answers:`, and `assert:`.
|
|
17
19
|
|
|
18
20
|
This split matters: release ground truth (`baseline:` / `baselines/`, produced by `sync`) is
|
|
@@ -52,18 +54,16 @@ Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejec
|
|
|
52
54
|
(it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
|
|
53
55
|
`references/fidelity-and-answers.md`.
|
|
54
56
|
|
|
55
|
-
**Every tier models Cowork's
|
|
57
|
+
**Every tier models Cowork's LOCAL lane** — agent on the user's machine, shell rooted at
|
|
56
58
|
`/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
|
|
57
59
|
**remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
|
|
58
60
|
different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
|
|
59
|
-
reproduces it and none can — that container is not something a local tool can stand up.
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
[announces](https://support.claude.com/en/articles/15520349-use-claude-cowork-on-web-desktop-and-mobile) that new
|
|
63
|
-
tasks run in the cloud from 2026-10-06. Check the session's own lane
|
|
61
|
+
reproduces it and none can — that container is not something a local tool can stand up.
|
|
62
|
+
[From 2026-10-06 new Pro and Max tasks run in the cloud](https://support.claude.com/en/articles/15520349-use-claude-cowork-on-web-desktop-and-mobile);
|
|
63
|
+
before then no setting reliably decided the lane. Check the session's own lane
|
|
64
64
|
([how](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md#which-lane-a-session-actually-ran-on)).
|
|
65
65
|
So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
|
|
66
|
-
asserting a **path, mount or
|
|
66
|
+
asserting a **path, mount, delivery mechanism or egress rule** is a claim about the local lane only. Declare
|
|
67
67
|
`lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
|
|
68
68
|
rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
|
|
69
69
|
in `run-record-replay.md`).
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "4.
|
|
20
|
+
(e.g. `version: "4.5.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.289 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
|
|
41
41
|
# served from .../claude-code-releases/rc/<commit>/. For some versions the stable path 404s
|
|
42
42
|
# (2.1.255); for others it returns 200 and serves a DIFFERENT BUILD UNDER THE SAME VERSION
|
|
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
82
82
|
GitHub-hosted runners, no token/Docker/agent:
|
|
83
83
|
|
|
84
84
|
```yaml
|
|
85
|
-
- run: npm i -g "cowork-harness@^4.
|
|
85
|
+
- run: npm i -g "cowork-harness@^4.5.0"
|
|
86
86
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
87
87
|
# no silent false-greens. WITHOUT --strict this
|
|
88
88
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -408,7 +408,7 @@ jobs:
|
|
|
408
408
|
with: { node-version: '24' }
|
|
409
409
|
- uses: actions/setup-python@v5
|
|
410
410
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
411
|
-
- run: npm i -g "cowork-harness@^4.
|
|
411
|
+
- run: npm i -g "cowork-harness@^4.5.0"
|
|
412
412
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
413
413
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
414
414
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -437,7 +437,7 @@ jobs:
|
|
|
437
437
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
438
438
|
fi
|
|
439
439
|
- if: steps.guard.outputs.live == 'true'
|
|
440
|
-
run: npm i -g "cowork-harness@^4.
|
|
440
|
+
run: npm i -g "cowork-harness@^4.5.0"
|
|
441
441
|
- if: steps.guard.outputs.live == 'true'
|
|
442
442
|
run: cowork-harness run scenarios/ --output-format json
|
|
443
443
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Debugging a run
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
|
|
4
4
|
|
|
5
5
|
## Part III — Debug
|
|
6
6
|
|
|
@@ -94,16 +94,14 @@ decide which assertions from *Assertions: two orthogonal axes* in `assertions-gu
|
|
|
94
94
|
bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
|
|
95
95
|
the current list rather than relying on a fixed enumeration here.
|
|
96
96
|
- **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
|
|
97
|
-
is held to.
|
|
98
|
-
|
|
99
|
-
[
|
|
100
|
-
|
|
97
|
+
is held to. `local` is the default and models the local lane, which new Pro and Max tasks do not use from
|
|
98
|
+
2026-10-06 (they run in the cloud); the lanes disagree about what *delivered* means. A live `local` run with an
|
|
99
|
+
environment-shaped assertion (a path, mount, delivery or egress key) prints one `[lane]` line on stderr saying
|
|
100
|
+
so, once per process; `--compact`/`--demo`, `CI` and `COWORK_HARNESS_NO_LANE_NOTICE=1` silence it. On `remote`,
|
|
101
101
|
location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
|
|
102
102
|
session end), `present_files` is NOT served, and `user_visible_artifact` /
|
|
103
103
|
`present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
|
|
104
|
-
to check a skill's delivery survives the cloud lane
|
|
105
|
-
[announces](https://support.claude.com/en/articles/15520349-use-claude-cowork-on-web-desktop-and-mobile) that
|
|
106
|
-
from 2026-10-06 new Pro and Max tasks run in the cloud). Orthogonal to `fidelity` — a `lane: remote`
|
|
104
|
+
to check a skill's delivery survives the cloud lane, where new Pro and Max tasks run from 2026-10-06. Orthogonal to `fidelity` — a `lane: remote`
|
|
107
105
|
scenario still runs locally.
|
|
108
106
|
- **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
|
|
109
107
|
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). The full guide is
|
|
4
4
|
[docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
|
|
5
5
|
need while running it.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -207,7 +207,11 @@ permissive behaviour is deliberately what the scenario is about.
|
|
|
207
207
|
|
|
208
208
|
A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
|
|
209
209
|
a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
|
|
210
|
-
current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
|
|
210
|
+
current baseline, not what each scenario pins, so it can report ready seconds before the run fails. Its
|
|
211
|
+
staged-agent row ends with what Desktop has staged against that pin: staged and pinned; a newer agent staged
|
|
212
|
+
(upgrade cowork-harness, or run `cowork-harness sync` if you maintain the baseline); this Desktop older than the
|
|
213
|
+
pin; the pinned agent not staged (staging may be withheld, or no task has booted the VM
|
|
214
|
+
since the update); or none found.
|
|
211
215
|
|
|
212
216
|
To keep the exact pin, **recover the pinned ELF**: re-download that version from the release channel,
|
|
213
217
|
check its sha256 against the baseline's, and set `COWORK_AGENT_BINARY` to it
|
|
@@ -298,7 +302,7 @@ up often enough to spell out:
|
|
|
298
302
|
behavior). See [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § "Still skipped on replay" and [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) § "Which
|
|
299
303
|
assertions survive replay."
|
|
300
304
|
- **`present_files` assertions can't verify off the tiers that serve the tool.** `no_scratchpad_leak` and
|
|
301
|
-
`present_files_called` check the `present_files` delivery path — the
|
|
305
|
+
`present_files_called` check the `present_files` delivery path — the local lane's tool (not used by new Pro and Max tasks from 2026-10-06); remote
|
|
302
306
|
Cowork delivers via the agent-native `SendUserFile` instead, so never hardcode a delivery tool name in
|
|
303
307
|
a SKILL.md (Gotcha 24 in `gotchas.md`). **The harness** serves `present_files` on `container` **and `hostloop`**
|
|
304
308
|
— not `microvm`/`protocol`. `present_files_called` works at both; `no_scratchpad_leak` stays
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Gotchas
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). The full "✓ passed ≠ correct" landmine catalog.
|
|
4
4
|
|
|
5
5
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
6
6
|
|
|
@@ -270,8 +270,8 @@ authorable). Reach for this list when debugging a run's behavior, that one while
|
|
|
270
270
|
`--output-format json`). *Fix:* re-record against the pinned agent.
|
|
271
271
|
|
|
272
272
|
24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
|
|
273
|
-
lane, and an agent only sees the one for the surface it is on. The
|
|
274
|
-
emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
|
|
273
|
+
lane, and an agent only sees the one for the surface it is on. The local-lane sandbox this harness
|
|
274
|
+
emulates (not used by new Pro and Max tasks from 2026-10-06) is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
|
|
275
275
|
Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
|
|
276
276
|
optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
|
|
277
277
|
other — and probing a remote session makes this harness look like it emulates the wrong tool under the
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Recipe 7 — Climb a skill with `/claude-api hillclimb` and the harness as its runner
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). It needs a `cowork-harness` whose
|
|
4
4
|
`hillclimb --help` lists `--skill` (help goes to stderr). This page is the loop's procedure, step by step, in the order of the
|
|
5
5
|
`/claude-api hillclimb` guide. Every mechanic (flags, refusals, the gate, `regrade`, `freeze-ref`, exit codes,
|
|
6
6
|
row keys) is in [`hillclimb.md`](hillclimb.md); the setup and the full list of differences from the guide's own
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `hillclimb` — the runner for a `/claude-api hillclimb` loop
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). It needs a `cowork-harness` whose
|
|
4
4
|
`hillclimb --help` lists `--skill` (help goes to stderr). The command reference is
|
|
5
5
|
[docs/cli.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md); this is the part a loop needs
|
|
6
6
|
while it runs. It covers `run`, `check`, `state-template`, `freeze-ref` and `regrade`.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Measurement
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
|
|
4
4
|
|
|
5
5
|
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Run, record and lock
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
|
|
4
4
|
|
|
5
5
|
## Part II — RUN, RECORD & LOCK
|
|
6
6
|
|
|
@@ -220,7 +220,7 @@ Recognize these before "fixing" a non-bug:
|
|
|
220
220
|
`markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
|
|
221
221
|
(`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
|
|
222
222
|
Desktop `2.9939.2` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
|
|
223
|
-
behind that sentence;
|
|
223
|
+
behind that sentence; 5 baselines have shipped since without a re-capture). The message says so ("likely a FALSE
|
|
224
224
|
NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
|
|
225
225
|
`COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
|
|
226
226
|
`allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
|
|
@@ -248,7 +248,7 @@ Recognize these before "fixing" a non-bug:
|
|
|
248
248
|
scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
|
|
249
249
|
**The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
|
|
250
250
|
can see them, and give the file tools an **absolute** path under the outputs directory the agent's prompt
|
|
251
|
-
names. On the
|
|
251
|
+
names. On the local lane's host loop (what Desktop ran when measured), against Desktop **2.7032.0 and later**,
|
|
252
252
|
the agent process runs outside the session (`/var/empty`), so a relative `Read`/`Write`/`Edit` — a bare
|
|
253
253
|
filename or `outputs/x.md` alike — is **refused** ("File is in a directory that is denied by your
|
|
254
254
|
permission settings."); only a pathless or relative `Grep`/`Glob` is redirected to outputs. (Before
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, replay class, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.
|
|
4
|
-
(baseline `desktop-2.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.5.0`
|
|
4
|
+
(baseline `desktop-2.26454.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` and `fidelity` are required:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `semantic_matches` — the full semantics
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). The one-line summary is in
|
|
4
4
|
[assertion-catalog.md](assertion-catalog.md); this is the whole contract, split out so the catalog stays within the
|
|
5
5
|
agent's single-read size.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 4.
|
|
5
|
+
Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -321,3 +321,125 @@ load-bearing gates with `--answer` once you know which fire.
|
|
|
321
321
|
`critique` finds what is wrong; `skill`/`run` checks it works. The loop's procedure, with `hillclimb run` as its
|
|
322
322
|
runner, is its own page: [`hillclimb-recipe.md`](hillclimb-recipe.md). The command reference is
|
|
323
323
|
[`hillclimb.md`](hillclimb.md).
|
|
324
|
+
|
|
325
|
+
## Recipe 8 — Goals the harness has no flag for
|
|
326
|
+
|
|
327
|
+
Also in [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md#recipes-for-goals-the-harness-has-no-flag-for).
|
|
328
|
+
Each recipe below uses only shipped flags and keys. Each says what it does **not** prove.
|
|
329
|
+
|
|
330
|
+
### Force a context compaction
|
|
331
|
+
|
|
332
|
+
Pin a session, run the task, compact it by hand, check that turn, then continue:
|
|
333
|
+
|
|
334
|
+
```bash
|
|
335
|
+
cowork-harness skill ./my-plugin "<the task>" --session-id compact-1
|
|
336
|
+
cowork-harness skill ./my-plugin "/compact" --session-id compact-1 --resume
|
|
337
|
+
jq -e '[.contextEvents[]? | select(.subtype=="compact_boundary")] | length > 0' <run-dir>/turns/2/result.json
|
|
338
|
+
cowork-harness skill ./my-plugin "<continue the task>" --session-id compact-1 --resume
|
|
339
|
+
cowork-harness trace <run-dir> # shows the latest turn: what the continued task did
|
|
340
|
+
```
|
|
341
|
+
|
|
342
|
+
The run dir is the one each turn's `[status]` line prints. A resumed session's dir holds one `turns/<n>/` per
|
|
343
|
+
turn, and `verify-run` refuses a dir with more than one turn, so the `compaction_occurred` assert cannot be checked
|
|
344
|
+
on it: read the `/compact` turn's own `result.json` instead (`jq` exits `0` when it recorded a compaction, `1` when it
|
|
345
|
+
did not). *Does not prove:* that the skill behaves as it would after an automatic compaction. A manual
|
|
346
|
+
`/compact` may not re-attach skills exactly as autocompact does (not verified), and re-attached skill text can come
|
|
347
|
+
back truncated, so a long `SKILL.md` may not return whole.
|
|
348
|
+
|
|
349
|
+
### Ablate one `SKILL.md` section
|
|
350
|
+
|
|
351
|
+
Copy the plugin, delete the section from the copy, and run the two as `eval` arms.
|
|
352
|
+
Size it first, at no cost:
|
|
353
|
+
|
|
354
|
+
```bash
|
|
355
|
+
cp -R ./my-plugin /tmp/nosec && $EDITOR /tmp/nosec/skills/<skill>/SKILL.md # remove the section
|
|
356
|
+
cowork-harness eval scenarios/ --arm full=./my-plugin --arm nosec=/tmp/nosec --dry-run --target-effect 30
|
|
357
|
+
```
|
|
358
|
+
|
|
359
|
+
*Does not prove:* which section drove a given action; it shows only whether removing it changes the graded outcome.
|
|
360
|
+
|
|
361
|
+
### Test a skill's parsing of a Desktop form reply
|
|
362
|
+
|
|
363
|
+
Desktop's elicitation form sends its answers as the next user
|
|
364
|
+
message, as one line. Send that line as a resumed turn: `cowork-harness skill ./my-plugin "<the reply line>" --session-id s --resume`. The format,
|
|
365
|
+
from Desktop's own form guide:
|
|
366
|
+
|
|
367
|
+
- one line: `<Title> details — Label: value · Label: value`, labels being the form's field names in sentence case;
|
|
368
|
+
- a multi-select value comma-joined; a short multi-line value flattened with ` / `; a value of 81–200 characters
|
|
369
|
+
in quotes;
|
|
370
|
+
- a value over 200 characters shown as `Label: (N chars — see below)`, and repeated in full after a
|
|
371
|
+
`--- Full content ---` line;
|
|
372
|
+
- a skipped form arrives as one fixed sentence saying it was skipped.
|
|
373
|
+
|
|
374
|
+
*Does not prove:* that the model would choose the form (the harness serves no `visualize` tools; see
|
|
375
|
+
[fidelity-gaps.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md#skill-argument-collection--the-elicitation-form-branch-is-not-reachable-here)),
|
|
376
|
+
or that a file the form attaches arrives.
|
|
377
|
+
|
|
378
|
+
### Resume the work in a new conversation
|
|
379
|
+
|
|
380
|
+
Keep the first run, export its outputs, and start a second scenario from
|
|
381
|
+
them:
|
|
382
|
+
|
|
383
|
+
```bash
|
|
384
|
+
cowork-harness run step1.yaml --keep
|
|
385
|
+
cowork-harness fixture export <run-dir> --out scenarios/step1-out
|
|
386
|
+
```
|
|
387
|
+
|
|
388
|
+
```yaml
|
|
389
|
+
# scenarios/step2.yaml
|
|
390
|
+
workspace_fixture: step1-out
|
|
391
|
+
assert:
|
|
392
|
+
- file_exists: {path: outputs/next.md, authored: true} # this conversation wrote it
|
|
393
|
+
- file_exists: {path: outputs/brief.md, authored: false} # carried over, not rewritten
|
|
394
|
+
```
|
|
395
|
+
|
|
396
|
+
Record step 2 with `--out` inside the same tree as the fixture: `record` refuses a fixture outside the cassette's git repository (outside git, outside the cassette's directory).
|
|
397
|
+
*Does not prove:* how Cowork treats a new task over the same folder on either lane (not verified); the second
|
|
398
|
+
conversation starts with no memory of the first.
|
|
399
|
+
|
|
400
|
+
### Assert a hook's JSON decision
|
|
401
|
+
|
|
402
|
+
A hook that decides by printing JSON and exiting 0 is not a block to
|
|
403
|
+
`hook_event_blocked`, which counts exit code 2 only. Read the hook's output, and what the agent got back:
|
|
404
|
+
|
|
405
|
+
```yaml
|
|
406
|
+
assert:
|
|
407
|
+
- hook_output_contains: {event: PreToolUse, stream: stdout, matches: '"permissionDecision"\s*:\s*"deny"'}
|
|
408
|
+
- tool_result_contains: "blocked by policy"
|
|
409
|
+
# updatedInput: the recorded call input is what the model sent; the rewrite shows in the paired result
|
|
410
|
+
- tool_called: {tool: Bash, result: {matches: 'outputs/archive/'}}
|
|
411
|
+
- hook_event_blocked: Stop # a Stop hook that blocks with exit 2
|
|
412
|
+
```
|
|
413
|
+
|
|
414
|
+
*Does not prove:* which tool a hook frame was about (frames do not name it), or that the model read the reason.
|
|
415
|
+
|
|
416
|
+
### Schema-check a written file
|
|
417
|
+
|
|
418
|
+
In the Python lane, pass a `jsonschema` check as the predicate:
|
|
419
|
+
|
|
420
|
+
```python
|
|
421
|
+
import jsonschema
|
|
422
|
+
def valid(doc):
|
|
423
|
+
jsonschema.validate(doc, SCHEMA) # raises with the failing path
|
|
424
|
+
return True
|
|
425
|
+
result.assert_artifact_json("outputs/cap.json", valid)
|
|
426
|
+
```
|
|
427
|
+
|
|
428
|
+
In a scenario, name the exact paths (`artifact_json` per file) and add `no_unexpected_files` so no other file slips
|
|
429
|
+
in. *Does not prove:* anything about a file whose name you did not list (no globbing).
|
|
430
|
+
|
|
431
|
+
### Hold a skill to an unattended host
|
|
432
|
+
|
|
433
|
+
Make any question fail the run, and give the answers in the prompt:
|
|
434
|
+
|
|
435
|
+
```yaml
|
|
436
|
+
fidelity: container
|
|
437
|
+
on_unanswered: fail # the default for `run`; say it anyway
|
|
438
|
+
prompt: "Build the weekly report. Assume: region = all, format = markdown; ask nothing."
|
|
439
|
+
assert:
|
|
440
|
+
- questions_count_max: 0
|
|
441
|
+
```
|
|
442
|
+
|
|
443
|
+
Leave `allow_stall` out, so ending on a question fails. `trace <run> --view questions` shows who answered each gate
|
|
444
|
+
(`answeredBy`). *Does not prove:* scheduled-task behaviour. Real scheduled tasks remove `AskUserQuestion` entirely and
|
|
445
|
+
tell the model no user is present; the harness models neither.
|
|
@@ -1383,7 +1383,7 @@ def lint_doc(doc, path, raw_lines, cassette_records=None):
|
|
|
1383
1383
|
"no `fidelity:` — the key is required (since 4.0.0); `run`, `record` and the loader "
|
|
1384
1384
|
"refuse this scenario.",
|
|
1385
1385
|
"Add a tier: `fidelity: container` keeps the pre-4.0 behaviour (it was the default) and "
|
|
1386
|
-
"models the VM-LOOP lane;
|
|
1386
|
+
"models the VM-LOOP lane; Cowork's local lane runs HOST-LOOP by default (gate 1143815894), so "
|
|
1387
1387
|
"`fidelity: hostloop` matches it and `fidelity: cowork` auto-picks the way Cowork does. "
|
|
1388
1388
|
"Switching tiers can COST you assertions: `no_scratchpad_leak` is container-only (an "
|
|
1389
1389
|
"error elsewhere) and `transcript_no_host_path` fails by design at hostloop/protocol. "
|