cowork-harness 2.4.0 → 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +8 -8
- package/.claude/skills/cowork-harness/references/ci-recipe.md +17 -17
- package/.claude/skills/cowork-harness/references/critique.md +72 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +6 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +20 -5
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +5 -2
- package/.claude/skills/cowork-harness/scripts/scenario.py +67 -2
- package/CHANGELOG.md +322 -0
- package/DESIGN.md +2 -2
- package/README.md +17 -17
- package/SPEC.md +14 -2
- package/baselines/desktop-1.40609.0.json +878 -0
- package/baselines/provisioning/rootfs-provisioning.json +32 -40
- package/dist/assert.js +42 -1
- package/dist/cli.js +16 -1
- package/dist/critique/command.js +202 -22
- package/dist/critique/evaluator.js +9 -3
- package/dist/critique/package-evidence.js +39 -5
- package/dist/run/cassette.js +27 -4
- package/dist/run/chat-result.js +2 -1
- package/dist/run/chat.js +15 -1
- package/dist/run/execute.js +46 -7
- package/dist/run/hook-events.js +51 -0
- package/dist/run/probe-dispatch.js +8 -2
- package/dist/run/provenance.js +6 -9
- package/dist/run/run.js +197 -8
- package/dist/run/skill-flag-surface.js +11 -0
- package/dist/run/tier-vacuous-tools.js +69 -0
- package/dist/run/tool-name-canonicalization.js +70 -0
- package/dist/run/verdict.js +17 -8
- package/dist/runtime/argv.js +16 -1
- package/dist/runtime/lima.js +37 -3
- package/dist/runtime/protocol.js +85 -13
- package/dist/scan.js +1 -0
- package/dist/sync/cowork-sync.js +39 -2
- package/dist/types.js +47 -6
- package/docs/cassette.md +4 -2
- package/docs/critique.md +52 -0
- package/docs/debugging.md +3 -1
- package/docs/discovery.md +1 -1
- package/docs/fidelity-gaps.md +78 -6
- package/docs/maintenance.md +1 -1
- package/docs/scenario.md +14 -8
- package/docs/subagents.md +5 -3
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/examples/sessions/l0-plugin-delivery.yaml +6 -0
- package/package.json +1 -1
- package/python/test_scenario_lint.py +1 -1
- package/schema/critique-report.json +41 -18
- package/schema/run-result.json +49 -3
- package/schema/scenario.schema.json +19 -5
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version:
|
|
7
|
-
tracks-harness: cowork-harness
|
|
6
|
+
version: 3.0.0
|
|
7
|
+
tracks-harness: cowork-harness 3.0.0 (baseline desktop-1.40609.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness
|
|
29
|
-
> `desktop-1.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.0` (baseline
|
|
29
|
+
> `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,13 +42,13 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.0"`. **Pin `@^3.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
49
49
|
— upgrade rather than work around it, since this skill's `file:line` pointers and flag names track the floor.
|
|
50
50
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
51
|
-
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
51
|
+
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
|
|
52
52
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
53
53
|
- **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
|
|
54
54
|
|
|
@@ -147,7 +147,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
|
|
|
147
147
|
| Tier | What it gives you | Use when |
|
|
148
148
|
|---|---|---|
|
|
149
149
|
| `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
|
|
150
|
-
| `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm
|
|
150
|
+
| `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
|
|
151
151
|
| `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
|
|
152
152
|
| `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
|
|
153
153
|
|
|
@@ -679,7 +679,7 @@ than a stuck `"running"`.)
|
|
|
679
679
|
### Place assertions in the right CI lane
|
|
680
680
|
|
|
681
681
|
CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
|
|
682
|
-
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@
|
|
682
|
+
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
|
|
683
683
|
PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
|
|
684
684
|
the four-stage pipeline.
|
|
685
685
|
|
|
@@ -1,23 +1,23 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
7
7
|
|
|
8
8
|
```yaml
|
|
9
|
-
- uses: yaniv-golan/cowork-harness@
|
|
9
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
10
10
|
with:
|
|
11
11
|
command: replay
|
|
12
12
|
path: cassettes/
|
|
13
|
-
version: "^
|
|
13
|
+
version: "^3" # hold the major; see below
|
|
14
14
|
```
|
|
15
15
|
|
|
16
|
-
**These recipes pin `version: "^
|
|
16
|
+
**These recipes pin `version: "^3"`.** The Action's `version` input *defaults* to `latest`, which means a
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "
|
|
20
|
+
(e.g. `version: "3.0.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.247 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
41
41
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
42
42
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -47,11 +47,11 @@ jobs:
|
|
|
47
47
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
48
48
|
# Background on the provenance chain: the "Agent-binary provenance" section of
|
|
49
49
|
# https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
|
|
50
|
-
- uses: yaniv-golan/cowork-harness@
|
|
50
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
51
51
|
with:
|
|
52
52
|
command: run
|
|
53
53
|
path: scenarios/
|
|
54
|
-
version: "^
|
|
54
|
+
version: "^3"
|
|
55
55
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
56
56
|
```
|
|
57
57
|
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.0.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -119,18 +119,18 @@ Action has no input for, and it creates a coupling nothing checks:
|
|
|
119
119
|
you have a reason:
|
|
120
120
|
|
|
121
121
|
```yaml
|
|
122
|
-
- uses: yaniv-golan/cowork-harness@
|
|
122
|
+
- uses: yaniv-golan/cowork-harness@v3
|
|
123
123
|
with:
|
|
124
124
|
command: lint
|
|
125
125
|
path: scenarios/
|
|
126
|
-
version: "^
|
|
127
|
-
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any
|
|
126
|
+
version: "^3" # holds the major
|
|
127
|
+
extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 3.x satisfies that
|
|
128
128
|
```
|
|
129
129
|
|
|
130
130
|
**If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
|
|
131
131
|
floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
|
|
132
132
|
recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
|
|
133
|
-
major instead — `version: "^
|
|
133
|
+
major instead — `version: "^3"`, which is what the steps above use — keeping the floor's intent while
|
|
134
134
|
stopping at the major boundary. An exact
|
|
135
135
|
pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
|
|
136
136
|
moment a recipe adopts a newer flag.
|
|
@@ -147,7 +147,7 @@ The split is not just about tokens — it decides **where each lane can run**:
|
|
|
147
147
|
Docker, no agent binary** — runs on a stock GitHub Actions runner. Evaluates **content** assertions —
|
|
148
148
|
`transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
|
|
149
149
|
`max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
|
|
150
|
-
`allow_permissive_auto_allow` / `allow_missing_capability` / `
|
|
150
|
+
`allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
|
|
151
151
|
`allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
|
|
152
152
|
`question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
153
153
|
(`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
|
|
@@ -322,7 +322,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
|
|
|
322
322
|
|
|
323
323
|
## GitHub Actions sketch
|
|
324
324
|
|
|
325
|
-
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@
|
|
325
|
+
The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v3` does
|
|
326
326
|
in one step (see the top of this doc) — reach for this form when you need independent per-command
|
|
327
327
|
gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
|
|
328
328
|
equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
|
|
@@ -342,7 +342,7 @@ jobs:
|
|
|
342
342
|
with: { node-version: '24' }
|
|
343
343
|
- uses: actions/setup-python@v5
|
|
344
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^
|
|
345
|
+
- run: npm i -g "cowork-harness@^3.0.0"
|
|
346
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +371,7 @@ jobs:
|
|
|
371
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
372
|
fi
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^
|
|
374
|
+
run: npm i -g "cowork-harness@^3.0.0"
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness
|
|
3
|
+
Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -34,6 +34,63 @@ computed over raw rows** — that exclusion also governs `stats --group-by skill
|
|
|
34
34
|
**total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
|
|
35
35
|
UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
|
|
36
36
|
|
|
37
|
+
## An exit-2 report — which turn failed, and whether it was really infrastructure
|
|
38
|
+
|
|
39
|
+
Exit 2 means no findings were produced, so the report's only job is to say what went wrong. Three fields
|
|
40
|
+
carry that; read all three before touching anything.
|
|
41
|
+
|
|
42
|
+
| Field | Meaning |
|
|
43
|
+
|---|---|
|
|
44
|
+
| `infraFailure` | the reason |
|
|
45
|
+
| `infraFailurePhase` | `task turn` (the graded run) or `reflection turn` (critique's own protocol turn) |
|
|
46
|
+
| `infraFailureKind` | why it failed — a harness `ErrCategory` (error envelope, exit 2/3) **or** a `resultErrorKind` (`usage_limit`/`transport`/`agent`) from a turn that RAN and errored (exit 1, top-level `error: null`). **Absent** = killed, or no envelope |
|
|
47
|
+
| `gradedErrorReason` | on a `taskResult: "error"` run (still gradeable, exit 0): why the GRADED turn errored, so a quota exhaustion is not read as a skill defect |
|
|
48
|
+
|
|
49
|
+
**Do NOT read "has a kind" as "the instrument is fine".** The CLI's top-level catch turns every
|
|
50
|
+
unexpected throw into category **`internal`** — Docker down, container start failure, missing staged
|
|
51
|
+
agent, harness bug — and `runtime` carries a refused run dir. Only **`unanswered`, `usage`, `boundary`**
|
|
52
|
+
and, from the result-row taxonomy, **`usage_limit`** (quota exhausted — retry after reset) and
|
|
53
|
+
**`transport`** (a tail-end drop) are the caller's problem. `agent` is not: for critique's own protocol
|
|
54
|
+
turn that IS the instrument breaking. The header encodes exactly that split and fails closed (an
|
|
55
|
+
unrecognized kind renders as infrastructure):
|
|
56
|
+
|
|
57
|
+
- `RUN FAILED (<turn>, <kind>): …` → ordinary, actionable, instrument healthy.
|
|
58
|
+
- `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>): …` → `internal`/`runtime`/`agent`/unknown/killed/no-envelope.
|
|
59
|
+
|
|
60
|
+
**A turn that exits 1 is not a crash.** It RAN and reported an errored result, with `error: null` and a
|
|
61
|
+
full `results[0]` — an exit-code-only reading of that path is what leaves an exhausted quota looking like
|
|
62
|
+
a broken instrument.
|
|
63
|
+
|
|
64
|
+
**Read the reason, not the category.** The reason carries the failed turn's own message *and* hint
|
|
65
|
+
verbatim. That matters most for `unanswered`, which is 36 distinct throw sites and only ONE of them is
|
|
66
|
+
"the skill asked an unscripted question" — the others are a mis-typed `--answer` label, malformed
|
|
67
|
+
`--answer-policy` YAML, a crashed or bad-JSON `--decider-cmd` helper, an out-of-set `--decider-llm`
|
|
68
|
+
reply, an unanswered dialog/elicit, even a self-declared harness bug. A remedy picked from the category
|
|
69
|
+
is wrong for nearly all of them; each site's own hint is written for its case. (Also note `--on-unanswered`
|
|
70
|
+
*conflicts* with `--decider-dir`/`--decider-cmd`, so it is not a blanket fallback.)
|
|
71
|
+
|
|
72
|
+
For the genuine unscripted-gate case: script it (`--answer`, `--answer-policy`), or, when the skill's
|
|
73
|
+
gates are LLM-authored and reworded every run so a literal regex will not match twice, use `--decider-llm`
|
|
74
|
+
(or the scenario's `on_unanswered: llm`).
|
|
75
|
+
|
|
76
|
+
## Which model was graded — `gradedModels`
|
|
77
|
+
|
|
78
|
+
**The two turns are a SUBPROCESS.** They inherit no model from whatever invoked `critique` — not your
|
|
79
|
+
session, not a project setting. With no `--model`, the graded run uses the spawned agent's own default,
|
|
80
|
+
which may not be the model you are otherwise working under, and nothing about the run announces it.
|
|
81
|
+
|
|
82
|
+
`gradedModels` (text header: `graded model(s):`) is read back from the graded turn's own `result.json` and
|
|
83
|
+
is the only record of which model produced the behaviour being graded — distinct from the evaluator's
|
|
84
|
+
resolved model, which is a **different workload with its own default** (`claude-opus-4-8`), reported
|
|
85
|
+
separately. An evaluator line naming a model you did not pass is therefore expected, not evidence your
|
|
86
|
+
`--model` was ignored. Pin with `--model <id>` whenever a critique will be compared against another, and
|
|
87
|
+
read `gradedModels` back to confirm it took.
|
|
88
|
+
|
|
89
|
+
It is **observed, not requested** — the ids come from the model stamped on the graded turn's assistant
|
|
90
|
+
messages, never from the flag. So `graded model(s): unknown` means no assistant message reached the run
|
|
91
|
+
(crash, kill, or a gate before the first reply); passing `--model` does not change that line. Past runs
|
|
92
|
+
can be checked without re-running: the same ids are in each kept run dir's `turns/1/result.json`.
|
|
93
|
+
|
|
37
94
|
## The report's item shape — no `title`, no `summary`
|
|
38
95
|
|
|
39
96
|
Each `items[]` entry's prose fields are **`idea`** and **`recommendedAction`** — there is no `title` field
|
|
@@ -80,6 +137,20 @@ false `already-covered` verdict. It is named in `corpusExcluded` instead, and an
|
|
|
80
137
|
specifically reports `skillMdStatus: "untracked"`, forcing the mechanical `already-covered` →
|
|
81
138
|
`not-adjudicable` downgrade. `git add` it (or commit before critiquing) if it should count as evidence.
|
|
82
139
|
|
|
140
|
+
## Read `referencesAccessed`, not `referencesRead`
|
|
141
|
+
|
|
142
|
+
`referencesRead` counts the **`Read` tool only**. An agent that reaches a reference with a `Bash cat`, a
|
|
143
|
+
`Grep` or a `Glob` leaves nothing in it, so its emptiness is **not** evidence the content went unread.
|
|
144
|
+
`referencesAccessed` is the wide signal — every file reached, with the channel each was reached through
|
|
145
|
+
(`read` / `grep` / `bash`) — and it is what the critique headline is computed from.
|
|
146
|
+
|
|
147
|
+
Two properties to carry: only the `read` channel is strong evidence the agent opened the file (a `bash`
|
|
148
|
+
entry means a command named the path); and detection **under-approximates** — a `cd` into the skill dir
|
|
149
|
+
then a bare relative `cat`, a heredoc body and a `$VAR`-built path are all invisible. So an absent path is
|
|
150
|
+
weak evidence, never proof. **Presence is the cannot-verify channel:** `[]` means the drive ran and saw
|
|
151
|
+
nothing (a real negative); an ABSENT field means there was no observable drive, and must never be read as
|
|
152
|
+
"none".
|
|
153
|
+
|
|
83
154
|
## `referencesRead` is main-agent-only — `noSkillFilesRead` is not
|
|
84
155
|
|
|
85
156
|
`result.json`'s top-level `referencesRead` lists **main-agent Reads only**. A dispatcher-style skill does
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -20,6 +20,11 @@ Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.379
|
|
|
20
20
|
- A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
|
|
21
21
|
true` — with no container around the native file tools, that combination gives the agent genuine,
|
|
22
22
|
software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
|
|
23
|
+
- A `protocol` scenario staging a plugin that declares runnable hooks needs `allow_host_hooks: true`
|
|
24
|
+
(`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
|
|
25
|
+
hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
|
|
26
|
+
Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
|
|
27
|
+
does not fall back to the default.
|
|
23
28
|
- **Set the tier in the scenario's `fidelity:` field — not a flag.** `--fidelity` is accepted only by
|
|
24
29
|
`skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
|
|
25
30
|
positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.0`
|
|
4
|
+
(baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -98,6 +98,17 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
|
|
|
98
98
|
# around hostloop's native file tools, that combination gives the
|
|
99
99
|
# agent genuine, software-checked-only host filesystem access.
|
|
100
100
|
# Read-only folders and folder-less runs need no opt-in.
|
|
101
|
+
|
|
102
|
+
allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
|
|
103
|
+
# declares runnable hooks (`<plugin>/hooks/hooks.json`): L0 passes
|
|
104
|
+
# --plugin-dir, so the CLI executes those hooks as NATIVE HOST
|
|
105
|
+
# processes under your account, with no container sandbox. A plugin
|
|
106
|
+
# that declares no hooks needs no opt-in, and a misplaced root-level
|
|
107
|
+
# `hooks.json` cannot execute so it does not trigger the gate.
|
|
108
|
+
# Use `--fidelity container` to run them sandboxed instead.
|
|
109
|
+
# NEEDS cowork-harness >= 3.0.0. The loader is a strict object, so an
|
|
110
|
+
# OLDER CLI does not default it — it hard-errors
|
|
111
|
+
# `Unrecognized key: "allow_host_hooks"` and exits 2.
|
|
101
112
|
```
|
|
102
113
|
|
|
103
114
|
Relative paths resolve from the file's own directory, so a scenario + session + referenced files
|
|
@@ -305,6 +316,10 @@ same set live from the schema.
|
|
|
305
316
|
| `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
|
|
306
317
|
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
|
|
307
318
|
| `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
|
|
319
|
+
| **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
|
|
320
|
+
| **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
|
|
321
|
+
| `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
|
|
322
|
+
| `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
|
|
308
323
|
| `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
|
|
309
324
|
| `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
|
|
310
325
|
| `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
|
|
@@ -358,7 +373,7 @@ same set live from the schema.
|
|
|
358
373
|
| `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
|
|
359
374
|
| `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
|
|
360
375
|
| `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
|
|
361
|
-
| `
|
|
376
|
+
| `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
|
|
362
377
|
| `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
|
|
363
378
|
| `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
|
|
364
379
|
| `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
|
|
@@ -400,7 +415,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
400
415
|
| `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
|
|
401
416
|
| `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
|
|
402
417
|
| `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
|
|
403
|
-
| `
|
|
418
|
+
| `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
|
|
404
419
|
| `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
|
|
405
420
|
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
406
421
|
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
@@ -439,7 +454,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
|
|
|
439
454
|
`skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
|
|
440
455
|
`max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
|
|
441
456
|
(`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
|
|
442
|
-
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `
|
|
457
|
+
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
|
|
443
458
|
`allow_stall` are also kept on replay, evaluated as no-op passes.
|
|
444
459
|
|
|
445
460
|
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness
|
|
5
|
+
Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"keys": [
|
|
4
4
|
"all_tasks_completed",
|
|
5
5
|
"allow_delete_in",
|
|
6
|
-
"
|
|
6
|
+
"allow_l0_host_config_contamination",
|
|
7
7
|
"allow_missing_capability",
|
|
8
8
|
"allow_outputs_delete",
|
|
9
9
|
"allow_permissive_auto_allow",
|
|
@@ -35,6 +35,7 @@
|
|
|
35
35
|
"no_hook_blocked",
|
|
36
36
|
"no_lost_write_back",
|
|
37
37
|
"no_mcp_error",
|
|
38
|
+
"no_observed_reference_access",
|
|
38
39
|
"no_path_denied",
|
|
39
40
|
"no_scratchpad_leak",
|
|
40
41
|
"no_skill_triggered",
|
|
@@ -46,6 +47,7 @@
|
|
|
46
47
|
"question_context",
|
|
47
48
|
"question_options",
|
|
48
49
|
"questions_count_max",
|
|
50
|
+
"reference_read",
|
|
49
51
|
"replay_protocol_fidelity",
|
|
50
52
|
"result",
|
|
51
53
|
"self_heal_ran",
|
|
@@ -81,6 +83,7 @@
|
|
|
81
83
|
"vm_path_denied"
|
|
82
84
|
],
|
|
83
85
|
"topLevelKeys": [
|
|
86
|
+
"allow_host_hooks",
|
|
84
87
|
"allow_host_writes",
|
|
85
88
|
"answers",
|
|
86
89
|
"assert",
|
|
@@ -99,7 +102,7 @@
|
|
|
99
102
|
],
|
|
100
103
|
"verdictModifierKeys": [
|
|
101
104
|
"allow_delete_in",
|
|
102
|
-
"
|
|
105
|
+
"allow_l0_host_config_contamination",
|
|
103
106
|
"allow_missing_capability",
|
|
104
107
|
"allow_outputs_delete",
|
|
105
108
|
"allow_permissive_auto_allow",
|
|
@@ -80,6 +80,10 @@ CONTENT_KEYS = {
|
|
|
80
80
|
"tool_result_not_matches",
|
|
81
81
|
"tool_called",
|
|
82
82
|
"tool_not_called",
|
|
83
|
+
# Replay re-derives these from the SAME frozen tool inputs the live run used (a cassette stores whole
|
|
84
|
+
# tool inputs), so they are content keys, not live-only.
|
|
85
|
+
"reference_read",
|
|
86
|
+
"no_observed_reference_access",
|
|
83
87
|
"subagent_tool_used",
|
|
84
88
|
"subagent_tool_absent",
|
|
85
89
|
"subagent_dispatched",
|
|
@@ -169,7 +173,7 @@ LANE_REMOTE_INCOMPATIBLE_KEYS = {"present_files_called", "no_scratchpad_leak", "
|
|
|
169
173
|
# verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
|
|
170
174
|
VERDICT_MODIFIER_KEYS = {
|
|
171
175
|
"allow_permissive_auto_allow",
|
|
172
|
-
"
|
|
176
|
+
"allow_l0_host_config_contamination",
|
|
173
177
|
"allow_missing_capability",
|
|
174
178
|
"allow_stall",
|
|
175
179
|
"allow_undelivered_deliverables",
|
|
@@ -258,7 +262,8 @@ _EMBEDDED_TOP_LEVEL_KEYS = {
|
|
|
258
262
|
"assert",
|
|
259
263
|
"skills", # opt-in skill-staleness hash scope
|
|
260
264
|
"requires_capabilities", # Fix 4b: scenario-level required-capability declaration (pre-flight gate)
|
|
261
|
-
"allow_host_writes",
|
|
265
|
+
"allow_host_writes",
|
|
266
|
+
"allow_host_hooks", # protocol consent: a staged plugin's hooks run as NATIVE HOST processes # hostloop native-split: consent for a writable connected folder (pre-run gate)
|
|
262
267
|
}
|
|
263
268
|
|
|
264
269
|
|
|
@@ -294,6 +299,8 @@ REGEX_KEYS = {
|
|
|
294
299
|
"hook_blocked",
|
|
295
300
|
"tool_result_matches",
|
|
296
301
|
"tool_result_not_matches",
|
|
302
|
+
"reference_read",
|
|
303
|
+
"no_observed_reference_access",
|
|
297
304
|
}
|
|
298
305
|
VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
|
|
299
306
|
VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
|
|
@@ -583,6 +590,44 @@ def lint_doc(doc, path, raw_lines):
|
|
|
583
590
|
# warns at run start, after authoring). Lint is deliberately STRICTER than the runtime: the docs
|
|
584
591
|
# declare the combination incompatible, so authoring it is a bug even if a tool-free run could
|
|
585
592
|
# accidentally pass. `cowork` gets a WARN naming the baseline-gate resolution dependency (the
|
|
593
|
+
# `tool_not_called` naming a tool the TIER does not serve can never be violated — it passes
|
|
594
|
+
# vacuously and verifies nothing. Expressible offline because the mapping is a harness constant
|
|
595
|
+
# (WORKSPACE_TOOL_ALIASES / VM_LOOP_TOOL_ALIASES), NOT a baseline read. Deliberately literals only,
|
|
596
|
+
# and deliberately a closed table: `--tools` gates the BUILT-IN set alone while every tier separately
|
|
597
|
+
# passes --mcp-config, so a session-MCP tool name is offered without appearing in any tool list and
|
|
598
|
+
# must never be flagged here. The harness refuses these at load; this catches them before any run.
|
|
599
|
+
_TIER_VACUOUS = {
|
|
600
|
+
"hostloop": {"Bash": "mcp__workspace__bash", "WebFetch": "mcp__workspace__web_fetch", "NotebookEdit": None},
|
|
601
|
+
"container": {"mcp__workspace__bash": "Bash"},
|
|
602
|
+
"microvm": {"mcp__workspace__bash": "Bash", "mcp__workspace__web_fetch": "WebFetch"},
|
|
603
|
+
# `protocol` absent on purpose: it passes no tool flags, so its surface is the operator's own
|
|
604
|
+
# host CLI registry — machine-dependent and about a different product.
|
|
605
|
+
}
|
|
606
|
+
# BOTH negative tool keys: `subagent_tool_absent` is judged against the tools sub-agents actually
|
|
607
|
+
# USED, not a per-dispatch declared list, so a tool the tier never serves makes it equally vacuous.
|
|
608
|
+
for _key in ("tool_not_called", "subagent_tool_absent"):
|
|
609
|
+
for _v in _assert_values(items, _key):
|
|
610
|
+
if not isinstance(_v, str) or "*" in _v or "?" in _v:
|
|
611
|
+
continue # a glob is not a literal claim about one tool
|
|
612
|
+
_repl = _TIER_VACUOUS.get(fidelity, {})
|
|
613
|
+
if _v not in _repl:
|
|
614
|
+
continue
|
|
615
|
+
_instead = _repl[_v]
|
|
616
|
+
findings.append(
|
|
617
|
+
Finding(
|
|
618
|
+
"WARN",
|
|
619
|
+
"tool-not-called-tier-vacuous",
|
|
620
|
+
f"`{_key}: {_v}` on `fidelity: {fidelity}` — that tier does not serve "
|
|
621
|
+
f"`{_v}` at all, so this can never be violated and verifies nothing.",
|
|
622
|
+
(
|
|
623
|
+
f"Assert `{_key}: {_instead}` instead — that is what the tier serves in its place."
|
|
624
|
+
if _instead
|
|
625
|
+
else f"The {fidelity} tier removes `{_v}` outright; drop this assertion."
|
|
626
|
+
),
|
|
627
|
+
path,
|
|
628
|
+
)
|
|
629
|
+
)
|
|
630
|
+
|
|
586
631
|
# linter stays offline — the message carries the gate fact instead of reading a baseline).
|
|
587
632
|
if "transcript_no_host_path" in assert_keys:
|
|
588
633
|
if fidelity in ("hostloop", "protocol"):
|
|
@@ -1016,6 +1061,26 @@ def lint_doc(doc, path, raw_lines):
|
|
|
1016
1061
|
)
|
|
1017
1062
|
)
|
|
1018
1063
|
|
|
1064
|
+
# Same shape for the reference-access pair: `reference_read: R` and `no_observed_reference_access: R`
|
|
1065
|
+
# with the IDENTICAL regex cannot both hold. Compared as raw pattern strings — two different regexes
|
|
1066
|
+
# that happen to match the same path are NOT a contradiction (the linter cannot know the paths), so
|
|
1067
|
+
# this only fires on the case that is unambiguously self-defeating.
|
|
1068
|
+
read_pats = {v for v in _assert_values(items, "reference_read") if isinstance(v, str)}
|
|
1069
|
+
unread_pats = {v for v in _assert_values(items, "no_observed_reference_access") if isinstance(v, str)}
|
|
1070
|
+
both_refs = sorted(read_pats & unread_pats)
|
|
1071
|
+
if both_refs:
|
|
1072
|
+
findings.append(
|
|
1073
|
+
Finding(
|
|
1074
|
+
"ERROR",
|
|
1075
|
+
"reference-access-contradiction",
|
|
1076
|
+
f"assert requires {both_refs} to be both accessed (`reference_read`) and never observed "
|
|
1077
|
+
"(`no_observed_reference_access`) — no run can satisfy that, so this would spend a run to fail.",
|
|
1078
|
+
"Drop whichever half the scenario does not mean. To check that ONE reference is reached while "
|
|
1079
|
+
"another is not, give the two keys different patterns.",
|
|
1080
|
+
path,
|
|
1081
|
+
)
|
|
1082
|
+
)
|
|
1083
|
+
|
|
1019
1084
|
# W: double-quoted regex with a backslash (raw-text scan — the parser already ate it)
|
|
1020
1085
|
findings.extend(_lint_regex_quoting(path, raw_lines))
|
|
1021
1086
|
|