cowork-harness 2.2.0 → 2.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +43 -17
- package/.claude/skills/cowork-harness/references/ci-recipe.md +15 -13
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +14 -12
- package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +32 -1
- package/CHANGELOG.md +366 -0
- package/DESIGN.md +3 -3
- package/README.md +41 -11
- package/baselines/desktop-1.37937.1.json +816 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +27 -0
- package/dist/assert.js +105 -3
- package/dist/cli.js +15 -4
- package/dist/hostloop/workspace-handler.js +12 -2
- package/dist/run/cassette.js +304 -78
- package/dist/run/execute.js +135 -8
- package/dist/run/trace-view.js +39 -1
- package/dist/runtime/container.js +53 -2
- package/dist/runtime/hostloop.js +52 -8
- package/dist/sync/cowork-sync.js +233 -16
- package/dist/types.js +40 -10
- package/docs/README.md +3 -0
- package/docs/cassette.md +36 -11
- package/docs/debugging.md +25 -0
- package/docs/fidelity-gaps.md +133 -6
- package/docs/maintenance.md +15 -1
- package/docs/scenario.md +66 -17
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +41 -46
- package/examples/replays/example-pdf-skill.cassette.json +164 -115
- package/examples/replays/hostloop-computer-links.cassette.json +62 -68
- package/examples/scenarios/example-pdf-skill.yaml +8 -1
- package/llms.txt +3 -1
- package/package.json +1 -1
- package/python/test_scenario_lint.py +39 -0
- package/schema/scenario.schema.json +29 -10
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 2.
|
|
7
|
-
tracks-harness: cowork-harness 2.
|
|
6
|
+
version: 2.4.0
|
|
7
|
+
tracks-harness: cowork-harness 2.4.0 (baseline desktop-1.37937.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -16,14 +16,17 @@ in the shell; this skill tells you how to author scenarios, pick a fidelity tier
|
|
|
16
16
|
path, place assertions in the right CI lane, and avoid the harness's "✓ passed ≠ actually correct"
|
|
17
17
|
traps.
|
|
18
18
|
|
|
19
|
+
`cowork-harness` is an unofficial, independent project — not affiliated with or endorsed by
|
|
20
|
+
Anthropic. Say so if a user asks what it is.
|
|
21
|
+
|
|
19
22
|
The single most important idea: **a green run is not automatically a correct run.** The harness has
|
|
20
23
|
several ways to no-op a check while still producing a green run (skip an assertion on replay — now
|
|
21
24
|
flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an empty egress
|
|
22
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
26
|
the highest-value part. Read it.
|
|
24
27
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.
|
|
26
|
-
> `desktop-1.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.4.0` (baseline
|
|
29
|
+
> `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
31
|
|
|
29
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
42
|
|
|
40
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.4.0"`. **Pin `@^2.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
46
|
|
|
44
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -144,9 +147,9 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
|
|
|
144
147
|
| Tier | What it gives you | Use when |
|
|
145
148
|
|---|---|---|
|
|
146
149
|
| `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
|
|
147
|
-
| `container` | Real sandbox + real default-deny egress (**default**) | Most functional + boundary tests. |
|
|
148
|
-
| `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity | Testing untrusted code escape, not network behavior. |
|
|
149
|
-
| `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with native Bash
|
|
150
|
+
| `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm`/`protocol` never offer it, so the same `tool_not_called` is vacuous there — moving a scenario between tiers can silently void a web-fetch assertion | Most functional + boundary tests. |
|
|
151
|
+
| `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
|
|
152
|
+
| `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
|
|
150
153
|
|
|
151
154
|
Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejects `--fidelity`
|
|
152
155
|
(it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
|
|
@@ -309,6 +312,7 @@ them by what you're trying to prove:
|
|
|
309
312
|
| no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
|
|
310
313
|
| a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
|
|
311
314
|
| the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
|
|
315
|
+
| the user was **told something specific** at a gate | `question_context: {when_question, matches}` — a regex over the question label + option labels + option **descriptions**. Reach for this when the wording may land in an option's `description`, which `question_asked` and `question_options` cannot see |
|
|
312
316
|
| a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
|
|
313
317
|
| every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
|
|
314
318
|
| a context compaction happened | `compaction_occurred: true` |
|
|
@@ -402,8 +406,15 @@ the discovery/encode/record dance entirely and answer gates **live during the re
|
|
|
402
406
|
`run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
|
|
403
407
|
`cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
|
|
404
408
|
refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
|
|
405
|
-
cassette-portability pre-flight below**, so it cannot green something a paid run would reject.
|
|
406
|
-
|
|
409
|
+
cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
|
|
410
|
+
guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
|
|
411
|
+
one a paid run would give. On a **directory** the path-dependent verdicts (host-inventory, cassette
|
|
412
|
+
portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
|
|
413
|
+
the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
|
|
414
|
+
takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
|
|
415
|
+
contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
|
|
416
|
+
real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
|
|
417
|
+
reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
|
|
407
418
|
|
|
408
419
|
**Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
|
|
409
420
|
Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
|
|
@@ -426,7 +437,12 @@ for a cassette that cannot verify staleness from its own location — recoverabl
|
|
|
426
437
|
|
|
427
438
|
**Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
|
|
428
439
|
scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
|
|
429
|
-
(and `verify-run`) read the gates + offered option
|
|
440
|
+
(and `verify-run`) read the gates + every offered option's **label and `description`** out of that run's
|
|
441
|
+
`events.jsonl` for free — a skill routinely puts the sentence the user is actually deciding on in a
|
|
442
|
+
`description`, and `question_context:` is the key that gates on it (`question_options:` compares labels
|
|
443
|
+
only). When a view renders no field you need, read `events.jsonl` directly rather than concluding the text
|
|
444
|
+
was never delivered — the views are a digest, and the record is wider (`jq` recipes in
|
|
445
|
+
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)). Iterate
|
|
430
446
|
your `answers:` against that kept run, then record once. **But the kept run is a snapshot:** if you change the
|
|
431
447
|
skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
|
|
432
448
|
current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
|
|
@@ -534,8 +550,15 @@ Recognize these before "fixing" a non-bug:
|
|
|
534
550
|
only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
|
|
535
551
|
**Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
|
|
536
552
|
scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
|
|
537
|
-
**The fix is lane-dependent.** On `lane: local`, write deliverables
|
|
538
|
-
|
|
553
|
+
**The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
|
|
554
|
+
can see them, but do **not** hardcode the literal prefix `outputs/`: on the desktop-local host-loop lane
|
|
555
|
+
(what production runs) the file tools are ALREADY rooted at `outputs/`, so `outputs/x.md` doubles to
|
|
556
|
+
`outputs/outputs/x.md` and the user never sees it — a **bare filename** is correct there. At
|
|
557
|
+
`fidelity: container`/`microvm` (VM-loop, the harness default) the base is the session root instead, so a
|
|
558
|
+
bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an explicit delivery. Addressing
|
|
559
|
+
a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
|
|
560
|
+
decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
|
|
561
|
+
[docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
|
|
539
562
|
nothing is delivered by location there, so only an explicit delivery counts. Assert
|
|
540
563
|
**`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
|
|
541
564
|
downloaded inputs) rather than a delivery gap.
|
|
@@ -821,7 +844,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
821
844
|
channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
|
|
822
845
|
reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
|
|
823
846
|
Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
|
|
824
|
-
string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` /
|
|
847
|
+
string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` / `question_context` /
|
|
825
848
|
`questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
|
|
826
849
|
old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
|
|
827
850
|
on unobserved delivery (absence of evidence is failure, not neutral).
|
|
@@ -1034,11 +1057,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
|
|
|
1034
1057
|
([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
|
|
1035
1058
|
→ "File delivery" has the binary-verified detail; repo-only.)
|
|
1036
1059
|
|
|
1037
|
-
25. **
|
|
1038
|
-
--allow-host-inventory-fixture`
|
|
1060
|
+
25. **Three host-inventory flags — two on `record`, one on `verify-cassettes`.** `record
|
|
1061
|
+
--allow-host-inventory-fixture` proceeds past the PRE-FLIGHT refusal when recording a host-inheriting
|
|
1039
1062
|
(`protocol`/`hostloop`/`cowork`-resolving-to-hostloop) cassette into a repo-visible path — otherwise
|
|
1040
1063
|
`record` refuses before it spends (freezing this machine's MCP servers/agents/account into a committed
|
|
1041
|
-
fixture is the risk).
|
|
1064
|
+
fixture is the risk). It bypasses that check and **nothing else**: the finished recording is still
|
|
1065
|
+
scanned, and a real finding still quarantines it, so you never have to audit the session by hand to
|
|
1066
|
+
pass it. Writing a recording the scan DID flag is the separate `record
|
|
1067
|
+
--allow-host-inventory-findings`. That pre-spend check **warns rather than refuses when the cassette already
|
|
1042
1068
|
exists** — refusing would fire on every `--rerecord-stale` pass — and it reads the tier and the
|
|
1043
1069
|
destination path, never the bytes. So `record` also scans the FINISHED recording, after redaction and
|
|
1044
1070
|
before the write: a `host-inventory`/`machine-inventory` finding on a repo-visible path is
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "2.
|
|
20
|
+
(e.g. `version: "2.4.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.246 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
41
41
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
42
42
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^2.
|
|
70
|
+
- run: npm i -g "cowork-harness@^2.4.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -149,7 +149,7 @@ The split is not just about tokens — it decides **where each lane can run**:
|
|
|
149
149
|
`max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
|
|
150
150
|
`allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
|
|
151
151
|
`allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
|
|
152
|
-
`questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
152
|
+
`question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
|
|
153
153
|
(`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
|
|
154
154
|
manifest. `file_absent` is in neither class — it is live/verify-run only.
|
|
155
155
|
**That list is illustrative, not the authoritative set** — more keys are replay-checkable than fit a
|
|
@@ -237,12 +237,14 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
237
237
|
(`hostloop`, `protocol`) has an **empty** redaction policy — that combination commits real host paths
|
|
238
238
|
the `path` scanner then hard-fails at `verify-cassettes` time. The always-on scanner remains the
|
|
239
239
|
universal net (container-tier recordings can trip it too).
|
|
240
|
-
- **Host-inheriting record refused by default — `--allow-host-inventory-fixture`
|
|
241
|
-
`protocol`/`hostloop`/`cowork`-resolving-to-hostloop record into a
|
|
242
|
-
freeze THIS machine's MCP server names, agents, and account metadata
|
|
243
|
-
`record` refuses **before the paid spawn**.
|
|
244
|
-
|
|
245
|
-
|
|
240
|
+
- **Host-inheriting record refused by default — `--allow-host-inventory-fixture` bypasses the
|
|
241
|
+
PRE-FLIGHT, not the scan.** A `protocol`/`hostloop`/`cowork`-resolving-to-hostloop record into a
|
|
242
|
+
repo-visible cassette path would freeze THIS machine's MCP server names, agents, and account metadata
|
|
243
|
+
into a committed fixture, so `record` refuses **before the paid spawn**. `--allow-host-inventory-fixture`
|
|
244
|
+
proceeds past that check — and nothing more: the finished recording is still scanned, and a real finding
|
|
245
|
+
still refuses the write and quarantines it. **You do not need to audit the session by hand to pass this
|
|
246
|
+
flag**; that precondition was never decidable by the operator, and the scan is the real gate. Writing a
|
|
247
|
+
recording the scan DID flag is the separate `--allow-host-inventory-findings`.
|
|
246
248
|
|
|
247
249
|
Two details that matter for a re-record loop. The pre-spend check **warns rather than refuses when the
|
|
248
250
|
cassette already exists**, deliberately: refusing there would fire on every `--rerecord-stale` pass and
|
|
@@ -340,7 +342,7 @@ jobs:
|
|
|
340
342
|
with: { node-version: '24' }
|
|
341
343
|
- uses: actions/setup-python@v5
|
|
342
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
343
|
-
- run: npm i -g "cowork-harness@^2.
|
|
345
|
+
- run: npm i -g "cowork-harness@^2.4.0"
|
|
344
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
345
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
346
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -369,7 +371,7 @@ jobs:
|
|
|
369
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
370
372
|
fi
|
|
371
373
|
- if: steps.guard.outputs.live == 'true'
|
|
372
|
-
run: npm i -g "cowork-harness@^2.
|
|
374
|
+
run: npm i -g "cowork-harness@^2.4.0"
|
|
373
375
|
- if: steps.guard.outputs.live == 'true'
|
|
374
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
375
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 2.
|
|
3
|
+
Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.4.0`
|
|
4
|
+
(baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -289,10 +289,10 @@ same set live from the schema.
|
|
|
289
289
|
| Assertion | Passes when |
|
|
290
290
|
|---|---|
|
|
291
291
|
| `result: success \| error` | the run ended with that status |
|
|
292
|
-
| `transcript_contains: <str>` | the assistant transcript includes the literal string |
|
|
293
|
-
| `transcript_not_contains: <str>` | it does not |
|
|
294
|
-
| `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose |
|
|
295
|
-
| `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) |
|
|
292
|
+
| `transcript_contains: <str>` | the assistant transcript includes the literal string **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
293
|
+
| `transcript_not_contains: <str>` | it does not **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
294
|
+
| `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
295
|
+
| `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
|
|
296
296
|
| `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
|
|
297
297
|
| `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
|
|
298
298
|
| `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
|
|
@@ -344,8 +344,9 @@ same set live from the schema.
|
|
|
344
344
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
345
345
|
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
346
346
|
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
347
|
-
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
|
|
348
|
-
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered —
|
|
347
|
+
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
348
|
+
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
349
|
+
| `question_context: {when_question?, matches}` | a regex over **everything the gate showed the user** — question label + every option label + every option **description**. Use it when the sentence you need to prove reached the founder may land in any of those fields: `question_asked` sees only the question text, `question_options` compares only labels, so a phrase delivered in an option's `description` is invisible to both. `when_question` narrows; omitting it searches every gate (NOT ambiguous here — this key asks whether the text was shown, not which gate offered which set). Ask-time payload only, never a `tool_result` (a producer that also writes the phrase to its own gate-state file would otherwise false-green it). Zero gates FAILS ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
|
|
349
350
|
| `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
|
|
350
351
|
| `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
|
|
351
352
|
| `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
|
|
@@ -369,8 +370,8 @@ same set live from the schema.
|
|
|
369
370
|
| `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
|
|
370
371
|
| `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
|
|
371
372
|
| `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
|
|
372
|
-
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) |
|
|
373
|
-
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** |
|
|
373
|
+
| `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
374
|
+
| `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
|
|
374
375
|
|
|
375
376
|
`expect_denied: [host, …]` adds one `egress_denied` per host. Run `cowork-harness assertions --list` for this
|
|
376
377
|
table from the live schema. Example: `artifact_json: { artifact: outputs/cap.json, path: me.run_id, equals: "r1" }`.
|
|
@@ -441,7 +442,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
|
|
|
441
442
|
modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
|
|
442
443
|
`allow_stall` are also kept on replay, evaluated as no-op passes.
|
|
443
444
|
|
|
444
|
-
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `questions_count_max`,
|
|
445
|
+
**Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
|
|
445
446
|
`gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`,
|
|
446
447
|
`path_denied`, `no_path_denied` (the latter three are also `fidelity: hostloop`-only — see the assertion
|
|
447
448
|
table). With `controlOut` present they evaluate; on an old
|
|
@@ -545,7 +546,8 @@ reference omits). Neither list is a strict superset of the other — reach for t
|
|
|
545
546
|
`egress_denied` is dropped. Both now warn loudly. → put egress/live-only checks on a live gate; one
|
|
546
547
|
concern per item; run the linter. (`LIVE_ONLY_KEYS`/`MANIFEST_KEYS` in `src/run/cassette.ts`.)
|
|
547
548
|
|
|
548
|
-
2. **Gate keys need a `controlOut` cassette.** `question_asked`, `
|
|
549
|
+
2. **Gate keys need a `controlOut` cassette.** `question_asked`, `question_options`, `question_context`,
|
|
550
|
+
`questions_count_max`,
|
|
549
551
|
`gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked` only evaluate on
|
|
550
552
|
replay with `controlOut`; on an old cassette they warn and are excluded (not passed).
|
|
551
553
|
`gate_answers_delivered` **fails on
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 2.
|
|
5
|
+
Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -17,7 +17,7 @@ a paid live re-record? Walk this tree — the answer is usually no:
|
|
|
17
17
|
`cowork-harness replay <cassette> --assert-from <scenario.yaml>`. Token-free, no re-record.
|
|
18
18
|
If the recording genuinely lacks the telemetry a key needs (very old cassettes), the key fails
|
|
19
19
|
**loud** as `evidence-unavailable` — that is correct behavior, not a bug; only then re-record.
|
|
20
|
-
2. **Gate keys** (`question_asked`, `question_options`, `questions_count_max`, `gate_answers_delivered`) on a cassette
|
|
20
|
+
2. **Gate keys** (`question_asked`, `question_options`, `question_context`, `questions_count_max`, `gate_answers_delivered`) on a cassette
|
|
21
21
|
**with `controlOut`** (any modern recording) → same token-free `--assert-from` path.
|
|
22
22
|
3. **Gate keys** on a **pre-`controlOut`** cassette → one re-record unlocks gate asserts for that
|
|
23
23
|
cassette permanently.
|
|
@@ -114,6 +114,7 @@ CONTENT_KEYS = {
|
|
|
114
114
|
GATE_KEYS = {
|
|
115
115
|
"question_asked",
|
|
116
116
|
"question_options",
|
|
117
|
+
"question_context",
|
|
117
118
|
"questions_count_max",
|
|
118
119
|
"gate_answers_delivered",
|
|
119
120
|
"gate_answer_count_min",
|
|
@@ -284,6 +285,7 @@ def _load_top_level_keys():
|
|
|
284
285
|
# every valid top-level scenario key (generated from the zod ScenarioObject schema; see _load_top_level_keys)
|
|
285
286
|
TOP_LEVEL_KEYS = _load_top_level_keys()
|
|
286
287
|
REGEX_KEYS = {
|
|
288
|
+
"matches",
|
|
287
289
|
"transcript_matches",
|
|
288
290
|
"transcript_not_matches",
|
|
289
291
|
"when_question",
|
|
@@ -486,6 +488,30 @@ def lint_doc(doc, path, raw_lines):
|
|
|
486
488
|
|
|
487
489
|
fidelity = (doc.get("fidelity") or "container")
|
|
488
490
|
lane = (doc.get("lane") or "local")
|
|
491
|
+
|
|
492
|
+
# W: no `fidelity:` — the default models the WRONG LANE.
|
|
493
|
+
# `container` (the schema default) models VM-loop; production runs HOST-LOOP, gate 1143815894 is
|
|
494
|
+
# force-ON in every shipped baseline. So an omitted key measures the scenario against a lane real
|
|
495
|
+
# users are not on: the file tools resolve a bare relative path differently, the shell starts
|
|
496
|
+
# somewhere else, and the offered tool set differs (measured 2026-08-27).
|
|
497
|
+
# Read the KEY, not the resolved value: `fidelity: container` is a deliberate choice and must not warn.
|
|
498
|
+
# DEPRECATION — `fidelity:` becomes REQUIRED in the next major; this is the warning window.
|
|
499
|
+
if "fidelity" not in doc:
|
|
500
|
+
findings.append(
|
|
501
|
+
Finding(
|
|
502
|
+
"WARN",
|
|
503
|
+
"fidelity-defaulted",
|
|
504
|
+
"no `fidelity:` — defaulting to `container`, which models the VM-LOOP lane. Production "
|
|
505
|
+
"runs HOST-LOOP by default (gate 1143815894), so this scenario is likely measured "
|
|
506
|
+
"against a lane your users are not on.",
|
|
507
|
+
"Name a tier: `fidelity: hostloop` to match production, `fidelity: cowork` to auto-pick "
|
|
508
|
+
"the way Cowork does, or `fidelity: container` to keep today's behaviour deliberately. "
|
|
509
|
+
"Switching tiers can COST you assertions: `no_scratchpad_leak` is container-only (an "
|
|
510
|
+
"error elsewhere) and `transcript_no_host_path` fails by design at hostloop/protocol. "
|
|
511
|
+
"The default is being removed — `fidelity:` becomes REQUIRED in the next major.",
|
|
512
|
+
path,
|
|
513
|
+
)
|
|
514
|
+
)
|
|
489
515
|
items = _assert_items(doc)
|
|
490
516
|
assert_keys = _all_assert_keys(items)
|
|
491
517
|
has_expect_denied = bool(doc.get("expect_denied"))
|
|
@@ -754,7 +780,11 @@ def lint_doc(doc, path, raw_lines):
|
|
|
754
780
|
# YAML 1.1 but are rejected by the loader as strings — see _numeric for the same dialect gap.)
|
|
755
781
|
delivered_values = _assert_values(items, "gate_answers_delivered")
|
|
756
782
|
if any(v is not False for v in delivered_values):
|
|
757
|
-
has_companion =
|
|
783
|
+
has_companion = (
|
|
784
|
+
"question_asked" in assert_keys
|
|
785
|
+
or "question_options" in assert_keys
|
|
786
|
+
or "question_context" in assert_keys
|
|
787
|
+
) or any(
|
|
758
788
|
(n := _numeric(v)) is not None and n >= 1 for v in _assert_values(items, "gate_answer_count_min")
|
|
759
789
|
)
|
|
760
790
|
if not has_companion:
|
|
@@ -837,6 +867,7 @@ def lint_doc(doc, path, raw_lines):
|
|
|
837
867
|
),
|
|
838
868
|
("question_asked", "question_asked" in assert_keys),
|
|
839
869
|
("question_options", "question_options" in assert_keys),
|
|
870
|
+
("question_context", "question_context" in assert_keys),
|
|
840
871
|
("gate_answers_delivered: false", any(v is False for v in _assert_values(items, "gate_answers_delivered"))),
|
|
841
872
|
],
|
|
842
873
|
"a delivered gate records at least one question, so requiring a gate to be present contradicts requiring zero questions",
|