cowork-harness 4.5.0 → 4.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (92) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +12 -8
  2. package/.claude/skills/cowork-harness/references/assertion-catalog-agents-skills-budgets.md +37 -0
  3. package/.claude/skills/cowork-harness/references/assertion-catalog-gates-hooks-modifiers.md +50 -0
  4. package/.claude/skills/cowork-harness/references/assertion-catalog-outcome-files-tools.md +35 -0
  5. package/.claude/skills/cowork-harness/references/assertion-catalog.md +10 -98
  6. package/.claude/skills/cowork-harness/references/assertions-guide.md +6 -4
  7. package/.claude/skills/cowork-harness/references/authoring.md +12 -3
  8. package/.claude/skills/cowork-harness/references/ci-recipe.md +7 -7
  9. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  10. package/.claude/skills/cowork-harness/references/debugging.md +8 -4
  11. package/.claude/skills/cowork-harness/references/eval.md +3 -2
  12. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +2 -1
  13. package/.claude/skills/cowork-harness/references/gotchas.md +1 -1
  14. package/.claude/skills/cowork-harness/references/hillclimb-recipe.md +1 -1
  15. package/.claude/skills/cowork-harness/references/hillclimb.md +4 -2
  16. package/.claude/skills/cowork-harness/references/measurement.md +1 -1
  17. package/.claude/skills/cowork-harness/references/run-record-replay.md +16 -6
  18. package/.claude/skills/cowork-harness/references/scenario-schema.md +23 -9
  19. package/.claude/skills/cowork-harness/references/semantic-judging.md +3 -3
  20. package/.claude/skills/cowork-harness/references/task-recipes.md +40 -17
  21. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +146 -0
  22. package/.claude/skills/cowork-harness/scripts/scenario.py +216 -11
  23. package/CHANGELOG.md +171 -0
  24. package/DESIGN.md +3 -3
  25. package/README.md +7 -5
  26. package/SPEC.md +14 -8
  27. package/baselines/desktop-2.26454.0.json +3 -0
  28. package/baselines/desktop-2.26454.2.json +1495 -0
  29. package/baselines/prompts/cowork-system-prompt-fingerprints.json +19 -1
  30. package/dist/agent/session.js +15 -3
  31. package/dist/answer-channel.js +111 -0
  32. package/dist/assert.js +664 -148
  33. package/dist/baseline.js +15 -0
  34. package/dist/cli.js +22 -5
  35. package/dist/decide/decider.js +18 -2
  36. package/dist/eval/classify.js +17 -1
  37. package/dist/eval/report.js +2 -0
  38. package/dist/eval/runs.js +1 -0
  39. package/dist/fixture/workspace.js +8 -1
  40. package/dist/gates-scripted.js +185 -0
  41. package/dist/glob.js +30 -0
  42. package/dist/run/artifacts.js +21 -6
  43. package/dist/run/cassette.js +160 -23
  44. package/dist/run/chat-result.js +1 -0
  45. package/dist/run/execute.js +104 -6
  46. package/dist/run/hook-events.js +35 -12
  47. package/dist/run/lane-notice.js +2 -2
  48. package/dist/run/renderer.js +4 -1
  49. package/dist/run/run-index.js +11 -2
  50. package/dist/run/run.js +32 -1
  51. package/dist/run/verdict.js +14 -2
  52. package/dist/run/verify-context.js +7 -0
  53. package/dist/runtime/argv.js +4 -2
  54. package/dist/runtime/container.js +2 -1
  55. package/dist/runtime/host-cli-probe.js +56 -0
  56. package/dist/runtime/microvm.js +2 -1
  57. package/dist/runtime/protocol.js +10 -4
  58. package/dist/session.js +23 -1
  59. package/dist/types.js +112 -8
  60. package/docs/README.md +1 -0
  61. package/docs/cassette.md +9 -6
  62. package/docs/ci.md +3 -3
  63. package/docs/cli.md +9 -8
  64. package/docs/companion-skill.md +2 -2
  65. package/docs/debugging.md +2 -1
  66. package/docs/eval.md +1 -0
  67. package/docs/fidelity-gaps.md +19 -6
  68. package/docs/headless-no-answer.md +119 -0
  69. package/docs/invariants.md +1 -1
  70. package/docs/maintenance.md +1 -1
  71. package/docs/run-status.md +4 -3
  72. package/docs/scenario.md +88 -24
  73. package/docs/session.md +9 -2
  74. package/docs/stats.md +4 -2
  75. package/examples/replays/README.md +1 -1
  76. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  77. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  78. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  79. package/examples/sessions/gated-probe-headless.yaml +11 -0
  80. package/examples/sessions/hook-decision-probe.yaml +8 -0
  81. package/llms.txt +1 -0
  82. package/package.json +1 -1
  83. package/python/README.md +2 -0
  84. package/python/cowork_harness.py +7 -0
  85. package/python/test_cowork_lane.py +23 -0
  86. package/python/test_scenario_lint.py +156 -1
  87. package/schema/cassette.v15.json +350 -0
  88. package/schema/run-result.json +10 -4
  89. package/schema/scenario.schema.json +288 -42
  90. package/schema/session.schema.json +11 -0
  91. package/scripts/check-versions.ts +2 -1
  92. package/scripts/gen-schema.ts +8 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did it make answers worse? (`eval`: paired A/B, pinned models) — or improving one round by round (`hillclimb`, the `/claude-api hillclimb` runner). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval / hillclimb commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 4.5.0
7
- tracks-harness: cowork-harness 4.5.0 (baseline desktop-2.26454.0)
6
+ version: 4.6.0
7
+ tracks-harness: cowork-harness 4.6.0 (baseline desktop-2.26454.2)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -26,8 +26,8 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
26
26
  full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
27
  Read them.
28
28
 
29
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.5.0` (baseline
30
- > `desktop-2.26454.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.6.0` (baseline
30
+ > `desktop-2.26454.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
31
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
32
32
 
33
33
  ## Preflight — make sure the harness can actually run
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
43
43
 
44
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
45
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
46
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.5.0"`. **Pin `@^4.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.6.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.6.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.6.0"`. **Pin `@^4.6.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
47
47
 
48
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
49
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -124,8 +124,9 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
124
124
  the gate it answers.
125
125
  4. **`replay` evaluates the FROZEN scenario.** Editing `scenarios/*.yaml` changes nothing on a plain
126
126
  `replay`: re-check an `assert:` edit with `replay --assert-from <file>`, and re-record for any other key.
127
- 5. **Some keys pass on absence.** `gate_answers_delivered` passes when no gate fired — pair it with
128
- `gate_answer_count_min: 1`. `tool_called` proves a tool ran, not that it was attempted.
127
+ 5. **Some keys pass on absence.** `gate_answers_delivered` and `gates_all_scripted` pass when no gate fired —
128
+ pair either with `gate_answer_count_min: 1` (or declare `questions_count_max: 0`). `tool_called` proves a
129
+ tool ran, not that it was attempted.
129
130
  6. **An untracked skill mounts empty.** `git add` a new skill before testing it, and commit before
130
131
  recording the cassette that locks it.
131
132
  7. **The tier decides what exists.** `protocol` has no sandbox and no egress, tool names differ per tier
@@ -158,7 +159,10 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
158
159
  | [`references/debugging.md`](references/debugging.md) | triage, `result.json` fields and `trace` views, `chat` |
159
160
  | [`references/gotchas.md`](references/gotchas.md) | the full "✓ passed ≠ correct" landmine catalog |
160
161
  | [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality, and goals with no flag (force a compaction, ablate a section, a form reply, resume in a new conversation, hook JSON decisions, schema checks, unattended runs) |
161
- | [`references/assertion-catalog.md`](references/assertion-catalog.md) | every `assert:` key's semantics, the verdict-signal table |
162
+ | [`references/assertion-catalog.md`](references/assertion-catalog.md) | the assertion catalog's index: conventions shared by every key, the verdict-signal table, and links to the per-family files below |
163
+ | [`references/assertion-catalog-outcome-files-tools.md`](references/assertion-catalog-outcome-files-tools.md) | per-key rows: outcome, transcript, files and artifacts, tools |
164
+ | [`references/assertion-catalog-agents-skills-budgets.md`](references/assertion-catalog-agents-skills-budgets.md) | per-key rows: sub-agents, skills and connectors, budgets, tasks, delivery |
165
+ | [`references/assertion-catalog-gates-hooks-modifiers.md`](references/assertion-catalog-gates-hooks-modifiers.md) | per-key rows: gates, hooks, path denial, verdict modifiers, egress, judged keys |
162
166
  | [`references/semantic-judging.md`](references/semantic-judging.md) | `semantic_matches` in full: what the judge reads, fork results, refusal reasons, provenance |
163
167
  | [`references/scenario-schema.md`](references/scenario-schema.md) | every YAML field, which keys survive `replay`, the `web_fetch` model |
164
168
  | [`references/fidelity-and-answers.md`](references/fidelity-and-answers.md) | tier semantics, answer paths, the determinism contract |
@@ -0,0 +1,37 @@
1
+ # Assertion catalog: sub-agents, skills and connectors, budgets, tasks, delivery
2
+
3
+ Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Part 2 of 3 of the per-key table. The conventions every key shares (an item with several
4
+ keys is an AND; glob vs regex name matching) and the verdict-signal table are in
5
+ [`assertion-catalog.md`](./assertion-catalog.md); the other parts are [`assertion-catalog-outcome-files-tools.md`](./assertion-catalog-outcome-files-tools.md), [`assertion-catalog-gates-hooks-modifiers.md`](./assertion-catalog-gates-hooks-modifiers.md).
6
+
7
+ | Assertion | Passes when |
8
+ |---|---|
9
+ | `subagent_tool_used: <glob>` | a sub-agent used a tool matching this glob (same `*`/`?`, anchored, case-sensitive semantics as `tool_called`, including the empty/regex-ish rejection) |
10
+ | `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
11
+ | `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
12
+ | `subagent_file_write: {path?, path_suffix?, tool?}` | a sub-agent-origin write attempt whose raw path equals `path` (exact) or ends with `path_suffix` has a paired non-error tool_result — the causal half of a delivery probe; requires one of `path`/`path_suffix`; `tool` defaults to Write/Edit/MultiEdit; content-class; tier-agnostic |
13
+ | `subagent_dispatch_healthy: {type?, delivered?, path?, path_suffix?, no_vm_paths?}` | **`fidelity: hostloop` only** — composite: selects dispatch(es) via `type` (same matching as `subagent_dispatched`; omit to require every dispatch) and, for EACH selected dispatch, checks it (not just any sub-agent) delivered a paired non-error write (`delivered`, default true — narrowed by `path`/`path_suffix`, same exact-vs-suffix precedence as `subagent_file_write`) and made no `/sessions` VM-path attempt (`no_vm_paths`, default true) — both scoped to that dispatch's OWN `parentToolUseId`, the per-dispatch correlation `subagent_file_write` (which matches ANY sub-agent write) cannot express; a `type` that matches no dispatch FAILS; content-class (`RunResult.fileToolAttempts` + `RunResult.toolResults`); any non-hostloop tier FAILS "cannot verify" |
14
+ | `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
15
+ | `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
16
+ | `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut. **Covers what the run dispatches** (`Agent`/`Task`, including `Agent(subagent_type:"fork")`), **not a `context: fork` skill** invoked through the `Skill` tool: that skill's own answer is never a dispatch — it comes back as the `Skill` tool result, which agent 2.1.284 builds as `Skill "<name>" completed (forked execution).`, a `Result:` line, then the answer. Assert on it with `tool_result_matches` anchored on that prefix, e.g. `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` (`[\s\S]*` because `.` stops at a newline; the match is case-insensitive, has no multiline flag so `^` is the start of the result, and sees the first 10,240 characters of each result — 32,768 for a top-level `Skill` result). This covers a **foreground** fork only: a backgrounded fork's result is the line `Skill "<name>" launched (forked execution, running in the background).`, which carries no answer |
17
+ | `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 in `scenario-schema.md` |
18
+ | `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked — via the `Skill` tool, or by a prompt starting `/<skill> …` / `/<plugin>:<skill> …` (the agent expands that itself with no `Skill` call; recorded as `slashInvokedSkills`) — evidence-unavailable (not a normal fail) if neither matched and the agent's init tools have no `Skill` tool, or the leading `/name` can't be resolved (ambiguous bare name, no skill inventory). A green here means the AGENT expanded the slash; real Cowork's app resolves a typed slash first and has refused a bare name that differs from its plugin's name (Desktop 2.19675.0, 2026-10-03, 4 runs), a refusal no run can see — pick the skill from the menu or name it like its plugin (`/<plugin>:<skill>` was not measured with a single copy installed) |
19
+ | `no_skill_triggered: <regex>` | no invoked skill id matched, counting a slash-command invocation as well as a `Skill` call — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent, the `Skill` tool is unobservable, or the prompt's leading `/name` can't be resolved |
20
+ | `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
21
+ | `connector_available: <regex>` | an MCP server/connector's name matched the regex (available, not necessarily used) — evidence-unavailable if `RunResult.context.mcpServers` is absent |
22
+ | `tool_available: <regex>` | a tool in the init manifest matched the regex (available, not necessarily called — see `tool_called` for invocation) — evidence-unavailable if `RunResult.context.tools` is absent. The `mcp__skills__*`/`mcp__plugins__*` discovery tools are modeled (as `alwaysLoad`) on `container`/`hostloop`/`cowork` — a miss there is a real absence; `microvm`/`protocol` declare no such server, so a miss on those two tiers means "not modeled at this tier", not "provably unavailable" |
23
+ | `skill_tool_used: {skill, tool}` | a tool whose name matches `tool` ran inside a skill-activation window whose `skillId` matches `skill` (`RunResult.skillActivity`) — evidence-unavailable if skill-activity telemetry is absent; heuristic for inline skills (a sticky, sequential window matching the agent's `activeSkill` scope, not an exact per-tool boundary). **Scope:** the window's tool counts **include sub-agent calls** made during it, so this key can't say which agent called (`subagent_tool_used` is the sub-agent-only claim), and it matches tool **names** only — never the path/args, so "did it read *this* file" isn't expressible (per-sub-agent reads are recorded at `subagents[].referencesRead`, readable but not assertable) |
24
+ | `max_cost_usd: <N>` | the run's SDK-reported cost is ≤ N USD (the agent session only — the `semantic_matches` judge and the LLM decider (`on_unanswered: llm` / `--decider-llm`) are separate model calls that are **not** included) — evidence-unavailable if cost telemetry is absent. **Replay asserts the frozen recording's cost, not fresh spend** — a real regression needs a live `run` |
25
+ | `max_tokens: <N>` | `usage.input_tokens + usage.output_tokens` ≤ N (cache tokens excluded) — same replay caveat as `max_cost_usd` |
26
+ | `tool_calls_max: <N>` | total top-level tool calls (sum of `toolCounts`) ≤ N — meaningfully replay-checkable (re-drive recomputes `toolCounts` deterministically) |
27
+ | `tool_no_error: <regex>` | no tool whose name matches the regex recorded any error (`RunResult.toolErrors[name].errors === 0` for every match) — **requires ≥1 matching tool call** (a regex matching nothing fails, so a typo can't silently pass); evidence-unavailable if tool-error telemetry is absent |
28
+ | `tool_no_error_if_called: <regex>` | like `tool_no_error` but passes vacuously when no tool matches the regex — the presence-free variant |
29
+ | `max_tool_errors: <N>` | total tool errors across all tools (sum of `RunResult.toolErrors[*].errors`) ≤ N — evidence-unavailable if tool-error telemetry is absent |
30
+ | `max_redundant_tool_calls: <N>` | total WASTED repeated tool calls (sum of `(count-1)` across every redundant `{name,args}` group in `RunResult.redundantToolCalls`) ≤ N — not the raw count of redundant groups; evidence-unavailable if redundant-call telemetry is absent |
31
+ | `max_turns: <N>` | the SDK-reported (or fallback-counted) turn count ≤ N — meaningfully replay-checkable (re-drive recounts turns deterministically, same as `tool_calls_max`) |
32
+ | `compaction_occurred: true` | a context-compaction boundary occurred (a `compact_boundary` system event was recorded) — lives in the stdout stream, so meaningfully replay-checkable; evidence-unavailable if context-event telemetry is absent. **Only `true` is valid** — omit to not require it |
33
+ | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
34
+ | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
35
+ | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
36
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, an empty macOS system dir from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
37
+ | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
@@ -0,0 +1,50 @@
1
+ # Assertion catalog: gates, hooks, path denial, verdict modifiers, egress, judged keys
2
+
3
+ Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Part 3 of 3 of the per-key table. The conventions every key shares (an item with several
4
+ keys is an AND; glob vs regex name matching) and the verdict-signal table are in
5
+ [`assertion-catalog.md`](./assertion-catalog.md); the other parts are [`assertion-catalog-outcome-files-tools.md`](./assertion-catalog-outcome-files-tools.md), [`assertion-catalog-agents-skills-budgets.md`](./assertion-catalog-agents-skills-budgets.md).
6
+
7
+ Under the session key `answer_channel: none` the question and gate keys below, `questions_count_max` included, are
8
+ refused at load: no question reaches the harness to grade.
9
+
10
+ | Assertion | Passes when |
11
+ |---|---|
12
+ | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
13
+ | `question_options: {when_question?, equals?: [..], contains?: [..], order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
14
+ | `question_context: {when_question?, matches}` | a regex over **everything the gate showed the user** — question label + every option label + every option **description**. Use it when the sentence you need to prove reached the founder may land in any of those fields: `question_asked` sees only the question text, `question_options` compares only labels, so a phrase delivered in an option's `description` is invisible to both. `when_question` narrows; omitting it searches every gate (NOT ambiguous here — this key asks whether the text was shown, not which gate offered which set). Ask-time payload only, never a `tool_result` (a producer that also writes the phrase to its own gate-state file would otherwise false-green it). Zero gates FAILS ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
15
+ | `question_option_count: {matches, exactly? \| min? \| max?, when_question?, case_sensitive?}` | counts option **labels** matching a regex on **every** selected sub-question (e.g. one `No changes — ` option per gate); zero asked FAILS. `case_sensitive` covers the whole pattern; single-quote regexes. Details: [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant. |
16
+ | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
17
+ | `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
18
+ | `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
19
+ | `gate_answer_count_min: <N>` | at least N AskUserQuestion gates fired AND were delivered non-error — presence companion to `gate_answers_delivered`'s vacuous-pass. **`: 0` asserts nothing** and does not satisfy that pairing; `>= 1` is **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
20
+ | `gates_all_scripted: true \| {include_permissions: true}` | every AskUserQuestion gate that fired was answered by a scripted rule (`answers:` / `--answer`) — not the LLM decider, `first`, or an external or human decider; a failure names each such gate and who answered it. A gate asked but never answered fails. **Evidence:** the run's decisions — a result with no decisions record, or an answer with no recognisable source, fails evidence-unavailable, never passes. On replay each gate is re-classified against the cassette's frozen `answers:`; that is evidence-unavailable when those answers were redacted, or when the cassette records that a live decider answered a gate yet its frozen rules cover every gate (they are not the rules it recorded with). **`{include_permissions: true}`** also checks tool permissions: a scripted `when_tool` rule or a fixed rule counts (a Read/Glob/Grep registry allow, a strict-parity deny, the harness's own path-gate deny, the fail-closed deny when nothing answered live — on replay that row means the recording held no answer, evidence-unavailable); a web_fetch deny that no answer source decided fails (`answered by abstain-fallback`), because real Cowork asks the user; cowork parity's permissive off-registry auto-allow never does, and a replayed web_fetch permission cannot be attributed (evidence-unavailable). Zero gates passes (nothing needed a person), so pair it with `gate_answer_count_min: >= 1` or `questions_count_max` to state whether gates were expected (`lint` warns `unpaired-gates-all-scripted`). |
21
+ | `hook_blocked: <regex>` | one of the harness's own PreToolUse hook callbacks blocked a tool whose name matches the regex (`RunResult.hookEvents`; a plugin's hook never reaches it: use `hook_event_blocked`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream). **Does not see the agent's own refusal**: on `hostloop` against Desktop 2.7032.0+ a relative `Read`/`Write`/`Edit` is denied by the agent's permission rules before any hook runs, so neither this key nor `path_denied` records it — assert it with `tool_result_contains: "denied by your permission settings"` |
22
+ | `no_hook_blocked: true` | no tool was blocked by the harness's own hook callbacks (a plugin's hook never reaches it: use `no_hook_event_blocked`) (distinguishes a real tool crash from an intentional hook block) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette. **Only `true` is valid** |
23
+ | `hook_event_fired: <HookEvent>` | a **command hook** for this event (a plugin's `hooks/hooks.json` or manifest hook — `Stop`, `SessionStart`, `PostToolUse`, …) ran: a `hook_response` system frame with that `hook_event` was recorded (`RunResult.contextEvents`). Any outcome counts. The harness passes `--include-hook-events` whenever a staged plugin declares hooks — that is what puts events other than SessionStart/Setup on the stream — so a recording made without it reports "never fired". Content-class, grades on replay. Recorded end-to-end for `Stop` ([stop-hook-probe.scenario.yaml](https://github.com/yaniv-golan/cowork-harness/blob/main/examples/probes/stop-hook-probe.scenario.yaml)) and for `PreToolUse` at `container` (deciding by JSON and by exit 2); the other names match the same frame but have not each been recorded |
24
+ | `hook_event_blocked: <HookEvent> \| {event, tool?, via?, min?, max?}` | that command hook **blocked**: a `hook_response` frame for the event denied — by exit code 2, or (object form only) by a JSON decision on stdout on a frame the agent marks `outcome: success` (exit 0, or an HTTP hook's 2xx status): a PreToolUse `permissionDecision: "deny"` or a top-level `decision: "block"`, which the agent treats alike. A JSON decision on any other frame, such as exit 1, is unreadable, since the frame does not show whether the agent applied it; a hook the agent cancelled (timed out or aborted) decided nothing. Only stdout that parses whole as a JSON object decides, since stdout is the hook's ordinary output too: a word like `deny`, prose, or JSON after other text is no decision. The bare event means at least one frame with exit code 2, the channel this key has always counted; a JSON deny is not counted there, and when the bare form fails it names any JSON deny it saw. The **object form** counts blocking frames by either channel, one per hook run (PreToolUse runs once per matching tool call, subagent calls included); `via: exit2` counts exit code 2 alone, `via: json` the JSON deny alone, and `via: any` (the default) either. `{event: Stop, max: 0}` is the per-event negative, and fails on a hook that denied by JSON alone. `tool` keeps only frames whose `hook_name` is `<event>:<tool>`: the tool that **fired** (or the SessionStart source), not the matcher in hooks.json. The shell is `Bash` at `container` and `mcp__workspace__bash` at `hostloop`, so the wrong name reports "never fired". An event whose frames carry no tool name (`Stop` is named `Stop`) reports evidence-unavailable for any `tool`. Fails naming the exit codes seen when the hook fired without blocking; fails "never fired" when no frame is in scope, `max: 0` included, so a disabled hook never passes. A frame whose outcome cannot be read (no exit code, a hook that started and never answered, stdout a redaction policy rewrote or the agent truncated, output the agent did not apply as a verdict) is evidence-unavailable when it could change the verdict. A Python hook whose script is missing (`python3 missing.py`) also exits 2, so it reads as a block. Cannot-verify when the run has no context events. Frames carry no plugin id: when a second staged plugin declares the event, or a `protocol` run can see hooks installed on the host, the run warns that the frame cannot be attributed to the plugin under test; the verdict does not change. Content-class |
25
+ | `no_hook_event_blocked: true \| {event, tool?}` | no command hook blocked: every `hook_response` frame (or every frame for `event`) neither exited 2 nor denied by JSON, the same rule as `hook_event_blocked`'s object form. **Never vacuous**: no frame in scope is evidence-unavailable, and `true` also needs a frame of an event other than SessionStart/Setup, because those stream even when hook events were not requested (they stream only when a staged plugin declares hooks). A frame whose outcome cannot be read (no exit code, a hook that started and never answered, stdout a redaction policy rewrote or the agent truncated, output the agent did not apply as a verdict) is evidence-unavailable. Frames carry no plugin id, so a second staged plugin's hook, or a host plugin's at `protocol`, counts as a frame here (the run warns). Distinct from `no_hook_blocked`, which reads the harness's **own** PreToolUse callbacks from `controlOut`. Content-class |
26
+ | `hook_decision: {event, decision, tool?, min?, max?}` | counts the `hook_response` frames for `event` whose decision is `decision`, read from the hook's stdout JSON on a frame the agent marks `outcome: success` (exit 0, or an HTTP hook's 2xx status): `hookSpecificOutput.permissionDecision` on `PreToolUse` and `PreModelSwitch`, where it overrides a top-level `decision`, and which the agent ignores on other events; for `PermissionRequest` `hookSpecificOutput.decision.behavior`; for `Elicitation` and `ElicitationResult` an `action: decline`, a deny; else the top-level `decision`; or from exit code 2, which is a deny. A JSON decision on any other frame, such as exit 1, is unreadable, since the frame does not show whether the agent applied it; a hook the agent cancelled (timed out or aborted) decided nothing. `decision` is one of `allow`, `deny`, `ask`, `defer`, as the agent applies them, plus two aliases: `block` means `deny` and `approve` means `allow`. So `deny` matches exit 2, `permissionDecision: "deny"` and `decision: "block"` alike, and the message shows which. Range and `tool` as for `hook_event_blocked` (default: at least one). Only stdout that parses whole as a JSON object decides; empty or non-JSON stdout is no decision, not an error, unless the agent rejected the frame's output, so `allow` never matches a hook that printed nothing. Fails "never fired" when no frame is in scope. A frame whose decision cannot be read (stdout a redaction policy rewrote so it does not parse, truncated output, a frame with no `stdout` field, a decision shape it does not model, a top-level `decision` other than `approve` or `block`, output the agent did not apply (`outcome: "error"` with exit 0, or stderr that opens with the agent's rejection of the JSON, its refusal to read an incomplete capture, or its failure to run the hook), no exit code, a hook that started and never answered, a `hookSpecificOutput` whose `hookEventName` is missing or names another event) is evidence-unavailable when it could change the verdict. Frames carry no plugin id: when a second staged plugin declares the event, or a `protocol` run can see hooks installed on the host, the run warns that the frame cannot be attributed to the plugin under test; the verdict does not change. `record` warns when its redaction policy rewrites hook stdout that this key, `hook_event_blocked` or `no_hook_event_blocked` reads, since the committed cassette's verdict could then only be evidence-unavailable. Content-class, grades on replay |
27
+ | `hook_output_contains: {event, stream?, text \| matches}` / `hook_output_not_contains: {…}` | a command hook **printed** (or never printed) a text on `stdout` / `stderr` / either (`stream`, default `any`) — some (or no) `hook_response` frame for `event` carries `text` (literal, case-sensitive) or `matches` (regex, case-insensitive, no multiline flag — `^`/`$` anchor the whole stream). For a hook that fails open on stderr while exiting 0. No frame for the event FAILS both. On a redacted stream, a literal miss and any `matches` result are evidence-unavailable for either key, while a literal hit outside a token counts; a miss is also evidence-unavailable on a missing field, a hook that started without a response, or output the agent truncated. Frames carry no plugin id (a second plugin, or at `protocol` without a sealed config dir a host-installed one, can answer). Content-class. Recorded end-to-end for `Stop` |
28
+ | `vm_path_denied: true` | **`fidelity: hostloop` only** — at least one recorded path denial (`RunResult.pathDenials`, any source) targeted a `/sessions` VM path — evidence-unavailable if path-denial telemetry is absent. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
29
+ | `path_denied: {tool?, path_matches?, source?, agent_scope?}` | **`fidelity: hostloop` only** — a path denial matching ALL given matchers (`tool` glob, `path_matches` regex, `source` ∈ pretooluse/can_use_tool/permission_denied, `agent_scope` ∈ main/subagent/any) was recorded. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify" |
30
+ | `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
31
+ | `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
32
+ | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
33
+ | `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
34
+ | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question or (after an `AskUserQuestion` gate) a closing request for input having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). On an open-ended `skill` / `probe-dispatch` run (no `assert:` block) pass `--allow-stall` instead. A helpful closing offer ("want me to run this through a structured pass?") also fails `stalled` — read the final message before believing it |
35
+ | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
36
+ | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Has an effect only on a baseline recording outputs `rw` (Desktop before 2.16120.0), where omitting `no_delete_in_outputs` does **not** permit deletes; on `rwd` (2.16120.0+) an outputs delete does not fail by default, so it is an accepted no-op (no warning), unless `no_delete_in_mounts` arms the outputs check, where it waives that check's signal as on `rw`. **Mutually exclusive** with `no_delete_in_outputs`. Silences `outputs_delete`, `outputs_delete_unconfirmed` and `outputs_diff_unavailable`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
37
+ | `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
38
+ | `transcript_no_host_path: true` | no host path (a path under a host home or system root: `Users`, `home`, `root`, the Cowork install dir `opt/cowork`, and the macOS `private/var`, `private/tmp`, `var/folders` and `Volumes` roots, each written here without its leading slash — also inside a `file://` or `computer://` link) leaked into model-visible text (a path that came verbatim from the scenario's input files, prompt, or declared plugins' or local skills' files is exempt; `scan.hostPathsFromInputs` counts such paths) — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
39
+ | `egress_denied: <host>` | the host was blocked by the egress proxy |
40
+ | `egress_allowed: <host>` | the host was allowed through |
41
+ | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
42
+ | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
43
+ | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer: the final result text, the transcript and the files the agent authored. The transcript is top-level `assistant_text` only — it excludes every `tool_use`/`tool_result` and sub-agent text unless opted in. Full semantics (evidence scope, fork results, refusals, judge provenance, cost): [semantic-judging.md](semantic-judging.md). |
44
+ | `semantic_pairwise: {refs?: [..], rubric?: [..], pass_if?, order?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge compares this run's judged document with a **frozen reference** — the same document an earlier run (usually the baseline) produced, written once by `ref freeze` and never regenerated — and answers win, tie, loss or `both_bad`. **The judged document is the one `semantic_matches` builds** (final message + transcript + authored files, same `evidence_files` / `include_subagent_text` / `include_fork_results` options), so the transcript **excludes every `tool_use`/`tool_result`** and a criterion about whether a tool was called cannot be judged from it. `refs` are reference stores relative to the scenario file; the run is judged against every one (under `hillclimb run`, against the flow's own references instead — only the baseline's gates `pass_if`, a later variant's is a metric recorded with `gate: false`), and `pass_if` (default `not_worse`) must hold against each gating one: `win`, `not_worse` (win or tie), or `any` (graded at all — a metric only). `both_bad` fails `win` and `not_worse`. Which output the judge sees first is a seeded coin per run, assert and reference (`order: both` judges both orders: a win in one and a loss in the other is position bias and scores as a tie; any other disagreement keeps the worse outcome; each order's own outcome is recorded as `pairwise[].orders.candidate_first` / `.ref_first`, absent for a single-order grade, and a judge that favours whichever output it sees first shows across runs as `candidate_first` winning more often than `ref_first`). The judge never sees the words "reference" or "baseline", and the stored reference is scrubbed with the grading process's secret set before it reads it (`pairwise[].refRedactions`). A reference must have been frozen with the same evidence options and for the same prompt; a missing, damaged, differently-scoped or other-prompt one **refuses the run before it spends anything**, as does a reference store inside any mounted source (the agent could read its own answer key). Unavailable evidence refuses with `semantic_matches`' typed reasons. Per-reference outcomes land in `assertions[].pairwise`. **LIVE-ONLY** — skipped on replay. |
45
+ | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud). A glob `artifact` (`*`, `?`, `**`; `[` is literal) needs `match: each` (every match) or `match: any` (≥1), e.g. `{artifact: outputs/artifacts/runs/*/run_status.json, match: each, path: status, equals: complete}`. It matches files under the user-visible roots only; uploaded inputs are not matched. Zero matches fail; >200 matches, a link (a matched file, or a symlinked directory on the way to one), a body-less match or an incomplete walk (over 32 levels below the glob's fixed part, over 20,000 entries, or an unreadable directory) is evidence-unavailable. Names match exactly, with no case or Unicode folding. `authored: true` applies to every match. A glob that can reach no user-visible root (an `uploads/` glob) fails and says why. At `record`, a match stored hash-only over the body cap, or an artifact walk that could not see the whole tree, is refused, as for a literal path. Fails on `lane: remote`, in every form: that lane's container filesystem is not locally observable, so there is no body to parse |
46
+ | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
47
+ | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
48
+
49
+ `expect_denied: [host, …]` adds one `egress_denied` per host. Run `cowork-harness assertions --list` for this
50
+ table from the live schema. Example: `artifact_json: { artifact: outputs/cap.json, path: me.run_id, equals: "r1" }`.
@@ -0,0 +1,35 @@
1
+ # Assertion catalog: outcome, transcript, files and artifacts, tools, tool results
2
+
3
+ Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Part 1 of 3 of the per-key table. The conventions every key shares (an item with several
4
+ keys is an AND; glob vs regex name matching) and the verdict-signal table are in
5
+ [`assertion-catalog.md`](./assertion-catalog.md); the other parts are [`assertion-catalog-agents-skills-budgets.md`](./assertion-catalog-agents-skills-budgets.md), [`assertion-catalog-gates-hooks-modifiers.md`](./assertion-catalog-gates-hooks-modifiers.md).
6
+
7
+ | Assertion | Passes when |
8
+ |---|---|
9
+ | `result: success \| error` | the run ended with that status |
10
+ | `transcript_contains: <str>` | the assistant transcript includes the literal string **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
11
+ | `transcript_not_contains: <str>` | it does not **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
12
+ | `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
13
+ | `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
14
+ | `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it. Object form `{path, authored}`: `authored: true` also requires that THIS run created or rewrote the file (an untouched pre-run file — e.g. from a `workspace_fixture` — fails; no pre-run manifest ⇒ evidence-unavailable); `authored: false` states that inheriting it is fine. On a `workspace_fixture` path one of the two is REQUIRED (refused at load otherwise) |
15
+ | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected. Takes the same `{path, authored}` object form as `file_exists` (and `artifact_text`/`artifact_json` take an `authored` field) |
16
+ | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. Checks on **every** baseline. Omitting it: on a baseline recording outputs `rw` (Desktop before 2.16120.0) a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored (`allow_outputs_delete: true` accepts an intended one); on `rwd` (2.16120.0+, including `latest`, where Cowork allows deletes in outputs) nothing checks outputs deletes, so author this key to keep the check. Detection is a post-run bash-command scan plus a per-turn filesystem diff of `outputs/`, not mount enforcement, so a green means none was *detected*. Fails when the diff proves the delete, when a delete in command/call position has an `outputs/` path as its own operand, or when the diff could not verify; a hit resting only on the detector's inference (e.g. a Python variable named `rm`) with a clean diff passes, and the `outputs_delete_unconfirmed` warn is still raised in the run output — see that code for the classes of real delete that land there |
17
+ | `no_delete_in_mounts: true` | no delete op touched `outputs` or any `rw` connected folder, except mounts waived by `allow_delete_in`. Covers `outputs` on **every** baseline, including those where Cowork allows an outputs delete (unless outputs is waived, authoring it arms the outputs check, so a filesystem-proven delete fails as the `outputs_delete` signal, which `allow_outputs_delete` waives). Production denies `unlink`/`rmdir` on a `rw` connected folder until approval, so `no_delete_in_outputs` covers only part of the rule. **only `true` is valid** |
18
+ | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
19
+ | `input_unmodified: <glob> \| [<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
20
+ | `self_heal_ran: <bool>` | a bash command the model wrote did (not) name a plugin under `/sessions/<id>/mnt/.local-plugins/` or `/sessions/<id>/mnt/.remote-plugins/` — the plugin-root self-heal path. It reads the command as the model wrote it, so a host plugin path that `hostloop`'s bash tool rewrote to the VM mount does not count |
21
+ | `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); evidence-unavailable on `lane: remote` / `preRunOrigin: remote-unavailable` |
22
+ | `artifact_text: {artifact, contains?: [..], not_contains?: [..], matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read. Fails on `lane: remote`: that lane's container filesystem is not locally observable, so there is no body to read |
23
+ | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
24
+ | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. `cowork-harness lint` reports it too (ERROR `scenario-invalid` — the wrapper runs the same loader); the bundled `scenario.py lint` run directly does NOT |
25
+ | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
26
+ | `tool_called: {tool: <glob> \| [..], input?, input_any?, result?, scope?, subagent_type?, count?}` | **object form** — asserts what a call CARRIED, where it RAN and what its paired result SAID, which no name glob and no `transcript_*` key can see (`transcript_*` reads top-level prose only, so it passes when the agent merely *says* it ran a command). `tool`: a glob or a list of globs (any-of; list both shells, `[Bash, mcp__workspace__bash]`, for a claim that must hold at hostloop too). `input: {<field>: <regex>}`: every named top-level input field must match (case-insensitive, unanchored; a missing field is no match; a non-string value is matched as JSON). `input_any: <regex>`: some top-level field matches. `result: {matches?, not_matches?, is_error?}`: predicates on the call's paired `tool_result` (by `toolUseId`, 10,240 characters of text — 32,768 for a top-level `Skill` result). `scope`: `main` (default — the main agent, including a `Skill`'s or `Agent(fork)`'s children, the set the string form reads), `subagent` (the parent is a sub-agent dispatch this run recorded, at any depth), or `any` (also calls whose parent is not a recorded dispatch, e.g. a `Skill` invoked inside a sub-agent). `subagent_type` (with `scope: subagent`): regex over the IMMEDIATE parent dispatch's type or description. `count: {min?, max?}` (default `min: 1`). **Fails closed:** an unpaired call never satisfies `result`; a truncated result or input that cannot settle a predicate, or a result.json without `toolCalls`, reports *evidence unavailable*, never a pass or a plain "not called". A red lists the calls it considered and names matches in another scope ("1 matching call in scope subagent — set `scope: any`"). `{tool: X}` alone is exactly `tool_called: X`. A cassette using this form stamps **v13**. |
27
+ | `tool_not_called: {tool: <glob> \| [..], input?, input_any?, result?, scope?, subagent_type?}` | **object form** — no call in scope satisfies every predicate (no `count`). Same fields and fail-closed rules as above: a candidate whose result is unpaired or truncated, or whose input was truncated, is *evidence unavailable*, never a pass. Refused at load when EVERY listed tool is one the tier does not serve. ⚠️ **Redaction hazard — this is the dangerous direction.** A committed cassette is redacted, which rewrites the recorded inputs: an `input` regex naming a literal the policy rewrites (a home path, an email) cannot see its target on replay. On replay, any candidate call whose field (or paired result, for `result.matches`) carries a redaction token is *evidence unavailable*, whatever the regex — the evaluator cannot know what the token replaced — so a negative check never passes over rewritten bytes. `record` warns — naming the redacted call and the ways out (narrow `scope`/`tool`, the string form, or live-only) — and refuses the write (the exact guard), and `lint` flags a regex naming a redactable literal (`tool-input-regex-redactable`) — a heuristic that reads `.cowork-redact.json` from the current and scenario directories only, not the cassette's. Match a part of the input redaction leaves alone (the verb and flags, a workspace-relative path), or keep the check on a live gate. Note also that the default `scope: main` does not see a sub-agent's call — use `scope: any` for "nothing anywhere ran X". |
28
+ | **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
29
+ | **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
30
+ | `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
31
+ | `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
32
+ | `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
33
+ | `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
34
+ | `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
35
+ | `tool_result_not_matches: <regex>` | the regex sibling of `tool_result_not_contains` — same fails-loud-on-absent-evidence semantics |
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog
2
2
 
3
- Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Every `assert:` key with its semantics, and the
3
+ Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Every `assert:` key with its semantics, and the
4
4
  verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
5
5
  the scenario and session YAML fields are there too.
6
6
 
@@ -14,109 +14,20 @@ and is not authorable (see gotcha 18 in [`scenario-schema.md`](./scenario-schema
14
14
 
15
15
  Looking for a key *by what you want to prove* (tool health, sub-agent work, panels, skill
16
16
  attribution, resources, diagnostics)? `assertions-guide.md`'s "goal → key" map is the by-purpose index into this
17
- table; the table below is the full per-key reference, and `cowork-harness assertions --list` prints the
17
+ table; the per-key tables in the three files below are the full reference, and `cowork-harness assertions --list` prints the
18
18
  same set live from the schema.
19
19
 
20
- | Assertion | Passes when |
21
- |---|---|
22
- | `result: success \| error` | the run ended with that status |
23
- | `transcript_contains: <str>` | the assistant transcript includes the literal string **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
24
- | `transcript_not_contains: <str>` | it does not **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
25
- | `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
26
- | `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
27
- | `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it. Object form `{path, authored}`: `authored: true` also requires that THIS run created or rewrote the file (an untouched pre-run file — e.g. from a `workspace_fixture` — fails; no pre-run manifest ⇒ evidence-unavailable); `authored: false` states that inheriting it is fine. On a `workspace_fixture` path one of the two is REQUIRED (refused at load otherwise) |
28
- | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected. Takes the same `{path, authored}` object form as `file_exists` (and `artifact_text`/`artifact_json` take an `authored` field) |
29
- | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. Checks on **every** baseline. Omitting it: on a baseline recording outputs `rw` (Desktop before 2.16120.0) a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored (`allow_outputs_delete: true` accepts an intended one); on `rwd` (2.16120.0+, including `latest`, where Cowork allows deletes in outputs) nothing checks outputs deletes, so author this key to keep the check. Detection is a post-run bash-command scan plus a per-turn filesystem diff of `outputs/`, not mount enforcement, so a green means none was *detected*. Fails when the diff proves the delete, when a delete in command/call position has an `outputs/` path as its own operand, or when the diff could not verify; a hit resting only on the detector's inference (e.g. a Python variable named `rm`) with a clean diff passes, and the `outputs_delete_unconfirmed` warn is still raised in the run output — see that code for the classes of real delete that land there |
30
- | `no_delete_in_mounts: true` | no delete op touched `outputs` or any `rw` connected folder, except mounts waived by `allow_delete_in`. Covers `outputs` on **every** baseline, including those where Cowork allows an outputs delete (unless outputs is waived, authoring it arms the outputs check, so a filesystem-proven delete fails as the `outputs_delete` signal, which `allow_outputs_delete` waives). Production denies `unlink`/`rmdir` on a `rw` connected folder until approval, so `no_delete_in_outputs` covers only part of the rule. **only `true` is valid** |
31
- | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
32
- | `input_unmodified: <glob> \| [<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
33
- | `self_heal_ran: <bool>` | a bash command the model wrote did (not) name a plugin under `/sessions/<id>/mnt/.local-plugins/` or `/sessions/<id>/mnt/.remote-plugins/` — the plugin-root self-heal path. It reads the command as the model wrote it, so a host plugin path that `hostloop`'s bash tool rewrote to the VM mount does not count |
34
- | `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); evidence-unavailable on `lane: remote` / `preRunOrigin: remote-unavailable` |
35
- | `artifact_text: {artifact, contains?: [..], not_contains?: [..], matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read |
36
- | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
37
- | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. `cowork-harness lint` reports it too (ERROR `scenario-invalid` — the wrapper runs the same loader); the bundled `scenario.py lint` run directly does NOT |
38
- | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
39
- | `tool_called: {tool: <glob> \| [..], input?, input_any?, result?, scope?, subagent_type?, count?}` | **object form** — asserts what a call CARRIED, where it RAN and what its paired result SAID, which no name glob and no `transcript_*` key can see (`transcript_*` reads top-level prose only, so it passes when the agent merely *says* it ran a command). `tool`: a glob or a list of globs (any-of; list both shells, `[Bash, mcp__workspace__bash]`, for a claim that must hold at hostloop too). `input: {<field>: <regex>}`: every named top-level input field must match (case-insensitive, unanchored; a missing field is no match; a non-string value is matched as JSON). `input_any: <regex>`: some top-level field matches. `result: {matches?, not_matches?, is_error?}`: predicates on the call's paired `tool_result` (by `toolUseId`, 10,240 characters of text — 32,768 for a top-level `Skill` result). `scope`: `main` (default — the main agent, including a `Skill`'s or `Agent(fork)`'s children, the set the string form reads), `subagent` (the parent is a sub-agent dispatch this run recorded, at any depth), or `any` (also calls whose parent is not a recorded dispatch, e.g. a `Skill` invoked inside a sub-agent). `subagent_type` (with `scope: subagent`): regex over the IMMEDIATE parent dispatch's type or description. `count: {min?, max?}` (default `min: 1`). **Fails closed:** an unpaired call never satisfies `result`; a truncated result or input that cannot settle a predicate, or a result.json without `toolCalls`, reports *evidence unavailable*, never a pass or a plain "not called". A red lists the calls it considered and names matches in another scope ("1 matching call in scope subagent — set `scope: any`"). `{tool: X}` alone is exactly `tool_called: X`. A cassette using this form stamps **v13**. |
40
- | `tool_not_called: {tool: <glob> \| [..], input?, input_any?, result?, scope?, subagent_type?}` | **object form** — no call in scope satisfies every predicate (no `count`). Same fields and fail-closed rules as above: a candidate whose result is unpaired or truncated, or whose input was truncated, is *evidence unavailable*, never a pass. Refused at load when EVERY listed tool is one the tier does not serve. ⚠️ **Redaction hazard — this is the dangerous direction.** A committed cassette is redacted, which rewrites the recorded inputs: an `input` regex naming a literal the policy rewrites (a home path, an email) cannot see its target on replay. On replay, any candidate call whose field (or paired result, for `result.matches`) carries a redaction token is *evidence unavailable*, whatever the regex — the evaluator cannot know what the token replaced — so a negative check never passes over rewritten bytes. `record` warns — naming the redacted call and the ways out (narrow `scope`/`tool`, the string form, or live-only) — and refuses the write (the exact guard), and `lint` flags a regex naming a redactable literal (`tool-input-regex-redactable`) — a heuristic that reads `.cowork-redact.json` from the current and scenario directories only, not the cassette's. Match a part of the input redaction leaves alone (the verb and flags, a workspace-relative path), or keep the check on a live gate. Note also that the default `scope: main` does not see a sub-agent's call — use `scope: any` for "nothing anywhere ran X". |
41
- | **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
42
- | **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
43
- | `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
44
- | `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
45
- | `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
46
- | `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
47
- | `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
48
- | `tool_result_not_matches: <regex>` | the regex sibling of `tool_result_not_contains` — same fails-loud-on-absent-evidence semantics |
49
- | `subagent_tool_used: <glob>` | a sub-agent used a tool matching this glob (same `*`/`?`, anchored, case-sensitive semantics as `tool_called`, including the empty/regex-ish rejection) |
50
- | `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
51
- | `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
52
- | `subagent_file_write: {path?, path_suffix?, tool?}` | a sub-agent-origin write attempt whose raw path equals `path` (exact) or ends with `path_suffix` has a paired non-error tool_result — the causal half of a delivery probe; requires one of `path`/`path_suffix`; `tool` defaults to Write/Edit/MultiEdit; content-class; tier-agnostic |
53
- | `subagent_dispatch_healthy: {type?, delivered?, path?, path_suffix?, no_vm_paths?}` | **`fidelity: hostloop` only** — composite: selects dispatch(es) via `type` (same matching as `subagent_dispatched`; omit to require every dispatch) and, for EACH selected dispatch, checks it (not just any sub-agent) delivered a paired non-error write (`delivered`, default true — narrowed by `path`/`path_suffix`, same exact-vs-suffix precedence as `subagent_file_write`) and made no `/sessions` VM-path attempt (`no_vm_paths`, default true) — both scoped to that dispatch's OWN `parentToolUseId`, the per-dispatch correlation `subagent_file_write` (which matches ANY sub-agent write) cannot express; a `type` that matches no dispatch FAILS; content-class (`RunResult.fileToolAttempts` + `RunResult.toolResults`); any non-hostloop tier FAILS "cannot verify" |
54
- | `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
55
- | `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
56
- | `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut. **Covers what the run dispatches** (`Agent`/`Task`, including `Agent(subagent_type:"fork")`), **not a `context: fork` skill** invoked through the `Skill` tool: that skill's own answer is never a dispatch — it comes back as the `Skill` tool result, which agent 2.1.284 builds as `Skill "<name>" completed (forked execution).`, a `Result:` line, then the answer. Assert on it with `tool_result_matches` anchored on that prefix, e.g. `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` (`[\s\S]*` because `.` stops at a newline; the match is case-insensitive, has no multiline flag so `^` is the start of the result, and sees the first 10,240 characters of each result — 32,768 for a top-level `Skill` result). This covers a **foreground** fork only: a backgrounded fork's result is the line `Skill "<name>" launched (forked execution, running in the background).`, which carries no answer |
57
- | `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 in `scenario-schema.md` |
58
- | `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked — via the `Skill` tool, or by a prompt starting `/<skill> …` / `/<plugin>:<skill> …` (the agent expands that itself with no `Skill` call; recorded as `slashInvokedSkills`) — evidence-unavailable (not a normal fail) if neither matched and the agent's init tools have no `Skill` tool, or the leading `/name` can't be resolved (ambiguous bare name, no skill inventory). A green here means the AGENT expanded the slash; real Cowork's app resolves a typed slash first and has refused a bare name that differs from its plugin's name (Desktop 2.19675.0, 2026-10-03, 4 runs), a refusal no run can see — pick the skill from the menu or name it like its plugin (`/<plugin>:<skill>` was not measured with a single copy installed) |
59
- | `no_skill_triggered: <regex>` | no invoked skill id matched, counting a slash-command invocation as well as a `Skill` call — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent, the `Skill` tool is unobservable, or the prompt's leading `/name` can't be resolved |
60
- | `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
61
- | `connector_available: <regex>` | an MCP server/connector's name matched the regex (available, not necessarily used) — evidence-unavailable if `RunResult.context.mcpServers` is absent |
62
- | `tool_available: <regex>` | a tool in the init manifest matched the regex (available, not necessarily called — see `tool_called` for invocation) — evidence-unavailable if `RunResult.context.tools` is absent. The `mcp__skills__*`/`mcp__plugins__*` discovery tools are modeled (as `alwaysLoad`) on `container`/`hostloop`/`cowork` — a miss there is a real absence; `microvm`/`protocol` declare no such server, so a miss on those two tiers means "not modeled at this tier", not "provably unavailable" |
20
+ The per-key table is split by assertion family so each file stays well within what one Read returns whole:
21
+
22
+ - [`assertion-catalog-outcome-files-tools.md`](./assertion-catalog-outcome-files-tools.md): result, transcript_*, file and artifact, delete and write-back, tool_called / tool_not_called, reference access, tool_result_*.
23
+ - [`assertion-catalog-agents-skills-budgets.md`](./assertion-catalog-agents-skills-budgets.md): subagent_*, dispatch, skills and connectors, cost / token / turn / tool-error budgets, tasks, scratchpad and present_files delivery.
24
+ - [`assertion-catalog-gates-hooks-modifiers.md`](./assertion-catalog-gates-hooks-modifiers.md): AskUserQuestion gates, hooks, path denial, allow_* verdict modifiers, host-path leaks, egress, MCP, RSS, semantic judging, artifact_json, computer:// links.
63
25
 
64
26
  **Name-matching styles differ by key (don't mix them up):** `tool_called`, `tool_not_called`,
65
27
  `subagent_tool_used`, `subagent_tool_absent` are **glob** (anchored, case-sensitive, `*`/`?`). `tool_available`,
66
28
  `skill_triggered`/`no_skill_triggered`, `skill_available`, `connector_available`, `skill_tool_used`,
67
29
  `subagent_type` are **regex** (unanchored, case-insensitive). So `tool_called: mcp__workspace__*` (glob) but
68
30
  `tool_available: mcp__workspace__.*` (regex) — a `.*` in a `tool_called` glob is a load-time schema error, not silently-matches-nothing.
69
- | `skill_tool_used: {skill, tool}` | a tool whose name matches `tool` ran inside a skill-activation window whose `skillId` matches `skill` (`RunResult.skillActivity`) — evidence-unavailable if skill-activity telemetry is absent; heuristic for inline skills (a sticky, sequential window matching the agent's `activeSkill` scope, not an exact per-tool boundary). **Scope:** the window's tool counts **include sub-agent calls** made during it, so this key can't say which agent called (`subagent_tool_used` is the sub-agent-only claim), and it matches tool **names** only — never the path/args, so "did it read *this* file" isn't expressible (per-sub-agent reads are recorded at `subagents[].referencesRead`, readable but not assertable) |
70
- | `max_cost_usd: <N>` | the run's SDK-reported cost is ≤ N USD (the agent session only — the `semantic_matches` judge and the LLM decider (`on_unanswered: llm` / `--decider-llm`) are separate model calls that are **not** included) — evidence-unavailable if cost telemetry is absent. **Replay asserts the frozen recording's cost, not fresh spend** — a real regression needs a live `run` |
71
- | `max_tokens: <N>` | `usage.input_tokens + usage.output_tokens` ≤ N (cache tokens excluded) — same replay caveat as `max_cost_usd` |
72
- | `tool_calls_max: <N>` | total top-level tool calls (sum of `toolCounts`) ≤ N — meaningfully replay-checkable (re-drive recomputes `toolCounts` deterministically) |
73
- | `tool_no_error: <regex>` | no tool whose name matches the regex recorded any error (`RunResult.toolErrors[name].errors === 0` for every match) — **requires ≥1 matching tool call** (a regex matching nothing fails, so a typo can't silently pass); evidence-unavailable if tool-error telemetry is absent |
74
- | `tool_no_error_if_called: <regex>` | like `tool_no_error` but passes vacuously when no tool matches the regex — the presence-free variant |
75
- | `max_tool_errors: <N>` | total tool errors across all tools (sum of `RunResult.toolErrors[*].errors`) ≤ N — evidence-unavailable if tool-error telemetry is absent |
76
- | `max_redundant_tool_calls: <N>` | total WASTED repeated tool calls (sum of `(count-1)` across every redundant `{name,args}` group in `RunResult.redundantToolCalls`) ≤ N — not the raw count of redundant groups; evidence-unavailable if redundant-call telemetry is absent |
77
- | `max_turns: <N>` | the SDK-reported (or fallback-counted) turn count ≤ N — meaningfully replay-checkable (re-drive recounts turns deterministically, same as `tool_calls_max`) |
78
- | `compaction_occurred: true` | a context-compaction boundary occurred (a `compact_boundary` system event was recorded) — lives in the stdout stream, so meaningfully replay-checkable; evidence-unavailable if context-event telemetry is absent. **Only `true` is valid** — omit to not require it |
79
- | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
80
- | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
81
- | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
82
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, an empty macOS system dir from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
83
- | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
84
- | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
85
- | `question_options: {when_question?, equals?: [..], contains?: [..], order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
86
- | `question_context: {when_question?, matches}` | a regex over **everything the gate showed the user** — question label + every option label + every option **description**. Use it when the sentence you need to prove reached the founder may land in any of those fields: `question_asked` sees only the question text, `question_options` compares only labels, so a phrase delivered in an option's `description` is invisible to both. `when_question` narrows; omitting it searches every gate (NOT ambiguous here — this key asks whether the text was shown, not which gate offered which set). Ask-time payload only, never a `tool_result` (a producer that also writes the phrase to its own gate-state file would otherwise false-green it). Zero gates FAILS ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
87
- | `question_option_count: {matches, exactly? \| min? \| max?, when_question?, case_sensitive?}` | counts option **labels** matching a regex on **every** selected sub-question (e.g. one `No changes — ` option per gate); zero asked FAILS. `case_sensitive` covers the whole pattern; single-quote regexes. Details: [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant. |
88
- | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
89
- | `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
90
- | `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
91
- | `gate_answer_count_min: <N>` | at least N AskUserQuestion gates fired AND were delivered non-error — presence companion to `gate_answers_delivered`'s vacuous-pass. **`: 0` asserts nothing** and does not satisfy that pairing; `>= 1` is **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
92
- | `hook_blocked: <regex>` | a PreToolUse hook blocked a tool whose name matches the regex (`RunResult.hookEvents`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream). **Does not see the agent's own refusal**: on `hostloop` against Desktop 2.7032.0+ a relative `Read`/`Write`/`Edit` is denied by the agent's permission rules before any hook runs, so neither this key nor `path_denied` records it — assert it with `tool_result_contains: "denied by your permission settings"` |
93
- | `no_hook_blocked: true` | no tool was hook-blocked during the run (distinguishes a real tool crash from an intentional hook block) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette. **Only `true` is valid** |
94
- | `hook_event_fired: <HookEvent>` | a **command hook** for this event (a plugin's `hooks/hooks.json` or manifest hook — `Stop`, `SessionStart`, `PostToolUse`, …) ran: a `hook_response` system frame with that `hook_event` was recorded (`RunResult.contextEvents`). Any outcome counts. The harness passes `--include-hook-events` whenever a staged plugin declares hooks — that is what puts events other than SessionStart/Setup on the stream — so a recording made without it reports "never fired". Content-class, grades on replay. Recorded end-to-end for `Stop` ([stop-hook-probe.scenario.yaml](https://github.com/yaniv-golan/cowork-harness/blob/main/examples/probes/stop-hook-probe.scenario.yaml)); the other names match the same frame but have not each been recorded |
95
- | `hook_event_blocked: <HookEvent>` | that command hook **blocked** at least once — a `hook_response` frame for the event carried `exit_code: 2`. Fails naming the exit codes seen when it fired without blocking (a frame with no `exit_code` is reported as such, never counted); fails "never fired" otherwise; cannot-verify when the run has no context events. Content-class |
96
- | `hook_output_contains: {event, stream?, text \| matches}` / `hook_output_not_contains: {…}` | a command hook **printed** (or never printed) a text on `stdout` / `stderr` / either (`stream`, default `any`) — some (or no) `hook_response` frame for `event` carries `text` (literal, case-sensitive) or `matches` (regex, case-insensitive, no multiline flag — `^`/`$` anchor the whole stream). For a hook that fails open on stderr while exiting 0. No frame for the event FAILS both. On a redacted stream, a literal miss and any `matches` result are evidence-unavailable for either key, while a literal hit outside a token counts; a miss is also evidence-unavailable on a missing field, a hook that started without a response, or output the agent truncated. Frames carry no plugin id (a second plugin, or at `protocol` without a sealed config dir a host-installed one, can answer). Content-class. Recorded end-to-end for `Stop` |
97
- | `vm_path_denied: true` | **`fidelity: hostloop` only** — at least one recorded path denial (`RunResult.pathDenials`, any source) targeted a `/sessions` VM path — evidence-unavailable if path-denial telemetry is absent. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
98
- | `path_denied: {tool?, path_matches?, source?, agent_scope?}` | **`fidelity: hostloop` only** — a path denial matching ALL given matchers (`tool` glob, `path_matches` regex, `source` ∈ pretooluse/can_use_tool/permission_denied, `agent_scope` ∈ main/subagent/any) was recorded. Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify" |
99
- | `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
100
- | `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
101
- | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
102
- | `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
103
- | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question or (after an `AskUserQuestion` gate) a closing request for input having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). On an open-ended `skill` / `probe-dispatch` run (no `assert:` block) pass `--allow-stall` instead. A helpful closing offer ("want me to run this through a structured pass?") also fails `stalled` — read the final message before believing it |
104
- | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
105
- | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Has an effect only on a baseline recording outputs `rw` (Desktop before 2.16120.0), where omitting `no_delete_in_outputs` does **not** permit deletes; on `rwd` (2.16120.0+) an outputs delete does not fail by default, so it is an accepted no-op (no warning), unless `no_delete_in_mounts` arms the outputs check, where it waives that check's signal as on `rw`. **Mutually exclusive** with `no_delete_in_outputs`. Silences `outputs_delete`, `outputs_delete_unconfirmed` and `outputs_diff_unavailable`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
106
- | `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
107
- | `transcript_no_host_path: true` | no host path (a path under a host home or system root: `Users`, `home`, `root`, the Cowork install dir `opt/cowork`, and the macOS `private/var`, `private/tmp`, `var/folders` and `Volumes` roots, each written here without its leading slash — also inside a `file://` or `computer://` link) leaked into model-visible text (a path that came verbatim from the scenario's input files, prompt, or declared plugins' or local skills' files is exempt; `scan.hostPathsFromInputs` counts such paths) — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
108
- | `egress_denied: <host>` | the host was blocked by the egress proxy |
109
- | `egress_allowed: <host>` | the host was allowed through |
110
- | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
111
- | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
112
- | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer: the final result text, the transcript and the files the agent authored. The transcript is top-level `assistant_text` only — it excludes every `tool_use`/`tool_result` and sub-agent text unless opted in. Full semantics (evidence scope, fork results, refusals, judge provenance, cost): [semantic-judging.md](semantic-judging.md). |
113
- | `semantic_pairwise: {refs?: [..], rubric?: [..], pass_if?, order?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge compares this run's judged document with a **frozen reference** — the same document an earlier run (usually the baseline) produced, written once by `ref freeze` and never regenerated — and answers win, tie, loss or `both_bad`. **The judged document is the one `semantic_matches` builds** (final message + transcript + authored files, same `evidence_files` / `include_subagent_text` / `include_fork_results` options), so the transcript **excludes every `tool_use`/`tool_result`** and a criterion about whether a tool was called cannot be judged from it. `refs` are reference stores relative to the scenario file; the run is judged against every one (under `hillclimb run`, against the flow's own references instead — only the baseline's gates `pass_if`, a later variant's is a metric recorded with `gate: false`), and `pass_if` (default `not_worse`) must hold against each gating one: `win`, `not_worse` (win or tie), or `any` (graded at all — a metric only). `both_bad` fails `win` and `not_worse`. Which output the judge sees first is a seeded coin per run, assert and reference (`order: both` judges both orders: a win in one and a loss in the other is position bias and scores as a tie; any other disagreement keeps the worse outcome; each order's own outcome is recorded as `pairwise[].orders.candidate_first` / `.ref_first`, absent for a single-order grade, and a judge that favours whichever output it sees first shows across runs as `candidate_first` winning more often than `ref_first`). The judge never sees the words "reference" or "baseline", and the stored reference is scrubbed with the grading process's secret set before it reads it (`pairwise[].refRedactions`). A reference must have been frozen with the same evidence options and for the same prompt; a missing, damaged, differently-scoped or other-prompt one **refuses the run before it spends anything**, as does a reference store inside any mounted source (the agent could read its own answer key). Unavailable evidence refuses with `semantic_matches`' typed reasons. Per-reference outcomes land in `assertions[].pairwise`. **LIVE-ONLY** — skipped on replay. |
114
- | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
115
- | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
116
- | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
117
-
118
- `expect_denied: [host, …]` adds one `egress_denied` per host. Run `cowork-harness assertions --list` for this
119
- table from the live schema. Example: `artifact_json: { artifact: outputs/cap.json, path: me.run_id, equals: "r1" }`.
120
31
 
121
32
  **Content correctness:** match the assertion to the deliverable. Prose → `transcript_matches`
122
33
  (regex, drift-tolerant) or `transcript_contains` (literal marker). `transcript_matches` is
@@ -129,7 +40,7 @@ dotted path.
129
40
 
130
41
  **VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; eleven
131
42
  are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
132
- `"success"`) and twelve are **warn**-severity (informational, never flip pass/fail). All twenty-three signal
43
+ `"success"`) and thirteen are **warn**-severity (informational, never flip pass/fail). All twenty-four signal
133
44
  codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
134
45
 
135
46
  | Code | Severity | Meaning |
@@ -156,8 +67,9 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
156
67
  | `undelivered_deliverables` | warn | The skill produced file(s) OUTSIDE every user-visible root and never delivered them, so they stay invisible to the user. Fires on every run without opting in, because the scenarios that most need it are the ones whose author never considered delivery. Silent when the evidence cannot answer the question (no workspace walk, a tier that runs no scratchpad walk, absent delivery telemetry, a resumed turn, or a lane where delivery is unobservable — see `delivery_unobservable`) — never a vacuous clean. **`lane: local` only**: on remote, delivery cannot be measured at all, so that lane reports `delivery_unobservable` instead of guessing. Opt out: `allow_undelivered_deliverables` |
157
68
  | `delivery_unobservable` | warn | `lane: remote` only — the run produced file(s), but whether any reached the user CANNOT be verified: nothing is delivered by location on that lane and the harness models no remote delivery tool (production uses the agent-native `SendUserFile`). The honest counterpart to `undelivered_deliverables`, which would otherwise fire on every remote run that writes anything — a signal that always fires carries no information. Mutually exclusive with it; quiet when the run produced nothing to deliver. A harness coverage gap, not a skill defect. Opt out: `allow_undelivered_deliverables` |
158
69
  | `partly_scripted_gate` | warn | A question batch (one `AskUserQuestion` with several sub-questions) that the scenario's `answers:` matched only PART of. Answers are delivered atomically, so the whole batch went to the `on_unanswered` fallback and the matched answers were NOT delivered — the fallback may contradict them. The message names the matched and unmatched sub-questions and who answered instead; `result.partlyScriptedGates` carries the full lists. Fires only on a partial match within one batch (a single-question gate no rule matched is the ordinary unanswered case). Re-derived on `replay` (from the cassette's frozen answers) and `verify-run` (from the current scenario). Fix: script every sub-question of the batch |
70
+ | `parked_at_question` | warn | The session declares `answer_channel: none` and the run ended on a question nobody can answer: its last message ends in `?`, even after tools ran (a finished answer ending on an offer warns too; a request without `?` does not fire). Takes the place of the `stalled` fail and never changes the verdict: with no answer channel, stopping at the question is the contract. Completion is judged from the file assertions such a scenario must carry (refused at load without one). No opt-out needed |
159
71
  | `exec_infra_error` | warn | Host-loop: one or more container `exec` calls failed for infrastructure reasons, so those tool calls returned an error to the agent. Warns rather than fails because the run's other evidence is intact — unlike `infra_error`, where a dead supervisor contaminates everything. Caveat: if *every* exec failed, the agent ran nothing and this still only warns — check `result.infraErrors` |
160
72
 
161
73
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
162
74
  overall run verdict and exit code — `assert result: success` alone won't catch it; check
163
- `result.verdict.signals[].severity` or the run's exit code. Only the twelve **warn** codes are truly benign.
75
+ `result.verdict.signals[].severity` or the run's exit code. Only the thirteen **warn** codes are truly benign.
@@ -1,6 +1,6 @@
1
1
  # Assertions guide
2
2
 
3
- Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
3
+ Tracks `cowork-harness 4.6.0` (baseline `desktop-2.26454.2`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog starts at `assertion-catalog.md`, which indexes the per-family files that hold every key's row.
4
4
 
5
5
  ### Assertions: two orthogonal axes
6
6
 
@@ -11,7 +11,7 @@ Conflating these is the **biggest landmine**. An assertion key has two independe
11
11
  not: match prose with `transcript_matches` / `transcript_contains` (stable lexical markers only —
12
12
  not semantic content the model paraphrases, which re-records red); check structured JSON with YAML
13
13
  `artifact_json` (or the [pytest lane](https://github.com/yaniv-golan/cowork-harness/blob/main/python/README.md) for complex predicates), not via a transcript substring.
14
- To check a command that RAN (not one the agent mentioned), use the object form `tool_called: {tool: <name>, input: {command: <regex>}, scope?, result?}` — see [assertion-catalog.md](./assertion-catalog.md).
14
+ To check a command that RAN (not one the agent mentioned), use the object form `tool_called: {tool: <name>, input: {command: <regex>}, scope?, result?}` — see [assertion-catalog-outcome-files-tools.md](./assertion-catalog-outcome-files-tools.md).
15
15
  - **Axis B — survives `replay`?** *Independent of Axis A.* On the token-free `replay` lane, only
16
16
  **content keys** evaluate; filesystem / egress keys are skipped (live-only) — loudly, via an
17
17
  `::warning::` annotation, not a silent no-op. A key
@@ -21,7 +21,7 @@ Getting Axis B wrong means a check that **does nothing in CI** — the harness w
21
21
  (an `::warning::` annotation, not a silent no-op — see the Axis B bullet above), and the bundled linter
22
22
  catches it before you push — run it (see *Scaffold a valid scenario, then lint before you push* in `authoring.md`).
23
23
 
24
- See `references/assertion-catalog.md` for the full assertion catalog, and `references/scenario-schema.md`'s
24
+ See `references/assertion-catalog.md` for the full assertion catalog (an index into its per-family files), and `references/scenario-schema.md`'s
25
25
  *Replay class* for which keys survive `replay`.
26
26
 
27
27
  #### Which assertion for which question (goal → key)
@@ -50,7 +50,9 @@ them by what you're trying to prove:
50
50
  | the user was **shown** the right choices, in order | `question_options: {when_question, equals: [..]}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
51
51
  | the user was **told something specific** at a gate | `question_context: {when_question, matches}` — a regex over the question label + option labels + option **descriptions**. Reach for this when the wording may land in an option's `description`, which `question_asked` and `question_options` cannot see |
52
52
  | every gate offered exactly one (or at most N) options of a kind | `question_option_count: {matches, exactly}` — counts the option LABELS matching a regex on EVERY sub-question asked (`when_question` narrows); zero sub-questions asked fails |
53
- | a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
53
+ | every gate was answered by a scripted rule (an unattended run that still asks) | `gates_all_scripted: true` with `gate_answer_count_min: 1` — fails naming any gate the LLM decider, `first`, or an external or human decider answered; `{include_permissions: true}` also fails cowork parity's permissive auto-allow (replay needs a `controlOut` cassette) |
54
+ | a **plugin's** command hook (any event: PreToolUse, Stop, …) blocked, didn't block, or decided | `hook_event_blocked: <event>` (exit 2) or `{event, tool?, via?, max: 0}`, `no_hook_event_blocked: true`, `hook_decision: {event, decision}` — stream content, no `controlOut` needed |
55
+ | the **harness's own** hook callbacks (the built-in Task hook, a custom hook bundle) blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) — a plugin's hook never reaches this list, so `no_hook_blocked` passes over a plugin block |
54
56
  | every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
55
57
  | a context compaction happened | `compaction_occurred: true` |
56
58
  | THIS run wrote the file (not just that it is there) | `file_exists: {path, authored: true}` (also `user_visible_artifact`, and `authored: true` on `artifact_text`/`artifact_json`) — an untouched pre-run file fails; needs the pre-run manifest, which `authored` arms |