cadet-agent 0.56.0 → 0.60.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -7,9 +7,9 @@ Cadet-Agent is **not a one-shot code generator**. It won't spit out a finished g
7
7
  ## Repository Layout
8
8
  - `.cadet/agent/core/` contains the shared Cadet-Agent framework documents.
9
9
  - `cadet-agent.md` is the thin global directive: identity, non-negotiable rules, workflow routing, hard-gate protocol, and skill dispatch.
10
- - `Harness.md` is the canonical harness contract: budgets, evidence-backed gates, retries, context tiers, tool routing, privacy, and escalation.
10
+ - `HarnessRuntime.md` is the lean runtime contract; `Harness.md` keeps the full harness rationale and reference.
11
11
  - `harness.schema.json` and `state.schema.json` are the machine-readable schemas for harness records and session state.
12
- - `skills/` contains scoped workflow-phase skills (PlanningReview, Requirements, Architecture, Spike, StoryBreakdown, TDD, Debugging, CodeReview, Resume, MCPSetup, AgentReviewer, Handoff, Reconciliation).
12
+ - `skills/` contains scoped workflow-phase skills (PlanningReview, Requirements, Architecture, DesignReview, Spike, StoryBreakdown, TDD, Debugging, CodeReview, VisualEvidence, Resume, MCPSetup, AgentReviewer, Handoff, Reconciliation).
13
13
  - `templates/` contains runtime templates for planning artifacts.
14
14
  - `.cadet/harness.json` holds repository-local budget/policy overrides (preserved by sync).
15
15
  - `.cadet/runs/` holds sanitized run ledgers (preserved by sync; no secrets or raw prompts by default).
@@ -37,11 +37,13 @@ Cadet-Agent provides the **same skills** across six IDEs — one canonical file
37
37
  | Planning Review | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
38
38
  | Requirements | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
39
39
  | Architecture | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
40
+ | Design Review | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
40
41
  | Spike | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
41
42
  | Story Breakdown | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
42
43
  | TDD | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
43
44
  | Debugging | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
44
45
  | Code Review | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
46
+ | Visual Evidence | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
45
47
  | Resume | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
46
48
  | MCP Setup | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
47
49
  | Reconciliation | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
@@ -118,12 +120,12 @@ flowchart TD
118
120
  REPORT["Report current phase,<br/>epics, stories & gates"]
119
121
  CR["🔍 Context Resolution<br/>classify change size,<br/>calibrate learner,<br/>detect policy"]
120
122
  REQ["📋 Requirements<br/>Given/When/Then criteria<br/>assumption audit"]
121
- PLANREV["🧩 Planning Review<br/>interview one question<br/>at a time · decision tree"]
123
+ PLANREV["🧩 Planning Review (skill)<br/>interview one question<br/>at a time · decision tree"]
122
124
  ARCH["🏗️ Architecture<br/>technical design,<br/>ADR decisions"]
123
125
  SPIKE["🧪 Spikes<br/>resolve unverified<br/>assumptions"]
124
126
  BREAKDOWN["📐 Story Breakdown<br/>epics → testable stories"]
125
127
  IMPL["🔨 Implementation<br/>TDD per story,<br/>red → green → refactor"]
126
- REVIEW["✅ Review<br/>hard gate: 17-step<br/>code review, security"]
128
+ REVIEW["✅ Review<br/>hard gate: 23-step<br/>code review, security"]
127
129
  VALIDATE["✔️ Validation<br/>acceptance criteria,<br/>design artifact sync"]
128
130
  CLOSED(["🎉 Closed"])
129
131
  NEXT_STORY{"More stories<br/>in epic?"}
@@ -140,13 +142,13 @@ flowchart TD
140
142
 
141
143
  REQ --> ARCH
142
144
  ARCH -->|"unverified assumptions"| SPIKE
143
- ARCH -->|"all assumptions resolved"| BREAKDOWN
145
+ ARCH -->|"all assumptions resolved<br/>gate: designReviewCompleted ✅ (opt-in)"| BREAKDOWN
144
146
  SPIKE -->|"spike complete"| ARCH
145
147
 
146
148
  BREAKDOWN --> IMPL
147
149
 
148
150
  IMPL -->|"story complete"| REVIEW
149
- REVIEW -->|"gate: codeReviewCompleted ✅<br/>gate: securityReviewPassed ✅<br/>gate: reachabilityAddressed ✅ (opt-in)"| VALIDATE
151
+ REVIEW -->|"gate: codeReviewCompleted ✅<br/>gate: securityReviewPassed ✅<br/>gate: acceptanceCriteriaValidated ✅<br/>gate: reachabilityAddressed ✅ (opt-in)<br/>gate: userPlaythroughConfirmed ✅ (opt-in)"| VALIDATE
150
152
  VALIDATE -->|"gate: designArtifactSyncConfirmed ✅"| NEXT_STORY
151
153
  NEXT_STORY -->|"yes"| IMPL
152
154
  NEXT_STORY -->|"no"| CLOSED
@@ -165,13 +167,15 @@ Use the `/cadet-resume` slash command to pick up where you left off. It reads `.
165
167
 
166
168
  ### Phase Gating
167
169
 
168
- Hard gates are enforced at every phase transition. The agent reads `.cadet/state.json → gates` before advancing and **blocks** the transition if any required gate is `false`. Gates cannot be skipped without an explicit user-directed exception recorded in `changeHistory`.
170
+ Hard gates are enforced at every phase transition. The agent reads `.cadet/state.json → gates` before advancing and **blocks** the transition if any required gate is `false`. Gates cannot be skipped without an explicit, user-directed exception recorded in `.cadet/state.json → gateExceptions` — bounded by `expiresAt`, and naming the work items it covers. A `v1`–`v3` exception is read out of `changeHistory`, its older home.
169
171
 
170
172
  | Transition | Required Gates |
171
173
  |---|---|
172
174
  | architectureComplete → story-breakdown | `designReviewCompleted` when `designReview.enabled` is set — the formal design review, recorded by `harness verify-design-review` |
173
175
  | implementation → review | `testsPassed`, `compileCheckConfirmed`, `unityAnalyzerClean`, `storyTrackingUpdated`, and `architectureFitnessPassed` when the project declares architecture checks and enables them |
174
- | review → validation | `codeReviewCompleted`, `securityReviewPassed`, `acceptanceCriteriaValidated`, and `reachabilityAddressed` when `reachability.enabled` is set |
176
+ | review → validation | `codeReviewCompleted`, `securityReviewPassed`, `acceptanceCriteriaValidated`, `reachabilityAddressed` when `reachability.enabled` is set, and `userPlaythroughConfirmed` when `userPlay.enabled` is set — a story declares `Play: required — <what the user does and what they see>` or `Play: deferred to <work item> — <why>`. A person's own record satisfies `required` (`harness play-form` writes the form, `harness confirm --artifact` records it); `harness verify-play` records `deferred`, and the deferral expires when the named work item closes |
177
+ | validation → closed | `designArtifactSyncConfirmed`, and `humanAcceptanceConfirmed` when `humanAcceptance.enabled` is set — a person's own record (`harness acceptance-form` writes the form, `harness confirm --artifact` records it); no command can produce it |
178
+
175
179
  ### Runtime context protocol (opt-in, and the framework's own claim discipline)
176
180
 
177
181
  `harness context plan` states what a phase requires (with a reason and a hash for each reference),
@@ -181,42 +185,67 @@ reported as it is — a run reported as `recorded` is never reported as `enforce
181
185
  a hook that declares it enforces context. A required reference that was never loaded, or that changed
182
186
  after the record, blocks the checkpoint. See `.cadet/agent/core/Harness.md` §2d.
183
187
 
184
- | validation → closed | `designArtifactSyncConfirmed`, and `humanAcceptanceConfirmed` when `humanAcceptance.enabled` is set — a person's own record (`harness acceptance-form` writes the form, `harness confirm --artifact` records it); no command can produce it |
185
-
186
188
  **`closed` is end-of-epic, not per-story.** `validation → closed` is taken only when no stories remain (`NEXT_STORY → no → CLOSED` above). When an epic still has stories, the next story re-enters from `validation → implementation` (`NEXT_STORY → yes → IMPL`). Do not close a story individually: `closed` is terminal, and there is no transition out of it.
187
189
 
188
- The full set of legal transitions is the three gated rows above **plus** the ungated forward edges (classification, planning progression, `story-breakdown → implementation`, and the `validation → implementation` next-story loop). Any transition outside that set is rejected with a named reason.
190
+ The full set of legal transitions is the three gated rows above **plus** the ungated forward edges (classification, planning progression, `story-breakdown → implementation`, and the `validation → implementation` next-story loop). Any transition outside that set is rejected with a named reason. **Planning Review is a skill, not a phase**: the agent dispatches it before requirements or architecture when the plan is fuzzy or contested, and it records no transition.
191
+
192
+ ### The response contract
193
+
194
+ The framework's only per-reply output is one line: `cadet-agent: ok`, or the problem in its place. `cadet-agent harness status` derives that line — read-only — from the state document and the run ledger, so it cannot claim a health nothing verified. `ok` means the record is readable, valid and fresh, and no recorded run stopped on a budget or failed to run. A gate that is unmet because the story is unfinished is normal, and is not reported. The rest of a reply carries the work: the decisions taken, what changed, what the checks show, and what is unverified or deferred.
189
195
 
190
196
  ### Harness
191
197
 
192
198
  Gates are backed by **evidence**, not assertion. Each claimed gate must have a fresh, non-superseded evidence record bound to the current work item, input tree hash, and acceptance criteria. The harness also bounds context, tokens, tool calls, retries, wall-clock time, cost, and archive sizes — and those bounds are enforced, not advisory.
193
199
 
194
- - Rules: `.cadet/agent/core/Harness.md`. Data contract: `docs/core/HarnessContract.md`.
200
+ - Runtime rules: `.cadet/agent/core/HarnessRuntime.md`. Full contract: `.cadet/agent/core/Harness.md`. Data contract: `docs/core/HarnessContract.md`.
195
201
  - Overrides: `.cadet/harness.json` (preserved by sync; conservative defaults in `src/harness/policy.mjs`).
196
202
  - Ledgers: `.cadet/runs/<runId>.json` (sanitized; artifacts are redacted before they are written; no secrets or raw prompts by default).
197
203
  - Transitions recompute the input tree hash from the evidence's relevant files, so editing a relevant file invalidates the evidence.
198
204
  - `harness verify` binds evidence to `--files` (or the working tree's changed files), and a `testsPassed` green result requires a prior red record.
205
+ - A `--files` binding may not name a path the recording command itself writes: `--files .cadet/state.json` is refused before anything runs, because the write that follows would stale the record it just wrote.
199
206
  - When Git is unavailable and no `--files` are given, verification blocks (`freshness-unavailable`) rather than recording unscoped evidence.
200
207
  - `state validate` rejects a `true` gate whose evidence is missing, stale, expired, superseded, or bound to another work item; evidence records are schema-validated in full (`command`, `result`, `criteriaHash`, and a freshness bound).
208
+ - `state validate` errors when a work item that a `storyCompletions` row records as finished, with evidence behind it, still reads `planned`. The remedy is `done` or `superseded`: a finished story must not be indistinguishable from one that was never started.
209
+ - A `changeHistory` entry is a pointer, not a retelling: it is limited to 400 characters, and `state compact` archives a longer entry in place rather than truncating it.
201
210
  - Evidence must include a UUID, work item, phase, gate, status, command/result, input-tree hash, criteria hash, relevant files, timestamp, and either `expiresAt` or `freshnessPolicy`.
202
211
  - **Evidence history does not live in `state.json`.** A v4 document keeps only the active work item's records inline; a closed work item's evidence is written into the commit that closes it, as `Cadet-*` trailers, and archived to `.cadet/archive/`. `evidenceCoverage` indexes what left, so the "a done story owns evidence" check still works offline. Cadet still never commits: `state seal` prepares a message file and you commit with `git commit -F`.
212
+ - **Two gates are human-owned: `humanAcceptanceConfirmed` and `userPlaythroughConfirmed`.** No
213
+ command can produce them, and `--command` is refused for them. `harness acceptance-form --epic <id>`
214
+ and `harness play-form --story <path>` write the form; `harness confirm --gate <gate> --artifact <form>`
215
+ records what the person wrote. A form still holding a placeholder is refused, and a form someone has
216
+ started is never overwritten.
203
217
  - Command output counts against the output budget; a configured cost budget cannot be satisfied by unmeasurable cost (the run is blocked, `budget-blocked`).
204
218
  - State and run ledgers are written atomically, so an interrupted write cannot truncate a record; persisted artifacts are redacted before hashing or writing.
205
219
  - Empty freshness coverage is an explicit policy decision: set `allowEmptyFreshness: true` in `.cadet/harness.json` only when unscoped evidence is acceptable.
206
220
 
207
221
  ```bash
222
+ cadet-agent state init --workflow-path large # write the first state document (validated before it lands)
208
223
  cadet-agent state validate # validate state against the schema (read-only)
209
224
  cadet-agent state validate --verify-sealed # also read evidence out of commit trailers
210
225
  cadet-agent state migrate # atomically upgrade v1 → the current version
211
226
  cadet-agent state migrate --to 4 # archive closed work items' evidence; build the index
212
227
  cadet-agent state compact --keep active # routine housekeeping on a v4 state
213
228
  cadet-agent state seal # write the active work item's evidence as commit trailers
229
+ cadet-agent state begin --epic <id> --story <id> # start a work item; archive the previous item's evidence
214
230
  cadet-agent state transition --to review # enforce the matrix + evidence
231
+ cadet-agent harness status # the health line: ok, or the problem (read-only)
215
232
  cadet-agent harness verify --gate testsPassed --files src/a.cs # bounded, classified loop
233
+ cadet-agent harness verify-acs --story <path> # derive coverage from the run report, then record the gate
234
+ cadet-agent harness verify-reachability --story <path> # check the declaration; run the project probe
235
+ cadet-agent harness verify-design-review --artifact <path> --files <design,requirements,ADRs>
236
+ cadet-agent harness verify-architecture # run the project's declared fitness checks
237
+ cadet-agent harness verify-play --story <path> # check a story's `Play:` declaration; record a deferral
238
+ cadet-agent harness play-form --story <path> # write a user-playthrough form for a person to fill
239
+ cadet-agent harness acceptance-form --epic <id> # write the epic's human-acceptance form
240
+ cadet-agent harness confirm --gate <gate> --artifact <form> # record a person's own account
241
+ cadet-agent harness changes # the files a story changed, with links (read-only)
216
242
  cadet-agent harness report # budget consumption and failures (no secrets)
217
243
  cadet-agent harness reconcile # reconcile the planning chain against state.json (read-only)
244
+ cadet-agent harness matrix-check # reconcile a TDD matrix against the test inventory (read-only)
245
+ cadet-agent harness context plan|record|validate # plan what a phase loads, record what it loaded, decide
218
246
  cadet-agent harness cleanup --older-than-ms <n> # apply the retention policy (bound required)
219
247
  cadet-agent harness capabilities # available CLI/Unity/MCP/hook/token/cost telemetry
248
+ cadet-agent harness capabilities --verify-host # probe the configured interception, per action
220
249
  ```
221
250
 
222
251
  Every command supports `--format human|json` and exits nonzero for invalid state, failed verification, budget exhaustion, stale evidence, or safety rejection.
@@ -332,13 +361,14 @@ If a specific game repository needs local conventions, add a policy file under `
332
361
 
333
362
  ## Package Output
334
363
  Running `./package-agent.ps1` produces `cadet-agent.zip` with this layout:
335
- - `.cadet/agent/core/` (including `Harness.md`, `harness.schema.json`, and `state.schema.json`)
364
+ - `.cadet/agent/core/` (including `HarnessRuntime.md`, `Harness.md`, `harness.schema.json`, and `state.schema.json`)
336
365
  - `.cadet/agent/core/skills/`
337
366
  - `.cadet/agent/core/templates/`
338
367
  - `.github/agents/cadet.agent.md`
339
368
  - `.github/agents/cadet-agent-reviewer.agent.md`
340
369
  - `.github/prompts/cadet-*.prompt.md`
341
- - `.github/hooks/`
370
+ - `.github/hooks/` (the Copilot `git-guard` hook and its scripts)
371
+ - `.githooks/pre-commit` (the portable Git hook; installed by you with `git config core.hooksPath .githooks`)
342
372
  - `.cursor/rules/cadet-agent.md`
343
373
  - `.cursor/rules/cadet-agent-reviewer.md`
344
374
  - `.continue/rules/cadet-agent.md`
@@ -346,6 +376,10 @@ Running `./package-agent.ps1` produces `cadet-agent.zip` with this layout:
346
376
  - `.continue/config.yaml`
347
377
  - `.claude/skills/cadet-agent/SKILL.md`
348
378
  - `.claude/skills/cadet-*/SKILL.md`
379
+ - `.agents/skills/cadet-agent/SKILL.md` (the cross-client root Deep Code and Hermes read)
380
+ - `.agents/skills/cadet-*/SKILL.md`
381
+ - `AGENTS.md` (create-only: an existing consumer copy is never overwritten)
382
+ - `.cadet/harness.json` (create-only: the new-consumer policy seed)
349
383
 
350
384
  ## Notes
351
385
  - `.cadet/agent/core/FrameworkManifest.json` defines the managed and preserved paths for packaged installs.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cadet-agent",
3
- "version": "0.56.0",
3
+ "version": "0.60.0",
4
4
  "description": "Cross-IDE agent framework for Unity/C# game-development — one-command install",
5
5
  "type": "module",
6
6
  "bin": {
package/src/cli.mjs CHANGED
@@ -24,9 +24,13 @@ import {
24
24
  parseReachabilityDeclaration, validateReachabilityDeclaration, collectWorkItems,
25
25
  findDeferralCycles, readSiblingDeclarations, normalizeWorkItemRef, describeReachabilityGaps,
26
26
  REACHABILITY_GATE, runCommand,
27
+ parsePlayDeclaration, validatePlayDeclaration, readSiblingPlayDeclarations, describePlayGaps,
28
+ USER_PLAY_GATE, buildPlayForm, parsePlayForm, writePlayForm, readPlayTemplate,
29
+ PLAY_FORM_GATE, PLAY_FORM_SUFFIX,
27
30
  createEvidence, newId, computeInputTreeHash, hashCriteria,
28
31
  collectDeclaredTestNames, reconcileTestNames,
29
32
  resolveCommand, describeCommand, describeAllCommands, checkUnattendedRequirements, COMMANDS,
33
+ selfBoundFiles,
30
34
  STATE_VERSION, sealedEvidence, recordEvidence, appendEvidence, sealWorkItem, toStateV4,
31
35
  isHistoryExternal, HISTORY_ENTRIES_KEPT, resetGatesForNewWorkItem,
32
36
  } from './harness/index.mjs';
@@ -72,6 +76,8 @@ function showHelp() {
72
76
  cadet-agent harness verify Run a bounded, classified verification loop
73
77
  cadet-agent harness verify-acs Verify declared AC↔test coverage against a test report
74
78
  cadet-agent harness verify-reachability Verify a story's declared reachability (opt-in)
79
+ cadet-agent harness verify-play Verify a story's Play: declaration and record the deferral (opt-in)
80
+ cadet-agent harness play-form Write a user-playthrough form for a story, filled in from state (opt-in)
75
81
  cadet-agent harness verify-design-review Check the design-review artifact and record the gate (opt-in)
76
82
  cadet-agent harness acceptance-form Write a human-acceptance form for an epic, filled in from state
77
83
  cadet-agent harness verify-architecture Run the project's declared architecture checks (opt-in)
@@ -99,8 +105,8 @@ function showHelp() {
99
105
  --expires-at ISO-8601 expiry bounding the confirmation (harness confirm)
100
106
  --environment key=value,... describing what was verified (harness confirm)
101
107
  --scope Comma-separated scope of the confirmation (harness confirm)
102
- --story Story markdown declaring the acceptance criteria or reachability (harness verify-acs|verify-reachability)
103
- --artifact Design-review artifact to check (verify-design-review), or the acceptance form to record from (confirm --gate humanAcceptanceConfirmed)
108
+ --story Story markdown declaring the acceptance criteria, reachability or play declaration (harness verify-acs|verify-reachability|verify-play|play-form)
109
+ --artifact Design-review artifact to check (verify-design-review), or the acceptance/playthrough form to record from (confirm --gate humanAcceptanceConfirmed|userPlaythroughConfirmed)
104
110
  --epic Epic a generated form belongs to (harness acceptance-form)
105
111
  --out Where to write the form (default: the epic's plan directory)
106
112
 
@@ -976,40 +982,73 @@ async function cmdHarness(opts) {
976
982
  // the record. `--witness` and `--limitations` no longer exist; a caller that passes
977
983
  // them is refused here rather than silently recorded, which is why the check is on
978
984
  // the missing artifact rather than on the flags.
985
+ const FORM_RECORDED_GATES = new Set([HUMAN_ACCEPTANCE_GATE, USER_PLAY_GATE]);
979
986
  if (gate === HUMAN_ACCEPTANCE_GATE && !opts.artifact) {
980
987
  fail(opts, `a human acceptance is recorded from a form: run "cadet-agent harness acceptance-form --epic <epicId>" to write one pre-filled from state, fill in its blank fields, then pass it back with --artifact <the form>.`, () => 1, { ok: false, gate, code: 'acceptance-form-required' });
981
988
  }
989
+ // A playthrough is recorded from a form too, and from nothing else. The gate asks whether a
990
+ // PERSON played the delivered work, so the only route is that person's own account of it —
991
+ // which is also why `harness verify-play` refuses a `required` declaration.
992
+ if (gate === USER_PLAY_GATE && !opts.artifact) {
993
+ fail(opts, `a playthrough is recorded from a form: run "cadet-agent harness play-form --story <path>" to write one pre-filled from state, play the work, fill in its blank fields, then pass it back with --artifact <the form>.`, () => 1, { ok: false, gate, code: 'play-form-required' });
994
+ }
982
995
 
983
- // `harness acceptance-form` writes the form pre-filled from state; this reads it back.
996
+ // `harness acceptance-form` and `harness play-form` write the form pre-filled from state;
997
+ // this reads it back.
984
998
  if (opts.artifact) {
985
- if (gate !== HUMAN_ACCEPTANCE_GATE) {
986
- fail(opts, `--artifact is only for ${HUMAN_ACCEPTANCE_GATE}: every other gate's evidence comes from its own command or from explicit fields.`, () => 1, { ok: false, gate, code: 'artifact-not-applicable' });
999
+ const isPlaythrough = gate === USER_PLAY_GATE;
1000
+ if (!FORM_RECORDED_GATES.has(gate)) {
1001
+ fail(opts, `--artifact is only for ${HUMAN_ACCEPTANCE_GATE} and ${USER_PLAY_GATE}: every other gate's evidence comes from its own command or from explicit fields.`, () => 1, { ok: false, gate, code: 'artifact-not-applicable' });
987
1002
  }
988
1003
  const conflicting = ['scope', 'environment'].filter((k) => opts[k]);
989
1004
  if (conflicting.length > 0) {
990
- fail(opts, `--artifact already carries the acceptance, so ${conflicting.map((k) => `--${k}`).join(' and ')} would be a second, competing source. Pass the artifact alone.`, () => 1, { ok: false, gate, code: 'artifact-conflicts-with-flags', conflicting });
1005
+ fail(opts, `--artifact already carries the record, so ${conflicting.map((k) => `--${k}`).join(' and ')} would be a second, competing source. Pass the artifact alone.`, () => 1, { ok: false, gate, code: 'artifact-conflicts-with-flags', conflicting });
991
1006
  }
992
1007
  let formText;
993
1008
  try {
994
1009
  formText = readFileSync(opts.artifact, 'utf-8');
995
1010
  } catch (err) {
996
- fail(opts, `the acceptance form could not be read (${err.message}).`, () => 1, { ok: false, gate, code: 'artifact-unreadable' });
997
- }
998
- const form = parseAcceptanceForm(formText);
999
- if (form.incomplete.length > 0) {
1000
- fail(opts, `the form still has unfilled fields: ${form.incomplete.join(', ')}. Fill them in ${opts.artifact} and run the command again — an acceptance nobody wrote down is not an acceptance.`, () => 1, { ok: false, gate, code: 'acceptance-form-incomplete', missing: form.incomplete });
1011
+ fail(opts, `the ${isPlaythrough ? 'playthrough' : 'acceptance'} form could not be read (${err.message}).`, () => 1, { ok: false, gate, code: 'artifact-unreadable' });
1001
1012
  }
1002
- if (state && form.epic && state.epics && !state.epics[form.epic]) {
1003
- fail(opts, `the form accepts epic "${form.epic}", which does not exist in state.json. Fix the form, or accept the epic the repository actually has.`, () => 1, { ok: false, gate, code: 'acceptance-form-epic-unknown', epic: form.epic });
1013
+
1014
+ if (isPlaythrough) {
1015
+ const form = parsePlayForm(formText);
1016
+ if (form.incomplete.length > 0) {
1017
+ fail(opts, `the form still has unfilled fields: ${form.incomplete.join(', ')}. Fill them in ${opts.artifact} and run the command again — a playthrough nobody described is not a playthrough.`, () => 1, { ok: false, gate, code: 'play-form-incomplete', missing: form.incomplete });
1018
+ }
1019
+ // The record binds to the ACTIVE work item, so a form that plays a different story is
1020
+ // refused rather than recorded against the one in flight — the same failure the
1021
+ // story-bound evidence rules exist to prevent, wearing a filled-in form.
1022
+ const activeStory = state?.activeWorkItem?.storyId || null;
1023
+ const formStory = String(form.story || '');
1024
+ if (activeStory && formStory && !formStory.includes(activeStory)) {
1025
+ fail(opts, `the form plays "${form.story}", but the active work item is "${activeStory}". Record a playthrough of the work in flight, or begin that story first.`, () => 1, { ok: false, gate, code: 'play-form-story-mismatch', story: form.story, active: activeStory });
1026
+ }
1027
+ opts.witness = form.witness;
1028
+ opts.limitations = form.limitations;
1029
+ opts.environment = form.environment || null;
1030
+ opts.scope = [form.story || activeStory || 'playthrough'];
1031
+ if (!opts.files || opts.files.length === 0) opts.files = form.fileList;
1032
+ opts.files = opts.files && opts.files.length ? opts.files : null;
1033
+ opts.filesGiven = opts.files !== null;
1034
+ opts.playedBy = form.player;
1035
+ } else {
1036
+ const form = parseAcceptanceForm(formText);
1037
+ if (form.incomplete.length > 0) {
1038
+ fail(opts, `the form still has unfilled fields: ${form.incomplete.join(', ')}. Fill them in ${opts.artifact} and run the command again — an acceptance nobody wrote down is not an acceptance.`, () => 1, { ok: false, gate, code: 'acceptance-form-incomplete', missing: form.incomplete });
1039
+ }
1040
+ if (state && form.epic && state.epics && !state.epics[form.epic]) {
1041
+ fail(opts, `the form accepts epic "${form.epic}", which does not exist in state.json. Fix the form, or accept the epic the repository actually has.`, () => 1, { ok: false, gate, code: 'acceptance-form-epic-unknown', epic: form.epic });
1042
+ }
1043
+ opts.witness = form.witness;
1044
+ opts.limitations = form.limitations;
1045
+ opts.environment = form.environment || null;
1046
+ opts.scope = [form.epic];
1047
+ if (!opts.files || opts.files.length === 0) opts.files = form.fileList;
1048
+ opts.files = opts.files && opts.files.length ? opts.files : null;
1049
+ opts.filesGiven = opts.files !== null;
1050
+ opts.acceptedBy = form.acceptor;
1004
1051
  }
1005
- opts.witness = form.witness;
1006
- opts.limitations = form.limitations;
1007
- opts.environment = form.environment || null;
1008
- opts.scope = [form.epic];
1009
- if (!opts.files || opts.files.length === 0) opts.files = form.fileList;
1010
- opts.files = opts.files && opts.files.length ? opts.files : null;
1011
- opts.filesGiven = opts.files !== null;
1012
- opts.acceptedBy = form.acceptor;
1013
1052
  }
1014
1053
 
1015
1054
  // No check for empty witness/limitations is needed here: the form is the only route
@@ -1679,6 +1718,179 @@ async function cmdHarness(opts) {
1679
1718
  return;
1680
1719
  }
1681
1720
 
1721
+ if (sub === 'verify-play') {
1722
+ refuseCommandOverride('harness verify-play', "the story's Play: declaration");
1723
+ // Mechanical play-declaration verification (contract v7 §1). A story states whether a
1724
+ // person can play its deliverable, or which work item will make it playable; this checks
1725
+ // that declaration against the work items that exist.
1726
+ //
1727
+ // WHY THIS COMMAND CANNOT SET THE GATE FOR A PLAYABLE STORY. The gate asks whether a
1728
+ // PERSON played the work. So this records the gate for a `deferred` declaration — the
1729
+ // declaration is the answer, exactly as a reachability deferral is — and REFUSES to
1730
+ // record it for a `required` one, pointing at the form. An agent that could answer
1731
+ // "the user played it" in a sentence would make the gate worthless, which is the whole
1732
+ // reason `userPlaythroughConfirmed` is human-owned in the gate registry.
1733
+ if (!opts.story) fail(opts, 'harness verify-play requires --story <path>');
1734
+ const storyPath = resolve(opts.targetDir, opts.story);
1735
+ const storyRel = relative(opts.targetDir, storyPath).replace(/\\/g, '/') || basename(storyPath);
1736
+ const { exists, state } = readState(opts.targetDir);
1737
+ assertExpectedPhase(opts, state);
1738
+ const playEnabled = policy.userPlay?.enabled === true;
1739
+ const workItemId = state ? workItemIdOf(state) : 'unscoped';
1740
+ const phase = state?.session?.currentPhase || 'implementation';
1741
+
1742
+ let playDeclaration;
1743
+ try {
1744
+ playDeclaration = parsePlayDeclaration(storyPath);
1745
+ } catch (err) {
1746
+ fail(opts, `cannot read story "${opts.story}": ${err.message}`, () => 1, { ok: false, code: 'story-unreadable', story: opts.story });
1747
+ }
1748
+
1749
+ const playWorkItems = exists ? collectWorkItems(state) : null;
1750
+ const playValidation = validatePlayDeclaration(playDeclaration, { workItems: playWorkItems, self: basename(storyPath) });
1751
+ const playSiblings = readSiblingPlayDeclarations(storyPath, { workItems: playWorkItems });
1752
+ const playGraph = playSiblings.length > 0
1753
+ ? playSiblings
1754
+ : [{ id: basename(storyPath), aliases: [], declaration: playDeclaration }];
1755
+ const playCycles = exists ? findDeferralCycles(playGraph) : [];
1756
+ const playGaps = describePlayGaps({ validation: playValidation, cycles: playCycles, story: opts.story });
1757
+
1758
+ // A playable story. The gate is NOT set here, and the refusal names the route that can.
1759
+ if (playValidation.ok && playValidation.code === 'required') {
1760
+ const formHint = `run "cadet-agent harness play-form --story ${opts.story}" to write one pre-filled from state, play the game, fill in the three blank fields, then pass it back with --artifact <the form>.`;
1761
+ fail(
1762
+ opts,
1763
+ `${USER_PLAY_GATE} is a person's record, and this story declares its deliverable playable: ${playDeclaration.instruction}. ${formHint}`,
1764
+ () => 1,
1765
+ { ok: false, gate: USER_PLAY_GATE, code: 'play-form-required', story: opts.story, declaration: playDeclaration },
1766
+ );
1767
+ }
1768
+
1769
+ const playOk = playValidation.ok && playCycles.length === 0;
1770
+
1771
+ if (!playEnabled) {
1772
+ if (opts.format === 'json') {
1773
+ emit(opts, '', { ok: playOk, story: opts.story, declaration: playDeclaration, play: playValidation, cycles: playCycles, gateSet: false, enabled: false });
1774
+ } else if (playOk) {
1775
+ console.log(`✅ Play declaration for ${opts.story}: ${playValidation.message}`);
1776
+ console.log(' userPlay.enabled is false — reported only, state.json unchanged.');
1777
+ } else {
1778
+ console.error(`⚠️ Play gaps in ${opts.story} (userPlay.enabled is false — reported only):`);
1779
+ for (const line of playGaps) console.error(line);
1780
+ }
1781
+ if (!playOk) process.exit(1);
1782
+ return;
1783
+ }
1784
+
1785
+ if (!playOk) {
1786
+ const detail = {
1787
+ ok: false,
1788
+ story: opts.story,
1789
+ declaration: playDeclaration,
1790
+ play: playValidation,
1791
+ cycles: playCycles,
1792
+ gateSet: false,
1793
+ code: playValidation.ok !== true ? playValidation.code : 'deferral-cycle',
1794
+ };
1795
+ if (opts.format === 'json') emit(opts, '', detail);
1796
+ else {
1797
+ console.error(`❌ Cannot set ${USER_PLAY_GATE} for ${opts.story}:`);
1798
+ for (const line of playGaps) console.error(line);
1799
+ }
1800
+ process.exit(1);
1801
+ }
1802
+
1803
+ const playAt = new Date();
1804
+ const playEvidence = createEvidence({
1805
+ evidenceId: newId(),
1806
+ workItemId,
1807
+ acceptanceCriterionId: null,
1808
+ phase,
1809
+ gate: USER_PLAY_GATE,
1810
+ status: 'passed',
1811
+ command: `harness verify-play --story ${opts.story}`,
1812
+ result: `play deferred (${playValidation.code})`,
1813
+ exitCode: 0,
1814
+ commit: opts.commit || null,
1815
+ inputTreeHash: computeInputTreeHash(opts.targetDir, [storyRel]),
1816
+ criteriaHash: hashCriteria([
1817
+ workItemId,
1818
+ playValidation.code,
1819
+ playDeclaration.deferTo || playDeclaration.instruction || '',
1820
+ ]),
1821
+ relevantFiles: [storyRel],
1822
+ createdAt: playAt,
1823
+ expiresAt: null,
1824
+ freshnessPolicy: { scope: 'story' },
1825
+ source: 'automated',
1826
+ });
1827
+
1828
+ const playLedger = new RunLedger({ targetDir: opts.targetDir, policy, runId: state?.activeRunId || null, workItemId, phase });
1829
+ playLedger.addEvidence(playEvidence);
1830
+ playLedger.addDecision({ kind: 'stop', reason: `play deferred (${playValidation.code})`, scope: 'declaration only' });
1831
+ playLedger.finalize({ status: 'ok' });
1832
+ const playLedgerPath = playLedger.persist();
1833
+
1834
+ if (exists) {
1835
+ const next = recordEvidence(state, playEvidence);
1836
+ writeState(opts.targetDir, next);
1837
+ }
1838
+
1839
+ if (opts.format === 'json') {
1840
+ emit(opts, '', { ok: true, story: opts.story, play: playValidation, cycles: playCycles, evidenceId: playEvidence.evidenceId, gateSet: exists, runId: playLedger.runId, path: playLedgerPath });
1841
+ } else {
1842
+ console.log(`✅ ${USER_PLAY_GATE} for ${opts.story}: ${playValidation.message}`);
1843
+ console.log(' The story defers the playthrough, so the declaration is the record — no form is needed.');
1844
+ console.log(` Ledger: ${playLedgerPath}`);
1845
+ }
1846
+ return;
1847
+ }
1848
+
1849
+ if (sub === 'play-form') {
1850
+ if (!opts.story) {
1851
+ fail(opts, 'harness play-form needs --story <path>: the form belongs to the story being played.', () => 1, { ok: false, code: 'story-required' });
1852
+ }
1853
+ const { exists, state } = readState(opts.targetDir);
1854
+ if (!exists) {
1855
+ fail(opts, 'No .cadet/state.json found. The form is generated from state, so there is nothing to fill it from yet.', () => 1, { ok: false, code: 'no-state' });
1856
+ }
1857
+ const storyPath = resolve(opts.targetDir, opts.story);
1858
+ const storyRel = relative(opts.targetDir, storyPath).replace(/\\/g, '/') || basename(storyPath);
1859
+ let playDeclaration = null;
1860
+ try {
1861
+ playDeclaration = parsePlayDeclaration(storyPath);
1862
+ } catch (err) {
1863
+ fail(opts, `cannot read story "${opts.story}": ${err.message}`, () => 1, { ok: false, code: 'story-unreadable', story: opts.story });
1864
+ }
1865
+ let playTemplate;
1866
+ try {
1867
+ playTemplate = readPlayTemplate(opts.targetDir);
1868
+ } catch (err) {
1869
+ fail(opts, `${err.message}. Run "cadet-agent sync" to restore it — the form is generated from that file so the template and the form cannot drift apart.`, () => 1, { ok: false, code: 'template-missing' });
1870
+ }
1871
+ const playText = buildPlayForm({
1872
+ template: playTemplate,
1873
+ state,
1874
+ storyPath,
1875
+ storyRel,
1876
+ targetDir: opts.targetDir,
1877
+ declaration: playDeclaration,
1878
+ });
1879
+ const playResult = writePlayForm(opts.targetDir, storyRel, playText, { out: opts.out });
1880
+ if (!playResult.written) {
1881
+ fail(opts, `harness play-form refuses to overwrite an existing form: ${playResult.reason} (${playResult.path}).`, () => 1, { ok: false, code: 'form-exists', path: playResult.path });
1882
+ }
1883
+ const playShown = playResult.path.slice(opts.targetDir.length + 1).replace(/\\/g, '/');
1884
+ if (opts.format === 'json') {
1885
+ emit(opts, '', { ok: true, path: playShown, story: opts.story, next: `cadet-agent harness confirm --gate ${PLAY_FORM_GATE} --artifact ${playShown}` });
1886
+ } else {
1887
+ console.log(`Playthrough form written: ${playShown}`);
1888
+ console.log(' Play the work, fill in the three unfilled fields, then run:');
1889
+ console.log(` cadet-agent harness confirm --gate ${PLAY_FORM_GATE} --artifact ${playShown}`);
1890
+ }
1891
+ return;
1892
+ }
1893
+
1682
1894
  if (sub === 'acceptance-form') {
1683
1895
  if (!opts.epicId) {
1684
1896
  fail(opts, 'harness acceptance-form needs --epic <epicId>: the form belongs to the epic being accepted.', () => 1, { ok: false, code: 'epic-required' });
@@ -2363,6 +2575,41 @@ export async function run(argv) {
2363
2575
  const opts = parseArgs(argv);
2364
2576
  const commandKey = resolveCommand(argv);
2365
2577
 
2578
+ // A command may not bind its own outputs as the evidence it records.
2579
+ //
2580
+ // A record's `inputTreeHash` covers the files `--files` names, and `harness verify`,
2581
+ // `harness confirm`, `harness verify-acs`, `harness verify-reachability`,
2582
+ // `harness verify-design-review` and `harness verify-architecture` then write the ledger
2583
+ // and `.cadet/state.json`. Binding one of those makes the record stale at the instant it
2584
+ // is created — the write it describes changes a file the hash covers — so the next
2585
+ // `state transition --dry-run` refuses the boundary for a record the harness itself just
2586
+ // wrote. Refused here, in the dispatcher, rather than at each of the four sites that
2587
+ // resolve `--files`: the registry already declares what each command writes, so the check
2588
+ // covers a new command and a new flag the moment they are registered, which is the same
2589
+ // reason `--dry-run` is driven from here (contract C13).
2590
+ if (commandKey && opts.files && opts.files.length) {
2591
+ const selfBound = selfBoundFiles(commandKey, opts.files);
2592
+ if (selfBound.length > 0) {
2593
+ const named = selfBound.join(', ');
2594
+ fail(
2595
+ opts,
2596
+ `--files names ${selfBound.length === 1 ? 'a path' : 'paths'} that "${commandKey}" writes itself: ${named}. `
2597
+ + 'Bind the files this gate JUDGES — sources, tests, and the story or epic markdown — never the ledger '
2598
+ + `or state document the evidence is recorded in: "${commandKey}" writes ${named} as it records the `
2599
+ + 'record, so a binding to it reads as stale the moment it is written ("input tree hash changed since '
2600
+ + 'the evidence was recorded"), and the gate can never be fresh.',
2601
+ () => 1,
2602
+ {
2603
+ ok: false,
2604
+ command: commandKey,
2605
+ code: 'self-bound-files',
2606
+ files: selfBound,
2607
+ writes: COMMANDS[commandKey].writes || [],
2608
+ },
2609
+ );
2610
+ }
2611
+ }
2612
+
2366
2613
  // Global `--dry-run`, driven by the registry rather than by each handler.
2367
2614
  //
2368
2615
  // This is the structural fix for the class of bug where a mutating command
@@ -165,6 +165,24 @@ export const COMMANDS = {
165
165
  writes: ['.cadet/runs/**', '.cadet/state.json'],
166
166
  unattended: true,
167
167
  },
168
+ 'harness verify-play': {
169
+ mutates: true,
170
+ summary: 'Verify a story\'s Play: declaration; record the gate for a deferral, refuse it for a playable story.',
171
+ // Same posture as verify-reachability: it records evidence for its gate, so it writes
172
+ // the ledger and state. It writes no artifact of its own — the declaration lives in the
173
+ // story, and a playable story's record comes from the person's form instead.
174
+ writes: ['.cadet/runs/**', '.cadet/state.json'],
175
+ unattended: true,
176
+ },
177
+ 'harness play-form': {
178
+ mutates: true,
179
+ summary: 'Write a user-playthrough form for a story, pre-filled from state.',
180
+ // One file, and only when it does not exist: a form is a person's worksheet once they
181
+ // have touched it, and regenerating it would discard what they wrote. It never writes
182
+ // state — recording the gate is `harness confirm --artifact`.
183
+ writes: ['.cadet/agent/project-plans/**'],
184
+ unattended: true,
185
+ },
168
186
  'harness acceptance-form': {
169
187
  mutates: true,
170
188
  summary: 'Write a human-acceptance form for an epic, pre-filled from state.',
@@ -229,6 +247,63 @@ export function readOnlyCommands() {
229
247
  return Object.entries(COMMANDS).filter(([, c]) => !c.mutates).map(([k]) => k);
230
248
  }
231
249
 
250
+ /** Repository-relative form of a path: forward slashes, no leading `./`. */
251
+ function normaliseRepoPath(path) {
252
+ return String(path).replace(/\\/g, '/').replace(/^\.\//, '');
253
+ }
254
+
255
+ /**
256
+ * Does one of a command's `writes` patterns cover a path?
257
+ *
258
+ * The shipped table uses exactly three shapes, and the matcher supports those and no more:
259
+ * an exact path (`AGENTS.md`), a directory subtree (`<dir>/**`), and a root-level suffix
260
+ * wildcard (`*.coverage.json`). `**` crosses directory boundaries and `*` does not — the
261
+ * distinction a shell makes, and the reason a `*` that crossed directories would refuse a
262
+ * file the command never touches.
263
+ *
264
+ * Wildcards are replaced FIRST, with placeholders that survive the escape pass. Doing it the
265
+ * other way round turns `**` into a literal `\*\*` and every declaration stops matching — the
266
+ * defect this function shipped with for one run, caught by the generated test below it.
267
+ *
268
+ * Absolute paths are deliberately not translated: `--files` carries repository-relative paths
269
+ * by contract, so an absolute path is a different mistake and not this function's to guess at.
270
+ */
271
+ function writePatternMatches(pattern, path) {
272
+ const escaped = normaliseRepoPath(pattern)
273
+ .replace(/\*\*/g, '\u0000')
274
+ .replace(/\*/g, '\u0001')
275
+ .replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
276
+ .replace(/\u0000/g, '.*')
277
+ .replace(/\u0001/g, '[^/]*');
278
+
279
+ return new RegExp(`^${escaped}$`).test(normaliseRepoPath(path));
280
+ }
281
+
282
+ /**
283
+ * The `--files` entries a command declares it WRITES, and which must therefore never be
284
+ * bound as its evidence.
285
+ *
286
+ * Why this exists: a record's `inputTreeHash` covers the files it binds, and the command
287
+ * then writes its own outputs — the run ledger and `.cadet/state.json`. Binding one of those
288
+ * makes the record stale at the instant it is created, because the write it records changes
289
+ * a file the hash covers. Measured on a consumer on 2026-10-01: a `storyTrackingUpdated`
290
+ * record that bound the story, the epic and `.cadet/state.json` was refused by the very next
291
+ * `state transition --dry-run` with "input tree hash changed since the evidence was
292
+ * recorded", and the gate had to be re-recorded twice before the boundary was allowed.
293
+ *
294
+ * The answer is a refusal rather than a filter. Filtering would leave the record claiming a
295
+ * coverage it does not have, and the caller asked for a binding that provably cannot hold —
296
+ * the same choice `AGENT_OWNED_GATES` and the swallowed `--command` flag both resolved the
297
+ * loud way. It is derived from the registry rather than hardcoded, so a new command's
298
+ * outputs are covered the moment it is registered.
299
+ */
300
+ export function selfBoundFiles(commandKey, files = []) {
301
+ const writes = COMMANDS[commandKey]?.writes || [];
302
+ if (writes.length === 0 || !Array.isArray(files) || files.length === 0) return [];
303
+
304
+ return files.filter((file) => writes.some((pattern) => writePatternMatches(pattern, file)));
305
+ }
306
+
232
307
  /**
233
308
  * Resolve the command key for a parsed invocation.
234
309
  *