cadet-agent 0.54.0 → 0.60.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -7,9 +7,9 @@ Cadet-Agent is **not a one-shot code generator**. It won't spit out a finished g
7
7
  ## Repository Layout
8
8
  - `.cadet/agent/core/` contains the shared Cadet-Agent framework documents.
9
9
  - `cadet-agent.md` is the thin global directive: identity, non-negotiable rules, workflow routing, hard-gate protocol, and skill dispatch.
10
- - `Harness.md` is the canonical harness contract: budgets, evidence-backed gates, retries, context tiers, tool routing, privacy, and escalation.
10
+ - `HarnessRuntime.md` is the lean runtime contract; `Harness.md` keeps the full harness rationale and reference.
11
11
  - `harness.schema.json` and `state.schema.json` are the machine-readable schemas for harness records and session state.
12
- - `skills/` contains scoped workflow-phase skills (PlanningReview, Requirements, Architecture, Spike, StoryBreakdown, TDD, Debugging, CodeReview, Resume, MCPSetup, AgentReviewer, Handoff, Reconciliation).
12
+ - `skills/` contains scoped workflow-phase skills (PlanningReview, Requirements, Architecture, DesignReview, Spike, StoryBreakdown, TDD, Debugging, CodeReview, VisualEvidence, Resume, MCPSetup, AgentReviewer, Handoff, Reconciliation).
13
13
  - `templates/` contains runtime templates for planning artifacts.
14
14
  - `.cadet/harness.json` holds repository-local budget/policy overrides (preserved by sync).
15
15
  - `.cadet/runs/` holds sanitized run ledgers (preserved by sync; no secrets or raw prompts by default).
@@ -28,7 +28,7 @@ Cadet-Agent is **not a one-shot code generator**. It won't spit out a finished g
28
28
 
29
29
  ## Cross-IDE Support
30
30
 
31
- Cadet-Agent provides full workflow parity across six IDEs. The same 11 skills + reviewer are available in each:
31
+ Cadet-Agent provides the **same skills** across six IDEs — one canonical file per skill, read through thin per-host pointers. It does **not** provide equal *enforcement*, and it does not claim to: no host here can block what it has no API to intercept, so enforcement is measured per action and published. Run `cadet-agent harness capabilities --verify-host` for the measured matrix on your repository, or read [Host Interception](docs/core/HostInterception.md).
32
32
 
33
33
  | Feature | GitHub Copilot | Cursor | Continue | Claude Code | Deep Code | Hermes |
34
34
  |---|---|---|---|---|---|---|
@@ -37,16 +37,22 @@ Cadet-Agent provides full workflow parity across six IDEs. The same 11 skills +
37
37
  | Planning Review | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
38
38
  | Requirements | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
39
39
  | Architecture | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
40
+ | Design Review | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
40
41
  | Spike | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
41
42
  | Story Breakdown | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
42
43
  | TDD | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
43
44
  | Debugging | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
44
45
  | Code Review | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
46
+ | Visual Evidence | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
45
47
  | Resume | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
46
48
  | MCP Setup | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
47
49
  | Reconciliation | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
48
50
  | Reviewer mode | Agent picker | Rule toggle | `/cadet-agent-reviewer` | `/cadet-agent-reviewer` | `cadet-agent-reviewer` skill | `/cadet-agent-reviewer` |
49
- | Git guard | PreToolUse hook | Manual | Manual | Manual | `permissions.ask` (`mutate-git-log`) | Command approval policies |
51
+ | Git guard (declared; measured by `--verify-host`) | `native` — PreToolUse hook | `external` via the repository Git hook, else `advisory` | `external` / `advisory` | `advisory` until a hook is configured | `external` (declared: `permissions.ask`) | `external` (declared: approval policies) |
52
+
53
+ Every other row is a capability: the skills are the same, the interception is not, and the difference is
54
+ measured rather than assumed (see [Host Interception](docs/core/HostInterception.md)). The portable control
55
+ that works for every host is the repository Git hook: `git config core.hooksPath .githooks`.
50
56
 
51
57
  All adapters delegate to the canonical files under `.cadet/agent/core/` — no duplicated rules or skills. See `ADAPTERS.md` for the full inventory, `docs/guidance/DeepCode.md` for Deep Code setup, and `docs/guidance/Hermes.md` for Hermes setup.
52
58
 
@@ -114,12 +120,12 @@ flowchart TD
114
120
  REPORT["Report current phase,<br/>epics, stories & gates"]
115
121
  CR["🔍 Context Resolution<br/>classify change size,<br/>calibrate learner,<br/>detect policy"]
116
122
  REQ["📋 Requirements<br/>Given/When/Then criteria<br/>assumption audit"]
117
- PLANREV["🧩 Planning Review<br/>interview one question<br/>at a time · decision tree"]
123
+ PLANREV["🧩 Planning Review (skill)<br/>interview one question<br/>at a time · decision tree"]
118
124
  ARCH["🏗️ Architecture<br/>technical design,<br/>ADR decisions"]
119
125
  SPIKE["🧪 Spikes<br/>resolve unverified<br/>assumptions"]
120
126
  BREAKDOWN["📐 Story Breakdown<br/>epics → testable stories"]
121
127
  IMPL["🔨 Implementation<br/>TDD per story,<br/>red → green → refactor"]
122
- REVIEW["✅ Review<br/>hard gate: 17-step<br/>code review, security"]
128
+ REVIEW["✅ Review<br/>hard gate: 23-step<br/>code review, security"]
123
129
  VALIDATE["✔️ Validation<br/>acceptance criteria,<br/>design artifact sync"]
124
130
  CLOSED(["🎉 Closed"])
125
131
  NEXT_STORY{"More stories<br/>in epic?"}
@@ -136,13 +142,13 @@ flowchart TD
136
142
 
137
143
  REQ --> ARCH
138
144
  ARCH -->|"unverified assumptions"| SPIKE
139
- ARCH -->|"all assumptions resolved"| BREAKDOWN
145
+ ARCH -->|"all assumptions resolved<br/>gate: designReviewCompleted ✅ (opt-in)"| BREAKDOWN
140
146
  SPIKE -->|"spike complete"| ARCH
141
147
 
142
148
  BREAKDOWN --> IMPL
143
149
 
144
150
  IMPL -->|"story complete"| REVIEW
145
- REVIEW -->|"gate: codeReviewCompleted ✅<br/>gate: securityReviewPassed ✅<br/>gate: reachabilityAddressed ✅ (opt-in)"| VALIDATE
151
+ REVIEW -->|"gate: codeReviewCompleted ✅<br/>gate: securityReviewPassed ✅<br/>gate: acceptanceCriteriaValidated ✅<br/>gate: reachabilityAddressed ✅ (opt-in)<br/>gate: userPlaythroughConfirmed ✅ (opt-in)"| VALIDATE
146
152
  VALIDATE -->|"gate: designArtifactSyncConfirmed ✅"| NEXT_STORY
147
153
  NEXT_STORY -->|"yes"| IMPL
148
154
  NEXT_STORY -->|"no"| CLOSED
@@ -161,48 +167,85 @@ Use the `/cadet-resume` slash command to pick up where you left off. It reads `.
161
167
 
162
168
  ### Phase Gating
163
169
 
164
- Hard gates are enforced at every phase transition. The agent reads `.cadet/state.json → gates` before advancing and **blocks** the transition if any required gate is `false`. Gates cannot be skipped without an explicit user-directed exception recorded in `changeHistory`.
170
+ Hard gates are enforced at every phase transition. The agent reads `.cadet/state.json → gates` before advancing and **blocks** the transition if any required gate is `false`. Gates cannot be skipped without an explicit, user-directed exception recorded in `.cadet/state.json → gateExceptions` — bounded by `expiresAt`, and naming the work items it covers. A `v1`–`v3` exception is read out of `changeHistory`, its older home.
165
171
 
166
172
  | Transition | Required Gates |
167
173
  |---|---|
168
- | implementation → review | `testsPassed`, `compileCheckConfirmed`, `unityAnalyzerClean`, `storyTrackingUpdated` |
169
- | review → validation | `codeReviewCompleted`, `securityReviewPassed`, `acceptanceCriteriaValidated`, and `reachabilityAddressed` when `reachability.enabled` is set |
170
- | validation → closed | `designArtifactSyncConfirmed` |
174
+ | architectureComplete → story-breakdown | `designReviewCompleted` when `designReview.enabled` is set — the formal design review, recorded by `harness verify-design-review` |
175
+ | implementation → review | `testsPassed`, `compileCheckConfirmed`, `unityAnalyzerClean`, `storyTrackingUpdated`, and `architectureFitnessPassed` when the project declares architecture checks and enables them |
176
+ | review → validation | `codeReviewCompleted`, `securityReviewPassed`, `acceptanceCriteriaValidated`, `reachabilityAddressed` when `reachability.enabled` is set, and `userPlaythroughConfirmed` when `userPlay.enabled` is set — a story declares `Play: required — <what the user does and what they see>` or `Play: deferred to <work item> — <why>`. A person's own record satisfies `required` (`harness play-form` writes the form, `harness confirm --artifact` records it); `harness verify-play` records `deferred`, and the deferral expires when the named work item closes |
177
+ | validation → closed | `designArtifactSyncConfirmed`, and `humanAcceptanceConfirmed` when `humanAcceptance.enabled` is set — a person's own record (`harness acceptance-form` writes the form, `harness confirm --artifact` records it); no command can produce it |
178
+
179
+ ### Runtime context protocol (opt-in, and the framework's own claim discipline)
180
+
181
+ `harness context plan` states what a phase requires (with a reason and a hash for each reference),
182
+ `harness context record` captures what the host loaded and the level it can claim, and
183
+ `harness context validate` decides whether a context-complete checkpoint may be claimed. The level is
184
+ reported as it is — a run reported as `recorded` is never reported as `enforced`, and `enforced` needs
185
+ a hook that declares it enforces context. A required reference that was never loaded, or that changed
186
+ after the record, blocks the checkpoint. See `.cadet/agent/core/Harness.md` §2d.
171
187
 
172
188
  **`closed` is end-of-epic, not per-story.** `validation → closed` is taken only when no stories remain (`NEXT_STORY → no → CLOSED` above). When an epic still has stories, the next story re-enters from `validation → implementation` (`NEXT_STORY → yes → IMPL`). Do not close a story individually: `closed` is terminal, and there is no transition out of it.
173
189
 
174
- The full set of legal transitions is the three gated rows above **plus** the ungated forward edges (classification, planning progression, `story-breakdown → implementation`, and the `validation → implementation` next-story loop). Any transition outside that set is rejected with a named reason.
190
+ The full set of legal transitions is the three gated rows above **plus** the ungated forward edges (classification, planning progression, `story-breakdown → implementation`, and the `validation → implementation` next-story loop). Any transition outside that set is rejected with a named reason. **Planning Review is a skill, not a phase**: the agent dispatches it before requirements or architecture when the plan is fuzzy or contested, and it records no transition.
191
+
192
+ ### The response contract
193
+
194
+ The framework's only per-reply output is one line: `cadet-agent: ok`, or the problem in its place. `cadet-agent harness status` derives that line — read-only — from the state document and the run ledger, so it cannot claim a health nothing verified. `ok` means the record is readable, valid and fresh, and no recorded run stopped on a budget or failed to run. A gate that is unmet because the story is unfinished is normal, and is not reported. The rest of a reply carries the work: the decisions taken, what changed, what the checks show, and what is unverified or deferred.
175
195
 
176
196
  ### Harness
177
197
 
178
198
  Gates are backed by **evidence**, not assertion. Each claimed gate must have a fresh, non-superseded evidence record bound to the current work item, input tree hash, and acceptance criteria. The harness also bounds context, tokens, tool calls, retries, wall-clock time, cost, and archive sizes — and those bounds are enforced, not advisory.
179
199
 
180
- - Rules: `.cadet/agent/core/Harness.md`. Data contract: `docs/core/HarnessContract.md`.
200
+ - Runtime rules: `.cadet/agent/core/HarnessRuntime.md`. Full contract: `.cadet/agent/core/Harness.md`. Data contract: `docs/core/HarnessContract.md`.
181
201
  - Overrides: `.cadet/harness.json` (preserved by sync; conservative defaults in `src/harness/policy.mjs`).
182
202
  - Ledgers: `.cadet/runs/<runId>.json` (sanitized; artifacts are redacted before they are written; no secrets or raw prompts by default).
183
203
  - Transitions recompute the input tree hash from the evidence's relevant files, so editing a relevant file invalidates the evidence.
184
204
  - `harness verify` binds evidence to `--files` (or the working tree's changed files), and a `testsPassed` green result requires a prior red record.
205
+ - A `--files` binding may not name a path the recording command itself writes: `--files .cadet/state.json` is refused before anything runs, because the write that follows would stale the record it just wrote.
185
206
  - When Git is unavailable and no `--files` are given, verification blocks (`freshness-unavailable`) rather than recording unscoped evidence.
186
207
  - `state validate` rejects a `true` gate whose evidence is missing, stale, expired, superseded, or bound to another work item; evidence records are schema-validated in full (`command`, `result`, `criteriaHash`, and a freshness bound).
208
+ - `state validate` errors when a work item that a `storyCompletions` row records as finished, with evidence behind it, still reads `planned`. The remedy is `done` or `superseded`: a finished story must not be indistinguishable from one that was never started.
209
+ - A `changeHistory` entry is a pointer, not a retelling: it is limited to 400 characters, and `state compact` archives a longer entry in place rather than truncating it.
187
210
  - Evidence must include a UUID, work item, phase, gate, status, command/result, input-tree hash, criteria hash, relevant files, timestamp, and either `expiresAt` or `freshnessPolicy`.
188
211
  - **Evidence history does not live in `state.json`.** A v4 document keeps only the active work item's records inline; a closed work item's evidence is written into the commit that closes it, as `Cadet-*` trailers, and archived to `.cadet/archive/`. `evidenceCoverage` indexes what left, so the "a done story owns evidence" check still works offline. Cadet still never commits: `state seal` prepares a message file and you commit with `git commit -F`.
212
+ - **Two gates are human-owned: `humanAcceptanceConfirmed` and `userPlaythroughConfirmed`.** No
213
+ command can produce them, and `--command` is refused for them. `harness acceptance-form --epic <id>`
214
+ and `harness play-form --story <path>` write the form; `harness confirm --gate <gate> --artifact <form>`
215
+ records what the person wrote. A form still holding a placeholder is refused, and a form someone has
216
+ started is never overwritten.
189
217
  - Command output counts against the output budget; a configured cost budget cannot be satisfied by unmeasurable cost (the run is blocked, `budget-blocked`).
190
218
  - State and run ledgers are written atomically, so an interrupted write cannot truncate a record; persisted artifacts are redacted before hashing or writing.
191
219
  - Empty freshness coverage is an explicit policy decision: set `allowEmptyFreshness: true` in `.cadet/harness.json` only when unscoped evidence is acceptable.
192
220
 
193
221
  ```bash
222
+ cadet-agent state init --workflow-path large # write the first state document (validated before it lands)
194
223
  cadet-agent state validate # validate state against the schema (read-only)
195
224
  cadet-agent state validate --verify-sealed # also read evidence out of commit trailers
196
225
  cadet-agent state migrate # atomically upgrade v1 → the current version
197
226
  cadet-agent state migrate --to 4 # archive closed work items' evidence; build the index
198
227
  cadet-agent state compact --keep active # routine housekeeping on a v4 state
199
228
  cadet-agent state seal # write the active work item's evidence as commit trailers
229
+ cadet-agent state begin --epic <id> --story <id> # start a work item; archive the previous item's evidence
200
230
  cadet-agent state transition --to review # enforce the matrix + evidence
231
+ cadet-agent harness status # the health line: ok, or the problem (read-only)
201
232
  cadet-agent harness verify --gate testsPassed --files src/a.cs # bounded, classified loop
233
+ cadet-agent harness verify-acs --story <path> # derive coverage from the run report, then record the gate
234
+ cadet-agent harness verify-reachability --story <path> # check the declaration; run the project probe
235
+ cadet-agent harness verify-design-review --artifact <path> --files <design,requirements,ADRs>
236
+ cadet-agent harness verify-architecture # run the project's declared fitness checks
237
+ cadet-agent harness verify-play --story <path> # check a story's `Play:` declaration; record a deferral
238
+ cadet-agent harness play-form --story <path> # write a user-playthrough form for a person to fill
239
+ cadet-agent harness acceptance-form --epic <id> # write the epic's human-acceptance form
240
+ cadet-agent harness confirm --gate <gate> --artifact <form> # record a person's own account
241
+ cadet-agent harness changes # the files a story changed, with links (read-only)
202
242
  cadet-agent harness report # budget consumption and failures (no secrets)
203
243
  cadet-agent harness reconcile # reconcile the planning chain against state.json (read-only)
244
+ cadet-agent harness matrix-check # reconcile a TDD matrix against the test inventory (read-only)
245
+ cadet-agent harness context plan|record|validate # plan what a phase loads, record what it loaded, decide
204
246
  cadet-agent harness cleanup --older-than-ms <n> # apply the retention policy (bound required)
205
247
  cadet-agent harness capabilities # available CLI/Unity/MCP/hook/token/cost telemetry
248
+ cadet-agent harness capabilities --verify-host # probe the configured interception, per action
206
249
  ```
207
250
 
208
251
  Every command supports `--format human|json` and exits nonzero for invalid state, failed verification, budget exhaustion, stale evidence, or safety rejection.
@@ -318,13 +361,14 @@ If a specific game repository needs local conventions, add a policy file under `
318
361
 
319
362
  ## Package Output
320
363
  Running `./package-agent.ps1` produces `cadet-agent.zip` with this layout:
321
- - `.cadet/agent/core/` (including `Harness.md`, `harness.schema.json`, and `state.schema.json`)
364
+ - `.cadet/agent/core/` (including `HarnessRuntime.md`, `Harness.md`, `harness.schema.json`, and `state.schema.json`)
322
365
  - `.cadet/agent/core/skills/`
323
366
  - `.cadet/agent/core/templates/`
324
367
  - `.github/agents/cadet.agent.md`
325
368
  - `.github/agents/cadet-agent-reviewer.agent.md`
326
369
  - `.github/prompts/cadet-*.prompt.md`
327
- - `.github/hooks/`
370
+ - `.github/hooks/` (the Copilot `git-guard` hook and its scripts)
371
+ - `.githooks/pre-commit` (the portable Git hook; installed by you with `git config core.hooksPath .githooks`)
328
372
  - `.cursor/rules/cadet-agent.md`
329
373
  - `.cursor/rules/cadet-agent-reviewer.md`
330
374
  - `.continue/rules/cadet-agent.md`
@@ -332,6 +376,10 @@ Running `./package-agent.ps1` produces `cadet-agent.zip` with this layout:
332
376
  - `.continue/config.yaml`
333
377
  - `.claude/skills/cadet-agent/SKILL.md`
334
378
  - `.claude/skills/cadet-*/SKILL.md`
379
+ - `.agents/skills/cadet-agent/SKILL.md` (the cross-client root Deep Code and Hermes read)
380
+ - `.agents/skills/cadet-*/SKILL.md`
381
+ - `AGENTS.md` (create-only: an existing consumer copy is never overwritten)
382
+ - `.cadet/harness.json` (create-only: the new-consumer policy seed)
335
383
 
336
384
  ## Notes
337
385
  - `.cadet/agent/core/FrameworkManifest.json` defines the managed and preserved paths for packaged installs.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cadet-agent",
3
- "version": "0.54.0",
3
+ "version": "0.60.0",
4
4
  "description": "Cross-IDE agent framework for Unity/C# game-development — one-command install",
5
5
  "type": "module",
6
6
  "bin": {