@sylad/cadence 0.27.1 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -8,7 +8,7 @@
8
8
  {
9
9
  "name": "cadence",
10
10
  "description": "Session start and close rituals driven by a versioned plan (raf), and deliveries proven by their effect. Needs the cadence CLI (npm i -g @sylad/cadence).",
11
- "version": "0.27.1",
11
+ "version": "1.0.0",
12
12
  "source": "./",
13
13
  "author": {
14
14
  "name": "Sylvain Ladoire"
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "cadence",
3
3
  "description": "A repo-native working method: session start and close rituals driven by a versioned plan (raf), deliveries proven by their effect, and three reviewer agents (UX, code, QA).",
4
- "version": "0.27.1",
4
+ "version": "1.0.0",
5
5
  "author": { "name": "Sylvain Ladoire" },
6
6
  "homepage": "https://github.com/Sylad/cadence",
7
7
  "repository": "https://github.com/Sylad/cadence",
package/README.md CHANGED
@@ -24,7 +24,7 @@ Code prompt that shows the session's context, quota, cost and the orchestrate wa
24
24
 
25
25
  ## What's new
26
26
 
27
- **0.27.1**: `cadence orchestrate` no longer leaves a session's child processes alive: a dev server started in its own process group, or orphaned by a shell that exits (`nohup srv &`), is killed with the session through its process tree and its `CADENCE_SESSION` mark, a tmux server, `screen`, `gpg-agent` or `dirmngr` started on demand being spared; a session that exits with code 0 while a descendant holds its output open is no longer reported as « délai dépassé » (L83); the `ux-reviewer` and `qa-reviewer` agents carry the briefs' Playwright rule outside `cadence orchestrate` (L84). **0.27.0**: a compliant review of `cadence orchestrate` settles the open sub-tasks that commits of the lot cite (`feat(L3/t2): …`, comma lists included), closed in the same plan commit as the verdict or, on a read-only plan, returned as `[sous-tâche clore]` proposals, so `raf done` no longer refuses a lot over sub-tasks whose work is done; on a foreign-format plan the ids are read as the plan reads them (L82); the test suite stays green with two full suites running in parallel (L150). **0.26.0**: `cadence orchestrate` raises the floor of a lot's budget from 200 k to 450 k tokens, enough to pay a write pass, its review, one fix pass and the short review that follows, so a lot of 0.5 or 1 day is no longer handed back at its second or third pass; a `--continue` wave draws fewer lots as a result (L128). **0.25.0**: the project line of the plans' progress in the `cadence-hud` band names the lots in progress after their count, three at most then `…` (L153, `cadence-hud` 0.4.2); the progress follows the folder the session was launched from, so a `cd` into a sub-project no longer shrinks the band, and the session context keeps the `session.start` folder (L154). **0.24.0**: `cadence orchestrate --drop <project:lot>` and `--stop-after-current` steer a live wave from another terminal, `--resume --drop` removes a lot before replaying a stopped wave (L79); the final table of a wave says « livrable jusqu'à <sha> » for a ready lot stacked under a lot handed back, with `git push origin <sha>:main && cadence deliver --sha <sha>` (L76); a version commit (version fields, `CHANGELOG.md`, the README's « What's new ») no longer asks for a new code review (L72). **0.23.0**: the progress bar of each project in the `cadence-hud` band is coloured by the share of lots done, red below 33 %, orange below 66 %, green beyond (L152, `cadence-hud` 0.4.1); `cadence deliver` under a wave lock refuses only when a lot of the wave still works in that very repository, a linked worktree included (L127). **0.22.0**: `cadence orchestrate --continue` draws the next ready lot of the plan itself, in the declared priority, until the budget, the time window or a question stops it (L147); `cadence session context` prints the context of the session, which the `cadence-hud` band publishes, and the `lead` skill chains lots without the human while it stays under 60 % (L148); the band shows the progress of each project's plan, and `cadence lead tour --json` gains `progress` (L149, `cadence-hud` 0.4.0); the final review of a lot is always played, the lot's own budget bounding the writing passes only, with 65 k tokens reserved for it (L145); the cause « tests rouges après … » names the pass whose tests were red (L151). **0.21.0**: `cadence orchestrate --status` and `cadence lead tour` print the free session slots, which the `lead` skill reads (L141); `orchestrate.effort` sets the effort level of each pass of a wave (L137); `docs.sync` makes `raf check`, `lead tour` and the review brief report a document that does not follow the code (L143); `docs.articles` opens, after a green `cadence deliver`, a lot `Article <project> à rafraîchir` in the neighbouring repository (L144). **0.20.0**: `cadence-hud` no longer misleads once the work is over: `/hud cmd` lists the background commands it counts (id, tool, origin, age, name), any other argument answers `argument inconnu` instead of hiding the band, commands with no end seen for an hour or no record go to a grey `N cmd sans fin vue`, a finished subagent takes its commands with it, two races that left a `1 cmd` counted forever are closed, and a finished wave shows grey without its budget then disappears after 30 minutes (L142); the `qa-reviewer` and `ux-reviewer` agents are read-only for real, their frontmatter listing Read, Grep, Glob, Bash and the Playwright tools (L140). **0.19.1**: `cadence-hud`'s `N cmd` counter goes back down when a Monitor expires, when the background commands a subagent started end with it, and at session start or plugin reload (L135); its `tool.call` and `prompt.submit` guard hooks carry a `.catch` on registration, which clears the `claude plugin validate` warning (L136). **0.19.0**: `cadence orchestrate` lifts a repository's lock and `pre-push` guard as soon as every lot of the wave that touches it is finished, so `cadence deliver` works there while the wave continues elsewhere (L132); a review that leaves a repository dirty hands back that lot only, suspends the lots still queued in that repository and keeps the other repositories running, its verdict reported instead of lost (L133); `cadence session close` re-run without a `session start` keeps the window since the last opening (L131); `cadence-hud` draws the lot rows of a wave in aligned columns on every surface (L134). **0.18.0**: `cadence session next` prints what it records (issue #6); `cadence session close` covers the period since the last `session start` (or the last close) and says so in its title, `--since` still wins (issue #7). **0.17.0**: `cadence-hud` shows an `N cmd` segment next to the agents, the background commands of the session (Bash run in the background and Monitor, subagents' included), tracked from the tool results and ended by the task's notification or `TaskStop`. **0.16.0**: a lot can declare **neighbouring repositories** (`repos:` key of the lot, `{ path, cite }` for a repository shared between projects): `cadence orchestrate` refuses, locks and guards them, gives them to every session as `--add-dir`, and `raf commits`, `raf show` and the review verdict read the commits that cite the lot there too. **0.15.0**: `hook.autostart` in `cadence.yaml` (`warn` by default, `refuse` or `start`): a pre-commit hook that refuses, or starts in the same commit, a lot still todo that the commit cites; `cadence session start|close --all [--depth n]`, one section per project under a folder. **0.14.0**: `cadence lead tour`, the lead's morning table of every project in a folder, printed by a program (one line per project: lots in progress, gaps, last notes, next ready lot, repository) in place of one subagent per project; the `lead` skill calls it. **0.13.1**: `cadence-hud` takes back the segments that fit once a big one has dropped (reset times, then the per-model consumption, now short: `opus 1.2M sonnet 800k`); `CADENCE_HUD_AMBIGUOUS=1` for a terminal that draws `▰ ▱ ⚙ │ ↻` in one cell. **0.13.0**: `cadence orchestrate --status --watch` (a wave's table redrawn in the terminal, which
27
+ **1.0.0**: the first major release, a stable contract: the commands and options, the keys of `cadence.yaml`, the plan file and the names of the skills and agents follow semantic versioning, and the README's *Stability* section lists what is guaranteed and what is not (L164). **0.28.0**: `orchestrate.precheck: local` runs the « deliverable already present? » pre-check of `cadence orchestrate` on the local model (`claude-local`, Ollama), off the Anthropic quota, Sonnet staying the default and the fallback (L146); the `qa-reviewer` agent walks only the pages a lot touched, with a **Scope** line in its report (L123); the sub-tasks a review closes are also read in a lot's neighbouring repositories (L156) and a commit's sub-task list is read with the plan's left guard (L155); an unstable test is fixed (L90). **0.27.1**: `cadence orchestrate` no longer leaves a session's child processes alive: a dev server started in its own process group, or orphaned by a shell that exits (`nohup srv &`), is killed with the session through its process tree and its `CADENCE_SESSION` mark, a tmux server, `screen`, `gpg-agent` or `dirmngr` started on demand being spared; a session that exits with code 0 while a descendant holds its output open is no longer reported as « délai dépassé » (L83); the `ux-reviewer` and `qa-reviewer` agents carry the briefs' Playwright rule outside `cadence orchestrate` (L84). **0.27.0**: a compliant review of `cadence orchestrate` settles the open sub-tasks that commits of the lot cite (`feat(L3/t2): …`, comma lists included), closed in the same plan commit as the verdict or, on a read-only plan, returned as `[sous-tâche clore]` proposals, so `raf done` no longer refuses a lot over sub-tasks whose work is done; on a foreign-format plan the ids are read as the plan reads them (L82); the test suite stays green with two full suites running in parallel (L150). **0.26.0**: `cadence orchestrate` raises the floor of a lot's budget from 200 k to 450 k tokens, enough to pay a write pass, its review, one fix pass and the short review that follows, so a lot of 0.5 or 1 day is no longer handed back at its second or third pass; a `--continue` wave draws fewer lots as a result (L128). **0.25.0**: the project line of the plans' progress in the `cadence-hud` band names the lots in progress after their count, three at most then `…` (L153, `cadence-hud` 0.4.2); the progress follows the folder the session was launched from, so a `cd` into a sub-project no longer shrinks the band, and the session context keeps the `session.start` folder (L154). **0.24.0**: `cadence orchestrate --drop <project:lot>` and `--stop-after-current` steer a live wave from another terminal, `--resume --drop` removes a lot before replaying a stopped wave (L79); the final table of a wave says « livrable jusqu'à <sha> » for a ready lot stacked under a lot handed back, with `git push origin <sha>:main && cadence deliver --sha <sha>` (L76); a version commit (version fields, `CHANGELOG.md`, the README's « What's new ») no longer asks for a new code review (L72). **0.23.0**: the progress bar of each project in the `cadence-hud` band is coloured by the share of lots done, red below 33 %, orange below 66 %, green beyond (L152, `cadence-hud` 0.4.1); `cadence deliver` under a wave lock refuses only when a lot of the wave still works in that very repository, a linked worktree included (L127). **0.22.0**: `cadence orchestrate --continue` draws the next ready lot of the plan itself, in the declared priority, until the budget, the time window or a question stops it (L147); `cadence session context` prints the context of the session, which the `cadence-hud` band publishes, and the `lead` skill chains lots without the human while it stays under 60 % (L148); the band shows the progress of each project's plan, and `cadence lead tour --json` gains `progress` (L149, `cadence-hud` 0.4.0); the final review of a lot is always played, the lot's own budget bounding the writing passes only, with 65 k tokens reserved for it (L145); the cause « tests rouges après … » names the pass whose tests were red (L151). **0.21.0**: `cadence orchestrate --status` and `cadence lead tour` print the free session slots, which the `lead` skill reads (L141); `orchestrate.effort` sets the effort level of each pass of a wave (L137); `docs.sync` makes `raf check`, `lead tour` and the review brief report a document that does not follow the code (L143); `docs.articles` opens, after a green `cadence deliver`, a lot `Article <project> à rafraîchir` in the neighbouring repository (L144). **0.20.0**: `cadence-hud` no longer misleads once the work is over: `/hud cmd` lists the background commands it counts (id, tool, origin, age, name), any other argument answers `argument inconnu` instead of hiding the band, commands with no end seen for an hour or no record go to a grey `N cmd sans fin vue`, a finished subagent takes its commands with it, two races that left a `1 cmd` counted forever are closed, and a finished wave shows grey without its budget then disappears after 30 minutes (L142); the `qa-reviewer` and `ux-reviewer` agents are read-only for real, their frontmatter listing Read, Grep, Glob, Bash and the Playwright tools (L140). **0.19.1**: `cadence-hud`'s `N cmd` counter goes back down when a Monitor expires, when the background commands a subagent started end with it, and at session start or plugin reload (L135); its `tool.call` and `prompt.submit` guard hooks carry a `.catch` on registration, which clears the `claude plugin validate` warning (L136). **0.19.0**: `cadence orchestrate` lifts a repository's lock and `pre-push` guard as soon as every lot of the wave that touches it is finished, so `cadence deliver` works there while the wave continues elsewhere (L132); a review that leaves a repository dirty hands back that lot only, suspends the lots still queued in that repository and keeps the other repositories running, its verdict reported instead of lost (L133); `cadence session close` re-run without a `session start` keeps the window since the last opening (L131); `cadence-hud` draws the lot rows of a wave in aligned columns on every surface (L134). **0.18.0**: `cadence session next` prints what it records (issue #6); `cadence session close` covers the period since the last `session start` (or the last close) and says so in its title, `--since` still wins (issue #7). **0.17.0**: `cadence-hud` shows an `N cmd` segment next to the agents, the background commands of the session (Bash run in the background and Monitor, subagents' included), tracked from the tool results and ended by the task's notification or `TaskStop`. **0.16.0**: a lot can declare **neighbouring repositories** (`repos:` key of the lot, `{ path, cite }` for a repository shared between projects): `cadence orchestrate` refuses, locks and guards them, gives them to every session as `--add-dir`, and `raf commits`, `raf show` and the review verdict read the commits that cite the lot there too. **0.15.0**: `hook.autostart` in `cadence.yaml` (`warn` by default, `refuse` or `start`): a pre-commit hook that refuses, or starts in the same commit, a lot still todo that the commit cites; `cadence session start|close --all [--depth n]`, one section per project under a folder. **0.14.0**: `cadence lead tour`, the lead's morning table of every project in a folder, printed by a program (one line per project: lots in progress, gaps, last notes, next ready lot, repository) in place of one subagent per project; the `lead` skill calls it. **0.13.1**: `cadence-hud` takes back the segments that fit once a big one has dropped (reset times, then the per-model consumption, now short: `opus 1.2M sonnet 800k`); `CADENCE_HUD_AMBIGUOUS=1` for a terminal that draws `▰ ▱ ⚙ │ ↻` in one cell. **0.13.0**: `cadence orchestrate --status --watch` (a wave's table redrawn in the terminal, which
28
28
  stops with the wave) and the lead skill that names the wave and follows its journal; lot ids with `/`
29
29
  (sub-tasks of a read-only plan) accepted by `orchestrate`; `cadence-hud`'s first line measured in
30
30
  terminal cells and kept to ASCII; `publish.yml` on the v5 actions. **0.12.0**: the `cadence-hud` plugin (a status band above the Claude Code prompt),
@@ -406,7 +406,7 @@ No gate and no command here: the QA review comes **after** a delivery, and
406
406
  page shows or what it is served (screen, API, data source, configuration of
407
407
  either) — in practice every delivery except docs-, plan- or tests-only ones: a
408
408
  backend-only lot can empty a page without touching a screen, and the agent then
409
- starts with the pages that call the changed endpoints. The `qa-reviewer` agent
409
+ walks the pages that call the changed endpoints. It also walks the pages that consume the services the lot changed; when the diff changes no route and no screen, or cannot be linked to pages, it walks every page and says so. The `qa-reviewer` agent
410
410
  opens each page of the running app in a real browser and judges it from the
411
411
  user's side. A page can be empty while everything else is green — no code
412
412
  changed, a data source went down upstream, the unit tests replace the network,
@@ -925,7 +925,8 @@ older `claude` is refused at launch unless every pass is `default`). Defaults: `
925
925
  `implement` and `fix` medium, `review` high (`review-small` follows `review`), `ux` high. A project overrides any of them
926
926
  in `cadence.yaml`; the levels are `low`, `medium`, `high`, `xhigh`, `max`, and `default` sends no `--effort` (the
927
927
  level of the session, as before this key). The format-retry session (`--resume`) keeps the level of its pass, and
928
- `--dry-run` shows the flag in each command line.
928
+ `--dry-run` shows the flag in each command line, except for a `precheck: local` session: `claude-local` is launched
929
+ without `--effort` (see below), so `effort.precheck` has no effect on it and the `precheck (local)` line does not show it.
929
930
 
930
931
  ```yaml
931
932
  orchestrate:
@@ -944,6 +945,20 @@ with proofs), `partiel` or `non`: on `oui` no implementation session is opened a
944
945
  implementation brief and a warning; on `non` — or an unreadable report, which only adds a warning — the wave goes on. A
945
946
  lot that already has commits is never pre-checked (resuming it is legitimate). `orchestrate.precheck: false` turns it off.
946
947
 
948
+ `orchestrate.precheck: local` (L146) runs the pre-check on the local model of the dev machine instead of Sonnet: the
949
+ orchestrator launches `claude-local` with the same brief, the `precheck-reader` prompt and tools, no quota used and nothing counted in the
950
+ lot's budget (the step shows `local` as its model). `claude-local` is **not shipped with cadence**: it is a personal wrapper (a script on the `PATH`, `qwen3-coder:30b` by default) around
951
+ `claude` pointed at Ollama. `CADENCE_CLAUDE_LOCAL_BIN` names another binary, which must take the prompt as its **first argument**
952
+ and add itself `-p`, `--bare`, `--strict-mcp-config` and `--model` (cadence passes none of them), then accept `--output-format`,
953
+ `--json-schema`, `--tools`, `--session-id`, `--permission-mode`, `--add-dir` and `--disallowedTools`. A bare `claude` does not fit.
954
+ Without such a binary every local check fails and is replayed on Sonnet: keep `true`. Sonnet stays the safety net: a local session with no result (Ollama
955
+ down, timeout, no structured report) or an unreadable report is replayed on Sonnet with a warning, and a local `oui` — the
956
+ only answer that hands the lot back without an implementation — is confirmed by Sonnet before it counts. A local `partiel`
957
+ or `non` is taken as is. The default stays `true` (Sonnet): switch to `local` only after comparing verdicts and durations
958
+ against Sonnet on a few lots, and go back to `true` the first time the local model gets a `partiel` wrong. `--dry-run`
959
+ prints the step as `precheck (local)` with its `claude-local` arguments. Those arguments carry neither `--model` nor `--effort` nor `--agents`: `effort.precheck` applies only to the Sonnet
960
+ session (the fallback, or the confirmation of a local `oui`).
961
+
947
962
  The `precheck-reader` agent does not set `omitClaudeMd` (Claude Code 2.1.271), on purpose: measured on a project with a
948
963
  297-line `CLAUDE.md` (haiku, `--agents` + `--agent`, same prompt, with and without the flag), the session context is
949
964
  identical (22 779 tokens, the project `CLAUDE.md` still answered when asked) — the flag only applies to an agent run as a
@@ -970,7 +985,7 @@ committed, a short re-review of that commit runs instead): `ready`, verdict reco
970
985
  the untreated minors returned as proposals. Same when the budget is exhausted right after a compliant
971
986
  review with minors: it concludes on that review instead of staying suspended.
972
987
 
973
- **Sub-tasks covered by commits (L82)**: when the review concludes compliant, the open sub-tasks of the lot that a work commit of the lot cites (`feat(L3/t2): …`; a commit that cites only the lot covers none) are closed in the same plan commit as the verdict, with the warning `sous-tâches closes : L3/t2 (sha)`, so `raf done` does not refuse the lot over finished work. A read-only plan is never written: each such sub-task is returned as a proposal `[sous-tâche clore] L3/t2 — couverte par <sha>`, for you to close with the project's tool. On a plan in a foreign format (read-only, `cadence.yaml` `plan:`), sub-task ids are read as the plan reads them: `B53/t4-ux1` covers `t4-ux1` only, not `t4`, and a list such as `feat(B53/t4-ux1,t5)` covers each of its sub-tasks. A non-compliant review closes nothing.
988
+ **Sub-tasks covered by commits (L82)**: when the review concludes compliant, the open sub-tasks of the lot that a work commit of the lot cites (`feat(L3/t2): …`, in the project or in one of the lot's neighbouring repositories; a commit that cites only the lot covers none) are closed in the same plan commit as the verdict, with the warning `sous-tâches closes : L3/t2 (sha)`, so `raf done` does not refuse the lot over finished work. A read-only plan is never written: each such sub-task is returned as a proposal `[sous-tâche clore] L3/t2 — couverte par <sha>`, for you to close with the project's tool. On a plan in a foreign format (read-only, `cadence.yaml` `plan:`), sub-task ids are read as the plan reads them: `B53/t4-ux1` covers `t4-ux1` only, not `t4`, and a list such as `feat(B53/t4-ux1,t5)` covers each of its sub-tasks. The lot id is not read inside another id (`fix(E-A2/t3,t4, A2/t1)` covers `A2/t1` only for lot `A2`), and a list may repeat the lot id (`fix(L1/t1, L1/t2,t3)` covers `t1`, `t2` and `t3`). A non-compliant review closes nothing.
974
989
 
975
990
  **Choices, not questions**: the author brief tells the session to decide minor interpretation questions itself and to list them under `choix` in its report; the reviewer receives that list to re-read, and the final table prints each one (`choix fait : …`). A session stops with a question only on a real blocker: a decision that changes the scope or the architecture or is costly to undo, AND that the plan, its notes and CLAUDE.md do not settle; everything else is a choice.
976
991
 
@@ -1068,8 +1083,7 @@ a time (above); `Agent`, `git push`, `cadence deliver`, `raf done|review|ux` are
1068
1083
  temporary `pre-push` hook, installed for the duration of the wave and removed at its end (or, for a repository, as soon as all the lots that touch it are finished), refuses any push
1069
1084
  from a session (`CADENCE_ORCHESTRATED` is in their environment; a repository that already has another
1070
1085
  `pre-push` hook is refused before anything starts — `pushurl` is never touched); after every session the
1071
- upstream ref and `git ls-remote` are compared with the "before", and a review that changed `HEAD` or the
1072
- tree (screenshots left at the root, a commit) is an incident of that lot only: the lot is handed back to the lead, the lots of the other repositories carry on while the lots still queued behind it in the same repository are suspended (resumable, nothing started on a dirty tree), and a verdict the review had already produced is reported in the outcome and the lot's warnings (never recorded in the plan). A push or a removed guard hook still stops the wave. `raf done|ux|review` and `cadence deliver` refuse when
1086
+ upstream ref and `git ls-remote` are compared with the "before", and a read-only step that left only UNTRACKED files (screenshots written at the root by a script with a relative path) is not an incident: the files are moved to `.cadence/runs/<wave>/<project>--<lot>/stray/` (kept in their tree; for a neighbour repository, under `stray/<its relative path, `/` replaced by `_`>/`, e.g. `stray/.._gitops/`), a warning naming them goes to the lot's warnings and the wave log, the verdict is kept and the lot continues (L162; a move that is impossible, or an untracked file that vanished during the step, stays an incident); a review that changed `HEAD` or a tracked file (a commit, a modified file) is an incident of that lot only: the lot is handed back to the lead, the lots of the other repositories carry on while the lots still queued behind it in the same repository are suspended (resumable, nothing started on a dirty tree), and a verdict the review had already produced is reported in the outcome and the lot's warnings (never recorded in the plan). A push or a removed guard hook still stops the wave. `raf done|ux|review` and `cadence deliver` refuse when
1073
1087
  `CADENCE_ORCHESTRATED` is set.
1074
1088
 
1075
1089
  **Budget**: the wave counts input + cache writes + output tokens (default 2 M); cache reads are kept and
@@ -1137,7 +1151,7 @@ News instruction (`cadence news new <lot>`, factual user-side text, a screenshot
1137
1151
  orchestrate:
1138
1152
  test: npm test # run by the orchestrator after a work step (optional)
1139
1153
  build: npm run build # run after the tests; both results go to the reviewer, who does not redo them (optional)
1140
- precheck: true # default: before the first implementation of a lot with no commit, a read-only Sonnet session checks whether the deliverable is already in the repository (see below); false skips it
1154
+ precheck: true # default: before the first implementation of a lot with no commit, a read-only Sonnet session checks whether the deliverable is already in the repository (see below); false skips it, local runs it on claude-local (Ollama) with Sonnet as fallback
1141
1155
  effort: { review: xhigh } # effort level per pass (see « Effort level per pass »): precheck, implement, fix, review, ux
1142
1156
  ux: http://localhost:4200 # a URL, a launch command, or { command, url, timeout? } — for the UX review (see below)
1143
1157
  permissionMode: auto # default
@@ -1280,12 +1294,14 @@ repository with `cadence skills install` (to `.claude/skills/cadence-*` and
1280
1294
  runs a build whose output is used live, and never edits code. A README or usage
1281
1295
  documentation that does not follow the lot's change is a *major* finding.
1282
1296
  - **qa-reviewer** (agent): any web app (read-only like `ux-reviewer`: same tool list, no `Edit`, no `Write`); given a repository and a base URL (and
1283
- optionally a lot id, to start with the pages it touched — for a backend-only
1284
- lot, those that call the changed endpoints), it opens each page of
1297
+ optionally a lot id, to walk only the pages it touched — those that call the
1298
+ changed endpoints, those of the screens it changed, the pages that consume the services the lot changed,
1299
+ and the home page; a backend-only lot stays in scope, and when the lot changes no route
1300
+ and no screen, or the endpoints cannot be mapped to pages, every page is walked; the others are named « Not walked »), it opens each page of
1285
1301
  the project's expectations file in a real browser at 1440 and 390 px and
1286
1302
  measures: expected content present and non-empty, no error or missing-data
1287
1303
  message, every API call answered 2xx with a non-empty body, no console error,
1288
- no broken content image. Findings are defects (a line of the expectations
1304
+ no broken content image. It reads the page as text first and takes a capture only for a gap. Findings are defects (a line of the expectations
1289
1305
  broken, or a universal check failing with a visible effect, with or without an
1290
1306
  expectations file), suspects (it looks like missing or wrong data and no
1291
1307
  expectation settles it) or noise (a console error or a failed request with no
@@ -1358,13 +1374,15 @@ A version exists in three places and is published in two; a release does all of
1358
1374
  in that commit, not before the review. A commit that touches any other file
1359
1375
  — or another field of a manifest, or a dependency in `package-lock.json` — is work like any other and has
1360
1376
  to be reviewed.
1361
- 2. `git tag v<version> && git push origin main v<version>` — the tag starts `.github/workflows/publish.yml`,
1377
+ 2. Run the plugin evaluation, `npm run eval:plugin` (billed, see [Evaluating the plugin](#evaluating-the-plugin-before-a-release)),
1378
+ on that version commit; a case below the threshold stops the release until you have read it in the report.
1379
+ 3. `git tag v<version> && git push origin main v<version>` — the tag starts `.github/workflows/publish.yml`,
1362
1380
  which publishes to npm through Trusted Publishing (OIDC, no token stored anywhere): it checks the tag
1363
1381
  matches `package.json`, `.claude-plugin/plugin.json` and `.claude-plugin/marketplace.json` and that `CHANGELOG.md` has a `## [x.y.z]` section for it (no section, no publication), then `npm publish --provenance`, where `prepublishOnly` runs the type-check and the
1364
1382
  tests and `prepare` builds `dist/`; a red suite stops the publication; then it creates the GitHub release with that CHANGELOG section as its text. The trusted publisher is declared
1365
1383
  once on npmjs.com (package settings → Trusted Publisher → GitHub Actions, `Sylad/cadence`, `publish.yml`).
1366
- 3. Watch the run: `gh run watch` (or `gh run list --workflow publish.yml`).
1367
- 4. Check the effect: `npm view @sylad/cadence version` answers the new version.
1384
+ 4. Watch the run: `gh run watch` (or `gh run list --workflow publish.yml`).
1385
+ 5. Check the effect: `npm view @sylad/cadence version` answers the new version.
1368
1386
  A run is safe to re-run, and two runs for one tag queue instead of racing (`concurrency` per ref, never cancelling the one that publishes):
1369
1387
  a version already on npm skips `npm publish`, and a GitHub release that is missing is created (`--verify-tag`)
1370
1388
  while an existing one is left alone. If the package is on npm but the release is still missing, re-run the job;
@@ -1373,7 +1391,61 @@ A version exists in three places and is published in two; a release does all of
1373
1391
  The Claude Code plugin is read from the repository, so pushing `main` is what updates it; npm is what
1374
1392
  `npx @sylad/cadence` and a global install read, and only the tag publishes there. A missing tag, or a red
1375
1393
  publish run, leaves npm behind without any other error — 0.3.0 and 0.4.0 were never published — hence
1376
- step 4.
1394
+ step 5.
1395
+
1396
+ ### Evaluating the plugin (before a release)
1397
+
1398
+ `evals/` holds an evaluation suite for the plugin itself, run by `claude plugin eval` (Claude Code 2.1.263 and later). Each case is
1399
+ a folder with a `prompt.md` and `graders/*.md`, and replays a trap found by hand: `raf-done-sans-revue` (`raf done` refused on a lot whose
1400
+ commits have no code review, with no `--force` to get round it), `deliver-sha` (`git push origin <sha>:main` then
1401
+ `cadence deliver --sha <sha>` for a pushed commit that is not `HEAD`), `orchestrate-id-avec-slash` (`maritime-atlas:Q4/accueil-4-ux12@haiku`
1402
+ keeps its `/`) and `livrer-sh-sans-bloc` (a `./livrer.sh` without a `deliver:` block in `cadence.yaml` is not
1403
+ run by `cadence deliver`). Every case also runs a **baseline without the plugin**, so the report gives the score of each arm and the
1404
+ difference: a trap the baseline already avoids proves nothing about the plugin. The cases are read-only questions (tools `Read`, `Glob`, `Grep`,
1405
+ `Skill`, no scaffold script), each graded by a regular expression and by a model-judged criterion.
1406
+
1407
+ ```sh
1408
+ npm run eval:plugin # claude plugin eval . --max-cost-usd 5 --no-publish; 4 cases × 3 runs × 2 arms = 24 runs
1409
+ claude plugin eval . --case deliver-sha --runs 1 --trust-plugin --no-publish --max-cost-usd 1 # one case, one run (the first run asks to trust the plugin)
1410
+ ```
1411
+
1412
+ **Every run is billed** (a full `claude` session on your own credential, plus the judge): the suite is run at the release, once the
1413
+ version commit is ready, never in `npm test`, in `prepublishOnly` or in CI, and `--max-cost-usd 5` aborts a runaway. A case below
1414
+ the threshold (1.0 by default, `--threshold`) makes the command exit 1: read the case in the report before deciding it is the plugin
1415
+ and not the wording of the case. `npm test` only checks that the
1416
+ files are well formed (`test/plugin-evals.test.ts`); `evals/` is not part of the npm package.
1417
+
1418
+ ## Stability
1419
+
1420
+ From 1.0.0 cadence follows [semantic versioning](https://semver.org): what is listed under
1421
+ *Guaranteed* only changes incompatibly in a major version; a minor or a patch never breaks it.
1422
+
1423
+ **Guaranteed**
1424
+
1425
+ - **Commands and options** of `raf` and `cadence` (`cadence raf …` is `raf …`): `init`, `add`, `start`, `done`, `drop`,
1426
+ `note`, `did`, `public`, `ux`, `review`, `commits`, `show`, `list`, `now`, `ignore`, `check`, `gantt`, `hook`;
1427
+ `cadence news`, `session` (`start`, `close`, `next`, `context`), `lead tour`, `deliver`, `verify`, `orchestrate`,
1428
+ `skills`; their options and exit codes as this README documents them. A new command or option is a minor change.
1429
+ - **The keys of `cadence.yaml`**: `plan`, `news`, `hook`, `session`, `deliver`, `orchestrate`, `docs`, `qa`, `priority`
1430
+ and the keys each of them takes. A key is not renamed nor removed, nor does its meaning change; a new key is a minor change.
1431
+ - **The plan file** (`docs/plan/raf.yaml`): `project`, `prefix`, `since`, `ignore`, `lots` (with their
1432
+ fields, `tasks`, `notes`, `repos`) and `acknowledged`, and the plan formats `plan:` in `cadence.yaml` can read.
1433
+ The `version` key is kept in the file but ignored (`src/plan.ts` does not read it): it is not a schema version.
1434
+ - **The names of the skills** (`session-start`, `session-close`, `deliver`, `lead`) **and of the agents**
1435
+ (`ux-reviewer`, `code-reviewer`, `qa-reviewer`).
1436
+
1437
+ **Not guaranteed** (may change in any release)
1438
+
1439
+ - The **text of the outputs**: wording, layout, order of the lines and columns of every report, table and message
1440
+ (the `--json` outputs and the exit codes are part of the commands above). In a `--json` output a new field is a
1441
+ minor change, so consumers must tolerate unknown fields; removing or renaming a field is a major change.
1442
+ - The **`precheck-reader`** agent: internal to `cadence orchestrate`, not meant to be called by hand.
1443
+ - The **`cadence-hud` band**: its segments, colours and layout; it is versioned on its own.
1444
+ - The **wave journals** under `.cadence/runs` and the files `cadence orchestrate` keeps there.
1445
+ - The wording of the briefs and of the instructions inside the skills and the agents.
1446
+
1447
+ There is no schema version and no migration: a file written by an earlier release keeps working, and a key a
1448
+ release adds is optional.
1377
1449
 
1378
1450
  ## License
1379
1451
 
@@ -15,9 +15,7 @@ that a defect. You do: a players page with no players is a defect, whatever the
15
15
  ## Inputs
16
16
 
17
17
  The absolute path of the repository and the base URL of the app — deployed, or a local server the
18
- caller started. Optionally a lot id: then start with the pages that lot touched (its title and
19
- notes in the plan, and `raf commits <id>`, tell which) — when the lot touched only the backend,
20
- the pages that call the changed endpoints — and walk the others after. If the path or the URL is
18
+ caller started. Optionally a lot id: then walk only the pages that lot touched, as « Scope of a lot » below says. If the path or the URL is
21
19
  missing, or the URL does not answer, say so and stop.
22
20
 
23
21
  ## Method
@@ -46,6 +44,11 @@ missing, or the URL does not answer, say so and stop.
46
44
  and **390 px** wide. Walk within the bounded pass below. Let it settle: after `load`, wait a fixed few seconds, scroll through the
47
45
  page (lazy images), wait again — never for network idle, which streams and polling never reach.
48
46
  Then measure:
47
+ - read the page as text first: its rendered text and structure (`innerText`, or the
48
+ accessibility snapshot), the counts of the selectors named by the expectations, the responses
49
+ you listened to. Compare that text with the expectations, and take a capture only for a gap —
50
+ a `shows:` line unmet, a `never:` text found, a failed call — as evidence of that gap: a page
51
+ that matches its expectations gets none;
49
52
  - browser state, once before the first page, in a profile already used (a persistent context, not a fresh one — a `userDataDir` reserved for QA and kept between passes, never the user's own browser profile): compare the bundle the page loaded (its script URL) with the one `index.html` references, re-read without cache — a returning visitor still holds the old one, so a difference is stated in the report — then clear the cache and measure;
50
53
  - the expected content is present and non-empty — name the selector or the text found and its
51
54
  count (`.player-card` ×14), not "the list looks fine";
@@ -96,6 +99,32 @@ missing, or the URL does not answer, say so and stop.
96
99
  false, or an error is shown to the user), *major* (secondary content missing or wrong, a section
97
100
  silently dropped after a failed or empty API call, a broken content image), *minor* (noise). A broken line of the expectations with no visible loss on the page (the content is on screen by another path) is *minor* too.
98
101
 
102
+ ## Scope of a lot
103
+
104
+ With a lot id, walk only the pages the lot touched — the cost of a pass is the pages walked, and
105
+ a delivery rarely changes them all:
106
+
107
+ - the pages that call the changed endpoints (`raf commits <id>` and the diff tell which routes
108
+ changed; the `api:` lines of the expectations, and the calls you see a page make, tell which
109
+ pages use them);
110
+ - the pages of the screens the lot changed (its title and notes in the plan, and
111
+ the components in its commits, name them);
112
+ - the home page.
113
+
114
+ A backend-only lot is kept in scope: it can empty a page without touching a screen, and the pages
115
+ that call its endpoints are exactly the ones to walk. It also touches the pages that consume the
116
+ services the lot changed — grep for their callers, from the changed service up to the endpoint
117
+ that uses it, and walk the pages of those endpoints. If a backend lot changes no route (a service, a
118
+ data source or a configuration changed under routes that keep their path) and no screen, the scope
119
+ cannot come from the routes: the fallback is mandatory: walk every page and say so in the report.
120
+ If you cannot link the diff to any page, or a backend part of it in a lot that also changes
121
+ screens (a shared helper, a utility, a data source), the fallback is mandatory too: walk every page
122
+ and say so in the report. A lot that changes only screens keeps the reduced scope:
123
+ the pages of those screens, plus the home page. If you cannot tell which pages a changed
124
+ endpoint feeds, walk every page and say so — a scope you cannot establish reduces nothing. Pages
125
+ outside the scope are named under « Not walked », never counted as checked. Without a lot id,
126
+ walk every page.
127
+
99
128
  ## Bounded pass
100
129
 
101
130
  The pass has a time budget: the one the caller names, otherwise 15 minutes. You keep the count
@@ -117,6 +146,7 @@ from the first page.
117
146
 
118
147
  A short report:
119
148
 
149
+ - **Scope**, one line: the lot id (or "none: every page"), the pages walked of pages in the expectations (or routes discovered), a « Not walked » list naming each page not walked and why, and the captures taken — the figures to compare from one pass to the next.
120
150
  - **Pages checked N/N**, with the base URL and the date and time of the run, and the two widths. The
121
151
  second N is every page of the expectations (or every route discovered): a page you could not open
122
152
  is counted and named, never dropped. A page counts as checked when both widths were measured; a
package/dist/audit.js CHANGED
@@ -192,14 +192,15 @@ export function lotWork(plan, root, lotId) {
192
192
  /**
193
193
  * Sous-tâches encore ouvertes d'un lot que citent ses commits de travail (`feat(L3/t2)`), avec le plus récent des commits qui
194
194
  * les citent (L82) : le travail est fait, il ne reste que le plan à tenir. Un commit qui ne cite que le lot ne couvre rien.
195
+ * `neighbourWork` : les commits du lot dans ses dépôts voisins (`repoWork`, L62), lus après ceux du projet (L156).
195
196
  */
196
- export function coveredTasks(plan, root, lotId) {
197
+ export function coveredTasks(plan, root, lotId, neighbourWork = []) {
197
198
  const lot = plan.lots().find((l) => l.id === lotId);
198
199
  if (!lot)
199
200
  return [];
200
201
  const open = new Set(lot.tasks.filter((t) => isOpen(t.status)).map((t) => t.id));
201
202
  const found = new Map();
202
- for (const c of lotWork(plan, root, lotId)) {
203
+ for (const c of [...lotWork(plan, root, lotId), ...neighbourWork]) {
203
204
  const cited = citedRefs(c, plan.refs).filter((r) => r.lot === lotId && r.task);
204
205
  // `fix(L3/t1,t2)` : le motif des références ne lit que la première sous-tâche d'une liste à virgule.
205
206
  // Même texte que citedRefs : la portée quand elle cite des lots, sinon le message entier (le corps ne compte pas à côté d'une portée).
@@ -208,7 +209,7 @@ export function coveredTasks(plan, root, lotId) {
208
209
  const scope = scopeOf(c.subject);
209
210
  const text = scope !== null && plan.refs(scope).length > 0 ? scope : `${c.subject}\n${c.body}`;
210
211
  const element = '[\\w-]+(?:\\.[\\w-]+)*';
211
- const listed = cited.length > 0 ? [...text.matchAll(new RegExp(`(?<![\\w/])${escapeRe(lotId)}/(${element}(?:\\s*,\\s*${element})*)`, 'g'))].flatMap((m) => m[1].split(/\s*,\s*/)).flatMap((e) => plan.refs(`${lotId}/${e}`).filter((r) => r.lot === lotId && r.task).map((r) => r.task)) : [];
212
+ const listed = cited.length > 0 ? [...text.matchAll(new RegExp(`(?<![\\w/.-])${escapeRe(lotId)}/(${element}(?:\\s*,\\s*(?:${escapeRe(lotId)}/)?${element})*)`, 'g'))].flatMap((m) => m[1].split(/\s*,\s*/)).map((e) => (e.startsWith(`${lotId}/`) ? e.slice(lotId.length + 1) : e)).flatMap((e) => plan.refs(`${lotId}/${e}`).filter((r) => r.lot === lotId && r.task).map((r) => r.task)) : [];
212
213
  for (const task of [...cited.map((r) => r.task), ...listed])
213
214
  if (open.has(task) && !found.has(task))
214
215
  found.set(task, c.sha);
package/dist/config.js CHANGED
@@ -226,7 +226,7 @@ const REVIEW_KEYS = ['threshold', 'light', 'full'];
226
226
  const ORCH_KEYS = ['start', 'verdict', 'test', 'build', 'ux', 'precheck', 'review', 'effort', 'permissionMode', 'addDirs', 'timeouts'];
227
227
  /** Clé `orchestrate:` de cadence.yaml. Absente : les défauts (auto, 45 min d'implémentation, 25 min de revue). */
228
228
  export function readOrchestrateConfig(file) {
229
- const config = { precheck: true, review: { threshold: 0.25, light: 'sonnet', full: 'opus' }, effort: { precheck: 'low', implement: 'medium', fix: 'medium', review: 'high', ux: 'high' }, permissionMode: 'auto', addDirs: [], timeouts: { work: 45 * 60_000, review: 25 * 60_000 } };
229
+ const config = { precheck: true, precheckLocal: false, review: { threshold: 0.25, light: 'sonnet', full: 'opus' }, effort: { precheck: 'low', implement: 'medium', fix: 'medium', review: 'high', ux: 'high' }, permissionMode: 'auto', addDirs: [], timeouts: { work: 45 * 60_000, review: 25 * 60_000 } };
230
230
  if (!existsSync(file))
231
231
  return config;
232
232
  let raw;
@@ -261,9 +261,10 @@ export function readOrchestrateConfig(file) {
261
261
  config[k] = o[k].trim();
262
262
  }
263
263
  if (o.precheck != null) {
264
- if (typeof o.precheck !== 'boolean')
265
- throw bad('precheck : true ou false attendu');
266
- config.precheck = o.precheck;
264
+ if (typeof o.precheck !== 'boolean' && o.precheck !== 'local')
265
+ throw bad('precheck : true, false ou local attendu');
266
+ config.precheck = o.precheck !== false;
267
+ config.precheckLocal = o.precheck === 'local';
267
268
  }
268
269
  if (o.review != null) {
269
270
  if (!isObject(o.review))
@@ -15,7 +15,7 @@ import { loadTemplates, newsText, objective, renderBrief } from './briefs.js';
15
15
  import { candidates, parsePriority, parseUntil, readPriority, stopReason } from './continue.js';
16
16
  import { Budget, DROPPED, MAX_PASSES, countInterrupted, handBackDroppedQuestions, needsPrecheck } from './cycle.js';
17
17
  import { canInstallPrePush, installPrePush, removePrePush, snapshot } from './guard.js';
18
- import { buildArgs, killSessions, mcpServersFor, readAgents, realClaude } from './launch.js';
18
+ import { buildArgs, buildLocalArgs, killSessions, mcpServersFor, readAgents, realClaude } from './launch.js';
19
19
  import { activeLock, REPO_LOCK, releaseLock, takeLock } from './lock.js';
20
20
  import { runPool } from './pool.js';
21
21
  import { schemaFor } from './schemas.js';
@@ -292,7 +292,7 @@ function dryRun(lots, io, deps, budget, id) {
292
292
  }
293
293
  const steps = [{ kind: 'implement', model: l.model }];
294
294
  if (needsPrecheck(env.config, env.loadPlan(), l.repo, l.lot, l.repos))
295
- steps.unshift({ kind: 'precheck', model: 'sonnet' });
295
+ steps.unshift({ kind: 'precheck', model: env.config.precheckLocal ? 'local' : 'sonnet' });
296
296
  const { light, full } = env.config.review;
297
297
  if (l.small)
298
298
  steps.push({ kind: l.visible ? 'review-small' : 'review', model: l.light ? light : full });
@@ -312,8 +312,9 @@ function dryRun(lots, io, deps, budget, id) {
312
312
  const brief = renderBrief(s.kind, vars, deps.templatesDir);
313
313
  writeFileSync(file, brief);
314
314
  const playwright = !!mcpServersFor(s.kind, l.visible, '', l.small).playwright;
315
- const args = buildArgs({ kind: s.kind, sessionId: '<uuid>', brief: '<brief>', model: s.model, effort: env.config.effort[s.kind === 'review-small' ? 'review' : s.kind], schema: schemaFor(s.kind), agent: s.kind === 'implement' ? undefined : s.kind === 'ux' ? 'ux-reviewer' : s.kind === 'precheck' ? 'precheck-reader' : 'code-reviewer', cwd: l.repo, wave: id, permissionMode: env.config.permissionMode, addDirs: [...env.config.addDirs, ...(playwright && s.kind === 'implement' ? [pwDir] : []), ...neighbourDirs(l)], timeoutMs: 0, mcpConfig: '<mcp>', playwright }, agents).map((a) => (a.startsWith('{') ? '<json>' : a));
316
- io.out(` ${s.kind} : claude ${args.join(' ')}`);
315
+ const local = s.model === 'local';
316
+ const args = (local ? buildLocalArgs : buildArgs)({ kind: s.kind, sessionId: '<uuid>', brief: '<brief>', model: s.model === 'local' ? 'sonnet' : s.model, local, effort: env.config.effort[s.kind === 'review-small' ? 'review' : s.kind], schema: schemaFor(s.kind), agent: s.kind === 'implement' ? undefined : s.kind === 'ux' ? 'ux-reviewer' : s.kind === 'precheck' ? 'precheck-reader' : 'code-reviewer', cwd: l.repo, wave: id, permissionMode: env.config.permissionMode, addDirs: [...env.config.addDirs, ...(playwright && s.kind === 'implement' ? [pwDir] : []), ...neighbourDirs(l)], timeoutMs: 0, mcpConfig: '<mcp>', playwright }, agents).map((a) => (a.startsWith('{') ? '<json>' : a));
317
+ io.out(` ${s.kind} : ${local ? 'claude-local' : 'claude'} ${args.join(' ')}`);
317
318
  io.out(` brief : ${file}`);
318
319
  const servers = Object.keys(mcpServersFor(s.kind, l.visible, '', l.small));
319
320
  io.out(` mcp : ${servers.length ? `${servers.join(', ')} (captures dans le dossier de la vague : <vague>/${lotSlug(l.project, l.lot)}/playwright)` : 'aucun'}`);
@@ -1020,7 +1021,7 @@ export function realOrchestrateDeps(env) {
1020
1021
  // Lues ici, une fois : ni les sessions ni les commandes du projet ne reçoivent les variables de relance.
1021
1022
  env = withoutLaunchVars({ ...env });
1022
1023
  return {
1023
- claude: realClaude(bin, env),
1024
+ claude: realClaude(bin, env, env.CADENCE_CLAUDE_LOCAL_BIN || 'claude-local'),
1024
1025
  claudeInfo: () => {
1025
1026
  try {
1026
1027
  const version = execFileSync(bin, ['--version'], { env, encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'], timeout: 20_000 }).trim();
@@ -2,7 +2,7 @@ import { execFileSync, spawn } from 'node:child_process';
2
2
  import { toolBinOf, withoutLaunchVars } from './snapshot.js';
3
3
  import { randomUUID } from 'node:crypto';
4
4
  import { dirname, join, relative } from 'node:path';
5
- import { writeFileSync } from 'node:fs';
5
+ import { copyFileSync, existsSync, mkdirSync, renameSync, rmSync, writeFileSync } from 'node:fs';
6
6
  import { coveredTasks, lotCommits, lotWork } from '../audit.js';
7
7
  import { readDocsConfig } from '../config.js';
8
8
  import { docSyncBrief, docSyncGaps, filesOf } from '../docsync.js';
@@ -39,8 +39,8 @@ export class Budget {
39
39
  * session (`--session-id`). Une étape déjà comptée (`tokens`) ne l'est jamais deux fois. Vrai quand quelque chose a été compté.
40
40
  */
41
41
  export function countInterrupted(claudeHome, repo, step, budget) {
42
- if (!claudeHome || !step.sessionId || step.tokens)
43
- return false;
42
+ if (!claudeHome || !step.sessionId || step.tokens || step.model === 'local')
43
+ return false; // le modèle local n'a rien pris au quota
44
44
  const spent = journalTokens(claudeHome, repo, step.sessionId);
45
45
  if (!spent)
46
46
  return false;
@@ -293,7 +293,9 @@ function reviewModel(c, kind) {
293
293
  return c.lot.light && c.lot.pass === 0 && (kind === 'review' || kind === 'review-small') ? light : full;
294
294
  }
295
295
  /** Une session : budget et quota vérifiés avant, état écrit avant et après, journal gardé, contrôles du dépôt après. */
296
- async function session(c, kind) {
296
+ /** Rendu par `session(…, true)` quand l'étape locale n'a rien donné (Ollama éteint, délai, rapport sans structure) : l'appelant se replie sur Sonnet (L146). */
297
+ const LOCAL_FAILED = Symbol('local-failed');
298
+ async function session(c, kind, local = false) {
297
299
  const w = c.wave;
298
300
  const l = c.lot;
299
301
  if (dropped(c))
@@ -305,7 +307,7 @@ async function session(c, kind) {
305
307
  // Le budget du lot borne l'écriture, jamais la revue (L145) : une revue est jouée dès que le code est à relire (sauf tests rouges après une écriture, cf. REVIEW_RESERVE), la passe fix réserve son coût.
306
308
  if (write && (lotOver(l) || (kind === 'fix' && !fixAffordable(l))))
307
309
  return overBudget(c, kind === 'fix' && !lotOver(l));
308
- const model = write ? l.model : kind === 'precheck' ? 'sonnet' : reviewModel(c, kind); // le contrôle préalable ne fait que lire : pas d'Opus
310
+ const model = write ? l.model : local ? 'local' : kind === 'precheck' ? 'sonnet' : reviewModel(c, kind); // le contrôle préalable ne fait que lire : pas d'Opus
309
311
  const effort = c.config.effort[kind === 'review-small' ? 'review' : kind];
310
312
  const before = await snapshot(l.repo);
311
313
  const neighbours = l.repos ?? [];
@@ -335,8 +337,9 @@ async function session(c, kind) {
335
337
  kind,
336
338
  sessionId,
337
339
  brief,
338
- model,
340
+ model: local ? 'sonnet' : model,
339
341
  effort,
342
+ local,
340
343
  schema: schemaFor(kind),
341
344
  agent: kind === 'ux' ? 'ux-reviewer' : kind === 'precheck' ? 'precheck-reader' : write ? undefined : 'code-reviewer',
342
345
  cwd: l.repo,
@@ -359,7 +362,7 @@ async function session(c, kind) {
359
362
  step.ended = new Date().toISOString();
360
363
  const dir = w.store.lotDir(l.project, l.lot);
361
364
  const base = `${n}-${kind}`;
362
- if (outcome.kind !== 'ok') {
365
+ if (outcome.kind !== 'ok' && !local) {
363
366
  // Une session en échec, au quota ou tuée au délai a consommé des tokens : ceux de sa sortie, sinon ceux de son journal.
364
367
  const spent = outcome.tokens ? { tokens: outcome.tokens, sessionId: outcome.sessionId } : w.claudeHome ? journalTokens(w.claudeHome, l.repo, outcome.sessionId ?? sessionId) : null;
365
368
  if (spent) {
@@ -372,6 +375,20 @@ async function session(c, kind) {
372
375
  step.peakContext = peakContext(w.claudeHome, l.repo, spent.sessionId);
373
376
  }
374
377
  }
378
+ if (local && outcome.kind !== 'ok') {
379
+ // Hors quota Anthropic : rien au budget (ni sortie, ni journal), et jamais d'arrêt du lot — Sonnet reprend le contrôle.
380
+ const cause = outcome.kind === 'failed' ? outcome.cause : outcome.message;
381
+ writeFileSync(join(dir, `${base}.json`), outcome.kind === 'failed' ? outcome.stdout : '');
382
+ if (outcome.kind === 'failed')
383
+ writeFileSync(join(dir, `${base}.err`), outcome.stderr);
384
+ step.status = 'failed';
385
+ step.cause = cause;
386
+ step.report = `${base}.json`;
387
+ delete step.tokens;
388
+ l.warnings.push(`contrôle préalable : modèle local sans résultat (${cause.slice(0, 200)}), repli sur Sonnet`);
389
+ save(c);
390
+ return LOCAL_FAILED;
391
+ }
375
392
  if (outcome.kind === 'failed') {
376
393
  if (outcome.firstStdout !== undefined)
377
394
  writeFileSync(join(dir, `${base}.first.json`), outcome.firstStdout);
@@ -399,12 +416,13 @@ async function session(c, kind) {
399
416
  writeFileSync(join(dir, `${base}.json`), JSON.stringify({ ...res, structured: res.structured }, null, 2));
400
417
  step.report = `${base}.json`;
401
418
  step.sessionId = res.sessionId;
402
- step.tokens = res.tokens;
419
+ step.tokens = local ? { ...res.tokens, counted: 0 } : res.tokens; // local : hors quota, hors budget
403
420
  if (res.formattingRetry) {
404
421
  step.formatRetry = true;
405
422
  w.log(`${lotKey(l.project, l.lot)} · session ${n} ${kind} : rapport sans sortie structurée, relance de mise en forme`);
406
423
  }
407
- w.budget.add(res.tokens);
424
+ if (!local)
425
+ w.budget.add(res.tokens);
408
426
  w.saveWave();
409
427
  if (w.claudeHome)
410
428
  step.peakContext = peakContext(w.claudeHome, l.repo, res.sessionId);
@@ -429,20 +447,91 @@ async function session(c, kind) {
429
447
  w.saveWave();
430
448
  return stop(c, 'failed', `incident : ${w.incident}`);
431
449
  }
450
+ // Étape en lecture seule qui n'a laissé que des fichiers NON SUIVIS (captures d'un script en chemin relatif, L162) : déplacés dans le dossier du lot, avertissement, le verdict est gardé.
451
+ if (!write && b.head === a.head && b.tracked.join() === a.tracked.join() && b.untracked.join() !== a.untracked.join() && b.untracked.every((f) => a.untracked.includes(f))) {
452
+ const fresh = a.untracked.filter((f) => !b.untracked.includes(f)).map(unquotePath);
453
+ const stray = join(dir, 'stray');
454
+ const to = where ? join(stray, relative(l.repo, path).replace(/[\\/]/g, '_')) : stray;
455
+ const moved = moveStray(path, fresh, to);
456
+ if (moved.length) {
457
+ const text = `${kind} : ${moved.length} fichier(s) non suivi(s) laissé(s)${where} par une étape en lecture seule, déplacé(s) dans ${to} : ${moved.join(', ')}`;
458
+ c.lot.warnings.push(text);
459
+ w.log(`${lotKey(l.project, l.lot)} · ${text}`);
460
+ }
461
+ if (moved.length === fresh.length)
462
+ continue;
463
+ // Déplacement partiel : l'incident ne nomme que ce qui reste dans le dépôt.
464
+ return handBack(c, kind, res.structured, step, where, path, b, a, fresh.filter((f) => !moved.includes(f)));
465
+ }
432
466
  if (!write && (b.head !== a.head || b.tracked.join() !== a.tracked.join() || b.untracked.join() !== a.untracked.join())) {
433
- // Incident du LOT (L133) : la revue a laissé des traces (captures, commit), ni push ni garde supprimée — les lots des autres dépôts continuent.
434
- step.status = 'failed';
435
- step.cause = `le dépôt${where} a changé pendant une revue`;
436
- const traces = [...trackedPaths(a).filter((f) => !trackedPaths(b).includes(f)), ...a.untracked.filter((f) => !b.untracked.includes(f))];
437
- (w.dirtyRepos ??= new Map()).set(path, lotKey(l.project, l.lot));
438
- const what = traces.length ? ` (${traces.join(', ')})` : b.head !== a.head ? ' (commit)' : '';
439
- return stop(c, 'handed-back', `incident : ${kind} de ${lotKey(l.project, l.lot)} a modifié le dépôt${where}${what}${lostVerdict(c, kind, res.structured)}`);
467
+ return handBack(c, kind, res.structured, step, where, path, b, a);
440
468
  }
441
469
  }
442
470
  step.status = 'ok';
443
471
  save(c);
444
472
  return { step, report: res.structured, before, after, others };
445
473
  }
474
+ /** Incident du LOT (L133) : la revue a laissé des traces (captures, commit), ni push ni garde supprimée — les lots des autres dépôts continuent. */
475
+ function handBack(c, kind, report, step, where, path, b, a, left) {
476
+ const l = c.lot;
477
+ const w = c.wave;
478
+ step.status = 'failed';
479
+ step.cause = `le dépôt${where} a changé pendant une revue`;
480
+ const traces = left ?? [...trackedPaths(a).filter((f) => !trackedPaths(b).includes(f)), ...a.untracked.filter((f) => !b.untracked.includes(f))];
481
+ (w.dirtyRepos ??= new Map()).set(path, lotKey(l.project, l.lot));
482
+ const what = traces.length ? ` (${traces.join(', ')})` : b.head !== a.head ? ' (commit)' : '';
483
+ return stop(c, 'handed-back', `incident : ${kind} de ${lotKey(l.project, l.lot)} a modifié le dépôt${where}${what}${lostVerdict(c, kind, report)}`);
484
+ }
485
+ /** Décode un chemin de `git status --porcelain` : entre guillemets, avec les échappements du C-quoting (`\303\251`, `\"`, `\\`, `\t`…). */
486
+ function unquotePath(f) {
487
+ if (!/^".*"$/s.test(f))
488
+ return f;
489
+ const bytes = [];
490
+ const esc = { a: 7, b: 8, t: 9, n: 10, v: 11, f: 12, r: 13 };
491
+ const body = f.slice(1, -1);
492
+ for (let i = 0; i < body.length; i++) {
493
+ const ch = body[i];
494
+ if (ch !== '\\') {
495
+ bytes.push(...Buffer.from(ch));
496
+ continue;
497
+ }
498
+ const oct = /^[0-7]{3}/.exec(body.slice(i + 1));
499
+ if (oct) {
500
+ bytes.push(parseInt(oct[0], 8));
501
+ i += 3;
502
+ }
503
+ else {
504
+ const n = body[++i];
505
+ bytes.push(...(n in esc ? [esc[n]] : Buffer.from(n)));
506
+ }
507
+ }
508
+ return Buffer.from(bytes).toString('utf8');
509
+ }
510
+ /** Déplace des fichiers non suivis (chemins relatifs au dépôt, décodés) sous `to`, arborescence gardée ; renvoie ceux qui l'ont été (s'arrête au premier échec). */
511
+ function moveStray(repo, files, to) {
512
+ const moved = [];
513
+ for (const rel of files) {
514
+ try {
515
+ const from = join(repo, rel);
516
+ let dest = join(to, rel);
517
+ for (let i = 1; existsSync(dest); i++)
518
+ dest = join(to, `${rel}.${i}`);
519
+ mkdirSync(dirname(dest), { recursive: true });
520
+ try {
521
+ renameSync(from, dest);
522
+ }
523
+ catch {
524
+ copyFileSync(from, dest); // autre volume
525
+ rmSync(from);
526
+ }
527
+ moved.push(rel);
528
+ }
529
+ catch {
530
+ break;
531
+ }
532
+ }
533
+ return moved;
534
+ }
446
535
  /** Verdict d'une revue interrompue par un incident du lot : rapporté au lead plutôt que perdu (L133). Rien n'est enregistré dans le plan. */
447
536
  function lostVerdict(c, kind, report) {
448
537
  if (kind === 'precheck')
@@ -523,18 +612,44 @@ export function needsPrecheck(config, plan, repo, lot, repos = []) {
523
612
  */
524
613
  async function precheck(c) {
525
614
  const l = c.lot;
526
- const done = await session(c, 'precheck');
527
- if (!done)
528
- return;
529
- let rep;
530
- try {
531
- rep = checkShape(done.report, PRECHECK_SCHEMA);
615
+ let rep = null;
616
+ let localYes = false;
617
+ if (c.config.precheckLocal) {
618
+ // Modèle local (L146) : un échec, un rapport illisible ou un « oui » (qui rendrait le lot sans implémentation) rend la main à Sonnet.
619
+ const done = await session(c, 'precheck', true);
620
+ if (!done)
621
+ return;
622
+ if (done !== LOCAL_FAILED) {
623
+ try {
624
+ rep = checkShape(done.report, PRECHECK_SCHEMA);
625
+ }
626
+ catch (e) {
627
+ l.warnings.push(`contrôle préalable : rapport du modèle local illisible (${e.message}), repli sur Sonnet`);
628
+ }
629
+ if (rep?.dejaPresent === 'oui') {
630
+ localYes = true;
631
+ rep = null;
632
+ }
633
+ }
532
634
  }
533
- catch (e) {
534
- l.warnings.push(`contrôle préalable illisible, implémentation lancée : ${e.message}`);
535
- l.next = 'implement';
536
- save(c);
537
- return;
635
+ if (!rep) {
636
+ const done = await session(c, 'precheck');
637
+ if (!done)
638
+ return;
639
+ try {
640
+ rep = checkShape(done.report, PRECHECK_SCHEMA);
641
+ }
642
+ catch (e) {
643
+ l.warnings.push(`contrôle préalable illisible, implémentation lancée : ${e.message}`);
644
+ l.next = 'implement';
645
+ save(c);
646
+ return;
647
+ }
648
+ }
649
+ if (localYes) {
650
+ l.warnings.push(rep.dejaPresent === 'oui'
651
+ ? 'contrôle préalable : « oui » du modèle local, confirmé par Sonnet'
652
+ : `contrôle préalable : le modèle local a dit « oui », Sonnet dit « ${rep.dejaPresent} » — le local s'est trompé, repasser orchestrate.precheck à true`);
538
653
  }
539
654
  const proofs = rep.preuves.length ? ` (${rep.preuves.join(' ; ')})` : '';
540
655
  // « oui » sans preuve ne se vérifie pas : il vaut « partiel », l'implémentation part avec le constat.
@@ -857,7 +972,7 @@ async function conclude(c, code, minorNote = '') {
857
972
  }
858
973
  l.constats = [];
859
974
  // Sous-tâches ouvertes que des commits du lot citent (L82) : la revue conforme vaut pour elles, sinon `raf done` refuserait le lot.
860
- const covered = coveredTasks(plan, l.repo, l.lot);
975
+ const covered = coveredTasks(plan, l.repo, l.lot, neighbours.flatMap((r) => repoWork(plan, r, l.lot)));
861
976
  if (!plan.readonly) {
862
977
  plan.recordReview(l.lot, verdict, c.wave.today, newer, neighbours.length ? shas : undefined);
863
978
  for (const k of covered)
@@ -60,6 +60,24 @@ export function buildArgs(spec, agents) {
60
60
  args.push('--strict-mcp-config', '--mcp-config', spec.mcpConfig);
61
61
  return args;
62
62
  }
63
+ /**
64
+ * Arguments de `claude-local` (L146) : le prompt vient en premier, le script ajoute lui-même `--bare`, `--model` et le
65
+ * serveur Ollama. Le harnais `--bare` ne charge pas `--agents` : le prompt de l'agent est joint au brief et ses outils
66
+ * passent par `--tools`. Pas de `--effort`, pas de MCP (aucune étape locale n'en charge).
67
+ */
68
+ export function buildLocalArgs(spec, agents) {
69
+ const a = spec.agent ? agents[spec.agent] : undefined;
70
+ if (spec.agent && !a)
71
+ throw new RafError(`agent introuvable dans le paquet : ${spec.agent}`);
72
+ const args = [a ? `${a.prompt}\n\n${spec.brief}` : spec.brief, '--output-format', 'json', '--json-schema', JSON.stringify(spec.schema)];
73
+ if (a?.tools)
74
+ args.push('--tools', a.tools.join(','));
75
+ args.push('--session-id', spec.sessionId, '--permission-mode', spec.permissionMode);
76
+ for (const d of spec.addDirs)
77
+ args.push('--add-dir', d);
78
+ args.push('--disallowedTools', ...DISALLOWED);
79
+ return args;
80
+ }
63
81
  /** Définitions d'agents du paquet (`agents/*.md`) au format de --agents. */
64
82
  export function readAgents(dir) {
65
83
  const out = {};
@@ -107,9 +125,11 @@ export function buildRetryArgs(spec, sessionId, agents) {
107
125
  * s'ajoutent à ceux de la session ; toujours rien après elle, c'est l'échec habituel.
108
126
  */
109
127
  export async function runSession(spec, deps) {
110
- const launch = (args) => deps.claude(args, { cwd: spec.cwd, env: { CADENCE_ORCHESTRATED: spec.wave, ...(spec.nodeBin || spec.toolBin ? { PATH: [spec.toolBin, spec.nodeBin, process.env.PATH ?? ''].filter(Boolean).join(delimiter) } : {}) }, timeoutMs: spec.timeoutMs, onSpawn: deps.onSpawn });
111
- const out = await launch(buildArgs(spec, deps.agents));
128
+ const launch = (args) => deps.claude(args, { local: spec.local, cwd: spec.cwd, env: { CADENCE_ORCHESTRATED: spec.wave, ...(spec.nodeBin || spec.toolBin ? { PATH: [spec.toolBin, spec.nodeBin, process.env.PATH ?? ''].filter(Boolean).join(delimiter) } : {}) }, timeoutMs: spec.timeoutMs, onSpawn: deps.onSpawn });
129
+ const out = await launch(spec.local ? buildLocalArgs(spec, deps.agents) : buildArgs(spec, deps.agents));
112
130
  const first = classify(out);
131
+ if (spec.local)
132
+ return first; // pas de relance de mise en forme : l'appelant se replie sur Sonnet
113
133
  const missing = first.kind === 'failed' && !out.timedOut && out.code === 0 ? lacksStructuredOutput(out.stdout) : null;
114
134
  if (!missing)
115
135
  return first;
@@ -175,13 +195,13 @@ export function killSessions() {
175
195
  marks.clear();
176
196
  }
177
197
  /** Vrai lanceur : `claude` (ou CADENCE_CLAUDE_BIN) dans son propre groupe de processus, tué en bloc au délai. */
178
- export function realClaude(bin, base = process.env) {
198
+ export function realClaude(bin, base = process.env, localBin = 'claude-local') {
179
199
  return (args, opts) => new Promise((resolve) => {
180
200
  // Marque héritée par tous les descendants, même orphelins rattachés à init (un shell qui sort aussitôt après
181
201
  // `nohup srv &` échappe à tout relevé de l'arbre) : à la fin de la session, tout ce qui la porte est tué.
182
202
  const mark = randomUUID();
183
203
  marks.add(mark);
184
- const child = spawn(bin, args, { cwd: opts.cwd, env: { ...base, ...opts.env, [SESSION_MARK_VAR]: mark }, stdio: ['ignore', 'pipe', 'pipe'], detached: true });
204
+ const child = spawn(opts.local ? localBin : bin, args, { cwd: opts.cwd, env: { ...base, ...opts.env, [SESSION_MARK_VAR]: mark }, stdio: ['ignore', 'pipe', 'pipe'], detached: true });
185
205
  const pid = child.pid;
186
206
  if (pid === undefined) {
187
207
  marks.delete(mark);
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@sylad/cadence",
3
- "version": "0.27.1",
3
+ "version": "1.0.0",
4
4
  "description": "A small, repo-native working method: a versioned plan linked to your commits, a changelog with screenshots, session rituals and deliveries proven by their effect.",
5
5
  "license": "MIT",
6
6
  "author": "Sylvain Ladoire",
@@ -26,6 +26,7 @@
26
26
  "build": "tsc -p tsconfig.build.json",
27
27
  "test": "vitest run",
28
28
  "typecheck": "tsc --noEmit",
29
+ "eval:plugin": "claude plugin eval . --max-cost-usd 5 --no-publish",
29
30
  "prepare": "npm run build",
30
31
  "prepublishOnly": "npm run typecheck && npm test"
31
32
  },
@@ -57,8 +57,7 @@ Deliver one project at a time: the one the human names, or ask.
57
57
  6. After a green delivery that changes what a page shows or what it is served (screen, API, data
58
58
  source, configuration of either) — in practice every delivery except docs-, plan- or tests-only
59
59
  ones — have the `qa-reviewer` agent walk the delivered app in a real browser, whether the lot
60
- is `visible` or not: give it the repository path, the base URL and the lot id. When the lot
61
- touched only the backend, the agent starts with the pages that call the changed endpoints. It
60
+ is `visible` or not: give it the repository path, the base URL and the lot id. With a lot id, the agent walks only the pages the lot touched (the pages that call the changed endpoints, the pages of the changed screens, the pages that consume the services the lot changed, the home page; when the diff changes no route and no screen, or cannot be linked to pages, it walks every page and says so) and names the others under « Not walked »: they are not checked. It
62
61
  checks each page against `docs/qa/expectations.md` — what the user must find there — and
63
62
  reports a page left empty, an error shown, an API call that failed or came back empty: what
64
63
  the checks of `cadence.yaml` do not see. Bring its blocking findings to the human. It is not a
@@ -169,8 +169,7 @@ After a green delivery that changes what a page shows or what it is served (scre
169
169
  source, configuration of either) — in practice every delivery except docs-, plan- or tests-only ones
170
170
  — have the `qa-reviewer` agent check the delivered app, as a fresh subagent: give it the absolute
171
171
  path of the project, the base URL of the delivered app and the lot id. The lot need not be
172
- `visible`: a backend-only lot can empty a page without changing a screen. When the lot touched only
173
- the backend, the agent starts with the pages that call the changed endpoints. It walks the pages in
172
+ `visible`: a backend-only lot can empty a page without changing a screen. With a lot id, the agent walks only the pages the lot touched (the pages that call the changed endpoints, the pages of the changed screens, the pages that consume the services the lot changed, the home page; when the diff changes no route and no screen, or cannot be linked to pages, it walks every page and says so) and names the others under « Not walked »: they are not checked. It walks the pages in
174
173
  a real browser against the project's expectations (`docs/qa/expectations.md`: per page, what the
175
174
  user must find there) and returns measured findings; it reads only, and never logs in. Bring its
176
175
  blocking findings back to the human — a page whose main content is missing, or that shows an error,