@sylad/cadence 0.27.1 → 0.28.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/.claude-plugin/plugin.json +1 -1
- package/README.md +25 -8
- package/agents/qa-reviewer.md +33 -3
- package/dist/audit.js +4 -3
- package/dist/config.js +5 -4
- package/dist/orchestrate/command.js +6 -5
- package/dist/orchestrate/cycle.js +64 -20
- package/dist/orchestrate/launch.js +24 -4
- package/package.json +1 -1
- package/skills/deliver/SKILL.md +1 -2
- package/skills/lead/SKILL.md +1 -2
|
@@ -8,7 +8,7 @@
|
|
|
8
8
|
{
|
|
9
9
|
"name": "cadence",
|
|
10
10
|
"description": "Session start and close rituals driven by a versioned plan (raf), and deliveries proven by their effect. Needs the cadence CLI (npm i -g @sylad/cadence).",
|
|
11
|
-
"version": "0.
|
|
11
|
+
"version": "0.28.0",
|
|
12
12
|
"source": "./",
|
|
13
13
|
"author": {
|
|
14
14
|
"name": "Sylvain Ladoire"
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cadence",
|
|
3
3
|
"description": "A repo-native working method: session start and close rituals driven by a versioned plan (raf), deliveries proven by their effect, and three reviewer agents (UX, code, QA).",
|
|
4
|
-
"version": "0.
|
|
4
|
+
"version": "0.28.0",
|
|
5
5
|
"author": { "name": "Sylvain Ladoire" },
|
|
6
6
|
"homepage": "https://github.com/Sylad/cadence",
|
|
7
7
|
"repository": "https://github.com/Sylad/cadence",
|
package/README.md
CHANGED
|
@@ -24,7 +24,7 @@ Code prompt that shows the session's context, quota, cost and the orchestrate wa
|
|
|
24
24
|
|
|
25
25
|
## What's new
|
|
26
26
|
|
|
27
|
-
**0.27.1**: `cadence orchestrate` no longer leaves a session's child processes alive: a dev server started in its own process group, or orphaned by a shell that exits (`nohup srv &`), is killed with the session through its process tree and its `CADENCE_SESSION` mark, a tmux server, `screen`, `gpg-agent` or `dirmngr` started on demand being spared; a session that exits with code 0 while a descendant holds its output open is no longer reported as « délai dépassé » (L83); the `ux-reviewer` and `qa-reviewer` agents carry the briefs' Playwright rule outside `cadence orchestrate` (L84). **0.27.0**: a compliant review of `cadence orchestrate` settles the open sub-tasks that commits of the lot cite (`feat(L3/t2): …`, comma lists included), closed in the same plan commit as the verdict or, on a read-only plan, returned as `[sous-tâche clore]` proposals, so `raf done` no longer refuses a lot over sub-tasks whose work is done; on a foreign-format plan the ids are read as the plan reads them (L82); the test suite stays green with two full suites running in parallel (L150). **0.26.0**: `cadence orchestrate` raises the floor of a lot's budget from 200 k to 450 k tokens, enough to pay a write pass, its review, one fix pass and the short review that follows, so a lot of 0.5 or 1 day is no longer handed back at its second or third pass; a `--continue` wave draws fewer lots as a result (L128). **0.25.0**: the project line of the plans' progress in the `cadence-hud` band names the lots in progress after their count, three at most then `…` (L153, `cadence-hud` 0.4.2); the progress follows the folder the session was launched from, so a `cd` into a sub-project no longer shrinks the band, and the session context keeps the `session.start` folder (L154). **0.24.0**: `cadence orchestrate --drop <project:lot>` and `--stop-after-current` steer a live wave from another terminal, `--resume --drop` removes a lot before replaying a stopped wave (L79); the final table of a wave says « livrable jusqu'à <sha> » for a ready lot stacked under a lot handed back, with `git push origin <sha>:main && cadence deliver --sha <sha>` (L76); a version commit (version fields, `CHANGELOG.md`, the README's « What's new ») no longer asks for a new code review (L72). **0.23.0**: the progress bar of each project in the `cadence-hud` band is coloured by the share of lots done, red below 33 %, orange below 66 %, green beyond (L152, `cadence-hud` 0.4.1); `cadence deliver` under a wave lock refuses only when a lot of the wave still works in that very repository, a linked worktree included (L127). **0.22.0**: `cadence orchestrate --continue` draws the next ready lot of the plan itself, in the declared priority, until the budget, the time window or a question stops it (L147); `cadence session context` prints the context of the session, which the `cadence-hud` band publishes, and the `lead` skill chains lots without the human while it stays under 60 % (L148); the band shows the progress of each project's plan, and `cadence lead tour --json` gains `progress` (L149, `cadence-hud` 0.4.0); the final review of a lot is always played, the lot's own budget bounding the writing passes only, with 65 k tokens reserved for it (L145); the cause « tests rouges après … » names the pass whose tests were red (L151). **0.21.0**: `cadence orchestrate --status` and `cadence lead tour` print the free session slots, which the `lead` skill reads (L141); `orchestrate.effort` sets the effort level of each pass of a wave (L137); `docs.sync` makes `raf check`, `lead tour` and the review brief report a document that does not follow the code (L143); `docs.articles` opens, after a green `cadence deliver`, a lot `Article <project> à rafraîchir` in the neighbouring repository (L144). **0.20.0**: `cadence-hud` no longer misleads once the work is over: `/hud cmd` lists the background commands it counts (id, tool, origin, age, name), any other argument answers `argument inconnu` instead of hiding the band, commands with no end seen for an hour or no record go to a grey `N cmd sans fin vue`, a finished subagent takes its commands with it, two races that left a `1 cmd` counted forever are closed, and a finished wave shows grey without its budget then disappears after 30 minutes (L142); the `qa-reviewer` and `ux-reviewer` agents are read-only for real, their frontmatter listing Read, Grep, Glob, Bash and the Playwright tools (L140). **0.19.1**: `cadence-hud`'s `N cmd` counter goes back down when a Monitor expires, when the background commands a subagent started end with it, and at session start or plugin reload (L135); its `tool.call` and `prompt.submit` guard hooks carry a `.catch` on registration, which clears the `claude plugin validate` warning (L136). **0.19.0**: `cadence orchestrate` lifts a repository's lock and `pre-push` guard as soon as every lot of the wave that touches it is finished, so `cadence deliver` works there while the wave continues elsewhere (L132); a review that leaves a repository dirty hands back that lot only, suspends the lots still queued in that repository and keeps the other repositories running, its verdict reported instead of lost (L133); `cadence session close` re-run without a `session start` keeps the window since the last opening (L131); `cadence-hud` draws the lot rows of a wave in aligned columns on every surface (L134). **0.18.0**: `cadence session next` prints what it records (issue #6); `cadence session close` covers the period since the last `session start` (or the last close) and says so in its title, `--since` still wins (issue #7). **0.17.0**: `cadence-hud` shows an `N cmd` segment next to the agents, the background commands of the session (Bash run in the background and Monitor, subagents' included), tracked from the tool results and ended by the task's notification or `TaskStop`. **0.16.0**: a lot can declare **neighbouring repositories** (`repos:` key of the lot, `{ path, cite }` for a repository shared between projects): `cadence orchestrate` refuses, locks and guards them, gives them to every session as `--add-dir`, and `raf commits`, `raf show` and the review verdict read the commits that cite the lot there too. **0.15.0**: `hook.autostart` in `cadence.yaml` (`warn` by default, `refuse` or `start`): a pre-commit hook that refuses, or starts in the same commit, a lot still todo that the commit cites; `cadence session start|close --all [--depth n]`, one section per project under a folder. **0.14.0**: `cadence lead tour`, the lead's morning table of every project in a folder, printed by a program (one line per project: lots in progress, gaps, last notes, next ready lot, repository) in place of one subagent per project; the `lead` skill calls it. **0.13.1**: `cadence-hud` takes back the segments that fit once a big one has dropped (reset times, then the per-model consumption, now short: `opus 1.2M sonnet 800k`); `CADENCE_HUD_AMBIGUOUS=1` for a terminal that draws `▰ ▱ ⚙ │ ↻` in one cell. **0.13.0**: `cadence orchestrate --status --watch` (a wave's table redrawn in the terminal, which
|
|
27
|
+
**0.28.0**: `orchestrate.precheck: local` runs the « deliverable already present? » pre-check of `cadence orchestrate` on the local model (`claude-local`, Ollama), off the Anthropic quota, Sonnet staying the default and the fallback (L146); the `qa-reviewer` agent walks only the pages a lot touched, with a **Scope** line in its report (L123); the sub-tasks a review closes are also read in a lot's neighbouring repositories (L156) and a commit's sub-task list is read with the plan's left guard (L155); an unstable test is fixed (L90). **0.27.1**: `cadence orchestrate` no longer leaves a session's child processes alive: a dev server started in its own process group, or orphaned by a shell that exits (`nohup srv &`), is killed with the session through its process tree and its `CADENCE_SESSION` mark, a tmux server, `screen`, `gpg-agent` or `dirmngr` started on demand being spared; a session that exits with code 0 while a descendant holds its output open is no longer reported as « délai dépassé » (L83); the `ux-reviewer` and `qa-reviewer` agents carry the briefs' Playwright rule outside `cadence orchestrate` (L84). **0.27.0**: a compliant review of `cadence orchestrate` settles the open sub-tasks that commits of the lot cite (`feat(L3/t2): …`, comma lists included), closed in the same plan commit as the verdict or, on a read-only plan, returned as `[sous-tâche clore]` proposals, so `raf done` no longer refuses a lot over sub-tasks whose work is done; on a foreign-format plan the ids are read as the plan reads them (L82); the test suite stays green with two full suites running in parallel (L150). **0.26.0**: `cadence orchestrate` raises the floor of a lot's budget from 200 k to 450 k tokens, enough to pay a write pass, its review, one fix pass and the short review that follows, so a lot of 0.5 or 1 day is no longer handed back at its second or third pass; a `--continue` wave draws fewer lots as a result (L128). **0.25.0**: the project line of the plans' progress in the `cadence-hud` band names the lots in progress after their count, three at most then `…` (L153, `cadence-hud` 0.4.2); the progress follows the folder the session was launched from, so a `cd` into a sub-project no longer shrinks the band, and the session context keeps the `session.start` folder (L154). **0.24.0**: `cadence orchestrate --drop <project:lot>` and `--stop-after-current` steer a live wave from another terminal, `--resume --drop` removes a lot before replaying a stopped wave (L79); the final table of a wave says « livrable jusqu'à <sha> » for a ready lot stacked under a lot handed back, with `git push origin <sha>:main && cadence deliver --sha <sha>` (L76); a version commit (version fields, `CHANGELOG.md`, the README's « What's new ») no longer asks for a new code review (L72). **0.23.0**: the progress bar of each project in the `cadence-hud` band is coloured by the share of lots done, red below 33 %, orange below 66 %, green beyond (L152, `cadence-hud` 0.4.1); `cadence deliver` under a wave lock refuses only when a lot of the wave still works in that very repository, a linked worktree included (L127). **0.22.0**: `cadence orchestrate --continue` draws the next ready lot of the plan itself, in the declared priority, until the budget, the time window or a question stops it (L147); `cadence session context` prints the context of the session, which the `cadence-hud` band publishes, and the `lead` skill chains lots without the human while it stays under 60 % (L148); the band shows the progress of each project's plan, and `cadence lead tour --json` gains `progress` (L149, `cadence-hud` 0.4.0); the final review of a lot is always played, the lot's own budget bounding the writing passes only, with 65 k tokens reserved for it (L145); the cause « tests rouges après … » names the pass whose tests were red (L151). **0.21.0**: `cadence orchestrate --status` and `cadence lead tour` print the free session slots, which the `lead` skill reads (L141); `orchestrate.effort` sets the effort level of each pass of a wave (L137); `docs.sync` makes `raf check`, `lead tour` and the review brief report a document that does not follow the code (L143); `docs.articles` opens, after a green `cadence deliver`, a lot `Article <project> à rafraîchir` in the neighbouring repository (L144). **0.20.0**: `cadence-hud` no longer misleads once the work is over: `/hud cmd` lists the background commands it counts (id, tool, origin, age, name), any other argument answers `argument inconnu` instead of hiding the band, commands with no end seen for an hour or no record go to a grey `N cmd sans fin vue`, a finished subagent takes its commands with it, two races that left a `1 cmd` counted forever are closed, and a finished wave shows grey without its budget then disappears after 30 minutes (L142); the `qa-reviewer` and `ux-reviewer` agents are read-only for real, their frontmatter listing Read, Grep, Glob, Bash and the Playwright tools (L140). **0.19.1**: `cadence-hud`'s `N cmd` counter goes back down when a Monitor expires, when the background commands a subagent started end with it, and at session start or plugin reload (L135); its `tool.call` and `prompt.submit` guard hooks carry a `.catch` on registration, which clears the `claude plugin validate` warning (L136). **0.19.0**: `cadence orchestrate` lifts a repository's lock and `pre-push` guard as soon as every lot of the wave that touches it is finished, so `cadence deliver` works there while the wave continues elsewhere (L132); a review that leaves a repository dirty hands back that lot only, suspends the lots still queued in that repository and keeps the other repositories running, its verdict reported instead of lost (L133); `cadence session close` re-run without a `session start` keeps the window since the last opening (L131); `cadence-hud` draws the lot rows of a wave in aligned columns on every surface (L134). **0.18.0**: `cadence session next` prints what it records (issue #6); `cadence session close` covers the period since the last `session start` (or the last close) and says so in its title, `--since` still wins (issue #7). **0.17.0**: `cadence-hud` shows an `N cmd` segment next to the agents, the background commands of the session (Bash run in the background and Monitor, subagents' included), tracked from the tool results and ended by the task's notification or `TaskStop`. **0.16.0**: a lot can declare **neighbouring repositories** (`repos:` key of the lot, `{ path, cite }` for a repository shared between projects): `cadence orchestrate` refuses, locks and guards them, gives them to every session as `--add-dir`, and `raf commits`, `raf show` and the review verdict read the commits that cite the lot there too. **0.15.0**: `hook.autostart` in `cadence.yaml` (`warn` by default, `refuse` or `start`): a pre-commit hook that refuses, or starts in the same commit, a lot still todo that the commit cites; `cadence session start|close --all [--depth n]`, one section per project under a folder. **0.14.0**: `cadence lead tour`, the lead's morning table of every project in a folder, printed by a program (one line per project: lots in progress, gaps, last notes, next ready lot, repository) in place of one subagent per project; the `lead` skill calls it. **0.13.1**: `cadence-hud` takes back the segments that fit once a big one has dropped (reset times, then the per-model consumption, now short: `opus 1.2M sonnet 800k`); `CADENCE_HUD_AMBIGUOUS=1` for a terminal that draws `▰ ▱ ⚙ │ ↻` in one cell. **0.13.0**: `cadence orchestrate --status --watch` (a wave's table redrawn in the terminal, which
|
|
28
28
|
stops with the wave) and the lead skill that names the wave and follows its journal; lot ids with `/`
|
|
29
29
|
(sub-tasks of a read-only plan) accepted by `orchestrate`; `cadence-hud`'s first line measured in
|
|
30
30
|
terminal cells and kept to ASCII; `publish.yml` on the v5 actions. **0.12.0**: the `cadence-hud` plugin (a status band above the Claude Code prompt),
|
|
@@ -406,7 +406,7 @@ No gate and no command here: the QA review comes **after** a delivery, and
|
|
|
406
406
|
page shows or what it is served (screen, API, data source, configuration of
|
|
407
407
|
either) — in practice every delivery except docs-, plan- or tests-only ones: a
|
|
408
408
|
backend-only lot can empty a page without touching a screen, and the agent then
|
|
409
|
-
|
|
409
|
+
walks the pages that call the changed endpoints. It also walks the pages that consume the services the lot changed; when the diff changes no route and no screen, or cannot be linked to pages, it walks every page and says so. The `qa-reviewer` agent
|
|
410
410
|
opens each page of the running app in a real browser and judges it from the
|
|
411
411
|
user's side. A page can be empty while everything else is green — no code
|
|
412
412
|
changed, a data source went down upstream, the unit tests replace the network,
|
|
@@ -925,7 +925,8 @@ older `claude` is refused at launch unless every pass is `default`). Defaults: `
|
|
|
925
925
|
`implement` and `fix` medium, `review` high (`review-small` follows `review`), `ux` high. A project overrides any of them
|
|
926
926
|
in `cadence.yaml`; the levels are `low`, `medium`, `high`, `xhigh`, `max`, and `default` sends no `--effort` (the
|
|
927
927
|
level of the session, as before this key). The format-retry session (`--resume`) keeps the level of its pass, and
|
|
928
|
-
`--dry-run` shows the flag in each command line
|
|
928
|
+
`--dry-run` shows the flag in each command line, except for a `precheck: local` session: `claude-local` is launched
|
|
929
|
+
without `--effort` (see below), so `effort.precheck` has no effect on it and the `precheck (local)` line does not show it.
|
|
929
930
|
|
|
930
931
|
```yaml
|
|
931
932
|
orchestrate:
|
|
@@ -944,6 +945,20 @@ with proofs), `partiel` or `non`: on `oui` no implementation session is opened a
|
|
|
944
945
|
implementation brief and a warning; on `non` — or an unreadable report, which only adds a warning — the wave goes on. A
|
|
945
946
|
lot that already has commits is never pre-checked (resuming it is legitimate). `orchestrate.precheck: false` turns it off.
|
|
946
947
|
|
|
948
|
+
`orchestrate.precheck: local` (L146) runs the pre-check on the local model of the dev machine instead of Sonnet: the
|
|
949
|
+
orchestrator launches `claude-local` with the same brief, the `precheck-reader` prompt and tools, no quota used and nothing counted in the
|
|
950
|
+
lot's budget (the step shows `local` as its model). `claude-local` is **not shipped with cadence**: it is a personal wrapper (a script on the `PATH`, `qwen3-coder:30b` by default) around
|
|
951
|
+
`claude` pointed at Ollama. `CADENCE_CLAUDE_LOCAL_BIN` names another binary, which must take the prompt as its **first argument**
|
|
952
|
+
and add itself `-p`, `--bare`, `--strict-mcp-config` and `--model` (cadence passes none of them), then accept `--output-format`,
|
|
953
|
+
`--json-schema`, `--tools`, `--session-id`, `--permission-mode`, `--add-dir` and `--disallowedTools`. A bare `claude` does not fit.
|
|
954
|
+
Without such a binary every local check fails and is replayed on Sonnet: keep `true`. Sonnet stays the safety net: a local session with no result (Ollama
|
|
955
|
+
down, timeout, no structured report) or an unreadable report is replayed on Sonnet with a warning, and a local `oui` — the
|
|
956
|
+
only answer that hands the lot back without an implementation — is confirmed by Sonnet before it counts. A local `partiel`
|
|
957
|
+
or `non` is taken as is. The default stays `true` (Sonnet): switch to `local` only after comparing verdicts and durations
|
|
958
|
+
against Sonnet on a few lots, and go back to `true` the first time the local model gets a `partiel` wrong. `--dry-run`
|
|
959
|
+
prints the step as `precheck (local)` with its `claude-local` arguments. Those arguments carry neither `--model` nor `--effort` nor `--agents`: `effort.precheck` applies only to the Sonnet
|
|
960
|
+
session (the fallback, or the confirmation of a local `oui`).
|
|
961
|
+
|
|
947
962
|
The `precheck-reader` agent does not set `omitClaudeMd` (Claude Code 2.1.271), on purpose: measured on a project with a
|
|
948
963
|
297-line `CLAUDE.md` (haiku, `--agents` + `--agent`, same prompt, with and without the flag), the session context is
|
|
949
964
|
identical (22 779 tokens, the project `CLAUDE.md` still answered when asked) — the flag only applies to an agent run as a
|
|
@@ -970,7 +985,7 @@ committed, a short re-review of that commit runs instead): `ready`, verdict reco
|
|
|
970
985
|
the untreated minors returned as proposals. Same when the budget is exhausted right after a compliant
|
|
971
986
|
review with minors: it concludes on that review instead of staying suspended.
|
|
972
987
|
|
|
973
|
-
**Sub-tasks covered by commits (L82)**: when the review concludes compliant, the open sub-tasks of the lot that a work commit of the lot cites (`feat(L3/t2):
|
|
988
|
+
**Sub-tasks covered by commits (L82)**: when the review concludes compliant, the open sub-tasks of the lot that a work commit of the lot cites (`feat(L3/t2): …`, in the project or in one of the lot's neighbouring repositories; a commit that cites only the lot covers none) are closed in the same plan commit as the verdict, with the warning `sous-tâches closes : L3/t2 (sha)`, so `raf done` does not refuse the lot over finished work. A read-only plan is never written: each such sub-task is returned as a proposal `[sous-tâche clore] L3/t2 — couverte par <sha>`, for you to close with the project's tool. On a plan in a foreign format (read-only, `cadence.yaml` `plan:`), sub-task ids are read as the plan reads them: `B53/t4-ux1` covers `t4-ux1` only, not `t4`, and a list such as `feat(B53/t4-ux1,t5)` covers each of its sub-tasks. The lot id is not read inside another id (`fix(E-A2/t3,t4, A2/t1)` covers `A2/t1` only for lot `A2`), and a list may repeat the lot id (`fix(L1/t1, L1/t2,t3)` covers `t1`, `t2` and `t3`). A non-compliant review closes nothing.
|
|
974
989
|
|
|
975
990
|
**Choices, not questions**: the author brief tells the session to decide minor interpretation questions itself and to list them under `choix` in its report; the reviewer receives that list to re-read, and the final table prints each one (`choix fait : …`). A session stops with a question only on a real blocker: a decision that changes the scope or the architecture or is costly to undo, AND that the plan, its notes and CLAUDE.md do not settle; everything else is a choice.
|
|
976
991
|
|
|
@@ -1137,7 +1152,7 @@ News instruction (`cadence news new <lot>`, factual user-side text, a screenshot
|
|
|
1137
1152
|
orchestrate:
|
|
1138
1153
|
test: npm test # run by the orchestrator after a work step (optional)
|
|
1139
1154
|
build: npm run build # run after the tests; both results go to the reviewer, who does not redo them (optional)
|
|
1140
|
-
precheck: true # default: before the first implementation of a lot with no commit, a read-only Sonnet session checks whether the deliverable is already in the repository (see below); false skips it
|
|
1155
|
+
precheck: true # default: before the first implementation of a lot with no commit, a read-only Sonnet session checks whether the deliverable is already in the repository (see below); false skips it, local runs it on claude-local (Ollama) with Sonnet as fallback
|
|
1141
1156
|
effort: { review: xhigh } # effort level per pass (see « Effort level per pass »): precheck, implement, fix, review, ux
|
|
1142
1157
|
ux: http://localhost:4200 # a URL, a launch command, or { command, url, timeout? } — for the UX review (see below)
|
|
1143
1158
|
permissionMode: auto # default
|
|
@@ -1280,12 +1295,14 @@ repository with `cadence skills install` (to `.claude/skills/cadence-*` and
|
|
|
1280
1295
|
runs a build whose output is used live, and never edits code. A README or usage
|
|
1281
1296
|
documentation that does not follow the lot's change is a *major* finding.
|
|
1282
1297
|
- **qa-reviewer** (agent): any web app (read-only like `ux-reviewer`: same tool list, no `Edit`, no `Write`); given a repository and a base URL (and
|
|
1283
|
-
optionally a lot id, to
|
|
1284
|
-
|
|
1298
|
+
optionally a lot id, to walk only the pages it touched — those that call the
|
|
1299
|
+
changed endpoints, those of the screens it changed, the pages that consume the services the lot changed,
|
|
1300
|
+
and the home page; a backend-only lot stays in scope, and when the lot changes no route
|
|
1301
|
+
and no screen, or the endpoints cannot be mapped to pages, every page is walked; the others are named « Not walked »), it opens each page of
|
|
1285
1302
|
the project's expectations file in a real browser at 1440 and 390 px and
|
|
1286
1303
|
measures: expected content present and non-empty, no error or missing-data
|
|
1287
1304
|
message, every API call answered 2xx with a non-empty body, no console error,
|
|
1288
|
-
no broken content image. Findings are defects (a line of the expectations
|
|
1305
|
+
no broken content image. It reads the page as text first and takes a capture only for a gap. Findings are defects (a line of the expectations
|
|
1289
1306
|
broken, or a universal check failing with a visible effect, with or without an
|
|
1290
1307
|
expectations file), suspects (it looks like missing or wrong data and no
|
|
1291
1308
|
expectation settles it) or noise (a console error or a failed request with no
|
package/agents/qa-reviewer.md
CHANGED
|
@@ -15,9 +15,7 @@ that a defect. You do: a players page with no players is a defect, whatever the
|
|
|
15
15
|
## Inputs
|
|
16
16
|
|
|
17
17
|
The absolute path of the repository and the base URL of the app — deployed, or a local server the
|
|
18
|
-
caller started. Optionally a lot id: then
|
|
19
|
-
notes in the plan, and `raf commits <id>`, tell which) — when the lot touched only the backend,
|
|
20
|
-
the pages that call the changed endpoints — and walk the others after. If the path or the URL is
|
|
18
|
+
caller started. Optionally a lot id: then walk only the pages that lot touched, as « Scope of a lot » below says. If the path or the URL is
|
|
21
19
|
missing, or the URL does not answer, say so and stop.
|
|
22
20
|
|
|
23
21
|
## Method
|
|
@@ -46,6 +44,11 @@ missing, or the URL does not answer, say so and stop.
|
|
|
46
44
|
and **390 px** wide. Walk within the bounded pass below. Let it settle: after `load`, wait a fixed few seconds, scroll through the
|
|
47
45
|
page (lazy images), wait again — never for network idle, which streams and polling never reach.
|
|
48
46
|
Then measure:
|
|
47
|
+
- read the page as text first: its rendered text and structure (`innerText`, or the
|
|
48
|
+
accessibility snapshot), the counts of the selectors named by the expectations, the responses
|
|
49
|
+
you listened to. Compare that text with the expectations, and take a capture only for a gap —
|
|
50
|
+
a `shows:` line unmet, a `never:` text found, a failed call — as evidence of that gap: a page
|
|
51
|
+
that matches its expectations gets none;
|
|
49
52
|
- browser state, once before the first page, in a profile already used (a persistent context, not a fresh one — a `userDataDir` reserved for QA and kept between passes, never the user's own browser profile): compare the bundle the page loaded (its script URL) with the one `index.html` references, re-read without cache — a returning visitor still holds the old one, so a difference is stated in the report — then clear the cache and measure;
|
|
50
53
|
- the expected content is present and non-empty — name the selector or the text found and its
|
|
51
54
|
count (`.player-card` ×14), not "the list looks fine";
|
|
@@ -96,6 +99,32 @@ missing, or the URL does not answer, say so and stop.
|
|
|
96
99
|
false, or an error is shown to the user), *major* (secondary content missing or wrong, a section
|
|
97
100
|
silently dropped after a failed or empty API call, a broken content image), *minor* (noise). A broken line of the expectations with no visible loss on the page (the content is on screen by another path) is *minor* too.
|
|
98
101
|
|
|
102
|
+
## Scope of a lot
|
|
103
|
+
|
|
104
|
+
With a lot id, walk only the pages the lot touched — the cost of a pass is the pages walked, and
|
|
105
|
+
a delivery rarely changes them all:
|
|
106
|
+
|
|
107
|
+
- the pages that call the changed endpoints (`raf commits <id>` and the diff tell which routes
|
|
108
|
+
changed; the `api:` lines of the expectations, and the calls you see a page make, tell which
|
|
109
|
+
pages use them);
|
|
110
|
+
- the pages of the screens the lot changed (its title and notes in the plan, and
|
|
111
|
+
the components in its commits, name them);
|
|
112
|
+
- the home page.
|
|
113
|
+
|
|
114
|
+
A backend-only lot is kept in scope: it can empty a page without touching a screen, and the pages
|
|
115
|
+
that call its endpoints are exactly the ones to walk. It also touches the pages that consume the
|
|
116
|
+
services the lot changed — grep for their callers, from the changed service up to the endpoint
|
|
117
|
+
that uses it, and walk the pages of those endpoints. If a backend lot changes no route (a service, a
|
|
118
|
+
data source or a configuration changed under routes that keep their path) and no screen, the scope
|
|
119
|
+
cannot come from the routes: the fallback is mandatory: walk every page and say so in the report.
|
|
120
|
+
If you cannot link the diff to any page, or a backend part of it in a lot that also changes
|
|
121
|
+
screens (a shared helper, a utility, a data source), the fallback is mandatory too: walk every page
|
|
122
|
+
and say so in the report. A lot that changes only screens keeps the reduced scope:
|
|
123
|
+
the pages of those screens, plus the home page. If you cannot tell which pages a changed
|
|
124
|
+
endpoint feeds, walk every page and say so — a scope you cannot establish reduces nothing. Pages
|
|
125
|
+
outside the scope are named under « Not walked », never counted as checked. Without a lot id,
|
|
126
|
+
walk every page.
|
|
127
|
+
|
|
99
128
|
## Bounded pass
|
|
100
129
|
|
|
101
130
|
The pass has a time budget: the one the caller names, otherwise 15 minutes. You keep the count
|
|
@@ -117,6 +146,7 @@ from the first page.
|
|
|
117
146
|
|
|
118
147
|
A short report:
|
|
119
148
|
|
|
149
|
+
- **Scope**, one line: the lot id (or "none: every page"), the pages walked of pages in the expectations (or routes discovered), a « Not walked » list naming each page not walked and why, and the captures taken — the figures to compare from one pass to the next.
|
|
120
150
|
- **Pages checked N/N**, with the base URL and the date and time of the run, and the two widths. The
|
|
121
151
|
second N is every page of the expectations (or every route discovered): a page you could not open
|
|
122
152
|
is counted and named, never dropped. A page counts as checked when both widths were measured; a
|
package/dist/audit.js
CHANGED
|
@@ -192,14 +192,15 @@ export function lotWork(plan, root, lotId) {
|
|
|
192
192
|
/**
|
|
193
193
|
* Sous-tâches encore ouvertes d'un lot que citent ses commits de travail (`feat(L3/t2)`), avec le plus récent des commits qui
|
|
194
194
|
* les citent (L82) : le travail est fait, il ne reste que le plan à tenir. Un commit qui ne cite que le lot ne couvre rien.
|
|
195
|
+
* `neighbourWork` : les commits du lot dans ses dépôts voisins (`repoWork`, L62), lus après ceux du projet (L156).
|
|
195
196
|
*/
|
|
196
|
-
export function coveredTasks(plan, root, lotId) {
|
|
197
|
+
export function coveredTasks(plan, root, lotId, neighbourWork = []) {
|
|
197
198
|
const lot = plan.lots().find((l) => l.id === lotId);
|
|
198
199
|
if (!lot)
|
|
199
200
|
return [];
|
|
200
201
|
const open = new Set(lot.tasks.filter((t) => isOpen(t.status)).map((t) => t.id));
|
|
201
202
|
const found = new Map();
|
|
202
|
-
for (const c of lotWork(plan, root, lotId)) {
|
|
203
|
+
for (const c of [...lotWork(plan, root, lotId), ...neighbourWork]) {
|
|
203
204
|
const cited = citedRefs(c, plan.refs).filter((r) => r.lot === lotId && r.task);
|
|
204
205
|
// `fix(L3/t1,t2)` : le motif des références ne lit que la première sous-tâche d'une liste à virgule.
|
|
205
206
|
// Même texte que citedRefs : la portée quand elle cite des lots, sinon le message entier (le corps ne compte pas à côté d'une portée).
|
|
@@ -208,7 +209,7 @@ export function coveredTasks(plan, root, lotId) {
|
|
|
208
209
|
const scope = scopeOf(c.subject);
|
|
209
210
|
const text = scope !== null && plan.refs(scope).length > 0 ? scope : `${c.subject}\n${c.body}`;
|
|
210
211
|
const element = '[\\w-]+(?:\\.[\\w-]+)*';
|
|
211
|
-
const listed = cited.length > 0 ? [...text.matchAll(new RegExp(`(?<![\\w
|
|
212
|
+
const listed = cited.length > 0 ? [...text.matchAll(new RegExp(`(?<![\\w/.-])${escapeRe(lotId)}/(${element}(?:\\s*,\\s*(?:${escapeRe(lotId)}/)?${element})*)`, 'g'))].flatMap((m) => m[1].split(/\s*,\s*/)).map((e) => (e.startsWith(`${lotId}/`) ? e.slice(lotId.length + 1) : e)).flatMap((e) => plan.refs(`${lotId}/${e}`).filter((r) => r.lot === lotId && r.task).map((r) => r.task)) : [];
|
|
212
213
|
for (const task of [...cited.map((r) => r.task), ...listed])
|
|
213
214
|
if (open.has(task) && !found.has(task))
|
|
214
215
|
found.set(task, c.sha);
|
package/dist/config.js
CHANGED
|
@@ -226,7 +226,7 @@ const REVIEW_KEYS = ['threshold', 'light', 'full'];
|
|
|
226
226
|
const ORCH_KEYS = ['start', 'verdict', 'test', 'build', 'ux', 'precheck', 'review', 'effort', 'permissionMode', 'addDirs', 'timeouts'];
|
|
227
227
|
/** Clé `orchestrate:` de cadence.yaml. Absente : les défauts (auto, 45 min d'implémentation, 25 min de revue). */
|
|
228
228
|
export function readOrchestrateConfig(file) {
|
|
229
|
-
const config = { precheck: true, review: { threshold: 0.25, light: 'sonnet', full: 'opus' }, effort: { precheck: 'low', implement: 'medium', fix: 'medium', review: 'high', ux: 'high' }, permissionMode: 'auto', addDirs: [], timeouts: { work: 45 * 60_000, review: 25 * 60_000 } };
|
|
229
|
+
const config = { precheck: true, precheckLocal: false, review: { threshold: 0.25, light: 'sonnet', full: 'opus' }, effort: { precheck: 'low', implement: 'medium', fix: 'medium', review: 'high', ux: 'high' }, permissionMode: 'auto', addDirs: [], timeouts: { work: 45 * 60_000, review: 25 * 60_000 } };
|
|
230
230
|
if (!existsSync(file))
|
|
231
231
|
return config;
|
|
232
232
|
let raw;
|
|
@@ -261,9 +261,10 @@ export function readOrchestrateConfig(file) {
|
|
|
261
261
|
config[k] = o[k].trim();
|
|
262
262
|
}
|
|
263
263
|
if (o.precheck != null) {
|
|
264
|
-
if (typeof o.precheck !== 'boolean')
|
|
265
|
-
throw bad('precheck : true ou
|
|
266
|
-
config.precheck = o.precheck;
|
|
264
|
+
if (typeof o.precheck !== 'boolean' && o.precheck !== 'local')
|
|
265
|
+
throw bad('precheck : true, false ou local attendu');
|
|
266
|
+
config.precheck = o.precheck !== false;
|
|
267
|
+
config.precheckLocal = o.precheck === 'local';
|
|
267
268
|
}
|
|
268
269
|
if (o.review != null) {
|
|
269
270
|
if (!isObject(o.review))
|
|
@@ -15,7 +15,7 @@ import { loadTemplates, newsText, objective, renderBrief } from './briefs.js';
|
|
|
15
15
|
import { candidates, parsePriority, parseUntil, readPriority, stopReason } from './continue.js';
|
|
16
16
|
import { Budget, DROPPED, MAX_PASSES, countInterrupted, handBackDroppedQuestions, needsPrecheck } from './cycle.js';
|
|
17
17
|
import { canInstallPrePush, installPrePush, removePrePush, snapshot } from './guard.js';
|
|
18
|
-
import { buildArgs, killSessions, mcpServersFor, readAgents, realClaude } from './launch.js';
|
|
18
|
+
import { buildArgs, buildLocalArgs, killSessions, mcpServersFor, readAgents, realClaude } from './launch.js';
|
|
19
19
|
import { activeLock, REPO_LOCK, releaseLock, takeLock } from './lock.js';
|
|
20
20
|
import { runPool } from './pool.js';
|
|
21
21
|
import { schemaFor } from './schemas.js';
|
|
@@ -292,7 +292,7 @@ function dryRun(lots, io, deps, budget, id) {
|
|
|
292
292
|
}
|
|
293
293
|
const steps = [{ kind: 'implement', model: l.model }];
|
|
294
294
|
if (needsPrecheck(env.config, env.loadPlan(), l.repo, l.lot, l.repos))
|
|
295
|
-
steps.unshift({ kind: 'precheck', model: 'sonnet' });
|
|
295
|
+
steps.unshift({ kind: 'precheck', model: env.config.precheckLocal ? 'local' : 'sonnet' });
|
|
296
296
|
const { light, full } = env.config.review;
|
|
297
297
|
if (l.small)
|
|
298
298
|
steps.push({ kind: l.visible ? 'review-small' : 'review', model: l.light ? light : full });
|
|
@@ -312,8 +312,9 @@ function dryRun(lots, io, deps, budget, id) {
|
|
|
312
312
|
const brief = renderBrief(s.kind, vars, deps.templatesDir);
|
|
313
313
|
writeFileSync(file, brief);
|
|
314
314
|
const playwright = !!mcpServersFor(s.kind, l.visible, '', l.small).playwright;
|
|
315
|
-
const
|
|
316
|
-
|
|
315
|
+
const local = s.model === 'local';
|
|
316
|
+
const args = (local ? buildLocalArgs : buildArgs)({ kind: s.kind, sessionId: '<uuid>', brief: '<brief>', model: s.model === 'local' ? 'sonnet' : s.model, local, effort: env.config.effort[s.kind === 'review-small' ? 'review' : s.kind], schema: schemaFor(s.kind), agent: s.kind === 'implement' ? undefined : s.kind === 'ux' ? 'ux-reviewer' : s.kind === 'precheck' ? 'precheck-reader' : 'code-reviewer', cwd: l.repo, wave: id, permissionMode: env.config.permissionMode, addDirs: [...env.config.addDirs, ...(playwright && s.kind === 'implement' ? [pwDir] : []), ...neighbourDirs(l)], timeoutMs: 0, mcpConfig: '<mcp>', playwright }, agents).map((a) => (a.startsWith('{') ? '<json>' : a));
|
|
317
|
+
io.out(` ${s.kind} : ${local ? 'claude-local' : 'claude'} ${args.join(' ')}`);
|
|
317
318
|
io.out(` brief : ${file}`);
|
|
318
319
|
const servers = Object.keys(mcpServersFor(s.kind, l.visible, '', l.small));
|
|
319
320
|
io.out(` mcp : ${servers.length ? `${servers.join(', ')} (captures dans le dossier de la vague : <vague>/${lotSlug(l.project, l.lot)}/playwright)` : 'aucun'}`);
|
|
@@ -1020,7 +1021,7 @@ export function realOrchestrateDeps(env) {
|
|
|
1020
1021
|
// Lues ici, une fois : ni les sessions ni les commandes du projet ne reçoivent les variables de relance.
|
|
1021
1022
|
env = withoutLaunchVars({ ...env });
|
|
1022
1023
|
return {
|
|
1023
|
-
claude: realClaude(bin, env),
|
|
1024
|
+
claude: realClaude(bin, env, env.CADENCE_CLAUDE_LOCAL_BIN || 'claude-local'),
|
|
1024
1025
|
claudeInfo: () => {
|
|
1025
1026
|
try {
|
|
1026
1027
|
const version = execFileSync(bin, ['--version'], { env, encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'], timeout: 20_000 }).trim();
|
|
@@ -39,8 +39,8 @@ export class Budget {
|
|
|
39
39
|
* session (`--session-id`). Une étape déjà comptée (`tokens`) ne l'est jamais deux fois. Vrai quand quelque chose a été compté.
|
|
40
40
|
*/
|
|
41
41
|
export function countInterrupted(claudeHome, repo, step, budget) {
|
|
42
|
-
if (!claudeHome || !step.sessionId || step.tokens)
|
|
43
|
-
return false;
|
|
42
|
+
if (!claudeHome || !step.sessionId || step.tokens || step.model === 'local')
|
|
43
|
+
return false; // le modèle local n'a rien pris au quota
|
|
44
44
|
const spent = journalTokens(claudeHome, repo, step.sessionId);
|
|
45
45
|
if (!spent)
|
|
46
46
|
return false;
|
|
@@ -293,7 +293,9 @@ function reviewModel(c, kind) {
|
|
|
293
293
|
return c.lot.light && c.lot.pass === 0 && (kind === 'review' || kind === 'review-small') ? light : full;
|
|
294
294
|
}
|
|
295
295
|
/** Une session : budget et quota vérifiés avant, état écrit avant et après, journal gardé, contrôles du dépôt après. */
|
|
296
|
-
|
|
296
|
+
/** Rendu par `session(…, true)` quand l'étape locale n'a rien donné (Ollama éteint, délai, rapport sans structure) : l'appelant se replie sur Sonnet (L146). */
|
|
297
|
+
const LOCAL_FAILED = Symbol('local-failed');
|
|
298
|
+
async function session(c, kind, local = false) {
|
|
297
299
|
const w = c.wave;
|
|
298
300
|
const l = c.lot;
|
|
299
301
|
if (dropped(c))
|
|
@@ -305,7 +307,7 @@ async function session(c, kind) {
|
|
|
305
307
|
// Le budget du lot borne l'écriture, jamais la revue (L145) : une revue est jouée dès que le code est à relire (sauf tests rouges après une écriture, cf. REVIEW_RESERVE), la passe fix réserve son coût.
|
|
306
308
|
if (write && (lotOver(l) || (kind === 'fix' && !fixAffordable(l))))
|
|
307
309
|
return overBudget(c, kind === 'fix' && !lotOver(l));
|
|
308
|
-
const model = write ? l.model : kind === 'precheck' ? 'sonnet' : reviewModel(c, kind); // le contrôle préalable ne fait que lire : pas d'Opus
|
|
310
|
+
const model = write ? l.model : local ? 'local' : kind === 'precheck' ? 'sonnet' : reviewModel(c, kind); // le contrôle préalable ne fait que lire : pas d'Opus
|
|
309
311
|
const effort = c.config.effort[kind === 'review-small' ? 'review' : kind];
|
|
310
312
|
const before = await snapshot(l.repo);
|
|
311
313
|
const neighbours = l.repos ?? [];
|
|
@@ -335,8 +337,9 @@ async function session(c, kind) {
|
|
|
335
337
|
kind,
|
|
336
338
|
sessionId,
|
|
337
339
|
brief,
|
|
338
|
-
model,
|
|
340
|
+
model: local ? 'sonnet' : model,
|
|
339
341
|
effort,
|
|
342
|
+
local,
|
|
340
343
|
schema: schemaFor(kind),
|
|
341
344
|
agent: kind === 'ux' ? 'ux-reviewer' : kind === 'precheck' ? 'precheck-reader' : write ? undefined : 'code-reviewer',
|
|
342
345
|
cwd: l.repo,
|
|
@@ -359,7 +362,7 @@ async function session(c, kind) {
|
|
|
359
362
|
step.ended = new Date().toISOString();
|
|
360
363
|
const dir = w.store.lotDir(l.project, l.lot);
|
|
361
364
|
const base = `${n}-${kind}`;
|
|
362
|
-
if (outcome.kind !== 'ok') {
|
|
365
|
+
if (outcome.kind !== 'ok' && !local) {
|
|
363
366
|
// Une session en échec, au quota ou tuée au délai a consommé des tokens : ceux de sa sortie, sinon ceux de son journal.
|
|
364
367
|
const spent = outcome.tokens ? { tokens: outcome.tokens, sessionId: outcome.sessionId } : w.claudeHome ? journalTokens(w.claudeHome, l.repo, outcome.sessionId ?? sessionId) : null;
|
|
365
368
|
if (spent) {
|
|
@@ -372,6 +375,20 @@ async function session(c, kind) {
|
|
|
372
375
|
step.peakContext = peakContext(w.claudeHome, l.repo, spent.sessionId);
|
|
373
376
|
}
|
|
374
377
|
}
|
|
378
|
+
if (local && outcome.kind !== 'ok') {
|
|
379
|
+
// Hors quota Anthropic : rien au budget (ni sortie, ni journal), et jamais d'arrêt du lot — Sonnet reprend le contrôle.
|
|
380
|
+
const cause = outcome.kind === 'failed' ? outcome.cause : outcome.message;
|
|
381
|
+
writeFileSync(join(dir, `${base}.json`), outcome.kind === 'failed' ? outcome.stdout : '');
|
|
382
|
+
if (outcome.kind === 'failed')
|
|
383
|
+
writeFileSync(join(dir, `${base}.err`), outcome.stderr);
|
|
384
|
+
step.status = 'failed';
|
|
385
|
+
step.cause = cause;
|
|
386
|
+
step.report = `${base}.json`;
|
|
387
|
+
delete step.tokens;
|
|
388
|
+
l.warnings.push(`contrôle préalable : modèle local sans résultat (${cause.slice(0, 200)}), repli sur Sonnet`);
|
|
389
|
+
save(c);
|
|
390
|
+
return LOCAL_FAILED;
|
|
391
|
+
}
|
|
375
392
|
if (outcome.kind === 'failed') {
|
|
376
393
|
if (outcome.firstStdout !== undefined)
|
|
377
394
|
writeFileSync(join(dir, `${base}.first.json`), outcome.firstStdout);
|
|
@@ -399,12 +416,13 @@ async function session(c, kind) {
|
|
|
399
416
|
writeFileSync(join(dir, `${base}.json`), JSON.stringify({ ...res, structured: res.structured }, null, 2));
|
|
400
417
|
step.report = `${base}.json`;
|
|
401
418
|
step.sessionId = res.sessionId;
|
|
402
|
-
step.tokens = res.tokens;
|
|
419
|
+
step.tokens = local ? { ...res.tokens, counted: 0 } : res.tokens; // local : hors quota, hors budget
|
|
403
420
|
if (res.formattingRetry) {
|
|
404
421
|
step.formatRetry = true;
|
|
405
422
|
w.log(`${lotKey(l.project, l.lot)} · session ${n} ${kind} : rapport sans sortie structurée, relance de mise en forme`);
|
|
406
423
|
}
|
|
407
|
-
|
|
424
|
+
if (!local)
|
|
425
|
+
w.budget.add(res.tokens);
|
|
408
426
|
w.saveWave();
|
|
409
427
|
if (w.claudeHome)
|
|
410
428
|
step.peakContext = peakContext(w.claudeHome, l.repo, res.sessionId);
|
|
@@ -523,18 +541,44 @@ export function needsPrecheck(config, plan, repo, lot, repos = []) {
|
|
|
523
541
|
*/
|
|
524
542
|
async function precheck(c) {
|
|
525
543
|
const l = c.lot;
|
|
526
|
-
|
|
527
|
-
|
|
528
|
-
|
|
529
|
-
|
|
530
|
-
|
|
531
|
-
|
|
544
|
+
let rep = null;
|
|
545
|
+
let localYes = false;
|
|
546
|
+
if (c.config.precheckLocal) {
|
|
547
|
+
// Modèle local (L146) : un échec, un rapport illisible ou un « oui » (qui rendrait le lot sans implémentation) rend la main à Sonnet.
|
|
548
|
+
const done = await session(c, 'precheck', true);
|
|
549
|
+
if (!done)
|
|
550
|
+
return;
|
|
551
|
+
if (done !== LOCAL_FAILED) {
|
|
552
|
+
try {
|
|
553
|
+
rep = checkShape(done.report, PRECHECK_SCHEMA);
|
|
554
|
+
}
|
|
555
|
+
catch (e) {
|
|
556
|
+
l.warnings.push(`contrôle préalable : rapport du modèle local illisible (${e.message}), repli sur Sonnet`);
|
|
557
|
+
}
|
|
558
|
+
if (rep?.dejaPresent === 'oui') {
|
|
559
|
+
localYes = true;
|
|
560
|
+
rep = null;
|
|
561
|
+
}
|
|
562
|
+
}
|
|
532
563
|
}
|
|
533
|
-
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
|
|
564
|
+
if (!rep) {
|
|
565
|
+
const done = await session(c, 'precheck');
|
|
566
|
+
if (!done)
|
|
567
|
+
return;
|
|
568
|
+
try {
|
|
569
|
+
rep = checkShape(done.report, PRECHECK_SCHEMA);
|
|
570
|
+
}
|
|
571
|
+
catch (e) {
|
|
572
|
+
l.warnings.push(`contrôle préalable illisible, implémentation lancée : ${e.message}`);
|
|
573
|
+
l.next = 'implement';
|
|
574
|
+
save(c);
|
|
575
|
+
return;
|
|
576
|
+
}
|
|
577
|
+
}
|
|
578
|
+
if (localYes) {
|
|
579
|
+
l.warnings.push(rep.dejaPresent === 'oui'
|
|
580
|
+
? 'contrôle préalable : « oui » du modèle local, confirmé par Sonnet'
|
|
581
|
+
: `contrôle préalable : le modèle local a dit « oui », Sonnet dit « ${rep.dejaPresent} » — le local s'est trompé, repasser orchestrate.precheck à true`);
|
|
538
582
|
}
|
|
539
583
|
const proofs = rep.preuves.length ? ` (${rep.preuves.join(' ; ')})` : '';
|
|
540
584
|
// « oui » sans preuve ne se vérifie pas : il vaut « partiel », l'implémentation part avec le constat.
|
|
@@ -857,7 +901,7 @@ async function conclude(c, code, minorNote = '') {
|
|
|
857
901
|
}
|
|
858
902
|
l.constats = [];
|
|
859
903
|
// Sous-tâches ouvertes que des commits du lot citent (L82) : la revue conforme vaut pour elles, sinon `raf done` refuserait le lot.
|
|
860
|
-
const covered = coveredTasks(plan, l.repo, l.lot);
|
|
904
|
+
const covered = coveredTasks(plan, l.repo, l.lot, neighbours.flatMap((r) => repoWork(plan, r, l.lot)));
|
|
861
905
|
if (!plan.readonly) {
|
|
862
906
|
plan.recordReview(l.lot, verdict, c.wave.today, newer, neighbours.length ? shas : undefined);
|
|
863
907
|
for (const k of covered)
|
|
@@ -60,6 +60,24 @@ export function buildArgs(spec, agents) {
|
|
|
60
60
|
args.push('--strict-mcp-config', '--mcp-config', spec.mcpConfig);
|
|
61
61
|
return args;
|
|
62
62
|
}
|
|
63
|
+
/**
|
|
64
|
+
* Arguments de `claude-local` (L146) : le prompt vient en premier, le script ajoute lui-même `--bare`, `--model` et le
|
|
65
|
+
* serveur Ollama. Le harnais `--bare` ne charge pas `--agents` : le prompt de l'agent est joint au brief et ses outils
|
|
66
|
+
* passent par `--tools`. Pas de `--effort`, pas de MCP (aucune étape locale n'en charge).
|
|
67
|
+
*/
|
|
68
|
+
export function buildLocalArgs(spec, agents) {
|
|
69
|
+
const a = spec.agent ? agents[spec.agent] : undefined;
|
|
70
|
+
if (spec.agent && !a)
|
|
71
|
+
throw new RafError(`agent introuvable dans le paquet : ${spec.agent}`);
|
|
72
|
+
const args = [a ? `${a.prompt}\n\n${spec.brief}` : spec.brief, '--output-format', 'json', '--json-schema', JSON.stringify(spec.schema)];
|
|
73
|
+
if (a?.tools)
|
|
74
|
+
args.push('--tools', a.tools.join(','));
|
|
75
|
+
args.push('--session-id', spec.sessionId, '--permission-mode', spec.permissionMode);
|
|
76
|
+
for (const d of spec.addDirs)
|
|
77
|
+
args.push('--add-dir', d);
|
|
78
|
+
args.push('--disallowedTools', ...DISALLOWED);
|
|
79
|
+
return args;
|
|
80
|
+
}
|
|
63
81
|
/** Définitions d'agents du paquet (`agents/*.md`) au format de --agents. */
|
|
64
82
|
export function readAgents(dir) {
|
|
65
83
|
const out = {};
|
|
@@ -107,9 +125,11 @@ export function buildRetryArgs(spec, sessionId, agents) {
|
|
|
107
125
|
* s'ajoutent à ceux de la session ; toujours rien après elle, c'est l'échec habituel.
|
|
108
126
|
*/
|
|
109
127
|
export async function runSession(spec, deps) {
|
|
110
|
-
const launch = (args) => deps.claude(args, { cwd: spec.cwd, env: { CADENCE_ORCHESTRATED: spec.wave, ...(spec.nodeBin || spec.toolBin ? { PATH: [spec.toolBin, spec.nodeBin, process.env.PATH ?? ''].filter(Boolean).join(delimiter) } : {}) }, timeoutMs: spec.timeoutMs, onSpawn: deps.onSpawn });
|
|
111
|
-
const out = await launch(buildArgs(spec, deps.agents));
|
|
128
|
+
const launch = (args) => deps.claude(args, { local: spec.local, cwd: spec.cwd, env: { CADENCE_ORCHESTRATED: spec.wave, ...(spec.nodeBin || spec.toolBin ? { PATH: [spec.toolBin, spec.nodeBin, process.env.PATH ?? ''].filter(Boolean).join(delimiter) } : {}) }, timeoutMs: spec.timeoutMs, onSpawn: deps.onSpawn });
|
|
129
|
+
const out = await launch(spec.local ? buildLocalArgs(spec, deps.agents) : buildArgs(spec, deps.agents));
|
|
112
130
|
const first = classify(out);
|
|
131
|
+
if (spec.local)
|
|
132
|
+
return first; // pas de relance de mise en forme : l'appelant se replie sur Sonnet
|
|
113
133
|
const missing = first.kind === 'failed' && !out.timedOut && out.code === 0 ? lacksStructuredOutput(out.stdout) : null;
|
|
114
134
|
if (!missing)
|
|
115
135
|
return first;
|
|
@@ -175,13 +195,13 @@ export function killSessions() {
|
|
|
175
195
|
marks.clear();
|
|
176
196
|
}
|
|
177
197
|
/** Vrai lanceur : `claude` (ou CADENCE_CLAUDE_BIN) dans son propre groupe de processus, tué en bloc au délai. */
|
|
178
|
-
export function realClaude(bin, base = process.env) {
|
|
198
|
+
export function realClaude(bin, base = process.env, localBin = 'claude-local') {
|
|
179
199
|
return (args, opts) => new Promise((resolve) => {
|
|
180
200
|
// Marque héritée par tous les descendants, même orphelins rattachés à init (un shell qui sort aussitôt après
|
|
181
201
|
// `nohup srv &` échappe à tout relevé de l'arbre) : à la fin de la session, tout ce qui la porte est tué.
|
|
182
202
|
const mark = randomUUID();
|
|
183
203
|
marks.add(mark);
|
|
184
|
-
const child = spawn(bin, args, { cwd: opts.cwd, env: { ...base, ...opts.env, [SESSION_MARK_VAR]: mark }, stdio: ['ignore', 'pipe', 'pipe'], detached: true });
|
|
204
|
+
const child = spawn(opts.local ? localBin : bin, args, { cwd: opts.cwd, env: { ...base, ...opts.env, [SESSION_MARK_VAR]: mark }, stdio: ['ignore', 'pipe', 'pipe'], detached: true });
|
|
185
205
|
const pid = child.pid;
|
|
186
206
|
if (pid === undefined) {
|
|
187
207
|
marks.delete(mark);
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@sylad/cadence",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.28.0",
|
|
4
4
|
"description": "A small, repo-native working method: a versioned plan linked to your commits, a changelog with screenshots, session rituals and deliveries proven by their effect.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": "Sylvain Ladoire",
|
package/skills/deliver/SKILL.md
CHANGED
|
@@ -57,8 +57,7 @@ Deliver one project at a time: the one the human names, or ask.
|
|
|
57
57
|
6. After a green delivery that changes what a page shows or what it is served (screen, API, data
|
|
58
58
|
source, configuration of either) — in practice every delivery except docs-, plan- or tests-only
|
|
59
59
|
ones — have the `qa-reviewer` agent walk the delivered app in a real browser, whether the lot
|
|
60
|
-
is `visible` or not: give it the repository path, the base URL and the lot id.
|
|
61
|
-
touched only the backend, the agent starts with the pages that call the changed endpoints. It
|
|
60
|
+
is `visible` or not: give it the repository path, the base URL and the lot id. With a lot id, the agent walks only the pages the lot touched (the pages that call the changed endpoints, the pages of the changed screens, the pages that consume the services the lot changed, the home page; when the diff changes no route and no screen, or cannot be linked to pages, it walks every page and says so) and names the others under « Not walked »: they are not checked. It
|
|
62
61
|
checks each page against `docs/qa/expectations.md` — what the user must find there — and
|
|
63
62
|
reports a page left empty, an error shown, an API call that failed or came back empty: what
|
|
64
63
|
the checks of `cadence.yaml` do not see. Bring its blocking findings to the human. It is not a
|
package/skills/lead/SKILL.md
CHANGED
|
@@ -169,8 +169,7 @@ After a green delivery that changes what a page shows or what it is served (scre
|
|
|
169
169
|
source, configuration of either) — in practice every delivery except docs-, plan- or tests-only ones
|
|
170
170
|
— have the `qa-reviewer` agent check the delivered app, as a fresh subagent: give it the absolute
|
|
171
171
|
path of the project, the base URL of the delivered app and the lot id. The lot need not be
|
|
172
|
-
`visible`: a backend-only lot can empty a page without changing a screen.
|
|
173
|
-
the backend, the agent starts with the pages that call the changed endpoints. It walks the pages in
|
|
172
|
+
`visible`: a backend-only lot can empty a page without changing a screen. With a lot id, the agent walks only the pages the lot touched (the pages that call the changed endpoints, the pages of the changed screens, the pages that consume the services the lot changed, the home page; when the diff changes no route and no screen, or cannot be linked to pages, it walks every page and says so) and names the others under « Not walked »: they are not checked. It walks the pages in
|
|
174
173
|
a real browser against the project's expectations (`docs/qa/expectations.md`: per page, what the
|
|
175
174
|
user must find there) and returns measured findings; it reads only, and never logs in. Bring its
|
|
176
175
|
blocking findings back to the human — a page whose main content is missing, or that shows an error,
|