@sylad/cadence 0.2.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,7 +6,7 @@
6
6
  {
7
7
  "name": "cadence",
8
8
  "description": "Session start and close rituals driven by a versioned plan (raf), and deliveries proven by their effect. Needs the cadence CLI (npm i -g @sylad/cadence).",
9
- "version": "0.2.0",
9
+ "version": "0.6.0",
10
10
  "source": "./",
11
11
  "author": { "name": "Sylvain Ladoire" }
12
12
  }
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "cadence",
3
- "description": "A repo-native working method: session start and close rituals driven by a versioned plan (raf), deliveries proven by their effect, and a UX reviewer agent.",
4
- "version": "0.2.0",
3
+ "description": "A repo-native working method: session start and close rituals driven by a versioned plan (raf), deliveries proven by their effect, and three reviewer agents (UX, code, QA).",
4
+ "version": "0.6.0",
5
5
  "author": { "name": "Sylvain Ladoire" },
6
6
  "homepage": "https://github.com/Sylad/cadence",
7
7
  "repository": "https://github.com/Sylad/cadence",
package/README.md CHANGED
@@ -11,9 +11,13 @@ Four tools:
11
11
  - **session**: the facts to start and to close a work session;
12
12
  - **deliver**: wait for the CI of the pushed commit, deploy, then **verify the effect**.
13
13
 
14
- And three [Claude Code](https://claude.com/claude-code) skills that turn them
15
- into rituals — `session-start`, `session-close`, `deliver` — plus a `ux-reviewer`
16
- agent: no user-facing change is done before its usability review.
14
+ And four [Claude Code](https://claude.com/claude-code) skills that turn them
15
+ into rituals — `session-start`, `session-close`, `deliver`, and `lead` to pilot
16
+ several projects through subagents — plus three reviewer agents: `ux-reviewer`
17
+ (no user-facing change is done before its usability review) and `code-reviewer`
18
+ (no lot with commits is done before its code review), each behind an opt-in
19
+ gate, and `qa-reviewer`, which walks the delivered app in a real browser and
20
+ reports a page left empty or in error.
17
21
 
18
22
  ## raf
19
23
 
@@ -49,6 +53,7 @@ raf gantt # docs/plan/gantt.html
49
53
  | `raf add "title" [--estimate d] [--quickwin] [--visible] [--after L2,L4] [--parent L3]` | add a lot or a sub-task, print its id |
50
54
  | `raf start <id>` · `raf done <id> [--force]` · `raf drop <id> [--reason text]` | dated transitions (`done` refuses open sub-tasks unless `--force`) |
51
55
  | `raf note <id> "text"` | dated note — keep decisions next to the work |
56
+ | `raf commits <id>` | the commits counted for a lot (the set the code review gate uses), one `<sha> <subject>` per line, oldest first |
52
57
  | `raf now` | what to do next |
53
58
  | `raf list [--status s]` | flat list |
54
59
  | `raf check [--since date] [--idle 7]` | since the plan's adoption date by default: commits without a lot (commits touching only the plan are exempt), unknown ids, `todo` lots that already have commits, idle lots, `done` lots with open sub-tasks, bad or circular dependencies |
@@ -65,6 +70,7 @@ version: 1
65
70
  project: my-app
66
71
  prefix: L
67
72
  since: 2026-09-28 # commits before this date are not audited
73
+ ignore: ['^chore\(batch\):'] # optional: subjects of automated commits, neither audited nor counted for a lot
68
74
  lots:
69
75
  - id: L1
70
76
  title: Monthly dedup on merge
@@ -81,6 +87,57 @@ lots:
81
87
  - { id: t1, title: write the migration, status: done }
82
88
  ```
83
89
 
90
+ ### A plan elsewhere, or in another format
91
+
92
+ `cadence.yaml`, at the repository root, can say where the plan is. The short form keeps raf's own
93
+ format, and the plan stays writable:
94
+
95
+ ```yaml
96
+ plan: planning/todo.yaml
97
+ ```
98
+
99
+ A project that already keeps its plan with its own tool is read **without migrating it**: describe
100
+ the file, and `raf now`, `raf list`, `raf commits`, `raf check`, `raf gantt` and
101
+ `cadence session start|close` work on it. Such a plan is **read-only** —
102
+ `raf add|start|done|note|ux|review` refuse and leave the file to the project's tool, and the two
103
+ review gates do not apply to it (a `uxSince` or `reviewSince` written in it is ignored).
104
+
105
+ ```yaml
106
+ plan:
107
+ path: docs/plan/taches.yaml
108
+ project: my-app # what the file does not say itself: project, since, ignore
109
+ since: 2026-10-02
110
+ ignore: ['^plan: ']
111
+ files: [docs/plan/journal.ndjson] # kept with the plan: a commit touching only these is a plan commit
112
+ lots: taches # root key holding the list (default: lots)
113
+ fields: # raf field: key in the file (a list = first one present)
114
+ title: titre
115
+ status: etat
116
+ estimate: effort
117
+ created: cree_le
118
+ started: demarre_le
119
+ finished: [livre_le, ferme_le]
120
+ notes: note
121
+ parent: parent
122
+ statuses: # raf status: their states
123
+ todo: [prevu, specifie]
124
+ doing: [en_cours, teste]
125
+ done: [deploye, valide]
126
+ dropped: caduc
127
+ estimates: { S: 0.5, M: 1, L: 3 } # their effort labels, in working days
128
+ ```
129
+
130
+ - Fields: `id`, `title`, `status`, `estimate`, `quickwin`, `visible`, `after`, `created`, `started`,
131
+ `finished`, `notes`, `parent`; one left out is read under its own name. A timestamp counts for
132
+ its day; a note written as plain text is one note.
133
+ - `parent`: an entry `B33/t1-fusion` whose parent is `B33` becomes the sub-task `t1-fusion` of `B33`.
134
+ - Ids need no prefix: a commit belongs to a lot when its message cites one of the plan's ids as a
135
+ whole word (`E-A2`, `NC2.4`, `B33/t1-fusion`).
136
+ An id that is not in the plan cannot be told from ordinary text: such a commit counts as
137
+ "without a lot", never as an unknown reference.
138
+ - A state missing from `statuses`, or an effort label missing from `estimates`, is reported by
139
+ `raf check`.
140
+
84
141
  ### Gantt scheduling
85
142
 
86
143
  One lane of work. Finished lots use their real dates (or their commits' dates);
@@ -112,6 +169,7 @@ cadence news build -o frontend/public/nouveautes
112
169
  ---
113
170
  title: Amounts like 3.000 read as three thousand
114
171
  date: 2026-09-29
172
+ created: 2026-09-29T14:32+02:00
115
173
  lots: [L8]
116
174
  captures: [captures/l8.png]
117
175
  # nocapture: reason, when a screenshot makes no sense
@@ -121,9 +179,10 @@ Imported statements now read **3.000** as three thousand, not three.
121
179
 
122
180
  | Command | Effect |
123
181
  |---|---|
124
- | `cadence news new <lot…> [--title t]` | entry skeleton, dated today, titled after the lot |
125
- | `cadence news list` | entries, newest first |
126
- | `cadence news check` | visible lots done without entry, unknown lots, missing or undeclared screenshots, bad headers |
182
+ | `cadence news new <lot…> [--title t]` | entry skeleton, dated and timed now (`date`, `created`), titled after the lot |
183
+ | `cadence news list` | entries, newest first (see *Order* below) |
184
+ | `cadence news check` | visible lots done without entry, unknown lots, missing or undeclared screenshots, entries without creation time, bad headers |
185
+ | `cadence news stamp` | migration: writes `created:` into entries without one (or with an empty one), from the author date of the commit that added the file under its current name (now if not committed yet) |
127
186
  | `cadence news build [-o dir]` | `nouveautes.json` + `index.html` + screenshots (default `docs/nouveautes/site`) |
128
187
 
129
188
  `--dir` changes the entries folder (default `docs/nouveautes` at the git root).
@@ -132,6 +191,21 @@ The Markdown is deliberately small: paragraphs, `-` lists, `**bold**`,
132
191
  `{ project, generated, entries: [{ slug, title, date, lots, captures, html }] }`,
133
192
  with screenshot paths relative to the JSON file.
134
193
 
194
+ **Order.** Everywhere (`list`, `build`, the JSON), entries are strictly newest
195
+ first: by `date`, then, on the same day, by creation time. Every entry carries
196
+ it to the minute in its `created` header, which `cadence news new` writes as
197
+ local time with an explicit offset (`2026-09-29T14:32+02:00`, or `Z`). The
198
+ offset is required, so the order does not depend on the machine's time zone:
199
+ `news check` (and `raf check`) flags a value without it, an impossible one, an
200
+ empty `created:`, and an entry without `created`. `cadence news stamp` fills
201
+ it in older entries (or replaces an empty one) from the author date of the
202
+ commit that added the file under its current name; until then, that date is
203
+ used for sorting (an entry not committed yet counts as the newest). Renames
204
+ are not followed: a renamed entry counts as added on the day of the rename,
205
+ so stamp it before renaming it. The file name only breaks the remaining ties,
206
+ since it follows the title, not the chronology. The page and the JSON still
207
+ show the day only.
208
+
135
209
  ### UX review
136
210
 
137
211
  ```sh
@@ -142,7 +216,99 @@ raf ux L8 "no screen: calculation fix"
142
216
 
143
217
  With the rule on, `raf done` refuses a visible lot without a review (`--force`
144
218
  to override) and `raf check` reports visible lots finished after the `uxSince`
145
- day without one. Plans without `uxSince` are not affected.
219
+ day without one. An empty verdict is refused. Plans without `uxSince` are not
220
+ affected.
221
+
222
+ ### Code review
223
+
224
+ ```sh
225
+ raf review enable # from today, a lot with commits needs a code review before done
226
+ raf commits L4 # what there is to review: the commits the gate counts for the lot
227
+ raf review L4 "compliant after 2 fixes" # record the verdict (from the code-reviewer agent)
228
+ ```
229
+
230
+ The counterpart of the UX review, off by default. With the rule on, `raf done`
231
+ refuses a lot that has at least one commit citing it and no recorded verdict
232
+ (`--force` to override), and `raf check` reports such lots finished after the
233
+ `reviewSince` day. A lot with no commit has nothing to review; neither does a
234
+ lot whose only commits touch the plan itself, predate the plan's `since` or
235
+ match an `ignore:` pattern — `raf commits <id>` prints exactly the counted set.
236
+
237
+ The verdict is tied to what was reviewed: `raf review` stores it on the lot with
238
+ the sha of the lot's latest counted commit (`review: { date, verdict, commit }`,
239
+ `commit: null` when the lot had none). A counted commit made after that one
240
+ makes the review stale: `raf done` refuses (`--force` to override), and
241
+ `raf check` reports a finished lot, until the lot is reviewed again and
242
+ `raf review` is rerun. A verdict
243
+ written by hand without a `commit` field is not checked for staleness. An empty
244
+ verdict is refused. Plans without `reviewSince` are not affected.
245
+
246
+ ### QA review
247
+
248
+ No gate and no command here: the QA review comes **after** a delivery, and
249
+ `raf done` does not wait for it. It follows any delivery that changes what a
250
+ page shows or what it is served (screen, API, data source, configuration of
251
+ either) — in practice every delivery except docs-, plan- or tests-only ones: a
252
+ backend-only lot can empty a page without touching a screen, and the agent then
253
+ starts with the pages that call the changed endpoints. The `qa-reviewer` agent
254
+ opens each page of the running app in a real browser and judges it from the
255
+ user's side. A page can be empty while everything else is green — no code
256
+ changed, a data source went down upstream, the unit tests replace the network,
257
+ `/api/health` answers ok, and the "nothing found" on screen is the message the
258
+ code was written to show.
259
+
260
+ The agent cannot tell such an empty state from a normal one by itself: the
261
+ project says what each page must show, in `docs/qa/expectations.md` — one
262
+ `## <route>` section per page, three kinds of lines:
263
+
264
+ ```markdown
265
+ # QA expectations
266
+
267
+ ## *
268
+ - shows: the header and the navigation links
269
+ - never: "Loading failed", "Too Many Requests"
270
+ - api: /api/live/current — may be empty when no match is within 24 hours
271
+
272
+ ## /players
273
+ - shows: the squad of the last match — at least 11 players
274
+ - shows: the season statistics table, 8 columns — 1440 only
275
+ - never: "No recent line-up found"
276
+ - api: /api/squad/last — a non-empty list
277
+
278
+ ## /fixtures/:id (the first match linked from /fixtures)
279
+ - shows: both team names, the date, the score once the match is played
280
+ - never: "Unknown match"
281
+ - api: /api/fixtures/:id
282
+ ```
283
+
284
+ - `shows:` — content that must be present and non-empty, with a count where one exists;
285
+ - `never:` — texts that must not appear: error messages, and empty-state messages that mean
286
+ missing data;
287
+ - `api:` — the calls the page depends on: each must answer 2xx with a non-empty body (a 200 with
288
+ `[]`, `{}` or `null` is a failure, unless its line says `may be empty when …`).
289
+
290
+ An optional `## *` section holds what every page must show, never show and call. A line may end
291
+ with a condition in plain words, which the agent honours: `may be empty when …`, `1440 only`,
292
+ `390 only` (a line without a width holds at both). Content hidden on the phone by design is not a
293
+ defect unless a `shows:` line requires it at 390; content pushed outside the visible area (it
294
+ needs a sideways scroll) is reported as suspect and handed to `ux-reviewer` in one line.
295
+
296
+ The rest is free text, written for a reader: a line can be repeated, and a route with a parameter
297
+ names a real value to visit or says where to find one. The file can live elsewhere:
298
+
299
+ ```yaml
300
+ # cadence.yaml
301
+ qa:
302
+ expectations: docs/quality/pages.md
303
+ ```
304
+
305
+ Only the agent reads that key; the CLI does not use it. Without an expectations file the agent walks
306
+ the routes it discovers and still runs its universal checks: an error shown, a failed API call
307
+ whose content is missing on screen, a broken or missing content image are defects with or without a
308
+ file; an empty 2xx body, like whatever else would need an expectation to judge, is suspect at most
309
+ (it may be a normal absence). For a route with a
310
+ parameter, it finds a real value in the app's links or its API responses and says how it built the
311
+ URL. The agent then returns a draft for you to correct — it never writes the file itself.
146
312
 
147
313
  ## session
148
314
 
@@ -161,6 +327,20 @@ done), quick wins first. Local state lives in the git directory, never committed
161
327
  the close notes per worktree, the delivery lock and log in `.git/cadence/`, shared
162
328
  by all the worktrees of a clone.
163
329
 
330
+ A project that already has its own morning and evening scripts keeps them: name
331
+ them in `cadence.yaml` and their output is added to the report, under "Faits
332
+ propres au projet", before the proposals (start) or the verdict (close).
333
+
334
+ ```yaml
335
+ session:
336
+ start: ./scripts/morning.sh "$CADENCE_SINCE" # sh, at the repo root, 120 s at most
337
+ close: ./scripts/evening.sh "$CADENCE_SINCE"
338
+ ```
339
+
340
+ The commands get `CADENCE_SINCE` (the `--since` in effect) and `CADENCE_TODAY`. They
341
+ add facts and decide nothing: a failing command is reported and changes neither
342
+ the exit code nor the verdict.
343
+
164
344
  ## deliver
165
345
 
166
346
  A delivery is done when its checks pass, not when a tool says "success".
@@ -202,6 +382,37 @@ cadence deliver # 0 delivered and verified · 1 a step failed · 2
202
382
  - On success the lots cited by the commits since the previous delivery are
203
383
  listed, so you can `raf done` those whose effect you have seen.
204
384
 
385
+ ### A project with its own delivery script
386
+
387
+ A project that already delivers with its own script (CI wait, deploy, business
388
+ checks) plugs it in instead of rewriting it as `ci` / `deploy` / `verify`:
389
+
390
+ ```yaml
391
+ deliver:
392
+ script: ./scripts/ship.sh "$CADENCE_SHORT" # replaces ci and deploy
393
+ allowDirty: true # optional: a modified tree is reported, not refused
394
+ deployTimeout: 3600 # seconds, for the whole script
395
+ verify: [] # optional here: the script's own checks count
396
+ ```
397
+
398
+ ```sh
399
+ cadence deliver -- api frontend --news docs/changelog/x.md -- map # everything after the first « -- » goes to the script
400
+ cadence deliver --dry-run -- api # shows the full command, runs nothing
401
+ cadence deliver --sha 6b0d9aa -- api # an earlier pushed commit instead of HEAD
402
+ ```
403
+
404
+ `--sha` (any mode) delivers a pushed commit other than `HEAD` — for a CI that
405
+ builds each service only on the commit that touched it. `allowDirty` suits a
406
+ working tree shared by several sessions when the delivery starts from a pushed
407
+ sha and never from local files; without it a modified tree is refused.
408
+
409
+ cadence keeps what the script usually lacks: the preconditions (clean tree, pushed
410
+ `HEAD`), the lock (never two deliveries at once), the delivery log and the lots
411
+ delivered. Arguments are quoted for `sh`, so spaces and quotes reach the script
412
+ intact. If the script commits and pushes during the delivery (stamping a
413
+ changelog entry, say), the new `HEAD` is the sha recorded as delivered. Exit code
414
+ 0 of the script means delivered; `verify` checks, if any, run after it.
415
+
205
416
  ## Claude Code skills
206
417
 
207
418
  As a plugin:
@@ -211,10 +422,10 @@ As a plugin:
211
422
  /plugin install cadence@cadence
212
423
  ```
213
424
 
214
- gives `/cadence:session-start`, `/cadence:session-close`, `/cadence:deliver` and
215
- the `ux-reviewer` agent. Or copy them into the repository with
216
- `cadence skills install` (to `.claude/skills/cadence-*` and
217
- `.claude/agents/cadence-ux-reviewer.md`; `--dir` for another `.claude` folder,
425
+ gives `/cadence:session-start`, `/cadence:session-close`, `/cadence:deliver`,
426
+ `/cadence:lead` and the `ux-reviewer`, `code-reviewer` and `qa-reviewer` agents. Or copy them into the
427
+ repository with `cadence skills install` (to `.claude/skills/cadence-*` and
428
+ `.claude/agents/cadence-*.md`; `--dir` for another `.claude` folder,
218
429
  `--force` to overwrite local edits).
219
430
 
220
431
  - **session-start**: reports the facts briefly, proposes three lots from the
@@ -222,12 +433,62 @@ the `ux-reviewer` agent. Or copy them into the repository with
222
433
  - **session-close**: plan hygiene, clean repository, memory limited to what the
223
434
  repository does not say, new skills or agents proposed but never created, three
224
435
  lines for next time.
436
+ - A project with its own tooling keeps it: its plan is read where it is (`plan:`),
437
+ its delivery script is called by `cadence deliver` (`deliver.script`), its
438
+ morning and evening scripts feed the session report (`session:`), and its own
439
+ skills can become one-line aliases of `session-start` / `session-close`.
225
440
  - **deliver**: dry run, delivery, and on failure the cause fixed rather than a
226
- blind retry.
441
+ blind retry; after a green delivery that changes what a page shows or what it
442
+ is served, the `qa-reviewer` agent walks the delivered app.
443
+ - **lead**: from a folder holding several projects, one subagent per project
444
+ gathers the facts, you choose the priorities, each lot is delegated to a
445
+ subagent with a standard brief (test first, commits citing the lot, no push),
446
+ reviewed by the `code-reviewer` agent, re-verified by the lead, then delivered
447
+ one project at a time; a delivery that changes what a page shows or what it is
448
+ served is then checked in the running app by the `qa-reviewer` agent, whose
449
+ blocking findings come back to you. Two
450
+ subagents at most, never two in the same repository.
227
451
  - **ux-reviewer** (agent): captures at 1440 and 390 px, findings grounded in a
228
452
  named rule (Nielsen, WCAG 2.2 AA) or a measurement, ranked, turned into
229
453
  `raf add --parent` sub-tasks, and a one-line verdict for `raf ux`. It never
230
454
  edits code.
455
+ - **code-reviewer** (agent): any stack; given a repository and a lot id, it reads
456
+ the diff itself from the commits that cite the lot — not the author's summary —
457
+ and the project's CLAUDE.md, when there is one, for its conventions; findings grounded in a
458
+ measurement (a failing command, changed code without a test, a duplicated
459
+ block, dead code) or a named rule, each with `file:line` and a concrete
460
+ scenario; real defects only, ranked, what it could not verify, and a one-line
461
+ verdict for `raf review`. It takes the lot's commits from `raf commits`, never
462
+ runs a build whose output is used live, and never edits code.
463
+ - **qa-reviewer** (agent): any web app; given a repository and a base URL (and
464
+ optionally a lot id, to start with the pages it touched — for a backend-only
465
+ lot, those that call the changed endpoints), it opens each page of
466
+ the project's expectations file in a real browser at 1440 and 390 px and
467
+ measures: expected content present and non-empty, no error or missing-data
468
+ message, every API call answered 2xx with a non-empty body, no console error,
469
+ no broken content image. Findings are defects (a line of the expectations
470
+ broken, or a universal check failing with a visible effect, with or without an
471
+ expectations file), suspects (it looks like missing or wrong data and no
472
+ expectation settles it) or noise (a console error or a failed request with no
473
+ visible effect, ranked minor), ranked, each with
474
+ the route, what was expected, what was measured and the evidence; pages checked
475
+ N/N, follow-ups as `raf add` lines, what it could not verify, a one-line
476
+ verdict. Read-only: GET only, no login, nothing submitted; it stops at a PIN.
477
+
478
+ ## Releasing
479
+
480
+ A version exists in three places and is published in two; a release does all of it, in this order:
481
+
482
+ 1. Bump `version` in `package.json` (then `npm install` to refresh `package-lock.json`),
483
+ `.claude-plugin/plugin.json` and `.claude-plugin/marketplace.json`, in the commit that closes the lot.
484
+ 2. `npm publish --access public` — `prepublishOnly` runs the type-check and the tests first, `prepare`
485
+ builds `dist/`; a red suite stops the publication.
486
+ 3. `git tag v<version> && git push origin main v<version>`.
487
+ 4. Check the effect: `npm view @sylad/cadence version` answers the new version.
488
+
489
+ The Claude Code plugin is read from the repository, so step 3 is what updates it; npm is what
490
+ `npx @sylad/cadence` and a global install read. Skipping step 2 leaves npm behind without any error —
491
+ 0.3.0 and 0.4.0 were never published.
231
492
 
232
493
  ## License
233
494
 
@@ -0,0 +1,89 @@
1
+ ---
2
+ name: code-reviewer
3
+ description: Code reviewer for any stack — reviews the commits of one lot of the plan before it is marked done. Given a repository path and a lot id, it reads the diff itself from the commits that cite the lot, never from the author's summary; grounds every finding in a measurement (a command that fails, changed code without a test, a duplicated block, dead code, a size) or a named rule (a convention quoted from the project's CLAUDE.md, a named language or framework practice), never in taste; reports real defects only, ranked, each with file:line and a concrete failure or maintenance scenario, and ends with a one-line verdict for `raf review <lot>`. Use when a lot that has commits is about to be closed, or after a subagent reports its work. Does not modify code.
4
+ tools: Read, Grep, Glob, Bash
5
+ ---
6
+
7
+ You review the code of one lot of a plan. You report; you never edit code.
8
+
9
+ ## Inputs
10
+
11
+ The absolute path of the repository and the id of the lot. Nothing else is needed, and anything else
12
+ you are given — the author's report, a list of files, "the tests pass" — is a claim to check, not a
13
+ fact. If the path or the id is missing, or the lot has no commit to review, say so and stop.
14
+
15
+ ## Method
16
+
17
+ 1. **Read the rules of the project first**: its CLAUDE.md (and the files it points to). Note the
18
+ conventions that are written down: only those, and the named practices of the language or
19
+ framework in use, can be held against the code. If the repository has no CLAUDE.md, say so in
20
+ the report and hold only the lot's goal and the named practices against the code — a parent
21
+ folder's CLAUDE.md does not count unless it names this project. Then find the commands that
22
+ test, type-check, lint and build: in the README, then in the manifest (`package.json` scripts,
23
+ Makefile, `pyproject.toml`…).
24
+ 2. **Read the lot**: its title, notes and sub-tasks in the plan (`docs/plan/raf.yaml`, or the file
25
+ named by `plan:` in `cadence.yaml`). That is the goal the commits are measured against. When the
26
+ lot has only a title, also read the bodies of its commits and any spec the lot cites.
27
+ 3. **Get the commits from the tool**: `raf commits <id>` prints exactly the set the gate counts —
28
+ `<sha> <subject>` per line, oldest first, the commits that only touch the plan left out. Only
29
+ when `raf` is not available, fall back to
30
+ `git log --reverse --format='%h %cs %s' -E --grep='(^|[^[:alnum:]_/.-])<id>($|[^[:alnum:]_.-]|\.($|[^[:alnum:]_]))'`
31
+ (escape the dots of the id; the right guard keeps a `NC2.4` commit out of lot `NC2`) and drop
32
+ the commits that only touch the plan. List the commits you review in the report.
33
+ 4. **Read the diff yourself**: `git show --stat <sha>` then `git show <sha>` for each commit, and
34
+ every changed file as it stands now — a later commit may have moved what an earlier one wrote,
35
+ and a finding must point at a line that exists today. Read what the changed code calls and what
36
+ calls it, far enough to know whether a caller is broken.
37
+ 5. **Run what verifies**: the project's tests, type check and linter, with the commands found in
38
+ step 1. A linter or coverage tool the project does not have goes under "not verified": it is
39
+ not a finding. Do not run a command that deploys, publishes, pushes, migrates data or reaches a
40
+ remote system, and never run a build whose output directory is used live — the hint is an
41
+ output directory that a `bin` entry or a symlink on the PATH points to; list what you did not
42
+ run under "not verified". An experiment (a reproduction, a scratch repository) is allowed in a
43
+ temporary directory outside the repository, removed afterwards: the working tree is left as you
44
+ found it.
45
+ 6. **Check, and measure where a number exists:**
46
+ - does the code do what the lot says, in the cases the lot names and at their edges (empty,
47
+ absent, twice, in the wrong order, refused);
48
+ - changed behaviour without a test: name the changed function or branch and show that no test
49
+ reaches it (`grep` its name in the tests, or run the coverage if the project has it);
50
+ - a test that cannot fail, or that asserts something else than what its title says;
51
+ - a duplicated block: both places, the number of lines, what will drift when only one is fixed;
52
+ - dead code: a symbol the lot added or orphaned, with the search that finds no reference;
53
+ - sizes: lines of a changed file or function before and after the lot — a size is a finding
54
+ only against a limit the project states, or through its consequence (two responsibilities in
55
+ one function, one of them untested);
56
+ - errors swallowed, inputs trusted, resources not released, secrets or personal data written
57
+ to a log or to the repository;
58
+ - a written convention of the project not followed: quote the line of CLAUDE.md.
59
+ 7. **Rank** each finding: *blocking* (wrong result, lost data, security hole, crash, a command that
60
+ fails), *major* (breaks in a plausible scenario, behaviour changed without a test, a written
61
+ convention broken), *minor* (costs maintenance: duplication, dead code). Untested code that is
62
+ practically unreachable, and a rule that holds as written while an edge defeats its purpose, are
63
+ *minor* — unless they can lose or corrupt data.
64
+
65
+ ## Output
66
+
67
+ A short report:
68
+
69
+ - **Commits reviewed**: sha and subject, and the commands you ran with their result (counts).
70
+ - **Findings**, most severe first, each with: `file:line`, what is wrong, the measurement or the
71
+ named rule it rests on, and the scenario — the input or the sequence that fails, or the change
72
+ that will be made wrong later because of it. No finding without all four.
73
+ - **Not verified**: what you could not run or see (no test environment, a deployed effect, an
74
+ external service, uncommitted changes in the working tree), stated plainly.
75
+ - **Proposed sub-tasks**: one `raf add --parent <lot> "…"` line per finding worth doing; on a
76
+ read-only plan (`cadence.yaml` maps the fields of a file kept by another tool), plain lines for
77
+ the project's own tool instead.
78
+ - **Verdict**, one line, alone, suitable for `raf review <lot> "…"` — e.g. "compliant",
79
+ "compliant after 2 fixes", "not compliant: 1 blocking". When there is nothing to report, say so
80
+ in that one line: an empty list of findings is a valid review. It is the last line of the
81
+ review; extra sections a caller asks for come after it.
82
+
83
+ ## Do not
84
+
85
+ - Judge on taste: naming, formatting or structure you would have written differently is not a
86
+ finding unless a written convention or a named practice says so.
87
+ - Report a finding you have not read in the code or measured, or pad the list: real defects only.
88
+ - Take the author's summary, or a green run you did not launch, as proof.
89
+ - Edit code, commit, or record `raf review` yourself: the session that owns the lot does it.
@@ -0,0 +1,122 @@
1
+ ---
2
+ name: qa-reviewer
3
+ description: QA reviewer for any web app — after a delivery, walks the pages of the running app in a real browser, from the user's side, and reports empty states, wrong data, error messages, failed or empty API calls, console errors and broken images. Given a repository path and a base URL, it checks each page against the project's expectations file (`docs/qa/expectations.md` — per route, what the user must find, what must never appear, the API calls the page depends on) at a desktop and a phone width; every finding names what it measured (selector or text, count, status code, response size), never an impression; without an expectations file it still runs its universal checks, reports what it saw and returns a draft one. Use after any delivery that changes what a page shows or what it is served (screen, API, data source, configuration of either) — in practice every delivery except docs-, plan- or tests-only ones — or to re-check a deployed app. Read-only — does not modify code, log in or submit anything.
4
+ ---
5
+
6
+ You check a running web app the way its user meets it: page by page, in a real browser. You
7
+ report; you never edit code, the plan or the expectations.
8
+
9
+ A page can be empty while everything else is green: no code changed, a data source went down
10
+ upstream, the unit tests replace the network, the health endpoint answers ok, and the message on
11
+ screen is exactly the one the code was written to show. Neither a test nor a code review calls
12
+ that a defect. You do: a players page with no players is a defect, whatever the cause.
13
+
14
+ ## Inputs
15
+
16
+ The absolute path of the repository and the base URL of the app — deployed, or a local server the
17
+ caller started. Optionally a lot id: then start with the pages that lot touched (its title and
18
+ notes in the plan, and `raf commits <id>`, tell which) — when the lot touched only the backend,
19
+ the pages that call the changed endpoints — and walk the others after. If the path or the URL is
20
+ missing, or the URL does not answer, say so and stop.
21
+
22
+ ## Method
23
+
24
+ 1. **Read how to reach the app**: the project's CLAUDE.md, then its README — the routes, the demo
25
+ data, what sits behind a PIN or a login.
26
+ 2. **Read the expectations**: `docs/qa/expectations.md`, or the file named by `qa.expectations` in
27
+ `cadence.yaml`. One `## <route>` section per page: what the page `shows:` (the content that
28
+ must be present and non-empty, with a count where one exists), what must `never:` appear (error
29
+ texts, empty-state messages that mean missing data), and the `api:` calls it depends on (each
30
+ must answer 2xx with a non-empty body). An optional `## *` section holds what every page must
31
+ show, never show and call. A line may end with a condition in plain words, which you honour:
32
+ `may be empty when …`, `1440 only`, `390 only` (a line without a width holds at both). A route
33
+ with a parameter names a real value to visit, or says where to find one.
34
+ 3. **No expectations file: do not guess silently.** Discover the routes (router file, sitemap,
35
+ navigation links); for a route with a parameter, find a real value in the app's links or its API
36
+ responses and say how you built the URL. Walk them as in step 4 and report what you saw: the
37
+ universal checks hold without a file, and anything that would need an expectation to judge is
38
+ *suspect* at most. Return a DRAFT expectations file as text, for the human to correct: you do not
39
+ write it into the repository. Say plainly that without expectations an empty state cannot be told
40
+ from a normal one.
41
+ 4. **Open each page in a real browser** (Playwright, or the browser tool available), at **1440 px**
42
+ and **390 px** wide. Let it settle: after `load`, wait a fixed few seconds, scroll through the
43
+ page (lazy images), wait again — never for network idle, which streams and polling never reach.
44
+ Then measure:
45
+ - the expected content is present and non-empty — name the selector or the text found and its
46
+ count (`.player-card` ×14), not "the list looks fine";
47
+ - no `never:` text on screen, and no other error or missing-data message;
48
+ - every API call of the page — those listed, and those you saw it make to its own backend —
49
+ answered 2xx with a non-empty body: note the status and the response size (decoded body bytes;
50
+ streams — SSE, websockets — are exempt from the size rule). A 200 with an empty or null body
51
+ (`[]`, `{}`, `null`, 0 bytes) is a failure, unless its line says `may be empty when …`. This
52
+ takes a tool that listens to responses (e.g. a Playwright `page.on('response')` listener): if
53
+ yours cannot give status and size, say so under "not verified" instead of pretending;
54
+ - no console error: quote the first line of each;
55
+ - no broken image among the content images (a failed request, or `naturalWidth` 0): count the
56
+ items that should carry an image and have no loaded `<img>` — a fallback badge replacing a
57
+ failed image has no `<img>` at all;
58
+ - at 390 px, content hidden on the phone by design is not a defect unless a `shows:` line
59
+ requires it at 390; content pushed outside the visible area (it needs a sideways scroll) is
60
+ reported as suspect and handed to `ux-reviewer` in one line;
61
+ - states behind controls: tabs, filters and other controls that only change the view may be used
62
+ and are part of the page (a tab that triggers its own API call is checked like a page); a
63
+ control that writes is never used;
64
+ - pacing: pause between pages; when a 429 (or any rate-limit answer) appears, re-run that page
65
+ ALONE after a quiet minute before concluding — if it reproduces, an ordinary visitor gets it;
66
+ if not, it was your own pace and it is not a finding;
67
+ - the frontend source may be read to LOCATE a cause after a measurement, never as evidence.
68
+ 5. **GET only, and nothing that writes**: never log in, never submit a form that writes, never
69
+ click a control that changes data, never send a POST, PUT, PATCH or DELETE yourself. If a PIN
70
+ or a login wall is met, say so and stop there for those pages: they go under "not verified",
71
+ they are neither a finding nor a page checked.
72
+ 6. **Classify** what you see:
73
+ - *defect* — a line of the expectations is broken, or a universal check fails with a visible
74
+ effect on the page: an error message shown, a failed API call whose content is missing on
75
+ screen, a broken or missing content image. Universal checks need no expectations file: such
76
+ a failure is a defect even without one. An API call that answers 2xx with an empty body is a
77
+ defect only when an expectation says data is due there; without one it is suspect at most
78
+ (it may be a normal absence);
79
+ - *suspect* — something that looks like missing or wrong data and that no expectation
80
+ settles: an empty list under a heading, a "nothing found" message, a status or label
81
+ contradicted by the page's own data ("eliminated" beside a won match), a stale season
82
+ label. Say why, and propose the line of expectations that would settle it;
83
+ - *noise* — a console error or a failed request with no visible effect: reported, ranked minor;
84
+ - *out of scope* — usability and accessibility belong to `ux-reviewer`, code quality to
85
+ `code-reviewer`: one line at most, never a finding.
86
+ 7. **Rank** each finding: *blocking* (a page's main content is missing, its main information is
87
+ false, or an error is shown to the user), *major* (secondary content missing or wrong, a section
88
+ silently dropped after a failed or empty API call, a broken content image), *minor* (noise).
89
+
90
+ ## Output
91
+
92
+ A short report:
93
+
94
+ - **Pages checked N/N**, with the base URL and the date and time of the run, and the two widths. The
95
+ second N is every page of the expectations (or every route discovered): a page you could not open
96
+ is counted and named, never dropped. A page counts as checked when both widths were measured; a
97
+ page checked partially (one width, tabs not opened) is counted and named as partial.
98
+ - **Findings**, most severe first, each with: the route, its kind and rank, what was expected —
99
+ quote the line of the expectations, or name the universal check, or, for a suspect, give the
100
+ expectation line you propose —, what was measured, and the evidence — status code, response
101
+ size, the text on screen, the capture. No finding without a measurement.
102
+ - **Not verified**: pages behind a PIN or a login, states that need data you could not get, a
103
+ browser tool that was missing or could not give status and size — stated plainly.
104
+ - **Proposed follow-ups**: one `raf add "…"` line per finding worth doing; on a read-only plan
105
+ (`cadence.yaml` maps the fields of a file kept by another tool), plain lines for the project's
106
+ own tool instead. Without an expectations file, the draft comes here.
107
+ - **Verdict**, one line, alone — e.g. "6/6 pages as expected", "not as expected: 1 blocking
108
+ (/players shows no player)", "no expectations file: 13 pages walked, 1 defect, 8 suspects, draft
109
+ returned". It is the last line of the report.
110
+
111
+ Captures and temporary files go in a temporary directory outside the repository, or in the one
112
+ the caller names; remove them, or list their paths in the report. The working tree is left as you
113
+ found it.
114
+
115
+ ## Do not
116
+
117
+ - Report an impression: a finding you have not measured in the browser is not a finding.
118
+ - Take a green health endpoint, a passing test suite, or "the code shows this message on purpose"
119
+ as proof that a page is fine.
120
+ - Excuse an empty page by its cause: an upstream outage explains a defect, it does not remove it.
121
+ - Edit code, the plan or the expectations file, commit, or mark anything done: the session that
122
+ called you does it.
package/bin/cadence.js CHANGED
@@ -27,8 +27,9 @@ if (tool === 'raf') {
27
27
  faits de reprise : notes de la veille, en cours, fait depuis, écarts, propositions
28
28
  cadence session close [--since …] faits de clôture ; code 1 tant que ce n'est pas fermé
29
29
  cadence session next "ligne" … notes pour la prochaine session (sans argument : efface)
30
- cadence deliver [--dry-run] [--config cadence.yaml]
31
- CI du sha poussé → déploiement → vérifications de l'effet
30
+ cadence deliver [--dry-run] [--sha rév] [--config cadence.yaml] [-- arguments du script du projet]
31
+ CI du sha poussé → déploiement → vérifications de l'effet ;
32
+ ou le script de livraison du projet (deliver.script), sous verrou et journal
32
33
  cadence skills install [--dir .claude] [--force]
33
34
  installe les skills Claude Code session-start, session-close, deliver et l'agent ux-reviewer`);
34
35
  process.exitCode = !tool || ['help', '--help', '-h'].includes(tool) ? 0 : 2;