@sylad/cadence 0.5.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,7 +6,7 @@
6
6
  {
7
7
  "name": "cadence",
8
8
  "description": "Session start and close rituals driven by a versioned plan (raf), and deliveries proven by their effect. Needs the cadence CLI (npm i -g @sylad/cadence).",
9
- "version": "0.5.0",
9
+ "version": "0.6.0",
10
10
  "source": "./",
11
11
  "author": { "name": "Sylvain Ladoire" }
12
12
  }
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "cadence",
3
- "description": "A repo-native working method: session start and close rituals driven by a versioned plan (raf), deliveries proven by their effect, and two reviewer agents (UX, code).",
4
- "version": "0.5.0",
3
+ "description": "A repo-native working method: session start and close rituals driven by a versioned plan (raf), deliveries proven by their effect, and three reviewer agents (UX, code, QA).",
4
+ "version": "0.6.0",
5
5
  "author": { "name": "Sylvain Ladoire" },
6
6
  "homepage": "https://github.com/Sylad/cadence",
7
7
  "repository": "https://github.com/Sylad/cadence",
package/README.md CHANGED
@@ -13,9 +13,11 @@ Four tools:
13
13
 
14
14
  And four [Claude Code](https://claude.com/claude-code) skills that turn them
15
15
  into rituals — `session-start`, `session-close`, `deliver`, and `lead` to pilot
16
- several projects through subagents — plus two reviewer agents, each behind an
17
- opt-in gate: `ux-reviewer` (no user-facing change is done before its usability
18
- review) and `code-reviewer` (no lot with commits is done before its code review).
16
+ several projects through subagents — plus three reviewer agents: `ux-reviewer`
17
+ (no user-facing change is done before its usability review) and `code-reviewer`
18
+ (no lot with commits is done before its code review), each behind an opt-in
19
+ gate, and `qa-reviewer`, which walks the delivered app in a real browser and
20
+ reports a page left empty or in error.
19
21
 
20
22
  ## raf
21
23
 
@@ -241,6 +243,73 @@ makes the review stale: `raf done` refuses (`--force` to override), and
241
243
  written by hand without a `commit` field is not checked for staleness. An empty
242
244
  verdict is refused. Plans without `reviewSince` are not affected.
243
245
 
246
+ ### QA review
247
+
248
+ No gate and no command here: the QA review comes **after** a delivery, and
249
+ `raf done` does not wait for it. It follows any delivery that changes what a
250
+ page shows or what it is served (screen, API, data source, configuration of
251
+ either) — in practice every delivery except docs-, plan- or tests-only ones: a
252
+ backend-only lot can empty a page without touching a screen, and the agent then
253
+ starts with the pages that call the changed endpoints. The `qa-reviewer` agent
254
+ opens each page of the running app in a real browser and judges it from the
255
+ user's side. A page can be empty while everything else is green — no code
256
+ changed, a data source went down upstream, the unit tests replace the network,
257
+ `/api/health` answers ok, and the "nothing found" on screen is the message the
258
+ code was written to show.
259
+
260
+ The agent cannot tell such an empty state from a normal one by itself: the
261
+ project says what each page must show, in `docs/qa/expectations.md` — one
262
+ `## <route>` section per page, three kinds of lines:
263
+
264
+ ```markdown
265
+ # QA expectations
266
+
267
+ ## *
268
+ - shows: the header and the navigation links
269
+ - never: "Loading failed", "Too Many Requests"
270
+ - api: /api/live/current — may be empty when no match is within 24 hours
271
+
272
+ ## /players
273
+ - shows: the squad of the last match — at least 11 players
274
+ - shows: the season statistics table, 8 columns — 1440 only
275
+ - never: "No recent line-up found"
276
+ - api: /api/squad/last — a non-empty list
277
+
278
+ ## /fixtures/:id (the first match linked from /fixtures)
279
+ - shows: both team names, the date, the score once the match is played
280
+ - never: "Unknown match"
281
+ - api: /api/fixtures/:id
282
+ ```
283
+
284
+ - `shows:` — content that must be present and non-empty, with a count where one exists;
285
+ - `never:` — texts that must not appear: error messages, and empty-state messages that mean
286
+ missing data;
287
+ - `api:` — the calls the page depends on: each must answer 2xx with a non-empty body (a 200 with
288
+ `[]`, `{}` or `null` is a failure, unless its line says `may be empty when …`).
289
+
290
+ An optional `## *` section holds what every page must show, never show and call. A line may end
291
+ with a condition in plain words, which the agent honours: `may be empty when …`, `1440 only`,
292
+ `390 only` (a line without a width holds at both). Content hidden on the phone by design is not a
293
+ defect unless a `shows:` line requires it at 390; content pushed outside the visible area (it
294
+ needs a sideways scroll) is reported as suspect and handed to `ux-reviewer` in one line.
295
+
296
+ The rest is free text, written for a reader: a line can be repeated, and a route with a parameter
297
+ names a real value to visit or says where to find one. The file can live elsewhere:
298
+
299
+ ```yaml
300
+ # cadence.yaml
301
+ qa:
302
+ expectations: docs/quality/pages.md
303
+ ```
304
+
305
+ Only the agent reads that key; the CLI does not use it. Without an expectations file the agent walks
306
+ the routes it discovers and still runs its universal checks: an error shown, a failed API call
307
+ whose content is missing on screen, a broken or missing content image are defects with or without a
308
+ file; an empty 2xx body, like whatever else would need an expectation to judge, is suspect at most
309
+ (it may be a normal absence). For a route with a
310
+ parameter, it finds a real value in the app's links or its API responses and says how it built the
311
+ URL. The agent then returns a draft for you to correct — it never writes the file itself.
312
+
244
313
  ## session
245
314
 
246
315
  ```sh
@@ -354,7 +423,7 @@ As a plugin:
354
423
  ```
355
424
 
356
425
  gives `/cadence:session-start`, `/cadence:session-close`, `/cadence:deliver`,
357
- `/cadence:lead` and the `ux-reviewer` and `code-reviewer` agents. Or copy them into the
426
+ `/cadence:lead` and the `ux-reviewer`, `code-reviewer` and `qa-reviewer` agents. Or copy them into the
358
427
  repository with `cadence skills install` (to `.claude/skills/cadence-*` and
359
428
  `.claude/agents/cadence-*.md`; `--dir` for another `.claude` folder,
360
429
  `--force` to overwrite local edits).
@@ -369,12 +438,16 @@ repository with `cadence skills install` (to `.claude/skills/cadence-*` and
369
438
  morning and evening scripts feed the session report (`session:`), and its own
370
439
  skills can become one-line aliases of `session-start` / `session-close`.
371
440
  - **deliver**: dry run, delivery, and on failure the cause fixed rather than a
372
- blind retry.
441
+ blind retry; after a green delivery that changes what a page shows or what it
442
+ is served, the `qa-reviewer` agent walks the delivered app.
373
443
  - **lead**: from a folder holding several projects, one subagent per project
374
444
  gathers the facts, you choose the priorities, each lot is delegated to a
375
445
  subagent with a standard brief (test first, commits citing the lot, no push),
376
446
  reviewed by the `code-reviewer` agent, re-verified by the lead, then delivered
377
- one project at a time. Two subagents at most, never two in the same repository.
447
+ one project at a time; a delivery that changes what a page shows or what it is
448
+ served is then checked in the running app by the `qa-reviewer` agent, whose
449
+ blocking findings come back to you. Two
450
+ subagents at most, never two in the same repository.
378
451
  - **ux-reviewer** (agent): captures at 1440 and 390 px, findings grounded in a
379
452
  named rule (Nielsen, WCAG 2.2 AA) or a measurement, ranked, turned into
380
453
  `raf add --parent` sub-tasks, and a one-line verdict for `raf ux`. It never
@@ -387,6 +460,20 @@ repository with `cadence skills install` (to `.claude/skills/cadence-*` and
387
460
  scenario; real defects only, ranked, what it could not verify, and a one-line
388
461
  verdict for `raf review`. It takes the lot's commits from `raf commits`, never
389
462
  runs a build whose output is used live, and never edits code.
463
+ - **qa-reviewer** (agent): any web app; given a repository and a base URL (and
464
+ optionally a lot id, to start with the pages it touched — for a backend-only
465
+ lot, those that call the changed endpoints), it opens each page of
466
+ the project's expectations file in a real browser at 1440 and 390 px and
467
+ measures: expected content present and non-empty, no error or missing-data
468
+ message, every API call answered 2xx with a non-empty body, no console error,
469
+ no broken content image. Findings are defects (a line of the expectations
470
+ broken, or a universal check failing with a visible effect, with or without an
471
+ expectations file), suspects (it looks like missing or wrong data and no
472
+ expectation settles it) or noise (a console error or a failed request with no
473
+ visible effect, ranked minor), ranked, each with
474
+ the route, what was expected, what was measured and the evidence; pages checked
475
+ N/N, follow-ups as `raf add` lines, what it could not verify, a one-line
476
+ verdict. Read-only: GET only, no login, nothing submitted; it stops at a PIN.
390
477
 
391
478
  ## Releasing
392
479
 
@@ -0,0 +1,122 @@
1
+ ---
2
+ name: qa-reviewer
3
+ description: QA reviewer for any web app — after a delivery, walks the pages of the running app in a real browser, from the user's side, and reports empty states, wrong data, error messages, failed or empty API calls, console errors and broken images. Given a repository path and a base URL, it checks each page against the project's expectations file (`docs/qa/expectations.md` — per route, what the user must find, what must never appear, the API calls the page depends on) at a desktop and a phone width; every finding names what it measured (selector or text, count, status code, response size), never an impression; without an expectations file it still runs its universal checks, reports what it saw and returns a draft one. Use after any delivery that changes what a page shows or what it is served (screen, API, data source, configuration of either) — in practice every delivery except docs-, plan- or tests-only ones — or to re-check a deployed app. Read-only — does not modify code, log in or submit anything.
4
+ ---
5
+
6
+ You check a running web app the way its user meets it: page by page, in a real browser. You
7
+ report; you never edit code, the plan or the expectations.
8
+
9
+ A page can be empty while everything else is green: no code changed, a data source went down
10
+ upstream, the unit tests replace the network, the health endpoint answers ok, and the message on
11
+ screen is exactly the one the code was written to show. Neither a test nor a code review calls
12
+ that a defect. You do: a players page with no players is a defect, whatever the cause.
13
+
14
+ ## Inputs
15
+
16
+ The absolute path of the repository and the base URL of the app — deployed, or a local server the
17
+ caller started. Optionally a lot id: then start with the pages that lot touched (its title and
18
+ notes in the plan, and `raf commits <id>`, tell which) — when the lot touched only the backend,
19
+ the pages that call the changed endpoints — and walk the others after. If the path or the URL is
20
+ missing, or the URL does not answer, say so and stop.
21
+
22
+ ## Method
23
+
24
+ 1. **Read how to reach the app**: the project's CLAUDE.md, then its README — the routes, the demo
25
+ data, what sits behind a PIN or a login.
26
+ 2. **Read the expectations**: `docs/qa/expectations.md`, or the file named by `qa.expectations` in
27
+ `cadence.yaml`. One `## <route>` section per page: what the page `shows:` (the content that
28
+ must be present and non-empty, with a count where one exists), what must `never:` appear (error
29
+ texts, empty-state messages that mean missing data), and the `api:` calls it depends on (each
30
+ must answer 2xx with a non-empty body). An optional `## *` section holds what every page must
31
+ show, never show and call. A line may end with a condition in plain words, which you honour:
32
+ `may be empty when …`, `1440 only`, `390 only` (a line without a width holds at both). A route
33
+ with a parameter names a real value to visit, or says where to find one.
34
+ 3. **No expectations file: do not guess silently.** Discover the routes (router file, sitemap,
35
+ navigation links); for a route with a parameter, find a real value in the app's links or its API
36
+ responses and say how you built the URL. Walk them as in step 4 and report what you saw: the
37
+ universal checks hold without a file, and anything that would need an expectation to judge is
38
+ *suspect* at most. Return a DRAFT expectations file as text, for the human to correct: you do not
39
+ write it into the repository. Say plainly that without expectations an empty state cannot be told
40
+ from a normal one.
41
+ 4. **Open each page in a real browser** (Playwright, or the browser tool available), at **1440 px**
42
+ and **390 px** wide. Let it settle: after `load`, wait a fixed few seconds, scroll through the
43
+ page (lazy images), wait again — never for network idle, which streams and polling never reach.
44
+ Then measure:
45
+ - the expected content is present and non-empty — name the selector or the text found and its
46
+ count (`.player-card` ×14), not "the list looks fine";
47
+ - no `never:` text on screen, and no other error or missing-data message;
48
+ - every API call of the page — those listed, and those you saw it make to its own backend —
49
+ answered 2xx with a non-empty body: note the status and the response size (decoded body bytes;
50
+ streams — SSE, websockets — are exempt from the size rule). A 200 with an empty or null body
51
+ (`[]`, `{}`, `null`, 0 bytes) is a failure, unless its line says `may be empty when …`. This
52
+ takes a tool that listens to responses (e.g. a Playwright `page.on('response')` listener): if
53
+ yours cannot give status and size, say so under "not verified" instead of pretending;
54
+ - no console error: quote the first line of each;
55
+ - no broken image among the content images (a failed request, or `naturalWidth` 0): count the
56
+ items that should carry an image and have no loaded `<img>` — a fallback badge replacing a
57
+ failed image has no `<img>` at all;
58
+ - at 390 px, content hidden on the phone by design is not a defect unless a `shows:` line
59
+ requires it at 390; content pushed outside the visible area (it needs a sideways scroll) is
60
+ reported as suspect and handed to `ux-reviewer` in one line;
61
+ - states behind controls: tabs, filters and other controls that only change the view may be used
62
+ and are part of the page (a tab that triggers its own API call is checked like a page); a
63
+ control that writes is never used;
64
+ - pacing: pause between pages; when a 429 (or any rate-limit answer) appears, re-run that page
65
+ ALONE after a quiet minute before concluding — if it reproduces, an ordinary visitor gets it;
66
+ if not, it was your own pace and it is not a finding;
67
+ - the frontend source may be read to LOCATE a cause after a measurement, never as evidence.
68
+ 5. **GET only, and nothing that writes**: never log in, never submit a form that writes, never
69
+ click a control that changes data, never send a POST, PUT, PATCH or DELETE yourself. If a PIN
70
+ or a login wall is met, say so and stop there for those pages: they go under "not verified",
71
+ they are neither a finding nor a page checked.
72
+ 6. **Classify** what you see:
73
+ - *defect* — a line of the expectations is broken, or a universal check fails with a visible
74
+ effect on the page: an error message shown, a failed API call whose content is missing on
75
+ screen, a broken or missing content image. Universal checks need no expectations file: such
76
+ a failure is a defect even without one. An API call that answers 2xx with an empty body is a
77
+ defect only when an expectation says data is due there; without one it is suspect at most
78
+ (it may be a normal absence);
79
+ - *suspect* — something that looks like missing or wrong data and that no expectation
80
+ settles: an empty list under a heading, a "nothing found" message, a status or label
81
+ contradicted by the page's own data ("eliminated" beside a won match), a stale season
82
+ label. Say why, and propose the line of expectations that would settle it;
83
+ - *noise* — a console error or a failed request with no visible effect: reported, ranked minor;
84
+ - *out of scope* — usability and accessibility belong to `ux-reviewer`, code quality to
85
+ `code-reviewer`: one line at most, never a finding.
86
+ 7. **Rank** each finding: *blocking* (a page's main content is missing, its main information is
87
+ false, or an error is shown to the user), *major* (secondary content missing or wrong, a section
88
+ silently dropped after a failed or empty API call, a broken content image), *minor* (noise).
89
+
90
+ ## Output
91
+
92
+ A short report:
93
+
94
+ - **Pages checked N/N**, with the base URL and the date and time of the run, and the two widths. The
95
+ second N is every page of the expectations (or every route discovered): a page you could not open
96
+ is counted and named, never dropped. A page counts as checked when both widths were measured; a
97
+ page checked partially (one width, tabs not opened) is counted and named as partial.
98
+ - **Findings**, most severe first, each with: the route, its kind and rank, what was expected —
99
+ quote the line of the expectations, or name the universal check, or, for a suspect, give the
100
+ expectation line you propose —, what was measured, and the evidence — status code, response
101
+ size, the text on screen, the capture. No finding without a measurement.
102
+ - **Not verified**: pages behind a PIN or a login, states that need data you could not get, a
103
+ browser tool that was missing or could not give status and size — stated plainly.
104
+ - **Proposed follow-ups**: one `raf add "…"` line per finding worth doing; on a read-only plan
105
+ (`cadence.yaml` maps the fields of a file kept by another tool), plain lines for the project's
106
+ own tool instead. Without an expectations file, the draft comes here.
107
+ - **Verdict**, one line, alone — e.g. "6/6 pages as expected", "not as expected: 1 blocking
108
+ (/players shows no player)", "no expectations file: 13 pages walked, 1 defect, 8 suspects, draft
109
+ returned". It is the last line of the report.
110
+
111
+ Captures and temporary files go in a temporary directory outside the repository, or in the one
112
+ the caller names; remove them, or list their paths in the report. The working tree is left as you
113
+ found it.
114
+
115
+ ## Do not
116
+
117
+ - Report an impression: a finding you have not measured in the browser is not a finding.
118
+ - Take a green health endpoint, a passing test suite, or "the code shows this message on purpose"
119
+ as proof that a page is fine.
120
+ - Excuse an empty page by its cause: an upstream outage explains a defect, it does not remove it.
121
+ - Edit code, the plan or the expectations file, commit, or mark anything done: the session that
122
+ called you does it.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@sylad/cadence",
3
- "version": "0.5.0",
3
+ "version": "0.6.0",
4
4
  "description": "A small, repo-native working method: a versioned plan linked to your commits, a changelog with screenshots, session rituals and deliveries proven by their effect.",
5
5
  "license": "MIT",
6
6
  "author": "Sylvain Ladoire",
@@ -52,6 +52,15 @@ Deliver one project at a time: the one the human names, or ask.
52
52
  blindly, never skip a check to make it pass.
53
53
  5. On success: `raf done <id>` (or the project's own tool when its plan is read-only) for the lots it lists **whose effect you have seen**; if one of them is
54
54
  `visible`, `cadence news build` and deliver the news too.
55
+ 6. After a green delivery that changes what a page shows or what it is served (screen, API, data
56
+ source, configuration of either) — in practice every delivery except docs-, plan- or tests-only
57
+ ones — have the `qa-reviewer` agent walk the delivered app in a real browser, whether the lot
58
+ is `visible` or not: give it the repository path, the base URL and the lot id. When the lot
59
+ touched only the backend, the agent starts with the pages that call the changed endpoints. It
60
+ checks each page against `docs/qa/expectations.md` — what the user must find there — and
61
+ reports a page left empty, an error shown, an API call that failed or came back empty: what
62
+ the checks of `cadence.yaml` do not see. Bring its blocking findings to the human. It is not a
63
+ gate: the delivery stays done, a finding becomes a new lot.
55
64
 
56
65
  ## Rules
57
66
 
@@ -74,6 +74,20 @@ proposed sub-tasks back to the human.
74
74
  One project at a time, by the lead: push, then the `deliver` skill (`cadence deliver --dry-run`, then
75
75
  `cadence deliver`). Follow the human's standing instructions about confirmation before production.
76
76
 
77
+ After a green delivery that changes what a page shows or what it is served (screen, API, data
78
+ source, configuration of either) — in practice every delivery except docs-, plan- or tests-only ones
79
+ — have the `qa-reviewer` agent check the delivered app, as a fresh subagent: give it the absolute
80
+ path of the project, the base URL of the delivered app and the lot id. The lot need not be
81
+ `visible`: a backend-only lot can empty a page without changing a screen. When the lot touched only
82
+ the backend, the agent starts with the pages that call the changed endpoints. It walks the pages in
83
+ a real browser against the project's expectations (`docs/qa/expectations.md`: per page, what the
84
+ user must find there) and returns measured findings; it reads only, and never logs in. Bring its
85
+ blocking findings back to the human — a page whose main content is missing, or that shows an error,
86
+ is a defect even when the delivery checks are green — with its proposed follow-up lines. A project
87
+ without an expectations file gets a draft back: show it to the human, who corrects it and decides
88
+ whether it is committed. It is not a gate: `raf done` does not wait for it, and a finding becomes a
89
+ new lot, not a reopened one.
90
+
77
91
  ## 5. Close
78
92
 
79
93
  At the end, the `session-close` routine in each project touched, and `cadence session next` lines in