@mrciphersmith/keryx 0.2.73 → 0.2.75
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli.js +33536 -33010
- package/package.json +1 -1
- package/src/gdskills/bundled/skills/core/reviewer-skill-creator/SKILL.md +214 -0
- package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +48 -1
- package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/input-contract.schema.json +70 -4
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/review-clean-code/SKILL.md +33 -1
- package/src/gdskills/bundled/skills/review/review-layout/SKILL.md +217 -0
- package/src/gdskills/bundled/skills/review/review-logic/SKILL.md +26 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +304 -2
- package/src/gdskills/bundled/skills/review/review-pr-feedback/SKILL.md +644 -113
- package/src/gdskills/bundled/skills/review/review-pr-feedback/input-contract.schema.json +79 -0
- package/src/gdskills/bundled/skills/review/review-pr-feedback/output-contract.schema.json +375 -0
- package/src/gdskills/bundled/skills/review/review-testing-practices/SKILL.md +111 -1
- package/src/gdskills/bundled/skills/review/review-verifier/SKILL.md +25 -1
|
@@ -0,0 +1,217 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: review-layout
|
|
3
|
+
model_tier: standard
|
|
4
|
+
description: |
|
|
5
|
+
Use when reviewing changes to rendered layout: flex and grid sizing, intrinsic vs
|
|
6
|
+
extrinsic sizing, overflow and collapse, logical properties and RTL, and sensitivity
|
|
7
|
+
to translated text length. Every finding is measured in a real layout engine, never
|
|
8
|
+
reasoned from class names in a DOM stub. Dispatched by review-orchestrator for
|
|
9
|
+
--layout, --all, or when the diff touches sizing/spacing utilities on a rendered
|
|
10
|
+
element.
|
|
11
|
+
NOT for: whether a class exists or comes from the theme (review-frontend-conventions),
|
|
12
|
+
naming and readability (review-style), re-render cost (review-performance), React or
|
|
13
|
+
MobX structure (review-frontend).
|
|
14
|
+
triggers:
|
|
15
|
+
- "review layout"
|
|
16
|
+
- "layout review"
|
|
17
|
+
- "does this render correctly"
|
|
18
|
+
- "dispatched by review-orchestrator"
|
|
19
|
+
metadata:
|
|
20
|
+
author: "MrCipherSmith"
|
|
21
|
+
version: "1.0.0"
|
|
22
|
+
category: "review"
|
|
23
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
24
|
+
license: "MIT"
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
# Review — Layout
|
|
28
|
+
|
|
29
|
+
Reviewer for what the CSS actually renders.
|
|
30
|
+
|
|
31
|
+
This lane exists because nobody else owns it. `review-frontend` owns React and
|
|
32
|
+
MobX structure; `review-style` explicitly excludes everything but naming and
|
|
33
|
+
readability; `review-frontend-conventions` owns whether a class is allowed;
|
|
34
|
+
`review-performance` owns re-render cost. **None of them answers whether the box
|
|
35
|
+
is the size the author thinks it is** — which is why a bar that collapses to 0px
|
|
36
|
+
and a score that sits 12px off centre both shipped through a full review round.
|
|
37
|
+
|
|
38
|
+
Every finding here is a claim about pixels, so every finding here is measured or
|
|
39
|
+
it is not filed.
|
|
40
|
+
|
|
41
|
+
---
|
|
42
|
+
|
|
43
|
+
## Scope
|
|
44
|
+
|
|
45
|
+
Dispatched when scope A touches, on a **rendered** element:
|
|
46
|
+
|
|
47
|
+
- sizing: `w-*`, `h-*`, `min-w-*`, `max-w-*`, `basis-*`, `flex-*`, `grid-cols-*`,
|
|
48
|
+
`w-fit`, `min-w-0`, `truncate`, `line-clamp-*`
|
|
49
|
+
- spacing that participates in centring or the box model: `p*`, `m*`, `gap-*`, and
|
|
50
|
+
their logical forms `ps-*`, `pe-*`, `ms-*`, `me-*`
|
|
51
|
+
- `overflow-*`, `position`, `z-*`, `aspect-*`
|
|
52
|
+
- inline `style` carrying a computed dimension
|
|
53
|
+
- any container whose children render translated text
|
|
54
|
+
|
|
55
|
+
Out of scope: whether a class exists or comes from the theme, and whether it is
|
|
56
|
+
well named.
|
|
57
|
+
|
|
58
|
+
---
|
|
59
|
+
|
|
60
|
+
## The one method that makes this lane real
|
|
61
|
+
|
|
62
|
+
**A DOM stub loads no stylesheet.** Under happy-dom or jsdom,
|
|
63
|
+
`toHaveClass("w-fit")` observes a string and never a pixel, and
|
|
64
|
+
`getBoundingClientRect` returns zeros for everything. A layout claim reasoned from
|
|
65
|
+
class names is worth nothing, and filing one is how this lane degenerates into a
|
|
66
|
+
class-name spellchecker.
|
|
67
|
+
|
|
68
|
+
Verify one of two ways, and put the numbers in the finding:
|
|
69
|
+
|
|
70
|
+
1. The project's real-CSS tier — the Storybook / Playwright / component-visual
|
|
71
|
+
suite. Find its script in `package.json` rather than assuming a name.
|
|
72
|
+
2. A direct probe that mounts the component in a real engine and reads computed
|
|
73
|
+
CSS off the node.
|
|
74
|
+
|
|
75
|
+
Report a table of measurements, not an adjective. `117.6px` is a finding; "the bar
|
|
76
|
+
may be too short" is not.
|
|
77
|
+
|
|
78
|
+
```
|
|
79
|
+
host width longest label bar width
|
|
80
|
+
560px "Controls" 256.0px
|
|
81
|
+
560px <longest ar-AE> 182.6px
|
|
82
|
+
280px "Controls" 0.0px <- collapses
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
**If you cannot measure, you cannot file.** Say which command you tried, what
|
|
86
|
+
stopped you, and what would settle it. An unmeasured layout observation is `info`,
|
|
87
|
+
marked unverified — that is the correct outcome, not a reason to reach for a class
|
|
88
|
+
name instead.
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
## Checklist
|
|
93
|
+
|
|
94
|
+
### Intrinsic sizing traps
|
|
95
|
+
|
|
96
|
+
- [ ] A `flex-1` / `basis-0` column inside a `w-fit` (or `max-content`) parent has
|
|
97
|
+
**no basis of its own**: its width is whatever the siblings' content leaves.
|
|
98
|
+
Measure it, then measure it again with the longest translated label the
|
|
99
|
+
catalog contains.
|
|
100
|
+
- [ ] `min-w-0` removes the automatic minimum that stops a flex item shrinking past
|
|
101
|
+
its content. Combined with a content-sized parent it permits **0px**. Find the
|
|
102
|
+
host width at which the element reaches zero and state it.
|
|
103
|
+
- [ ] `w-fit` + `max-w-full` is an overflow fix, not a sizing floor. Replacing a
|
|
104
|
+
removed `min-w-*` with `min-w-0` trades an overflow bug for a collapse bug.
|
|
105
|
+
- [ ] A `min-w-*` on a child plus `w-fit` at the call sites gives the card a hard
|
|
106
|
+
floor. Compute it, and check it against the narrowest supported host.
|
|
107
|
+
|
|
108
|
+
### Box model and centring
|
|
109
|
+
|
|
110
|
+
- [ ] Padding inside a fixed-width box shifts centred content by half the asymmetric
|
|
111
|
+
padding. `w-48` (192px) with `pe-6` (24px) centres in 168px — 12px toward the
|
|
112
|
+
inline start.
|
|
113
|
+
- [ ] Directional padding that is a gutter in one state is an uncompensated offset in
|
|
114
|
+
the state where the neighbour is absent. Check **every** conditional render of
|
|
115
|
+
that sibling, including the empty and never-run states.
|
|
116
|
+
|
|
117
|
+
### Logical properties and RTL
|
|
118
|
+
|
|
119
|
+
- [ ] `ps-*` / `pe-*` / `ms-*` / `me-*` flip under RTL: an offset toward the start in
|
|
120
|
+
LTR is toward the end in RTL. State both.
|
|
121
|
+
- [ ] A physical gradient, shadow, or transform in a component that otherwise uses
|
|
122
|
+
logical properties does **not** flip. Name the direction it points in each.
|
|
123
|
+
- [ ] Text length differs by locale. A component sized by its own labels renders
|
|
124
|
+
**different geometry for identical data** across locales. Measure with the
|
|
125
|
+
longest real translation in the catalog, never the English one.
|
|
126
|
+
|
|
127
|
+
### Truncation and overflow
|
|
128
|
+
|
|
129
|
+
- [ ] `truncate` / `line-clamp` needs a constrained width to do anything at all.
|
|
130
|
+
- [ ] `overflow-hidden` on a container with a percentage-width child silently clips
|
|
131
|
+
rather than scrolls.
|
|
132
|
+
|
|
133
|
+
---
|
|
134
|
+
|
|
135
|
+
## Iron Laws
|
|
136
|
+
|
|
137
|
+
### This reviewer's own
|
|
138
|
+
|
|
139
|
+
1. **A layout finding without a measurement is `info`.** Name the host width, the
|
|
140
|
+
element width or the offset in pixels, and the command or probe that produced it.
|
|
141
|
+
2. **Measure the worst real case, not the convenient one.** The longest translation
|
|
142
|
+
in the catalog, the narrowest supported host, the state where the sibling is
|
|
143
|
+
absent, the locale that flips direction.
|
|
144
|
+
3. **Never flag a theoretical narrow viewport.** If no supported host reaches the
|
|
145
|
+
collapse width, the finding is `info` and says what would settle it. A breakpoint
|
|
146
|
+
nobody ships is not a trigger.
|
|
147
|
+
|
|
148
|
+
### Shared laws (every reviewer)
|
|
149
|
+
|
|
150
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
151
|
+
name the input, call, or condition that reaches the code, you have an
|
|
152
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
153
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
154
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
155
|
+
pattern because it is often wrong elsewhere.
|
|
156
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
157
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
158
|
+
finding hide the other nine problems.
|
|
159
|
+
|
|
160
|
+
Severity levels are defined once, in `review-orchestrator/SKILL.md` →
|
|
161
|
+
**Severity (canonical)**. This reviewer keeps no rubric of its own; below is only
|
|
162
|
+
where its recurring conditions land under that one.
|
|
163
|
+
|
|
164
|
+
| Condition | Severity | Why, under the canonical rubric |
|
|
165
|
+
|---|---|---|
|
|
166
|
+
| An element renders at 0px while its label shows, at a supported host width | `major` | Named trigger (that width), named outcome (the control is invisible) |
|
|
167
|
+
| Identical data renders different geometry per locale | `major` | Named trigger (the locale), named outcome (a different picture of the same number) |
|
|
168
|
+
| Content clipped, or the page forced to scroll horizontally, at a supported width | `major` | Named trigger and outcome |
|
|
169
|
+
| A cosmetic offset a user would not notice — a few px off centre | `minor` | The code behaves correctly; the cost is aesthetic |
|
|
170
|
+
| A measured-but-unreachable collapse, or any claim you could not measure | `info` | Shared law 1 |
|
|
171
|
+
|
|
172
|
+
---
|
|
173
|
+
|
|
174
|
+
## Orchestrated Review Contract
|
|
175
|
+
|
|
176
|
+
When dispatched by `review-orchestrator`, follow the provided
|
|
177
|
+
`reviewer-input.schema.json` payload. Return a `REVIEW_RESULT` object compatible
|
|
178
|
+
with `skills/review/review-orchestrator/reviewer-finding.schema.json`, then a
|
|
179
|
+
concise markdown summary. Prefix finding ids `LY-`.
|
|
180
|
+
|
|
181
|
+
### Class scope — required for `blocker` and `major`
|
|
182
|
+
|
|
183
|
+
Every `blocker` and `major` carries `class_scope`: **every** element with the same
|
|
184
|
+
sizing shape, and **how you enumerated them** — the search you ran, or the
|
|
185
|
+
component whose call sites derive the set.
|
|
186
|
+
|
|
187
|
+
```yaml
|
|
188
|
+
class_scope:
|
|
189
|
+
sites: ["src/dq/components/DqScoreCard.tsx:111", "src/dq/components/DqTrendCard.tsx:88"]
|
|
190
|
+
enumeration_method: "grep for `min-w-0 flex-1` under a w-fit ancestor; 4 cards, 2 collapse"
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
"I checked the others" is not an enumeration method. A single-entry `sites` list is
|
|
194
|
+
a claim that the class has exactly one member — make it deliberately.
|
|
195
|
+
|
|
196
|
+
```markdown
|
|
197
|
+
### [LY-NNN] Title
|
|
198
|
+
|
|
199
|
+
- **Severity**: blocker | major | minor | info
|
|
200
|
+
- **File**: path/to/Component.tsx:line
|
|
201
|
+
- **Problem**: what the box does
|
|
202
|
+
- **Measurement**: host width, element width or offset, and the command or probe
|
|
203
|
+
- **Why it matters**: what the user sees, in which state and which locale
|
|
204
|
+
- **Fix**: the concrete sizing change
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
---
|
|
208
|
+
|
|
209
|
+
## Red Flags
|
|
210
|
+
|
|
211
|
+
| Rationalization | Why it is wrong |
|
|
212
|
+
|----------------|-----------------|
|
|
213
|
+
| "`toHaveClass('w-fit')` passes, so the width is right." | The stub loads no CSS. That assertion cannot fail on a layout bug. |
|
|
214
|
+
| "This will obviously overflow on mobile." | Measure it, or name the supported host that reaches it. Otherwise `info`. |
|
|
215
|
+
| "The class names look fine." | This lane does not review class names. That is `review-frontend-conventions`. |
|
|
216
|
+
| "I could not run the visual tier, so I reasoned it out." | Then it is `info`, marked unverified. Reasoning does not become measurement by being careful. |
|
|
217
|
+
| "RTL is handled, the component uses logical properties." | Check the gradients, shadows and transforms too. Those are the ones that do not flip. |
|
|
@@ -107,6 +107,32 @@ Do not review unrelated files.
|
|
|
107
107
|
- [ ] Incorrect boolean logic (double negation, De Morgan's law violations)
|
|
108
108
|
- [ ] Unreachable branches or dead code that hides a bug
|
|
109
109
|
|
|
110
|
+
### The Change That Does Nothing
|
|
111
|
+
|
|
112
|
+
The diff adds a prop, guard, field or branch, and nothing can reach it. The code is
|
|
113
|
+
correct — which is why every reviewer asking "is this correct" passes it — and the
|
|
114
|
+
work the change was written to do is not done.
|
|
115
|
+
|
|
116
|
+
- [ ] A prop passed to a component that cannot act on it: the parent unmounts the
|
|
117
|
+
child rather than disabling it, or the render site the prop guards is gated by
|
|
118
|
+
a condition that is false whenever the prop is true.
|
|
119
|
+
- [ ] A field added to a type, a `Pick`, or a payload and never read. Search for the
|
|
120
|
+
reads before accepting it as plumbing for a later change; if it is, say so.
|
|
121
|
+
- [ ] A guard against a value its producer cannot emit — `=== undefined` against a
|
|
122
|
+
backend that stamps `0`, a status the producer writes in one place the route
|
|
123
|
+
never reaches.
|
|
124
|
+
- [ ] State set on one path and cleared on none, or cleared only on paths the
|
|
125
|
+
triggering interaction cannot take.
|
|
126
|
+
- [ ] A branch whose condition is unsatisfiable given the call sites in this repo.
|
|
127
|
+
|
|
128
|
+
Prove it by **negative enumeration**: name the complete candidate set and why each
|
|
129
|
+
member fails to apply, in `class_scope.enumeration_method` with `sites: []`. See
|
|
130
|
+
**Negative enumeration** in `review-orchestrator/SKILL.md` → Finding Format.
|
|
131
|
+
|
|
132
|
+
`minor` when the change is merely dead. `major` when its deadness means the defect
|
|
133
|
+
it was written to fix is still live — the usual case when the change answers an
|
|
134
|
+
earlier review round, because there the finding is recorded as closed and is not.
|
|
135
|
+
|
|
110
136
|
### Null / Undefined / Optional Chaining
|
|
111
137
|
|
|
112
138
|
- [ ] Accessing property on value that can be `null` or `undefined` without guard
|
|
@@ -33,7 +33,7 @@ triggers:
|
|
|
33
33
|
- "review --mobx-store"
|
|
34
34
|
metadata:
|
|
35
35
|
author: "MrCipherSmith"
|
|
36
|
-
version: "1.
|
|
36
|
+
version: "1.9.0"
|
|
37
37
|
category: "review"
|
|
38
38
|
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
39
39
|
license: "MIT"
|
|
@@ -66,7 +66,7 @@ Review Orchestrator Progress:
|
|
|
66
66
|
- [ ] Step 11: Sort by severity, deduplicate, emit unified report
|
|
67
67
|
- [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
|
|
68
68
|
- [ ] Step 13: Report the stage counts: dropped by pre-filter, refuted by the verifier, retained
|
|
69
|
-
- [ ] Step 14: AFTER THE FINAL ROUND ONLY — answer every external comment once, `keryx review comments reply --final`
|
|
69
|
+
- [ ] Step 14: AFTER THE FINAL ROUND ONLY — answer every external comment once, `keryx review comments reply --final` — never against a pull request the dispatch named as the caller's
|
|
70
70
|
```
|
|
71
71
|
|
|
72
72
|
Step 0 runs on **every** round. Step 14 runs **once**, after the last one. They are
|
|
@@ -394,6 +394,41 @@ problem, mark its outcome `escalate: true`: it leaves the reply queue, is report
|
|
|
394
394
|
to the operator immediately, and the command exits non-zero. Answering a blocking
|
|
395
395
|
question at the end answers the wrong question late.
|
|
396
396
|
|
|
397
|
+
### A round never answers a pull request another skill is answering
|
|
398
|
+
|
|
399
|
+
`review-pr-feedback` is the entry point for the other direction of this pipe: a
|
|
400
|
+
human or a bot has already reviewed pull request `#A`, and someone wants those
|
|
401
|
+
comments interpreted, checked against the code, fixed and answered. Under its
|
|
402
|
+
`--fix` mode it dispatches `flow-orchestrator`, which opens a **second** pull
|
|
403
|
+
request `#B` carrying the fix, based on `#A`'s own head branch, and dispatches
|
|
404
|
+
**this** orchestrator on every round against `#B`.
|
|
405
|
+
|
|
406
|
+
`#B` is its own conversation. Collect and reply on it exactly as always: someone
|
|
407
|
+
reviewing the fix deserves an answer from the run that made the fix, and its
|
|
408
|
+
record is filed under `#B`'s own number, so nothing about `#A` is touched.
|
|
409
|
+
|
|
410
|
+
The rule is about the other pull request:
|
|
411
|
+
|
|
412
|
+
- **Never run a reply pass against a pull request the dispatch named as the
|
|
413
|
+
caller's.** When `constraints` says the caller owns the reply for `#A` — or
|
|
414
|
+
the target resolves onto a pull request another skill declared — collect if the
|
|
415
|
+
round needs the record and stop there. `review-pr-feedback` answers `#A` once,
|
|
416
|
+
after the merge, from the outcomes its own verdicts produced.
|
|
417
|
+
|
|
418
|
+
The harm is not a duplicate. `collectPrComments` skips a comment whose record
|
|
419
|
+
carries a `reply_url` as `already-handled`, rescuing it only when somebody else
|
|
420
|
+
posts later in the thread — so a round that answered mid-loop writes that record
|
|
421
|
+
FIRST, and the post-merge reply citing the merge SHA is then skipped as already
|
|
422
|
+
answered. The reviewer keeps the interim answer, which by then has stopped being
|
|
423
|
+
true, and `replies.posted` counts only what went out. A suppressed correction is
|
|
424
|
+
worse than a duplicate, because nothing shows it is missing.
|
|
425
|
+
- **Absent such a constraint this orchestrator owns the reply**, as it always
|
|
426
|
+
has. A top-level review of a pull request is the normal case; a caller holding
|
|
427
|
+
the conversation is the exception, and the exception declares itself.
|
|
428
|
+
- The reply is the only thing that moves. Collection, severity classification,
|
|
429
|
+
the refusal to let the verifier refute an external comment, and the cap
|
|
430
|
+
exemption are unchanged in both shapes.
|
|
431
|
+
|
|
397
432
|
---
|
|
398
433
|
|
|
399
434
|
## Everything written to GitHub is brief
|
|
@@ -427,12 +462,65 @@ Required content:
|
|
|
427
462
|
- Git/PR metadata: repo, branch, base, head, merge-base, PR number/URL when available.
|
|
428
463
|
- Scope summary: changed files grouped by domain, high-risk files, generated/ignored files.
|
|
429
464
|
- Requirements: issue URL, linked task docs, acceptance criteria extracted from `context_doc` when available.
|
|
465
|
+
- **The PR's own description**, fetched not assumed: `gh pr view <n> --json title,body`. See below.
|
|
466
|
+
- **Cross-repo contracts the diff depends on**, each pinned to the SHA you read it at. See below.
|
|
430
467
|
- Rules: matched repository rules and convention docs by path.
|
|
431
468
|
- **Memory: accepted project memory intersecting the changed paths.** See below — this step is required, not best-effort.
|
|
432
469
|
- Decisions: why each reviewer was selected or skipped.
|
|
433
470
|
- Token policy: effective budget, truncation decisions, files summarized instead of fully inlined.
|
|
434
471
|
- Legacy/profile reviewer availability and selection state.
|
|
435
472
|
|
|
473
|
+
### The PR description is evidence, not decoration
|
|
474
|
+
|
|
475
|
+
Fetch the body into `review_context.pr.body` and hand it to every reviewer. It is
|
|
476
|
+
the author's own statement of what the change does, and it is checkable against
|
|
477
|
+
the diff.
|
|
478
|
+
|
|
479
|
+
**A description that promises an approach the diff does not take is a `minor`
|
|
480
|
+
finding** — file it, do not merely mention it. The failure this catches is
|
|
481
|
+
specific: over a multi-round review the code moves and the description does not,
|
|
482
|
+
so by round three the PR body describes an approach that was deleted in round one.
|
|
483
|
+
Whoever reads the merge commit a year later reads the description, not the rounds.
|
|
484
|
+
|
|
485
|
+
It is `minor` and not `major` because nothing at runtime is wrong. It is a finding
|
|
486
|
+
and not a pleasantry because unfiled housekeeping is raised again every round and
|
|
487
|
+
fixed in none — three consecutive rounds of "still worth rewriting" is the
|
|
488
|
+
recorded outcome of leaving it out of the findings array.
|
|
489
|
+
|
|
490
|
+
Same class, same severity: an approach that depends on another repository's change
|
|
491
|
+
being deployed first, with no deploy note saying so.
|
|
492
|
+
|
|
493
|
+
### Cross-repo contracts are read, not assumed
|
|
494
|
+
|
|
495
|
+
When the diff consumes a contract owned by another service — a payload shape, a
|
|
496
|
+
status enum, a timeout, a permission, a default — **read the producer** and pin
|
|
497
|
+
the SHA:
|
|
498
|
+
|
|
499
|
+
```yaml
|
|
500
|
+
cross_repo:
|
|
501
|
+
- repo: vantage-backend
|
|
502
|
+
sha: f5219d5d4
|
|
503
|
+
reason: "DQ report payload: which halves are null vs 0"
|
|
504
|
+
facts:
|
|
505
|
+
- "sqlScore/schemaDriftScore stay null when the half is absent"
|
|
506
|
+
- "sqlWeightPercentage/schemaDriftWeightPercentage init to 0.0 and are set unconditionally"
|
|
507
|
+
```
|
|
508
|
+
|
|
509
|
+
Put the facts in `review_context` so every reviewer shares one reading, and so a
|
|
510
|
+
later round can tell a contract that moved from a reviewer that misread it.
|
|
511
|
+
|
|
512
|
+
The rule that follows for reviewers, and that belongs in the dispatch prompt:
|
|
513
|
+
|
|
514
|
+
> **A claim about another service's behaviour, with no `file:line` at a pinned SHA
|
|
515
|
+
> behind it, is `info`.** Not `major`, however confident. The two failure modes are
|
|
516
|
+
> symmetric and both recorded: a finding asserted from the consuming side that the
|
|
517
|
+
> producer disproves, and a defect missed because the producer's actual default was
|
|
518
|
+
> assumed rather than read.
|
|
519
|
+
|
|
520
|
+
If the other repository is not available to you, say so in `cross_repo` and leave
|
|
521
|
+
the dependent findings at `info`. An unavailable producer is a result. Assuming
|
|
522
|
+
one is not.
|
|
523
|
+
|
|
436
524
|
### Memory (required)
|
|
437
525
|
|
|
438
526
|
Run, once, per review:
|
|
@@ -571,6 +659,62 @@ Recording the enumeration matters as much as doing it: a round that searched and
|
|
|
571
659
|
found nothing is a different fact from a round that never searched, and only the
|
|
572
660
|
recorded list distinguishes them.
|
|
573
661
|
|
|
662
|
+
#### Every prior finding leaves the round with a disposition
|
|
663
|
+
|
|
664
|
+
A fix round that reports only new findings is unreadable: the author cannot tell
|
|
665
|
+
which of their fixes landed. Close the loop explicitly — one line per prior
|
|
666
|
+
finding, in the report, before the new findings:
|
|
667
|
+
|
|
668
|
+
| Disposition | Meaning |
|
|
669
|
+
|---|---|
|
|
670
|
+
| `closed` | Checked against the code, not against the commit message, and the defect is gone |
|
|
671
|
+
| `open` | The fix does not reach the defect; say what is still true |
|
|
672
|
+
| `partial` | One site of the class was fixed and the enumeration named others |
|
|
673
|
+
| `regressed` | The fix removed this defect and introduced another — file the new one separately |
|
|
674
|
+
| `withdrawn` | The finding was wrong. See below |
|
|
675
|
+
|
|
676
|
+
**Check the code, not the commit message.** A commit titled *"report an abandoned
|
|
677
|
+
sync as abandoned"* is a claim; the disposition is whether the branch it renamed
|
|
678
|
+
is reachable and pinned. On a recorded round, two such commits asserted behaviour
|
|
679
|
+
on lines no test could reach.
|
|
680
|
+
|
|
681
|
+
#### A fix is a change, and changes get reviewed
|
|
682
|
+
|
|
683
|
+
The most expensive class in a multi-round review is **the defect the fix
|
|
684
|
+
introduced**. It is systematically under-found, for a structural reason: the fix
|
|
685
|
+
arrives framed as the answer to a finding, so it is read as an answer rather than
|
|
686
|
+
as new code. It is new code.
|
|
687
|
+
|
|
688
|
+
So scope A of a fix round includes the fix, reviewed on its own merits, and the
|
|
689
|
+
report carries its own section:
|
|
690
|
+
|
|
691
|
+
```markdown
|
|
692
|
+
## Regressions the fixes introduced
|
|
693
|
+
<[F-NNN] — the finding it was answering, and the new defect it created>
|
|
694
|
+
```
|
|
695
|
+
|
|
696
|
+
Recorded shapes, all from fixes that correctly closed the finding they answered:
|
|
697
|
+
an early return added to stop a fall-through, which then skipped the work the
|
|
698
|
+
caller needed; a persistence call added to save expanded state, which then
|
|
699
|
+
persisted the broken state on the error path too. Both were closed correctly and
|
|
700
|
+
both shipped a new bug in the same commit.
|
|
701
|
+
|
|
702
|
+
#### Withdrawing your own earlier finding
|
|
703
|
+
|
|
704
|
+
A finding from a previous round that this round disproves is **withdrawn**,
|
|
705
|
+
explicitly, at the top of the report, with the evidence — before any new finding.
|
|
706
|
+
|
|
707
|
+
This is not a courtesy. An uncorrected wrong finding costs the author a fix they
|
|
708
|
+
did not need, and it stays in `prior_findings` steering later rounds. Withdrawal
|
|
709
|
+
is also the one self-correction this pipeline permits, and it is permitted because
|
|
710
|
+
it is asymmetric: it *deletes* a claim, so it cannot inflate the finding count,
|
|
711
|
+
which is the failure mode that removed the re-scoring pass from Wave C.
|
|
712
|
+
|
|
713
|
+
State what made the original claim wrong, in one sentence, and if the same
|
|
714
|
+
reasoning error has now happened twice in one review, say that too. A reviewer
|
|
715
|
+
that names its own recurring error is calibrating; one that quietly drops a
|
|
716
|
+
finding is hiding a result.
|
|
717
|
+
|
|
574
718
|
### Step 1: Determine Review Mode
|
|
575
719
|
|
|
576
720
|
Before anything else, determine whether the request is **diff mode** or **path mode**:
|
|
@@ -766,6 +910,7 @@ If the repository has local convention docs such as `CLAUDE.md`, `AGENTS.md`,
|
|
|
766
910
|
| `**/*.test.*`, `**/*.spec.*`, `**/*.integration.test.*`, `**/*.msw.ts`, `src/test/**`, `test/**`, `e2e/**` | `review-testing-practices` |
|
|
767
911
|
| `src/core/**`, `core/**`, `shared/**`, `foundation/**` | `review-core-boundaries` |
|
|
768
912
|
| `src/core/flow/**`, `src/graph/**`, `src/shared/flow/**` | `review-flow-graph` |
|
|
913
|
+
| A `.tsx`/`.jsx`/`.css`/`.scss` hunk that adds or changes a sizing or spacing utility on a rendered element — `w-*`, `min-w-*`, `max-w-*`, `flex-*`, `basis-*`, `grid-cols-*`, `truncate`, `line-clamp-*`, `overflow-*`, `p*`/`m*`/`gap-*` and their logical `ps-*`/`pe-*`/`ms-*`/`me-*` forms, or an inline `style` carrying a dimension | `review-layout` |
|
|
769
914
|
|
|
770
915
|
These convention reviewers are additive: keep the generic reviewers selected by normal detection,
|
|
771
916
|
then add the matching convention pass. Deduplicate reviewer names before dispatch.
|
|
@@ -797,6 +942,84 @@ reviewer wrongly skipped hides a real defect, and that asymmetry is not close.
|
|
|
797
942
|
Record the exclusions with their reasons alongside the pre-filter drops. A
|
|
798
943
|
reviewer silently absent from a report reads as "it had nothing to say".
|
|
799
944
|
|
|
945
|
+
### Path gate — the second filter, same asymmetry
|
|
946
|
+
|
|
947
|
+
Stack scoping asks *does this repository have the thing*. The path gate asks *does
|
|
948
|
+
this diff*. A reviewer whose path triggers match **no file in scope A** is not
|
|
949
|
+
dispatched; it goes to `Skipped reviewers` with reason `no-matching-paths`.
|
|
950
|
+
|
|
951
|
+
This is worth its own step because `--all` is the common case and it is where the
|
|
952
|
+
waste is: a run of sixteen reviewers over a fourteen-file frontend diff dispatched
|
|
953
|
+
`review-core-boundaries` and `review-flow-graph` against a diff containing no
|
|
954
|
+
`src/core/**` file at all. Two agents, full prompt each, guaranteed empty.
|
|
955
|
+
|
|
956
|
+
Three bounds keep it from deleting coverage:
|
|
957
|
+
|
|
958
|
+
- **Only path-triggered reviewers.** A reviewer selected by an explicit flag, or
|
|
959
|
+
one whose scope is the whole change rather than a path set — `review-logic`,
|
|
960
|
+
`review-regression`, `review-verifier` — is never path-gated.
|
|
961
|
+
- **Scope A only.** Scope B is the blast radius; it is *supposed* to name files the
|
|
962
|
+
diff never touched.
|
|
963
|
+
- **Ambiguity includes.** If you cannot decide whether a path matches, dispatch.
|
|
964
|
+
Same asymmetry as stack scoping: a needless reviewer costs tokens, a missing one
|
|
965
|
+
costs a defect.
|
|
966
|
+
|
|
967
|
+
Record every gated reviewer with the trigger set that found nothing. "Skipped:
|
|
968
|
+
no matching paths" is a result the reader can check; an absent reviewer is not.
|
|
969
|
+
|
|
970
|
+
### Project-local reviewers — ask, do not assume
|
|
971
|
+
|
|
972
|
+
The routing table below names the reviewers **keryx ships**. A project may also
|
|
973
|
+
define its own, and before this step existed they were invisible to every round:
|
|
974
|
+
a team could write a reviewer, register it, and watch nothing dispatch it.
|
|
975
|
+
|
|
976
|
+
So the reviewer set is asked for, not recited:
|
|
977
|
+
|
|
978
|
+
```bash
|
|
979
|
+
keryx review reviewers --json
|
|
980
|
+
```
|
|
981
|
+
|
|
982
|
+
It returns two halves. `bundled` is the installed gdskills review tree — the set
|
|
983
|
+
this project's install profile actually produced, which is not everything keryx
|
|
984
|
+
ships. `project` is every project-skill under module `review`, living at
|
|
985
|
+
`.metaproject/project-skills/review/<name>/` so that it sits beside
|
|
986
|
+
`.metaproject/skills/gdskills/review/<name>/` and needs no other marking.
|
|
987
|
+
|
|
988
|
+
**Dispatch the project half alongside the bundled one.** A project reviewer is a
|
|
989
|
+
reviewer: it returns `REVIEW_RESULT`, its findings merge with everyone else's,
|
|
990
|
+
it is bound by the canonical severity rubric and the shared laws, and it is
|
|
991
|
+
verified in Wave C like any other. It is not advisory and not a second-class
|
|
992
|
+
pass — a team that wrote down how it reviews has said something about this
|
|
993
|
+
codebase that no shipped reviewer knows.
|
|
994
|
+
|
|
995
|
+
Two things it does NOT get:
|
|
996
|
+
|
|
997
|
+
- **No exemption from the contract.** A project reviewer whose output does not
|
|
998
|
+
conform is handled by the Sub-Agent Report Quality Gate exactly as a bundled
|
|
999
|
+
one would be. Local authorship is not evidence.
|
|
1000
|
+
- **No self-verification.** The never-self-verify rule is about the actor, not
|
|
1001
|
+
the origin.
|
|
1002
|
+
|
|
1003
|
+
#### `drift` — the source moved, the reviewer did not
|
|
1004
|
+
|
|
1005
|
+
A project reviewer built from an external file — a rules file, a review profile,
|
|
1006
|
+
a conventions doc — records where it came from and the hash of that file at
|
|
1007
|
+
import. `keryx review reviewers` re-reads the source and reports:
|
|
1008
|
+
|
|
1009
|
+
| `drift` | Meaning | What to do this round |
|
|
1010
|
+
|---|---|---|
|
|
1011
|
+
| `none` | No external source; written here | Nothing |
|
|
1012
|
+
| `clean` | Source matches the import | Nothing |
|
|
1013
|
+
| `changed` | Source has moved on since import | Dispatch it, and say so in the report |
|
|
1014
|
+
| `missing` | Source can no longer be read | Dispatch it, and say so in the report |
|
|
1015
|
+
|
|
1016
|
+
**A drifted reviewer still runs.** It is a reviewer built from an older version
|
|
1017
|
+
of its source, which is a fact about provenance, not a defect in its findings —
|
|
1018
|
+
suppressing it would trade real coverage for tidiness. Record the drift in
|
|
1019
|
+
`review_context` and name it once in the report, so the next person knows the
|
|
1020
|
+
profile is due a re-read. Never file it as a finding against the code under
|
|
1021
|
+
review: it is a fact about the review, not about the diff.
|
|
1022
|
+
|
|
800
1023
|
### Convention Reviewer Confirmation
|
|
801
1024
|
|
|
802
1025
|
When convention reviewers are auto-detected and the user did not explicitly pass
|
|
@@ -880,6 +1103,7 @@ Skipped reviewers:
|
|
|
880
1103
|
| `--testing-practices` | `review-testing-practices` |
|
|
881
1104
|
| `--core-boundaries` | `review-core-boundaries` |
|
|
882
1105
|
| `--flow-graph` | `review-flow-graph` |
|
|
1106
|
+
| `--layout` | `review-layout` |
|
|
883
1107
|
| `--all` | all reviewers above (including `review-clean-code`, `review-highload`, applicable legacy/profile reviewers, and project convention reviewers when local convention docs exist) |
|
|
884
1108
|
| `--verify` | `review-verifier`, AFTER all others; checks the consolidated findings by running something. Delete-only. |
|
|
885
1109
|
| (auto) | detected from diff file extensions — see Auto-detection table |
|
|
@@ -912,6 +1136,20 @@ Dispatch selected reviewers in parallel when independent. Use waves when token b
|
|
|
912
1136
|
3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
|
|
913
1137
|
exist, `--verify` is set, or the PR is high-risk. See below.
|
|
914
1138
|
|
|
1139
|
+
**Execution belongs to both B and C, and they execute for opposite reasons.**
|
|
1140
|
+
Wave C runs a command to *decide the fate of a finding that exists*; it is
|
|
1141
|
+
delete-only and can produce nothing. Wave B runs commands to *make findings* —
|
|
1142
|
+
`review-testing-practices` deletes each gate the diff added and records whether
|
|
1143
|
+
the suite goes red, and `review-layout` measures rendered geometry in a real
|
|
1144
|
+
engine.
|
|
1145
|
+
|
|
1146
|
+
The consequence is directional and easy to get wrong: **a surviving mutation is
|
|
1147
|
+
not something Wave C can hand you.** If Wave B skips its mutation pass, that
|
|
1148
|
+
finding class is unreachable for the entire round, and no amount of verification
|
|
1149
|
+
downstream recovers it. When a testing reviewer returns without a mutation table,
|
|
1150
|
+
treat it the way you would treat a reviewer that returned without findings *and*
|
|
1151
|
+
without evidence: ask once, then record that the pass did not run.
|
|
1152
|
+
|
|
915
1153
|
### Wave C — verification, and what it replaced
|
|
916
1154
|
|
|
917
1155
|
Wave C used to run `review-strict`: a meta-pass that re-read the consolidated
|
|
@@ -1178,6 +1416,42 @@ is a claim that the class has exactly one member — make it deliberately, becau
|
|
|
1178
1416
|
`minor` and `info` may omit it: enumerating the class for every low-severity
|
|
1179
1417
|
observation is theatre, not rigour.
|
|
1180
1418
|
|
|
1419
|
+
#### Negative enumeration — the empty set is also a class scope
|
|
1420
|
+
|
|
1421
|
+
Some of the strongest findings assert that something is **absent**: nothing clears
|
|
1422
|
+
this state, nothing reads this field, this prop can never fire, no test builds this
|
|
1423
|
+
shape. Those are claims about a set being empty, and an empty set is exactly as
|
|
1424
|
+
enumerable as a full one — the same `class_scope` machinery carries it.
|
|
1425
|
+
|
|
1426
|
+
```yaml
|
|
1427
|
+
class_scope:
|
|
1428
|
+
sites: []
|
|
1429
|
+
enumeration_method: "all 5 clearSubmitExecutionState() call sites are onUnmount,
|
|
1430
|
+
setActiveControl, setActiveTemplate, submitActiveControl, resetActiveControlForm;
|
|
1431
|
+
none reachable from a check-section edit, and there is no reaction/autorun"
|
|
1432
|
+
```
|
|
1433
|
+
|
|
1434
|
+
The `enumeration_method` is the whole finding here, so it carries more weight than
|
|
1435
|
+
usual: it must name the complete candidate set and why each member fails to apply.
|
|
1436
|
+
"I looked and did not find one" is not that. A search that returned nothing, with
|
|
1437
|
+
the query, is.
|
|
1438
|
+
|
|
1439
|
+
This form covers a class of defect nothing else in this pipeline reaches:
|
|
1440
|
+
|
|
1441
|
+
- **The inert change.** The diff adds a prop, a guard, a field or a branch that
|
|
1442
|
+
nothing can reach — a `disabled` prop on a component its parent unmounts instead
|
|
1443
|
+
of disabling; a field added to two type `Pick`s and never read; a `=== undefined`
|
|
1444
|
+
guard against a producer that never emits `undefined`. The diff looks like it
|
|
1445
|
+
does something and does nothing, and the reviewer that checks whether the change
|
|
1446
|
+
is *correct* passes it, because it is.
|
|
1447
|
+
- **The missing clear.** State is set on one path and no path unsets it.
|
|
1448
|
+
- **The unreachable branch.** A value the code handles that its producer cannot
|
|
1449
|
+
emit.
|
|
1450
|
+
|
|
1451
|
+
An inert change is `minor` when it is merely dead, and `major` when its deadness
|
|
1452
|
+
means the bug it was added to fix is still live — the second is the common case
|
|
1453
|
+
when the change was written to answer an earlier round's finding.
|
|
1454
|
+
|
|
1181
1455
|
All findings from all sub-reviewers must be normalized to this format before consolidation:
|
|
1182
1456
|
|
|
1183
1457
|
```markdown
|
|
@@ -1272,10 +1546,38 @@ STATUS: DONE | DONE_WITH_CONCERNS
|
|
|
1272
1546
|
## Minor & Info
|
|
1273
1547
|
<[F-NNN] findings with severity=minor or info>
|
|
1274
1548
|
|
|
1549
|
+
## Checked and cleared
|
|
1550
|
+
<Required whenever a reviewer tested a hypothesis and it did not hold. One line
|
|
1551
|
+
each: the defect that was looked for, and the evidence that rules it out. Not a
|
|
1552
|
+
list of virtues — a list of hypotheses that died, so no later round spends a
|
|
1553
|
+
reviewer re-raising them.>
|
|
1554
|
+
|
|
1275
1555
|
## Positive Notes
|
|
1276
1556
|
<Optional. Highlight things done well. Keep brief.>
|
|
1277
1557
|
```
|
|
1278
1558
|
|
|
1559
|
+
### Why `Checked and cleared` is separate from `Positive Notes`
|
|
1560
|
+
|
|
1561
|
+
They read alike and do opposite work. "Both entry points read the score from the
|
|
1562
|
+
same report" is a virtue; it tells a later round nothing, because no round was
|
|
1563
|
+
going to claim otherwise. "The drift bar cannot contradict the drift badge —
|
|
1564
|
+
`schema-drift-table.ts:61` always seeds `ddl` and `distributed` is a primitive
|
|
1565
|
+
boolean that is always serialized, so a backend match implies a `ddlTextDiffers`
|
|
1566
|
+
match" is a **retired hypothesis**: it names the bug that was hunted and the fact
|
|
1567
|
+
that kills it.
|
|
1568
|
+
|
|
1569
|
+
The cost of omitting them is paid in rounds. A plausible-but-wrong finding that
|
|
1570
|
+
was investigated and dropped in round 1 is investigated again in round 2 by a
|
|
1571
|
+
different reviewer, and the author answers it twice. Writing the negative down
|
|
1572
|
+
once ends that loop; it is also the only artifact that distinguishes *checked and
|
|
1573
|
+
clean* from *never looked at*, which is the distinction an approving verdict rests
|
|
1574
|
+
on.
|
|
1575
|
+
|
|
1576
|
+
Entries come from the reviewers, not from you: a reviewer that dropped a candidate
|
|
1577
|
+
returns it with the evidence, and consolidation collects them. A reviewer that
|
|
1578
|
+
returns zero findings and zero cleared hypotheses has told you nothing about the
|
|
1579
|
+
code, and should be asked once what it examined.
|
|
1580
|
+
|
|
1279
1581
|
---
|
|
1280
1582
|
|
|
1281
1583
|
## Skill Learning Handoff
|