@mrciphersmith/keryx 0.2.69 → 0.2.71
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli.js +11136 -4863
- package/docs/README.md +54 -0
- package/docs/requirements/shared-agent-context/README.md +104 -0
- package/package.json +3 -2
- package/src/gdgraph/build-lang.test.ts +10 -3
- package/src/gdgraph/build.ts +54 -9
- package/src/gdgraph/import-kind.test.ts +205 -0
- package/src/gdgraph/query.ts +6 -1
- package/src/gdgraph/types.ts +34 -0
- package/src/gdskills/bundled/rules/core/model-selection.mdc +184 -31
- package/src/gdskills/bundled/rules/core/skills-storage-workflow.mdc +36 -0
- package/src/gdskills/bundled/rules/core/subagent-status-protocol.md +27 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +159 -20
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.codex.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.cursor.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.opencode.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.zed.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.codex.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.cursor.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.md +22 -3
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.opencode.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.zed.md +20 -2
- package/src/gdskills/bundled/skills/planning/autodoc-analyst/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-architect/SKILL.md +3 -1
- package/src/gdskills/bundled/skills/planning/autodoc-assembler/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-orchestrator/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-scanner/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-writer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/docpack-orchestrator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/docpack-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/metaproject-security/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/review-architecture/SKILL.md +37 -10
- package/src/gdskills/bundled/skills/review/review-backend/SKILL.md +48 -14
- package/src/gdskills/bundled/skills/review/review-clean-code/SKILL.md +49 -12
- package/src/gdskills/bundled/skills/review/review-core-boundaries/SKILL.md +34 -2
- package/src/gdskills/bundled/skills/review/review-flow-graph/SKILL.md +33 -2
- package/src/gdskills/bundled/skills/review/review-frontend/SKILL.md +70 -29
- package/src/gdskills/bundled/skills/review/review-frontend-conventions/SKILL.md +34 -3
- package/src/gdskills/bundled/skills/review/review-highload/SKILL.md +49 -15
- package/src/gdskills/bundled/skills/review/review-logic/SKILL.md +39 -11
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +659 -64
- package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-finding.schema.json +7 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/verification-claim.schema.json +78 -0
- package/src/gdskills/bundled/skills/review/review-performance/SKILL.md +43 -13
- package/src/gdskills/bundled/skills/review/review-pr-feedback/SKILL.md +8 -2
- package/src/gdskills/bundled/skills/review/review-regression/SKILL.md +185 -0
- package/src/gdskills/bundled/skills/review/review-security-code/SKILL.md +44 -13
- package/src/gdskills/bundled/skills/review/review-style/SKILL.md +26 -6
- package/src/gdskills/bundled/skills/review/review-testing-practices/SKILL.md +35 -3
- package/src/gdskills/bundled/skills/review/review-verifier/SKILL.md +276 -0
- package/src/gdskills/contracts/review-finding.schema.json +119 -1
- package/src/gdskills/contracts/subagent-dispatch.schema.json +59 -3
- package/src/gdskills/bundled/skills/review/review-strict/SKILL.md +0 -328
|
@@ -4,10 +4,9 @@ description: |
|
|
|
4
4
|
Use when: a code review is requested and the user does not explicitly name a specialized reviewer.
|
|
5
5
|
Handles "review", "code review", "review PR", "review --frontend", "review --backend",
|
|
6
6
|
"review --architecture", "review --security", "review --performance", "review --style",
|
|
7
|
-
"review --
|
|
7
|
+
"review --verify", "review --project-conventions", "review --legacy-profiles", "review --all". Routes to specialized reviewers in parallel and
|
|
8
8
|
consolidates findings into one unified report.
|
|
9
9
|
NOT for: running a single specialized reviewer — invoke it directly by name instead.
|
|
10
|
-
version: "1.6.0"
|
|
11
10
|
triggers:
|
|
12
11
|
- "review"
|
|
13
12
|
- "code review"
|
|
@@ -18,11 +17,10 @@ triggers:
|
|
|
18
17
|
- "review --security"
|
|
19
18
|
- "review --performance"
|
|
20
19
|
- "review --style"
|
|
21
|
-
- "review --
|
|
20
|
+
- "review --verify"
|
|
22
21
|
- "review --all"
|
|
23
22
|
- "review --clean-code"
|
|
24
23
|
- "review --highload"
|
|
25
|
-
- "review --greptile"
|
|
26
24
|
- "review --project-conventions"
|
|
27
25
|
- "review --frontend-conventions"
|
|
28
26
|
- "review --testing-practices"
|
|
@@ -35,10 +33,10 @@ triggers:
|
|
|
35
33
|
- "review --mobx-store"
|
|
36
34
|
metadata:
|
|
37
35
|
author: "MrCipherSmith"
|
|
38
|
-
version: "1.
|
|
36
|
+
version: "1.8.0"
|
|
39
37
|
category: "review"
|
|
38
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
40
39
|
license: "MIT"
|
|
41
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
42
40
|
---
|
|
43
41
|
|
|
44
42
|
# Review Orchestrator
|
|
@@ -53,26 +51,109 @@ unified report sorted by severity. It does not perform any review logic itself.
|
|
|
53
51
|
|
|
54
52
|
```
|
|
55
53
|
Review Orchestrator Progress:
|
|
54
|
+
- [ ] Step 0: On a PR target, collect external comments — `keryx review comments collect`
|
|
56
55
|
- [ ] Step 1: Build Review Context Pack (PR metadata, scope, rules, context_doc summary)
|
|
57
56
|
- [ ] Step 2: Detect review mode (diff mode vs. path mode)
|
|
58
|
-
- [ ] Step 3:
|
|
57
|
+
- [ ] Step 3: Build the bounded scope with `keryx review scope` — never by hand
|
|
58
|
+
- [ ] Step 3b: On a deep round, compute scope B with `keryx review blast-radius` — never by browsing — and KEEP the `--json` file; `review ingest --blast-radius <file>` is refused without it
|
|
59
59
|
- [ ] Step 4: Parse flags / auto-detect domain from scope
|
|
60
|
-
- [ ] Step 5: Ask user to confirm optional convention
|
|
61
|
-
- [ ] Step 6: Plan sub-agent dispatch
|
|
60
|
+
- [ ] Step 5: Ask user to confirm optional convention reviewers (legacy/profile reviewers are flag-only, never prompted)
|
|
61
|
+
- [ ] Step 6: Plan sub-agent dispatch and token budgets, and compute each dispatch's model with `keryx review tier` — never by hand
|
|
62
62
|
- [ ] Step 7: Stage 1 gate - spec compliance check (if issue/task provided)
|
|
63
63
|
- [ ] Step 8: Dispatch selected reviewers in PARALLEL with reviewer-input schema
|
|
64
64
|
- [ ] Step 9: Collect reviewer-finding schema results and handle NEEDS_CONTEXT
|
|
65
|
-
- [ ] Step 10:
|
|
65
|
+
- [ ] Step 10: Wave C — dispatch `review-verifier` over the consolidated findings
|
|
66
66
|
- [ ] Step 11: Sort by severity, deduplicate, emit unified report
|
|
67
|
+
- [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
|
|
68
|
+
- [ ] Step 13: Report the stage counts: dropped by pre-filter, refuted by the verifier, retained
|
|
69
|
+
- [ ] Step 14: AFTER THE FINAL ROUND ONLY — answer every external comment once, `keryx review comments reply --final`
|
|
67
70
|
```
|
|
68
71
|
|
|
72
|
+
Step 0 runs on **every** round. Step 14 runs **once**, after the last one. They are
|
|
73
|
+
two commands for that reason: a caller that already runs collection per round
|
|
74
|
+
would carry the posting along with it, and the reviewer would get six replies to
|
|
75
|
+
one comment.
|
|
76
|
+
|
|
77
|
+
### Step 6 — the model is computed, not chosen
|
|
78
|
+
|
|
79
|
+
Before dispatching each reviewer, run `keryx review tier` with the signals you
|
|
80
|
+
already hold (`--scope`, `--findings`, `--diff-lines`, `--fix-attempt`,
|
|
81
|
+
`--verifier`, `--security`, `--forced-strategy-change`) and paste the `model`
|
|
82
|
+
block it prints into that dispatch.
|
|
83
|
+
|
|
84
|
+
Do NOT assign the tier by reading the table in `rules/core/model-selection.mdc`.
|
|
85
|
+
Working it out in your head is exactly the mechanical step that rule moves into
|
|
86
|
+
code — and it is the step that was documented as running for a whole release
|
|
87
|
+
while nothing called it.
|
|
88
|
+
|
|
89
|
+
The command names no model. It ranks whatever your provider reports at runtime
|
|
90
|
+
and places the tiers relative to your own model; when it cannot rank anything it
|
|
91
|
+
prints `inherit: true` and exits 0, which means the dispatch runs on the session
|
|
92
|
+
model. That is a correct answer, not a failure — never "fix" it by writing a
|
|
93
|
+
model id into a dispatch.
|
|
94
|
+
|
|
95
|
+
---
|
|
96
|
+
|
|
97
|
+
## Step 12 — the `keryx:findings` block
|
|
98
|
+
|
|
99
|
+
The prose report is for a human. **`keryx review ingest` reads the fenced block,
|
|
100
|
+
not the prose**, and a round that emits only prose cannot be used to build the
|
|
101
|
+
next round's input — the reviewer's `confidence`, `evidence`, `impact` and
|
|
102
|
+
`suggested_fix` do not survive rendering, and no regex recovers what was never
|
|
103
|
+
written down.
|
|
104
|
+
|
|
105
|
+
So every report ends with one fenced block whose info string carries
|
|
106
|
+
`keryx:findings`:
|
|
107
|
+
|
|
108
|
+
````text
|
|
109
|
+
```json keryx:findings
|
|
110
|
+
[ { …one object per finding, conforming to review-finding.schema.json… } ]
|
|
111
|
+
```
|
|
112
|
+
````
|
|
113
|
+
|
|
114
|
+
Rules:
|
|
115
|
+
|
|
116
|
+
- **Exactly one block per report.** It is the array the reviewers returned,
|
|
117
|
+
carried through — not re-derived from the prose above it. A second block is a
|
|
118
|
+
hard error naming both offsets: reviewers are consolidated by merging their
|
|
119
|
+
findings into one array, never by concatenating one block each.
|
|
120
|
+
- The fence may be indented up to three spaces (CommonMark), which is what
|
|
121
|
+
happens when the block is nested under a list item. Beyond that it is not a
|
|
122
|
+
fence and ingest will not see it.
|
|
123
|
+
- Each object conforms to **`review-finding.schema.json`** — the same contract
|
|
124
|
+
`prior_findings[].finding` is validated against, which is why the block
|
|
125
|
+
round-trips. Unknown per-finding properties are **dropped, not rejected**:
|
|
126
|
+
ingest writes exactly the properties that contract names, so anything else you
|
|
127
|
+
put on a finding is silently discarded rather than flagged. Pipeline triage
|
|
128
|
+
fields (`classification`, `flow_relevance`) are not finding properties at all;
|
|
129
|
+
they are the orchestrator's judgement, not the reviewer's, and are recorded in
|
|
130
|
+
`decisions.md`.
|
|
131
|
+
- `reviewer` is the reviewer that actually produced the finding. Never the
|
|
132
|
+
orchestrator's own name — that is the field whose loss made round 2
|
|
133
|
+
unconstructible.
|
|
134
|
+
- **A block that is present but unusable fails loudly.** Ingest refuses a block
|
|
135
|
+
it cannot parse, and equally refuses one that parses to something other than
|
|
136
|
+
an array of findings (or a single `{ reviewer, findings }` result) — `null`
|
|
137
|
+
included. It never falls back to parsing the prose, because a silent fallback
|
|
138
|
+
would reintroduce exactly the lossy path this replaces while the report still
|
|
139
|
+
visibly carries the structured array.
|
|
140
|
+
|
|
141
|
+
A report without the block is still readable by a human and still ingestible by
|
|
142
|
+
the legacy Markdown path — but it is a **legacy** report. Four fields the prose
|
|
143
|
+
does not carry (`impact`, `suggested_fix`, `evidence`, `confidence`) are written
|
|
144
|
+
with an explicit `not recorded:` provenance where the report supplies nothing,
|
|
145
|
+
and `confidence` is stamped `low` because a regex over prose is a low-confidence
|
|
146
|
+
derivation whatever the reviewer believed. Such a round **can** still seed a fix
|
|
147
|
+
round — that is the point of keeping the parser — but it seeds one that knows
|
|
148
|
+
which of its inputs were recovered and which were never written down.
|
|
149
|
+
|
|
69
150
|
---
|
|
70
151
|
|
|
71
152
|
## Input Contract
|
|
72
153
|
|
|
73
154
|
| Field | Type | Required | Description |
|
|
74
155
|
|-------|------|----------|-------------|
|
|
75
|
-
| `flags` | string[] | no | One or more of: `--frontend`, `--backend`, `--architecture`, `--security`, `--performance`, `--style`, `--clean-code`, `--highload`, `--project-conventions`, `--frontend-conventions`, `--testing-practices`, `--core-boundaries`, `--flow-graph`, `--legacy-profiles`, `--code-ai`, `--b091`, `--code-style`, `--mobx-store`, `--
|
|
156
|
+
| `flags` | string[] | no | One or more of: `--frontend`, `--backend`, `--architecture`, `--security`, `--performance`, `--style`, `--clean-code`, `--highload`, `--project-conventions`, `--frontend-conventions`, `--testing-practices`, `--core-boundaries`, `--flow-graph`, `--legacy-profiles`, `--code-ai`, `--b091`, `--code-style`, `--mobx-store`, `--verify`, `--all` |
|
|
76
157
|
| `path` | string | no | File or directory path to review (e.g., `src/stores/`, `src/components/UserCard.tsx`). Activates **path mode** — reviews the files at this path directly, not a git diff. |
|
|
77
158
|
| `commit_range` | string | no | Explicit commit hash or range (e.g., `abc123..HEAD`). Overrides merge-base detection. Ignored in path mode. |
|
|
78
159
|
| `issue_url` | string | no | GitHub issue or task URL. If provided, Stage 1 gate checks spec compliance before dispatching reviewers. |
|
|
@@ -81,6 +162,8 @@ Review Orchestrator Progress:
|
|
|
81
162
|
| `token_budget` | object | no | Optional budget controls: `{total, per_reviewer, diff_max_chars, file_max_chars}`. |
|
|
82
163
|
| `model_strategy` | string | no | `current`, `ask`, or `adaptive`. Default: `current`; do not switch models unless user or automation allows it. |
|
|
83
164
|
| `managed_review` | object | no | Optional managed review mode: `{mode, target, target_ref, flow_id, reviewers}` where mode is `lightweight`, `attach-review`, `review-flow`, or `ingest`. |
|
|
165
|
+
| `verification_mode` | string | no | `off`, `annotate`, or `filter`. Default `annotate` — verdicts are recorded and nothing is removed. See Wave C. |
|
|
166
|
+
| `pr_comments` | object | no | `{enabled, max_replies_total, max_sentences_per_reply}`. Defaults: enabled when a PR exists, `30`, `2`. Collect every round, reply once at the end. See External PR comments. |
|
|
84
167
|
|
|
85
168
|
---
|
|
86
169
|
|
|
@@ -111,10 +194,66 @@ Runtime CLI surface:
|
|
|
111
194
|
keryx review attach --flow <id> --target <kind> --ref <ref>
|
|
112
195
|
keryx review start --target <kind> --ref <ref>
|
|
113
196
|
keryx review ingest --report <path> [--flow <id>] --ref <ref>
|
|
197
|
+
[--verifications <file>] [--verification-mode off|annotate|filter]
|
|
198
|
+
[--scope <scope.json>] [--blast-radius <blast-radius.json>]
|
|
199
|
+
[--refuted <file>]
|
|
200
|
+
# --blast-radius is REQUIRED whenever this round dispatched a
|
|
201
|
+
# scope-B reviewer (review-regression). See below.
|
|
114
202
|
keryx review status <review-id-or-path>
|
|
115
203
|
keryx review complete <review-id-or-path>
|
|
204
|
+
[--finding <id> --disposition <state> --evidence <ref>]...
|
|
116
205
|
```
|
|
117
206
|
|
|
207
|
+
**An unrecognised option is refused, not ignored.** A misspelling used to be
|
|
208
|
+
accepted with exit 0, so `review complete --disposition ...` printed
|
|
209
|
+
`status: closed` and wrote nothing at all.
|
|
210
|
+
|
|
211
|
+
`--verifications` takes what `review-verifier` returned. `--scope` takes the
|
|
212
|
+
whole `--json` output of `keryx review scope`, so the package records what the
|
|
213
|
+
pre-filter dropped, **with a reason per drop**, as well as what the verifier
|
|
214
|
+
refuted. Without it — and with no `## Pre-filter scope` block already in the
|
|
215
|
+
package — the record says **`not recorded`** for that stage rather than `0`;
|
|
216
|
+
"dropped nothing" and "never ran" are different facts.
|
|
217
|
+
|
|
218
|
+
`--blast-radius` is **not optional on any round that dispatched
|
|
219
|
+
`review-regression`** — which is every `recommended` and `full` round, because
|
|
220
|
+
`review-regression` is in both profiles. AC3 is enforced in code: an ingest
|
|
221
|
+
carrying a scope-B finding with no blast-radius record is **refused**, and the
|
|
222
|
+
round is not recordable until the set is supplied. Pass the `--json` output of
|
|
223
|
+
`keryx review blast-radius` — the same record you computed at dispatch time to
|
|
224
|
+
bound the round.
|
|
225
|
+
|
|
226
|
+
Two channels, and both work:
|
|
227
|
+
|
|
228
|
+
- `--blast-radius <file>` — always correct, and the one to use.
|
|
229
|
+
- `.metaproject/reviews/blast-radius.json` (or
|
|
230
|
+
`.metaproject/flows/<flow>/reviews/blast-radius.json` for a flow-attached
|
|
231
|
+
round) — the handoff slot, read when no flag is passed. The ingest **consumes**
|
|
232
|
+
it: it is copied into the package it screened and removed from the slot, so a
|
|
233
|
+
later round cannot be screened against a set nobody recomputed.
|
|
234
|
+
|
|
235
|
+
Do **not** write the record inside the package directory before the ingest. A
|
|
236
|
+
round that names no `--review-id` has not allocated that directory yet, and
|
|
237
|
+
creating it makes the ingest take the next free name instead.
|
|
238
|
+
|
|
239
|
+
`--refuted` takes the findings this round **raised and then dismissed**, in the
|
|
240
|
+
same finding shape, each carrying `disposition: {state, evidence}` with one of
|
|
241
|
+
the `dismissed-*` states. Without it a package keeps only the survivors of an
|
|
242
|
+
unlogged triage — which is why precision measured over the recorded corpus
|
|
243
|
+
returns 100% whatever the reviewers actually got right.
|
|
244
|
+
|
|
245
|
+
`review complete --finding <id> --disposition <state> --evidence <ref>` records
|
|
246
|
+
what became of a finding when the fix round closes; the triple is repeatable, one
|
|
247
|
+
group per finding. States: `unknown`, `acted-on`, `dismissed-incorrect`,
|
|
248
|
+
`dismissed-wont-fix`, `dismissed-out-of-scope`, `dismissed-deprioritised`.
|
|
249
|
+
Everything except `unknown` must cite where the outcome is written down — the
|
|
250
|
+
commit, the test, the decision. A recorded state and its citation cannot be
|
|
251
|
+
overwritten by a later close; record a correction as a new round.
|
|
252
|
+
|
|
253
|
+
**Closing a fix round without dispositions leaves every finding reading
|
|
254
|
+
`unknown`.** That is not "the reviewers were right"; it is "nobody wrote down
|
|
255
|
+
what happened", and it is the single reason the precision figure cannot be read.
|
|
256
|
+
|
|
118
257
|
Managed modes:
|
|
119
258
|
|
|
120
259
|
- `lightweight`: report-only; no flow or managed review artifacts are created.
|
|
@@ -138,6 +277,147 @@ When attaching to a flow, resolve the flow by explicit `flow_id`, PR URL, issue
|
|
|
138
277
|
URL, or branch metadata. Never mutate `.metaproject/flows/*/flow.json` from
|
|
139
278
|
review code; Task Manager state changes remain owned by `keryx flow`.
|
|
140
279
|
|
|
280
|
+
---
|
|
281
|
+
|
|
282
|
+
## External PR comments
|
|
283
|
+
|
|
284
|
+
A human or a bot reviews our pull request. Before this existed, nothing collected
|
|
285
|
+
it, nothing fixed it and nothing answered it — silence was the behaviour for every
|
|
286
|
+
comment, and to the person who wrote it silence is indistinguishable from
|
|
287
|
+
disagreement.
|
|
288
|
+
|
|
289
|
+
**Collect every round. Answer once, at the end.** Both halves are mechanical and
|
|
290
|
+
both live in the CLI; the judgement — is this comment right, what do we say — stays
|
|
291
|
+
here.
|
|
292
|
+
|
|
293
|
+
```text
|
|
294
|
+
keryx review comments collect --repo <owner/repo> --pr <n> --sha <head-sha>
|
|
295
|
+
[--self <login>] [--round <n>] [--out <findings.json>] [--json]
|
|
296
|
+
keryx review comments reply --repo <owner/repo> --pr <n> --outcomes <file|->
|
|
297
|
+
--sha <head-sha> --final [--dry-run]
|
|
298
|
+
[--max-replies <n>] [--max-sentences <n>] [--max-chars <n>]
|
|
299
|
+
[--flow-link <url>]
|
|
300
|
+
```
|
|
301
|
+
|
|
302
|
+
Add `--fixtures <dir>` to either to run the whole loop against JSON on disk —
|
|
303
|
+
no token, no network, nothing posted. Use it to see what a reply pass would say
|
|
304
|
+
before it says it.
|
|
305
|
+
|
|
306
|
+
`--sha` is the commit you collected against, and it is required. The completion
|
|
307
|
+
gate compares it to the pull request's head: a collection that ran before the
|
|
308
|
+
comments arrived is **stale**, and a gate that could not tell the difference
|
|
309
|
+
would pass a flow with unanswered reviewers on it while printing
|
|
310
|
+
`0 outstanding`. A record with no SHA reads as "cannot be shown current", never
|
|
311
|
+
as "fresh".
|
|
312
|
+
|
|
313
|
+
### What collection does, so you do not do it by hand
|
|
314
|
+
|
|
315
|
+
- Reads **all three** sources: inline review comments, review submissions and
|
|
316
|
+
their bodies, and PR-level discussion.
|
|
317
|
+
- **A bot reviewer is a reviewer.** CodeRabbit, Greptile and Copilot comments go
|
|
318
|
+
down exactly the same path as a human's. The bot flag is recorded so a report
|
|
319
|
+
can say who spoke; nothing filters on it.
|
|
320
|
+
- Excludes our own identity, and comments already answered — **unless** the thread
|
|
321
|
+
has a newer reply from somebody else, which makes the comment new again.
|
|
322
|
+
- Everything filtered is listed with its reason. A filter that removes silently
|
|
323
|
+
reads as "nobody commented".
|
|
324
|
+
|
|
325
|
+
### Severity is classified, never invented
|
|
326
|
+
|
|
327
|
+
A comment on a review whose state is `CHANGES_REQUESTED` starts at **`major`**.
|
|
328
|
+
Everything else starts at **`minor`**. There is no third rule and no model call.
|
|
329
|
+
|
|
330
|
+
When the classifying fact is missing — an inline comment whose parent review was
|
|
331
|
+
not returned, or a review state GitHub does not document — the comment is **not
|
|
332
|
+
dropped and the severity is not guessed**: it takes the `minor` floor and carries
|
|
333
|
+
`basis: unclassified` naming what was missing. A derived `minor` and a defaulted
|
|
334
|
+
one are different claims, and a record that cannot tell them apart is the
|
|
335
|
+
`dismissed-out-of-scope: 0` failure in a new field.
|
|
336
|
+
|
|
337
|
+
You may **lower** a severity only by assigning a terminal disposition with a
|
|
338
|
+
reason. You may never silently drop an external comment.
|
|
339
|
+
|
|
340
|
+
### The verifier cannot refute an external comment
|
|
341
|
+
|
|
342
|
+
An external finding enters the same fix loop as an internal one with one
|
|
343
|
+
exception: **a `refuted` verdict does not remove it and does not dismiss it.** A
|
|
344
|
+
human asked a question; a machine deciding the question was invalid is not an
|
|
345
|
+
answer. `keryx review ingest` turns that verdict into the disposition
|
|
346
|
+
`answered-disagree`, keeps the finding, and records the reclaim in `scope.md`.
|
|
347
|
+
`answered-disagree` still owes a reply explaining why.
|
|
348
|
+
|
|
349
|
+
The per-reviewer findings cap does not truncate external comments either, for the
|
|
350
|
+
same reason: the cap drops silently, and an external comment may not be dropped
|
|
351
|
+
silently.
|
|
352
|
+
|
|
353
|
+
### Replying — once, at the end, briefly
|
|
354
|
+
|
|
355
|
+
The reply pass runs **after the final round and before the completion gate**, so
|
|
356
|
+
every reply states a settled outcome rather than an intention. `keryx review
|
|
357
|
+
comments reply` refuses without `--final`; it is not a reminder you can skip.
|
|
358
|
+
|
|
359
|
+
| Outcome | Reply is |
|
|
360
|
+
|---|---|
|
|
361
|
+
| `acted-on` | one sentence naming what changed, plus the commit SHA |
|
|
362
|
+
| `answered-disagree` | one or two sentences on why not, and a link to the flow's journal entry |
|
|
363
|
+
| `dismissed-out-of-scope` / `dismissed-deprioritised` | one sentence, and where it was recorded instead |
|
|
364
|
+
|
|
365
|
+
Rules, all of them enforced in code rather than asked for here:
|
|
366
|
+
|
|
367
|
+
- **At most two sentences per comment.** A longer reply is CUT to two and the
|
|
368
|
+
remainder is replaced by a link — the long version is not reachable from the
|
|
369
|
+
command's output. A truncation with no link to point at is refused outright: the
|
|
370
|
+
conclusion posted and the explanation nowhere is worse than either alternative.
|
|
371
|
+
- A fenced code block in a reply is refused. Link, do not paste.
|
|
372
|
+
- Replies go **in the thread**. A review submission body and a PR-level comment
|
|
373
|
+
have no thread — GitHub offers no reply endpoint for either — so those become one
|
|
374
|
+
top-level comment that names what it answers.
|
|
375
|
+
- **Never resolve or hide a thread we did not open.** Replying is ours; resolving
|
|
376
|
+
is the reviewer's call, and auto-resolving is how a bot silences a human. The
|
|
377
|
+
resolve, hide, minimise and dismiss endpoints are unreachable through the port
|
|
378
|
+
this command uses, GraphQL included.
|
|
379
|
+
- Exactly **one** reply per comment, and one disposition. A round that changed
|
|
380
|
+
nothing for a comment still gets a reply saying so, with a terminal disposition
|
|
381
|
+
— `unknown` is refused, because it is what an unanswered comment already reads
|
|
382
|
+
as.
|
|
383
|
+
- Capped at **30** replies. Beyond it, one summary comment and a backlog reported
|
|
384
|
+
by id.
|
|
385
|
+
- Handling is durable: `.metaproject/reviews/pr-comments/<owner>__<repo>__<n>.json`
|
|
386
|
+
records id, thread, author, url, first-seen round, handled-at, sha, disposition
|
|
387
|
+
and reply url, written after **every** post. A resumed session answers nobody
|
|
388
|
+
twice.
|
|
389
|
+
|
|
390
|
+
**The trade-off, stated rather than hidden:** a reviewer who comments early waits
|
|
391
|
+
until the end. That is deliberate — answering with a work-in-progress state that
|
|
392
|
+
later changes is worse. If a comment **blocks** progress rather than reporting a
|
|
393
|
+
problem, mark its outcome `escalate: true`: it leaves the reply queue, is reported
|
|
394
|
+
to the operator immediately, and the command exits non-zero. Answering a blocking
|
|
395
|
+
question at the end answers the wrong question late.
|
|
396
|
+
|
|
397
|
+
---
|
|
398
|
+
|
|
399
|
+
## Everything written to GitHub is brief
|
|
400
|
+
|
|
401
|
+
One rule, applied to every outward surface: **PR bodies, PR comments, review
|
|
402
|
+
replies, issue comments, and commit messages going to a PR.**
|
|
403
|
+
|
|
404
|
+
- Lead with the conclusion. No preamble, no restating the question, no apology,
|
|
405
|
+
no summary of the flow.
|
|
406
|
+
- Say what changed and where. **Link, do not paste.**
|
|
407
|
+
- The reasoning, the evidence, the rejected alternatives and the round history live
|
|
408
|
+
in the flow package — `journal.md`, `context.md`, the review artifacts — which is
|
|
409
|
+
durable, searchable, and costs a reader nothing to skip.
|
|
410
|
+
- A GitHub artifact that needs more than a short paragraph is a signal that the
|
|
411
|
+
detail belongs in the flow with a link out, **not** that the paragraph should
|
|
412
|
+
grow.
|
|
413
|
+
- No orchestrator-written PR comment or reply exceeds two sentences without
|
|
414
|
+
carrying a link to the artifact holding the detail. The reply pass enforces
|
|
415
|
+
this; for anything else you write outward, hold yourself to it.
|
|
416
|
+
|
|
417
|
+
This is deliberately asymmetric: **verbose in the flow, terse on GitHub.** The
|
|
418
|
+
flow is written for whoever resumes the work; GitHub is read by someone who did
|
|
419
|
+
not ask for our reasoning and is reading between other tasks.
|
|
420
|
+
|
|
141
421
|
## Review Context Pack
|
|
142
422
|
|
|
143
423
|
Before routing reviewers, build a compact `review_context` object. This is the shared source of truth for all sub-agents and must follow `skills/review-orchestrator/review-context.schema.json`.
|
|
@@ -238,7 +518,7 @@ If the platform supports assigning models to sub-agents and the user/automation
|
|
|
238
518
|
|---|---|---|
|
|
239
519
|
| simple | cheaper/faster coding model | `review-style`, `review-clean-code`, docs-only convention checks, legacy/profile checks |
|
|
240
520
|
| normal | current/default model | `review-frontend`, `review-backend`, `review-testing-practices`, convention reviewers |
|
|
241
|
-
| complex | strongest available coding/reasoning model | `review-logic`, `review-architecture`, `review-security-code`, `review-highload`,
|
|
521
|
+
| complex | strongest available coding/reasoning model | `review-logic`, `review-architecture`, `review-security-code`, `review-highload`, strict synthesis |
|
|
242
522
|
|
|
243
523
|
Rules:
|
|
244
524
|
- Do not silently change model class when `model_strategy` is `current`.
|
|
@@ -297,34 +577,152 @@ Before anything else, determine whether the request is **diff mode** or **path m
|
|
|
297
577
|
|
|
298
578
|
See shared script: `skills/shared/git-merge-base.md`
|
|
299
579
|
|
|
300
|
-
Run the script to determine `BASE_SHA`, then:
|
|
580
|
+
Run the script to determine `BASE_SHA`, then let the pre-filter build the scope:
|
|
301
581
|
|
|
302
582
|
```bash
|
|
303
|
-
|
|
304
|
-
|
|
583
|
+
keryx review scope --ref "${BASE_SHA}" --json > scope.json # KEEP THIS FILE
|
|
584
|
+
keryx review scope --ref "${BASE_SHA}" --scoped-diff # what reviewers get
|
|
305
585
|
```
|
|
306
586
|
|
|
587
|
+
**Keep `scope.json` until the round is ingested, and pass it as `--scope`.** That
|
|
588
|
+
file is how the drop list reaches the review record. `--append
|
|
589
|
+
"<review-package>/scope.md"` also writes it and is still supported — it now
|
|
590
|
+
REPLACES an existing `## Pre-filter scope` block rather than adding a second, and
|
|
591
|
+
`review ingest` carries any block it finds forward verbatim rather than
|
|
592
|
+
overwriting it — but `--scope scope.json` is the supported path, because it is
|
|
593
|
+
the one that does not depend on running two commands against the same file in the
|
|
594
|
+
right order.
|
|
595
|
+
|
|
596
|
+
**Do not run `git diff` yourself, and do not decide what to leave out.** The
|
|
597
|
+
pre-filter is deterministic code with no model call: it drops generated,
|
|
598
|
+
lockfile, snapshot, vendored and minified paths, drops whitespace-only and
|
|
599
|
+
comment-only change blocks, and bounds every retained change to ±20 lines of
|
|
600
|
+
context (`--context <n>`) instead of the whole file. Dropping a lockfile needs no
|
|
601
|
+
judgement, so it does not get one.
|
|
602
|
+
|
|
603
|
+
The record carries the retained scope **and every drop with its reason**. Both
|
|
604
|
+
halves are required: a scope that shrank without saying so reads afterwards as
|
|
605
|
+
"we reviewed everything". Note that `--scope` takes the WHOLE `--json` document —
|
|
606
|
+
handing over only its `counts` object is refused, because eight integers carry no
|
|
607
|
+
reason for any individual drop.
|
|
608
|
+
|
|
609
|
+
Use `.files` from `scope.json` for the auto-detection table below. A dropped path
|
|
610
|
+
must not select a reviewer, and **neither may a blast-radius path**: scope B is
|
|
611
|
+
under regression check, so a `.tsx` file that only appears there must not pull in
|
|
612
|
+
`review-frontend`. Reviewer selection is driven by the scope-A file list alone.
|
|
613
|
+
|
|
307
614
|
Scope is limited to **changes introduced in the current branch since merge-base**.
|
|
308
615
|
|
|
309
616
|
---
|
|
310
617
|
|
|
618
|
+
### Scope B — the blast radius (deep rounds)
|
|
619
|
+
|
|
620
|
+
Everything above is **scope A**: the change, bounded. It answers *is this change
|
|
621
|
+
correct?* It does not answer *did this change break something that was working*,
|
|
622
|
+
and those are different questions — only the first has ever been asked here.
|
|
623
|
+
|
|
624
|
+
A deep round dispatches under **both**. Scope B is computed, never browsed:
|
|
625
|
+
|
|
626
|
+
```bash
|
|
627
|
+
keryx review blast-radius --ref "${BASE_SHA}" --json > blast-radius.json # KEEP THIS FILE
|
|
628
|
+
keryx review blast-radius --ref "${BASE_SHA}" --brief # what a scope-B reviewer is told
|
|
629
|
+
```
|
|
630
|
+
|
|
631
|
+
**KEEP THIS FILE** is not advice. The ingest at Step 12 is **refused** if a
|
|
632
|
+
scope-B finding arrives without it — pass it back as
|
|
633
|
+
`review ingest ... --blast-radius blast-radius.json`.
|
|
634
|
+
|
|
635
|
+
It walks `gdgraph affected` outward from every changed file, ranks by edge
|
|
636
|
+
distance, keeps distance ≤ 2, cuts at 40 files closest-first, and adds a changed
|
|
637
|
+
file's naming-related tests when the graph did not already reach them. Requires a
|
|
638
|
+
built graph — run `keryx gdgraph build` if it refuses.
|
|
639
|
+
|
|
640
|
+
**Do not pick the files yourself, and do not widen it.** "Review the
|
|
641
|
+
functionality so nothing breaks" naively means "review the whole repository every
|
|
642
|
+
round", which is unaffordable *and* actively harmful: review quality decays as
|
|
643
|
+
context grows — measured F1 0.65 at round 2 falling to 0.29 at round 10. An
|
|
644
|
+
unbounded scope B makes later rounds worse than earlier ones.
|
|
645
|
+
|
|
646
|
+
The bounds are measured on this repository, not guessed: at depth 2 the set is a
|
|
647
|
+
median of 19 files (p90 65); depth 3 buys eight more in the median and doubles
|
|
648
|
+
the p90. The 40-file cap fires on 25% of commits and removes only hop-2 entries
|
|
649
|
+
on all but 2 of 80, so it almost never costs a direct dependent — and when it
|
|
650
|
+
does, it says so.
|
|
651
|
+
|
|
652
|
+
**Record the whole thing.** `--out "<review-package>/blast-radius.md"` writes the
|
|
653
|
+
set, the depth, and **every file the cap removed**. A truncation nobody can see
|
|
654
|
+
reads afterwards as "we checked everything", which is the claim this pipeline
|
|
655
|
+
exists to stop making. An empty radius is reported as `unresolved`, not as clean:
|
|
656
|
+
the graph indexes code, so a change to a skill, a rule or a schema has no blast
|
|
657
|
+
radius at all and that is a different fact from "nothing depends on it".
|
|
658
|
+
|
|
659
|
+
#### The scope-B question, and what is rejected
|
|
660
|
+
|
|
661
|
+
> Does this change break an existing behaviour **at these sites**?
|
|
662
|
+
|
|
663
|
+
Nothing else. The blast-radius set is **under regression check, not under
|
|
664
|
+
review**. A finding about style, naming or architecture in code the change did
|
|
665
|
+
not touch is refused **by the orchestrator in code** — not discouraged here —
|
|
666
|
+
under three rules, every one of them a fact about the claim rather than about who
|
|
667
|
+
made it:
|
|
668
|
+
|
|
669
|
+
| Rule | Refused because |
|
|
670
|
+
|---|---|
|
|
671
|
+
| `outside-set` | the file is neither in the computed set nor in the changed set; the reviewer went browsing |
|
|
672
|
+
| `non-regression-severity` | below `major`. Under the canonical rubric `minor` states the code behaves correctly and `info` names neither trigger nor outcome; neither can be a claim that something broke |
|
|
673
|
+
| `no-link-to-change` | nothing in the finding names a changed file, module or symbol. A regression claim says THE CHANGE broke this site |
|
|
674
|
+
|
|
675
|
+
Rejections are **recorded, not deleted** — raise the observation under scope A or
|
|
676
|
+
as a separate review. Pass `--brief` output verbatim into the scope-B dispatch:
|
|
677
|
+
the code rejection is the enforcement, but a reviewer told afterwards has already
|
|
678
|
+
spent the round producing findings that will all be refused.
|
|
679
|
+
|
|
680
|
+
`class_scope` on a scope-B finding names the **caller that breaks**, not the
|
|
681
|
+
changed line, because that is the site a human has to look at.
|
|
682
|
+
|
|
683
|
+
#### When it is recomputed
|
|
684
|
+
|
|
685
|
+
| Round | Scope A | Scope B |
|
|
686
|
+
|---|---|---|
|
|
687
|
+
| 1 (first after the draft PR) | yes | yes |
|
|
688
|
+
| 2..N | yes | recomputed only if the changed-file set moved |
|
|
689
|
+
| final | yes | **yes, always** |
|
|
690
|
+
|
|
691
|
+
Do not decide this by memory:
|
|
692
|
+
|
|
693
|
+
```bash
|
|
694
|
+
keryx review blast-radius --ref "${BASE_SHA}" --previous blast-radius.json [--final]
|
|
695
|
+
```
|
|
696
|
+
|
|
697
|
+
It prints the decision and the reason, and reuses the previous record when
|
|
698
|
+
nothing moved. The final round recomputes whatever the file set did — otherwise a
|
|
699
|
+
fix introduced in round 3 gets no regression check at all, and the round that
|
|
700
|
+
certifies the flow is the one that checked the least.
|
|
701
|
+
|
|
702
|
+
---
|
|
703
|
+
|
|
311
704
|
### Path Mode
|
|
312
705
|
|
|
313
|
-
When a path or target is named, collect the files
|
|
706
|
+
When a path or target is named, collect the candidate files:
|
|
314
707
|
|
|
315
708
|
```bash
|
|
316
709
|
# If a directory path is given:
|
|
317
710
|
find <path> -type f \( -name "*.ts" -o -name "*.tsx" -o -name "*.js" -o -name "*.jsx" \) | sort
|
|
318
711
|
|
|
319
|
-
# If a file path is given:
|
|
320
|
-
cat <file>
|
|
321
|
-
|
|
322
712
|
# If a module name is given (e.g. "UserStore", "pipelines module"):
|
|
323
713
|
find . -type f -name "*<name>*" \( -name "*.ts" -o -name "*.tsx" \)
|
|
324
714
|
# Also check common locations: src/stores/, src/modules/, src/components/
|
|
325
715
|
```
|
|
326
716
|
|
|
327
|
-
|
|
717
|
+
Then put the list through the same exclusions before reading anything:
|
|
718
|
+
|
|
719
|
+
```bash
|
|
720
|
+
keryx review scope --path "src/a.ts,src/b.ts" --json > scope.json
|
|
721
|
+
```
|
|
722
|
+
|
|
723
|
+
Pass the full **file contents** of the paths it **retained** to sub-reviewers, and
|
|
724
|
+
read none of the ones it dropped. Set `SCOPE_MODE: path`. Path mode has no hunks
|
|
725
|
+
and therefore no context window; the drop list is recorded exactly the same way.
|
|
328
726
|
|
|
329
727
|
**Reviewer behavior in path mode:** reviewers check the entire file content — not just added lines. All findings apply to the current state of the code, not only to changes.
|
|
330
728
|
|
|
@@ -351,7 +749,7 @@ If the repository has local convention docs such as `CLAUDE.md`, `AGENTS.md`,
|
|
|
351
749
|
|
|
352
750
|
| File pattern | Reviewers appended |
|
|
353
751
|
|---|---|
|
|
354
|
-
| `src/**/*.
|
|
752
|
+
| `src/**/*.tsx`, `*.stories.tsx`, or a `.ts`/`.js` change in a repo where `package.json` declares `react`/`react-dom`/`mobx`/`mobx-react`/`mobx-react-lite` as a dependency | `review-frontend-conventions` |
|
|
355
753
|
| `**/*.test.*`, `**/*.spec.*`, `**/*.integration.test.*`, `**/*.msw.ts`, `src/test/**`, `test/**`, `e2e/**` | `review-testing-practices` |
|
|
356
754
|
| `src/core/**`, `core/**`, `shared/**`, `foundation/**` | `review-core-boundaries` |
|
|
357
755
|
| `src/core/flow/**`, `src/graph/**`, `src/shared/flow/**` | `review-flow-graph` |
|
|
@@ -359,6 +757,33 @@ If the repository has local convention docs such as `CLAUDE.md`, `AGENTS.md`,
|
|
|
359
757
|
These convention reviewers are additive: keep the generic reviewers selected by normal detection,
|
|
360
758
|
then add the matching convention pass. Deduplicate reviewer names before dispatch.
|
|
361
759
|
|
|
760
|
+
### Stack scoping — run it after detection, before dispatch
|
|
761
|
+
|
|
762
|
+
The tables above select reviewers by **file shape**. A `.ts` file looks the same
|
|
763
|
+
whether or not the repository has React in it, so those tables will happily
|
|
764
|
+
dispatch a React/MobX conventions reviewer at a Bun CLI with no frontend — which
|
|
765
|
+
is exactly what happened here, on every review, for months.
|
|
766
|
+
|
|
767
|
+
So the selected set is filtered once more, by what the repository actually
|
|
768
|
+
declares:
|
|
769
|
+
|
|
770
|
+
```bash
|
|
771
|
+
keryx review stack --json
|
|
772
|
+
```
|
|
773
|
+
|
|
774
|
+
It reads `package.json` and reports, per reviewer, `include` or `exclude` with a
|
|
775
|
+
reason. A reviewer carrying `metadata.stack_requires` is dispatched when **any**
|
|
776
|
+
tag it names is present — matching what `keryx review stack` actually computes,
|
|
777
|
+
and failing toward inclusion rather than away from it.
|
|
778
|
+
|
|
779
|
+
**Its failure mode is to include, never to skip.** A missing, unparsable or
|
|
780
|
+
unexpected manifest sets `uncertain`, and an uncertain detection marks every tag
|
|
781
|
+
present, so every reviewer runs. A reviewer that runs needlessly costs tokens; a
|
|
782
|
+
reviewer wrongly skipped hides a real defect, and that asymmetry is not close.
|
|
783
|
+
|
|
784
|
+
Record the exclusions with their reasons alongside the pre-filter drops. A
|
|
785
|
+
reviewer silently absent from a report reads as "it had nothing to say".
|
|
786
|
+
|
|
362
787
|
### Convention Reviewer Confirmation
|
|
363
788
|
|
|
364
789
|
When convention reviewers are auto-detected and the user did not explicitly pass
|
|
@@ -399,25 +824,15 @@ Legacy/profile reviewers are specialized review profiles that predate the review
|
|
|
399
824
|
| `--mobx-store` | `code-mobx-store-review` |
|
|
400
825
|
| `*.store.ts`, `makeObservable`, `observable`, `computed`, `action.bound` | suggest `code-mobx-store-review` as optional profile reviewer |
|
|
401
826
|
|
|
402
|
-
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
C) Skip legacy/profile reviewers (recommended unless you need these profiles)
|
|
412
|
-
|
|
413
|
-
Available:
|
|
414
|
-
- code-ai-review: strict AI review profile
|
|
415
|
-
- code-b091-review: b091-style strict logic profile
|
|
416
|
-
- code-style-review: legacy style/architecture profile
|
|
417
|
-
- code-mobx-store-review: MobX store/state profile (only if MobX/store files are present)
|
|
418
|
-
```
|
|
419
|
-
|
|
420
|
-
If the user chooses B, list only applicable reviewers and ask for exact names. If the review is part of `job-orchestrator`, use `reviewers` and `conditional_reviewers` automation settings when provided.
|
|
827
|
+
Legacy/profile reviewers are never auto-included and never prompted for — do not ask the user
|
|
828
|
+
about them. They are exempt from the finding contract (for example `code-ai-review` emits
|
|
829
|
+
free-prose Russian with no per-finding severity field, so its output cannot be normalised into
|
|
830
|
+
the unified report), which is why inclusion must be a deliberate, explicit act rather than a
|
|
831
|
+
default the user has to opt out of on every review. Dispatch them ONLY when the user passes one
|
|
832
|
+
of the flags in the Trigger table above, or when `job-orchestrator` provides `reviewers` /
|
|
833
|
+
`conditional_reviewers` automation settings that name them. The `code-mobx-store-review`
|
|
834
|
+
auto-suggestion (MobX/store files present) is informational only — list it in the Review Plan
|
|
835
|
+
Preview below, but do not dispatch it and do not ask about it without an explicit flag.
|
|
421
836
|
|
|
422
837
|
Review Plan Preview must include an `Optional legacy/profile reviewers` group and a `Skipped reviewers` group with reasons such as:
|
|
423
838
|
|
|
@@ -447,14 +862,13 @@ Skipped reviewers:
|
|
|
447
862
|
| `--style` | `review-style` |
|
|
448
863
|
| `--clean-code` | `review-clean-code` |
|
|
449
864
|
| `--highload` | `review-highload` |
|
|
450
|
-
| `--greptile` | `review-greptile` (codebase-aware; requires PR number) |
|
|
451
865
|
| `--project-conventions` | all generic convention reviewers: `review-frontend-conventions` + `review-testing-practices` + `review-core-boundaries` + `review-flow-graph` |
|
|
452
866
|
| `--frontend-conventions` | `review-frontend-conventions` |
|
|
453
867
|
| `--testing-practices` | `review-testing-practices` |
|
|
454
868
|
| `--core-boundaries` | `review-core-boundaries` |
|
|
455
869
|
| `--flow-graph` | `review-flow-graph` |
|
|
456
|
-
| `--all` | all reviewers above (including `review-clean-code`, `review-highload`, applicable legacy/profile reviewers, project convention reviewers when local convention docs exist
|
|
457
|
-
| `--
|
|
870
|
+
| `--all` | all reviewers above (including `review-clean-code`, `review-highload`, applicable legacy/profile reviewers, and project convention reviewers when local convention docs exist) |
|
|
871
|
+
| `--verify` | `review-verifier`, AFTER all others; checks the consolidated findings by running something. Delete-only. |
|
|
458
872
|
| (auto) | detected from diff file extensions — see Auto-detection table |
|
|
459
873
|
|
|
460
874
|
Multiple flags may be combined. Example: `review --backend --security` dispatches
|
|
@@ -482,7 +896,85 @@ Dispatch selected reviewers in parallel when independent. Use waves when token b
|
|
|
482
896
|
|
|
483
897
|
1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
|
|
484
898
|
2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files.
|
|
485
|
-
3. Wave C -
|
|
899
|
+
3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
|
|
900
|
+
exist, `--verify` is set, or the PR is high-risk. See below.
|
|
901
|
+
|
|
902
|
+
### Wave C — verification, and what it replaced
|
|
903
|
+
|
|
904
|
+
Wave C used to run `review-strict`: a meta-pass that re-read the consolidated
|
|
905
|
+
findings and **adjusted their severity with no new evidence**, under an elevation
|
|
906
|
+
table biased 3:1 toward escalation. It was **removed, not improved**, and the
|
|
907
|
+
reason is measured rather than stylistic:
|
|
908
|
+
|
|
909
|
+
- **GPT-4 on GSM8K across self-correction rounds: 95.5 → 91.5 → 89.0.**
|
|
910
|
+
**GPT-3.5 on CommonSenseQA: 75.8 → 38.1.** Among the answers that changed,
|
|
911
|
+
correct → incorrect exceeded incorrect → correct (Huang et al., *Large Language
|
|
912
|
+
Models Cannot Self-Correct Reasoning Yet*, ICLR 2024, arXiv:2310.01798).
|
|
913
|
+
- **Self-Refine (arXiv:2303.17651): +49.2 on dialogue response generation, +0.2
|
|
914
|
+
on maths.** Self-refinement gains are on subjective tasks and vanish on
|
|
915
|
+
verifiable reasoning. Judging whether a null-guard is missing is verifiable
|
|
916
|
+
reasoning.
|
|
917
|
+
|
|
918
|
+
Re-scoring a finding by re-reading it is therefore not a rigour pass; it is a
|
|
919
|
+
coin flip weighted toward more findings. **Do not restore it because it looks
|
|
920
|
+
obviously useful — it looked obviously useful the first time.**
|
|
921
|
+
|
|
922
|
+
`review-verifier` occupies the slot and differs in exactly one way that matters:
|
|
923
|
+
**it runs something.** Verification that executes rejects 85–96% of false reports
|
|
924
|
+
against 4–15% unaided while finding 30–44% more true bugs (AnyPoC,
|
|
925
|
+
arXiv:2604.11950); Meta's TestGen-LLM funnel discards 75% of its own output
|
|
926
|
+
(75% build → 57% build and pass → 25% improve coverage) and the surviving quarter
|
|
927
|
+
reaches 73% human acceptance (arXiv:2402.09171).
|
|
928
|
+
|
|
929
|
+
It also **never votes.** 80+ agents unanimously endorsed a padding-oracle
|
|
930
|
+
vulnerability that did not exist, and a single empirical test killed it: consensus
|
|
931
|
+
cannot detect a hallucination its members share, so agreement between reviewers is
|
|
932
|
+
not evidence and must never be recorded as verification.
|
|
933
|
+
|
|
934
|
+
That rule is about agreement *standing in for* evidence. It is not a rule that
|
|
935
|
+
two verifiers may not both check the same finding: each claim is admitted on its
|
|
936
|
+
own — a named non-author, a real method, real evidence, with `reasoning` already
|
|
937
|
+
capped — and when two such claims reach the **same** verdict the merge records it,
|
|
938
|
+
naming both verifiers and carrying both pieces of evidence. Claims that
|
|
939
|
+
**disagree** still cancel, because there the only thing deciding the outcome
|
|
940
|
+
would be claim order.
|
|
941
|
+
|
|
942
|
+
Dispatch rules:
|
|
943
|
+
|
|
944
|
+
- Pass the consolidated findings, each carrying `global_id` and the **real**
|
|
945
|
+
originating `reviewer`. A finding whose `reviewer` is the orchestrator cannot be
|
|
946
|
+
routed away from its author, so the never-self-verify rule silently stops
|
|
947
|
+
applying — that field was hardcoded to `review-orchestrator` on all 83 recorded
|
|
948
|
+
findings and is fixed only from 0.2.70 onward.
|
|
949
|
+
- **A finding is never verified by the reviewer that raised it.** When only one
|
|
950
|
+
reviewer ran, its findings are simply left unverified; verifying them yourself
|
|
951
|
+
is worse than not verifying them. The merge compares the two names after
|
|
952
|
+
normalising case, surrounding whitespace, `_`/`-`, and a trailing `(model)`
|
|
953
|
+
annotation, so `review-logic `, `Review-Logic` and `review-logic (sonnet)` are
|
|
954
|
+
all the same actor. Do not try to route around it by respelling the name — the
|
|
955
|
+
comparison deliberately over-matches, because a refused claim only ever costs a
|
|
956
|
+
verdict while a missed self-verification costs the finding.
|
|
957
|
+
- The verifier returns `verification-claim.schema.json`. Merge it with
|
|
958
|
+
`keryx review ingest --verifications <file>`; do not apply verdicts by hand.
|
|
959
|
+
- **The verifier can only delete.** If it returns a severity, a new finding, or a
|
|
960
|
+
rewritten finding, the merge discards that whole claim and records the attempt.
|
|
961
|
+
Do not "help" by applying it.
|
|
962
|
+
|
|
963
|
+
### `verification_mode`
|
|
964
|
+
|
|
965
|
+
`off` | `annotate` | `filter`. **Default `annotate`, and it stays `annotate` for
|
|
966
|
+
one release.**
|
|
967
|
+
|
|
968
|
+
| Mode | What happens |
|
|
969
|
+
|---|---|
|
|
970
|
+
| `off` | No verification. Claims are refused rather than silently ignored. |
|
|
971
|
+
| `annotate` | Verdicts are recorded on the findings. **Nothing is removed.** A `refuted` finding is still reported, marked refuted. |
|
|
972
|
+
| `filter` | An applied `refuted` verdict removes the finding from the reported set and records it as `dismissed-incorrect`, with the verification evidence. |
|
|
973
|
+
|
|
974
|
+
`annotate` is the default so the drop rate is a **measured number** before it
|
|
975
|
+
costs a real finding. The risk is named rather than assumed away: SWE-agent keeps
|
|
976
|
+
its equivalent step opt-in because it sometimes rejects correct patches. Do not
|
|
977
|
+
switch a project to `filter` on the strength of one round.
|
|
486
978
|
|
|
487
979
|
### Agent Runtime Compatibility
|
|
488
980
|
|
|
@@ -524,24 +1016,6 @@ Each reviewer must return a `REVIEW_RESULT` object matching `skills/review-orche
|
|
|
524
1016
|
|
|
525
1017
|
**Important for path mode:** instruct each reviewer to check the **entire file**, not just changes. The scope report should say "Path: `<TARGET_PATH>`" instead of a branch/merge-base.
|
|
526
1018
|
|
|
527
|
-
### Greptile Reviewer
|
|
528
|
-
|
|
529
|
-
`review-greptile` runs in parallel with the other reviewers **when a PR number is available** (diff mode with a PR). It is excluded in path mode (no PR) unless `--greptile` is explicitly specified.
|
|
530
|
-
|
|
531
|
-
When dispatching `review-greptile`, pass additionally:
|
|
532
|
-
|
|
533
|
-
```
|
|
534
|
-
PR_NUMBER: <pr number>
|
|
535
|
-
REPO: <owner/repo>
|
|
536
|
-
REMOTE: github | gitlab
|
|
537
|
-
```
|
|
538
|
-
|
|
539
|
-
Greptile findings use `G-` prefixed IDs and are merged into the consolidated report under a dedicated section **"## Greptile (Codebase-Aware Findings)"** placed before the Blockers section. If Greptile identified cross-file impact not caught by other reviewers, those appear as additional blockers/majors.
|
|
540
|
-
|
|
541
|
-
**Auto-include Greptile when:** `--all` flag is used AND a PR number is resolvable from the current branch (`gh pr view` succeeds).
|
|
542
|
-
|
|
543
|
-
---
|
|
544
|
-
|
|
545
1019
|
## Scope Boundaries
|
|
546
1020
|
|
|
547
1021
|
| Concern | This skill | Use instead |
|
|
@@ -553,6 +1027,7 @@ Greptile findings use `G-` prefixed IDs and are merged into the consolidated rep
|
|
|
553
1027
|
| Security vulnerabilities | NO | `review-security-code` |
|
|
554
1028
|
| Performance anti-patterns | NO | `review-performance` |
|
|
555
1029
|
| Style / naming / import order | NO | `review-style` |
|
|
1030
|
+
| Checking whether a reported finding is real | NO | `review-verifier` |
|
|
556
1031
|
| Clean Code principles + SOLID at code level | NO | `review-clean-code` |
|
|
557
1032
|
| Concurrency, resource pools, caching, queues, idempotency | NO | `review-highload` |
|
|
558
1033
|
| Frontend repository conventions | NO | `review-frontend-conventions` |
|
|
@@ -576,6 +1051,93 @@ Before consolidation, validate every reviewer result:
|
|
|
576
1051
|
|
|
577
1052
|
---
|
|
578
1053
|
|
|
1054
|
+
## Severity (canonical)
|
|
1055
|
+
|
|
1056
|
+
**This is the only severity rubric in the review domain.** Reviewers do not carry
|
|
1057
|
+
their own. Ten private rubrics feeding one sort produce a ranking that means ten
|
|
1058
|
+
different things at once, and ranking is what an operator uses to decide what to
|
|
1059
|
+
read first. A reviewer may state which of *its* conditions land where; it may not
|
|
1060
|
+
redefine the levels.
|
|
1061
|
+
|
|
1062
|
+
### `blocker` — merge-blocking, and nothing else
|
|
1063
|
+
|
|
1064
|
+
Exactly four shapes. Nothing outside this list is a `blocker`, however strongly
|
|
1065
|
+
the reviewer feels about it:
|
|
1066
|
+
|
|
1067
|
+
1. **A crash** — the process, request, or render dies on an input the change
|
|
1068
|
+
admits.
|
|
1069
|
+
2. **Data loss or corruption** — something persisted, transmitted, or returned is
|
|
1070
|
+
destroyed or silently wrong.
|
|
1071
|
+
3. **An exploitable vulnerability** — an attacker action with a named entry point
|
|
1072
|
+
and a named impact.
|
|
1073
|
+
4. **An unimplemented acceptance criterion** — the change claims work the diff
|
|
1074
|
+
does not contain.
|
|
1075
|
+
|
|
1076
|
+
Everything else is at most `major`. "This will definitely cause problems later"
|
|
1077
|
+
is not one of the four. Neither is "this violates the architecture", "this fails
|
|
1078
|
+
the linter", or "this is how the last outage started".
|
|
1079
|
+
|
|
1080
|
+
### `major` / `minor` / `info` — the boundary test
|
|
1081
|
+
|
|
1082
|
+
Ask one question, and ask it of the **finding**, not of the code:
|
|
1083
|
+
|
|
1084
|
+
> **Does it name a trigger, and the observable outcome that trigger produces?**
|
|
1085
|
+
|
|
1086
|
+
- **`major`** — it does. There is an input, a call, a render, or a load level, and
|
|
1087
|
+
a resulting behaviour a user or a caller would call wrong: a wrong value, a lost
|
|
1088
|
+
update, a leak, a hang, a cost stated together with the frequency that makes it
|
|
1089
|
+
a cost. Not one of the four shapes above, so not merge-blocking — but the code
|
|
1090
|
+
does the wrong thing.
|
|
1091
|
+
- **`minor`** — it does not, and does not claim to. The code behaves correctly;
|
|
1092
|
+
the cost lands on whoever reads or edits it next, and the finding names that
|
|
1093
|
+
cost at a named site.
|
|
1094
|
+
- **`info`** — it names neither. An observation, a preference, or a risk with no
|
|
1095
|
+
demonstrated path.
|
|
1096
|
+
|
|
1097
|
+
The test is procedural on purpose. It is applied by reading the finding, so
|
|
1098
|
+
someone who did not write it — and has not read the code — reaches the same
|
|
1099
|
+
answer: look for the trigger and the outcome. Present → `major`. Absent, but a
|
|
1100
|
+
concrete maintenance cost is named → `minor`. Neither → `info`.
|
|
1101
|
+
|
|
1102
|
+
Two consequences, both previously decided differently in different files:
|
|
1103
|
+
|
|
1104
|
+
- A finding that **claims** runtime harm and cannot name the trigger is `info`,
|
|
1105
|
+
not `major`. It is not demoted to `minor`: `minor` is for findings that never
|
|
1106
|
+
claimed runtime harm at all. The two are different failures and stay
|
|
1107
|
+
distinguishable.
|
|
1108
|
+
- Severity is a property of the demonstrated outcome, never of the reviewer that
|
|
1109
|
+
found it. A security reviewer's unproven concern is `info` under the same test
|
|
1110
|
+
that puts a style reviewer's unproven concern there.
|
|
1111
|
+
- **And never of how crisply the finding is worded.** An outcome that costs a
|
|
1112
|
+
user, a caller or persisted state nothing is `minor` however precisely its
|
|
1113
|
+
trigger is named. Without this clause the test above rates prose quality: a
|
|
1114
|
+
cosmetic wording nit stated as "trigger X produces output Y" reads as `major`,
|
|
1115
|
+
while a real defect stated tersely reads as `info`. That is not academic — the
|
|
1116
|
+
findings cap truncates by severity, so the well-written typo would survive and
|
|
1117
|
+
the terse real defect would be cut.
|
|
1118
|
+
|
|
1119
|
+
The boundary this rubric does **not** draw is `major` against `major`. Two
|
|
1120
|
+
findings that both name a trigger and an outcome are the same severity even when
|
|
1121
|
+
one is obviously worse; the ordering inside a severity is the operator's, and
|
|
1122
|
+
inventing a fifth level to express it would put us back where we started.
|
|
1123
|
+
|
|
1124
|
+
### Shared laws (every reviewer)
|
|
1125
|
+
|
|
1126
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
1127
|
+
name the input, call, or condition that reaches the code, you have an
|
|
1128
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
1129
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
1130
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
1131
|
+
pattern because it is often wrong elsewhere.
|
|
1132
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
1133
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
1134
|
+
finding hide the other nine problems.
|
|
1135
|
+
|
|
1136
|
+
`review-security-code` carries a fourth — every security finding states its attack
|
|
1137
|
+
vector — which does not generalise and stays there.
|
|
1138
|
+
|
|
1139
|
+
---
|
|
1140
|
+
|
|
579
1141
|
## Finding Format
|
|
580
1142
|
|
|
581
1143
|
### Class scope — required for `blocker` and `major`
|
|
@@ -676,6 +1238,18 @@ STATUS: DONE | DONE_WITH_CONCERNS
|
|
|
676
1238
|
- minor: N
|
|
677
1239
|
- info: N
|
|
678
1240
|
|
|
1241
|
+
## Stage counts
|
|
1242
|
+
<!-- Required. State what each stage REMOVED, and never state it as a precision
|
|
1243
|
+
improvement: no precision baseline exists to improve on. The one measured
|
|
1244
|
+
from the review packages on disk was 53/53 = 100% — pinned there by
|
|
1245
|
+
construction, because nothing in that corpus could record a finding as
|
|
1246
|
+
wrong. Copy these from `scope.md`; do not re-count by hand. -->
|
|
1247
|
+
- dropped by pre-filter: <files>, <blocks>, <changed lines> (or `not recorded` if no scope was built)
|
|
1248
|
+
- verification mode: `<off | annotate | filter>`
|
|
1249
|
+
- verdicts: confirmed N, refuted N, unverifiable N, unverified N
|
|
1250
|
+
- refuted by the verifier: N (removed: N — always 0 outside `filter`)
|
|
1251
|
+
- retained: N
|
|
1252
|
+
|
|
679
1253
|
## Blockers (must fix before merge)
|
|
680
1254
|
<[F-NNN] findings with severity=blocker, sorted by file>
|
|
681
1255
|
|
|
@@ -744,13 +1318,22 @@ Publish this review report to the PR?
|
|
|
744
1318
|
|
|
745
1319
|
### Concise PR Comment
|
|
746
1320
|
|
|
747
|
-
The visible PR comment is for humans. It must be written in English only and stay
|
|
1321
|
+
The visible PR comment is for humans. It must be written in English only and stay
|
|
1322
|
+
concise, under the brevity rule above: **the summary is at most two sentences and
|
|
1323
|
+
carries a link to the artifact holding the detail.**
|
|
1324
|
+
|
|
1325
|
+
The finding rows below are a bounded exception, not a licence: they exist because
|
|
1326
|
+
a reviewer scanning a PR needs the blockers in front of them. Keep them to the
|
|
1327
|
+
`blocker` and `major` rows; everything at `minor` or below goes behind the
|
|
1328
|
+
`<details>` fold or, better, into the AI artifact and is linked. The full findings
|
|
1329
|
+
set, the round history and the reasoning belong in the flow package — pasting them
|
|
1330
|
+
here is the failure this rule names.
|
|
748
1331
|
|
|
749
1332
|
```markdown
|
|
750
1333
|
## AI Review Report
|
|
751
1334
|
|
|
752
1335
|
**Verdict:** REQUEST_CHANGES
|
|
753
|
-
**Summary:**
|
|
1336
|
+
**Summary:** At most two sentences: the overall risk and the main merge blocker. Detail: <link to the AI artifact or the flow package>.
|
|
754
1337
|
|
|
755
1338
|
| Severity | Area | Finding | Suggested Fix | Owner |
|
|
756
1339
|
|---|---|---|---|---|
|
|
@@ -922,6 +1505,18 @@ If absent, proceed normally — context is optional and non-blocking.
|
|
|
922
1505
|
| "Spec compliance can wait until after quality review" | Stage 1 gate exists because unimplemented requirements invalidate quality work |
|
|
923
1506
|
| "I'll deduplicate findings manually in my head" | Always normalize to [F-NNN] format before consolidation to avoid losing findings |
|
|
924
1507
|
| "Minor findings from one reviewer cancel out the major from another" | Each finding stands independently; severity is per-finding, not averaged |
|
|
1508
|
+
| "The reviewer's own severity table said blocker" | There are no reviewer tables. One rubric, in **Severity (canonical)** above; a reviewer that ships one is the defect this replaced |
|
|
1509
|
+
| "It's a security/architecture finding, so it's a blocker" | Severity is the demonstrated outcome, not the domain that found it. `blocker` is exactly the four shapes |
|
|
1510
|
+
| "It will definitely break something eventually, so blocker" | Name the trigger and the outcome. Named → `major`. Unnamed → `info`. "Eventually" is neither |
|
|
1511
|
+
| "A strict re-read of the findings will sharpen them" | That pass existed and was removed: self-correction without new evidence measured 95.5 → 91.5 → 89.0 on GSM8K and 75.8 → 38.1 on CommonSenseQA. Run something instead |
|
|
1512
|
+
| "Three reviewers agree, so the finding is verified" | Consensus is not evidence. 80+ agents unanimously endorsed a vulnerability that did not exist; one empirical test killed it |
|
|
1513
|
+
| "The verifier suggested a higher severity, I'll apply it" | It cannot suggest one. A claim carrying a severity is discarded whole and the attempt is recorded |
|
|
1514
|
+
| "This finding has no `verification`, so it can be dropped" | Absent means nobody checked. All 83 recorded findings are in that state; none of them is thereby wrong |
|
|
1515
|
+
| "Precision went up after the verifier landed" | There is no precision baseline to have gone up from. State stage counts: dropped, refuted, retained |
|
|
1516
|
+
| "I'll widen the blast radius, this change looks risky" | It is bounded because review quality decays with context: F1 0.65 at round 2 → 0.29 at round 10. Widening makes the later rounds worse, not safer |
|
|
1517
|
+
| "The blast radius came back empty, so nothing can break" | Empty and unresolved are different facts. The graph indexes code — a Markdown or JSON change has no radius at all, and the record says which one you got |
|
|
1518
|
+
| "The changed files are the same as last round, so scope B can be skipped on the final round" | The final round always recomputes. A fix landed in round 3 is the change; skipping means the certifying round checked the least |
|
|
1519
|
+
| "This scope-B file has an obvious naming problem, I'll report it" | Rejected in code: a naming problem is `minor` at best, and the floor is `major`. The set is under regression check, not under review — raise it under scope A |
|
|
925
1520
|
| "No flags means no reviewers" | No flags → run auto-detection; never produce an empty review |
|
|
926
1521
|
| "User named a module so I'll use diff mode" | Named module/component/store → path mode; diff mode is only for branch changes |
|
|
927
1522
|
| "Path mode should only show lines I'd flag in diff mode" | Path mode reviews the entire file — all findings apply, not just added lines |
|